Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems
☆75Aug 31, 2026Updated this week
Alternatives and similar repositories for KernelGen
Users that are interested in KernelGen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆324Updated this week
- FlagOS skills for model deployment, HW adaptation, train&infer, eval, kernel dev and perf tuning☆19Jul 18, 2026Updated last month
- A vLLM plugin built on the FlagOS unified multi-chip backend.☆88Updated this week
- FlagGems is an operator library for large language models implemented in the Triton Language.☆1,089Updated this week
- Review automated kernel generation in the era of LLMs☆307Jun 25, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆137Jun 29, 2026Updated 2 months ago
- ☆17Jun 10, 2025Updated last year
- MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation☆67Jul 8, 2026Updated last month
- ☆425Aug 26, 2026Updated last week
- ☆17Jul 18, 2026Updated last month
- KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels☆51Jun 1, 2026Updated 3 months ago
- ☆25Dec 11, 2024Updated last year
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆535Jul 15, 2026Updated last month
- ☆103Nov 22, 2025Updated 9 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Tmux sidebar for vibe coding. Manage sessions and monitor agents at a glance☆15Aug 25, 2026Updated last week
- ☆30Jun 18, 2026Updated 2 months ago
- ☆216Aug 26, 2026Updated last week
- Code for Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking, EMNLP 2022, https://aclan…☆14Mar 30, 2026Updated 5 months ago
- A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It …☆203Apr 22, 2026Updated 4 months ago
- PTO instruction set architecture☆68Updated this week
- Skills for writing tilelang and debugging with CUDA toolkits.☆137May 20, 2026Updated 3 months ago
- ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and agentized workflo…☆36Jun 19, 2026Updated 2 months ago
- [IJCAI 2024] CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning☆26Feb 1, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language☆362Aug 17, 2026Updated 2 weeks ago
- ☆34Mar 12, 2026Updated 5 months ago
- ☆263Updated this week
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 4 months ago
- DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery☆23Sep 24, 2025Updated 11 months ago
- ChipScope / ILA using XVC (XIlinx Virtual Cable Over PCIe) with a PR (Partial Reconfiguration) design Example.☆15Jun 1, 2017Updated 9 years ago
- Benchmark PyTorch Custom Operators☆14Jul 6, 2023Updated 3 years ago
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆151Nov 10, 2025Updated 9 months ago
- A benchmark of real-world DL kernel problems☆288Jul 15, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Jan 7, 2025Updated last year
- A crowd-powered database system, with SQL-like query interface, multi-goal optimization☆11Sep 4, 2017Updated 8 years ago
- RLVR training for LLM in CUDA/C++☆42Jul 25, 2026Updated last month
- An OpenCL-Based FPGA Accelerator for Compressed YOLOv2☆38May 27, 2021Updated 5 years ago
- [ICML 2026]☆18Jul 4, 2026Updated last month
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- The code repository for "Multi-layer Rehearsal Feature Augmentation for Class-Incremental Learning" (ICML24)☆12Jun 7, 2024Updated 2 years ago