KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
☆1,281Mar 24, 2026Updated 6 months ago
Alternatives and similar repositories for KernelBench
Users that are interested in KernelBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆579Sep 8, 2026Updated last month
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆143Jun 14, 2025Updated last year
- Building the Virtuous Cycle for AI-driven LLM Systems☆288Sep 25, 2026Updated 2 weeks ago
- ☆107Nov 22, 2025Updated 10 months ago
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.☆370Oct 2, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Tile primitives for speedy kernels☆3,746Sep 12, 2026Updated 3 weeks ago
- Distributed Compiler and Optimized Parallel Kernels☆1,559Sep 18, 2026Updated 3 weeks ago
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,542Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,556Updated this week
- A benchmark of real-world DL kernel problems☆308Jul 15, 2026Updated 2 months ago
- Ship correct and fast LLM kernels to PyTorch☆155Jan 14, 2026Updated 8 months ago
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆210Mar 29, 2026Updated 6 months ago
- ☆497Aug 26, 2026Updated last month
- A Quirky Assortment of CuTe Kernels☆1,166Sep 17, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆156Aug 18, 2025Updated last year
- Automated High-Performance GPU Kernel Generation☆139Jun 1, 2026Updated 4 months ago
- Review automated kernel generation in the era of LLMs☆315Jun 25, 2026Updated 3 months ago
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆960Updated this week
- Accelerating MoE with IO and Tile-aware Optimizations☆777Aug 29, 2026Updated last month
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆495Sep 17, 2026Updated 3 weeks ago
- Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.☆1,304Oct 1, 2026Updated last week
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆149Sep 14, 2026Updated 3 weeks ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆8,516Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation☆1,337Jul 8, 2026Updated 3 months ago
- Kernels, of the mega variety :)☆843May 26, 2026Updated 4 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,580Mar 19, 2026Updated 6 months ago
- DeeperGEMM: crazy optimized version☆85May 5, 2025Updated last year
- We invite you to visit and follow our new repository at https://github.com/microsoft/TileFusion. TiledCUDA is a highly efficient kernel …☆192Jan 28, 2025Updated last year
- Fast low-bit matmul kernels in Triton☆488Oct 1, 2026Updated last week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming☆823Updated this week
- DeepGEMM: clean and efficient BLAS kernel library on GPU☆8,874Sep 30, 2026Updated last week
- Speed of Light Analysis for ML Model Runtime☆131Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- cuTile is a programming model for writing parallel kernels for NVIDIA GPUs☆2,154Updated this week
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,546Sep 23, 2026Updated 2 weeks ago
- A kernel library written in tilelang☆1,934Sep 30, 2026Updated last week
- ☆360Updated this week
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,369Aug 28, 2025Updated last year
- Puzzles for learning Triton☆2,625Apr 1, 2026Updated 6 months ago
- ☆32Jul 2, 2025Updated last year