MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation
☆66Jul 8, 2026Updated 3 weeks ago
Alternatives and similar repositories for MultiKernelBench
Users that are interested in MultiKernelBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Ascend operator generation☆31Jun 17, 2026Updated last month
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- Review automated kernel generation in the era of LLMs☆276Jun 25, 2026Updated last month
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,163Mar 24, 2026Updated 4 months ago
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆196Mar 29, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems☆68Jul 18, 2026Updated last week
- Triton language and compiler for Ascend NPU☆118Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆491Jul 15, 2026Updated 2 weeks ago
- ☆28Jun 18, 2026Updated last month
- LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development☆76Mar 31, 2026Updated 3 months ago
- kernelboard is the webapp for https://www.gpumode.com☆18Updated this week
- Generating Efficient AI-Centric Kernels☆138Updated this week
- Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development☆27May 7, 2026Updated 2 months ago
- ☆23Feb 14, 2026Updated 5 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization☆57Jun 18, 2026Updated last month
- NVFP4 Flash-Attention 4 on BlackWell☆31Updated this week
- SGLang kernel library for NPU☆170Updated this week
- Torq compiler sources☆53Jul 22, 2026Updated last week
- ☆19Jun 13, 2025Updated last year
- Building the Virtuous Cycle for AI-driven LLM Systems☆261May 1, 2026Updated 2 months ago
- A benchmark of real-world DL kernel problems☆265Jul 15, 2026Updated 2 weeks ago
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆55Updated this week
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆143May 19, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆315Nov 3, 2025Updated 8 months ago
- A "standard library" of Triton kernels.☆26Oct 2, 2025Updated 9 months ago
- ☆57Mar 15, 2025Updated last year
- Democratizing AlphaFold3: an PyTorch reimplementation to accelerate protein structure prediction☆22May 24, 2025Updated last year
- Automated High-Performance GPU Kernel Generation☆120Jun 1, 2026Updated last month
- Ascend TileLang adapter☆340Updated this week
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- ☆10Mar 8, 2025Updated last year
- FlagGems is an operator library for large language models implemented in the Triton Language.☆1,057Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Unlimited Vector Extension with Data Streaming Support☆12Nov 25, 2024Updated last year
- ☆101Nov 22, 2025Updated 8 months ago
- B站-数电的ppt☆11Feb 19, 2024Updated 2 years ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated last year
- Throughput-oriented multi-turn inference engine for KernelBench [ICML '25]☆24May 27, 2025Updated last year
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆252Updated this week
- ☆19Updated this week