MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation
☆67Jul 8, 2026Updated last month
Alternatives and similar repositories for MultiKernelBench
Users that are interested in MultiKernelBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆138Jun 14, 2025Updated last year
- Ascend operator generation☆35Jun 17, 2026Updated 2 months ago
- Review automated kernel generation in the era of LLMs☆291Jun 25, 2026Updated last month
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,202Mar 24, 2026Updated 4 months ago
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆203Mar 29, 2026Updated 4 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems☆72Jul 18, 2026Updated last month
- Triton language and compiler for Ascend NPU☆138Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆514Jul 15, 2026Updated last month
- LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development☆82Mar 31, 2026Updated 4 months ago
- kernelboard is the webapp for https://www.gpumode.com☆18Jul 27, 2026Updated 3 weeks ago
- Ship correct and fast LLM kernels to PyTorch☆153Jan 14, 2026Updated 7 months ago
- ☆22Jun 29, 2026Updated last month
- Generating Efficient AI-Centric Kernels☆153Updated this week
- ☆25Feb 14, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization☆57Jul 28, 2026Updated 3 weeks ago
- NVFP4 Flash-Attention 4 on BlackWell☆42Jul 23, 2026Updated 3 weeks ago
- SGLang kernel library for NPU☆173Updated this week
- ☆11Nov 1, 2023Updated 2 years ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆267May 1, 2026Updated 3 months ago
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆68Aug 8, 2026Updated last week
- A benchmark of real-world DL kernel problems☆279Jul 15, 2026Updated last month
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆144Updated this week
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆316Nov 3, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- SICP Online Judge, consisting of a server, a react web interface and a modified Ok client.☆12Dec 5, 2022Updated 3 years ago
- ☆57Mar 15, 2025Updated last year
- ☆21Jun 12, 2026Updated 2 months ago
- Automated High-Performance GPU Kernel Generation☆130Jun 1, 2026Updated 2 months ago
- Ascend TileLang adapter☆352Updated this week
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- FlagGems is an operator library for large language models implemented in the Triton Language.☆1,079Updated this week
- Unlimited Vector Extension with Data Streaming Support☆12Nov 25, 2024Updated last year
- ☆102Nov 22, 2025Updated 8 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- PyTorch distributed training from scratch (for educational purposes only)☆22Apr 12, 2025Updated last year
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated 2 years ago
- Throughput-oriented multi-turn inference engine for KernelBench [ICML '25]☆24May 27, 2025Updated last year
- SIMPLER MAGIC: Synthesis and In-memory MaPping of Logic Execution in a single Row for Memristor Aided loGIC☆13Dec 5, 2019Updated 6 years ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆264Updated this week
- An attempt to migrate Karpathy's llm.c to safe rust.☆13Jun 4, 2024Updated 2 years ago
- Large language models designed for formal theorem proving through tool-integrated reasoning.☆33Aug 13, 2025Updated last year