☆103Nov 22, 2025Updated 9 months ago
Alternatives and similar repositories for robust-kbench
Users that are interested in robust-kbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,227Mar 24, 2026Updated 5 months ago
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆321Nov 3, 2025Updated 10 months ago
- Automated High-Performance GPU Kernel Generation☆137Jun 1, 2026Updated 3 months ago
- Official Repo of CudaForge☆87Dec 2, 2025Updated 9 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆280May 1, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Ship correct and fast LLM kernels to PyTorch☆154Jan 14, 2026Updated 7 months ago
- Automated bottleneck detection and solution orchestration☆22Feb 24, 2026Updated 6 months ago
- Tinker ↔ KernelBench Integration enabling RL for GPU Kernel Generation☆31Mar 5, 2026Updated 5 months ago
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆138Jun 14, 2025Updated last year
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆147Aug 13, 2026Updated 3 weeks ago
- The official repository of ALE-Bench☆217Aug 26, 2026Updated last week
- Throughput-oriented multi-turn inference engine for KernelBench [ICML '25]☆24May 27, 2025Updated last year
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆207Mar 29, 2026Updated 5 months ago
- Optimizing diffusion for production-ready speeds☆40Jan 10, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Generating Efficient AI-Centric Kernels☆167Updated this week
- ☆68Jul 14, 2025Updated last year
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆38Dec 24, 2025Updated 8 months ago
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.☆367Updated this week
- NVFP4 Flash-Attention 4 on BlackWell☆53Jul 23, 2026Updated last month
- ☆23Jul 28, 2026Updated last month
- ☆634May 24, 2026Updated 3 months ago
- Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X☆80Feb 11, 2026Updated 6 months ago
- ☆25Feb 14, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Python package for rematerialization-aware gradient checkpointing☆27Oct 31, 2023Updated 2 years ago
- AlgoTune is a NeurIPS 2025 benchmark made up of 154 math, physics, and computer science problems. The goal is write code that solves each…☆116Jun 24, 2026Updated 2 months ago
- A benchmark of real-world DL kernel problems☆289Jul 15, 2026Updated last month
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆90Jul 12, 2026Updated last month
- [NeurIPS 2025 (spotlight)] HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs☆16Dec 17, 2025Updated 8 months ago
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆28Jun 30, 2026Updated 2 months ago
- [ICLR 2025] "Training LMs on Synthetic Edit Sequences Improves Code Synthesis" (Piterbarg, Pinto, Fergus)☆19Feb 11, 2025Updated last year
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 30, 2026Updated last month
- Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems☆75Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution 🧬☆1,367Aug 21, 2026Updated 2 weeks ago
- [MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization☆68Jul 28, 2026Updated last month
- ☆155Aug 18, 2025Updated last year
- [COLM-LLA 2026] The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficien…☆24Aug 23, 2026Updated last week
- ☆13Jul 15, 2024Updated 2 years ago
- Official implementation of Paper "System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving"☆30Apr 17, 2026Updated 4 months ago
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning☆478Mar 30, 2026Updated 5 months ago