[MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization
☆57Jul 28, 2026Updated 2 weeks ago
Alternatives and similar repositories for AccelOpt
Users that are interested in AccelOpt are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Development☆17Jan 6, 2026Updated 7 months ago
- A Feishu/Lark AI agent bot☆15Feb 27, 2026Updated 5 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆265May 1, 2026Updated 3 months ago
- ☆69Jul 14, 2026Updated 3 weeks ago
- Can AI Agents Build Bespoke Systems?☆90Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- Nex Venus Communication Library☆75Nov 17, 2025Updated 8 months ago
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆511Jul 15, 2026Updated 3 weeks ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,196Mar 24, 2026Updated 4 months ago
- ☆47Updated this week
- a size profiler for cuda binary☆71Jan 15, 2026Updated 6 months ago
- Benchmark PyTorch Custom Operators☆14Jul 6, 2023Updated 3 years ago
- A DL compiler fuzzer☆15Nov 1, 2024Updated last year
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Jul 17, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆189Updated this week
- ☆14Apr 24, 2024Updated 2 years ago
- DeeperGEMM: crazy optimized version☆86May 5, 2025Updated last year
- ☆13Dec 9, 2024Updated last year
- A Triton JIT runtime and ffi provider in C++☆38Aug 7, 2026Updated last week
- ☆16Jun 15, 2026Updated last month
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆75Updated this week
- Generating Efficient AI-Centric Kernels☆154Updated this week
- HeteroHalide: From Image Processing DSL to Efficient FPGA Acceleration☆15Sep 14, 2020Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆33Mar 12, 2026Updated 5 months ago
- Vivado HLS study notes, courses, documents.☆12Dec 7, 2019Updated 6 years ago
- ☆17Updated this week
- Manages vllm-nccl dependency☆18Jun 3, 2024Updated 2 years ago
- Project showing how to develop NKI kernels for Llama 3.2 1B inference☆21May 29, 2025Updated last year
- MLSys competition for the best MOE NKI kernels☆48May 29, 2026Updated 2 months ago
- ☆19Nov 21, 2025Updated 8 months ago
- ☆10May 16, 2024Updated 2 years ago
- Code for the paper "A Boolean Task Algebra For Reinforcement Learning"☆11Dec 8, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- some docs for rookies in nics-efc☆22Mar 17, 2022Updated 4 years ago
- ☆102Nov 22, 2025Updated 8 months ago
- A benchmark of real-world DL kernel problems☆277Jul 15, 2026Updated 3 weeks ago
- ☆16May 27, 2026Updated 2 months ago
- MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation☆66Jul 8, 2026Updated last month
- Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs☆101Jul 26, 2026Updated 2 weeks ago
- ☆17Apr 16, 2026Updated 3 months ago