Building the Virtuous Cycle for AI-driven LLM Systems
☆261May 1, 2026Updated 2 months ago
Alternatives and similar repositories for flashinfer-bench
Users that are interested in flashinfer-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernels☆178Apr 26, 2026Updated 3 months ago
- A benchmark of real-world DL kernel problems☆263Jul 15, 2026Updated last week
- a size profiler for cuda binary☆71Jan 15, 2026Updated 6 months ago
- Open ABI and FFI for Machine Learning Systems☆436Updated this week
- Compact and Agent-Native MoE Training System☆299Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Automated High-Performance GPU Kernel Generation☆120Jun 1, 2026Updated last month
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆490Jul 15, 2026Updated last week
- Ship correct and fast LLM kernels to PyTorch☆151Jan 14, 2026Updated 6 months ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,156Mar 24, 2026Updated 4 months ago
- Speed of Light Analysis for ML Model Runtime☆108Jun 10, 2026Updated last month
- ☆32Jul 2, 2025Updated last year
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆21Nov 28, 2025Updated 7 months ago
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,390Updated this week
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆67Jun 24, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated 11 months ago
- ☆65Apr 26, 2025Updated last year
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆218Dec 24, 2025Updated 7 months ago
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆106Apr 7, 2026Updated 3 months ago
- ☆314Jun 9, 2026Updated last month
- FlashInfer: Kernel Library for LLM Serving☆6,032Updated this week
- An Optimizer for Nvidia Compilers.☆110Jul 3, 2026Updated 3 weeks ago
- Distributed Compiler based on Triton for Parallel Systems☆1,498Updated this week
- A Quirky Assortment of CuTe Kernels☆1,070Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Accelerating MoE with IO and Tile-aware Optimizations☆732Jul 4, 2026Updated 3 weeks ago
- mKernel: fast multi-node, multi-GPU fused kernels☆255Jun 21, 2026Updated last month
- NVIDIA cuTile learn☆169Dec 9, 2025Updated 7 months ago
- FLA but cuTile☆27Apr 17, 2026Updated 3 months ago
- Tile-based language built for AI computation across all scales☆176Jun 16, 2026Updated last month
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆75Dec 11, 2025Updated 7 months ago
- ☆52May 19, 2025Updated last year
- ☆775Jun 2, 2026Updated last month
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆137Jun 14, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆44Jul 1, 2026Updated 3 weeks ago
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆489Jul 5, 2026Updated 2 weeks ago
- Skills for writing tilelang and debugging with CUDA toolkits.☆133May 20, 2026Updated 2 months ago
- CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.☆535Updated this week
- ☆37Aug 7, 2025Updated 11 months ago
- DeeperGEMM: crazy optimized version☆86May 5, 2025Updated last year
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling☆32Updated this week