A tutorial on modern GPU programming for machine learning systems
☆1,290Sep 3, 2026Updated 2 weeks ago
Alternatives and similar repositories for modern-gpu-programming-for-mlsys
Users that are interested in modern-gpu-programming-for-mlsys are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High-performance GPU kernels written in TIRx.☆103Updated this week
- A kernel library written in tilelang☆1,773Apr 23, 2026Updated 4 months ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,444Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,460Updated this week
- Open sources book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.☆11,984Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.☆1,053Sep 14, 2026Updated last week
- Compact and Agent-Native MoE Training System☆352Updated this week
- ☆458Aug 26, 2026Updated 3 weeks ago
- High Performance LLM Inference Operator Library☆1,164Sep 7, 2026Updated 2 weeks ago
- Distributed Compiler and Optimized Parallel Kernels☆1,547Updated this week
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,104May 17, 2026Updated 4 months ago
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference