a personal collection of my notes for ml sys
☆113Jul 31, 2026Updated this week
Alternatives and similar repositories for ml-systems-notes
Users that are interested in ml-systems-notes are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- a LLM inference engine to run on consumer hardware☆47Apr 15, 2026Updated 3 months ago
- writing really fast kernels☆19Jul 15, 2026Updated 2 weeks ago
- Repository for GPU related kernels for learning/testing purposes☆19May 27, 2026Updated 2 months ago
- Companion code for The Physics of LLM Inference book☆26Apr 21, 2026Updated 3 months ago
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆43May 17, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆19Apr 26, 2026Updated 3 months ago
- a minimal paged attention implementation☆20Jan 30, 2026Updated 6 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 2 months ago
- A Python DSL for Apple Metal GPU compute☆25Jul 23, 2026Updated last week
- A blazingly fast bombparty game bot☆18Oct 10, 2025Updated 9 months ago
- ☆30Sep 16, 2023Updated 2 years ago
- ☆92Dec 16, 2025Updated 7 months ago
- ☆49Jul 14, 2026Updated 3 weeks ago
- Educational WIP☆73Feb 16, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average…☆148Jun 10, 2026Updated last month
- This Repo Contains all the materials of Pattern Recognition and Neural Networks course done at IISc and recorded for NPTEL - Mathematical…☆31Mar 17, 2026Updated 4 months ago
- An educational distributed training and inference library for neural nets using local computing☆74Jun 10, 2026Updated last month
- A PyTorch implementation of the GPT-OSS-20B architecture. All components are coded from scratch: RoPE with YaRN, RMSNorm, SwiGLU with cla…☆238Dec 2, 2025Updated 8 months ago
- ☆25Apr 4, 2026Updated 4 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 5 months ago
- Curated collection of AI inference engineering resources — LLM serving, GPU kernels, quantization, distributed inference, and production …☆259Feb 4, 2026Updated 6 months ago
- Stas' Python Cookbook - Python recipes that I use daily☆57Jul 2, 2026Updated last month
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.☆33Apr 2, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This Repo consists the python note books of IITM - Mathematical Foundations for Generative AI Course,☆401Feb 17, 2026Updated 5 months ago
- ☆83Jul 25, 2026Updated last week
- ☆17Apr 23, 2026Updated 3 months ago
- ☆12Jun 14, 2024Updated 2 years ago
- a simple c++ inference engine for gpt based architecture☆40Dec 10, 2025Updated 7 months ago
- Quantized LLM training in pure CUDA/C++.☆252Jul 23, 2026Updated last week
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆106Updated this week
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,278Apr 27, 2026Updated 3 months ago
- ☆19Mar 29, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A cross-platform, safe, pure-Rust graphics API.☆17May 18, 2026Updated 2 months ago
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tables☆22May 18, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- ☆26Jul 9, 2026Updated 3 weeks ago
- Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)☆86May 29, 2026Updated 2 months ago
- ☆27Dec 31, 2025Updated 7 months ago
- A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.☆78Updated this week