The Art of Debugging Open Book
☆1,728Sep 3, 2026Updated last month
Alternatives and similar repositories for the-art-of-debugging
Users that are interested in the-art-of-debugging are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Machine Learning Engineering Open Book☆19,080Updated this week
- GPU programming related news and material links☆2,353Jun 15, 2026Updated 3 months ago
- ML/DL Math and Method notes☆68Dec 2, 2023Updated 2 years ago
- Material for gpu-mode lectures☆6,676Sep 9, 2026Updated 3 weeks ago
- Minimalistic large language model 3D-parallelism training☆2,828Sep 23, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Python tools☆14Oct 22, 2023Updated 2 years ago
- Solve puzzles. Improve your pytorch.☆4,343Jul 15, 2024Updated 2 years ago
- What would you do with 1000 H100s...☆1,196Jan 10, 2024Updated 2 years ago
- Tile primitives for speedy kernels☆3,736Sep 12, 2026Updated 3 weeks ago
- Deep learning for dummies. All the practical details and useful utilities that go into working with real models.☆854Aug 12, 2026Updated last month
- A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.☆604Aug 14, 2026Updated last month
- An ML Systems Onboarding list☆1,136Feb 19, 2026Updated 7 months ago
- Puzzles for learning Triton☆2,620Apr 1, 2026Updated 6 months ago
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,315Aug 26, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Solve puzzles. Learn CUDA.☆12,505Sep 1, 2024Updated 2 years ago
- Pragmatic approach to parsing import profiles for CI's☆12Jul 1, 2024Updated 2 years ago
- A PyTorch native platform for training generative AI models☆5,775Updated this week
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,207May 17, 2026Updated 4 months ago
- Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.☆6,259Aug 22, 2025Updated last year
- PyTorch Single Controller☆1,078Updated this week
- Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and…☆2,032Updated this week
- ☆359Updated this week
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard a…☆2,149Dec 3, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- MoE training for Me and You and maybe other people☆399Mar 15, 2026Updated 6 months ago
- Quantized LLM training in pure CUDA/C++.☆261Aug 31, 2026Updated last month
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆3,946Sep 12, 2026Updated 3 weeks ago
- ☆592Jul 11, 2024Updated 2 years ago
- Development repository for the Triton language and compiler☆20,286Updated this week
- A playbook for systematically maximizing the performance of deep learning models.☆30,349Jun 18, 2024Updated 2 years ago
- My solutions for Advanced Python Mastery (course by @dabeaz)☆11Jan 29, 2024Updated 2 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,718Updated this week
- Nano vLLM☆15,689Apr 26, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Minimal example scripts of the Hugging Face Trainer, focused on staying under 150 lines☆196May 6, 2024Updated 2 years ago
- FlashInfer: Kernel Library for LLM Serving☆6,532Updated this week
- Puffing up reinforcement learning☆6,492Sep 13, 2026Updated 2 weeks ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆252Jun 29, 2026Updated 3 months ago
- My learning notes for ML SYS.☆7,421Sep 20, 2026Updated last week
- Efficient Triton Kernels for LLM Training☆6,643Updated this week
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,518Sep 23, 2026Updated last week