Slides, notes, and materials for the workshop
☆344Jun 1, 2024Updated 2 years ago
Alternatives and similar repositories for gpu-optimization-workshop
Users that are interested in gpu-optimization-workshop are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Slides and recordings of talks hosted by our community☆21Jun 21, 2024Updated 2 years ago
- ☆80Jun 25, 2024Updated 2 years ago
- GPU programming related news and material links☆2,336Jun 15, 2026Updated 3 months ago
- Material for gpu-mode lectures☆6,648Sep 9, 2026Updated 2 weeks ago
- PyTorch native quantization for training and inference☆2,986Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆47Jun 1, 2023Updated 3 years ago
- Fine-tune an LLM to perform batch inference and online serving.☆124May 29, 2025Updated last year
- An ML Systems Onboarding list☆1,133Feb 19, 2026Updated 7 months ago
- ☆92Feb 29, 2024Updated 2 years ago
- this is a rust project☆12Jan 19, 2023Updated 3 years ago
- Puzzles for learning Triton☆2,609Apr 1, 2026Updated 5 months ago
- Puzzles for exploring transformers☆403May 4, 2023Updated 3 years ago
- See https://github.com/cuda-mode/triton-index/ instead!☆11May 8, 2024Updated 2 years ago
- Deep learning for dummies. All the practical details and useful utilities that go into working with real models.☆853Aug 12, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- What would you do with 1000 H100s...☆1,198Jan 10, 2024Updated 2 years ago
- Unconditional music synthesis using a diffusion model in the STFT domain☆12May 31, 2022Updated 4 years ago
- Because it's there.☆16Sep 22, 2024Updated 2 years ago
- Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.☆6,255Aug 22, 2025Updated last year
- Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"☆11Mar 31, 2024Updated 2 years ago
- Solve puzzles. Learn CUDA.☆12,485Sep 1, 2024Updated 2 years ago
- Cataloging released Triton kernels.☆312Sep 9, 2025Updated last year
- https://huyenchip.com/ml-interviews-book/☆4,770Mar 21, 2025Updated last year
- Machine Learning Engineering Open Book☆19,039Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A PyTorch native platform for training generative AI models☆5,764Updated this week
- ☆182Sep 18, 2026Updated last week
- Tutorial Kubernetes Operator☆16Aug 14, 2021Updated 5 years ago
- Tile primitives for speedy kernels☆3,726Sep 12, 2026Updated last week
- Lite weight wrapper for the independent implementation of SPLADE++ models for search & retrieval pipelines. Models and Library created by…☆35Aug 24, 2024Updated 2 years ago
- Full finetuning of large language models without large memory requirements☆92Sep 22, 2025Updated last year
- Efficient Triton Kernels for LLM Training☆6,632Updated this week
- Implementation of a multimodal diffusion transformer in Pytorch☆108Jun 24, 2024Updated 2 years ago
- ring-attention experiments☆176Oct 17, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- LLM training in simple, raw C/CUDA☆31,052Jun 26, 2025Updated last year
- ☆69May 23, 2025Updated last year
- ☆356Updated this week
- Solve puzzles. Improve your pytorch.☆4,339Jul 15, 2024Updated 2 years ago
- Accessible large language models via k-bit quantization for PyTorch.☆8,499Sep 7, 2026Updated 2 weeks ago
- A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.☆604Aug 14, 2026Updated last month
- Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and…☆1,993Sep 17, 2026Updated last week