This repository is a curated collection of resources, tutorials, and practical examples designed to guide you through the journey of mastering CUDA programming. Whether you're just starting or looking to optimize and scale your GPU-accelerated applications.
☆467Feb 22, 2025Updated last year
Alternatives and similar repositories for cuda-learning
Users that are interested in cuda-learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated list of resources for learning and exploring Triton, OpenAI's programming language for writing efficient GPU code.☆496Mar 10, 2025Updated last year
- A 120-day CUDA learning plan covering daily concepts, exercises, pitfalls, and references (including “Programming Massively Parallel Proc…☆950Mar 29, 2025Updated last year
- GPU Kernels☆225Apr 27, 2025Updated last year
- ☆441Apr 10, 2025Updated last year
- 100 days of building GPU kernels!☆626Apr 27, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆46May 24, 2025Updated last year
- Building GPT ...☆18Dec 1, 2024Updated last year
- This repository documents my 100-day journey of learning and writing CUDA kernels.☆35Mar 29, 2026Updated 4 months ago
- Complete solutions to the Programming Massively Parallel Processors Edition 4☆833Jun 18, 2025Updated last year
- Puzzles for learning Triton☆2,562Apr 1, 2026Updated 4 months ago
- ☆3,914Mar 11, 2026Updated 5 months ago
- ☆11Aug 4, 2025Updated last year
- Package for data-driven and phenomenological gravitational waveform models☆12Aug 4, 2026Updated last week
- Minimalistic, hackable PyTorch implementation of SimSiam in ~400 lines. Achieves good performance on ImageNet with ResNet50. Features dis…☆22Nov 25, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Apply GPU in ML and DL☆69Mar 23, 2026Updated 4 months ago
- Verifying Cairo Programs in SP1☆14Oct 16, 2024Updated last year
- From Minimal GEMM to Everything☆234Jul 9, 2026Updated last month
- Learnings and programs related to CUDA☆439Jun 29, 2025Updated last year
- Astronomy Research + JAX Meeting 2024☆13Oct 23, 2024Updated last year
- GPU programming related news and material links☆2,274Jun 15, 2026Updated 2 months ago
- Exploring how optimizations for GEMMs work☆36Feb 28, 2026Updated 5 months ago
- Learn CUDA with PyTorch☆366Jun 1, 2026Updated 2 months ago
- Code for the study "Gravitational wave inference on a numerical-relativity simulation of a black hole merger beyond general relativity"☆11Jan 10, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆15Feb 23, 2025Updated last year
- ☆14Mar 29, 2026Updated 4 months ago
- GWInferno: Gravitational-Wave Hierarchical Inference with NumPyro☆21Jun 23, 2026Updated last month
- ☆355Updated this week
- Fastest kernels written from scratch☆591Sep 18, 2025Updated 10 months ago
- Performing parameter estimation on gravitational wave data with machine learning☆21May 4, 2026Updated 3 months ago
- learningggggggg 🐳☆635Apr 2, 2025Updated last year
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming☆796Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,280Aug 26, 2025Updated 11 months ago
- ☆259Jan 2, 2025Updated last year
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 8 months ago
- Port of Karpathy's micrograd in pure C. Micrograd is a tiny scalar-valued autograd engine and a neural net library on top of it with PyTo…☆36Jul 27, 2024Updated 2 years ago
- Material for gpu-mode lectures☆6,441Jun 15, 2026Updated 2 months ago
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆36Mar 20, 2026Updated 4 months ago
- Computing solutions to the frequency-domain radial Teukolsky equation with the Generalized Sasaki-Nakamura (GSN) formalism in julia☆27Jul 16, 2026Updated 3 weeks ago