This repository is a curated collection of resources, tutorials, and practical examples designed to guide you through the journey of mastering CUDA programming. Whether you're just starting or looking to optimize and scale your GPU-accelerated applications.
☆461Feb 22, 2025Updated last year
Alternatives and similar repositories for cuda-learning
Users that are interested in cuda-learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated list of resources for learning and exploring Triton, OpenAI's programming language for writing efficient GPU code.☆495Mar 10, 2025Updated last year
- A 120-day CUDA learning plan covering daily concepts, exercises, pitfalls, and references (including “Programming Massively Parallel Proc…☆941Mar 29, 2025Updated last year
- GPU Kernels☆225Apr 27, 2025Updated last year
- ☆440Apr 10, 2025Updated last year
- 100 days of building GPU kernels!☆616Apr 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆46May 24, 2025Updated last year
- Building GPT ...☆18Dec 1, 2024Updated last year
- This repository documents my 100-day journey of learning and writing CUDA kernels.☆35Mar 29, 2026Updated 3 months ago
- Complete solutions to the Programming Massively Parallel Processors Edition 4☆816Jun 18, 2025Updated last year
- Puzzles for learning Triton☆2,540Apr 1, 2026Updated 3 months ago
- ☆3,867Mar 11, 2026Updated 4 months ago
- ☆11Aug 4, 2025Updated 11 months ago
- Package for data-driven and phenomenological gravitational waveform models☆12Jul 6, 2026Updated 2 weeks ago
- Minimalistic, hackable PyTorch implementation of SimSiam in ~400 lines. Achieves good performance on ImageNet with ResNet50. Features dis…☆22Nov 25, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Apply GPU in ML and DL☆68Mar 23, 2026Updated 4 months ago
- A Wadler–Lindig pretty printer for Python☆48Apr 20, 2026Updated 3 months ago
- From Minimal GEMM to Everything☆227Jul 9, 2026Updated 2 weeks ago
- Verifying Cairo Programs in SP1☆14Oct 16, 2024Updated last year
- Learnings and programs related to CUDA☆439Jun 29, 2025Updated last year
- Astronomy Research + JAX Meeting 2024☆13Oct 23, 2024Updated last year
- GPU programming related news and material links☆2,239Jun 15, 2026Updated last month
- Exploring how optimizations for GEMMs work☆36Feb 28, 2026Updated 4 months ago
- ☆15Feb 23, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Awesome things around client-side GPU ecosystems.☆18Mar 23, 2026Updated 4 months ago
- ☆14Mar 29, 2026Updated 3 months ago
- GWInferno: Gravitational-Wave Hierarchical Inference with NumPyro☆21Jun 23, 2026Updated last month
- ☆350Jul 16, 2026Updated last week
- Fastest kernels written from scratch☆586Sep 18, 2025Updated 10 months ago
- Performing parameter estimation on gravitational wave data with machine learning☆21May 4, 2026Updated 2 months ago
- learningggggggg 🐳☆634Apr 2, 2025Updated last year
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming☆780Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,258Aug 26, 2025Updated 11 months ago
- ☆257Jan 2, 2025Updated last year
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 8 months ago
- Port of Karpathy's micrograd in pure C. Micrograd is a tiny scalar-valued autograd engine and a neural net library on top of it with PyTo…☆36Jul 27, 2024Updated last year
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆36Mar 20, 2026Updated 4 months ago
- Computing solutions to the frequency-domain radial Teukolsky equation with the Generalized Sasaki-Nakamura (GSN) formalism in julia☆27Jul 16, 2026Updated last week
- Open deep learning compiler stack for cpu, gpu and specialized accelerators☆20Updated this week