A comprehensive hands-on project for learning GPU programming with CUDA and HIP, covering fundamental concepts through advanced optimization techniques.
☆38Nov 20, 2025Updated 8 months ago
Alternatives and similar repositories for gpu-programming-101
Users that are interested in gpu-programming-101 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Dynamo Workshop☆19Nov 7, 2025Updated 9 months ago
- An end-to-end pipeline to optimize and host LLM for 100K parallel queries☆37Jul 6, 2025Updated last year
- Source code for the CPU-Free model - a fully autonomous execution model for multi-GPU applications that completely excludes the involveme…☆21Apr 25, 2024Updated 2 years ago
- This repository provides the code for applying Contrastive Learning Penalty Loss (CLPL) and Mixture of Experts (MoE) to the BGE-M3 text e…☆11Dec 27, 2024Updated last year
- KFunca: A minimalist, high-performance GPU-based automatic differentiation framework☆31Aug 14, 2025Updated 11 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Microbenchmark that unveals the mechanisms behind power readings reported by nvidia-smi on your NVIDIA GPU.☆15Dec 12, 2024Updated last year
- Companion code for Grokking Megakernels: fuse an entire LLM forward pass into a single CUDA kernel☆23Feb 9, 2026Updated 6 months ago
- learn socket api☆14Jun 14, 2024Updated 2 years ago
- Three.js gradient pills effect using TSL and WebGPU☆16Jan 20, 2026Updated 6 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated 11 months ago
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated last year
- Jacta: A Versatile Planner for Learning Dexterous and Whole-body Manipulation☆15Nov 10, 2025Updated 9 months ago
- BERT&RoBERTa预训练代码,tensorflow和torch两种版本实现☆13Feb 8, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.☆17Aug 19, 2025Updated 11 months ago
- Self-Augmented Robot Trajectory: Efficient Imitation Learning via Safe Self-augmentation with Demonstrator-annotated Precision☆15Nov 12, 2025Updated 9 months ago
- Options pricing, Greeks, strategy P&L, volatility surfaces, and scenario analysis.☆14Mar 29, 2025Updated last year
- Can LLMs Write Correct and Efficient GPU Communication Code?☆59Jul 7, 2026Updated last month
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆195Updated this week
- Python library to compress LitGPT models for resource efficient inference.☆16Jul 31, 2026Updated last week
- Challenging myself to learn CUDA (Basics → Intermediate) these 100 days.☆37Mar 2, 2026Updated 5 months ago
- A programmable, explicit world model on Godot — worlds are pure JSON run by a fixed primitives + interpreter engine. Built by Claude, for…☆20Jun 30, 2026Updated last month
- Synthetic data generation for evaluating LLM symbolic and logic reasoning☆23Mar 6, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 3.34× faster inference on Apple Silicon — native MLX port of DFlash speculative decoding☆19Apr 11, 2026Updated 4 months ago
- Build an LLM from scratch with MAX☆65Updated this week
- A step by step implementation of building an AI agent that plays 3d shooting game☆22Jul 16, 2025Updated last year
- A Beginner's Guide to Monetizing Your Python AI Chatbot☆19Apr 22, 2025Updated last year
- A simple implementation of Llama 1, 2. Llama Architecture built from scratch using PyTorch all the models are built from scratch that inc…☆14May 6, 2024Updated 2 years ago
- Official repo for "Generative Point Tracking with Flow Matching".☆24Oct 22, 2025Updated 9 months ago
- A Tiny, Pure Python implementation of Gradient Boosted Trees.☆14Dec 28, 2022Updated 3 years ago
- Minimal implementation of a Byte Pair Encoding (BPE) tokenizer in Zig☆15Apr 7, 2025Updated last year
- The Docker Compose file for quick setup a monitoring solution based on InfluxDB and Grafana.☆13Feb 8, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A Compute Express Link (CXL) Benchmark Suite☆22Feb 12, 2025Updated last year
- Polygon Studio is a 3D modeling app for Snap's AR Spectacles.☆24Oct 16, 2024Updated last year
- Inference Llama/Llama2/Llama3 Modes in NumPy☆21Nov 22, 2023Updated 2 years ago
- ML tools that we use internally and which you may find useful too.☆26Apr 27, 2022Updated 4 years ago
- TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages☆27Feb 25, 2025Updated last year
- A curated list of Large Language Model resources, covering model training, serving, fine-tuning, and building LLM applications.☆17Nov 8, 2024Updated last year
- ☆16Feb 6, 2024Updated 2 years ago