Recreating PyTorch from scratch (C/C++, CUDA, NCCL and Python, with multi-GPU support and automatic differentiation!)
☆168Nov 25, 2025Updated 10 months ago
Alternatives and similar repositories for PyNorch
Users that are interested in PyNorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆47May 24, 2025Updated last year
- A really tiny autograd engine☆99May 26, 2025Updated last year
- A simple implementation of Llama 1, 2. Llama Architecture built from scratch using PyTorch all the models are built from scratch that inc…☆14May 6, 2024Updated 2 years ago
- Contrastive Reinforcement Learning☆69Updated this week
- Implementation of FlashAttention (FA1-FA4) in PyTorch for educational and algorithmic clarity☆225Apr 12, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Trained a 114 million Parameter LLM from Scratch.☆19Jul 21, 2024Updated 2 years ago
- Reference implementation of Mistral AI 7B v0.1 model.☆30Dec 25, 2023Updated 2 years ago
- ☆94Jul 5, 2024Updated 2 years ago
- LLM training parallelisms (DP, FSDP, TP, PP) in pure C☆29Jan 27, 2026Updated 8 months ago
- 6th place solution☆22Feb 28, 2023Updated 3 years ago
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,311Aug 26, 2025Updated last year
- Optimize tensor program fast with Felix, a gradient descent autotuner.☆33Mar 5, 2026Updated 6 months ago
- Comprehensive CUDA tutorials for Maths & ML with examples☆237Jun 11, 2025Updated last year
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆16Jul 4, 2026Updated 2 months ago
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.☆21Jan 24, 2025Updated last year
- High-Performance FP32 GEMM on CUDA devices☆129Jan 21, 2025Updated last year
- 🥪 Mess portal where owners can set their weekly menu, price, time, and students can purchase their desired coupons, with a QR code syste…☆12Jun 2, 2023Updated 3 years ago
- 📑 Dive into Big Model Training☆116Dec 1, 2022Updated 3 years ago
- ☆17Apr 5, 2019Updated 7 years ago
- A practical skill that turns a repository into a beginner-friendly, single-file architecture map with layered dependency and framework-fl…☆26Mar 11, 2026Updated 6 months ago
- A Easy-to-understand TensorOp Matmul Tutorial☆453Mar 5, 2026Updated 6 months ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆26Jan 30, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 使用 cutlass 实现 flash-attention 精简版,具有教学意义☆59Aug 12, 2024Updated 2 years ago
- flash attention tutorial written in python, triton, cuda, cutlass☆535Jan 20, 2026Updated 8 months ago
- rl from zero pretrain, can it be done? yes.☆296Sep 28, 2025Updated last year
- Find, list, and inspect processes from Go (golang).☆10Feb 4, 2018Updated 8 years ago
- Code for the article series on building a Python compiler and interpreter☆14Feb 13, 2025Updated last year
- A PyTorch implementation of the GPT-OSS-20B architecture. All components are coded from scratch: RoPE with YaRN, RMSNorm, SwiGLU with cla…☆237Dec 2, 2025Updated 9 months ago
- A collection of reusable, high-performance, well-documented, thorough-tested layers and models in Jax☆23Jun 8, 2025Updated last year
- A simple Python tool to measure the performance of ONNX models.☆27Sep 15, 2024Updated 2 years ago
- A zero-dependency ML framework in C with a modern Python API for full control over execution and memory.☆707Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implement custom operators in PyTorch with cuda/c++☆77Jan 1, 2023Updated 3 years ago
- ☆121Mar 18, 2026Updated 6 months ago
- To create a primary vision-based self-driving car in GTA5.☆13Oct 18, 2018Updated 7 years ago
- Code for ACL22 short Paper "Hierarchical Curriculum Learning for AMR Parsing"☆13Jun 1, 2022Updated 4 years ago
- Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O☆601Sep 13, 2025Updated last year
- Repository for the COLM 2025 paper SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths☆19Jul 10, 2025Updated last year
- ☆11Jan 24, 2025Updated last year