minimal pytorch 4D parallelism
☆74Feb 16, 2026Updated 7 months ago
Alternatives and similar repositories for heiretsu
Users that are interested in heiretsu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Oct 10, 2025Updated last year
- NanoGPT-speedrunning for the poor T4 enjoyers☆72Apr 22, 2025Updated last year
- MoE training for Me and You and maybe other people☆400Mar 15, 2026Updated 6 months ago
- Well documented examples of running distributed training jobs on Modal☆35Sep 25, 2026Updated 2 weeks ago
- ☆47May 24, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆25May 23, 2025Updated last year
- Instant Neural Graphics Primitives from scratch, zero dependencies. Learning by doing.☆10Aug 18, 2023Updated 3 years ago
- a fun and educational take on vLLM☆229Jan 25, 2026Updated 8 months ago
- An experiment to see if chatgpt can improve the output of the stanford alpaca dataset☆12Mar 29, 2023Updated 3 years ago
- Repository for GPU related kernels for learning/testing purposes☆21May 27, 2026Updated 4 months ago
- rl from zero pretrain, can it be done? yes.☆297Sep 28, 2025Updated last year
- Hand-Rolled GPU communications library☆95Nov 25, 2025Updated 10 months ago
- Code companion for the RL Post-Training Handbook - training reasoning models on a single GPU☆20Jan 30, 2026Updated 8 months ago
- Program that produces retail/wholesale trade statistics using machine learning + visualisations☆17Aug 21, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimal and highly hackable implementation of Looped Transformers with GPT☆25Mar 8, 2026Updated 7 months ago
- Course on Flash-attention in Triton☆104Feb 9, 2026Updated 8 months ago
- Educational distributed training and inference library for heterogeneous local hardware: Mac minis, Raspberry Pis and GPUs over plain Pyt…☆87Updated this week
- Alternative approach for Adaptive Computation Time in TensorFlow☆19Mar 24, 2023Updated 3 years ago
- Inline PTX Assembly in CUDA example☆15Updated this week
- implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880☆378Feb 17, 2026Updated 7 months ago
- CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs☆256Updated this week
- XmodelLM☆38Nov 19, 2024Updated last year
- ☆25Jan 28, 2026Updated 8 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆263Nov 24, 2025Updated 10 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆20Apr 18, 2026Updated 5 months ago
- Use Muon optimizer instead of AdamW.☆50Mar 2, 2026Updated 7 months ago
- WhaleSearch 🐳 - semantic search engine. (SSE)☆16Jan 4, 2025Updated last year
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆15Aug 31, 2023Updated 3 years ago
- Tensor-Slayer : Manipulate weights and tensors of LLMs to achieve performance upgrades and introduce a novel inferenceless mechanistic in…☆28May 27, 2025Updated last year
- ☆32Nov 4, 2024Updated last year
- ☆13Apr 16, 2025Updated last year
- UCSD ECE277 GPU Programming coursework: GPU-accelerated reinforcement learning on CUDA C with Nsight System☆15Aug 17, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,321Aug 26, 2025Updated last year
- 100M tokens. Infinite compute. Lowest val loss wins.☆560Sep 15, 2026Updated 3 weeks ago
- Ludic – an LLM-RL library for the era of experience☆67Aug 9, 2026Updated 2 months ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,237May 17, 2026Updated 4 months ago
- The best ChatGPT that $100 can buy.☆59Oct 1, 2026Updated last week
- a LLM inference engine to run on consumer hardware☆48Apr 15, 2026Updated 5 months ago
- a personal collection of my notes for ml sys☆124Sep 1, 2026Updated last month