minimal pytorch 4D parallelism
☆73Feb 16, 2026Updated 7 months ago
Alternatives and similar repositories for heiretsu
Users that are interested in heiretsu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Oct 10, 2025Updated 11 months ago
- NanoGPT-speedrunning for the poor T4 enjoyers☆72Apr 22, 2025Updated last year
- MoE training for Me and You and maybe other people☆400Mar 15, 2026Updated 6 months ago
- Well documented examples of running distributed training jobs on Modal☆34Updated this week
- Companion code for Grokking Megakernels: fuse an entire LLM forward pass into a single CUDA kernel☆30Feb 9, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆25May 23, 2025Updated last year
- An experiment to see if chatgpt can improve the output of the stanford alpaca dataset☆12Mar 29, 2023Updated 3 years ago
- Repository for GPU related kernels for learning/testing purposes☆21May 27, 2026Updated 3 months ago
- rl from zero pretrain, can it be done? yes.☆296Sep 28, 2025Updated 11 months ago
- Hand-Rolled GPU communications library☆94Nov 25, 2025Updated 9 months ago
- Minimal and highly hackable implementation of Looped Transformers with GPT☆25Mar 8, 2026Updated 6 months ago
- PyTorch Lightning based framework to run experiments for self-supervised learning tasks.☆10Feb 14, 2020Updated 6 years ago
- A repository of prompts and Python scripts for intelligent transformation of raw text into diverse formats.☆33May 29, 2023Updated 3 years ago
- An educational distributed training and inference library for neural nets using local computing☆80Jun 10, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- RL models to play Sokoban. The fastest recipe wins.☆31Jul 25, 2026Updated last month
- implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880☆377Feb 17, 2026Updated 7 months ago
- CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs☆250Updated this week
- XmodelLM☆38Nov 19, 2024Updated last year
- ☆25Jan 28, 2026Updated 7 months ago
- ☆259Nov 24, 2025Updated 9 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 5 months ago
- Use Muon optimizer instead of AdamW.☆50Mar 2, 2026Updated 6 months ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆15Aug 31, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Tensor-Slayer : Manipulate weights and tensors of LLMs to achieve performance upgrades and introduce a novel inferenceless mechanistic in…☆28May 27, 2025Updated last year
- ☆32Nov 4, 2024Updated last year
- ☆13Apr 16, 2025Updated last year
- UCSD ECE277 GPU Programming coursework: GPU-accelerated reinforcement learning on CUDA C with Nsight System☆14Aug 17, 2021Updated 5 years ago
- Simple Video Summarization using Text-to-Segment Anything (Florence2 + SAM2) This project provides a video processing tool that utilizes…☆10Feb 20, 2025Updated last year
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,305Aug 26, 2025Updated last year
- 100M tokens. Infinite compute. Lowest val loss wins.☆541Updated this week
- Ludic – an LLM-RL library for the era of experience☆67Aug 9, 2026Updated last month
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,096May 17, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The best ChatGPT that $100 can buy.☆58Updated this week
- a LLM inference engine to run on consumer hardware☆48Apr 15, 2026Updated 5 months ago
- a whirlwind tour to deep learning and deep learning systems☆83Updated this week
- Efficient non-uniform quantization with GPTQ for GGUF☆66Sep 17, 2025Updated last year
- Official repository for Parallax (Parameterized Local Linear Attention)☆69Jul 30, 2026Updated last month
- video description generation vision-language model☆22Jan 21, 2025Updated last year
- RLM for coding agent☆102Feb 19, 2026Updated 7 months ago