Minimalistic 4D-parallelism distributed training framework for education purpose
☆2,305Aug 26, 2025Updated last year
Alternatives and similar repositories for picotron
Users that are interested in picotron are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minimalistic large language model 3D-parallelism training☆2,825Updated this week
- ☆259Nov 24, 2025Updated 9 months ago
- A PyTorch native platform for training generative AI models☆5,746Updated this week
- The simplest, fastest repository for training/finetuning small-sized VLMs.☆5,027Oct 27, 2025Updated 10 months ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,093May 17, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Efficient Triton Kernels for LLM Training☆6,618Updated this week
- NanoGPT (124M) in 90 seconds☆5,803Updated this week
- Nano vLLM☆15,505Apr 26, 2026Updated 4 months ago
- Tile primitives for speedy kernels☆3,716Updated this week
- FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference.☆359Nov 2, 2025Updated 10 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,499Updated this week
- Implementing DeepSeek R1's GRPO algorithm from scratch☆1,897Apr 18, 2025Updated last year
- 🚀 Efficient implementations for emerging model architectures☆5,764Updated this week
- Our library for RL environments + evals☆4,631Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,487Updated this week
- My learning notes for ML SYS.☆7,369Sep 10, 2026Updated last week
- FlashInfer: Kernel Library for LLM Serving☆6,443Updated this week
- Simple MPI implementation for prototyping or learning☆329Aug 6, 2025Updated last year
- Puzzles for learning Triton☆2,599Apr 1, 2026Updated 5 months ago
- What would you do with 1000 H100s...☆1,198Jan 10, 2024Updated 2 years ago
- Distributed Compiler and Optimized Parallel Kernels☆1,546Updated this week
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆2,937Updated this week
- Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.☆4,766Jul 18, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- PyTorch native quantization for training and inference☆2,980Updated this week
- Ongoing research training transformer models at scale☆17,939Updated this week
- Material for gpu-mode lectures☆6,615Sep 9, 2026Updated last week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,140Updated this week
- Agentic RL Training at Scale☆2,051Updated this week
- Helpful tools and examples for working with flex-attention☆1,247Updated this week
- Train transformer language models with reinforcement learning.☆19,339Updated this week
- ☆1,312May 20, 2026Updated 3 months ago
- Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.☆6,253Aug 22, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on H…☆3,540Updated this week
- Everything about the SmolLM and SmolVLM family of models☆3,894Updated this week
- Ring attention implementation with flash attention☆1,057Sep 10, 2025Updated last year
- A scalable asynchronous reinforcement learning implementation with in-flight weight updates.☆433Aug 5, 2026Updated last month
- GPU programming related news and material links☆2,331Jun 15, 2026Updated 3 months ago
- 🚀 Efficiently (pre)training foundation models with native PyTorch features, including FSDP for training and SDPA implementation of Flash…☆288Nov 24, 2025Updated 9 months ago
- Build compute kernels and load them from the Hub.☆745Updated this week