🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
☆4,301Sep 3, 2026Updated this week
Alternatives and similar repositories for hands-on-modern-rl
Users that are interested in hands-on-modern-rl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, Mo…☆199Updated this week
- Harness engineering beginner tutorial, from 0 to 1☆14,918Aug 26, 2026Updated last week
- slime is an LLM post-training framework for RL Scaling.☆8,404Updated this week
- 🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based,…☆4,561Jul 31, 2026Updated last month
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,336Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- 🧠 Train a 64M-parameter LLM from scratch in just 2h!☆59,506Updated this week
- This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."☆17,702Aug 10, 2026Updated 3 weeks ago
- My learning notes for ML SYS.☆7,269Aug 19, 2026Updated 2 weeks ago
- Nano vLLM☆15,333Apr 26, 2026Updated 4 months ago
- Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.☆11,900Updated this week
- Sutskever 30 implementations inspired by https://papercode.vercel.app/ | For Agents, use https://github.com/pageman/Sutskever-Agent | Pol…☆4,546Mar 15, 2026Updated 5 months ago
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asy…☆9,979Aug 13, 2026Updated 3 weeks ago
- Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1☆76,266Aug 26, 2026Updated last week
- Awesome List for Agentic RL☆1,834Aug 28, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,733Updated this week
- Bub it. Build it. A hook-first runtime for agents that live alongside people.☆1,599Updated this week
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,385Nov 13, 2025Updated 9 months ago
- This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models …☆292Aug 24, 2026Updated 2 weeks ago
- Implement a ChatGPT-like LLM in PyTorch from scratch, step by step☆104,525Sep 1, 2026Updated last week
- OpenClaw-RL: Train any agent simply by talking☆5,670May 23, 2026Updated 3 months ago
- The best ChatGPT that $100 can buy.☆57,847Aug 2, 2026Updated last month
- An agentic-first RL framework for research (9k lines).☆1,025Updated this week
- 🛠️ Awesome tools & guides for harness engineering.☆4,003Aug 19, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 📚 从零开始构建大模型☆33,576Aug 8, 2026Updated last month
- SGLang is a high-performance serving framework for large language models and multimodal models.☆35,594Updated this week
- AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI☆102,701Updated this week
- awesome autoresearch list☆726Updated this week
- ☆387May 4, 2026Updated 4 months ago
- AI agents running research on single-GPU nanochat training automatically☆95,351Mar 26, 2026Updated 5 months ago
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)☆74,619Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,816Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆91,170Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)☆25,013Jul 19, 2026Updated last month
- 本人的科研经验☆13,964Jun 6, 2026Updated 3 months ago
- A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models☆793Aug 30, 2026Updated last week
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL…☆15,525Updated this week
- 👀 Train a 65M-parameter VLM from scratch in just 2h!☆8,580Aug 6, 2026Updated last month
- Implement a reasoning LLM in PyTorch from scratch, step by step☆5,171Updated this week
- Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/g…☆12,405Jun 16, 2026Updated 2 months ago