🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
☆4,457Sep 3, 2026Updated 3 weeks ago
Alternatives and similar repositories for hands-on-modern-rl
Users that are interested in hands-on-modern-rl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Harness engineering beginner tutorial, from 0 to 1☆16,373Aug 26, 2026Updated last month
- slime is an LLM post-training framework for RL Scaling.☆8,547Updated this week
- 🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based,…☆4,619Jul 31, 2026Updated last month
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,653Updated this week
- 🧠 Train a 64M-parameter LLM from scratch in just 2h!☆62,739Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."☆17,901Sep 21, 2026Updated last week
- My learning notes for ML SYS.☆7,407Sep 20, 2026Updated last week
- Nano vLLM☆15,636Apr 26, 2026Updated 5 months ago
- Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.☆12,010Updated this week
- Sutskever 30 implementations inspired by https://papercode.vercel.app/ | For Agents, use https://github.com/pageman/Sutskever-Agent | Pol…☆4,663Mar 15, 2026Updated 6 months ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,182May 17, 2026Updated 4 months ago
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asy…☆10,045Sep 17, 2026Updated last week
- Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1☆77,675Aug 26, 2026Updated last month
- Awesome List for Agentic RL☆1,850Sep 15, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,796Updated this week
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,454Nov 13, 2025Updated 10 months ago
- Bub it. Build it. A tiny agent runtime, composable with plugins.☆1,674Updated this week
- This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models …☆292Aug 24, 2026Updated last month
- OpenClaw-RL: Train any agent simply by talking☆5,705May 23, 2026Updated 4 months ago
- Implement a ChatGPT-like LLM in PyTorch from scratch, step by step☆105,652Updated this week
- The best ChatGPT that $100 can buy.☆58,286Sep 7, 2026Updated 3 weeks ago
- 🛠️ Awesome tools & guides for harness engineering.☆4,175Aug 19, 2026Updated last month
- 📚 从零开始构建大模型☆34,081Aug 8, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- An agentic-first and HuggingFace-native RL framework for research (9k lines).☆1,156Updated this week
- SGLang is a high-performance serving framework for large language models and multimodal models.☆36,481Updated this week
- AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI☆109,724Updated this week
- ☆389May 4, 2026Updated 4 months ago
- awesome autoresearch list☆764Updated this week
- AI agents running research on single-GPU nanochat training automatically☆96,869Mar 26, 2026Updated 6 months ago
- Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)☆75,082Sep 14, 2026Updated 2 weeks ago
- Democratizing Reinforcement Learning for LLMs☆5,841Sep 12, 2026Updated 2 weeks ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆92,771Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)☆25,115Jul 19, 2026Updated 2 months ago
- 本人的科研经验☆14,330Sep 13, 2026Updated 2 weeks ago
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL…☆15,735Updated this week
- A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models☆850Aug 30, 2026Updated 3 weeks ago
- 👀 Train a 65M-parameter VLM from scratch in just 2h!☆8,692Updated this week
- Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/g…☆13,074Jun 16, 2026Updated 3 months ago
- Implement a reasoning LLM in PyTorch from scratch, step by step☆5,300Updated this week