[AAAI 2026] - Official repo for paper: "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't"
☆292Mar 11, 2026Updated 5 months ago
Alternatives and similar repositories for open-rs
Users that are interested in open-rs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] Reinforcement Learning for Reasoning in Large Language Models with One Training Example☆445Mar 11, 2026Updated 5 months ago
- Simple RL training for reasoning☆3,874Dec 23, 2025Updated 8 months ago
- Scalable RL solution for advanced reasoning of language models☆1,872Mar 18, 2025Updated last year
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆54Jul 15, 2025Updated last year
- Understanding R1-Zero-Like Training: A Critical Perspective☆1,274Aug 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Repo for Open-Reasoner-Zero☆2,100Jun 2, 2025Updated last year
- The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.☆450Jul 11, 2025Updated last year
- Official code for the paper: DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models☆24Jan 6, 2026Updated 7 months ago
- Democratizing Reinforcement Learning for LLMs☆5,808Aug 24, 2026Updated last week
- [TMLR] Process Reward Models That Think☆91Jul 30, 2026Updated last month
- ☆39Nov 18, 2025Updated 9 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning☆405Mar 30, 2026Updated 5 months ago
- Repo of paper "Free Process Rewards without Process Labels"☆172Mar 14, 2025Updated last year
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning☆265May 14, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Recipes to train the self-rewarding reasoning LLMs.☆231Mar 2, 2025Updated last year
- A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning☆301Sep 25, 2025Updated 11 months ago
- An Open-source RL System from ByteDance Seed and Tsinghua AIR☆1,861May 11, 2025Updated last year
- ☆767Dec 23, 2025Updated 8 months ago
- [COLM 2025] LIMO: Less is More for Reasoning☆1,084Jul 30, 2025Updated last year
- Official repository for ACL 2025 paper "ProcessBench: Identifying Process Errors in Mathematical Reasoning"☆192May 20, 2025Updated last year
- A series of technical report on Slow Thinking with LLM