[NeurIPS 2025] RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
☆57Oct 23, 2025Updated 9 months ago
Alternatives and similar repositories for RL-Tango
Users that are interested in RL-Tango are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- Emergent Hierarchical Reasoning in LLMs/VLMs through Reinforcement Learning [ICLR26]☆64Apr 11, 2026Updated 3 months ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 8 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning☆401Mar 30, 2026Updated 3 months ago
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆189Jul 23, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The code for paper "EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning"☆40Jul 13, 2026Updated 2 weeks ago
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆229Nov 27, 2025Updated 8 months ago
- Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization☆82Dec 25, 2025Updated 7 months ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated 11 months ago
- [ICML 2026 Spotlight] Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback☆70Jun 3, 2026Updated last month
- [ICLR 2026] Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs☆46May 20, 2025Updated last year
- A repository of OpenDecoder framework: Open Large Language Model Decoding to Incorporate Document Quality in RAG (WWW 2026)☆26Jan 27, 2026Updated 6 months ago
- Benchmark Test-Time Scaling of General LLM Agents☆20Apr 14, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation (EVOL-RL).☆51Mar 31, 2026Updated 3 months ago
- ☆29Jan 31, 2026Updated 5 months ago
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search in ACL'25☆99Jun 16, 2025Updated last year
- [COLM 2025] Code for Paper: Learning Adaptive Parallel Reasoning with Language Models☆145Dec 17, 2025Updated 7 months ago
- CE-GPPO: Controlling Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning☆16Jan 23, 2026Updated 6 months ago
- ☆139Jun 11, 2025Updated last year
- [ICLR 2026] Learning to Reason without External Rewards☆420Jan 26, 2026Updated 6 months ago
- ☆18Jul 31, 2025Updated 11 months ago
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆34Aug 11, 2025Updated 11 months ago
- ☆56Jul 7, 2025Updated last year
- [COLM 2025] "C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing"☆21Apr 9, 2025Updated last year
- Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"☆65Jan 5, 2026Updated 6 months ago
- ☆82Jun 8, 2026Updated last month
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆15Apr 22, 2026Updated 3 months ago
- Self-Hinting Language Models Enhance Reinforcement Learning☆27Mar 28, 2026Updated 4 months ago
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation. A system-level optimization for scalable LLM tra…☆60Oct 11, 2025Updated 9 months ago
- Official Implementation of the paper "Jointly Reinforcing Diversity and Quality in Language Model Generations"☆61May 8, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 3 months ago
- ☆34Mar 18, 2026Updated 4 months ago
- Socratic-Zero is a fully autonomous framework that generates high-quality training data for mathematical reasoning☆37Oct 26, 2025Updated 9 months ago
- This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DP…☆32Dec 5, 2024Updated last year
- ☆44Mar 31, 2026Updated 3 months ago
- ☆275May 14, 2025Updated last year
- ☆111Jul 6, 2026Updated 3 weeks ago