[ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)
☆76Jun 29, 2026Updated 3 weeks ago
Alternatives and similar repositories for RPG
Users that are interested in RPG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of TBA for async LLM post-training.☆31Nov 5, 2025Updated 8 months ago
- The code for paper "EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning"☆40Jul 13, 2026Updated last week
- [ICML 2025] Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment (https://arxiv.org/abs/2410.02197)☆43Jun 15, 2026Updated last month
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆54Jul 15, 2025Updated last year
- ☆32Oct 8, 2025Updated 9 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Self-Hinting Language Models Enhance Reinforcement Learning☆26Mar 28, 2026Updated 3 months ago
- The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.☆443Jul 11, 2025Updated last year
- Reinforcing General Reasoning without Verifiers☆102Jun 24, 2025Updated last year
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.