Accelerating RL for LLM Reasoning with Optimal Advantage Regression
☆41May 30, 2025Updated last year
Alternatives and similar repositories for A-PO
Users that are interested in A-PO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆30Feb 24, 2026Updated 6 months ago
- A simple implementation of ReasonGenRM.☆19Apr 21, 2025Updated last year
- ☆31May 20, 2026Updated 4 months ago
- ☆83Jun 8, 2026Updated 3 months ago
- The official code release for Q#: Provably Optimal Distributional RL for LLM Post-Training☆21Mar 4, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Learning from Mixed Rollouts: Logit Fusion as a Bridge Between Imitation and Exploration☆18Feb 24, 2026Updated 6 months ago
- Optimizing Review Generation Through Prompt Generation☆17Apr 15, 2024Updated 2 years ago
- ☆22Jun 4, 2025Updated last year
- ☆31Mar 2, 2023Updated 3 years ago
- Source code to accompany research paper on training multi token prediction language models using self-distillation.☆42Feb 21, 2026Updated 7 months ago
- Code and data used in the paper: "Training on Incorrect Synthetic Data via RL Scales LLM Math Reasoning Eight-Fold"☆32Jun 16, 2024Updated 2 years ago
- Arbitrary Entropy Policy Optimization: Entropy Is Controllable in Reinforcement Fine-tuning☆18Jan 19, 2026Updated 8 months ago
- Official implementation of TBA for async LLM post-training.☆32Nov 5, 2025Updated 10 months ago
- ☆30Jan 31, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [TACL, EMNLP 2025 Oral] Code, datasets, and checkpoints for the paper "CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Thr…☆35Dec 5, 2025Updated 9 months ago
- ☆15Apr 25, 2025Updated last year
- [ICML 2025] Official code of "AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization"☆32Jan 10, 2026Updated 8 months ago
- Fully open reproduction of DeepSeek-R1☆11Mar 24, 2025Updated last year
- ☆15Apr 26, 2025Updated last year