Codes for the paper "BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping" by Zhiheng Xi et al.
☆94Jan 29, 2026Updated 5 months ago
Alternatives and similar repositories for BAPO
Users that are interested in BAPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Nex General Agentic Data Pipeline, an end-to-end pipeline for generating high-quality agentic training data.☆36Nov 19, 2025Updated 8 months ago
- HTML Agent based on NexAU☆16Nov 20, 2025Updated 8 months ago
- Nex Agent for Agent is a meta-agent system that automatically creates specialized AI agents based on natural language requirements.☆29Nov 18, 2025Updated 8 months ago
- NexRL is an ultra-loosely-coupled LLM post-training framework.☆114Updated this week
- ☆116Dec 5, 2025Updated 7 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- NexDR (Nex Deep Research), a leading deep research agent that autonomously investigates complex topics and generates rich, structured rep…☆36Dec 4, 2025Updated 7 months ago
- Use the tokenizer in parallel to achieve superior acceleration☆20Mar 21, 2024Updated 2 years ago
- Code and resources for the NeurIPS 2025 Paper "BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset" by Zhiheng X…☆18Oct 14, 2025Updated 9 months ago
- [EMNLP 2025] Distill Visual Chart Reasoning Ability from LLMs to MLLMs☆61Aug 25, 2025Updated 11 months ago
- Safety-J: Evaluating Safety with Critique☆16Jul 28, 2024Updated last year
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement …☆44Aug 6, 2025Updated 11 months ago
- ☆15Apr 17, 2026Updated 3 months ago
- Python SDK for Weaver.☆17Updated this week
- Implementation of the ICML 2024 paper "Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning" pr…☆116Feb 9, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning☆34May 25, 2026Updated 2 months ago
- Archer2.0 evolves from its predecessor by introducing ASPO, which overcomes fundamental PPO-Clip limitations to prevent premature converg…☆31Oct 10, 2025Updated 9 months ago
- Towards Systematic Measurement for Long Text Quality☆39Sep 5, 2024Updated last year
- ☆48Mar 25, 2025Updated last year
- ☆59Sep 2, 2024Updated last year
- code for Scaling Laws of RoPE-based Extrapolation☆73Oct 16, 2023Updated 2 years ago
- Pre-trained, Scalable, High-performance Reward Models via Policy Discriminative Learning.☆166Sep 23, 2025Updated 10 months ago
- ☆135May 12, 2026Updated 2 months ago
- We introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that S…☆315Jun 21, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ACL 2024] Unveiling Linguistic Regions in Large Language Models☆34Jun 9, 2024Updated 2 years ago
- [AAAI 2024] LLMEval Phase I dataset — 17 categories, 453 questions, 2186 annotators for Chinese LLM evaluation☆114May 21, 2026Updated 2 months ago
- RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment☆18Dec 19, 2024Updated last year
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning.☆24Oct 7, 2025Updated 9 months ago
- [AAAI 2024] LLMEval Phase II dataset — professional domain evaluation across 12 academic disciplines☆71May 21, 2026Updated 2 months ago
- a survey of long-context LLMs from four perspectives, architecture, infrastructure, training, and evaluation☆62Mar 31, 2025Updated last year
- Repository of the paper ''CritiQ: Mining Data Quality Criteria from Human Preferences". Code for CritiQ Flow & Training CritiQ Scorer.☆22Dec 11, 2025Updated 7 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning☆401Mar 30, 2026Updated 3 months ago
- ☆65Mar 30, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhi…☆817May 30, 2026Updated last month
- ☆13Jul 14, 2024Updated 2 years ago
- ☆44Sep 19, 2024Updated last year
- Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcemen…☆820Feb 15, 2026Updated 5 months ago
- The official repository of paper "Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models''☆113Aug 15, 2025Updated 11 months ago
- RL with Experience Replay☆58Jul 27, 2025Updated 11 months ago
- [ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning☆29Oct 14, 2025Updated 9 months ago