Arbitrary Entropy Policy Optimization: Entropy Is Controllable in Reinforcement Fine-tuning
☆17Jan 19, 2026Updated 7 months ago
Alternatives and similar repositories for AEPO
Users that are interested in AEPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- RL with Experience Replay☆59Jul 27, 2025Updated last year
- This is the official repository of the paper "Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Schedulin…☆15Jul 27, 2025Updated last year
- [ICML 2026] ZwZ model family: SOTA fine-grained perception performace; ZoomBench: a new challenging perception benchmark☆187May 4, 2026Updated 3 months ago
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement …☆44Aug 6, 2025Updated last year
- The official repository for guided jailbreak benchmark☆32Jul 28, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A curated list of cutting-edge research papers and resources on Long Chain-of-Thought (CoT) Reasoning with Tools.☆46Dec 17, 2025Updated 8 months ago
- ☆16Sep 4, 2025Updated 11 months ago
- ☆16Jun 25, 2025Updated last year
- Official PyTorch implementation of "Multisize Dataset Condensation" (ICLR'24 Oral)☆17Apr 18, 2024Updated 2 years ago
- ☆16Aug 19, 2026Updated last week
- The official repository of paper "Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models''☆112Aug 15, 2025Updated last year
- On the Feasibility of Cross-Task Transfer with Model-Based Reinforcement Learning☆16Apr 30, 2023Updated 3 years ago
- ☆21May 28, 2026Updated 3 months ago
- ☆45May 29, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆18Feb 2, 2026Updated 6 months ago
- 🎭 Explore the magic of face swapping with "Painted Skin"! 🎭 Upload a source picture or video, select a target picture, and let the fun …☆19Nov 9, 2023Updated 2 years ago
- (ICLR 2025) AgentRefine: Enhancing Agent Generalization through Refinement Tuning☆20Nov 22, 2025Updated 9 months ago
- ☆27May 26, 2026Updated 3 months ago
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆18Oct 20, 2025Updated 10 months ago
- [ACM MM 2025] PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning☆18Jun 6, 2026Updated 2 months ago
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- ☆17May 31, 2023Updated 3 years ago
- ☆23May 14, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Benchmarking Open-Ended Inference Optimization by AI Agents☆43Jul 6, 2026Updated last month
- [ICML 2026] GUIEvalKit: Open-source Evaluation Toolkit for GUI Agents☆25Feb 26, 2026Updated 6 months ago
- ☆30Sep 4, 2024Updated last year
- "DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion" code repository☆18Jan 30, 2023Updated 3 years ago
- Pytorch Implementation of MuZero for gym environment. It support any Discrete , Box and Box2D configuration for the action space and obse…☆19Jan 24, 2023Updated 3 years ago
- [ICLR 2026 Blogpost Track Poster] JustRL: Scaling a 1.5B LLM with a Simple RL Recipe☆296Jun 29, 2026Updated 2 months ago
- [ACL 2025] Adaptive Retrieval without Self-Knowledge? Bringing Uncertainty Back Home☆21May 17, 2025Updated last year
- Collect the awesome works evolved around reasoning models like O1/R1 in visual domain☆55Jul 21, 2025Updated last year
- Llemma formal2formal (tactic prediction) theorem proving experiments☆20Oct 17, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2026 main] PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records☆25Apr 11, 2026Updated 4 months ago
- The code of CIKM 2023 (Oral Presentation) : A Multi-Task Semantic Decomposition Framework with Task-specific Pre-training for Few-Shot NE…☆14Jul 19, 2024Updated 2 years ago
- ☆17Jul 12, 2025Updated last year
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning☆26Feb 8, 2026Updated 6 months ago
- DCPO: Dynamic Adaptive Clipping for RL☆50Apr 1, 2026Updated 4 months ago
- [ICLR2026] DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs☆23Oct 19, 2025Updated 10 months ago
- ☆14Dec 18, 2024Updated last year