[EMNLP 2025] The official implementation for paper "Agentic-R1: Distilled Dual-Strategy Reasoning"
☆104Apr 21, 2026Updated 3 months ago
Alternatives and similar repositories for DualDistill
Users that are interested in DualDistill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generat…☆22Jul 12, 2026Updated last week
- [ICML2025] Official Repo for Paper "Optimizing Temperature for Language Models with Multi-Sample Inference"☆23Feb 16, 2025Updated last year
- [ICLR2026] The official repository for the CodeGym project: "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym"☆35Oct 14, 2025Updated 9 months ago
- ☆72Oct 23, 2025Updated 9 months ago
- [NeurIPS'25 Spotlight] ARM: Adaptive Reasoning Model☆68Apr 6, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Repository for "Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators"☆12Mar 25, 2025Updated last year
- Official Repo for SwS: A Weakness-driven Problem Synthesis Framework in RL for LLM Reasoning☆42Nov 11, 2025Updated 8 months ago
- ☆45Dec 15, 2025Updated 7 months ago
- [NeurIPS 2025] Official implementation of "Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning"☆32Oct 20, 2025Updated 9 months ago
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆189Jul 23, 2025Updated last year
- Official code and dataset for our paper: RefineBench: Evaluating Refinement Capability of Language Models via Checklists☆17Dec 1, 2025Updated 7 months ago
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated 11 months ago
- Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks☆271May 5, 2025Updated last year
- A curated list of cutting-edge research papers and resources on Long Chain-of-Thought (CoT) Reasoning with Tools.☆46Dec 17, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆228Nov 27, 2025Updated 7 months ago
- A Framework for Decoupling and Assessing the Capabilities of VLMs☆44Jun 28, 2024Updated 2 years ago
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 10 months ago
- RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.☆2,753Apr 14, 2026Updated 3 months ago
- ☆35Jul 8, 2026Updated 2 weeks ago
- COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play☆30Jul 11, 2026Updated last week
- [NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"☆31Jun 4, 2026Updated last month
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 3 weeks ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning☆401Mar 30, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆17Jul 12, 2025Updated last year
- ☆90Aug 16, 2025Updated 11 months ago
- [EMNLP 2024 Findings] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs☆29May 22, 2025Updated last year
- [ICCV 2025] MMReason, MLLMs, step by step, reasoning benchmark, AGI☆15Apr 25, 2026Updated 2 months ago
- [ICLR 2025] BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval☆206Sep 13, 2025Updated 10 months ago
- The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]☆189Jun 5, 2025Updated last year
- ☆27Jun 2, 2026Updated last month
- [ACL 2025 Main] Official Repository for "Evaluating Language Models as Synthetic Data Generators"☆41Dec 13, 2024Updated last year
- Official implementation of the paper "From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large L…☆55Jun 24, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆44Sep 19, 2024Updated last year
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models☆55Sep 19, 2025Updated 10 months ago
- ☆352May 24, 2025Updated last year
- ☆811Jun 9, 2025Updated last year
- ☆27Aug 31, 2025Updated 10 months ago
- ☆111Jul 6, 2026Updated 2 weeks ago
- ☆28Sep 11, 2025Updated 10 months ago