Agent-RRM: Exploring Reasoning Reward Model for Agents
☆70Mar 17, 2026Updated 4 months ago
Alternatives and similar repositories for Reagent
Users that are interested in Reagent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Unlocking Iterative Reasoning for Any Image Editor☆111Jan 18, 2026Updated 6 months ago
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆72Apr 8, 2026Updated 3 months ago
- The official repo of "MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data"☆66Mar 27, 2026Updated 3 months ago
- ☆55Dec 10, 2025Updated 7 months ago
- Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs☆45Jun 17, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Source Code for our ICLR'26 paper☆17Feb 22, 2026Updated 5 months ago
- 🐧 Unify-Agent: An end-to-end unified multimodal agent for faithful, knowledge-grounded image generation.☆86May 2, 2026Updated 2 months ago
- 🔥 OneThinker: All-in-one Reasoning Model for Image and Video [CVPR 2026]☆463Feb 28, 2026Updated 4 months ago
- [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)☆1,090Jul 13, 2026Updated last week
- ☆18Mar 16, 2026Updated 4 months ago
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 3 months ago
- [ICML 2026 Spotlight] Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback☆70Jun 3, 2026Updated last month
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use☆31Nov 4, 2025Updated 8 months ago
- 🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diver…☆254May 19, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆86May 2, 2026Updated 2 months ago
- ☆155Nov 17, 2025Updated 8 months ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 4 months ago
- Gen-Searcher: Reinforcing Agentic Search for Image Generation☆376Apr 7, 2026Updated 3 months ago
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward☆94Aug 8, 2025Updated 11 months ago
- The official implementation of "EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis".☆176Feb 12, 2026Updated 5 months ago
- GISA: A Benchmark for General Information-Seeking Assistant☆36Mar 20, 2026Updated 4 months ago
- ☆45Jan 19, 2026Updated 6 months ago
- [ICML 2026] Multimodal deep-research MLLM and benchmark. The first long-horizon multimodal deep-research MLLM, extending the number of re…☆657Jun 8, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆15Jul 1, 2026Updated 3 weeks ago
- [NeurIPS 2025] Implementation for the paper "The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning"☆165Mar 2, 2026Updated 4 months ago
- HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches☆40Oct 9, 2025Updated 9 months ago
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images☆52Apr 28, 2026Updated 2 months ago
- ☆33Jun 30, 2026Updated 3 weeks ago
- ☆68Aug 14, 2025Updated 11 months ago
- [NAACL 2024] CoE-SQL: In-Context Learning for Multi-Turn Text-to-SQL with Chain-of-Editions☆13May 7, 2024Updated 2 years ago
- The implementation for SIGIR 2026: Learning to Retrieve from Agent Trajectories.☆55Jul 14, 2026Updated last week
- Including 12+ cutting-edge agent systems across multiple research directions☆35Nov 10, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 🚀 Text2Grad: Converting natural language feedback into gradient signals for precise model optimization. Revolutionizing RLHF with span-l…☆37Feb 6, 2026Updated 5 months ago
- SimKO: Simple Pass@K Policy Optimization☆31Oct 24, 2025Updated 8 months ago
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated 11 months ago
- ☆18Feb 2, 2026Updated 5 months ago
- Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"☆354Updated this week
- MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning☆68Jun 14, 2026Updated last month
- [ICLR 2026] Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents☆128Jul 14, 2026Updated last week