Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
☆106Jul 23, 2026Updated last month
Alternatives and similar repositories for SSP
Users that are interested in SSP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Feb 14, 2026Updated 6 months ago
- ☆30Jun 9, 2026Updated 2 months ago
- ☆23Jun 16, 2026Updated 2 months ago
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR☆22Apr 7, 2026Updated 5 months ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 🪐 Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback☆25Jan 29, 2026Updated 7 months ago
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆18Oct 20, 2025Updated 10 months ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆25Jun 26, 2026Updated 2 months ago
- [ICLR 2026] Tree Search for LLM Agent Reinforcement Learning☆402Jan 26, 2026Updated 7 months ago
- Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework☆350Jan 17, 2026Updated 7 months ago
- ☆16May 18, 2026Updated 3 months ago
- [ICLR26]GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning☆181Jan 29, 2026Updated 7 months ago
- [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)☆1,114Aug 20, 2026Updated 2 weeks ago
- HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurren…☆76May 23, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Domain-specific preference (DSP) data and customized RM fine-tuning.☆26Mar 7, 2024Updated 2 years ago
- [ICLR 2026] Geometric-Mean Policy Optimization☆105Jan 26, 2026Updated 7 months ago
- {DeepL, Google, WMT-Best, davinci-003, turbo, gpt-4} × {En-De, En-Cs, En-Ru, En-Zh, De-Fr, En-Ja, Uk-En, Uk-Cs, En-Hr, En-Ha, En-Is}☆14Jun 18, 2023Updated 3 years ago
- An Open-Source Large-Scale Reinforcement Learning Project for Search Agents☆610Nov 26, 2025Updated 9 months ago
- [IROS 2023] Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition☆21Jul 12, 2025Updated last year
- ☆27Jun 10, 2025Updated last year
- ☆44Oct 28, 2025Updated 10 months ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆15Jul 19, 2026Updated last month
- The official repository of "A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Eva…☆292Jul 30, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Codes for our paper "AgentMonitor: A Plug-and-Play Framework for Predictive and Secure Multi-Agent Systems"☆14Dec 13, 2024Updated last year
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding☆24Jul 6, 2026Updated 2 months ago
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆91Dec 8, 2025Updated 8 months ago
- Official codebase of the paper -- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills☆275May 1, 2026Updated 4 months ago
- ☆47Jan 19, 2026Updated 7 months ago
- [ECCV 2026] Promsa: Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering☆93Jul 7, 2026Updated 2 months ago
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use☆31Nov 4, 2025Updated 10 months ago
- OmniGAIA: Towards Native Omni-Modal AI Agents☆145Apr 2, 2026Updated 5 months ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,375Nov 13, 2025Updated 9 months ago
- [ICLR 2026] FASA: FREQUENCY-AWARE SPARSE ATTENTION☆22Aug 18, 2026Updated 2 weeks ago
- [EMNLP25] Official code for "POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation…☆38Nov 11, 2025Updated 9 months ago
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- [ACL-2026] MMSearch-R1 is an end-to-end RL framework that enables LMMs to perform on-demand, multi-turn search with real-world multimodal…☆485Apr 7, 2026Updated 5 months ago
- Official implementation of MATPO: Multi-Agent Tool-Integrated Policy Optimization.☆83Oct 31, 2025Updated 10 months ago
- An awesome & curated list of anything that might be useful for computer science students☆13Mar 27, 2023Updated 3 years ago