Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
☆106Jul 23, 2026Updated 3 weeks ago
Alternatives and similar repositories for SSP
Users that are interested in SSP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Feb 14, 2026Updated 6 months ago
- ☆29Jun 9, 2026Updated 2 months ago
- ☆22Jun 16, 2026Updated 2 months ago
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR☆21Apr 7, 2026Updated 4 months ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆25Dec 30, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🪐 Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback☆23Jan 29, 2026Updated 6 months ago
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆18Oct 20, 2025Updated 9 months ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆24Jun 26, 2026Updated last month
- [ICLR 2026] Tree Search for LLM Agent Reinforcement Learning☆395Jan 26, 2026Updated 6 months ago
- Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework☆340Jan 17, 2026Updated 7 months ago
- ☆16May 18, 2026Updated 3 months ago
- [ICLR26]GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning☆181Jan 29, 2026Updated 6 months ago
- [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)☆1,108Jul 13, 2026Updated last month
- HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurren…☆73May 23, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Domain-specific preference (DSP) data and customized RM fine-tuning.☆26Mar 7, 2024Updated 2 years ago
- [ICLR 2026] Geometric-Mean Policy Optimization☆104Jan 26, 2026Updated 6 months ago
- {DeepL, Google, WMT-Best, davinci-003, turbo, gpt-4} × {En-De, En-Cs, En-Ru, En-Zh, De-Fr, En-Ja, Uk-En, Uk-Cs, En-Hr, En-Ha, En-Is}☆14Jun 18, 2023Updated 3 years ago
- An Open-Source Large-Scale Reinforcement Learning Project for Search Agents☆608Nov 26, 2025Updated 8 months ago
- [IROS 2023] Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition☆21Jul 12, 2025Updated last year
- ☆27Jun 10, 2025Updated last year
- ☆44Oct 28, 2025Updated 9 months ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆16Jul 19, 2026Updated 3 weeks ago
- The official repository of "A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Eva…☆285Jul 30, 2026Updated 2 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Codes for our paper "AgentMonitor: A Plug-and-Play Framework for Predictive and Secure Multi-Agent Systems"☆14Dec 13, 2024Updated last year
- Official codebase of the paper -- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills☆243May 1, 2026Updated 3 months ago
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆91Dec 8, 2025Updated 8 months ago
- [ECCV 2026] Promsa: Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering☆91Jul 7, 2026Updated last month
- ☆46Jan 19, 2026Updated 6 months ago
- Official repository for ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use☆31Nov 4, 2025Updated 9 months ago
- OmniGAIA: Towards Native Omni-Modal AI Agents☆145Apr 2, 2026Updated 4 months ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆15Jun 28, 2025Updated last year
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,302Nov 13, 2025Updated 9 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [ICLR 2026] FASA: FREQUENCY-AWARE SPARSE ATTENTION☆21Mar 1, 2026Updated 5 months ago
- [EMNLP25] Official code for "POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation…☆38Nov 11, 2025Updated 9 months ago
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- [ACL-2026] MMSearch-R1 is an end-to-end RL framework that enables LMMs to perform on-demand, multi-turn search with real-world multimodal…☆478Apr 7, 2026Updated 4 months ago
- Official implementation of MATPO: Multi-Agent Tool-Integrated Policy Optimization.☆83Oct 31, 2025Updated 9 months ago
- An awesome & curated list of anything that might be useful for computer science students☆13Mar 27, 2023Updated 3 years ago
- ☆18Apr 18, 2025Updated last year