Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
☆66Apr 24, 2026Updated 4 months ago
Alternatives and similar repositories for AutoResearchBench
Users that are interested in AutoResearchBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Executive Memory for Coherent Long-Horizon Reasoning!☆86Jan 14, 2026Updated 7 months ago
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated last month
- Including 12+ cutting-edge agent systems across multiple research directions☆36Nov 10, 2025Updated 10 months ago
- [ICLR 2026] EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling☆258Mar 20, 2026Updated 5 months ago
- A generalist autonomous research agent — runs experiments, researches, and iteratively optimizes, autonomously.☆1,065Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆20Jan 18, 2026Updated 7 months ago
- [ACL 2025 Oral] 🔥🔥 MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval☆249Nov 6, 2025Updated 10 months ago
- From Prompt Injection to Persistent Control: Defending Agentic Workspaces Against Trojan Backdoors☆19Jun 1, 2026Updated 3 months ago
- ☆24Jul 23, 2025Updated last year
- [ACL 2025] AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark☆167Mar 29, 2026Updated 5 months ago
- ☆13Nov 26, 2021Updated 4 years ago
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- ☆33Apr 29, 2026Updated 4 months ago
- Repo for WWW 2022 paper: Progressively Optimized Bi-Granular Document Representation for Scalable Embedding Based Retrieval☆16Mar 1, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Evaluation code and datasets for the ACL 2024 paper, VISTA: Visualized Text Embedding for Universal Multi-Modal Retrieval. The original c…☆48Nov 16, 2024Updated last year
- SPAR: multi-agent scholarly retrieval with query decomposition, query evolution, and citation-aware exploration.☆28Aug 25, 2026Updated 2 weeks ago
- A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse☆64Jul 10, 2026Updated 2 months ago
- ☆12Oct 28, 2024Updated last year
- ☆135Sep 3, 2026Updated last week
- ☆13Jun 18, 2019Updated 7 years ago
- Baselines for all tasks from Long Code Arena benchmarks 🏟️☆38Mar 30, 2025Updated last year
- ☆74Feb 22, 2023Updated 3 years ago
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆14Mar 18, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)☆355May 28, 2026Updated 3 months ago
- ☆19Aug 28, 2025Updated last year
- An all-in-one framework for Ad-hoc Information Retrieval.☆18Apr 3, 2024Updated 2 years ago
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆86Updated this week
- [EMNLP'26 Findings] OPD-Evolver☆44Jun 17, 2026Updated 2 months ago
- ☆25Jul 24, 2023Updated 3 years ago
- ☆43May 12, 2026Updated 4 months ago
- [ACL'26 Findings] Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets☆20Jun 27, 2026Updated 2 months ago
- ☆18Feb 20, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Benchmarking Language Agents Under Controllable and Extreme Context Growth☆60Apr 29, 2026Updated 4 months ago
- [ICLR 2025] Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning (SASR)☆12Aug 26, 2025Updated last year
- [ACL 2026] CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy☆21Oct 10, 2025Updated 11 months ago
- FeedbackQA: Improving Question Answering Post-Deployment with Interactive Feedback☆12Jul 13, 2022Updated 4 years ago
- ☆16May 15, 2025Updated last year
- ☆11Jun 6, 2023Updated 3 years ago
- DLLM-Searcher has been accepted by SIGIR 2026! 🥳☆34Jan 23, 2026Updated 7 months ago