Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
☆60Apr 24, 2026Updated 3 months ago
Alternatives and similar repositories for AutoResearchBench
Users that are interested in AutoResearchBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Executive Memory for Coherent Long-Horizon Reasoning!☆86Jan 14, 2026Updated 6 months ago
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated 3 weeks ago
- ☆155Nov 17, 2025Updated 8 months ago
- Including 12+ cutting-edge agent systems across multiple research directions☆35Nov 10, 2025Updated 8 months ago
- Advancing search on top of AI agents☆32Jun 9, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 2022 USTC 011705 (OSH) Course Project of Runikraft Group☆13Jul 22, 2022Updated 4 years ago
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆231Updated this week
- ☆20Jan 18, 2026Updated 6 months ago
- [ACL 2025 Oral] 🔥🔥 MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval☆248Nov 6, 2025Updated 8 months ago
- From Prompt Injection to Persistent Control: Defending Agentic Workspaces Against Trojan Backdoors☆19Jun 1, 2026Updated 2 months ago
- [ACL 2025] AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark☆167Mar 29, 2026Updated 4 months ago
- 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark☆266Apr 13, 2026Updated 3 months ago
- ☆13Nov 26, 2021Updated 4 years ago
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆30Apr 29, 2026Updated 3 months ago
- Repo for WWW 2022 paper: Progressively Optimized Bi-Granular Document Representation for Scalable Embedding Based Retrieval☆16Mar 1, 2022Updated 4 years ago
- ☆27Jul 23, 2025Updated last year
- A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse☆60Jul 10, 2026Updated 3 weeks ago
- ☆13Oct 28, 2024Updated last year
- GISA: A Benchmark for General Information-Seeking Assistant☆37Mar 20, 2026Updated 4 months ago
- ☆126Updated this week
- Baselines for all tasks from Long Code Arena benchmarks 🏟️☆38Mar 30, 2025Updated last year
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)☆325May 28, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation. A typed knowledge graph unifies data synthe…☆22Jul 8, 2026Updated 3 weeks ago
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆75May 14, 2026Updated 2 months ago
- An all-in-one framework for Ad-hoc Information Retrieval.☆18Apr 3, 2024Updated 2 years ago
- ☆42Jun 17, 2026Updated last month
- ☆15May 7, 2022Updated 4 years ago
- ☆25Jul 24, 2023Updated 3 years ago
- IKEA: Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent☆71May 13, 2025Updated last year
- [ACL'26 Findings] Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets☆20Jun 27, 2026Updated last month
- ☆41May 12, 2026Updated 2 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆16Feb 20, 2026Updated 5 months ago
- Benchmarking Language Agents Under Controllable and Extreme Context Growth☆51Apr 29, 2026Updated 3 months ago
- AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents☆105May 5, 2026Updated 2 months ago
- [ACL 2026] CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy☆21Oct 10, 2025Updated 9 months ago
- About Code release for "FlashBias: Fast Computation of Attention with Bias" (NeurIPS 2025), https://arxiv.org/abs/2505.12044☆33Nov 17, 2025Updated 8 months ago
- FeedbackQA: Improving Question Answering Post-Deployment with Interactive Feedback☆12Jul 13, 2022Updated 4 years ago
- ☆15May 15, 2025Updated last year