Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
☆64Apr 24, 2026Updated 4 months ago
Alternatives and similar repositories for AutoResearchBench
Users that are interested in AutoResearchBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Executive Memory for Coherent Long-Horizon Reasoning!☆86Jan 14, 2026Updated 7 months ago
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated last month
- ☆154Nov 17, 2025Updated 9 months ago
- Including 12+ cutting-edge agent systems across multiple research directions☆36Nov 10, 2025Updated 9 months ago
- [ICLR 2026] EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling☆257Mar 20, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Advancing search on top of AI agents☆33Jun 9, 2026Updated 2 months ago
- A general memory system for agents, powered by deep-research☆859Mar 14, 2026Updated 5 months ago
- A generalist autonomous research agent — runs experiments, researches, and iteratively optimizes, autonomously.☆1,030Updated this week
- 2022 USTC 011705 (OSH) Course Project of Runikraft Group☆13Jul 22, 2022Updated 4 years ago
- ☆20Jan 18, 2026Updated 7 months ago
- From Prompt Injection to Persistent Control: Defending Agentic Workspaces Against Trojan Backdoors☆19Jun 1, 2026Updated 2 months ago
- ☆24Jul 23, 2025Updated last year
- [ACL 2025] AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark☆167Mar 29, 2026Updated 4 months ago
- 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark☆268Apr 13, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Talk to research papers like talking to authors - Python package with AI agent for arXiv papers☆771Aug 12, 2026Updated last week
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- ☆31Apr 29, 2026Updated 3 months ago
- SPAR: multi-agent scholarly retrieval with query decomposition, query evolution, and citation-aware exploration.☆27Updated this week
- A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse☆62Jul 10, 2026Updated last month
- ☆13Oct 28, 2024Updated last year
- Bulk download PDFs and TEI XML files from OpenAlex☆29Mar 18, 2026Updated 5 months ago
- GISA: A Benchmark for General Information-Seeking Assistant☆37Mar 20, 2026Updated 5 months ago
- ☆132Aug 1, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Jun 18, 2019Updated 7 years ago
- Baselines for all tasks from Long Code Arena benchmarks 🏟️☆38Mar 30, 2025Updated last year
- ☆74Feb 22, 2023Updated 3 years ago
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆15Mar 18, 2026Updated 5 months ago
- Prolog Implementation in Python☆12Dec 28, 2017Updated 8 years ago
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)☆340May 28, 2026Updated 2 months ago
- Evaluate the performance of the oversampling method KMeans-SMOTE☆17Mar 30, 2019Updated 7 years ago
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆83Aug 14, 2026Updated last week
- An all-in-one framework for Ad-hoc Information Retrieval.☆18Apr 3, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆42Jun 17, 2026Updated 2 months ago
- ☆25Jul 24, 2023Updated 3 years ago
- IKEA: Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent☆71May 13, 2025Updated last year
- [ACL'26 Findings] Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets☆20Jun 27, 2026Updated last month
- ☆41May 12, 2026Updated 3 months ago
- ☆17Feb 20, 2026Updated 6 months ago
- Benchmarking Language Agents Under Controllable and Extreme Context Growth☆54Apr 29, 2026Updated 3 months ago