DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research systems and human experts. It does so by decomposing expert-written reports into hierarchical rubrics covering presentation, analysis, and evidence, and using them to evaluate model-generated.
☆94Sep 11, 2026Updated 3 weeks ago
Alternatives and similar repositories for DeepResearch-Bench-II
Users that are interested in DeepResearch-Bench-II are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents☆837Updated this week
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆23Oct 8, 2025Updated 11 months ago
- A deep research framework☆37Apr 21, 2026Updated 5 months ago
- [ACL 2026] DR-Arena: an Automated Evaluation Framework for Deep Research Agents☆18Jul 8, 2026Updated 2 months ago
- you.com's framework for evaluating deep research systems.☆77May 15, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- WideSearch: Benchmarking Agentic Broad Info-Seeking☆154Oct 9, 2025Updated 11 months ago
- Code repository for ICLR 2026 paper "ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents" (https://ww…☆31Feb 10, 2026Updated 7 months ago
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆51Jul 6, 2026Updated 2 months ago
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆123Sep 29, 2025Updated last year
- [SIGKDD 2024] Rethinking Fair Graph Neural Networks from Re-balancing☆10Jul 15, 2024Updated 2 years ago
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval And Synthesis For SLMs☆64Oct 7, 2025Updated 11 months ago
- The Source Code for DR3-Eval☆40Aug 12, 2026Updated last month
- Official eval scripts for JobBench☆56Sep 14, 2026Updated 2 weeks ago
- Baidu Qianfan Deep Research☆39Jun 8, 2026Updated 3 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆93Dec 8, 2025Updated 9 months ago
- Step-DeepResearch☆573Mar 24, 2026Updated 6 months ago
- GISA: A Benchmark for General Information-Seeking Assistant☆37Updated this week
- ☆35May 16, 2025Updated last year
- ThinkDepth.ai Deep Research☆189Jan 5, 2026Updated 8 months ago
- ☆13Jan 17, 2017Updated 9 years ago
- Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery☆72Apr 24, 2026Updated 5 months ago
- UniScientist is designed to advance universal scientific research intelligence through a unified paradigm☆172Mar 14, 2026Updated 6 months ago
- ☆23Jun 18, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆164May 14, 2025Updated last year
- Official Code: TheWebConf 2022 Compact Graph Structure Learning via Mutual Information Compression☆25Mar 17, 2024Updated 2 years ago
- "DeepResearch-Eval: An End-to-End Evaluation Framework for DeepResearch Systems"☆50Oct 16, 2025Updated 11 months ago
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry☆54Jul 21, 2026Updated 2 months ago
- [NeurIPS 2026] $OneMillion-Bench: How Far are Language Agents from Human Experts?☆53Updated this week
- LLM Safeguarding with Internal Representations☆21Apr 27, 2026Updated 5 months ago
- Fast multi-modal decision model☆450Updated this week
- Latent-variable Synchronous Context-Free Grammar Toolkit☆10Sep 30, 2014Updated 12 years ago
- [NeurIPS 2024] Fast Best-of-N Decoding via Speculative Rejection☆56Oct 29, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆21Oct 29, 2025Updated 11 months ago
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL☆349Jun 17, 2026Updated 3 months ago
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆75Apr 8, 2026Updated 5 months ago
- ☆66Jan 31, 2026Updated 8 months ago
- Kimi K2 Thinking Agentic Search Unofficial Implementation☆15Nov 9, 2025Updated 10 months ago
- [CVPRW & CLIC 2022 - Perceptual Metrics] Image Quality Assessment with Transformers and Multi-Metric Fusion Modules☆13Apr 27, 2023Updated 3 years ago
- Orpheus TTS Server with streaming support (TTFB ~160ms)☆26Sep 21, 2025Updated last year