Code repository for ICLR 2026 paper "ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents" (https://www.arxiv.org/abs/2511.07685)
☆29Feb 10, 2026Updated 6 months ago
Alternatives and similar repositories for researchrubrics
Users that are interested in researchrubrics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2026] Skill-Targeted Adaptive Training☆26Mar 12, 2026Updated 5 months ago
- Open source codebase for PRBench☆20Jan 15, 2026Updated 7 months ago
- code for EACL2024-main:Generative Dense Retrieval: Memory Can Be a Burden☆32Jan 19, 2024Updated 2 years ago
- [ACL 2026] DR-Arena: an Automated Evaluation Framework for Deep Research Agents☆18Jul 8, 2026Updated last month
- ☆32Mar 17, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [EMNLP 2023] Question Answering as Programming for Solving Time-Sensitive Questions☆12Dec 18, 2023Updated 2 years ago
- Official implementation for paper "How Far Are We from Genuinely Useful Deep Research Agents?"☆66Dec 10, 2025Updated 8 months ago
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆83Aug 14, 2026Updated 2 weeks ago
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry☆53Jul 21, 2026Updated last month
- Code for EACL 26 Findings paper "I-MCTS: Enhancing Agentic AutoML via Introspective Monte Carlo Tree Search"☆13Jan 28, 2026Updated 7 months ago
- [NeurIPS 2025] Official PyTorch implementation of paper "Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression".☆16Oct 24, 2025Updated 10 months ago
- [SIGIR 2025] Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph☆17Jun 6, 2025Updated last year
- Code for the paper - Controlling Dialogue Generation with Semantic Exemplars (Naacl 2021) A semantic exemplar based retrieve-refine appro…☆18Mar 26, 2021Updated 5 years ago
- INSCIT: Information-Seeking Conversations with Mixed-Initiative Interactions☆16Jan 21, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A lightweight kmer-based algorithm for designing diagnostic CRISPR assays using genome data.☆13Aug 19, 2024Updated 2 years ago
- R3: Robust Rubric-Agnostic Reward Models☆23Jul 12, 2025Updated last year
- mechanical stage☆12Aug 6, 2023Updated 3 years ago
- [WWW 2026 Oral] MoE-CL:Self-Evolving LLMs via Continual Instruction Tuning☆21Dec 1, 2025Updated 8 months ago
- Metrics for evaluating biological sequence design☆16Jul 22, 2026Updated last month
- Verifying the optimization phases of the GraalVM compiler☆15Jul 30, 2026Updated last month
- ☆10Mar 11, 2024Updated 2 years ago
- ☆22Oct 12, 2024Updated last year
- ☆16Dec 10, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- KuaiSearch PERKS☆12Nov 16, 2021Updated 4 years ago
- ☆12Oct 7, 2020Updated 5 years ago
- ☆19Jan 29, 2026Updated 7 months ago
- ☆100Jun 27, 2024Updated 2 years ago
- ☆71Jun 24, 2025Updated last year
- Sutracli is an AI-powered code manager for coding agents. It spawns agents for multiple projects, connects repos through cross-indexing, …☆29Nov 7, 2025Updated 9 months ago
- Code and data release for FEABench: Evaluating Language Models on Multiphysics Reasoning Ability. [MATH-AI workshop, NeurIPS 2024]☆14May 7, 2025Updated last year
- Templates and examples for ACL and EMNLP conference posters.☆15Oct 5, 2024Updated last year
- Step-DeepResearch☆571Mar 24, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Open-source Large Language Models are Strong Zero-shot Query Likelihood Models for Document Ranking☆17Oct 26, 2023Updated 2 years ago
- Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research☆700Jun 17, 2026Updated 2 months ago
- The omegaUp sandbox☆15Updated this week
- ☆10Feb 21, 2024Updated 2 years ago
- This is the public repository of AAAI 2024 paper "Is a Large Language Model a Good Annotator for Event Extraction"☆10Feb 16, 2024Updated 2 years ago
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- ☆16Jul 31, 2025Updated last year