Live Deep Research Bench. A challenging, objective benchmark for deep research tasks.
☆20Oct 16, 2025Updated 9 months ago
Alternatives and similar repositories for livedrbench
Users that are interested in livedrbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code release for "CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning", ICLR 2025☆34Apr 21, 2025Updated last year
- An agentic evaluation framework☆22Feb 11, 2026Updated 5 months ago
- ☆13Aug 26, 2024Updated last year
- EANN(Pytorch)☆10Mar 12, 2022Updated 4 years ago
- A comprehensive benchmark for evaluating deep research agents on academic survey tasks☆56Sep 4, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆14Sep 1, 2025Updated 11 months ago
- We introduce EfficientRAG, an efficient retriever for multi-hop question answering. EfficientRAG iteratively generates new queries withou…☆17Mar 4, 2025Updated last year
- ☆13Sep 26, 2024Updated last year
- ☆10Mar 4, 2025Updated last year
- Neuropathic Pain Diagnosis Simulator☆14Jul 6, 2023Updated 3 years ago
- StAtutory Reasoning Assessment☆17Dec 8, 2022Updated 3 years ago
- Multilingual RAG benchmark.☆11Nov 22, 2024Updated last year
- 微博谣言检测, 前端Vue,后端Flask☆13Jul 14, 2020Updated 6 years ago
- ☆13Oct 15, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for paper ”Language Versatilists vs. Specialists: An Empirical Revisiting on Multilingual Transfer Ability“☆15Jun 13, 2023Updated 3 years ago
- ☆22Apr 4, 2025Updated last year
- ☆10Jul 10, 2023Updated 3 years ago
- ☆12Oct 17, 2022Updated 3 years ago
- Code for BYOP [CVPR 2023]☆11Sep 25, 2023Updated 2 years ago
- Code for Quantifying Ignorance in Individual-Level Causal-Effect Estimates under Hidden Confounding☆25Dec 6, 2022Updated 3 years ago
- Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark containing 100K image-text pairs for robust image-text …☆16Nov 20, 2025Updated 8 months ago
- The paper list of multilingual pre-trained models (Continual Updated).☆25Jun 18, 2024Updated 2 years ago
- ☆15Nov 10, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A practical bilingual guide to staying safe and prepared at conferences in Brazil / 巴西参会实用攻略与自救指南☆17Apr 22, 2026Updated 3 months ago
- ☆14Mar 2, 2023Updated 3 years ago
- [ACL 2025] Agentic Knowledgeable Self-awareness☆93Jun 15, 2025Updated last year
- EMNLP 2022: "MABEL: Attenuating Gender Bias using Textual Entailment Data" https://arxiv.org/abs/2210.14975☆38Dec 14, 2023Updated 2 years ago
- Get the URL from a web shortcut file☆15Aug 14, 2021Updated 4 years ago
- ☆22Nov 1, 2025Updated 9 months ago
- The official implementation of the paper "Self-Updatable Large Language Models by Integrating Context into Model Parameters"☆15May 18, 2025Updated last year
- Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization☆22Dec 13, 2024Updated last year
- [ICML 24] Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space☆17Aug 9, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆26Nov 8, 2022Updated 3 years ago
- [ACM MM 2024] Pytorch Code for the paper "Robust Variational Contrastive Learning for Partially View-unaligned Clustering"☆17Feb 7, 2026Updated 6 months ago
- MASSW is a comprehensive text dataset on Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 152,000 peer-review…☆22May 16, 2025Updated last year
- ☆22Jun 16, 2025Updated last year
- [WACV 2025] Uniform Attention Maps: Enhancing Image Fidelity in Reconstruction and Editing☆17Mar 16, 2025Updated last year
- ☆21Oct 31, 2024Updated last year
- [ICLR 2026] A framework to "create benchmarks" and "evaluate AI co-scientists" in experimental data-driven real-world scientific research…☆16Feb 16, 2026Updated 5 months ago