Discovering Data-driven Hypotheses in the Wild
☆157Jun 9, 2025Updated last year
Alternatives and similar repositories for discoverybench
Users that are interested in discoverybench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆154Jul 18, 2026Updated 3 weeks ago
- ☆130Aug 1, 2026Updated 2 weeks ago
- A virtual environment for developing and evaluating automated scientific discovery agents.☆218Mar 10, 2025Updated last year
- ☆154Aug 3, 2026Updated last week
- Official code for NeurIPS 2025 paper "AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise"☆201Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2024] <Large Language Models for Automated Open-domain Scientific Hypotheses Discovery>. It has also received the best poster award …☆47Oct 28, 2024Updated last year
- ☆58Apr 4, 2025Updated last year
- ☆45Aug 20, 2025Updated 11 months ago
- [EMNLP 2024 Findings] Benchmarking Language Model Agents for Data-Driven Science☆35Oct 25, 2024Updated last year
- Evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology☆126Sep 27, 2025Updated 10 months ago
- Repo housing the open sourced code for the ai2 scholar qa app and also the corresponding library☆280Jun 25, 2026Updated last month
- BioDiscoveryAgent is an LLM-based AI agent for closed-loop design of genetic perturbation experiments☆119Jul 6, 2025Updated last year
- [EMNLP'25] AutoSDT is a fully automatic pipeline to collect data-driven scientific coding tasks to train co-scientist models.☆21Aug 11, 2025Updated last year
- ScienceWorld is a text-based virtual environment centered around accomplishing tasks from the standardized elementary science curriculum.☆378Aug 4, 2026Updated last week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆33Updated this week
- Dataset and annotations for ASSETS 2022 publication☆13Oct 6, 2022Updated 3 years ago
- Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data☆48Feb 18, 2025Updated last year
- Official code for TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents☆46Nov 30, 2025Updated 8 months ago
- Automated Hypothesis Testing with Agentic Sequential Falsifications☆284May 14, 2025Updated last year
- EmbedGEM: A framework to evaluate the utility of embeddings for genetic discovery☆23Oct 3, 2024Updated last year
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models☆31Jul 13, 2025Updated last year
- Repository containing dataset, models and code associated with the CHIME project☆18Aug 22, 2024Updated last year
- ☆10Nov 6, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Jun 1, 2024Updated 2 years ago
- ☆11Feb 11, 2020Updated 6 years ago
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification☆11Nov 15, 2023Updated 2 years ago
- ☆10Oct 2, 2024Updated last year
- [NeurIPS 2024] OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI☆106Mar 6, 2025Updated last year
- Benchmarking LLMs with Challenging Tasks from Real Users☆256Nov 3, 2024Updated last year
- ☆12Jan 7, 2025Updated last year
- A benchmark that challenges language models to code solutions for scientific problems☆217Updated this week
- Dataset and source code for LearningQ: A Large-scale Dataset for Educational Question Generation (ICWSM 2018).☆65Jun 16, 2020Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- CauSight: Learning to Supersense for Visual Causal Discovery☆21Mar 11, 2026Updated 5 months ago
- S2ORC: The Semantic Scholar Open Research Corpus: https://www.aclweb.org/anthology/2020.acl-main.447/☆1,080Apr 26, 2024Updated 2 years ago
- Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike stat…☆543Updated this week
- AIRA-dojo: a framework for developing and evaluating AI research agents☆157Apr 14, 2026Updated 4 months ago
- Code/data for MARG (multi-agent review generation)☆64Mar 5, 2026Updated 5 months ago
- Ready-to-use agent skills that clone the thinking of founders, philosophers, and scientists into your agent. Generated with K-Dense-AI/mi…☆101Updated this week
- Code release for "CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning", ICLR 2025☆34Apr 21, 2025Updated last year