Discovering Data-driven Hypotheses in the Wild
☆161Jun 9, 2025Updated last year
Alternatives and similar repositories for discoverybench
Users that are interested in discoverybench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆173Jul 18, 2026Updated 2 months ago
- ☆138Sep 3, 2026Updated 3 weeks ago
- ☆21Jan 29, 2026Updated 7 months ago
- A virtual environment for developing and evaluating automated scientific discovery agents.☆222Mar 10, 2025Updated last year
- ☆154Aug 3, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official code for NeurIPS 2025 paper "AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise"☆209Aug 13, 2026Updated last month
- [ACL 2024] <Large Language Models for Automated Open-domain Scientific Hypotheses Discovery>. It has also received the best poster award …☆47Oct 28, 2024Updated last year
- ☆58Apr 4, 2025Updated last year
- ☆51Aug 20, 2025Updated last year
- [EMNLP 2024 Findings] Benchmarking Language Model Agents for Data-Driven Science☆38Oct 25, 2024Updated last year
- [EMNLP'25] AutoSDT is a fully automatic pipeline to collect data-driven scientific coding tasks to train co-scientist models.☆22Aug 11, 2025Updated last year
- A curated list of papers on LLMs and agents for scientific research and development☆99Dec 11, 2024Updated last year
- ScienceWorld is a text-based virtual environment centered around accomplishing tasks from the standardized elementary science curriculum.☆394Aug 20, 2026Updated last month
- ☆39Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data☆49Feb 18, 2025Updated last year
- Official code for TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents☆49Nov 30, 2025Updated 9 months ago
- ☆10Nov 6, 2024Updated last year
- ☆10Jun 1, 2024Updated 2 years ago
- ☆11Feb 11, 2020Updated 6 years ago
- ☆10Oct 2, 2024Updated last year
- [NeurIPS 2024] OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI☆106Mar 6, 2025Updated last year
- This repository explores a variety of data visualization techniques, with a particular focus on applications in the hospitality domain. I…☆41Oct 16, 2025Updated 11 months ago
- Benchmarking LLMs with Challenging Tasks from Real Users☆257Nov 3, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A benchmark that challenges language models to code solutions for scientific problems☆231Updated this week
- Top Life Sciences open-source software☆24Jun 9, 2024Updated 2 years ago
- Dataset and source code for LearningQ: A Large-scale Dataset for Educational Question Generation (ICWSM 2018).☆65Jun 16, 2020Updated 6 years ago
- Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike stat…☆557Aug 26, 2026Updated 3 weeks ago
- Deep Data Research. Seek More, See Beyond.☆19Feb 6, 2026Updated 7 months ago
- AIRA-dojo: a framework for developing and evaluating AI research agents☆172Apr 14, 2026Updated 5 months ago
- Code/data for MARG (multi-agent review generation)☆64Mar 5, 2026Updated 6 months ago
- ☆19Apr 14, 2026Updated 5 months ago
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆146May 6, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Headway - Selenium Maven TestNG POM Data Driven Framework☆18Jul 2, 2025Updated last year
- Official implementation of Panacea: A foundation model for clinical trial design, recruitment, search, and summarization.☆21Dec 24, 2024Updated last year
- Code and data used to create and evaluate LLM4Mat-Bench☆35Jan 31, 2026Updated 7 months ago
- Generate Python docstrings automatically with LLM and syntax trees☆20Jun 13, 2025Updated last year
- Benchmark agents on BioML tasks☆77Sep 14, 2025Updated last year
- Interpretable Neural Predictions with Differentiable Binary Variables☆85May 7, 2021Updated 5 years ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,752Apr 24, 2026Updated 5 months ago