Discovering Data-driven Hypotheses in the Wild
☆160Jun 9, 2025Updated last year
Alternatives and similar repositories for discoverybench
Users that are interested in discoverybench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆168Jul 18, 2026Updated last month
- ☆134Aug 1, 2026Updated last month
- ☆22Jan 29, 2026Updated 7 months ago
- A virtual environment for developing and evaluating automated scientific discovery agents.☆219Mar 10, 2025Updated last year
- ☆155Aug 3, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official code for NeurIPS 2025 paper "AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise"☆205Aug 13, 2026Updated 3 weeks ago
- [ACL 2024] <Large Language Models for Automated Open-domain Scientific Hypotheses Discovery>. It has also received the best poster award …☆47Oct 28, 2024Updated last year
- ☆58Apr 4, 2025Updated last year
- ☆48Aug 20, 2025Updated last year
- [EMNLP 2024 Findings] Benchmarking Language Model Agents for Data-Driven Science☆37Oct 25, 2024Updated last year
- Repo housing the open sourced code for the ai2 scholar qa app and also the corresponding library☆280Jun 25, 2026Updated 2 months ago
- BioDiscoveryAgent is an LLM-based AI agent for closed-loop design of genetic perturbation experiments☆125Jul 6, 2025Updated last year
- [EMNLP'25] AutoSDT is a fully automatic pipeline to collect data-driven scientific coding tasks to train co-scientist models.☆21Aug 11, 2025Updated last year
- A curated list of papers on LLMs and agents for scientific research and development☆99Dec 11, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ScienceWorld is a text-based virtual environment centered around accomplishing tasks from the standardized elementary science curriculum.☆385Aug 20, 2026Updated 2 weeks ago
- ☆38Updated this week
- Dataset and annotations for ASSETS 2022 publication☆13Oct 6, 2022Updated 3 years ago
- Official code for TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents☆47Nov 30, 2025Updated 9 months ago
- Automated Hypothesis Testing with Agentic Sequential Falsifications☆289May 14, 2025Updated last year
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models☆31Jul 13, 2025Updated last year
- Repository containing dataset, models and code associated with the CHIME project☆18Aug 22, 2024Updated 2 years ago
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification☆11Nov 15, 2023Updated 2 years ago
- Reproducible and flexible LLM evaluations for scientific reasoning.☆31Jul 23, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- This repository explores a variety of data visualization techniques, with a particular focus on applications in the hospitality domain. I…☆41Oct 16, 2025Updated 10 months ago
- Benchmarking LLMs with Challenging Tasks from Real Users☆256Nov 3, 2024Updated last year
- ☆12Jan 7, 2025Updated last year
- A benchmark that challenges language models to code solutions for scientific problems☆224Updated this week
- CauSight: Learning to Supersense for Visual Causal Discovery☆21Mar 11, 2026Updated 5 months ago
- S2ORC: The Semantic Scholar Open Research Corpus: https://www.aclweb.org/anthology/2020.acl-main.447/☆1,083Apr 26, 2024Updated 2 years ago
- Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike stat…☆550Aug 26, 2026Updated last week
- If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions☆17Apr 4, 2024Updated 2 years ago
- AIRA-dojo: a framework for developing and evaluating AI research agents☆164Apr 14, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code/data for MARG (multi-agent review generation)☆64Mar 5, 2026Updated 5 months ago
- Ready-to-use agent skills that clone the thinking of founders, philosophers, and scientists into your agent. Generated with K-Dense-AI/mi…☆120Aug 18, 2026Updated 2 weeks ago
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆145May 6, 2026Updated 3 months ago
- [ACL 2024, Main Conference] CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Fol…☆15Aug 7, 2024Updated 2 years ago
- Headway - Selenium Maven TestNG POM Data Driven Framework☆18Jul 2, 2025Updated last year
- Official implementation of Panacea: A foundation model for clinical trial design, recruitment, search, and summarization.☆21Dec 24, 2024Updated last year
- Code and data used to create and evaluate LLM4Mat-Bench☆34Jan 31, 2026Updated 7 months ago