☆126Aug 1, 2026Updated last week
Alternatives and similar repositories for asta-bench
Users that are interested in asta-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆153Aug 3, 2026Updated last week
- Discovering Data-driven Hypotheses in the Wild☆157Jun 9, 2025Updated last year
- ☆22Jan 29, 2026Updated 6 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆37May 31, 2026Updated 2 months ago
- frozen-in-time version of our Paper Finder agent for reproducing evaluation results☆244Mar 17, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A benchmark to evaluate search-augmented LLMs☆17Aug 28, 2025Updated 11 months ago
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆237Jul 28, 2026Updated last week
- Official code for NeurIPS 2025 paper "AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise"☆198Jul 2, 2026Updated last month
- ☆30Updated this week
- The official repo for the paper "Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness"☆33Jul 21, 2026Updated 3 weeks ago
- AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents☆108May 5, 2026Updated 3 months ago
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆74Jul 21, 2026Updated 2 weeks ago
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification☆11Nov 15, 2023Updated 2 years ago
- Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery☆61Apr 24, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15May 15, 2025Updated last year
- Forecasting high-impact research topics via machine learning on evolving knowledge graphs☆56Jul 19, 2026Updated 3 weeks ago
- [NeurIPS-2024] The offical Implementation of "Instruction-Guided Visual Masking"☆42Nov 15, 2024Updated last year
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 11 months ago
- Scientific articles using or citing Common Crawl data☆30Jul 8, 2026Updated last month
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆21Oct 8, 2025Updated 10 months ago
- Benchmark dataset for the evaluation of scientific article representations on the task of citation recommendation across various scientif…☆12Oct 21, 2022Updated 3 years ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 3 months ago
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆153Jul 18, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆27Jul 23, 2025Updated last year
- Official Implementation of the paper "Jointly Reinforcing Diversity and Quality in Language Model Generations"☆61May 8, 2026Updated 3 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- ☆163May 13, 2026Updated 2 months ago
- ☆41May 26, 2026Updated 2 months ago
- OpenAI Frontier Evals☆1,276Apr 21, 2026Updated 3 months ago
- Repository for the listwise reranker Rank-K☆16May 23, 2025Updated last year
- Repo housing the open sourced code for the ai2 scholar qa app and also the corresponding library☆281Jun 25, 2026Updated last month
- AI Assistance for Writing Scientific Alt Text☆14Feb 7, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- PhD/MBA-level human-annotated rubrics dataset across Physics, Chemistry, Finance and Consulting☆33Oct 30, 2025Updated 9 months ago
- ☆22Jan 2, 2026Updated 7 months ago
- Towards autonomous quantum physics research using LLM agents with access to intelligent tools☆19Jul 26, 2026Updated 2 weeks ago
- Generating Protein Variants with Different Generative Models (HMM, VAE, ESM-2, ProtGPT2)☆11Mar 14, 2024Updated 2 years ago
- ☆18May 15, 2023Updated 3 years ago
- ☆311Jul 1, 2026Updated last month
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆46Jul 6, 2026Updated last month