Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
☆628Sep 23, 2026Updated this week
Alternatives and similar repositories for terminal-bench-science
Users that are interested in terminal-bench-science are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Measuring and evolving with the frontier of agent work☆766Updated this week
- A curated list of awesome Harbor ecosystem projects☆53May 29, 2026Updated 3 months ago
- SWE-Marathon: an ultra long-horizon SWE benchmark☆164Updated this week
- ☆415Apr 30, 2026Updated 4 months ago
- ☆52Apr 1, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- open source SWE-Atlas☆71Aug 20, 2026Updated last month
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆147Feb 16, 2026Updated 7 months ago
- Agentic layer for ASTRA — dedicated skills, workflow execution, and HPC/container management for reproducible research☆24Sep 11, 2026Updated 2 weeks ago
- Framework for evaluating and improving agents☆5,571Updated this week
- A benchmark for evaluating AI agents on realistic business workflows☆305Aug 4, 2026Updated last month
- ☆123Apr 1, 2026Updated 5 months ago
- Data recipes and robust infrastructure for training AI agents☆295Updated this week
- Official eval scripts for JobBench☆54Sep 14, 2026Updated last week
- A scientific reasoning model, dataset, and reward functions for chemistry.☆168Oct 26, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆46Dec 8, 2024Updated last year
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆786Updated this week
- The IRI Facility API reference implementation (Python)☆25Updated this week
- ☆18Jul 25, 2025Updated last year
- Polygraph evaluates and compares groups of nucleic acid sequences based on their sequence and functional content for effective design of …☆40Mar 27, 2025Updated last year
- Source code for the paper 'Uncovering Neural Scaling Laws in Molecular Representation Learning' (NeurIPS 2023 Datasets and Benchmarks).☆14Dec 2, 2023Updated 2 years ago
- OpenTelemetry Benchmark - can AI trace your failed login?☆23Jul 14, 2026Updated 2 months ago
- Simple scVI implementation☆13Aug 31, 2026Updated 3 weeks ago
- This repository contains information on the creation, evaluation, and benchmark models for the L+M-24 Dataset. L+M-24 will be featured as…☆30Jan 23, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆148Mar 31, 2026Updated 5 months ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆45Jul 6, 2026Updated 2 months ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆567Updated this week
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,813Jul 23, 2026Updated 2 months ago
- Annotated sequence data☆11Feb 2, 2025Updated last year
- MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion (ACL 2025)☆37Jul 16, 2025Updated last year
- Automated Complex Generator☆16Dec 16, 2024Updated last year
- ☆14Oct 15, 2024Updated last year
- Reference implementation of "Ewald-based Long-Range Message Passing for Molecular Graphs" (ICML 2023)☆52May 4, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,346Updated this week
- This is a repository with examples to run inference endpoints on various ALCF clusters☆29Sep 15, 2026Updated last week
- This is the official implementation of the paper: Fractional Denoising for 3D Molecular Pre-training☆22Mar 30, 2024Updated 2 years ago
- E(n) Equivariant GNN in jax☆14Aug 31, 2023Updated 3 years ago
- Continual Learning Bench☆226Jul 19, 2026Updated 2 months ago
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆173Jul 18, 2026Updated 2 months ago
- Repository for "NucleoBench: A Large-Scale Benchmark of Neural Nucleic Acid Design Algorithms"☆27Sep 14, 2026Updated last week