☆23Jun 18, 2026Updated 2 months ago
Alternatives and similar repositories for terminal-bench-challenges
Users that are interested in terminal-bench-challenges are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A compact high-signal benchmark for evaluating frontier agents☆37Aug 3, 2026Updated last month
- Convert GitHub PRs into Harbor tasks☆83Jul 13, 2026Updated last month
- Measuring and evolving with the frontier of agent work☆673Updated this week
- A curated list of awesome Harbor ecosystem projects☆52May 29, 2026Updated 3 months ago
- Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains☆571Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Benchmark harness and code for "SWE-fficiency: Can Language Models Optimize Real World Repositories on Real World Workloads?"☆25Feb 17, 2026Updated 6 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 6 months ago
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆66Sep 3, 2026Updated last week
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 7 months ago
- OpenTelemetry Benchmark - can AI trace your failed login?☆21Jul 14, 2026Updated last month
- ☆122Apr 1, 2026Updated 5 months ago
- open source SWE-Atlas☆71Aug 20, 2026Updated 3 weeks ago
- Download Web-10K data by querying Bing Image Search☆10Feb 1, 2022Updated 4 years ago
- Python package for extractive NLP using the OpenAI API☆17Aug 28, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆41May 16, 2026Updated 3 months ago
- ☆10Dec 17, 2020Updated 5 years ago
- Official implementation of "BERTs are Generative In-Context Learners"☆32Mar 14, 2025Updated last year
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆74Jul 28, 2025Updated last year
- Realistic examples of building evals and optimizing agents with Harbor☆206Apr 23, 2026Updated 4 months ago
- ☆11Apr 24, 2023Updated 3 years ago
- ☆15Dec 12, 2024Updated last year
- SWE-Marathon: an ultra long-horizon SWE benchmark☆152Updated this week
- Samaya AI's FrontierFinance Benchmark Grader☆19Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆24Mar 18, 2025Updated last year
- Benchmarking Open-Ended Inference Optimization by AI Agents☆44Jul 6, 2026Updated 2 months ago
- ☆407Apr 30, 2026Updated 4 months ago
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the p…☆14Jan 26, 2025Updated last year
- JMLR Cover Letter Template☆10Dec 15, 2021Updated 4 years ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆18Updated this week
- ☆15Dec 23, 2022Updated 3 years ago
- Data recipes and robust infrastructure for training AI agents☆290Sep 2, 2026Updated last week
- A benchmark for LLMs on complicated tasks in the terminal☆2,576Jul 11, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LLM Benchmark problems for SWE tasks in julia☆15Nov 24, 2025Updated 9 months ago
- ☆12Mar 18, 2024Updated 2 years ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- ☆18Apr 15, 2024Updated 2 years ago
- [IROS'25] COCMT☆12Aug 14, 2025Updated last year
- Materials for the LLM Evals Workshop from Weights & BIases☆15Feb 24, 2025Updated last year
- The official implementation of the EMNLP 2023 paper "Paraphrase Types for Generation and Detection"☆12Oct 20, 2024Updated last year