☆20Jun 18, 2026Updated last month
Alternatives and similar repositories for terminal-bench-challenges
Users that are interested in terminal-bench-challenges are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A compact high-signal benchmark for evaluating frontier agents☆21Jul 28, 2026Updated last week
- Convert GitHub PRs into Harbor tasks☆73Jul 13, 2026Updated 3 weeks ago
- Measuring and evolving with the frontier of agent work☆436Updated this week
- A curated list of awesome Harbor ecosystem projects☆51May 29, 2026Updated 2 months ago
- Terminal-Bench Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal☆230Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Benchmark harness and code for "SWE-fficiency: Can Language Models Optimize Real World Repositories on Real World Workloads?"☆21Feb 17, 2026Updated 5 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆29Feb 20, 2026Updated 5 months ago
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆44Jun 25, 2026Updated last month
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 6 months ago
- OpenTelemetry Benchmark - can AI trace your failed login?☆20Jul 14, 2026Updated 2 weeks ago
- Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.☆58Jul 14, 2026Updated 2 weeks ago
- open source SWE-Atlas☆63Jul 20, 2026Updated 2 weeks ago
- Download Web-10K data by querying Bing Image Search☆10Feb 1, 2022Updated 4 years ago
- Python package for extractive NLP using the OpenAI API☆17Aug 28, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆38May 16, 2026Updated 2 months ago
- ☆12Mar 3, 2022Updated 4 years ago
- [CHIL 2024] Interpretation of Intracardiac Electrograms Through Textual Representations☆12Sep 4, 2024Updated last year
- ☆10Dec 17, 2020Updated 5 years ago
- Official implementation of "BERTs are Generative In-Context Learners"☆32Mar 14, 2025Updated last year
- Realistic examples of building evals and optimizing agents with Harbor☆158Apr 23, 2026Updated 3 months ago
- SWE-Marathon: an ultra long-horizon SWE benchmark☆125Updated this week
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆71Jul 28, 2025Updated last year
- ☆11Apr 24, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2025] On Evluating the Durability of Safegurads for Open-Weight LLMs☆13Jun 20, 2025Updated last year
- ☆14Dec 12, 2024Updated last year
- Samaya AI's FrontierFinance Benchmark Grader☆17Jul 16, 2026Updated 2 weeks ago
- Hasura GraphQL Engine on Render☆15Aug 28, 2023Updated 2 years ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Updated this week
- ☆357Apr 30, 2026Updated 3 months ago
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the p…☆13Jan 26, 2025Updated last year
- JMLR Cover Letter Template☆10Dec 15, 2021Updated 4 years ago
- ☆15Dec 23, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Evaluating Durability: Benchmark Insights into Multimodal Watermarking☆12Jun 7, 2024Updated 2 years ago
- Data recipes and robust infrastructure for training AI agents☆273Updated this week
- A benchmark for LLMs on complicated tasks in the terminal☆2,518Jul 11, 2026Updated 3 weeks ago
- LLM Benchmark problems for SWE tasks in julia☆15Nov 24, 2025Updated 8 months ago
- Adversarially Robust Generalization Just Requires More Unlabeled Data☆11Aug 8, 2019Updated 6 years ago
- 🌳 A compressed rank/select dictionary exploiting approximate linearity and repetitiveness.☆15Jun 28, 2022Updated 4 years ago
- ☆17May 31, 2026Updated 2 months ago