☆23Jun 18, 2026Updated 2 months ago
Alternatives and similar repositories for terminal-bench-challenges
Users that are interested in terminal-bench-challenges are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A compact high-signal benchmark for evaluating frontier agents☆25Aug 3, 2026Updated 2 weeks ago
- Measuring and evolving with the frontier of agent work☆533Updated this week
- A curated list of awesome Harbor ecosystem projects☆52May 29, 2026Updated 2 months ago
- Benchmark harness and code for "SWE-fficiency: Can Language Models Optimize Real World Repositories on Real World Workloads?"☆22Feb 17, 2026Updated 6 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 7 months ago
- OpenTelemetry Benchmark - can AI trace your failed login?☆20Jul 14, 2026Updated last month
- ☆122Apr 1, 2026Updated 4 months ago
- Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.☆58Jul 14, 2026Updated last month
- Download Web-10K data by querying Bing Image Search☆10Feb 1, 2022Updated 4 years ago
- Agentic Research and Evaluation Suite☆109Aug 12, 2026Updated last week
- Python package for extractive NLP using the OpenAI API☆17Aug 28, 2024Updated last year
- ☆38May 16, 2026Updated 3 months ago
- ☆10Dec 17, 2020Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of "BERTs are Generative In-Context Learners"☆32Mar 14, 2025Updated last year
- Realistic examples of building evals and optimizing agents with Harbor☆184Apr 23, 2026Updated 4 months ago
- SWE-Marathon: an ultra long-horizon SWE benchmark☆138Updated this week
- A Datasette instance for searching WebVid-10M☆15Sep 30, 2022Updated 3 years ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆73Jul 28, 2025Updated last year
- [ICLR 2025] On Evluating the Durability of Safegurads for Open-Weight LLMs☆13Jun 20, 2025Updated last year
- ☆15Dec 12, 2024Updated last year
- Hasura GraphQL Engine on Render☆15Aug 28, 2023Updated 2 years ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆24Mar 18, 2025Updated last year
- Benchmarking Open-Ended Inference Optimization by AI Agents☆41Jul 6, 2026Updated last month
- ☆15Jun 7, 2024Updated 2 years ago
- JMLR Cover Letter Template☆10Dec 15, 2021Updated 4 years ago
- [COLING 2025] Official repo of paper: "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jail…☆12Jul 26, 2024Updated 2 years ago
- Demo showing how to sync data with ElectricSQL from Postgres to Cloudflare's Workers KV☆17Aug 20, 2024Updated 2 years ago
- Data recipes and robust infrastructure for training AI agents☆282Updated this week
- LLM Benchmark problems for SWE tasks in julia☆15Nov 24, 2025Updated 9 months ago
- ☆21May 31, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆17Mar 30, 2025Updated last year
- Showing how to mock data for Posts and Authors example from graphql-tools docs☆14Dec 14, 2017Updated 8 years ago
- Demo Project for Live Stream public facing demo.☆11Dec 6, 2023Updated 2 years ago
- ☆13Mar 26, 2019Updated 7 years ago
- Complex Reasoning with ReAct and LangChain☆12Apr 24, 2024Updated 2 years ago
- NLSpec instruction following benchmark for https://factory.strongdm.ai/products/attractor☆20Feb 26, 2026Updated 5 months ago
- 基于Hadoop的k_means实现对NBA球队球风聚类☆10Apr 22, 2019Updated 7 years ago