SWE-Marathon: an ultra long-horizon SWE benchmark
☆138Aug 19, 2026Updated this week
Alternatives and similar repositories for swe-marathon
Users that are interested in swe-marathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Convert GitHub PRs into Harbor tasks☆77Jul 13, 2026Updated last month
- open source SWE-Atlas☆70Updated this week
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆215Aug 13, 2026Updated last week
- Measuring and evolving with the frontier of agent work☆533Updated this week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆528Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆23Jun 18, 2026Updated 2 months ago
- Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal☆273Updated this week
- ☆38May 16, 2026Updated 3 months ago
- Framework for evaluating and improving agents☆4,533Updated this week
- ☆18Aug 9, 2026Updated 2 weeks ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆163Updated this week
- Repository for results and data (coming soon!) for ClawsBench☆33Apr 8, 2026Updated 4 months ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆145Feb 16, 2026Updated 6 months ago
- Research infra for creating RL environments, post-training, and evals.☆331Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- MCP server that provides Manus-like capabilities☆39Apr 10, 2025Updated last year
- Can Language Models Rebuild Programs From Scratch?☆905Jul 26, 2026Updated 3 weeks ago
- A curated list of awesome Harbor ecosystem projects☆52May 29, 2026Updated 2 months ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Updated this week
- ☆162May 13, 2026Updated 3 months ago
- ☆122Apr 1, 2026Updated 4 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆508May 18, 2026Updated 3 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 6 months ago
- Measuring frontier coding agents on original, long-horizon engineering tasks☆1,474Aug 6, 2026Updated 2 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Terminal-Bench 2.1☆82Updated this week
- Benchmark harness and code for "SWE-fficiency: Can Language Models Optimize Real World Repositories on Real World Workloads?"☆22Feb 17, 2026Updated 6 months ago
- A benchmark for evaluating AI agents on realistic business workflows☆216Aug 4, 2026Updated 2 weeks ago
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆90Jul 12, 2026Updated last month
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 4 months ago
- ☆137Mar 31, 2026Updated 4 months ago
- Official Repository for "Training Versatile Coding Agents in Synthetic Environments"☆22Jan 11, 2026Updated 7 months ago
- MCP Atlas☆148Aug 12, 2026Updated last week
- A compact high-signal benchmark for evaluating frontier agents☆25Aug 3, 2026Updated 2 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆461Updated this week
- ☆81Jun 25, 2026Updated last month
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆72Jan 28, 2026Updated 6 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,550Jul 11, 2026Updated last month
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 7 months ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,708Jul 23, 2026Updated last month
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆749Updated this week