SWE-Marathon: an ultra long-horizon SWE benchmark
☆125Aug 1, 2026Updated this week
Alternatives and similar repositories for swe-marathon
Users that are interested in swe-marathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Convert GitHub PRs into Harbor tasks☆73Jul 13, 2026Updated 3 weeks ago
- open source SWE-Atlas☆63Jul 20, 2026Updated 2 weeks ago
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆202Jul 17, 2026Updated 2 weeks ago
- Measuring and evolving with the frontier of agent work☆436Updated this week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆484Jul 22, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Terminal-Bench Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal☆230Updated this week
- ☆20Jun 18, 2026Updated last month
- ☆38May 16, 2026Updated 2 months ago
- Framework for evaluating and improving agents☆3,786Updated this week
- ☆17Jul 11, 2026Updated 3 weeks ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆134Jul 12, 2026Updated 3 weeks ago
- ☆61May 21, 2025Updated last year
- Repository for results and data (coming soon!) for ClawsBench☆31Apr 8, 2026Updated 3 months ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆143Feb 16, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Research infra for creating RL environments, post-training, and evals☆309Updated this week
- MCP server that provides Manus-like capabilities☆39Apr 10, 2025Updated last year
- An agent for auditing repositories of traces for violations of safety properties. Automatically finds cheating (task-level gaming and har…☆15Jun 6, 2026Updated last month
- Can Language Models Rebuild Programs From Scratch?☆875Jul 26, 2026Updated last week
- A curated list of awesome Harbor ecosystem projects☆51May 29, 2026Updated 2 months ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Updated this week
- ☆152May 13, 2026Updated 2 months ago
- Academic papers and works related to SWE-bench and SWE-agents☆15Dec 8, 2025Updated 7 months ago
- ☆121Apr 1, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆491May 18, 2026Updated 2 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆29Feb 20, 2026Updated 5 months ago
- Measuring frontier coding agents on original, long-horizon engineering tasks☆1,305Jul 22, 2026Updated last week
- Terminal-Bench 2.1☆63Jul 22, 2026Updated last week
- Benchmark harness and code for "SWE-fficiency: Can Language Models Optimize Real World Repositories on Real World Workloads?"☆21Feb 17, 2026Updated 5 months ago
- A benchmark for evaluating AI agents on realistic business workflows☆170Jul 16, 2026Updated 2 weeks ago
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆90Jul 12, 2026Updated 3 weeks ago
- ☆136Mar 31, 2026Updated 4 months ago
- Official Repository for "Training Versatile Coding Agents in Synthetic Environments"☆22Jan 11, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- MCP Atlas☆137Updated this week
- A compact high-signal benchmark for evaluating frontier agents☆21Jul 28, 2026Updated last week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆449Updated this week
- ☆80Jun 25, 2026Updated last month
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆72Jan 28, 2026Updated 6 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,518Jul 11, 2026Updated 3 weeks ago
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 6 months ago