Convert GitHub PRs into Harbor tasks
☆82Jul 13, 2026Updated last month
Alternatives and similar repositories for SWE-gen
Users that are interested in SWE-gen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Marathon: an ultra long-horizon SWE benchmark☆150Updated this week
- ☆23Jun 18, 2026Updated 2 months ago
- ☆15Jan 14, 2026Updated 7 months ago
- ☆122Apr 1, 2026Updated 5 months ago
- ☆144Mar 31, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Data recipes and robust infrastructure for training AI agents☆287Updated this week
- Benchmarking Goal-Oriented Software Engineering☆206Jul 16, 2026Updated last month
- ☆82Jun 25, 2026Updated 2 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆63Jul 7, 2026Updated 2 months ago
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆80Mar 12, 2026Updated 5 months ago
- Measuring and evolving with the frontier of agent work☆636Updated this week
- Run SWE-bench evaluations remotely☆79Aug 14, 2025Updated last year
- 💻 SETA: Scaling Environments for Terminal Agents☆146Jul 28, 2026Updated last month
- Framework for evaluating and improving agents☆5,010Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment☆26Jun 30, 2026Updated 2 months ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆146Feb 16, 2026Updated 6 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆409Aug 24, 2025Updated last year
- Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains☆542Updated this week
- ☆401Apr 30, 2026Updated 4 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆518May 18, 2026Updated 3 months ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆18Aug 21, 2026Updated 2 weeks ago
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆760Aug 31, 2026Updated last week
- Advancing Small and Medium-sized Code Agents.☆18May 29, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 6 months ago
- Fork to run instances from SWE-rebench☆30Jun 3, 2026Updated 3 months ago
- Realistic examples of building evals and optimizing agents with Harbor☆203Apr 23, 2026Updated 4 months ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆181Aug 29, 2026Updated last week
- Benchmarking Open-Ended Inference Optimization by AI Agents☆43Jul 6, 2026Updated 2 months ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆732Jul 29, 2025Updated last year
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆91Updated this week
- Official implementation of the ACL Findings 2023 paper: Interpretable Automatic Fine-grained Inconsistency Detection in Text Summarizatio…☆15Jan 25, 2024Updated 2 years ago
- A compact high-signal benchmark for evaluating frontier agents☆27Aug 3, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Continual Learning Bench☆218Jul 19, 2026Updated last month
- A benchmark for LLMs on complicated tasks in the terminal☆2,569Jul 11, 2026Updated last month
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- Can Language Models Rebuild Programs From Scratch?☆921Jul 26, 2026Updated last month
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆28Jun 30, 2026Updated 2 months ago
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆23May 6, 2026Updated 4 months ago
- An easy-to-use manual to use the OpenBB terminal developed by 3 university students☆10Dec 11, 2022Updated 3 years ago