Convert GitHub PRs into Harbor tasks
☆72Jul 13, 2026Updated 2 weeks ago
Alternatives and similar repositories for SWE-gen
Users that are interested in SWE-gen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated list of awesome Harbor ecosystem projects☆48May 29, 2026Updated 2 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆71Jul 28, 2025Updated last year
- ☆119Apr 1, 2026Updated 3 months ago
- Data recipes and robust infrastructure for training AI agents☆271Updated this week
- Benchmarking Goal-Oriented Software Engineering☆193Jul 16, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆81Jun 25, 2026Updated last month
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆72Mar 12, 2026Updated 4 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆51Jul 7, 2026Updated 3 weeks ago
- 💻 SETA: Scaling Environments for Terminal Agents☆128Updated this week
- Framework for evaluating and improving agents☆3,611Updated this week
- New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment☆24Jun 30, 2026Updated 3 weeks ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆143Feb 16, 2026Updated 5 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆398Aug 24, 2025Updated 11 months ago
- ☆349Apr 30, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆488May 18, 2026Updated 2 months ago
- Recovery-Bench is a benchmark for evaluating the capability of LLM agents to recover from mistakes☆27Jun 17, 2026Updated last month
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆716Updated this week
- Advancing Small and Medium-sized Code Agents.☆17May 29, 2026Updated 2 months ago
- Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play☆15Jun 3, 2026Updated last month
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆29Feb 20, 2026Updated 5 months ago
- Fork to run instances from SWE-rebench☆31Jun 3, 2026Updated last month
- Realistic examples of building evals and optimizing agents with Harbor☆148Apr 23, 2026Updated 3 months ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆34Jul 6, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆130Jul 12, 2026Updated 2 weeks ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆712Jul 29, 2025Updated last year
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆83Jun 13, 2026Updated last month
- Official implementation of the ACL Findings 2023 paper: Interpretable Automatic Fine-grained Inconsistency Detection in Text Summarizatio…☆15Jan 25, 2024Updated 2 years ago
- A compact high-signal benchmark for evaluating frontier agents☆21Updated this week
- Continual Learning Bench☆189Jul 19, 2026Updated last week
- A benchmark for LLMs on complicated tasks in the terminal☆2,495Jul 11, 2026Updated 2 weeks ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- Can Language Models Rebuild Programs From Scratch?☆865Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆25Jun 30, 2026Updated 3 weeks ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆472Jul 22, 2026Updated last week
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 2 months ago
- An easy-to-use manual to use the OpenBB terminal developed by 3 university students☆10Dec 11, 2022Updated 3 years ago
- ☆36May 16, 2026Updated 2 months ago
- Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?☆54Updated this week
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆197Jul 17, 2026Updated last week