Convert GitHub PRs into Harbor tasks
☆76Jul 13, 2026Updated last month
Alternatives and similar repositories for SWE-gen
Users that are interested in SWE-gen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Marathon: an ultra long-horizon SWE benchmark☆136Updated this week
- A curated list of awesome Harbor ecosystem projects☆52May 29, 2026Updated 2 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆73Jul 28, 2025Updated last year
- ☆137Mar 31, 2026Updated 4 months ago
- ☆122Apr 1, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Data recipes and robust infrastructure for training AI agents☆281Updated this week
- Benchmarking Goal-Oriented Software Engineering☆201Jul 16, 2026Updated last month
- ☆80Jun 25, 2026Updated last month
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆56Jul 7, 2026Updated last month
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆78Mar 12, 2026Updated 5 months ago
- ☆37Apr 2, 2026Updated 4 months ago
- Measuring and evolving with the frontier of agent work☆507Updated this week
- Run SWE-bench evaluations remotely☆79Aug 14, 2025Updated last year
- 💻 SETA: Scaling Environments for Terminal Agents☆136Jul 28, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Framework for evaluating and improving agents☆4,366Updated this week
- Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal☆264Updated this week
- New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment☆24Jun 30, 2026Updated last month
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆145Feb 16, 2026Updated 6 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆406Aug 24, 2025Updated 11 months ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Jul 30, 2026Updated 2 weeks ago
- ☆379Apr 30, 2026Updated 3 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆505May 18, 2026Updated 3 months ago
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆746Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 5 months ago
- Fork to run instances from SWE-rebench☆31Jun 3, 2026Updated 2 months ago
- Realistic examples of building evals and optimizing agents with Harbor☆177Apr 23, 2026Updated 3 months ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆40Jul 6, 2026Updated last month
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆722Jul 29, 2025Updated last year
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆87Jun 13, 2026Updated 2 months ago
- Official implementation of the ACL Findings 2023 paper: Interpretable Automatic Fine-grained Inconsistency Detection in Text Summarizatio…☆15Jan 25, 2024Updated 2 years ago
- A compact high-signal benchmark for evaluating frontier agents☆25Aug 3, 2026Updated 2 weeks ago
- MCP server that provides Manus-like capabilities☆39Apr 10, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Continual Learning Bench☆202Jul 19, 2026Updated 3 weeks ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,541Jul 11, 2026Updated last month
- Can Language Models Rebuild Programs From Scratch?☆898Jul 26, 2026Updated 3 weeks ago
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆28Jun 30, 2026Updated last month
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆518Updated this week
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 3 months ago
- An easy-to-use manual to use the OpenBB terminal developed by 3 university students☆10Dec 11, 2022Updated 3 years ago