Convert GitHub PRs into Harbor tasks
☆88Jul 13, 2026Updated 2 months ago
Alternatives and similar repositories for SWE-gen
Users that are interested in SWE-gen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Marathon: an ultra long-horizon SWE benchmark☆163Updated this week
- ☆23Jun 18, 2026Updated 3 months ago
- A curated list of awesome Harbor ecosystem projects☆54May 29, 2026Updated 3 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆74Jul 28, 2025Updated last year
- ☆15Jan 14, 2026Updated 8 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆148Mar 31, 2026Updated 5 months ago
- Data recipes and robust infrastructure for training AI agents☆295Updated this week
- Benchmarking Goal-Oriented Software Engineering☆212Jul 16, 2026Updated 2 months ago
- ☆84Jun 25, 2026Updated 3 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆68Updated this week
- ☆39Apr 2, 2026Updated 5 months ago
- Run SWE-bench evaluations remotely☆82Aug 14, 2025Updated last year
- 💻 SETA: Scaling Environments for Terminal Agents☆156Jul 28, 2026Updated 2 months ago
- Framework for evaluating and improving agents☆5,637Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Measuring and evolving with the frontier of agent work☆787Updated this week
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆147Feb 16, 2026Updated 7 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆412Aug 24, 2025Updated last year
- Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains☆638Updated this week
- ☆416Apr 30, 2026Updated 4 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆531Updated this week
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆18Sep 15, 2026Updated last week
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆789Updated this week
- Advancing Small and Medium-sized Code Agents.☆17May 29, 2026Updated 3 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play☆18Updated this week
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 7 months ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆47Jul 6, 2026Updated 2 months ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆746Jul 29, 2025Updated last year
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆95Sep 15, 2026Updated last week
- Official implementation of the ACL Findings 2023 paper: Interpretable Automatic Fine-grained Inconsistency Detection in Text Summarizatio…☆15Jan 25, 2024Updated 2 years ago
- A compact high-signal benchmark for evaluating frontier agents☆39Aug 3, 2026Updated last month
- Continual Learning Bench☆226Jul 19, 2026Updated 2 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,595Jul 11, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Can Language Models Rebuild Programs From Scratch?☆934Sep 18, 2026Updated last week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆570Updated this week
- An easy-to-use manual to use the OpenBB terminal developed by 3 university students☆11Dec 11, 2022Updated 3 years ago
- ☆44May 16, 2026Updated 4 months ago
- Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?☆63Updated this week
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆230Aug 13, 2026Updated last month
- OpenTelemetry Benchmark - can AI trace your failed login?☆23Jul 14, 2026Updated 2 months ago