Convert GitHub PRs into Harbor tasks
☆72Jul 13, 2026Updated 2 weeks ago
Alternatives and similar repositories for SWE-gen
Users that are interested in SWE-gen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Marathon: an ultra long-horizon SWE benchmark☆117Updated this week
- ☆19Jun 18, 2026Updated last month
- A curated list of awesome Harbor ecosystem projects☆48May 29, 2026Updated 2 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆71Jul 28, 2025Updated last year
- ☆15Jan 14, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆135Mar 31, 2026Updated 3 months ago
- ☆119Apr 1, 2026Updated 3 months ago
- Data recipes and robust infrastructure for training AI agents☆272Updated this week
- Benchmarking Goal-Oriented Software Engineering☆195Jul 16, 2026Updated last week
- ☆81Jun 25, 2026Updated last month
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆73Mar 12, 2026Updated 4 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆51Jul 7, 2026Updated 3 weeks ago
- ☆37Apr 2, 2026Updated 3 months ago
- Measuring and evolving with the frontier of agent work☆416Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Run SWE-bench evaluations remotely☆78Aug 14, 2025Updated 11 months ago
- Framework for evaluating and improving agents☆3,647Updated this week
- New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment☆24Jun 30, 2026Updated 3 weeks ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆143Feb 16, 2026Updated 5 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆398Aug 24, 2025Updated 11 months ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆15Jun 18, 2026Updated last month
- ☆349Apr 30, 2026Updated 2 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆488May 18, 2026Updated 2 months ago
- Recovery-Bench is a benchmark for evaluating the capability of LLM agents to recover from mistakes☆27Jun 17, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆717Updated this week
- Advancing Small and Medium-sized Code Agents.☆17May 29, 2026Updated 2 months ago
- Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play☆15Jun 3, 2026Updated last month
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆29Feb 20, 2026Updated 5 months ago
- Fork to run instances from SWE-rebench☆31Jun 3, 2026Updated last month
- Realistic examples of building evals and optimizing agents with Harbor☆149Apr 23, 2026Updated 3 months ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆130Jul 12, 2026Updated 2 weeks ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆712Jul 29, 2025Updated last year
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆83Jun 13, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A compact high-signal benchmark for evaluating frontier agents☆21Updated this week
- MCP server that provides Manus-like capabilities☆39Apr 10, 2025Updated last year
- Continual Learning Bench☆189Jul 19, 2026Updated last week
- A benchmark for LLMs on complicated tasks in the terminal☆2,499Jul 11, 2026Updated 2 weeks ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- Can Language Models Rebuild Programs From Scratch?☆870Updated this week
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆25Jun 30, 2026Updated 3 weeks ago