Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.
☆85Mar 12, 2026Updated 6 months ago
Alternatives and similar repositories for SWE-rebench-V2
Users that are interested in SWE-rebench-V2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fork to run instances from SWE-rebench☆31Jun 3, 2026Updated 3 months ago
- ☆94Jul 21, 2026Updated 2 months ago
- Convert GitHub PRs into Harbor tasks☆85Jul 13, 2026Updated 2 months ago
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents☆335Jul 13, 2025Updated last year
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆195May 12, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training☆105Feb 28, 2026Updated 6 months ago
- Toolkit for measuring Claude Code and Codex performance over time against a baseline using SWEbench-lite dataset **No API key required fo…☆32Nov 22, 2025Updated 10 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆530Updated this week
- ☆210Mar 16, 2026Updated 6 months ago
- ☆177May 13, 2026Updated 4 months ago
- ☆21Oct 6, 2023Updated 2 years ago
- Public repository for the Remote Labor Index (RLI)☆77Nov 3, 2025Updated 10 months ago
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆230Aug 13, 2026Updated last month
- ☆23Jun 18, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆37Feb 12, 2026Updated 7 months ago
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆784Updated this week
- Automate the build, execution and test of GitHub repositories across programming languages and operating systems.☆146Updated this week
- Data recipes and robust infrastructure for training AI agents☆295Updated this week
- ☆17Sep 11, 2025Updated last year
- ☆10Jul 14, 2024Updated 2 years ago
- ☆10Jun 23, 2018Updated 8 years ago
- ☆20Apr 9, 2025Updated last year
- [NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!☆243Sep 10, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- PyTorch implementation of FAIR's paper "End-to-End Memory Network", NIPS 2015☆12Oct 19, 2017Updated 8 years ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆66Updated this week
- ☆11Oct 2, 2023Updated 2 years ago
- Include ML DL RL, knowledge and code☆12Feb 12, 2023Updated 3 years ago
- ☆12Apr 19, 2024Updated 2 years ago
- Realistic examples of building evals and optimizing agents with Harbor☆219Apr 23, 2026Updated 5 months ago
- ☆90Jun 19, 2026Updated 3 months ago
- dMel: Speech Tokenization Made Simple☆20Sep 11, 2026Updated last week
- SlopCodeBench: Measuring Code Erosion Under Iterative Specification Refinement☆208Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution☆105Sep 24, 2025Updated last year
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving☆362Dec 18, 2025Updated 9 months ago
- Python package to augment multilingual data☆15Feb 15, 2023Updated 3 years ago
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- ☆138May 8, 2025Updated last year
- ☆15Aug 1, 2026Updated last month
- A curated list of papers and resources on byte-based large language models (LLMs) — models that operate directly on raw bytes.☆17Jul 12, 2025Updated last year