Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.
☆71Mar 12, 2026Updated 4 months ago
Alternatives and similar repositories for SWE-rebench-V2
Users that are interested in SWE-rebench-V2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆134Mar 31, 2026Updated 3 months ago
- Fork to run instances from SWE-rebench☆30Jun 3, 2026Updated last month
- Convert GitHub PRs into Harbor tasks☆72Jul 13, 2026Updated last week
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents☆308Jul 13, 2025Updated last year
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆183May 12, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Toolkit for measuring Claude Code and Codex performance over time against a baseline using SWEbench-lite dataset **No API key required fo…☆32Nov 22, 2025Updated 8 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆485May 18, 2026Updated 2 months ago
- Official Code of MEnvAgent☆23Feb 3, 2026Updated 5 months ago
- [ICLR2026🔥Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving☆15Feb 26, 2026Updated 4 months ago
- ☆198Mar 16, 2026Updated 4 months ago
- ☆144May 13, 2026Updated 2 months ago
- ☆21Oct 6, 2023Updated 2 years ago
- A simple extendable markdown extension for the php twig template engine.☆11Jul 31, 2023Updated 2 years ago
- Public repository for the Remote Labor Index (RLI)☆75Nov 3, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A compact high-signal benchmark for evaluating frontier agents☆19Updated this week
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆189Updated this week
- Example code using the DSPy framework.☆20May 30, 2024Updated 2 years ago
- ☆18Jun 18, 2026Updated last month
- ☆34Feb 12, 2026Updated 5 months ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆708Jul 29, 2025Updated 11 months ago
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆710Updated this week
- ☆29Jul 3, 2026Updated 2 weeks ago
- Data recipes and robust infrastructure for training AI agents☆264Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Sep 11, 2025Updated 10 months ago
- Modified Arena-Hard-Auto LLM evaluation toolkit with an emphasis on Russian language☆47Mar 20, 2025Updated last year
- ☆17Apr 9, 2025Updated last year
- Ongoing research project for code&math LLMs☆32Jul 4, 2025Updated last year
- Repo for Anonymous purpose, pls don't distribute☆10Oct 2, 2024Updated last year
- PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance☆14May 15, 2024Updated 2 years ago
- ☯️ AllenNLP training configurations for promising models on Named Entity Recognition. (BiLSTM-CRF, BiLSTM-CNN-CRF, BERT, BERT-CRF)☆15Nov 26, 2020Updated 5 years ago
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models☆60Nov 26, 2024Updated last year
- [NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!☆210Jun 11, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Realistic examples of building evals and optimizing agents with Harbor☆141Apr 23, 2026Updated 3 months ago
- This repository provides open-source code for sparse continuous distributions and corresponding Fenchel-Young losses.☆15May 10, 2023Updated 3 years ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆47Jul 7, 2026Updated 2 weeks ago
- Using stochastic gradient descent (SGD) with explicit and implicit updates to fit large-scale statistical models.☆16Aug 21, 2014Updated 11 years ago
- Repo2Run is an LLM-based agent that automates environment configuration by generating error-free Dockerfiles for Python repositories.☆195Jun 10, 2026Updated last month
- ☆82Jun 19, 2026Updated last month
- dMel: Speech Tokenization Made Simple☆19May 13, 2025Updated last year