Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.
☆78Mar 12, 2026Updated 5 months ago
Alternatives and similar repositories for SWE-rebench-V2
Users that are interested in SWE-rebench-V2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆137Mar 31, 2026Updated 4 months ago
- Fork to run instances from SWE-rebench☆31Jun 3, 2026Updated 2 months ago
- Convert GitHub PRs into Harbor tasks☆76Jul 13, 2026Updated last month
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆186May 12, 2026Updated 3 months ago
- ☆98Feb 28, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Toolkit for measuring Claude Code and Codex performance over time against a baseline using SWEbench-lite dataset **No API key required fo…☆32Nov 22, 2025Updated 8 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆500May 18, 2026Updated 2 months ago
- Linter for Claude Code agents, commands, and skills☆16Updated this week
- [ICLR2026🔥Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving☆15Feb 26, 2026Updated 5 months ago
- ☆37Apr 2, 2026Updated 4 months ago
- ☆207Mar 16, 2026Updated 4 months ago
- ☆156May 13, 2026Updated 3 months ago
- A simple extendable markdown extension for the php twig template engine.☆11Jul 31, 2023Updated 3 years ago
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆212Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning☆14Jun 19, 2024Updated 2 years ago
- Example code using the DSPy framework.☆20May 30, 2024Updated 2 years ago
- ☆22Jun 18, 2026Updated last month
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆740Updated this week
- Automate the build, execution and test of GitHub repositories across programming languages and operating systems.☆134Jun 16, 2026Updated last month
- ☆29Jul 3, 2026Updated last month
- Data recipes and robust infrastructure for training AI agents☆279Updated this week
- ☆16Sep 11, 2025Updated 11 months ago
- ☆10Jul 14, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆10Jun 23, 2018Updated 8 years ago
- ☆20Apr 9, 2025Updated last year
- Ongoing research project for code&math LLMs☆31Jul 4, 2025Updated last year
- Repo for Anonymous purpose, pls don't distribute☆10Oct 2, 2024Updated last year
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models☆60Nov 26, 2024Updated last year
- [NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!☆220Jun 11, 2026Updated 2 months ago
- Realistic examples of building evals and optimizing agents with Harbor☆171Apr 23, 2026Updated 3 months ago
- This repository provides open-source code for sparse continuous distributions and corresponding Fenchel-Young losses.☆15May 10, 2023Updated 3 years ago
- PyTorch implementation of FAIR's paper "End-to-End Memory Network", NIPS 2015☆12Oct 19, 2017Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆55Jul 7, 2026Updated last month
- Include ML DL RL, knowledge and code☆12Feb 12, 2023Updated 3 years ago
- ☆12Apr 19, 2024Updated 2 years ago
- Using stochastic gradient descent (SGD) with explicit and implicit updates to fit large-scale statistical models.☆16Aug 21, 2014Updated 11 years ago
- Repo2Run is an LLM-based agent that automates environment configuration by generating error-free Dockerfiles for Python repositories.☆197Jun 10, 2026Updated 2 months ago
- A Python SDK for Open Reward Standard servers and clients☆17Mar 24, 2026Updated 4 months ago
- ☆87Jun 19, 2026Updated last month