Run SWE-bench evaluations remotely
☆83Aug 14, 2025Updated last year
Alternatives and similar repositories for sb-cli
Users that are interested in sb-cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆798Updated this week
- Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.☆617Updated this week
- ☆138May 8, 2025Updated last year
- ☆13Mar 5, 2025Updated last year
- Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.☆284Sep 3, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Benchmarking Goal-Oriented Software Engineering☆215Jul 16, 2026Updated 2 months ago
- Docker image registry for SWE-bench, created by Epoch AI.☆20Aug 21, 2025Updated last year
- Agentless Lite: RAG-based SWE-Bench software engineering scaffold☆49Apr 15, 2025Updated last year
- A curated list of awesome Harbor ecosystem projects☆55May 29, 2026Updated 4 months ago
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,994Sep 18, 2026Updated 3 weeks ago
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents☆337Jul 13, 2025Updated last year
- ☆43Jan 8, 2025Updated last year
- ☆105Jul 17, 2024Updated 2 years ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆748Jul 29, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 8 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆74Jul 28, 2025Updated last year
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆202Sep 28, 2026Updated last week
- macOS session replay — capture, encode, and upload screen recordings for user analytics☆18Aug 5, 2026Updated 2 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆541Sep 22, 2026Updated 2 weeks ago
- CLI tool to extend the GitHub Copilot CLI to accept more selectable models☆18Oct 11, 2025Updated 11 months ago
- [ISSTA'24] A Large-Scale Dataset Capable of Enhancing the Prowess of Large Language Models for Program Testing☆12Jan 7, 2025Updated last year
- The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—b…☆8,343Updated this week
- ☆44May 16, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Landing page + leaderboard for SWE-Bench benchmark☆15Sep 1, 2026Updated last month
- [NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!☆253Sep 25, 2026Updated 2 weeks ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆147Feb 16, 2026Updated 7 months ago
- ☆28Jun 2, 2026Updated 4 months ago
- Artifact for TOSEM Submission: GiantRepair☆12Jun 26, 2024Updated 2 years ago
- ☆14Sep 12, 2024Updated 2 years ago
- ☆15Apr 26, 2025Updated last year
- ☆10Sep 10, 2023Updated 3 years ago
- SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution☆105Sep 24, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- For our ISSTA'23 paper ACETest: Automated Constraint Extraction for Testing Deep Learning Operators☆17Apr 28, 2026Updated 5 months ago
- Implementation of <Model Merging with Functional Dual Anchors>☆47Nov 23, 2025Updated 10 months ago
- ☆27Aug 16, 2025Updated last year
- A complete Lean 4 formalization of the Kakeya set problem over finite fields☆23Dec 16, 2025Updated 9 months ago
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving☆362Dec 18, 2025Updated 9 months ago
- ☆12Nov 2, 2021Updated 4 years ago
- Enhancing AI Software Engineering with Repository-level Code Graph☆305Apr 1, 2025Updated last year