Run SWE-bench evaluations remotely
☆78Aug 14, 2025Updated 11 months ago
Alternatives and similar repositories for sb-cli
Users that are interested in sb-cli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆709Jul 13, 2026Updated last week
- Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.☆554Updated this week
- ☆139May 8, 2025Updated last year
- ☆13Mar 5, 2025Updated last year
- Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.☆274Mar 29, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Benchmarking Goal-Oriented Software Engineering☆189Updated this week
- Agentless Lite: RAG-based SWE-Bench software engineering scaffold☆49Apr 15, 2025Updated last year
- A manager for terminal+IDE projects☆23Jul 7, 2026Updated 2 weeks ago
- ☆12Jan 31, 2024Updated 2 years ago
- ☆17Apr 9, 2025Updated last year
- A curated list of awesome Harbor ecosystem projects☆46May 29, 2026Updated last month
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,459Apr 1, 2026Updated 3 months ago
- Convert GitHub PRs into Harbor tasks☆71Jul 13, 2026Updated last week
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents☆307Jul 13, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A simple, fast local viewer for your Claude Code sessions.☆17Oct 8, 2025Updated 9 months ago
- ☆28Jun 2, 2026Updated last month
- [ACL25] FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation☆57Jan 28, 2026Updated 5 months ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆122Jul 12, 2026Updated last week
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 5 months ago
- Community themes for Claude Code. Use tweakcc to apply them.☆16Updated this week
- [NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!☆209Jun 11, 2026Updated last month
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆485May 18, 2026Updated 2 months ago
- The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—b…☆5,914Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Gateway for inspecting and auditing your MCP servers☆23Dec 6, 2025Updated 7 months ago
- ☆36May 16, 2026Updated 2 months ago
- [NeurIPS'25] Official codebase for "SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution"☆712Mar 16, 2025Updated last year
- Turn any Office style document to markdown☆23Jul 10, 2026Updated last week
- Using cloudflare workers and DOs to make a https tunnel that scales☆27Nov 22, 2025Updated 7 months ago
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆142Feb 16, 2026Updated 5 months ago
- Run a real Linux desktop for AI agents.☆28Jun 16, 2026Updated last month
- A complete Lean 4 formalization of the Kakeya set problem over finite fields☆19Dec 16, 2025Updated 7 months ago
- Prolog implemented in Python☆12Sep 6, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆14Jun 26, 2025Updated last year
- Artifact for TOSEM Submission: GiantRepair☆12Jun 26, 2024Updated 2 years ago
- ☆11Sep 10, 2023Updated 2 years ago
- ☆13Sep 12, 2024Updated last year
- Implementation of <Model Merging with Functional Dual Anchors>☆46Nov 23, 2025Updated 7 months ago
- SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution☆105Sep 24, 2025Updated 9 months ago
- For our ISSTA'23 paper ACETest: Automated Constraint Extraction for Testing Deep Learning Operators☆17Apr 28, 2026Updated 2 months ago