Evaluation harness for OpenHands V1.
☆121Sep 4, 2026Updated last week
Alternatives and similar repositories for benchmarks
Users that are interested in benchmarks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A clean, modular SDK for building AI agents with OpenHands V1.☆1,100Updated this week
- Lightweight OpenHands CLI in a binary executable☆252Updated this week
- Public registry for OpenHands extensions.☆144Updated this week
- ☆16Nov 1, 2025Updated 10 months ago
- ☆17Aug 30, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Jun 11, 2026Updated 3 months ago
- ☆38Mar 7, 2026Updated 6 months ago
- Agent computer interface for AI software engineer.☆135Apr 16, 2026Updated 4 months ago
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆193May 12, 2026Updated 4 months ago
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆765Updated this week
- [ACL25' Findings] SWE-Dev is an SWE agent with a scalable test case construction pipeline.☆66Jul 21, 2025Updated last year
- Official implementation for the paper, StackEval: Benchmarking LLMs in Coding Assistance, https://arxiv.org/abs/2412.05288☆22Oct 30, 2024Updated last year
- ALAS: Autonomous Learning Agent System☆19Aug 14, 2025Updated last year
- ☆18Mar 28, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆39Dec 24, 2025Updated 8 months ago
- Continual Memorization of Factoids in Large Language Models☆12Nov 20, 2024Updated last year
- ☆48Dec 12, 2024Updated last year
- ☆13Mar 5, 2025Updated last year
- Continual harness optimization☆71Aug 6, 2026Updated last month
- Scalable, cloud-native infrastructure for evaluating AI agents across any benchmark.☆37Updated this week
- The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—b…☆7,450Updated this week
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,829Sep 2, 2026Updated last week
- SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution☆105Sep 24, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents☆333Jul 13, 2025Updated last year
- Advances and Frontiers of LLM-based Issue Resolution in Software Engineering A Comprehensive Survey☆87Sep 2, 2026Updated last week
- ☆20Dec 14, 2024Updated last year
- ☆15Nov 18, 2025Updated 9 months ago
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?☆522May 18, 2026Updated 3 months ago
- ☆94Jul 21, 2026Updated last month
- Framework for evaluating and improving agents☆5,158Updated this week
- ☆23Nov 26, 2025Updated 9 months ago
- ☆18Sep 23, 2025Updated 11 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Benchmarking Complex Instruction-Following with Multiple Constraints Composition (NeurIPS 2024 Datasets and Benchmarks Track)☆104Feb 20, 2025Updated last year
- ☆20Aug 9, 2026Updated last month
- Codebase for Linguistic Collapse: Neural Collapse in (Large) Language Models [NeurIPS 2024] [arXiv:2405.17767]☆18Apr 14, 2025Updated last year
- Inspect AI interface to Harbor tasks☆22Updated this week
- Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.☆588Updated this week
- ☆28Jul 14, 2025Updated last year
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 11 months ago