FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research
☆197Jul 17, 2026Updated last week
Alternatives and similar repositories for frontier-swe
Users that are interested in frontier-swe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Marathon: an ultra long-horizon SWE benchmark☆117Updated this week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆475Jul 22, 2026Updated last week
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 2 months ago
- Official eval scripts for JobBench☆32Jul 18, 2026Updated last week
- ☆145May 13, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆83Jun 13, 2026Updated last month
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆50May 30, 2026Updated last month
- Continual Learning Bench☆189Jul 19, 2026Updated last week
- Can Language Models Rebuild Programs From Scratch?☆870Updated this week
- [ICML 2026] RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments☆226Apr 30, 2026Updated 2 months ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆443Updated this week
- Official Implementation of ACL2023: Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span …☆14Aug 25, 2023Updated 2 years ago
- ☆24Jan 22, 2026Updated 6 months ago
- ☆81Jun 25, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆90Jul 12, 2026Updated 2 weeks ago
- ☆13Feb 7, 2023Updated 3 years ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,163Mar 24, 2026Updated 4 months ago
- ☆87Jul 21, 2026Updated last week
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,469Apr 17, 2026Updated 3 months ago
- The best ChatGPT that $100 can buy.☆56Updated this week
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆717Updated this week
- Personal solutions to the Triton Puzzles☆21Jul 18, 2024Updated 2 years ago
- Framework for evaluating and improving agents☆3,647Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Measuring frontier coding agents on original, long-horizon engineering tasks☆1,262Jul 22, 2026Updated last week
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Gym-Anything: Turn any Software into an Agent Environment☆264Updated this week
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆1,809Updated this week
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆288Updated this week
- Measuring and evolving with the frontier of agent work☆416Updated this week
- Repository of paper "Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis" (ACL 2025 Main)☆19Jul 19, 2025Updated last year
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆73Mar 12, 2026Updated 4 months ago
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆56Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LongCLI-Bench's official repository☆44May 25, 2026Updated 2 months ago
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆183May 12, 2026Updated 2 months ago
- ☆18Mar 10, 2023Updated 3 years ago
- Convert GitHub PRs into Harbor tasks☆72Jul 13, 2026Updated 2 weeks ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆35May 31, 2026Updated last month
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆24Updated this week
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆712Jul 29, 2025Updated last year