FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research
☆231Aug 13, 2026Updated last month
Alternatives and similar repositories for frontier-swe
Users that are interested in frontier-swe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆580Updated this week
- SWE-Marathon: an ultra long-horizon SWE benchmark☆164Sep 25, 2026Updated last week
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆25May 6, 2026Updated 4 months ago
- ☆178May 13, 2026Updated 4 months ago
- Official eval scripts for JobBench☆56Sep 14, 2026Updated 2 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆95Sep 15, 2026Updated 2 weeks ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆80Updated this week
- Continual Learning Bench☆226Jul 19, 2026Updated 2 months ago
- Can Language Models Rebuild Programs From Scratch?☆943Updated this week
- [ICML 2026] RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments☆236Apr 30, 2026Updated 5 months ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆491Aug 18, 2026Updated last month
- Official Implementation of ACL2023: Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span …☆14Aug 25, 2023Updated 3 years ago
- ☆27Jan 22, 2026Updated 8 months ago
- ☆84Jun 25, 2026Updated 3 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆92Updated this week
- ☆13Feb 7, 2023Updated 3 years ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,273Mar 24, 2026Updated 6 months ago
- ☆95Jul 21, 2026Updated 2 months ago
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,522Apr 17, 2026Updated 5 months ago
- The best ChatGPT that $100 can buy.☆59Updated this week
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆792Updated this week
- Personal solutions to the Triton Puzzles☆22Jul 18, 2024Updated 2 years ago
- Framework for evaluating and improving agents☆5,791Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Measuring frontier coding agents on original, long-horizon engineering tasks☆1,785Aug 26, 2026Updated last month
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Gym-Anything: Turn any Software into an Agent Environment☆288Sep 26, 2026Updated last week
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆322Updated this week
- Repository of paper "Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis" (ACL 2025 Main)☆20Jul 19, 2025Updated last year
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆85Mar 12, 2026Updated 6 months ago
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆3,039Updated this week
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆197May 12, 2026Updated 4 months ago
- ☆18Mar 10, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LongCLI-Bench's official repository☆48Aug 14, 2026Updated last month
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆84Updated this week
- Measuring and evolving with the frontier of agent work☆834Updated this week
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆42May 31, 2026Updated 4 months ago
- Efficient retrieval head analysis with triton flash attention that supports topK probability☆13Jun 15, 2024Updated 2 years ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆748Jul 29, 2025Updated last year
- [ESEC/FSE'23] Hue: A User-Adaptive Parser for Hybrid Logs☆10Aug 24, 2023Updated 3 years ago