FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research
☆229Aug 13, 2026Updated last month
Alternatives and similar repositories for frontier-swe
Users that are interested in frontier-swe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆554Updated this week
- SWE-Marathon: an ultra long-horizon SWE benchmark☆152Updated this week
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆23May 6, 2026Updated 4 months ago
- Official eval scripts for JobBench☆52Updated this week
- ☆173May 13, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆52May 30, 2026Updated 3 months ago
- Continual Learning Bench☆221Jul 19, 2026Updated last month
- Can Language Models Rebuild Programs From Scratch?☆924Updated this week
- [ICML 2026] RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments☆232Apr 30, 2026Updated 4 months ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆482Aug 18, 2026Updated 3 weeks ago
- Official Implementation of ACL2023: Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span …☆14Aug 25, 2023Updated 3 years ago
- ☆27Jan 22, 2026Updated 7 months ago
- ☆82Jun 25, 2026Updated 2 months ago
- [NeurIPS '25] GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents☆91Jul 12, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆13Feb 7, 2023Updated 3 years ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,239Mar 24, 2026Updated 5 months ago
- ☆94Jul 21, 2026Updated last month
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,504Apr 17, 2026Updated 4 months ago
- The best ChatGPT that $100 can buy.☆58Sep 3, 2026Updated last week
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆765Updated this week
- Personal solutions to the Triton Puzzles☆22Jul 18, 2024Updated 2 years ago
- Framework for evaluating and improving agents☆5,158Updated this week
- Measuring frontier coding agents on original, long-horizon engineering tasks☆1,668Aug 26, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Gym-Anything: Turn any Software into an Agent Environment☆282Updated this week
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆313Updated this week
- Repository of paper "Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis" (ACL 2025 Main)☆20Jul 19, 2025Updated last year
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆2,810Updated this week
- LongCLI-Bench's official repository☆47Aug 14, 2026Updated 3 weeks ago
- Measuring and evolving with the frontier of agent work☆673Updated this week
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆83Mar 12, 2026Updated 6 months ago
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks☆193May 12, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆18Mar 10, 2023Updated 3 years ago
- kernelbench.com — GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.☆77Updated this week
- Convert GitHub PRs into Harbor tasks☆83Jul 13, 2026Updated 2 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆40May 31, 2026Updated 3 months ago
- Efficient retrieval head analysis with triton flash attention that supports topK probability☆13Jun 15, 2024Updated 2 years ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆735Jul 29, 2025Updated last year
- [ESEC/FSE'23] Hue: A User-Adaptive Parser for Hybrid Logs☆10Aug 24, 2023Updated 3 years ago