FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale
☆50May 30, 2026Updated 2 months ago
Alternatives and similar repositories for FrontierSmith
Users that are interested in FrontierSmith are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆302Updated this week
- A simple SQL parser based on Apache Calcite.☆14May 8, 2026Updated 3 months ago
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 3 months ago
- [ESEC/FSE'23] Hue: A User-Adaptive Parser for Hybrid Logs☆10Aug 24, 2023Updated 2 years ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 30, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A toolkit for testing and improving named entity recognition [ESEC/FSE'23]☆11Aug 31, 2023Updated 2 years ago
- Visualization synthesis☆15May 12, 2021Updated 5 years ago
- Programmable chat templates for LLM training and inference.☆149Updated this week
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆145Feb 16, 2026Updated 6 months ago
- Official Codebase: LT2: Linear-Time Looped Transformers.☆52Jul 27, 2026Updated 3 weeks ago
- [ACL'25] UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench☆36Aug 12, 2025Updated last year
- Implementation and explorations into DiscoRL, Discovering state-of-the-art reinforcement learning algorithms, David Silver's last work at…☆21Jun 13, 2026Updated 2 months ago
- Spatialyze: A Geospatial Video Analytic System with Spatial-Aware Optimizations☆13Mar 3, 2025Updated last year
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆215Aug 13, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Continual Learning Bench☆205Jul 19, 2026Updated last month
- ☆23Jun 2, 2026Updated 2 months ago
- The best ChatGPT that $100 can buy.☆58Updated this week
- Web page for "🍅HumanTOMATO: Text-aligned Whole-body Motion Generation".☆15May 25, 2024Updated 2 years ago
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆87Jun 13, 2026Updated 2 months ago
- Repository for getting started with the OfficeQA Benchmark.☆180Aug 6, 2026Updated 2 weeks ago
- A toolkit for hybrid log parsing☆18Aug 23, 2023Updated 3 years ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 2 months ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆46Apr 1, 2026Updated 4 months ago
- ☆55Jun 25, 2026Updated last month
- ☆137Mar 31, 2026Updated 4 months ago
- A toolkit for Light Log Anomaly Detection [ICSE'24]☆23Feb 22, 2025Updated last year
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆616Updated this week
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆46Aug 4, 2026Updated 2 weeks ago
- ☆16May 23, 2026Updated 3 months ago
- A creative coding environment where Claude can express itself through generative art using p5.js. See tweet thread for examples: https://…☆13Feb 3, 2026Updated 6 months ago
- Implementation of RigidFormer, Learning Rigid Dynamics using Transformers☆32Aug 11, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Preview Code for Continuum Paper☆98Aug 13, 2026Updated last week
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆43May 25, 2026Updated 2 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆56Jul 7, 2026Updated last month
- ☆21Jun 9, 2025Updated last year
- ☆23Jun 16, 2026Updated 2 months ago
- Research artifacts from Recursive's automated AI research system☆221Jun 11, 2026Updated 2 months ago
- ☆48Aug 6, 2026Updated 2 weeks ago