FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale
☆50May 30, 2026Updated 2 months ago
Alternatives and similar repositories for FrontierSmith
Users that are interested in FrontierSmith are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆288Updated this week
- A simple SQL parser based on Apache Calcite.☆14May 8, 2026Updated 2 months ago
- [ICML 2026] SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark☆22May 6, 2026Updated 2 months ago
- [ESEC/FSE'23] Hue: A User-Adaptive Parser for Hybrid Logs☆10Aug 24, 2023Updated 2 years ago
- A toolkit for testing and improving named entity recognition [ESEC/FSE'23]☆11Aug 31, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Programmable chat templates for LLM training and inference.☆135Updated this week
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆143Feb 16, 2026Updated 5 months ago
- Official Codebase: LT2: Linear-Time Looped Transformers.☆50Jul 27, 2026Updated last week
- [ACL'25] UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench☆36Aug 12, 2025Updated 11 months ago
- Implementation and explorations into DiscoRL, Discovering state-of-the-art reinforcement learning algorithms, David Silver's last work at…☆21Jun 13, 2026Updated last month
- FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research☆202Jul 17, 2026Updated 2 weeks ago
- Continual Learning Bench☆191Jul 19, 2026Updated 2 weeks ago
- The best ChatGPT that $100 can buy.☆57Updated this week
- Web page for "🍅HumanTOMATO: Text-aligned Whole-body Motion Generation".☆15May 25, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"☆83Jun 13, 2026Updated last month
- ☆11Aug 7, 2023Updated 2 years ago
- Repository for getting started with the OfficeQA Benchmark.☆165Updated this week
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆34May 26, 2026Updated 2 months ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- ☆46Apr 1, 2026Updated 4 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- ☆54Jun 25, 2026Updated last month
- A toolkit for Light Log Anomaly Detection [ICSE'24]☆23Feb 22, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆45Jun 1, 2026Updated 2 months ago
- ☆16May 23, 2026Updated 2 months ago
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated 10 months ago
- A creative coding environment where Claude can express itself through generative art using p5.js. See tweet thread for examples: https://…☆13Feb 3, 2026Updated 6 months ago
- Implementation of RigidFormer, Learning Rigid Dynamics using Transformers☆29Jun 14, 2026Updated last month
- Preview Code for Continuum Paper☆93Jul 20, 2026Updated 2 weeks ago
- Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers☆42Jul 1, 2026Updated last month
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆41May 25, 2026Updated 2 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆51Jul 7, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆32Mar 11, 2026Updated 4 months ago
- ☆22Jun 16, 2026Updated last month
- Research artifacts from Recursive's automated AI research system☆194Jun 11, 2026Updated last month
- ☆42Jul 15, 2026Updated 2 weeks ago
- [ICSE 2026] [ACM SIGSOFT Distinguished Paper Award] SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Re…☆20Apr 22, 2025Updated last year
- Meshcapade support for Unreal Editor for Fortnite (UEFN)☆22Apr 17, 2024Updated 2 years ago
- Official code for the All is not Lost paper☆16Mar 27, 2026Updated 4 months ago