This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?"
☆1,430Jul 18, 2025Updated last year
Alternatives and similar repositories for SWELancer-Benchmark
Users that are interested in SWELancer-Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- OpenAI Frontier Evals☆1,309Apr 21, 2026Updated 5 months ago
- Agentless🐱: an agentless approach to automatically solve software development problems☆2,123Dec 22, 2024Updated last year
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆748Jul 29, 2025Updated last year
- [NeurIPS'25] Official codebase for "SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution"☆719Mar 16, 2025Updated last year
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,994Sep 18, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving☆362Dec 18, 2025Updated 9 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,769Apr 24, 2026Updated 5 months ago
- Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.☆284Sep 3, 2026Updated last month
- ☆143Jun 6, 2025Updated last year
- Democratizing Reinforcement Learning for LLMs☆5,860Updated this week
- SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersec…☆20,514Updated this week
- Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.☆22,043Apr 15, 2026Updated 5 months ago
- MLGym A New Framework and Benchmark for Advancing AI Research Agents☆626Aug 10, 2025Updated last year
- ☆4,655Apr 22, 2026Updated 5 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Commit0: Library Generation from Scratch☆194Feb 24, 2026Updated 7 months ago
- Renderer for the harmony response format to be used with gpt-oss☆4,523Apr 8, 2026Updated 6 months ago
- Sky-T1: Train your own O1 preview model within $450☆3,401Jul 12, 2025Updated last year
- ☆643Sep 1, 2025Updated last year
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents☆798Updated this week
- 👩⚖️ Agent-as-a-Judge: The Magic for Open-Endedness☆833Mar 28, 2026Updated 6 months ago
- Fully open reproduction of DeepSeek-R1☆26,476Oct 2, 2026Updated last week
- s1: Simple test-time scaling☆6,674Jun 25, 2025Updated last year
- [ICML 2025 Oral] CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction☆572May 6, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- 🙌 OpenHands: AI-Driven Development☆90,380Updated this week
- A benchmark for LLMs on complicated tasks in the terminal☆2,600Jul 11, 2026Updated 2 months ago
- Muon is Scalable for LLM Training☆1,548Aug 3, 2025Updated last year
- Minimal reproduction of DeepSeek R1-Zero