Toolkit for measuring Claude Code and Codex performance over time against a baseline using SWEbench-lite dataset **No API key required for Max or Pro subscribers**
☆32Nov 22, 2025Updated 7 months ago
Alternatives and similar repositories for claudecode_gemini_and_codex_swebench
Users that are interested in claudecode_gemini_and_codex_swebench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.☆71Mar 12, 2026Updated 4 months ago
- ☆18Dec 21, 2025Updated 7 months ago
- A survey and analysis of source code sandboxing☆17Jun 16, 2025Updated last year
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆143Feb 16, 2026Updated 5 months ago
- ☆21May 14, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code & data for ICLR 2024 spotlight paper: 🍯MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data☆43May 29, 2024Updated 2 years ago
- Formalization of IMO shortlist problems in Lean 4☆25May 2, 2026Updated 2 months ago
- DynAuditClaw — A security audit skill that dynamically discovers your OpenClaw agent's real configuration, designs targeted attack scenar…☆15Apr 6, 2026Updated 3 months ago
- ProofNet dataset ported into Lean 4☆31Jun 9, 2025Updated last year
- Code for paper "WildReward: Learning Reward Models from In-the-Wild Human Interactions"☆23Feb 26, 2026Updated 4 months ago
- A minimal language for Isabelle/HOL, designed for easing machine learning.☆29Updated this week
- OpenAI compatible API for open source LLMs☆16Oct 30, 2023Updated 2 years ago
- ☆15Jan 14, 2026Updated 6 months ago
- This is the repo for the paper A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression☆42Apr 23, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Page for the CVPR 2023 Tutorial - Efficient Neural Networks: From Algorithm Design to Practical Mobile Deployments☆12Jun 30, 2023Updated 3 years ago
- ☆25Feb 3, 2026Updated 5 months ago
- The data and code of the paper "Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction" (AAA…☆32Nov 26, 2025Updated 7 months ago
- A canonical source of GenAI energy benchmark and meausrements☆50Nov 29, 2025Updated 7 months ago
- Demonstration of DSPy optimization for Skill.md files☆15Dec 28, 2025Updated 6 months ago
- This repository contains the official implementation for the ECCV'22 paper, "SPIN: An Empirical Evaluation on Sharing Parameters of Isotr…☆20Sep 9, 2023Updated 2 years ago
- We trained high performing open source models on image scans of tissue biopsies to predict endoscopic categories in inflammatory bowel di…☆17Mar 3, 2023Updated 3 years ago
- [NIPS 2025 DB Spotlight] AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios☆39Dec 1, 2025Updated 7 months ago
- Try a tactic at each step in a Lean proof.☆38Jul 13, 2026Updated last week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for the paper "Reflectance-guided, contrast-accumulated histogram equalization" published in ICASSP 2020.☆10Sep 15, 2022Updated 3 years ago
- ☆21Feb 12, 2025Updated last year
- A Lean 4 Jupyter kernel via repl☆37Nov 19, 2024Updated last year
- Pytorch implementation of Mix-Shifting-MLP (MS-MLP)☆17Feb 16, 2022Updated 4 years ago
- ☆36Jan 10, 2025Updated last year
- ☆53Jun 13, 2025Updated last year
- Bencharking pipeline for evaluating Transcriptomic representations for perturbation tasks☆14Nov 5, 2024Updated last year
- ☆11Jan 25, 2019Updated 7 years ago
- Structural verification for graph databases. 5M vertices. 35 microseconds. Zero ML.☆18Jun 20, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆15Jan 9, 2026Updated 6 months ago
- ☆11Mar 20, 2023Updated 3 years ago
- Convert GitHub PRs into Harbor tasks☆72Jul 13, 2026Updated last week
- Cellmate is a sandboxing framework for BUAs that enforces strict boundaries on their behavior, ensuring safety even in the worst-case exe…☆30Jun 22, 2026Updated 3 weeks ago
- ☆15Apr 21, 2026Updated 3 months ago
- ZenStack Monorepo Demo☆11Feb 6, 2026Updated 5 months ago
- ☆12Jul 1, 2023Updated 3 years ago