A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.
☆164Aug 30, 2026Updated 2 weeks ago
Alternatives and similar repositories for autolab
Users that are interested in autolab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34May 15, 2026Updated 3 months ago
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts☆26Feb 23, 2024Updated 2 years ago
- A unified framework for vision-language environments with Gymnasium-compatible interface☆38Mar 17, 2026Updated 5 months ago
- Official eval scripts for JobBench☆52Updated this week
- ☆15May 15, 2026Updated 3 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research☆21Updated this week
- 开源我自己 — A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source y…☆25Apr 21, 2026Updated 4 months ago
- ☆115Updated this week
- Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.☆60Jul 14, 2026Updated last month
- Repository for TRACE: Capability-Targeted Agentic Training☆118Jul 12, 2026Updated 2 months ago
- ☆50Jun 7, 2025Updated last year
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆554Updated this week
- Reference code for the Meta-Harness paper.☆1,550Updated this week
- Benchmark harness for evaluating DSPy RLMs on data analysis tasks (InfiAgent-DABench)☆24Mar 22, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- Your efficient and accurate answer verification system for RL training.☆41Jun 23, 2025Updated last year
- Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and mu…☆978Updated this week
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated last year
- [Arxiv 2025] ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions☆45Jun 11, 2025Updated last year
- ☆27Feb 26, 2026Updated 6 months ago
- Continual Learning Bench☆221Jul 19, 2026Updated last month
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- An implementation of a Meta Harness for Hermes.☆114Jul 11, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICML 2026] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning☆34May 25, 2026Updated 3 months ago
- ☆17Jun 3, 2025Updated last year
- [ICML2025] The code and data of Paper: Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation☆166Oct 25, 2024Updated last year
- Kinetics: Rethinking Test-Time Scaling Laws☆86Jul 11, 2025Updated last year
- ☆17Oct 22, 2024Updated last year
- Aurora: Unified Video Editing with a Tool-Using Agent☆147Jun 16, 2026Updated 2 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- [Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minim…☆65Sep 22, 2025Updated 11 months ago
- 🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents☆124May 28, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆233Nov 27, 2025Updated 9 months ago
- ☆192Aug 17, 2026Updated 3 weeks ago
- An infinite canvas with AI-powered code terminals☆16Feb 15, 2026Updated 6 months ago
- Production focused port of RLMs that allows the LM to call its sub-lm with DSPy signatures. Define your inputs, outputs, and tools — the …☆16Apr 14, 2026Updated 4 months ago
- Understanding deep networks and large models.☆30Jan 23, 2026Updated 7 months ago
- 💧 Query mode for agents☆108May 11, 2026Updated 4 months ago
- ☆28May 10, 2026Updated 4 months ago