A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.
☆158Jun 17, 2026Updated last month
Alternatives and similar repositories for autolab
Users that are interested in autolab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32May 15, 2026Updated 2 months ago
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts☆26Feb 23, 2024Updated 2 years ago
- A unified framework for vision-language environments with Gymnasium-compatible interface☆35Mar 17, 2026Updated 4 months ago
- Official eval scripts for JobBench☆35Jul 18, 2026Updated 2 weeks ago
- ☆15May 15, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research☆20Jul 24, 2026Updated last week
- 开源我自己 — A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source y…☆25Apr 21, 2026Updated 3 months ago
- ☆47May 25, 2026Updated 2 months ago
- ☆81Updated this week
- Official codebase for the paper "WorldMemArena: Evaluating Multimodal Agent Memory Through Action–World Interaction"☆23May 29, 2026Updated 2 months ago
- Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.☆54Jul 14, 2026Updated 3 weeks ago
- Repository for TRACE: Capability-Targeted Agentic Training☆109Jul 12, 2026Updated 3 weeks ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆484Jul 22, 2026Updated last week
- Lossless-first prompt compression for JSON, YAML, CSV, and Markdown. Library, CLI, MCP server, desktop app, and browser extension.☆20Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Reference code for the Meta-Harness paper.☆1,357Jul 11, 2026Updated 3 weeks ago
- Benchmark harness for evaluating DSPy RLMs on data analysis tasks (InfiAgent-DABench)☆23Mar 22, 2026Updated 4 months ago
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- CORAL is a robust, lightweight infrastructure for multi-agent autonomous self-evolution, built for autoresearch. Works with Claude Code, …☆869Updated this week
- Your efficient and accurate answer verification system for RL training.☆42Jun 23, 2025Updated last year
- Aurora: Unified Video Editing with a Tool-Using Agent☆59Jun 16, 2026Updated last month
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 11 months ago
- [Arxiv 2025] ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions☆45Jun 11, 2025Updated last year
- Continual Learning Bench☆191Jul 19, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆27Feb 26, 2026Updated 5 months ago
- Official repository for AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning☆17Jul 24, 2025Updated last year
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- ☆14Jul 17, 2025Updated last year
- An implementation of a Meta Harness for Hermes.☆104Jul 11, 2026Updated 3 weeks ago
- [ICML 2026] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning☆34May 25, 2026Updated 2 months ago
- ☆17Jun 3, 2025Updated last year
- [ICML2025] The code and data of Paper: Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation☆163Oct 25, 2024Updated last year
- Kinetics: Rethinking Test-Time Scaling Laws☆87Jul 11, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆17Oct 22, 2024Updated last year
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆18Jul 4, 2025Updated last year
- ☆11Aug 27, 2017Updated 8 years ago
- [Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minim…☆63Sep 22, 2025Updated 10 months ago
- 🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents☆120May 28, 2026Updated 2 months ago
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆229Nov 27, 2025Updated 8 months ago
- ☆187Jul 14, 2026Updated 3 weeks ago