A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.
☆169Aug 30, 2026Updated last month
Alternatives and similar repositories for autolab
Users that are interested in autolab are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34May 15, 2026Updated 4 months ago
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts☆26Feb 23, 2024Updated 2 years ago
- A unified framework for vision-language environments with Gymnasium-compatible interface☆38Mar 17, 2026Updated 6 months ago
- ☆15May 15, 2026Updated 4 months ago
- Official eval scripts for JobBench☆56Sep 14, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research☆21Sep 8, 2026Updated 3 weeks ago
- 开源我自己 — A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source y…☆25Apr 21, 2026Updated 5 months ago
- ☆59May 25, 2026Updated 4 months ago
- CoDA is a multi-agent framework that turns natural language queries into publication-quality visualizations.☆38Mar 20, 2026Updated 6 months ago
- ☆120Updated this week
- Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.☆60Jul 14, 2026Updated 2 months ago
- Repository for TRACE: Capability-Targeted Agentic Training☆124Jul 12, 2026Updated 2 months ago
- ☆50Jun 7, 2025Updated last year
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆580Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Lossless-first prompt compression for JSON, YAML, CSV, and Markdown. Library, CLI, MCP server, desktop app, and browser extension.☆20Updated this week
- Official code for Meta-Harness (2603.28052)☆1,633Updated this week
- Benchmark harness for evaluating DSPy RLMs on data analysis tasks (InfiAgent-DABench)☆24Mar 22, 2026Updated 6 months ago
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- Your efficient and accurate answer verification system for RL training.☆41Jun 23, 2025Updated last year
- Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and mu…☆1,044Sep 8, 2026Updated 3 weeks ago
- Official Implementation of Trajectory-Refined Distillation☆37Jun 9, 2026Updated 3 months ago
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated last year
- ☆27Feb 26, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Continual Learning Bench☆226Jul 19, 2026Updated 2 months ago
- Official repository for AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning☆17Jul 24, 2025Updated last year
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- ☆14Jul 17, 2025Updated last year
- An implementation of a Meta Harness for Hermes.☆117Jul 11, 2026Updated 2 months ago
- [ICML 2026] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning☆34May 25, 2026Updated 4 months ago
- ☆18Jun 3, 2025Updated last year
- [ICML2025] The code and data of Paper: Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation☆167Oct 25, 2024Updated last year
- Kinetics: Rethinking Test-Time Scaling Laws☆87Jul 11, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Control LLM generation format efficiently. A simple version of microsoft/aici in vllm and transformers☆14Jun 7, 2024Updated 2 years ago
- ☆17Oct 22, 2024Updated last year
- [ICLR'26] Building a Foundational Guardrail for General Agentic Systems via Synthetic Data☆49Oct 26, 2025Updated 11 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- ☆11Aug 27, 2017Updated 9 years ago
- [Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minim…☆66Sep 22, 2025Updated last year
- 🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents☆124May 28, 2026Updated 4 months ago