Public repository containing METR's DVC pipeline for eval data analysis
☆331Mar 6, 2026Updated 6 months ago
Alternatives and similar repositories for eval-analysis-public
Users that are interested in eval-analysis-public are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Vivaria is METR's tool for running evaluations and conducting agent elicitation research.☆142May 18, 2026Updated 4 months ago
- ☆24Jul 28, 2026Updated last month
- ☆18Dec 10, 2025Updated 9 months ago
- ☆131Jul 28, 2026Updated last month
- Running UK AISI's Inspect in the Cloud☆25May 6, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆18Mar 10, 2026Updated 6 months ago
- ☆36Jun 4, 2025Updated last year
- Data visualization for Inspect AI large language model evalutions.☆22Jul 15, 2026Updated 2 months ago
- NeurIPS 2024: SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation☆13May 24, 2025Updated last year
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 7 months ago
- ☆14Aug 14, 2026Updated last month
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.☆241Aug 24, 2026Updated last month
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 10 months ago
- METR Task Standard☆197Feb 3, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆28Apr 1, 2026Updated 5 months ago
- Inspect: A framework for large language model evaluations☆2,850Updated this week
- An alignment auditing agent capable of quickly exploring alignment hypothesis☆1,346Updated this week
- Collection of evals for Inspect AI☆680Updated this week
- ☆75Jun 16, 2026Updated 3 months ago
- A repository that holds templates, examples, and tests to help external parties submit tasks to AISI that conform with the Autonomous Sys…☆11Jan 16, 2026Updated 8 months ago
- ☆57Sep 9, 2026Updated 2 weeks ago
- Shaping capabilities with token-level pretraining data filtering☆96Jan 28, 2026Updated 7 months ago
- ☆16Nov 14, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Public repository for the Remote Labor Index (RLI)☆77Nov 3, 2025Updated 10 months ago
- ☆91Nov 21, 2025Updated 10 months ago
- Generate Python docstrings automatically with LLM and syntax trees☆20Jun 13, 2025Updated last year
- ☆25Apr 23, 2024Updated 2 years ago
- ☆11Jun 2, 2021Updated 5 years ago
- The open-source AISI toolkit for sandboxing agentic evaluations☆36Aug 7, 2025Updated last year
- ☆28Jun 1, 2026Updated 3 months ago
- Benchmarking Dark Patterns in LLMs (ICLR 2025)☆18Mar 29, 2025Updated last year
- ☆311Jul 1, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A toolkit for describing model features and intervening on those features to steer behavior.☆262Mar 16, 2026Updated 6 months ago
- ☆16Jun 17, 2025Updated last year
- Programmatic memory for long-horizon LLM agents: the harness appends everything to one log, and the agent searches it with code. 97.4% on…☆457Aug 21, 2026Updated last month
- Cited 83-model x 49-benchmark LLM evaluation matrix with 18 matrix completion methods☆41Feb 25, 2026Updated 6 months ago
- [ICML '26] Code, Data and Red Teaming for ZeroBench☆70Aug 11, 2026Updated last month
- Inference API for many LLMs and other useful tools for empirical research☆137May 29, 2026Updated 3 months ago
- ☆47Jul 30, 2026Updated last month