Public repository containing METR's DVC pipeline for eval data analysis
☆311Mar 6, 2026Updated 5 months ago
Alternatives and similar repositories for eval-analysis-public
Users that are interested in eval-analysis-public are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Vivaria is METR's tool for running evaluations and conducting agent elicitation research.☆141May 18, 2026Updated 2 months ago
- ☆22Jul 28, 2026Updated 2 weeks ago
- ☆153Oct 16, 2025Updated 9 months ago
- ☆18Dec 10, 2025Updated 8 months ago
- ☆129Jul 28, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Running UK AISI's Inspect in the Cloud☆25May 6, 2026Updated 3 months ago
- ☆17Mar 10, 2026Updated 5 months ago
- ☆34Jun 4, 2025Updated last year
- Data visualization for Inspect AI large language model evalutions.☆21Jul 15, 2026Updated 3 weeks ago
- NeurIPS 2024: SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation☆13May 24, 2025Updated last year
- An Inspect extension for agentic cyber evaluations☆38Jun 18, 2026Updated last month
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆17Feb 23, 2026Updated 5 months ago
- ☆14Updated this week
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.☆223Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- METR Task Standard☆188Feb 3, 2025Updated last year
- ☆28Apr 1, 2026Updated 4 months ago
- Inspect: A framework for large language model evaluations☆2,547Updated this week
- Collection of evals for Inspect AI☆625Updated this week
- An alignment auditing agent capable of quickly exploring alignment hypothesis☆1,286Updated this week
- ☆74Jun 16, 2026Updated last month
- A repository that holds templates, examples, and tests to help external parties submit tasks to AISI that conform with the Autonomous Sys…☆11Jan 16, 2026Updated 6 months ago
- ☆53Jul 27, 2026Updated 2 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆39Updated this week
- Public repository for the Remote Labor Index (RLI)☆75Nov 3, 2025Updated 9 months ago
- ☆88Nov 21, 2025Updated 8 months ago
- Generate Python docstrings automatically with LLM and syntax trees☆20Jun 13, 2025Updated last year
- The open-source AISI toolkit for sandboxing agentic evaluations☆27Aug 7, 2025Updated last year
- ☆25Apr 23, 2024Updated 2 years ago
- ☆11Jun 2, 2021Updated 5 years ago
- Programmatic memory for long-horizon LLM agents: the harness appends everything to one log, and the agent searches it with code. 97.4% on…☆355Updated this week
- ☆28Jun 1, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Benchmarking Dark Patterns in LLMs (ICLR 2025)☆18Mar 29, 2025Updated last year
- ☆311Jul 1, 2026Updated last month
- A toolkit for describing model features and intervening on those features to steer behavior.☆253Mar 16, 2026Updated 4 months ago
- ☆16Jun 17, 2025Updated last year
- Optimally-weighted herding is Bayesian Quadrature☆18Jul 8, 2016Updated 10 years ago
- Cited 83-model x 49-benchmark LLM evaluation matrix with 18 matrix completion methods☆40Feb 25, 2026Updated 5 months ago
- Inference API for many LLMs and other useful tools for empirical research☆136May 29, 2026Updated 2 months ago