Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused.
☆96Jul 29, 2026Updated this week
Alternatives and similar repositories for every_eval_ever
Users that are interested in every_eval_ever are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20Jan 7, 2026Updated 6 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- An implementation of "Subspace Representations for Soft Set Operations and Sentence Similarities" (NAACL 2024)☆10May 31, 2024Updated 2 years ago
- James' cookbook of evaluations and finetuning experiments☆32Feb 19, 2026Updated 5 months ago
- Tool for exporting Apple Neural Engine-accelerated versions of transformers models on HuggingFace Hub.☆16Jun 25, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Collection of evals for Inspect AI☆607Updated this week
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- Source code of NAACL 2025 Findings "Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models"☆16Dec 16, 2025Updated 7 months ago
- ERRor ANnotation Toolkit: Automatically extract and classify grammatical errors in parallel original and corrected sentences.☆12Mar 23, 2023Updated 3 years ago
- ☆31Nov 11, 2025Updated 8 months ago
- Code repo for the model organisms and convergent directions of EM papers.☆72Sep 22, 2025Updated 10 months ago
- Code for "Transformer-Based Deep Survival Analysis"☆13May 27, 2022Updated 4 years ago
- ☆16Jul 20, 2026Updated last week
- ICLR 2025 Workshop & CHI 2025 SIG: "Bidirectional Human-AI Alignment"☆59Aug 6, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- utilities for batched llm calls with retries☆51Updated this week
- A python library built on top of UKAISI Inspect to support Control research, by Redwood☆30Updated this week
- Code for paper "Robustness of Bayesian Neural Networks to Gradient-Based Attacks"☆17Feb 26, 2024Updated 2 years ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆35May 31, 2026Updated last month
- The open-source AISI toolkit for sandboxing agentic evaluations☆26Aug 7, 2025Updated 11 months ago
- codebase release for EMNLP2023 paper publication☆19Sep 18, 2025Updated 10 months ago
- A package dedicated for running benchmark agreement testing☆19Sep 18, 2025Updated 10 months ago
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆269Updated this week
- ☆16Jul 7, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Dataset for Unified Editing, EMNLP 2023. This is a model editing dataset where edits are natural language phrases.☆24Sep 4, 2024Updated last year
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated last year
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.☆118Updated this week
- A Streamlit app to add structured tags to a dataset card☆23Jun 30, 2022Updated 4 years ago
- An implementation of GrASP (Shnarch et. al., 2017)☆24Aug 29, 2022Updated 3 years ago
- harvard-cs-2881-classroom-hw0-c2881-hw0 created by GitHub Classroom☆16Jul 26, 2025Updated last year
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆33Jan 23, 2025Updated last year
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆20Jan 19, 2025Updated last year
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆333Feb 8, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆38Jul 8, 2025Updated last year
- ☆18Aug 22, 2022Updated 3 years ago
- TypeScript port of ACE framework, written entirely by Claude Code running in a loop☆22Dec 4, 2025Updated 7 months ago
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆18Jul 4, 2025Updated last year
- a tool like just and entr. A minimal cross-platform, config-driven tool for running commands whenever files change☆15Jun 4, 2026Updated last month
- Inspect: A framework for large language model evaluations☆2,426Updated this week