Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused.
☆107Aug 27, 2026Updated this week
Alternatives and similar repositories for every_eval_ever
Users that are interested in every_eval_ever are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Deterministic-mode checks for LLM inference: measure run/batch variance, generate repro packs, and explain why outputs differ.☆20Aug 20, 2026Updated last week
- RowLang is a minimalistic esoteric programming language written as an analogy to rowing.☆20Jul 6, 2024Updated 2 years ago
- This codebase accompanies the paper "Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews""☆24May 7, 2026Updated 3 months ago
- ☆26Jan 7, 2026Updated 7 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 11 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- James' cookbook of evaluations and finetuning experiments☆35Feb 19, 2026Updated 6 months ago
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- ☆17Apr 25, 2026Updated 4 months ago
- Source code of NAACL 2025 Findings "Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models"☆16Dec 16, 2025Updated 8 months ago
- Transform JSON output from Senzing SDK for use with graph technologies, semantics, and downstream LLM integration☆20Aug 11, 2026Updated 2 weeks ago
- ☆40May 7, 2026Updated 3 months ago
- ☆25Jul 8, 2026Updated last month
- Starter code for semester project in Cloud Computing Architecture course at ETH Zurich☆13Mar 23, 2026Updated 5 months ago
- Code repo for the model organisms and convergent directions of EM papers.☆81Sep 22, 2025Updated 11 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for "Transformer-Based Deep Survival Analysis"☆13May 27, 2022Updated 4 years ago
- An audio-friendly ls, with a little something extra.☆11Feb 20, 2026Updated 6 months ago
- utilities for batched llm calls with retries☆51Jul 25, 2026Updated last month
- A python library built on top of UKAISI Inspect to support Control research, by Redwood☆30Updated this week
- Code for paper "Robustness of Bayesian Neural Networks to Gradient-Based Attacks"☆17Feb 26, 2024Updated 2 years ago
- codebase release for EMNLP2023 paper publication☆19Sep 18, 2025Updated 11 months ago
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆275Updated this week
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated last year
- Run-time validation of tensors for machine-learning systems.☆11Apr 8, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆34Jan 23, 2025Updated last year
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆21Jan 19, 2025Updated last year
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆338Feb 8, 2026Updated 6 months ago
- ☆14Jul 15, 2022Updated 4 years ago
- Quality Controlled Paraphrase Generation (ACL 2022)☆74Sep 17, 2025Updated 11 months ago
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- Official code for ICML 2024 paper on Persona In-Context Learning (PICLe)☆28Jun 27, 2024Updated 2 years ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- An official PyTorch implementation for CLIPPR☆30Jul 22, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Inspect: A framework for large language model evaluations☆2,660Updated this week
- Lab files of IBM's Qiskit Global Summer School 2020.☆18Sep 3, 2020Updated 5 years ago
- Delta Attention Residuals - supplementary code and pretrained models☆43May 20, 2026Updated 3 months ago
- A Python toolkit for analyzing machine learning models and datasets.☆79Jul 9, 2026Updated last month
- $100K or 100 Days: Trade-offs when Pre-Training with Academic Resources☆153Oct 2, 2025Updated 10 months ago
- UML model and code examples of design patterns for TypeScript. The model is created with Astah.☆15Aug 6, 2025Updated last year
- Implementation of the BatchTopK activation function for training sparse autoencoders (SAEs)☆67Jul 24, 2025Updated last year