Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused.
☆118Sep 10, 2026Updated last week
Alternatives and similar repositories for every_eval_ever
Users that are interested in every_eval_ever are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆28Jan 7, 2026Updated 8 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated last year
- An implementation of "Subspace Representations for Soft Set Operations and Sentence Similarities" (NAACL 2024)☆10May 31, 2024Updated 2 years ago
- Tool for exporting Apple Neural Engine-accelerated versions of transformers models on HuggingFace Hub.☆16Jun 25, 2026Updated 2 months ago
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Source code of NAACL 2025 Findings "Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models"☆16Dec 16, 2025Updated 9 months ago
- ☆25Jul 8, 2026Updated 2 months ago
- ☆33Nov 11, 2025Updated 10 months ago
- ICLR 2025 Workshop & CHI 2025 SIG: "Bidirectional Human-AI Alignment"☆59Aug 6, 2024Updated 2 years ago
- utilities for batched llm calls with retries☆52Updated this week
- A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp☆32Updated this week
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- A package dedicated for running benchmark agreement testing☆19Sep 18, 2025Updated last year
- Official PyTorch Implementation for the "Unsupervised Model Tree Heritage Recovery" paper (ICLR 2025).☆62Jul 1, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆275Updated this week
- ☆16Jul 7, 2026Updated 2 months ago
- Dataset for Unified Editing, EMNLP 2023. This is a model editing dataset where edits are natural language phrases.☆24Sep 4, 2024Updated 2 years ago
- Reproducible, flexible LLM evaluations☆395Mar 24, 2026Updated 5 months ago
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated 2 years ago
- A Streamlit app to add structured tags to a dataset card☆23Jun 30, 2022Updated 4 years ago
- Run-time validation of tensors for machine-learning systems.☆11Apr 8, 2021Updated 5 years ago
- An implementation of GrASP (Shnarch et. al., 2017)☆24Aug 29, 2022Updated 4 years ago
- harvard-cs-2881-classroom-hw0-c2881-hw0 created by GitHub Classroom☆18Jul 26, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆34Jan 23, 2025Updated last year
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆21Jan 19, 2025Updated last year
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆338Feb 8, 2026Updated 7 months ago
- ☆38Jul 8, 2025Updated last year
- Demo setups for ai-atlas-nexus☆16Jul 15, 2026Updated 2 months ago
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- Notes on Direct Preference Optimization☆28Apr 14, 2024Updated 2 years ago
- An ecosystem for model-driven IoT applications☆15Oct 23, 2025Updated 10 months ago
- [COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model☆26Nov 25, 2025Updated 9 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Inspect: A framework for large language model evaluations☆2,808Updated this week
- Delta Attention Residuals - supplementary code and pretrained models☆43May 20, 2026Updated 3 months ago
- A Python toolkit for analyzing machine learning models and datasets.☆78Jul 9, 2026Updated 2 months ago
- UML model and code examples of design patterns for TypeScript. The model is created with Astah.☆15Aug 6, 2025Updated last year
- Implementation of the BatchTopK activation function for training sparse autoencoders (SAEs)☆67Jul 24, 2025Updated last year
- Codebase for "On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback". This repo implements a generative multi-tur…☆26Dec 3, 2024Updated last year
- Generate synthetic labeled data for extremely low-resource languages using bilingual lexicons.☆20Oct 3, 2024Updated last year