Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
☆124Aug 28, 2026Updated this week
Alternatives and similar repositories for evalstats
Users that are interested in evalstats are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashSampling: Fast and Memory-Efficient Exact Sampling (https://huggingface.co/papers/2603.15854)☆87Aug 5, 2026Updated 3 weeks ago
- ☆50Mar 22, 2026Updated 5 months ago
- Low memory full parameter finetuning of LLMs☆54Jul 18, 2025Updated last year
- ☆38May 4, 2026Updated 3 months ago
- Assessing Disparate Impacts of Personalized Interventions: Identifiability and Bounds☆11Oct 28, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Dataset: BuzzFeed News “Trending” Strip, 2018–2023☆18May 24, 2023Updated 3 years ago
- ☆12Nov 1, 2019Updated 6 years ago
- Implementation of Recursive Language Model paper from scratch☆47Feb 10, 2026Updated 6 months ago
- ☆10Jun 1, 2022Updated 4 years ago
- ☆30Nov 8, 2024Updated last year
- ☆12Feb 18, 2025Updated last year
- The official repo for DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph☆19Oct 13, 2024Updated last year
- Meta-Reinforcement Learning with Self-Reflection☆34Mar 26, 2026Updated 5 months ago
- Companion code for The Physics of LLM Inference book☆27Apr 21, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open-source code for ''Individual Fairness for Graph Neural Networks: A Ranking based Approach''.☆11Jul 8, 2022Updated 4 years ago
- Replication Resources (code+data) for "Causal Claims in Economics" by Garg P. and Fetzer T. (2026)☆32Mar 11, 2026Updated 5 months ago
- A collection of lightweight interpretability scripts to understand how LLMs think☆91Mar 18, 2026Updated 5 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- Data-Driven operations management - https://d3group.github.io/ddop☆17Jun 17, 2024Updated 2 years ago
- ☆11Mar 11, 2026Updated 5 months ago
- ☆108Feb 27, 2026Updated 6 months ago
- ☆32Mar 11, 2026Updated 5 months ago
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆75Jul 21, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Structured close reading (or rather, close watching) transcripts of _almost _ every lesson in Jeremy Howard's Practical Deep Learning for…☆35Mar 9, 2026Updated 5 months ago
- Official repository of the paper MPMQA: Multimodal Question Answering on Product Manuals (AAAI 2023)☆20Nov 28, 2022Updated 3 years ago
- Inference API for many LLMs and other useful tools for empirical research☆137May 29, 2026Updated 3 months ago
- Example of Spin app written in Java using TeaVM-WASI and wit-bindgen☆12Dec 7, 2022Updated 3 years ago
- Miscellaneous Stata Commands☆13May 23, 2023Updated 3 years ago
- AI-Powered Thesis Review Tool☆18Aug 8, 2025Updated last year
- a language model with gated conv nets (implements https://arxiv.org/pdf/1612.08083v1.pdf)☆16Jan 5, 2021Updated 5 years ago
- Compiles AI agent traces and truns them into reusable context.☆97Jul 28, 2026Updated last month
- ☆66Feb 14, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Minimal and scalable research codebase in JAX, designed for rapid iteration on frontier research in LLM and other autoregressive models.☆565Updated this week
- ☆81Feb 18, 2026Updated 6 months ago
- Today I Learned☆11Jun 2, 2017Updated 9 years ago
- build and benchmark deep research☆244Mar 28, 2026Updated 5 months ago
- Taller de xaringan creado para R-Ladies Xalapa☆12Jan 6, 2022Updated 4 years ago
- ☆50Aug 11, 2026Updated 2 weeks ago
- Fraud detection related data and scripts to share with partners.☆27Mar 5, 2023Updated 3 years ago