Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
☆129Sep 30, 2026Updated last week
Alternatives and similar repositories for evalstats
Users that are interested in evalstats are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A projectional editor for JSON DSLs☆29Mar 30, 2024Updated 2 years ago
- ☆50Mar 22, 2026Updated 6 months ago
- Low memory full parameter finetuning of LLMs☆54Jul 18, 2025Updated last year
- ☆38May 4, 2026Updated 5 months ago
- Implementation of Recursive Language Model paper from scratch☆47Feb 10, 2026Updated 8 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆29Nov 8, 2024Updated last year
- Meta-Reinforcement Learning with Self-Reflection☆37Mar 26, 2026Updated 6 months ago
- Claude Code Skills by Sundial☆152Jul 15, 2026Updated 2 months ago
- Personal Quarto Revealjs Template☆12Jun 9, 2026Updated 4 months ago
- Automatically interact with SVG charts.☆21Sep 23, 2025Updated last year
- ☆37Apr 25, 2026Updated 5 months ago
- A collection of lightweight interpretability scripts to understand how LLMs think☆92Mar 18, 2026Updated 6 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆20Apr 18, 2026Updated 5 months ago
- Replication Resources (code+data) for "Causal Claims in Economics" by Garg P. and Fetzer T. (2026)☆37Mar 11, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆29Aug 31, 2026Updated last month
- Decennial and Intercensal Estimates of US Population by Age and Sex 1990-2019☆10Apr 12, 2020Updated 6 years ago
- Agentic Research and Evaluation Suite☆116Aug 12, 2026Updated last month
- ☆106Jul 16, 2026Updated 2 months ago
- 📊 🌐 🧑🏫 Website for graduate-level course on program evaluation and causal inference using R, built with Quarto☆10Jul 30, 2024Updated 2 years ago
- My personal website☆11May 30, 2026Updated 4 months ago
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆76Jul 21, 2026Updated 2 months ago
- Learning from Mixed Rollouts: Logit Fusion as a Bridge Between Imitation and Exploration☆18Feb 24, 2026Updated 7 months ago
- Programmatic memory for long-horizon LLM agents: the harness appends everything to one log, and the agent searches it with code. 97.4% on…☆455Aug 21, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Structured close reading (or rather, close watching) transcripts of _almost _ every lesson in Jeremy Howard's Practical Deep Learning for…☆35Mar 9, 2026Updated 7 months ago
- Official repository of the paper MPMQA: Multimodal Question Answering on Product Manuals (AAAI 2023)☆21Nov 28, 2022Updated 3 years ago
- Inference API for many LLMs and other useful tools for empirical research☆136May 29, 2026Updated 4 months ago
- Example of Spin app written in Java using TeaVM-WASI and wit-bindgen☆12Dec 7, 2022Updated 3 years ago
- Miscellaneous Stata Commands☆13May 23, 2023Updated 3 years ago
- Compiles AI agent traces and truns them into reusable context.☆97Sep 16, 2026Updated 3 weeks ago
- context-efficient terminal agent powered by an RLM☆61Feb 7, 2026Updated 8 months ago
- ☆66Feb 14, 2026Updated 7 months ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆308Apr 24, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 7 months ago
- ☆18May 30, 2025Updated last year
- A partially observable world for agents☆30Mar 17, 2026Updated 6 months ago
- Talk on writing reproducible manuscripts☆15Oct 8, 2020Updated 6 years ago
- ☆11Jun 8, 2023Updated 3 years ago
- Repository for public code and data associated with the paper "Fake News on Twitter During the 2016 U.S. Presidential Election☆12Dec 5, 2019Updated 6 years ago
- build and benchmark deep research☆248Mar 28, 2026Updated 6 months ago