ianarawjo / evalstatsView on GitHub
Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
124Aug 28, 2026Updated this week

Alternatives and similar repositories for evalstats

Users that are interested in evalstats are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?