Preprint
Jul 2026
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
Evalci, a pure-Python library that turns a per-item results table into a publication-ready claim, and re-analyzes a public comparison of nine language models'MMLU accuracy to find that 3 of the 8 adjacent leaderboard-rank gaps are not statistically significant after correcting for the 36 pairwise comparisons the ranking implies.
Shreyas Chandrahas
· 0 citations