Skip to content
Preprint

Statistical Methods for Multiple Language Model Comparison on a Shared Evaluation

Aug 2026 · 0 citations · 8 references
Mathematics

Abstract

A rigorous statistical treatment of two-model comparisons on the same evals can be achieved by paired t-tests, analyzing their standard errors and a clustering correction for correlated questions. Nevertheless, leaderboards, ablations studies, and hyperparameter sweeps, usually compare $K>2$ models simultaneously. In this paper, we present a single random-effects model for scores on shared evals. We leverage models as a fixed effects and questions (or question cluster) as random effects. We show that fitting it a classical ANOVA or a linear mixed model we can recover Miller's paired and clustered estimators for K=2, but we extend the results to any K and to unbalanced, clustered designs. We validate the described model in a simulation study and using a real-data application. We take six openly available language models scored on 1,497 shared MMLU-Pro questions across 14 subject clusters, and we show that pairwise ranking claims survive depending on properly accounting for both question-level pairing and multiple comparisons. By the end of the manuscript, we provide concrete recommendations for reporting multi-model eval results.

View source