Skip to content

Author

Juan Francisco Mandujano Reyes

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Statistical Methods for Multiple Language Model Comparison on a Shared Evaluation

A rigorous statistical treatment of two-model comparisons on the same evals can be achieved by paired t-tests, analyzing their standard errors and a clustering correction for correlated questions. Nevertheless, leaderboards, ablations studies, and hyperparameter sweeps, usually compare $K>2$ models simultaneously. In this paper, we present a single random-effects model for scores on shared evals. We leverage models as a fixed effects and questions (or question cluster) as random effects. We show that fitting it a classical ANOVA or a linear mixed model we can recover Miller's paired and clustered estimators for K=2, but we extend the results to any K and to unbalanced, clustered designs. We validate the described model in a simulation study and using a real-data application. We take six openly available language models scored on 1,497 shared MMLU-Pro questions across 14 subject clusters, and we show that pairwise ranking claims survive depending on properly accounting for both question-level pairing and multiple comparisons. By the end of the manuscript, we provide concrete recommendations for reporting multi-model eval results.

Juan Francisco Mandujano Reyes · 0 citations