Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Characterizing LLM performance via a Bayesian lens

While large language models (LLMs) are having a transformative impact on human society, evaluating them remains challenging. Standard benchmarks usually rely on single point estimates that obscure response stochasticity and variability in question difficulty. Here, we introduce a hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets. Our approach moves beyond single accuracy metrics by modeling correct responses binomially and decomposing performance variation into intra-question stochasticity (response variability for a given question) and inter-question heterogeneity (variation in difficulty across questions) using separate priors. The framework yields a probabilistic assessment, providing full posterior distributions and credible intervals for mean accuracy, inter-question heterogeneity, and mean intra-question response variability, enabling rigorous uncertainty quantification. We demonstrate its utility by evaluating multiple LLMs across diverse benchmarks, including under semantic perturbations like question rephrasing. This analysis reveals nuanced model robustness insights and uncovers distinct behaviors across model classes (e.g., reasoning vs. non-reasoning) often missed by traditional descriptive statistics. This methodology offers a statistically grounded and powerful Bayesian lens for analyzing LLM performance, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.

Giannis Manousaridis, John G. Samuelsson, B. Emir et al. · 0 citations