Skip to content

Speculative Evaluation of Stochastic LLMs

Sep 2026 · 0 citations · 30 references
Mathematics Computer Science

TL;DR

Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits, and outperforming hindsight-tuned empirical and independent Bayesian baselines.

Abstract

Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation with a Hierarchical Bayesian Neyman (HBN) policy with pilot size and stage weight jointly chosen ex ante. It runs a short uniform pilot, pools per-task success counts with a hierarchical Bayesian model, and uses posterior expectations of task-level sampling variances for exact positive-integer Neyman allocation. To mitigate the pilot synchronization barrier, HBN-async speculatively executes continuations from partial pilot feedback and retains those selected by the final allocation. Across six checkpoints and 18 benchmark groups, we evaluate 107 nondegenerate benchmark-checkpoint profiles. For rollout budgets of 8-64 per task, Speculative Evaluation reduces variance relative to Uniform by 12.8%-33.6% on average across profiles, outperforming hindsight-tuned empirical and independent Bayesian baselines. Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits.

View source

Similar papers

#machine learning Review Sep 2026

Tail-Influence Sampling for CVaR Policy Evaluation

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es...

P. Bourigault, Xiao-Tong Ji, Matthieu Zimmer et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We i...

Priyanath Maji, S. Chowdhury · 0 citations
Open access Aug 2026

Equal Budgets Change the Verdict: Finite-Sample Bias and a Matched-Budget Re-Examination of Diversity-Enhanced flowMC Ensembles

On two real Bayesian posteriors scored against long NUTS references, uniform pooling improved on the sample-count-matched single run under marginal JS, energy distance, and MMD2 in every recorded run, while the aggregation’s deficit is confined to the biased marginal metric.

Ming-Yu Shi, Fan Zhang · 0 citations
Preprint Aug 2026

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where uncertainty remains high, and stop where est...

Toby D. Pilditch · 1 citation
#machine learning Preprint Oct 2026

Metropolis-Hastings Dominates Importance Resampling for Policy Composition

Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted pr...

A. Kurennoy, R. Yarullin, Fergal Reid · 0 citations
Preprint Aug 2026

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Standard evaluation of large language models is challenged by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels, evaluating four models on three reasoning benchmarks, and finding three findings that argue for budget-conditioned evaluation protocols.

Rodrigo Guedes de Souza, Alison R. Panisson · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.