Skip to content

Next-token functional estimation

Sep 2026 · 0 citations · 45 references
Mathematics Computer Science

TL;DR

A leave-a-window-out estimator is proposed, which deletes a window of length $\tau$ after each index before forming the empirical measure and reduces to leave-one-out at $\tau = 1$.

Abstract

Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length $\tau$ after each index before forming the empirical measure and reduces to leave-one-out at $\tau = 1$. Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary $\beta$-mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.

View source

Similar papers

Preprint Oct 2026

Mean squared error reduction of (plug-in) look-ahead estimators for Markov chains

We consider the estimation of the invariant distribution $\pi$ of a Markov kernel on a discrete state space (or of the mean value $\pi(f)$ of a functional $f$ under $\pi$) when solving the invariance equation is not feasible. Look-ahead estimators exploit the knowledge of the kernel by propagating the empirical occupat...

Romain Azais, Benoît Henry, Henri Péchoux · 0 citations
#machine learning Preprint Sep 2026

Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation

For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this f...

Levin David Schwab · 0 citations
Preprint Aug 2026

Estimation of distribution functions, their jumps and interval probabilities under measurement error

We consider the classical additive measurement-error model $X=Y+Z$, where the latent random variable $Y$ has unknown distribution $F_Y$ and the error $Z$ has a known distribution. We develop direct estimators for three functionals of $F_Y$: (i) $F_Y(x)$ at continuity points; (ii) interval probabilities $F_Y(y)-F_Y(x)$...

K. Mynbaev, Carlos Martins-Filho, Chad Brown · 0 citations
Preprint Aug 2026

Gaussian-efficient testing by betting on the mean of bounded data

Given $[0,1]$-valued random variables $X_1,\dots,X_n$ such that $\mathbb{E}[X_i | X_1,\dots,X_{i-1}]= \mu$ for all $i$, we propose a new nonasymptotic confidence interval for $\mu$ that is obtained by inverting terminal e-values generated by a novel betting strategy. When the data are iid, its limiting width matches th...

Diego Martinez-Taboada, Aaditya Ramdas · 1 citation
Preprint Sep 2026

Empirical Bayes for compound adaptive experiments

We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal lik...

Karun Adusumilli, Jia-Ying Gu, Jun-Fan Tao · 0 citations
#machine learning Preprint Sep 2026

Extreme classification: beating chance with one training example from each class

We study a minimal classification problem: Given independent labeled observations $X\sim P$ and $Z\sim Q$ from two unknown distributions $P,Q$, and given an independent target $Y$ drawn with equal probability from $P$ or $Q$, can one classify $Y$ strictly better than chance whenever $P\neq Q$? The one-nearest-neighbor...

K. Bleakley, Aaditya Ramdas · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.