Skip to content
Preprint

Stochastic Autoregressive Learning

Aug 2026 · 1 citation
Computer Science

TL;DR

It is shown that stochastic autoregressive learning fundamentally differs from the deterministic theory, and that CoT learning at scale $\varepsilon$ is upper-bounded by base learning at scale $\varepsilon/M^2$, whereas e2e learning at scale $\varepsilon$ is upper-bounded, up to logarithmic factors, by $(M/\varepsilon) m_{CoT}(\Theta(\v

Abstract

Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning. This generalizes the deterministic autoregressive learning framework of Joshi et al., COLT 2025. In our model, one fixed generator assigns a Bernoulli next-token distribution to every prompt string. Starting from an input prompt, a token is sampled and appended to the prompt; the same generator is then applied again to this expanded prompt; this procedure is repeated for $M$ steps. Three forms of supervision are considered: base one-step samples, chain-of-thought (CoT) samples that reveal full random trajectories of length $M$, and end-to-end (e2e) samples that reveal only the final token of length $M$ trajectories. For a generator class, we study the minimum number of samples $m_{base}(\varepsilon),m_{CoT}(\varepsilon), m_{e2e}(\varepsilon)$, resp., required to learn the one-step probabilities in the base model, and the final-token probability in the CoT and e2e models, under squared loss error~$\varepsilon$. We show that stochastic autoregressive learning fundamentally differs from the deterministic theory. At scale $\varepsilon$, there is no universal comparison between the three learning tasks: both $m_{CoT}/m_{base}$ and $m_{e2e}/m_{CoT}$ can be made simultaneously arbitrarily larger than $M/\varepsilon$, the natural analogue for the existing deterministic results. Nevertheless, after altering scales, for every class, CoT learning at scale $\varepsilon$ is upper-bounded by base learning at scale $\varepsilon/M^2$, whereas e2e learning at scale $\varepsilon$ is upper-bounded, up to logarithmic factors, by $(M/\varepsilon) m_{CoT}(\Theta(\varepsilon))$. These dependencies and scales are essentially tight. We complement these bounds by studying dimension $d$ logistic functions in our model.

View source

Similar papers

#machine learning Preprint Sep 2026

Next-token functional estimation

A leave-a-window-out estimator is proposed, which deletes a window of length $\tau$ after each index before forming the empirical measure and reduces to leave-one-out at $\tau = 1$.

Milind Nakul, Vidya K. Muthukumar, A. Pananjady · 0 citations
#machine learning Preprint Aug 2026

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

A tractable teacher--student model where a stable latent linear RNN generates trajectories and a sketched linear recurrent student is trained by safeguarded full-batch WSD gradient descent on next-token prediction is studied.

Zi-Yan Chen, Zhong-Zhu Zhou, Pei-Lin Liu et al. · 0 citations
#machine learning Preprint Oct 2026

Learning to Learn a Language

We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequen...

Lennart Carstens-Behrens, Holger Fröhlich · 0 citations
Preprint Sep 2026

Empirical Bayes for compound adaptive experiments

We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal lik...

Karun Adusumilli, Jia-Ying Gu, Jun-Fan Tao · 0 citations
#machine learning Preprint Sep 2026

Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation

For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this f...

Levin David Schwab · 0 citations
#machine learning Preprint Oct 2026

LS-AR: Future-Predictive Latent Steering in Autoregressive LLMs

Standard autoregressive (AR) models process high-level task instructions, state history, and transient tokens within a single shared sequence of tokens. Consequently, they lack the architectural mechanisms needed to isolate macro-objectives from context noise. To overcome this single-channel limitation, we introduce La...

Anubha Gupta, Eduardo Pignatelli · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.