Skip to content

How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL

Sep 2026 · 0 citations · 69 references
Computer Science Mathematics

TL;DR

It is proved that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM, and isolates two missing links between internal update gaps and predictive cost.

Abstract

Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at least linearly in the natural confidence scale $L_K(q)$, while their categorical $D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$ vanishes at the same explicit witness. Along stationary HMM trajectories, the expected terminal KL between filtered posteriors also converges to zero at $H(q)=\lceil-\log(q)/c\rceil+1$. Typical blocks without switches drive both filters into a common confidence cone, where softmax curvature suppresses their disagreement; a single Gaussian maximal event controls adaptive noise. A sweep with equally spaced Gaussians over $K\in\{2,4,8\}$ illustrates the opposing trends, and binary controls at long horizons compare saturating and nonsaturating recurrences. The result isolates two missing links between internal update gaps and predictive cost: the contribution of separating states to expected loss and decoder sensitivity. Thus even an unbounded internal update gap does not by itself certify predictive failure. The construction is fixed in $K$ and does not provide a universal criterion for when compression is harmless or characterize when internal gaps must incur task loss.

View source

Similar papers

Preprint Aug 2026

Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration

A single, horizon-free algorithm that satisfies the optimal regret rate for every bounded proper loss and also adapt to every smooth proper loss, covering nondifferentiable losses and changes of the active simplex face.

Pahan Dewasurendra · 0 citations
#machine learning Preprint Sep 2026

The cost of useful natural gradient updates

What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family who...

Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome · 0 citations
Preprint Sep 2026

Prediction with Five Experts and Geometric Stopping: A Probabilistic Construction and Analytic Verification

It is shown that the alternating COMB control attains the limiting Hamiltonian maximum exactly when the two highest scores coincide and the third- and fourth-highest scores coincide, and that the alternating COMB control attains the limiting Hamiltonian maximum exactly when the two highest scores coincide.

Erhan Bayraktar, Ibrahim Ekren, N. Kolliopoulos · 2 citations
#machine learning Preprint Sep 2026

Sparse Priors for Efficient Distribution Learning

The results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on $n$ and introduces the class of sparse priors and defines the "Sparse Dimension" as a measure of sparsity of a prior over the space of all distributions.

Saumya Goyal, B. Póczos · 0 citations
Preprint Sep 2026

On Prior-to-Posterior Stability in the Wasserstein Metric for Bayesian Inverse Problems

Priors in Bayesian inverse problems are often approximated through discretization, hyperparameter estimation, or generative modeling. Understanding how prior approximation errors propagate to the posterior and subsequent predictions is therefore important. In this work, we study the stability of the prior-to-posterior...

Liang-Hao Cao · 0 citations
Preprint Aug 2026

Q-Learning with Stable Infinite-Dimensional Linear Function Approximation

Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture need not preserve the Bellman contraction. We develop a stable infinite-dimensional linear function approximation framework for Q-learning from a single Markovian behavior-policy trajectory. The learning variab...

Shengbo Wang · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.