Skip to content
Preprint

Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and Corrections

Aug 2026 · 0 citations · 11 references
Computer Science Mathematics

TL;DR

It is shown that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$.

Abstract

Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with probability one. For fixed discount and regularization parameters, we further show that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order $R\sqrt{\log(T/\delta)}$ at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$. We identify the proof error: different terminal times use different Gaussian mixing distributions, so the fixed-time mixtures do not form one supermartingale, and the stopping-time argument does not repair this failure. Finally, we show that the weighted inequality remains valid at each fixed deterministic time, give valid finite- and infinite-horizon corrections, and discuss consequences for downstream analyses.

View source

Similar papers

Preprint Aug 2026

A Finite-Sample Analysis of Quantile Temporal-Difference Learning

A global high-probability last-iterate guarantee for synchronous tabular QTD under general positive, nonincreasing step-size sequences and arbitrary initialization in the natural parameter range is established.

Zijie Cheng, Xiang Li, Yang Peng et al. · 0 citations
Preprint Sep 2026

Well-Posedness for SDEs with Logarithmical Critical Distributional Drifts

We study the stochastic differential equation $$d X_t=b(t,X_t)d t+\sqrt{2}d W_t$$ on $\mathbb R^d$, where $b$ is a time-dependent, divergence-free distributional drift of critical H\"older--Besov regularity $-1$, strengthened by an iterated-logarithmic correction. For every initial probability law, we construct a weak...

Zi-Kai Chen, Zi-Mo Hao, Xicheng Zhang · 0 citations
Preprint Sep 2026

Exact finite-sample inference for multi-mixed fractional Brownian motion with drift

In this paper we study a linear drift perturbed by a superposition of $m$ independent fractional Brownian motions with known Hurst parameters and a common scale, observed at $N$ equidistant times. Inference for such models is usually asymptotic; we show that here it is exact. We derive the maximum likelihood estimators...

Afrah Al-Harby, E. Mliki · 0 citations
#machine learning Preprint Sep 2026

Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling

This work introduces a variant of the Thompson Sampling algorithm that uses a fractional or $\alpha$-posterior instead of the standard posterior, and identifies general regularity conditions on the prior and reward distributions that enable a regret analysis of $\alpha$-TS without assuming any tractable approximation o...

Prateek Jaiswal, D. Pati, A. Bhattacharya et al. · 0 citations
#machine learning Preprint Sep 2026

The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives

We study the time-uniform convergence of the raw iterate of standard stochastic gradient descent (SGD) for unconstrained smooth convex objectives. We prove that, under standard noise assumptions, the time-uniform convergence rate gets arbitrarily close to $\sqrt{\log n / n}$ but never reaches it. More specifically, we...

Rui-Jie Li, Kang Chen, Tian-Yu Wang · 0 citations
Preprint Aug 2026

A Central Limit Theorem for Regularized M-Estimators

We prove a quantitative central limit theorem for linear functionals of regularized empirical-risk minimizers in the proportional-dimensional regime \(p=O(n)\). The data columns are independent, not necessarily identically distributed, and satisfy a uniform columnwise Poincar\'e inequality. Under uniform curvature and...

Cosme Louart · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.