It is shown that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$.
Abstract
Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with probability one. For fixed discount and regularization parameters, we further show that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order $R\sqrt{\log(T/\delta)}$ at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$. We identify the proof error: different terminal times use different Gaussian mixing distributions, so the fixed-time mixtures do not form one supermartingale, and the stopping-time argument does not repair this failure. Finally, we show that the weighted inequality remains valid at each fixed deterministic time, give valid finite- and infinite-horizon corrections, and discuss consequences for downstream analyses.
A global high-probability last-iterate guarantee for synchronous tabular QTD under general positive, nonincreasing step-size sequences and arbitrary initialization in the natural parameter range is established.
Zijie Cheng, Xiang Li, Yang Peng et al.· 0 citations
We study the stochastic differential equation $$d X_t=b(t,X_t)d t+\sqrt{2}d W_t$$ on $\mathbb R^d$, where $b$ is a time-dependent, divergence-free distributional drift of critical H\"older--Besov regularity $-1$, strengthened by an iterated-logarithmic correction. For every initial probability law, we construct a weak...
In this paper we study a linear drift perturbed by a superposition of $m$ independent fractional Brownian motions with known Hurst parameters and a common scale, observed at $N$ equidistant times. Inference for such models is usually asymptotic; we show that here it is exact. We derive the maximum likelihood estimators...
This work introduces a variant of the Thompson Sampling algorithm that uses a fractional or $\alpha$-posterior instead of the standard posterior, and identifies general regularity conditions on the prior and reward distributions that enable a regret analysis of $\alpha$-TS without assuming any tractable approximation o...
Prateek Jaiswal, D. Pati, A. Bhattacharya et al.· 0 citations
We study the time-uniform convergence of the raw iterate of standard stochastic gradient descent (SGD) for unconstrained smooth convex objectives. We prove that, under standard noise assumptions, the time-uniform convergence rate gets arbitrarily close to $\sqrt{\log n / n}$ but never reaches it. More specifically, we...
We prove a quantitative central limit theorem for linear functionals of regularized empirical-risk minimizers in the proportional-dimensional regime \(p=O(n)\). The data columns are independent, not necessarily identically distributed, and satisfy a uniform columnwise Poincar\'e inequality. Under uniform curvature and...
Cosme Louart· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.