Let $X_1,\ldots,X_n$ be independent Gaussian tensors in $\mathbb{R}^{d_1}\otimes\cdots\otimes\mathbb{R}^{d_k}$ with a common covariance matrix given by the Kronecker product of $k$ unknown positive-definite factors, and let $D=\prod_{a=1}^k d_a$ and $d_{\max}=\max_a d_a$. Franks et al. (2026) established condition-number-free guarantees for the tensor-normal maximum likelihood estimator under the sample-size condition $nD\gtrsim k^2 d_{\max}^3$ and asked whether the cubic dependence on $d_{\max}$ could be reduced to a quadratic one. We answer this question affirmatively. For $t\geq 1$, if $nD\geq C k^2 d_{\max}^2 t^2$, then with high probability the maximum likelihood estimator exists, is unique, and satisfies $d_{\rm FR}(\widehat\Theta,\Theta)\leq C t \sqrt{k} d_{\max}/\sqrt{n}$ and $d_{\rm FR}(\widehat\Theta_a,\Theta_a)\leq C t\sqrt{k d_a} d_{\max}/\sqrt{nD}$ for every mode $a$. For every mode $a$ with $d_a=d_{\max}$, we further establish the sharp Thompson-metric bound $d_{\rm op}(\widehat\Theta_a,\Theta_a)\leq C t d_{\max}/\sqrt{nD}$. These guarantees are uniform over the unknown covariance factors and require neither condition-number bounds nor sparsity assumptions. Gaussian submodel lower bounds match the full and largest-factor Fisher--Rao rates up to a factor of $\sqrt{k}$ and the largest-factor Thompson rate up to universal constants. Consequently, for fixed $k$, the quadratic dependence of the sample-size threshold on $d_{\max}$ is optimal. GPT-5.6 Sol and Claude Fable 5 were used to assist with proof development, verification, and manuscript preparation.
Let $X = [ \xi_1, \,\, \xi_2,...\,\, ,\xi_d]^\top$ be a zero-mean random vector of large dimension $d$ ($d \rightarrow \infty$) with (hidden) covariance matrix $M = (m_{ij})_{1 \leq i, j \leq d},$ where $m_{ij} = m_{ji} = \textbf{Cov}(\xi_i, \xi_j).$ Let $X_1, X_2, \dots, X_n$ be $n$ iid samples of $X$. Consider the sample covariance matrix $$\textstyle \tilde{M} := \frac{1}{n} \sum_{i=1}^{n} X_i X_i^\top.$$ In practice, one frequently uses the eigenvectors and eigenspaces of $\tilde M$ as estimators for those of $M$. A central task is to provide an error analysis for these estimators. In this paper, we provide an optimal error analysis, obtaining upper and lower bounds of matching order of magnitude, for a wide range of parameters $d$ and $n$, under mild assumptions on $M$. As corollaries, we obtain new necessary and sufficient conditions for the consistency of the estimators. In these conditions, we only require the number of samples $n$ to depend linearly on the effective rank of $M$, which can be much smaller than the dimension $d$.
Let $X=(X_1,\ldots,X_n)$ have independent coordinates with mean zero, variance one, and $\|X_i\|_{\psi_2}\le K$, and let $H_d=(\mathbb R^n)^{\otimes_2 d}$. Let $L>0$ and let $f:H_d\to\mathbb R$ be convex and $L$-Lipschitz. We prove that, for $0\le t\le c_KLn^{d/2}$, \[ \textsf{P}\left\{ \left\lvert f(X^{\otimes d})-\textsf{E}f(X^{\otimes d})\right\rvert>t \right\} \le C\exp\left[-c_K\mathcal I_{n,d}\left( \frac{t}{L n^{(d-1)/2}} \right)\right], \] where \[ \mathcal I_{n,d}(s)= \min\left\{ \frac{s^2}{d^2}, \frac{s^2}{d\log(e+nd/s^2)} \right\},\qquad s>0, \qquad \mathcal I_{n,d}(0)=0. \] The first rate is forced by changes in $\|X\|$. The second comes from changes of $X$ when its norm is nearly fixed. The proof constructs one coupling that controls both the coordinatewise conditional displacement and the mean squared Euclidean distance, and combines these bounds with a second-order estimate for $x\mapsto x^{\otimes d}$. The rate is minimax sharp, scale by scale, even when the subgaussian norms are bounded by an absolute constant. For bounded coordinates the logarithm in the second rate disappears.
Let $S=P_{cA_0}Q_{B_0}P_{cA_0}$ be the spatio-spectral concentration operator of bounded sets $cA_0,B_0\subset\mathbb{R}^d$, and let $\Lambda_\varepsilon=\#\{n:\varepsilon<\lambda_n(S)<1-\varepsilon\}$ be its plunge count. For $A_0$ and $B_0$ finite disjoint unions of bounded axis-parallel open boxes, we prove an explicit uniform upper bound on $\Lambda_\varepsilon$, valid for every $d\geq1$, $c>0$, and $0<\varepsilon<1/2$, with all constants written in terms of the side lengths. On the range $\alpha\geq4$, $c\geq2$, and $\alpha^{-c}<\varepsilon<1/2$, it gives $\Lambda_\varepsilon\leq Cc^{d-1}\log(1/\varepsilon)\log\!\bigl(\alpha c/\log(1/\varepsilon)\bigr)$. Kulikov and Dam Larsen previously proved this order on that range for a broader class; the present contribution is an independent proof and an explicit all-parameter estimate for product boxes. The proof uses a telescoping tensorization of $P_{(cA_0)^c}Q_{B_0}P_{cA_0}$ into $d$ elementary tensor operators, with one one-dimensional off-diagonal factor and $d-1$ localization factors. Schatten quasi-norms then multiply across tensor factors, and the single logarithm arises only from the normal direction. For the model cube pair, we also prove that, when $\varepsilon<4^{-d}$, $\Lambda_\varepsilon\geq M_a^d=\Omega((\log c)^d)$. Using an exact trace identity, an explicit cubic minorant, and the sine-kernel determinant asymptotics of Basor and Widom, we further obtain $\operatorname{Tr}((S-S^2)^m)=\beta_m\pi^{-2}\log c+O_m(1)$ for each fixed $m$, where $\beta_m=B(m,m)$, together with a two-sided fixed-depth window estimate of order $\log c$. The lower bound is not matching, and the fixed-order statements are not uniform in $m$.
We study online discrepancy minimization: vectors $v_1,\ldots,v_T\in\mathbb{R}^n$ arrive sequentially, and each must immediately be assigned a sign $x_t\in\{\pm1\}$, with the aim of minimizing $\|\sum_{t=1}^T x_t v_t\|_\infty$. We give a polynomial-time potential-based algorithm combining a regularization of the $\ell_\infty$-norm with restriction to an adaptively chosen coordinate set. For i.i.d. inputs with independent, symmetric, centered, unit-variance sub-Gaussian coordinates of sub-Gaussian norm at most $\sigma$, the algorithm achieves terminal discrepancy $O(\sigma^8\sqrt{n})$ with probability at least $1-\exp(-\Omega(\sigma^3\sqrt{n}))$. If the coordinates are independently masked by Bernoulli variables with mean $k/n$, where $k\gtrsim(\log n)^2$, the bound improves to $O(\sigma^8\sqrt{k})$, with failure probability $\exp(-\Omega(\sigma^3\sqrt{k}))$. Both guarantees hold for every prescribed finite horizon $T$, with no dependence on $T$. The dense result substantially generalizes a theorem of Bansal and Spencer (2020) for Rademacher inputs and gives an efficient $O(\sqrt{n})$ bound for Gaussian inputs, as conjectured by Gamarnik et al. (2022). When $T$ is polynomially larger than $n$, this is conditionally close to optimal: under worst-case hardness assumptions for standard approximate lattice problems, Vafa and Vaikuntanathan (2025) showed that no polynomial-time algorithm, even offline, can improve the $\sqrt{n}$ scale by a fixed polynomial factor in $T/n$.
The theorem below establishes the order for arbitrary real factors under the factorization contract of Arkhipov and Kalinin, who prove the matching lower order for factors with entries in $\{0,1\}$ and state the arbitrary-factor extension as open.
Let $R$ be an $m\times m$ correlation matrix satisfying $R-\mathbf{1}\mathbf{1}^{\mathsf T}/m\succeq0$, let $X\sim\mathcal{N}(0,R)$, and let $Z_1,\ldots,Z_m$ be independent standard Gaussian random variables. We prove $\max_i X_i\leq_{\mathrm{st}}\max_i Z_i$, with equality in distribution if and only if $R=I_m$. We use this comparison to resolve the Weak Simplex Conjecture: among $d+1$ equiprobable equal-energy signals in $\mathbb{R}^d$ transmitted over an additive white Gaussian noise channel, the regular simplex is the unique maximizer of the average probability of correct maximum-likelihood decoding at every signal-to-noise ratio. The same comparison proves the Simplex Mean Width Conjecture and gives the exact finite-energy performance of deterministic no-feedback AWGN codes with equiprobable messages, no restriction on the number of channel uses, and a maximal per-codeword energy constraint. The proof uses a Gaussian product inequality for log-concave functions whose first moments with respect to standard Gaussian measure vanish. A variational argument chooses one exponential tilt and one truncation endpoint in each coordinate so that this product inequality applies and a Gaussian change of measure returns all coordinates to the prescribed common threshold. A strict form of the product inequality also shows that, unless $R=I_m$, $\mathbb{P}\{X\leq c\mathbf{1}\}>\Phi(c)^m$ for every finite $c$, and hence gives the distributional equality statement. A Lean formalization is available at https://github.com/abhmul/weak-simplex-conjecture-lean.