Skip to content
Preprint

Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble

Aug 2026 · 1 citation · 32 references
Mathematics Economics

Abstract

The set of $n\times n$ correlation matrices, known as the elliptope, has volume decaying at the super-exponential rate $\exp\{-\tfrac14 n^2\log n\}$. We characterize where this vanishing volume concentrates. A uniform draw is entrywise close to the identity yet globally far from it and nearly singular: its maximum absolute correlation is of order $\sqrt{\log n/n}$, its Frobenius distance is asymptotic to $\sqrt n$, its empirical spectral distribution converges to the Marchenko-Pastur law with ratio one, and its smallest eigenvalue has the exact $\operatorname{Beta}(1,d)$ distribution, where $d=n(n-1)/2$, and is therefore of order $n^{-2}$. More generally, distinct off-diagonal entries are exactly pairwise independent under every $\operatorname{LKJ}(\eta)$ law. For the uniform law, this yields a Chen-Stein proof of the extreme-correlation point-process limit and an $O(n^{-1})$ total-variation bound for finite-dimensional exceedance counts relative to Poisson laws with their exact finite-$n$ means. We also identify two distinct scales: $\eta_n\asymp n$ alters the limiting spectrum, whereas $\eta_n\asymp n^2$ is needed to keep the Frobenius distance bounded. Finally, for a bounded, centered i.i.d. off-diagonal specification, projection to the nearest correlation matrix incurs a squared repair cost asymptotically at least one-half of the squared Frobenius norm of its off-diagonal part.

View source

Similar papers

Preprint Aug 2026

Non-Gaussian fluctuations for traces of squared sample correlation matrices in high dimensions

We provide limit theory for the trace of the squared sample correlation matrix $\mathbf R$, constructed from $n$ observations of a $p$-dimensional random vector with iid components. If the entries have finite fourth moment and $p$ and $n$ grow proportionally, it is known that $\operatorname{tr}({\mathbf R}^2)$ satisfies a central limit theorem (CLT) and the centering and scaling sequences are universal in the sense that they do not depend on the entry distribution. Under symmetry and regular variation assumption with index $\alpha$ and any growth rate of the dimension, we prove that the universal CLT remains valid for $\alpha>3$. For $\alpha<3$, we identify a critical dimension growth at which the fluctuations of $\operatorname{tr}({\mathbf R}^2)$ become non-Gaussian. Moreover, if the dimension $p$ grows faster and $\alpha\le 3$ we establish a non-universal CLT with norming sequences depending on the value of $\alpha$. Our findings are illustrated in a simulation study.

J. Heiny, Xuechun Hu, Felix J. Seo · 0 citations
Preprint Aug 2026

Sharp Berry-Esseen Bounds for the Log Determinant of a Gaussian Sample Correlation Matrix

Let $\widehat R$ be the Pearson sample correlation matrix formed from $n$ independent Gaussian observations in $p$ dimensions, and write $m=n-1\ge p$. Under the null correlation $R=I_p$, the classical independent beta product, exact cumulants, and full Fourier inversion yield, along every sequence $p\to\infty$ with $m\ge p$, a uniform first Edgeworth expansion for $\log\det\widehat R$, centered by its exact mean and scaled by its exact standard deviation. The expansion identifies the exact finite dimensional skewness correction and gives the sharp Kolmogorov equivalent $A_{m,p}/\{6\sqrt{2\pi}V_{m,p}^{3/2}\}$, where $V_{m,p}$ is the exact variance and $A_{m,p}$ is the absolute third cumulant. This equivalent unifies the square, fixed gap, growing gap, proportional, and dilute regimes; in the square regime the error has order $(\log p)^{-3/2}$ with an exact constant. For every positive definite population correlation matrix $R$, we prove a uniform finite sample Berry-Esseen bound that explicitly tracks population dependence. All theoretical results have exact or proved equivalent Lean 4 formulations whose declarations and dependencies are kernel checked.

Hongru Zhao · 0 citations
Preprint Aug 2026

Local Laws and Edge Universality for Noncentral Sample Covariance Matrices

We consider the real noncentral sample covariance matrices $\mathcal{W}=YY^\top$ with $Y=A+\Sigma^{1/2}X$. Here $A\in\mathbb{R}^{M\times N}$ is deterministic, $\Sigma$ is a deterministic positive definite population covariance matrix and $X\in\mathbb{R}^{M\times N}$ has independent centered entries with variance $N^{-1}$. We prove local laws near regular right edges down to optimal spectral scales without requiring the commutativity of $AA^\top$ and $\Sigma$. As a consequence, we obtain optimal eigenvalue rigidity at the rightmost regular edge and delocalization of the corresponding left and right singular vectors. We also show that, with high probability, there are no eigenvalues in the adjacent spectral gap beyond the optimal $N^{-2/3}$ edge scale, up to an arbitrarily small $N^\varepsilon$ loss. Finally, we establish edge universality at the rightmost regular edge: after centering and scaling, the largest eigenvalue converges to the Tracy--Widom distribution. The main technical ingredient is a stability analysis of the matrix Dyson equation (MDE) associated with the linearization of $Y$, whose self-energy operator does not satisfy the flatness condition of the general MDE theory. Exploiting the special block structure, we reduce the stability analysis exactly to a two-dimensional operator. This reduction yields regularity of the spectral density and square-root behavior at regular right edges, together with sharp stability bounds near such edges.

Can Hu, Jiang Hu, Zhidong Bai · 0 citations
Preprint Jul 2026

Spectra of high-dimensional Spearman correlation matrices under scale-mixture dependence

We study the asymptotic spectral properties of high-dimensional Spearman correlation matrices for scale-mixture data. We consider observations of the form $x_t=\sigma_t \xi_t \in \mathbb{R}^N,$ where the coordinates of $\xi_t$ are i.i.d.\ and the scalar mixture variable $\sigma_t$ is shared by all coordinates. Under natural symmetry assumptions, the coordinates of $x_t$ are pairwise uncorrelated in both the Pearson and Spearman sense. Nevertheless, they are not independent when the mixture variable is non-degenerate. We show that this higher-order dependence survives the rank transformation and leaves a nontrivial spectral signature. In the proportional regime $N/T\to q\in(0,\infty),$ the empirical spectral distribution of the Spearman correlation matrix converges almost surely to a generalized Mar\v{c}enko--Pastur law governed by the limiting distribution of an effective rank variance. We also formulate a broader latent-variable extension, which covers, in particular, some scale-mixture models with correlated directional components. We discuss solvable examples and numerical approximations, motivated in part by heavy-tailed data in robust multivariate statistics, econometrics, and finance.

J. Bouchaud, Pierre Bousseyroux, Tomas Espana et al. · 0 citations
Preprint Aug 2026

Critical tensor covariance at the Marchenko--Pastur threshold

Let $X$ be a centered, variance-one random variable with finite fourth moment, and form the principal degree-$d$ tensor feature vector of all square-free monomials in $n$ independent copies of $X$. For $m$ independent samples we determine the global spectral law of the sample covariance throughout the critical scale $d^2/n\to\lambda\in[0,\infty)$, with aspect ratio $p/m\to c$. For a fixed base distribution with finite fourth moment and $P(|X|=1)<1$, prior work gives ordinary Marchenko--Pastur convergence if and only if $d=o(\sqrt n)$. We identify the finite critical boundary: when $d^2/n\to\lambda\in(0,\infty)$, the tensor radius converges in quadratic Wasserstein distance to a lognormal law determined by the fourth moment, while all remaining bounded quadratic fluctuations vanish. A leave-one-out resolvent argument then yields almost-sure convergence of the empirical spectral distribution to a free compound-Poisson law driven by this endogenous lognormal jump. The limit reduces to Marchenko--Pastur when the fourth-moment excess or the overlap intensity vanishes. In the unit-modulus case, our estimates recover the sharp range $\min(d,n-d)=o(n)$ for uniform quadratic-form concentration and imply Marchenko--Pastur convergence throughout that range, with an explicit variance bound.

Xiaohui Xie · 0 citations
Preprint Aug 2026

An Exponential Lower Bound for the Permanent of Random Bernoulli Matrix

Let $M_n$ be an $n\times n$ matrix with independent uniform sign entries. We prove that there exist absolute constants $C,c>0$ such that, for all sufficiently large $n$, \[ \mathbb{P}\!\left( \left|\operatorname{Per}(M_n)\right| \ge e^{-Cn}\sqrt{n!} \right) \ge 1-n^{-c}. \] Our proof tracks the total squared permanent of minors under successive row exposure. Up to $k=\lfloor n/2\rfloor$, the total squared permanent grows deterministically via the Boolean lattice up-operator; for larger $k$, the row exposure increments are governed by positive semidefinite Rademacher quadratic forms. Therefore, we confirms the exponential scale lower bound suggested by Tao and Vu.

Yiming Chen · 0 citations