Linear regression is one of the simplest and most widely used tools to learn patterns from data: it fits a set of coefficients so that a linear combination of predictors best matches observed responses. The quality of the fit is measured by the residual sum of squares, the total squared mismatch between predictions and data, whose minimum defines the training loss. We consider Gaussian design and noise, with teacher coefficients independently drawn from a general distribution $p(\beta)$, and a general class of separable regularizers, including Ridge and Lasso. Using the zero-temperature replica method, we compute analytically the large-deviation statistics of the minimum training loss for large numbers $P$ of predictors and $N$ of observations, with $r=P/N$ fixed. The rate function we compute governs rare sample-to-sample fluctuations of the optimal loss. Extensive numerical simulations are in excellent agreement with our theory and clearly show a pronounced deviation from the Gaussian regime of typical fluctuations in the tails.
The error of regression using noisy function values naturally decomposes into an approximation error from the choice of model space and stochastic error from the noise on the observations. Statistical theory is mostly concerned with the stochastic error while approximation theory focuses on the approximation error, oft...
In a statistical factor model, principal components (or eigenvectors) of a sample covariance matrix serve as estimates of {\it principal directions}, the true drivers of co-movement of a collection of observed variables. We write the often substantial error in these estimates as a sum of two interpretable terms, which...
Alex Bernstein, Lisa R. Goldberg, Nicholas Gunther et al.· 0 citations
We investigate the least squares linear regression problem with random partial Discrete Fourier Transform (DFT) matrices, providing a rigorous analysis of the model's generalization error. By leveraging tools from random matrix theory, we derive exact non-asymptotic bounds for the risk of the Moore-Penrose estimator, w...
We study the mean function of longitudinal functional data, where each subject contributes a small number of complete profiles over a general domain, observed at random visit times. The mean is projected onto an orthonormal basis in the time direction, and each coefficient function is estimated by a weighted average of...
It is proved that both recover the infinite-autoregressive representation of the true process at a near-parametric rate in fixed dimension, so the truncation introduces no asymptotic bias.
This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the...
Zhi-Jun Liu, Dandan Jiang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.