It is formally established that GKRReg belongs to the family of redescending M-estimators, providing the theoretical foundation for the inferential procedures that follow and proposing a pairs bootstrap that re-estimates the kernel width hyper-parameter gamma^2 on every replicate, capturing variability that the sandwich ignores.
Abstract
The Gaussian Kernel Robust Regression method (GKRReg) is a robust regression estimator that iteratively re-weights observations via a Gaussian kernel so that outliers and leverage points receive near-zero weight, with convergence of the estimation algorithm theoretically guaranteed. Despite a thorough study of estimation, the original work leaves open the problem of statistical inference for the regression coefficients. We fill this gap with three contributions. First, we formally establish that GKRReg belongs to the family of redescending M-estimators, providing the theoretical foundation for the inferential procedures that follow. Second, we derive a closed-form analytic sandwich variance estimator based on the theory of generalised M-estimators, corresponding to the HC0 class of heteroskedasticity-robust covariance matrices; we show that a finite-sample correction analogous to HC3 requires the weighted hat matrix of the converged IRWLS step, and identify this as a direction for future work. Third, we propose a pairs bootstrap that re-estimates the kernel width hyper-parameter gamma^2 on every replicate, capturing variability that the sandwich ignores. All procedures are implemented in the R package gkrreg, which also provides four estimators for gamma^2 and an automatic data-driven selection procedure, comprehensive diagnostic plots, and six real datasets from the robust regression literature. Applications to real data sets and comparison with traditional robust regression models highlight the potential of the GKRReg and the usability of the R package.
We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.
J. M. Olea, Ryan Strong, Amilcar Velez et al.· 0 citations
Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance estimation, and higher-order moment estimation in a single framework. Our main result gives a precise $1+o_{\mathbb{P}}(1)$ expansion for the $L^2$ error at each admissible regularization parameter. The expansion consists of bias and three variance terms corresponding respectively to variation across the independent sample paths, latent signal variation at each sample point, and variation from measurement errors, identifying the refined error structure underlying functional data. As applications, we show that KRR attains the minimax rate for source smoothness $s \leq 2$ but becomes suboptimal in the sparse regime for $s>2$ due to saturation. A technical ingredient is a set of concentration inequalities for $U$-statistics suited to the dependent product structure of functional observations.
We introduce a robust nonparametric regression framework for functional covariates that combines functional principal component analysis (FPCA), marginal copula-scale normalization, bounded-score M-estimation, and multivariate Bernstein smoothing. The proposed procedure reduces the infinite-dimensional functional predictor to a low-dimensional score representation, transforms the retained scores onto the compact unit cube, and estimates a conditional M-functional through a smoothly aggregated system of local estimating equations. This construction is designed to accommodate nonlinear regression structure, heavy-tailed score distributions, and response contamination while limiting the influence of extreme observations. Under suitable regularity and undersmoothing conditions, we establish pointwise and uniform consistency, derive explicit convergence rates, and prove asymptotic normality. The limiting variance contains an explicit Bernstein concentration factor that plays a role analogous to the integrated squared kernel in classical nonparametric regression. The analysis also clarifies the interaction among the projection dimension, the Bernstein resolution, the empirical copula transformation, and the effective local sample size. The finite-sample performance of the method is examined through simulations involving heavy-tailed functional scores, Student-t errors, nonlinear regression effects, and increasing response contamination. The proposed estimator exhibits strong overall predictive performance and good robustness, with particularly favorable behavior under absolute-error criteria.
In survey sampling, the goal is to estimate finite population parameters such as totals, means, and proportions. At the estimation stage, it is common to have access to auxiliary information in the form of covariates known either in aggregate form or for each population unit. These covariates are often used, through models relating them to the variable of interest, to improve efficiency; this approach is known as model-assisted estimation. Modern applications increasingly involve settings where a large number of covariates are observed, sometimes of the same order as the sample size. While this setting offers greater modeling flexibility, it also creates important challenges for inference. In this article, we study variance estimation for the generalized regression (GREG) estimator in high-dimensional regimes. We derive new theoretical results that characterize the high-dimensional asymptotic bias of commonly used variance estimators, including those based on Taylor linearization. Furthermore, under suitable distributional assumptions on the covariates, we show that a cross-validated variance estimator is naturally asymptotically unbiased.
Inference for spectral edges of large covariance matrices is a fundamental problem in high-dimensional statistics. A major difficulty is that the largest non-spiked sample eigenvalues, which serve as natural estimators of the edge, fluctuate on the Tracy--Widom scale. Consequently, valid inference requires accurate centering by the deterministic spectral edge together with a precise scaling constant, both of which are often difficult to estimate in practice under general unknown population covariance structures. In this paper, we propose a bias-corrected multiplier bootstrap procedure for inference on the deterministic edge of the bulk spectrum. The key idea is to introduce a carefully calibrated multiplier perturbation that regularizes the edge fluctuation to a slightly larger scale at which Gaussian approximation becomes tractable. The resulting confidence interval is constructed directly from bootstrap eigenvalues, together with a data-driven recentering step that corrects the bootstrap-induced shift of the deterministic edge. On the theoretical side, we show that, after bias correction and rescaling, the largest few non-spiked bootstrap eigenvalues are asymptotically Gaussian conditionally on the data. Building on this result, we establish the asymptotic validity of the proposed confidence interval, whose length is only slightly larger than the Tracy--Widom scale, and prove vanishing coverage under alternatives in which additional spikes separate from the bulk at a local scale larger than $n^{-1/6}$. As a consequence, the same confidence interval yields a threshold-free estimator for the number of spikes, without requiring the spikes to be distinct or very large. Equivalently, the procedure yields a data-driven and theoretically justified cutoff for the scree plot.