A general design-assisted regression framework in which the estimating criterion depends on both the conditional model for $Y \mid \bfX$ and structured features of the covariate distribution, which improves estimation while preserving first-order prediction performance.
Abstract
We consider regression problems in which the marginal distribution of the covariates is informative for estimation and variable selection, rather than merely auxiliary. Motivated by random-design, high-dimensional, and latent-effect settings, we propose a general design-assisted regression framework in which the estimating criterion depends on both the conditional model for $Y \mid \bfX$ and structured features of the covariate distribution. The framework identifies two roles of design information: stabilizing weak design directions through quadratic regularization and correcting latent-effect distortion through nuisance augmentation. We establish oracle properties for the resulting estimator, separate the effects of stochastic error, shrinkage, and approximation, and compare it with a benchmark sparse procedure that ignores design information. These results show that the proposed framework improves estimation while preserving first-order prediction performance. Numerical studies and two real-data applications illustrate the practical impact of incorporating design information.
We study additive regression under an unknown and potentially non product design distribution, allowing the number of covariates to grow with the sample size. We consider a coupled class that separately controls the smoothness of the marginal densities and of each additive component multiplied by the corresponding marg...
In this work, we study the one-dimensional regression problem under random design and Gaussian errors. Our framework is very general: we make no prior assumptions about the design (which may be nonstationary and exhibit short or long-range dependence), nor do we assume that the errors are homoscedastic. We examine in d...
Emmanuel Caron Parte, J. Dedecker, B. Michel· 0 citations
Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for sc...
Jyotishka Ray Choudhury, K. A. Verchand, R. Samworth et al.· 0 citations
We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level....
When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the tra...
María Eugenia Riaño· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.