Skip to content
Preprint

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

Aug 2026 · 0 citations
Mathematics Computer Science

Abstract

Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under sub-Weibull noise, we establish selection consistency through approximation-error separation without relying on RIP-type conditions, together with nonasymptotic oracle risk bounds that remain valid under model misspecification. The same coding principle places heterogeneous classes on a common complexity scale at a small additional class-identification cost. This extension yields class--model recovery under suitable identifiability conditions and risk adaptation across classes. We further develop a complexity-guided search path that makes the computation--statistics trade-off explicit. Large penalties yield polynomial-size retained search regions with high probability, whereas smaller penalties sharpen the oracle risk benchmark. Numerical experiments illustrate stable support recovery and favorable estimation performance under strong dependence and model-class uncertainty.

View source

Similar papers

Preprint Jul 2026

Best Subset Selection in Linear Regression: Fixed-Design Error Bounds and Insights for Random Designs

We study exact support recovery by best subset selection in linear regression under a fixed design. For a known support size $s$ and ambient dimension $d$, we derive a non-asymptotic upper bound on the probability that best subset selection fails to recover the true support. The bound is expressed through a deterministic subset-separability parameter, which measures how well the true support can be distinguished from competing supports after projection. The result holds for all sample sizes $n$ exceeding a certain sufficient threshold which we state explicitly in terms of the signal-to-noise ratio, the subset separability of the realized design, and a logarithmic factor of order $\ln s + \ln(d - s)$. In contrast to random-design analyses, no full log-combinatorial term over the candidate support class appears. We discuss how such terms may reappear when the design is random and the separability parameter must be controlled uniformly over many competing subsets. The fixed-design formulation and the proof strategy also indicate settings in which the effective complexity of best subset selection may be reduced, for instance, under structured designs or restricted candidate subset classes.

M. Fedotov · 0 citations
Preprint Jul 2026

Ball-Codifference Screening for Heavy-Tailed Predictors

High-dimensional screening is commonly built on covariance, correlation, or least-squares measures. These summary measures can be unstable or even undefined, when predictors are sparse or have heavy-tailed distributions. Building on our recent work on extended codifference and the idea of Ball-covariance, we develop Ball-codifference for marginal screening in statistical modeling with heavy-tailed predictors and responses. The proposed statistic combines the rank-type geometry of random balls with the codifference as a dependency measure constructed based on the characteristic function, so it can be computed without requiring well-defined finite first or second moments. We define Ball-codifference and its normalized screening utility, and formulate a sure independence screening procedure. Large-sample normality follows from a bounded V-statistic and functional-delta-method argument under standard nondegeneracy and regularity conditions. Simulation studies under Gaussian and sub-Gaussian stable designs show that codifference-weighted Ball screening gives competitive or improved recovery of highly associated predictors, especially when tail heaviness is pronounced. Also, our data example illustrates that our variable screening method significantly improves prediction accuracy in linear regression.

Mohsen Rezapour, Vahed Maroufy · 0 citations
Preprint Jul 2026

Universally Optimal Robustness-Efficiency Tradeoffs for a General Class of Minimum Divergence Estimators

Balancing the efficiency of an estimator under ideal conditions against its robustness under contamination remains a central challenge in robust statistics. While minimum divergence methods offer a flexible alternative to traditional M-estimation, choosing the appropriate discrepancy measure has historically relied on heuristic or empirical justifications. This manuscript introduces a rigorous optimality criterion for this selection process. By investigating the comprehensive Generalized Alpha-Beta Divergence (GABD) family, we explicitly characterize the Pareto frontier dictating the lowest possible asymptotic variance for any strictly enforced asymptotic breakdown point. Our main theoretical results establish that the estimator achieving this mathematical optimum invariably falls within the extended $(\phi, \gamma)$-divergence class. Crucially, the derived optimal tuning parameter, $\phi^*$, given other parameters, depends solely on the desired breakdown threshold and is entirely invariant to both the assumed parametric model and the exact nature of the data contamination. Supported by comprehensive derivations of asymptotic normality, influence functions, and breakdown thresholds for both continuous and discrete settings, this work offers a unified, theoretical resolution to the long-standing problem of optimal divergence selection in robust inference.

Subhrajyoty Roy, Supratik Basu, Abhik Ghosh et al. · 1 citation
Open access Jul 2026

Adaptive Wrapped Robust Canonical Correlation Analysis in High-Dimensional Data

Classical canonical correlation analysis becomes numerically unstable when the number of variables is large relative to the sample size and is sensitive to contamination in observations or individual cells. This study develops an integrated robust and regularized procedure that combines bounded cellwise wrapping, shrinkage estimation of the joint correlation matrix, and robust reweighting in a low-dimensional canonical score space. The resulting observation weights enter a second regularized canonical correlation fit, so the final estimator remains well defined when the combined number of variables exceeds the sample size. The simulation study shows that relative estimation accuracy depends on the signal strength, contamination mechanism, and dimensional configuration. The proposed estimator is competitive in several moderate-signal settings and has a clear computational advantage, whereas the minimum regularized covariance determinant plug-in estimator provides lower estimation error in many high-signal configurations. An additional ultra-high-dimensional experiment demonstrates numerical feasibility with modest memory use but also reveals substantial attenuation, identifying a limitation of the present dense estimator. The results therefore support a regime-dependent interpretation rather than a claim of uniform superiority. The complete reproducible simulation workflow is provided.

Hasan Bulut, Müjgan Zobu, V. Saglam · 0 citations