Skip to content
Open access

Correlation-Sensitive Adaptive LASSO for High-Dimensional Data: A Redundancy-Aware Regularization Approach

Aug 2026 · Symmetry · 0 citations · 18 references

TL;DR

Overall, CDA-LASSO directly incorporates the internal correlation structure of the data into the penalty weights without requiring a predefined graphical structure and provides a practical methodological extension for more controlled and parsimonious variable selection in high-dimensional correlated settings.

Abstract

In multivariate statistical analysis, accurate modeling of the covariance structure is critical for high-dimensional data analysis, variable selection, and regularization. In high-dimensional settings, strong inter-variable correlation and redundancy are key factors limiting the performance of classical sparsity-based methods. While LASSO and its variants provide effective tools for coefficient shrinkage and variable selection, they may select redundant variables and produce unnecessarily complex models in highly correlated settings. In this study, a Correlation-Sensitive Adaptive LASSO (CDA-LASSO) method is proposed to address these limitations. The proposed approach is based on a hybrid weighting mechanism that makes the penalty term sensitive not only to initial coefficient magnitudes but also to the correlation structure between variables. This structure incorporates correlation-based redundancy information and imposes stronger penalties on predictors with higher directed redundancy scores. Under fixed-dimensional regularity conditions, the bounded correlation multiplier is shown to preserve the selection consistency and oracle limiting distribution of Adaptive LASSO. The method was evaluated through 14 high-dimensional simulation scenarios covering different sample sizes, dimensionalities, sparsity levels, correlation strengths, support structures, and normal or heavy-tailed errors. The results indicate that the Max and kMean variants generally reduce the false discovery rate and model size relative to LASSO and Elastic Net while maintaining broadly comparable predictive performance. Numerical improvements over Adaptive LASSO were also observed in several scenarios, although these differences were not uniformly statistically significant. Under very high correlation, reductions in false discoveries were sometimes accompanied by modest decreases in the true positive rate. The real-world Riboflavin analysis further showed that the CDA-LASSO variants produced smaller models than LASSO and Elastic Net while retaining comparable prediction errors. Overall, CDA-LASSO directly incorporates the internal correlation structure of the data into the penalty weights without requiring a predefined graphical structure and provides a practical methodological extension for more controlled and parsimonious variable selection in high-dimensional correlated settings.

Read PDF

Similar papers

Aug 2026

Inverse-based Lasso Regression: Novel Algorithms for Feature Selection and Multicollinearity Mitigation

Novel algorithms for estimating Lasso regression parameters by reformulating the problem as an inverse single-point optimization task are introduced, eliminating the need for explicit regularization parameter specification while maintaining robust feature selection capabilities and effective multicollinearity mitigatio...

Gribanova Ekaterina, Gerasimov Roman · 0 citations
Open access 2026

Variable Selection under Multicollinearity and Heavy-Tailed Errors: A Systematic Assessment of Classical and Penalized Methods

Variable selection in high-dimensional regression becomes particularly challenging when multicollinearity and heavy-tailed errors occur simultaneously. This study systematically evaluates the performance boundaries of classical, shrinkage, sparse, and robust regression methods under these conditions. A Monte Carlo simu...

B. Ofuru, Ijomah M. A., Nwakuya M. T. · 0 citations
#machine learning Preprint Sep 2026

A Parameter-Free Zeroth-Order Method with Covariance Matrix Adaptation and Effective Dimension

Zeroth-order optimization methods are essential for solving black-box problems where gradient information is unavailable or expensive to compute. This paper presents POEM-CMA, a novel parameter-free stochastic zeroth-order algorithm that extends the recent POEM method by integrating covariance matrix alignment and the...

Alexander Sholokhov, A. Rogozin · 0 citations
Preprint Aug 2026

Fast high-dimensional mean testing via logistic regression

This work proposes computationally efficient tests for equality of mean vectors of two or more high-dimensional populations by establishing an equivalence between equality of means and a zero population logistic regression parameter.

Sayan Das, Debraj Das, S. Dutta · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.