Skip to content
Preprint

Bridging extrinsic and intrinsic variable importance

Jul 2026 · 0 citations · 29 references
Mathematics

TL;DR

Simulations show that VIMP and MPLOCO agree most closely when the fitted learner is well aligned with the data-generating mechanism, and clarify when intrinsic and extrinsic importance can be interpreted similarly and when they provide complementary information.

Abstract

Variable importance may describe either intrinsic predictive information in a population or extrinsic importance for a fitted prediction rule. Quantifying the uncertainty in variable importance estimates is critical for interpretation. Methods for estimating intrinsic variable importance (we will refer to these as VIMP) and the minipatch leave-one-covariate-out procedure (MPLOCO) target intrinsic and extrinsic importance, respectively, and provide methods for computing standard errors. These two approaches have a shared structure, comparing prediction performance with and without features, but the relationship between them has not been formally characterized. We establish conditions under which the two perspectives align. Under squared-error loss, if the fitted full and reduced learners converge to their oracle counterparts sufficiently fast, then MPLOCO is asymptotically equivalent to VIMP. We provide further conditions extending this result to general loss functions and formalize grouped MPLOCO for potentially overlapping feature groups. Through simulations, we show that VIMP and MPLOCO agree most closely when the fitted learner is well aligned with the data-generating mechanism. In a high-dimensional grouped simulation, both procedures identified the signal-containing groups. In an analysis of HIV-1 VRC01 neutralization sensitivity, both methods placed the same three biologically relevant feature groups among their highest-ranked groups. These results clarify when intrinsic and extrinsic importance can be interpreted similarly and when they provide complementary information.

View source

Similar papers

Open access Aug 2026

Parameter-Independent Feature Ranking with Volume-Integrated Sharma–Mittal Entropy: Kernel-Based Estimation, Theoretical Properties and Empirical Validation

Feature selection is a critical step in regression problems where a large number of continuous explanatory variables explain the same target through different dependency structures. Classical filters may remain sensitive to a single form of dependence, a single scale, or a specific discretization scheme; generalized entropy measures, on the other hand, typically require the parameters to be fixed at a single point. This study proposes a framework that evaluates the Sharma–Mittal entropy volumetrically across a two-dimensional parameter region rather than for a single parameter pair. For the continuous target and explanatory variables, the marginal, joint, and conditional densities are obtained using a Gaussian kernel density estimation; the conditional entropy and information gain surfaces are integrated across the region Ω = [0.05, 0.95]2 in the α-β plane to define three indices: PICSME, which measures the conditional uncertainty volume; PIGSME, which measures the gain volume; and NIGSME, which is the ratio of this gain to the total entropy volume of the target. The method is supported by bandwidth consistency and the renormalization of conditional densities; thus, the issue of negative gain that can occur in the continuous variables is resolved, yielding positive and interpretable scores across all six datasets. It is formally demonstrated that the fact that the three indices produce the same ranking is not an empirical observation but rather the result of a monotonicity relationship valid under a fixed target entropy volume. The method is compared with Pearson and Spearman correlations, the Shannon information gain, mutual information, and random forest variable importance across six regression datasets (Airfoil Self-Noise, AirQualityUCI, BodyFat, Meteorology, Concrete, and WineQualityWhite) that differ in their sample size, dimensions, and application domain. The evaluation is not limited to ranking consistency; the out-of-sample prediction performance is measured using least-squares models on the top-k subsets, with rankings calculated from the training partition. The findings show that NIGSME exhibits a performance comparable to that of built-in filters, outperforms them on the Concrete and Meteorology datasets, and never ranks as the weakest method on any dataset. The results demonstrate that volumetric entropy metrics defined across the entire parameter space provide a feature-ranking tool that is independent of parameter selection for continuous variables.

Nida Oruç Ünal, Muzaffer Göztaş, Doğan Yıldız · 0 citations
Preprint Jul 2026

Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs

Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter importance that explicitly captures the interaction between model weights and first-order gradient information and identifies parameters that disproportionately influence model behavior, such as those responsible for collapse phenomena in LLMs. Across a range of models and settings, we show that WAG surfaces a tiny but critical subset of parameters whose modification leads to dramatic degradation in performance, a failure mode that existing importance metrics overlook. These findings reveal a previously underexplored interplay between weights and gradients, suggesting that parameter importance cannot be fully understood through either signal alone. The surprising effectiveness of WAG points to fundamental structural properties of trained networks and motivates new open questions about the role of zeroth-order and first-order information in deep learning. We demonstrate the practical utility of WAG across multiple applications, including expert allocation in mixture-of-expert architectures, parameter-specific unlearning, mixed-precision quantization, and layer selection for knowledge editing. Our results position WAG as a unified approach for analyzing, debugging, and controlling LLMs, and opens new directions for principled model-level interpretation.

Shrestha Datta, Hongfu Liu, Anshuman Chhabra · 2 citations · ⚡1
Preprint Aug 2026

Handling covariate shift by model averaging

Distributional mismatch between the data used to construct a statistical procedure and the population to which it is ultimately applied is pervasive in modern data analysis. We study covariate shift, a fundamental instance of this problem, and develop an adaptive importance-weighted model averaging method for prediction when labeled observations are available from a source distribution, whereas only unlabeled covariates are observed from the target distribution. Procedures fitted directly to the source sample generally optimize prediction risk under the source distribution and may therefore be suboptimal for target prediction. Importance weighting by the density ratio between the target and source covariate marginals provides a natural correction, but a small number of large density-ratio values can substantially inflate the variance of the resulting estimator in finite samples. We address this bias-variance trade-off by treating the degree of importance-weighting correction as a source of model uncertainty. Specifically, we construct a family of adaptive importance-weighted least-squares estimators by raising the estimated density ratio to a range of exponents, with the endpoints corresponding to ordinary least squares and standard importance-weighted least squares, and form a data-driven average over these candidates. Under model misspecification, the proposed model averaging estimator is shown to be asymptotically optimal relative to the infeasible best convex combination of the candidate estimators. Under correct specification, a diverging penalty is shown to make the selected weights concentrate near the ordinary least-squares endpoint. Simulations and a real-data application show that the proposed method achieves competitive target-prediction performance across the settings considered.

Yifan Zhang, Tianfa Xie, Xinyu Zhang · 0 citations
Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space.

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations
Preprint Aug 2026

Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single quantity -- estimable as the difference of two discriminators'held-out cross-entropies, D_CJS = CE(Z|X) - CE(Z|X,Y) -- is exactly the part of a dataset difference that covariate shift cannot explain. We propose the Conditional Jensen-Shannon Discrepancy (CJSD): with a task indicator Z, the chain rule I(Z;X,Y) = I(Z;X) + I(Z;Y|X) splits total task discrepancy exactly into a covariate axis and a functional axis, both estimable from two ordinary classifiers, with no task-specific predictors, generative models, or bootstrap surrogates. We prove a covariate-null property (the functional axis is exactly zero under pure covariate shift, however severe), a drift-mass law (D_CJS/ln2 equals the mass of the disagreement region for deterministic labels), a one-sided misspecification-control inequality (each direction of estimation error is bounded, unconditionally, by the excess risk of a single discriminator), and a fixed-measure metrization via an identifiability lemma. Empirically, on a ten-measure battery over 202 dataset pairs (synthetic, Electricity, Covertype), only the two conditional-information estimators -- CJSD and a kNN plug-in for the same estimand -- separate concept from covariate shift with AUC 1.0; the case for CJSD is the estimator: under controlled dimensionality scaling the kNN plug-in fails from d=64 while the discriminator route holds to d=256 with a swappable classifier, and it alone yields paired confidence intervals and sequential extensions from the same learned object. The same estimator audits the conditional fidelity of synthetic-data generators that marginal and joint QA metrics pass, detects annotation-guideline changes invisible to input-space monitors, and supports null-calibrated fairness audits.

Kentaro Oda · 1 citation