Skip to content
Preprint

Controlling for Omitted Variable Bias in Deep Neural Networks

Aug 2026 · 0 citations · 74 references
Mathematics Computer Science

TL;DR

This work introduces an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, and shows how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect.

Abstract

Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, and demonstrate consistent estimation of true effects. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data. Code is available at https://github.com/mpff/cocodeel.

View source

Similar papers

#machine learning Preprint Sep 2026

Weighted Spline-Expanded Networks with Distributional Balancing for Continuous Treatment Effects

Estimating causal effects with continuous treatments in observational studies is challenging due to confounding, model misspecification, and high-dimensional covariates. We propose the Weighted Spline-Expanded Network (WSENet), an end-to-end neural framework that addresses these challenges by combining covariate balanc...

Shu-Cheng Liu, Chan Park, Guan-Hua Chen · 0 citations
Preprint Aug 2026

Handling covariate shift by model averaging

Distributional mismatch between the data used to construct a statistical procedure and the population to which it is ultimately applied is pervasive in modern data analysis. We study covariate shift, a fundamental instance of this problem, and develop an adaptive importance-weighted model averaging method for predictio...

Yifan Zhang, Tianfa Xie, Xinyu Zhang · 0 citations
Preprint Sep 2026

Design-Assisted Regression

A general design-assisted regression framework in which the estimating criterion depends on both the conditional model for $Y \mid \bfX$ and structured features of the covariate distribution, which improves estimation while preserving first-order prediction performance.

S. Ye, Guan-Bo Wang, Cong Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

Theoretically, it is proved that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation.

Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al. · 0 citations
Preprint Oct 2026

Inference after data-driven control-unit selection in difference-in-differences with estimated covariance

In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper,Nakano and Hoshino (2016), and the present paper jointly prov...

Ryoya Nakano, Takahiro Hoshino · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.