Skip to content
Preprint

Emputation: Identification-Guided Neural Imputation Framework

Jul 2026 · 0 citations
Mathematics

TL;DR

It is shown that the population minimizer of the emputation risk recovers the target extrapolation distribution under a broad class of identification assumptions, including several missing-not-at-random assumptions.

Abstract

We propose Emputation, a deep generative framework for learning imputation models. Emputation targets the extrapolation distribution of missing variables given observed variables, and training is guided by specific missingness assumptions that guarantee identification of the target distribution. The training objective, called the emputation risk, is an energy-score-based risk in which the identification assumption determines how observed entries are masked and which observations contribute to training. The resulting framework enables direct conditional sampling for multiple imputation. We show that the population minimizer of the emputation risk recovers the target extrapolation distribution under a broad class of identification assumptions, including several missing-not-at-random assumptions. Simulations show strong performance under both pointwise and distributional evaluation metrics, and an application to an Alzheimer's disease dataset demonstrates its practical value.

View source

Similar papers

Open access Jul 2026

An end-to-end distribution- and task-aware parallel imputation method for improving target task performance

Most machine-learning imputation techniques treat missing values independently of the downstream task, resulting in suboptimal predictive performance. While some recent methods jointly train an imputer with a target predictor, they fail to produce diverse and context-sensitive imputations and suffer from training inefficiencies. Additionally, in many methods, missing values are ignored during training, and the imputers are trained only on observed data. To overcome these limitations, we propose an imputation method, named end-to-end task-aware parallel Imputation with class-wise prototypes (ETPI). It captures the class-conditional distributions of input data using a few proxies. For the missing entries in any sample, ETPI generates class-aware pseudo-targets based on the predictor loss and the estimated class-conditional distribution. By exposing the imputer to both observed data and optimal pseudo-targets during training, ETPI effectively leverages training information to fit the imputer model and aligns imputation with the objectives of the target task. Extensive experiments on classification and regression tasks show that ETPI outperforms other state-of-the-art methods. It also maintains high imputation quality even with limited training data or high missing rates, mainly due to the high quality of the generated pseudo-targets and the integration of imputation and prediction into a single end-to-end pipeline.

Karrar Al-Kaabi, Davood Zabihzadeh · 0 citations
Preprint Jul 2026

Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation

Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution. We recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data. Leveraging positive semi definite (PSD) kernel densities we obtain a convex empirical risk problem with closed form marginals, solvable by a Newton interior point method. The resulting PSD Impute model yields both single and multiple imputations from the same fitted density, enjoys statistical consistency with fast adaptive excess risk beating the curse of dimensionality for very regular probabilities. Preliminary experiments on one synthetic and eleven real world datasets already indicate competitive distributional accuracy compared with popular imputation baselines, suggesting strong practical promise.

A. Basteri, C. Ciliberto, Alessandro Rudi · 0 citations
Open access Aug 2026

A Copula-Tensor Neural Network Framework for High-Dimensional Causal Inference

Estimating conditional average treatment effects (CATEs) in high-dimensional causal inference problems remains challenging because complex nonlinear relationships, heterogeneous feature distributions, and dependence among covariates can limit the effectiveness of conventional machine learning approaches. To address this challenge, we propose a copula-enhanced neural learning framework that integrates empirical copula transformations, manifold-based feature augmentation, structured treatment–covariate interaction representations, and deep neural networks for flexible CATE estimation. The empirical copula transformation does not introduce additional dependence information; instead, it provides a rank-based feature representation that normalizes marginal distributions, reduces sensitivity to heterogeneous feature scales and extreme observations, and offers a dependence-aware representation for subsequent learning. The proposed framework is evaluated through Monte Carlo simulations under diverse data-generating mechanisms and a real-world application using the Criteo uplift dataset. The simulation study examines the contribution of individual model components through ablation experiments and compares the proposed approach with established causal learning methods. Results demonstrate that the proposed framework achieves competitive CATE estimation accuracy while providing stable policy evaluation based on Inverse Propensity Scoring (IPS) and Doubly Robust (DR) estimators. In the Criteo application, the proposed method exhibits predictive performance comparable to conventional neural-network approaches while producing more stable Doubly Robust policy value estimates. These findings suggest that copula-based feature representations combined with deep learning provide a flexible approach for heterogeneous treatment effect estimation, particularly in high-dimensional settings with complex covariate dependence. The benefits of the proposed framework depend on data characteristics, including sample size, dimension, dependence structure, and treatment assignment mechanisms.

Jong-Min Kim · 0 citations
Preprint Aug 2026

Inferential Evaluation of Surrogate-Derived Models under Covariate Shift

In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved. Evaluating its target performance is essential for determining whether decisions based on the model remain reliable, yet it is difficult when gold labels are scarce, and covariate distributions differ across data sources. We study a three-sample setting with a small gold-labeled source, a larger surrogate-labeled source, and an unlabeled target. Under conditional transportability, we evaluate the surrogate-derived model against the latent gold-standard outcome in the target population. We propose cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios. We also combine outcome-regression augmentation with a kernel correction for estimating the model near a threshold, accounting for uncertainty from all three samples. We establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC. Simulations assess bias, coverage, and sensitivity to bandwidth and relative sample sizes. A retrospective temporal validation on Chatbot Arena and a semi-synthetic ACS-Income study provide validation in real-world AI applications.

Long-Tian Shi, Molei Liu, Doudou Zhou · 0 citations
Review Open access Jul 2026

Systematic assessment of mixed imputation methods and explainable machine learning

The convergence of artificial intelligence and precision oncology is frequently hampered by the quality of real-world clinical data, particularly the pervasive challenge of missing values. This opinion review critically appraises the methodology and evidentiary framework of the study, which proposes a hybrid imputation architecture, HDI-MF-Gower, integrated with an extra trees classifier and Shaply Additive exPlanation interpretability for predicting survival outcomes following curative gastrectomy. We deconstruct the pivotal assumptions and potential sensitivities of their adaptive weighted similarity initialization. This design is engineered to provide a “warm start” aligned with the underlying data structure for iterative imputation, theoretically mitigating the risks of distributional distortion associated with simplistic initialization strategies. However, a primary boundary of the current evidence lies in the validation hierarchy; the reported validation relies predominantly on random splitting within a single-center cohort, lacking the robustness of temporal extrapolation or genuine external validation. Furthermore, statistical comparisons suggest that the performance differences between the proposed model and several robust ensemble baselines are not consistently distinguishable, making it difficult to attribute performance gains solely to the specific choice of the learner. We conclude that future research must construct a more rigorous evidence chain within multicenter and multimodal frameworks. Crucially, adherence to transparent reporting of a multivariable prediction model for individual prognosis or diagnosis + artificial intelligence guidelines - specifically regarding missing data mechanisms, sensitivity analyses, calibration and net benefit assessments, and the availability of reproducible materials - is essential to substantiate generalizable clinical utility.

Jiayao Li, Yan Zhao · 0 citations
Preprint Jul 2026

Amortized Inference for Sampling Distributions Where the Bootstrap Fails

Efron's bootstrap is the default tool for estimating the sampling distribution of a statistic, yet it is provably inconsistent for maxima of bounded-support distributions, means under infinite variance, extreme quantiles, and tail-index estimators. The classical remedies, the m-out-of-n bootstrap and subsampling, require rate corrections that depend on unknown parameters and behave erratically at realistic sample sizes. We propose an amortized alternative: a neural network is trained on simulated datasets drawn from a prior over a distribution family, using single independent draws of the root T_n - T(F) scored by the pinball loss, a proper scoring rule whose population minimizer is the posterior-predictive law of the root. At test time, a single forward pass maps one dataset of n = 200 observations to its full sampling-distribution estimate, from which confidence intervals follow directly. On four canonical bootstrap-failure problems (bounded-support maximum, alpha-stable mean, Pareto tail index, and 99% value-at-risk under tempered stable returns), the method attains nominal 95% coverage, beats every feasible classical method in Wasserstein distance to the true sampling distribution, and captures over 97% of the achievable improvement where the exact Bayes-optimal answer is computable. For the value-at-risk problem no distribution-free method can reach nominal coverage at all; the learned method attains 94.7%. A single universal network with a statistic token matches all four specialists, and on real daily market returns the unchanged model averages 0.87 coverage against 0.73 for the bootstrap, as predicted by our out-of-family analysis.

Akash Deep · 0 citations