Skip to content
Preprint

Contrast-Free ICA and Causal Inference via Wasserstein Distances to the Gaussian

Jul 2026 · 1 citation · 45 references
Mathematics

TL;DR

This work defines empirical plug-in estimators and prove distribution-free uniform convergence under finite-moment assumptions, before detailing three practical solvers: a Picard-style orthogonal optimizer for ICA, an exhaustive dynamic program for causal order search, and a greedy order search variant.

Abstract

We study the squared $2$-Wasserstein distance to the standard Gaussian as a non-Gaussianity criterion and use it for linear Independent Component Analysis (ICA) and causal discovery in Linear Non-Gaussian Acyclic Models (LiNGAM). Unlike commonly used parametric contrasts and approximations of information-theoretic quantities, this criterion requires no distributional regularity beyond finite second moments, involves neither approximation nor tuning parameters, and can be computed exactly and efficiently from empirical order statistics. Our analysis relies on a strict subadditivity property of the $2$-Wasserstein distance to the Gaussian. At the population level, we prove exact identification of the ICA unmixing matrix, up to signed permutation, and give an analogous characterization of causal orders through sequential least-squares residuals. We then define empirical plug-in estimators and prove distribution-free uniform convergence under finite-moment assumptions, before detailing three practical solvers: a Picard-style orthogonal optimizer for ICA, an exhaustive dynamic program for causal order search, and a greedy order search variant. Empirically, we demonstrate competitive performance for both tasks and provide open-source implementations for source separation and causal discovery.

View source

Similar papers

Preprint Jul 2026

Linear Independent Component Analysis via Optimal Transport

Linear Independent Component Analysis (ICA) recovers jointly independent source signals from their linear mixtures. To achieve this, classical ICA algorithms attempt to maximize non-Gaussianity, measured by negentropy, which is linked to independence by information theory. Because exact negentropy optimization is intractable, they rely on proxy contrast functions, such as fourth-order cumulants, and parametric log-likelihoods. We propose instead to measure non-Gaussianity using the squared Wasserstein distance $W_2^2$ to a standard Gaussian. We prove that the Wasserstein distance between a standard normal distribution and linear projections of the data is maximized when the projection recovers an independent component. Based on this observation, we propose the OT-ICA algorithm which finds this projection by gradient-based optimization. Empirical evaluation on simulated data shows that OT-ICA outperforms proxy-based methods for different distributions of the latent variables. Application to EEG artifact removal and econometric price discovery confirm OT-ICA can be used for applied ICA tasks without distributional assumptions.

Ashutosh Jha, M. Besserve, Simon Buchholz · 1 citation
Nov 2026

Statistical inference for Bures-Wasserstein flows

We develop a statistical framework for conducting inference on collections of time-varying covariance operators (covariance flows) over a general, possibly infinite dimensional, Hilbert space, and fully develop asymptotic theory based on interpretable and verifiable assumptions. We model the intrinsically non-linear structure of covariances by means of the Bures-Wasserstein metric geometry. We make use of the Riemannian-like structure induced by this metric to define a notion of mean and covariance of a random flow, and develop an associated Karhunen-Loève expansion. We then treat theproblem of estimation and construction of functional principal components from a finite collection of covariance flows, observed fully or irregularly. Our theoretical results are motivated by modern problems in functional data analysis, where one observes operator-valued random processes – for instance when analysing dynamic functional connectivity and fMRI data, or when analysing multiple functional time series in the frequency domain. Nevertheless, our framework is also novel in the finite-dimensions (matrix case), and we demonstrate what simplifications can be afforded then. We illustrate our methodology by means of simulations and data analyses.

Leonardo V. Santoro, Victor M. Panaretos · 0 citations
Preprint Aug 2026

Robust Scale Estimation in Additive Noise via Weighted Order Statistics

This manuscript develops a non-parametric and robust framework for estimating the scale of additive noise in weakly sparse systems. The method does not require independence, prescribed dependence, or temporal regularity of the noise sequence. We introduce a class of order-statistic estimators based on comparing the sorted observations with deterministic or random proxies generated from a reference noise distribution. This purely spatial approach avoids preliminary filtering or temporal decorrelation, and therefore preserves the sparsity structure of the latent signal. We establish non-asymptotic concentration inequalities for weighted loss functions, with bounds that separate the contribution of the signal from the discrepancy between the ordered noise and the proxy. We then control this proxy discrepancy in independent and correlated regimes, including heavy-tailed reference laws. Finally, we apply the method to high-frequency observations of continuous-time stochastic processes, obtaining scale estimators for fractional Brownian motion and stable L\'evy noise in the presence of lower-variation additive perturbations.

Jorge Ignacio Gonz'alez C'azares, Arturo Jaramillo · 0 citations
#machine learning Preprint Aug 2026

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

This work proposes SURE-Ridge, a non-iterative, closed-form estimator for equal variance linear Gaussian SEM, which achieves the lowest structural Hamming distance in the small-sample regime and the lowest run time across all sample sizes tested, compared with NOTEARS, DAGMA, and GBNSL baselines.

Sambit Mishra, Urbashi Mitra · 0 citations
Preprint Aug 2026

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning. The Widely Applicable Bayesian Information Criterion (WBIC) relies on local learning coefficients $\lambda$, which in the analytic case coincides with local Real Log Canonical Thresholds (RLCT) of the Kullback-Leibler divergence of the model, to capture correct marginal likelihood asymptotics. Exact computation of the learning coefficients has been limited to special cases, and only sampling-based estimation methods are generally applicable. We present the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks. Beyond providing ground truth to calibrate sampling-based estimators, exact computation reveals algebraic structure in learning coefficients that sampling cannot and out-speeds it in the shallow regime.

Grégoire Sergeant-Perthuis, Elias P. Tsigaridas, Jules Tsukahara Cqsb et al. · 0 citations
Preprint Jul 2026

Distance Profile Embedding for Independence and Conditional Independence Testing of Random Objects

Testing independence or conditional independence is fundamental to statistical inference, yet existing methods for non-Euclidean random objects often face a difficult trade-off between geometric flexibility and theoretical tractability. We introduce the Distance Profile Embedding (DPE), a novel representation that maps random objects from general metric spaces into a Hilbert space of square-integrable functions. We prove that this mapping is injective and preserves full distributional information without requiring isometric Hilbert embeddings or one-to-one correspondence conditions. Leveraging the DPE, we develop a unified framework for marginal and conditional independence testing of random objects that enjoys a rigorous asymptotic theory for both size and power. Notably, our framework is the first in the literature to accommodate object-valued conditioning variables when testing conditional independence, overcoming the Euclidean or Hilbertian constraints of existing methodologies. We facilitate the calculation of analytic $p$-values using closed-form asymptotic null distributions, which avoids the computational burden of permutation tests common in existing metric-based methods. The numerical properties of our methods are demonstrated through both simulations and two real-world applications involving gut microbiome compositions and global human mortality distributions, respectively.

W. Tan, Bing Li, Lingzhou Xue · 0 citations