Skip to content
Preprint

Approximating the null distribution of generalized distance covariance

Aug 2026 · 0 citations · 22 references
Mathematics

Abstract

The null distribution of distance covariance is usually approximated by permutation, which is prohibitive when very small p-values are needed, or by matching a few moments to a parametric family, which is inaccurate in the tails. A third option is to approximate the limiting distribution, a weighted sum of chi-square variables, directly through the spectra of the doubly centred distance matrices. This is used for kernel-based tests but has lacked a rigorous justification. We prove that the empirical spectra give a uniformly consistent approximation of the limiting null distribution, and hence an asymptotically valid test, for a general class of distances of negative type on separable metric spaces. The result covers the Hilbert-Schmidt independence criterion as a special case. We also give an adaptive algorithm that brackets the p-value from a partial eigendecomposition, reducing the cost from $O(n^3)$ to $O(k n^2)$, and a shrinkage correction matching the first two moments. In simulations, the proposed tests are the only non-Monte-Carlo procedures whose empirical type I error converges to the nominal level.

View source

Similar papers

Preprint Jul 2026

Distance Profile Embedding for Independence and Conditional Independence Testing of Random Objects

Testing independence or conditional independence is fundamental to statistical inference, yet existing methods for non-Euclidean random objects often face a difficult trade-off between geometric flexibility and theoretical tractability. We introduce the Distance Profile Embedding (DPE), a novel representation that maps random objects from general metric spaces into a Hilbert space of square-integrable functions. We prove that this mapping is injective and preserves full distributional information without requiring isometric Hilbert embeddings or one-to-one correspondence conditions. Leveraging the DPE, we develop a unified framework for marginal and conditional independence testing of random objects that enjoys a rigorous asymptotic theory for both size and power. Notably, our framework is the first in the literature to accommodate object-valued conditioning variables when testing conditional independence, overcoming the Euclidean or Hilbertian constraints of existing methodologies. We facilitate the calculation of analytic $p$-values using closed-form asymptotic null distributions, which avoids the computational burden of permutation tests common in existing metric-based methods. The numerical properties of our methods are demonstrated through both simulations and two real-world applications involving gut microbiome compositions and global human mortality distributions, respectively.

W. Tan, Bing Li, Lingzhou Xue · 0 citations
Preprint Jul 2026

Weak Information Geometry: Riemannian Structures from Distributional Inference Functions and Stein Discrepancies

The class of parametric statistical models that can be treated as Riemannian manifolds is considerably larger than the classical Fisher-Rao setting allows, once one works in the space of tempered distributions. A law is represented by a tempered distribution T in S'(R^k), while an instrument - a positive Schwartz kernel, a weak regular inference function, or a weak Stein representation - extracts information from the law without being part of it. Any instrument with full-rank sensitivity and positive-definite variability induces the Godambe information G = S^T V^{-1} S, a Riemannian metric on the parameter space; the Fisher-Rao manifold is recovered exactly when the score is an admissible instrument, and every Godambe metric is dominated by the Fisher metric in the Loewner order whenever the latter exists. Four examples lie outside the Fisher-Rao class for four different reasons: a location model built on the Cantor distribution (an undominated family - no likelihood, no score, and no Fisher information exist at all), the uniform scale model (parameter-dependent support), the shifted exponential model (transform-based inference), and a stratified finite mixture (a provably biased score in a dominated model); a lattice stochastic heat equation driven by alpha-stable noise provides a fifth, dynamical example, whose closed-form weak Godambe information stabilises at a rate governed by the spectral gap of the discrete Laplacian. Quadratic Stein discrepancies induce the same local geometry, and reproducing-kernel constructions generate a hierarchy of geometries. Because there is no canonical instrument, the model carries a family of Godambe metrics; we discuss the inferential, diagnostic, geometric, and computational roles of its members, and show that weak inferential separation (nonformation) appears geometrically as block-diagonality of the Godambe metric.

R. Labouriau · 1 citation
Preprint Aug 2026

The Sampling Distribution of the Log-Euclidean Distance Between Sample Correlation Matrices

Comparing correlation matrices across time or stress scenarios is critical in quantitative finance and multivariate statistics, yet sample estimation noise often obscures whether an observed distance reflects a true structural shift. We derive the asymptotic sampling distribution of the intrinsic off-log (log-Euclidean) distance between two independently estimated full-rank correlation matrices under the null hypothesis that their population correlation matrices coincide. Under general sampling with finite fourth moments, the scaled squared distance converges to a weighted sum of independent $\chi_1^2$ variables, with weights determined by the asymptotic covariance of the Generalized Fisher Transformation (GFT) coordinates. Under Gaussian sampling at independence, this simplifies to a parameter-free $4\chi_d^2$ law. To calibrate tail probabilities, we provide closed-form cumulant generating functions, Lugannani--Rice saddlepoint quantiles, and an explicit Chernoff envelope requiring no root-finding. The first moment of the limiting law establishes a simple rule of thumb for the baseline expected distance under the null hypothesis ($\operatorname E[d_{\mathrm{LE}}] \lesssim 2\sqrt{d/n}$ near independence), quantifying the average separation induced strictly by estimation error. We establish plug-in consistency, present an explicit Gaussian covariance factorization, compare the distance statistic with coordinate Wald tests, and characterize its local power.

A. Kuketayev · 0 citations
Aug 2026

Refining the Convergence Rate for $\phi$-Divergence Statistics in Multinomial Goodness-of-Fit Tests

We study asymptotic properties of statistics in goodness-of-fit tests under the simple null hypothesis of a multinomial distribution. It is also assumed that the statistics are based on $\phi$-divergences, the sample size $n$ tends to infinity, and the number of cells in the group $k$ is fixed. It is known that in this case the limiting law is the chi-square distribution, and according to the classical theory, the distribution function of the test statistic converges to the limiting one at rate $O(n^{-1/2})$. A distinctive feature of the present paper is the use of an approach in which the original problem is reduced to a well-known problem in number theory, namely, to the generalized Gauss problem on the number of “integer” points in an expanding convex set with smooth boundary. This makes it possible to refine the order of convergence to $ O(n^{-1+\alpha(k)})$, where $\alpha(k) > 0$ and $\alpha(k)$ decreases to zero with increasing $k$. Thus, a result previously known only for the family of power divergence statistics is extended to a wider class of statistics based on $\phi$-divergences.

V. Ulyanov, A. V. Sheryaev · 0 citations
Preprint Aug 2026

Exact Likelihood and Sampling for Riemannian Gaussian Distributions on Correlation Matrices

Correlation matrices arise when marginal scales are removed from covariance matrices, yet a normalized likelihood must account for both quotient distance and quotient volume. We propose a Riemannian Gaussian model for full-rank correlation matrices under quotient-affine geometry. The distribution is proper and has finite radial moments. We derive exact score and profiled-scale equations and recover Fisher-transformed Gaussian inference for two-dimensional matrices. In higher dimension, a curvature calculation shows that the normalizing constant can vary with the center. Exact maximum likelihood and Fr\'{e}chet estimation may therefore have different population targets. We develop chart-based methods for evaluating the normalizer, fitting the likelihood, and sampling. Numerical studies verify the analytic case and compare integration, estimation, and sampling procedures across dimensions and dispersion regimes. A rolling-finance application and a controlled prior study illustrate both the value and computational cost of the model. The method is most reliable in small to moderate dimensions, while proposal efficiency and numerical conditioning deteriorate near the boundary and at larger dispersion.

Kisung You · 0 citations
Preprint Jul 2026

Projective Maximum Entropy: Universality and Acceptance-Region Calibration

Maximum-entropy reference distributions are usually constructed on the normalized probability simplex. This formulation is less natural for unnormalized statistical models, in which positive multiples represent the same shape, and it does not directly explain how a prescribed admissible region should determine the deformation parameter of a bounded-support reference distribution. We formulate maximum entropy on the projective space of nonnegative measures and establish three results of statistical relevance. First, a universality theorem shows that every admissible monotone transform of the same normalized power functional has exactly the same optimizer under linear moment constraints. The result unifies the maximum-entropy implications of Tsallis and R\'enyi entropies, H\"older composite scores, pseudo-spherical scores, Bregman--H\"older constructions, and related homogeneous divergences without asserting a new distribution family. Second, the common optimizer is characterized as a $q$-exponential density; under mean and covariance constraints it is a compactly supported $q$-Gaussian for positive deformation and a Student-type density for negative deformation. Third, a prescribed Mahalanobis acceptance region with squared radius $R^2>d+2$ uniquely determines the deformation parameter $\gamma_R=2/(R^2-d-2)$. The resulting affine-equivariant reference density is the unique projective maximum-entropy solution, and its support coincides with the specified ellipsoid without an additional support constraint. This provides a principled method for constructing bounded-support statistical reference distributions from robust location and scatter estimates or from externally specified admissible regions.

H. Hino · 0 citations