Skip to content
Preprint

Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

Sep 2026 · 0 citations · 18 references
Mathematics

TL;DR

The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters that preserves the distributional information distinguishing the clusters.

Abstract

We consider dimensionality reduction for high-dimensional observations accompanied by a supplied partition into two or more clusters. The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters. For each coordinate, the proposed procedure compares the cluster-specific empirical distribution functions through a several-sample Kolmogorov-Smirnov separation statistic. We formalize the resulting marginal cluster support and establish simultaneous finite-sample concentration over all coordinates, explicit bounds for false inclusions and omissions, and exact support recovery when the minimum distributional separation dominates the high-dimensional stochastic error. We also quantify the dimension inflation induced by using an unadjusted testing level and give a familywise-error-controlled version. Under a conditional sufficiency condition, sure screening preserves the full-data posterior cluster probabilities, mutual information, and Bayes risk; an additional result characterizes robustness to imperfectly estimated cluster labels. The procedure is invariant to strictly increasing coordinate transformations and can retain low-variance cluster signals that principal components may discard. We further develop average dual information, a criterion combining partition agreement after transformation with structural coverage of cluster-relevant coordinates, and derive its basic properties and consistency. Simulations illustrate the theory, the interpretability of the selected coordinates, and the distinction between cluster-directed screening and variance-directed projection.

View source

Similar papers

Preprint Sep 2026

High-Dimensional Two-Sample Inference via Marginal Likelihood-Ratio Rank Statistics

We develop a rank-based framework for high-dimensional two-sample testing that detects marginal distributional differences beyond means and variances. Three marginal likelihood-ratio statistics generate SUM tests for widespread differences, MAX tests for concentrated departures, and Cauchy combinations for unknown sign...

Xiao-Xu Zhang, Long Feng · 0 citations
Preprint Aug 2026

Fast high-dimensional mean testing via logistic regression

This work proposes computationally efficient tests for equality of mean vectors of two or more high-dimensional populations by establishing an equivalence between equality of means and a zero population logistic regression parameter.

Sayan Das, Debraj Das, S. Dutta · 0 citations
Preprint Sep 2026

Nonparametric Identification of Latent Dimension under Heavy-Tailed and Mixed-Signals

Principal component analysis and factor analysis are foundational to the study of high-dimensional data, yet their efficacy depends entirely on correctly identifying the latent dimension $r$. While parallel analysis (PA) is widely regarded as the gold standard for this task, its performance suffers in high-dimensional...

Chetkar Jha · 0 citations
Preprint Aug 2026

High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. For observat...

Guoqing Zhang, Zhaixin Chen · 0 citations
Open access Sep 2026

Gap-Entropy Testing (GET): A Framework for Structural and Distributional Differences

The classical two-sample problem is framed as the detection of differences in marginal distributions, in terms of location, scale, or overall shape. In this paper, Gap–Entropy Testing (GET) is proposed, a novel nonparametric framework that extends the two-sample paradigm by adding a spacing-sensitive component to conve...

V. Karalis · 0 citations
Preprint Aug 2026

An Entropy-based Coefficient of Determination with Adjustment of Optimization Bias

Classical likelihood-ratio tests and $\Delta$AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment effect (ATE) quantify practical magnitude, their reliance on the expectation operat...

Long-Xian Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.