Skip to content

Selective Inference for Deep Clustering in Latent Spaces

Sep 2026 · 0 citations
Mathematics Computer Science

TL;DR

An SI framework for deep clustering with a fixed pretrained encoder that provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering and enables valid statistical testing of differences between clusters identified in the latent space.

Abstract

Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.

View source

Similar papers

Open access Aug 2026

Weighted Deep Embedded Clustering for Robust Representation Learning in Noisy Data

A Reliability-Based Deep Embedded Clustering (RDEC) approach that improves clustering reliability through a novel distance-aware weighting strategy, coupled with a novel Kullback–Leibler divergence objective function to focus on the most representative data instances and guide the clustering process.

Meaad Altwaimi, M. B. Ben Ismail, Ouiem Bchir · 0 citations
Open access Sep 2026

Robust deep subspace clustering based on unsupervised fusion learning

A flexible encoder based on low-rank embedding that adaptively captures the most informative components in the feature space of the input data while filtering out redundant or noisy information, effectively alleviating over-connected affinities is proposed.

Li Guo, Qian Wang · 0 citations
Sep 2026

Large-scale structured subspace clustering

A Large-scale Structured Subspace Clustering (LSSC) framework that integrates anchor-based reconstruction, locality-aware coefficient initialization, distance-weighted structure regularization, and coefficient regularization to learn a compact nonnegative sample-to-anchor representation is proposed, thereby providing f...

Mou-An Chen, Xue-Song Yin, Qi Huang et al. · 0 citations
Conference Open access Sep 2026

Salient-Residual Decoupled Multi-View Learning for Clustering

This paper proposes salient-residual decoupled multi-view learning for clustering, SRDMVC, introducing a novel decomposition-fusion iterative optimization, which separates the feature space into a salient space and a residual subspace effectively and fuses them using a novel attention mechanism.

Gao-Kai Wang, Yazhou Ren, Feng-Yu Zhang et al. · 0 citations
Open access Aug 2026

A generative model for dimensionality reduction with millions of features and few samples

It is demonstrated that it is feasible to train a deep generative model for dimensionality reduction with millions of features using few samples, which makes this type of generative model a more versatile alternative to standard methods for dimensionality reduction.

C. Pancotti, P. Fariselli, J. Meisner et al. · 0 citations
Preprint Aug 2026

Convergent Evolution in Neural Representation Space: Emergent Order in Deep Belief Networks

It is shown that layer-wise generative learning can spontaneously uncover and progressively amplify class-related structure in unlabeled data and improve average clustering can coexist with reduced accessibility for a few difficult class pairs.

Patrick Krauss, Achim Schilling, Andreas K. Maier et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.