Skip to content

Bracketing Uncertainty in Clustering Under the Manifold Hypothesis

Sep 2026 · 1 citation · 39 references
Mathematics Computer Science Biology

TL;DR

Manifold-Based Clustering (MBC) is proposed, which returns an explicit bracket interval to quantify the underlying data uncertainty, and suggests that ambiguity in cluster number is often intrinsic, and should be quantified rather than resolved.

Abstract

The manifold hypothesis suggests a natural criterion for clustering: partition data according to the manifold component from which each point is drawn. Whether two components are separable depends on a geometric tradeoff: the ambient separation between components versus the largest gap in sampling. In practice, this tradeoff is rarely assessed explicitly, leading standard methods to over-commit to a single clustering assignment even when the data do not support a unique answer. We formalize this tradeoff by combining intrinsic manifold geometry (volume growth and reach) with sample-level quantities (fill distance and density), yielding a threshold phenomenon for mutual-$k$-nearest-neighbor graphs: when the offset-to-fill ratio exceeds a conservative upper threshold, component separation is preserved; below a lower threshold, components fuse. The gap between these thresholds defines a geometric uncertainty zone in which the number of clusters is not identifiable from the data. Nevertheless, conventional approaches still seek one: sweeping parameters (an engineering approach) or fitting a generative mixture model (a model-based approach). Rather than forcing a single estimate of the number of clusters, we propose Manifold-Based Clustering (MBC), which returns an explicit bracket interval to quantify the underlying data uncertainty. This bracket acts as an empirically calibrated diagnostic: it narrows when a single resolution is supported, widens when multiple resolutions coexist, and collapses to one when no separated structure is detectable. Empirically, we find that many real datasets lie within the uncertainty zone rather than admitting one clear answer. Our results suggest that ambiguity in cluster number is often intrinsic, and should be quantified rather than resolved.

View source

Similar papers

#machine learning Preprint Sep 2026

Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle

DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimension (ID) benchmarks. Practical data, however, introduce neighborhood-relative noise and sample-amplitude heterogeneity...

Chih-Hsuan Huang, Chih-Wei Chen, Szu-Chi Chung · 0 citations
#machine learning Preprint Sep 2026

Riemannian Difference-of-Convex Optimization for K-Means Clustering

K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a difference-of-convex (DC) penalty and establish a global err...

Meng Xu, Bo Jiang, Han-Fu Zhang et al. · 0 citations
Preprint Aug 2026

Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms

Numerical experiments show that GW quantization opens up many modeling possibilities beyond normal clustering methods and that the introduced algorithm leads to useful numerical solutions with approximation quality often in line with theoretically optimal rates.

F. Beier, S. Eckstein · 0 citations
Preprint Sep 2026

Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters that preserves the distributional information distinguishing the clusters.

S. Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Unimodal Bases: Pullback Geometry for Multimodal Data

Data-driven Riemannian geometry provides nonlinear interpolation and geometric representations of high-dimensional data. For these operations to be statistically meaningful, paths between observations should preferentially traverse high-likelihood regions. Existing scalable pullback constructions typically use a unimod...

Honglei Brinkmann, Lucas Ng, Georgios Batzolis et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.