A novel framework, Active Learning via Subspace Ensembles and Similarity (ALSES), which effectively isolates informative variables and achieves superior classification accuracy with significantly fewer labeled instances, demonstrating its robustness in complex, noisy applications.
Abstract
Supervised learning in fields such as genomics and medical imaging is often hindered by the high cost of expert data annotation. Active learning addresses this bottleneck by iteratively selecting the most informative unlabeled samples for labeling. However, in high-dimensional environments, traditional diversity-based query strategies lose their effectiveness due to the degradation of global distance metrics. To address these challenges, this paper proposes a novel framework, Active Learning via Subspace Ensembles and Similarity (ALSES). Instead of relying on global distances, ALSES constructs a similarity matrix by sampling an ensemble of random feature subspaces. The subspaces are filtered based on their discriminative power, and pairwise sample similarities are aggregated using cluster co-occurrence. This structural representation is integrated into a hybrid batch selection strategy that balances model uncertainty and data representativeness. Extensive evaluations on simulated datasets and real-world high-dimensional cancer cohorts demonstrate that ALSES consistently outperforms standard active learning baselines. The framework effectively isolates informative variables and achieves superior classification accuracy with significantly fewer labeled instances, demonstrating its robustness in complex, noisy applications.
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
An Adaptive Ensemble Learning (AEL) framework that addresses this small-n-large-p regime through three coupled mechanisms, and was evaluated on six benchmark high-dimensional datasets containing between 617 and 12,600 features.
Porwal Rabins· International Journal of Inn...· 0 citations
This work proposes Cooperative Learning with a penalized Linear Mixed Model (CL-pLMM) for high-dimensional multiview data with a clustered structure and shows that its objective function can be represented as a penalized linear mixed-effects model applied to augmented data, allowing existing estimation procedures to be...
A flexible encoder based on low-rank embedding that adaptively captures the most informative components in the feature space of the input data while filtering out redundant or noisy information, effectively alleviating over-connected affinities is proposed.
Li Guo, Qian Wang· Journal of King Saud Univers...· 0 citations
A systematic cross-domain study of KLPCDA reveals consistent patterns in the interaction of the three core objectives variance, between-class, and within-class terms, providing a unified and interpretable understanding of their roles in stabilizing representations and enhancing discrimination in SSS settings.
Ling-Xiao Qu, Yan Pei· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.