Manifold-Constrained Hyper-Connections (mHC) is introduced, reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix, highlighting its effectiveness for robust speaker representation learning.
Abstract
Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle path, constraining representation capacity. We introduce Manifold-Constrained Hyper-Connections (mHC), reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix. By employing Sinkhorn-Knopp iterations, mHC ensures energy conservation by preserving signal intensity and feature mean, which stabi- lizes gradients and mitigates signal degradation in complex net- works. We evaluate mHC by replacing standard residual con- nections in backbones including ECAPA-TDNN, ResNet-34, Res2Net, and E-Res2Net. Extensive experiments on VoxCeleb1 demonstrate that mHC connections consistently enhance per- formance across all architectures, highlighting its effectiveness for robust speaker representation learning.
PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.
A. Shukla, R. Thakur, Aryan Das et al.· 0 citations
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations a...
G. Botté, Séverin Baroudi, Samir Sadok et al.· 0 citations
A self-supervised model that integrates a time-frequency decoupled (TF-D) stem with masked latent reconstruction with the strongest overall transfer results among the internal variants and a competitive balance among representation quality, encoder scale, and inference efficiency is proposed.
Jie Xu, Yu-Hao Dai, Zhi-Feng Wang· Signal, Image and Video Proc...· 0 citations
On near-domain benchmarks, static and dynamic approaches perform comparably, and on harder cross-corpus benchmarks with diverse synthesis methods, trajectory dynamics provide substantial gains, confirming that temporal physiological constraints carry detection signal beyond utterance-level statistics.
Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordin...
Results show that task-informed subspace construction can be effective for low-resource speech adaptation, and that adaptive rank allocation provides a robust way to improve parameter efficiency without increasing model capacity.
Asmee Mishra, Meng-Jie Qian, Brechtje Post et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.