Skip to content
Preprint

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

Manifold-Constrained Hyper-Connections (mHC) is introduced, reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix, highlighting its effectiveness for robust speaker representation learning.

Abstract

Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle path, constraining representation capacity. We introduce Manifold-Constrained Hyper-Connections (mHC), reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix. By employing Sinkhorn-Knopp iterations, mHC ensures energy conservation by preserving signal intensity and feature mean, which stabi- lizes gradients and mitigates signal degradation in complex net- works. We evaluate mHC by replacing standard residual con- nections in backbones including ECAPA-TDNN, ResNet-34, Res2Net, and E-Res2Net. Extensive experiments on VoxCeleb1 demonstrate that mHC connections consistently enhance per- formance across all architectures, highlighting its effectiveness for robust speaker representation learning.

View source

Similar papers

Preprint Aug 2026

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.

A. Shukla, R. Thakur, Aryan Das et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations a...

G. Botté, Séverin Baroudi, Samir Sadok et al. · 0 citations
Sep 2026

Self-supervised audio representation learning model based on time-frequency decoupling and masked reconstruction

A self-supervised model that integrates a time-frequency decoupled (TF-D) stem with masked latent reconstruction with the strongest overall transfer results among the internal variants and a competitive balance among representation quality, encoder scale, and inference efficiency is proposed.

Jie Xu, Yu-Hao Dai, Zhi-Feng Wang · 0 citations
Preprint Aug 2026

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

On near-domain benchmarks, static and dynamic approaches perform comparably, and on harder cross-corpus benchmarks with diverse synthesis methods, trajectory dynamics provide substantial gains, confirming that temporal physiological constraints carry detection signal beyond utterance-level statistics.

T. Weber · 0 citations
Preprint Sep 2026

Domain-Incremental Learning for Multi-Channel Replay Speech Detection

Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordin...

Michael Neri · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.