Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's langua...
Jun-Jie Li, Zheng Liang, Zhe Li et al.· 0 citations
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD un...
Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget and apply latent adversarial post-training inspired by Flow2GAN for few-step generation.
Manifold-Constrained Hyper-Connections (mHC) is introduced, reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix, highlighting its effectiveness for robust speaker representation learning.
Zezhong Jin, Xiaoyu Wang, Zhe Li et al.· 0 citations
Experiments on the MER2026-EmoPrefer Challenge dataset and the error-augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence.
Zilong Huang, Junyi Peng, Junjie Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.