Audio anti-spoofing systems increasingly combine self-supervised learning, parameter-efficient fine-tuning, and graph-attention-based backends. However, performance gains in such systems are often entangled with concurrent changes in the backbone, fine-tuning strategy, and training protocol, making the independent cont...
Hao-Yu Wang, Jing Yang, Chen-Yu Liu et al.· 0 citations
Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also...
Hong-Yu Jin, Wen-Da Zhang, Run-Qiu Fei et al.· 0 citations
An evidence-grounded generative SE framework that uses a deterministic estimate as imperfect evidence as imperfect evidence is proposed and SNR-Conditioned CSG (SNR-CSG), which maps a calibrated residual-SNR estimate to an utterance-level strength and constructs an adaptive grounded anchor is introduced.
Hao Shi, Yuan Gao, Zhao-Heng Ni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.