Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly s...
Rui Liu, Bhavin Jawade, Hao-Qi Li et al.· 0 citations
This work proposes AuEmoChat, a CSS framework for authentic emotion understanding and rendering, and develops AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion ca...
Zhenqi Jia, Yuan Zhao, Aruukhan et al.· arXiv.org· 1 citation
This work proposes FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone, and introduces a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's...
Yifan Hu, Shuwei He, Rui Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.