Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic...
Yu-Xiang Wang, Kun-Yu Feng, Yuan-Cheng Wang et al.· 0 citations
Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and tra...
Wan Lin, Li Wang, Jindong Wang et al.· arXiv.org· 1 citation
RecurTrace introduces Loop Memory Attention, which lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone.
Yu-Xiang Wang, Kun-Yu Feng, Ying-Da Shen et al.· 3 citations
MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.
Kun-Yu Feng, Yue Ma, Bing-Yuan Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.