Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel ca...
Eunji Shin, Kyudan Jung, Jihwan Kim et al.· 0 citations
SNAP, a speaker-nulling framework, is introduced that reduces speaker entanglement and encourages detectors to focus on artifact-related patterns, leading to state-of-the-art performance.
Kyudan Jung, Ji-Hoon Kim, Minwoo Lee et al.· arXiv.org· 0 citations
OmniACBench, a benchmark for evaluating context-grounded acoustic control in omni-modal models, is introduced and three common failure modes are identified-weak direct control, failed implicit inference, and failed multimodal grounding-providing insights for developing models that can verbalize responses effectively.
Seunghee Kim, B. Park, Kyudan Jung et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.