Long-form text-to-speech must retain requested words over extended generations, yet sentence-level training and evaluation can hide omissions and early stops. We introduce Balalaika-Longform, an open Russian corpus of 189 hours in continuous units of 30 seconds to 15 minutes. Long units and matched short windows suppor...
Nikita Vasiliev, Kirill Borodin, V. Kudryavtsev et al.· 0 citations
Speech anti-spoofing countermeasures degrade when the generator, codec or channel changes, and a common remedy is an auxiliary objective that shapes the embedding space; whether it does is invisible to EER, a pure ranking metric. We compare seven such objectives with cross-entropy over 113 runs on five corpora, AASIST3...
Ksenia Lysikova, Kirill Borodin, Maxim V. Maslov et al.· 0 citations
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses thi...
Maxim V. Maslov, Kirill Borodin, V. Kudryavtsev et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.