LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call this failure Generative Context Mis-anchor...
Shi-Bo Wang, Zicheng Zhang, Libo Wang et al.· 0 citations
Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capaci...
Hengyuan Zhang, Jing-Na Sun, Mei-Guang Jin et al.· arXiv.org· 1 citation
Results show that \method preserves stable appearance across prompt-conditioned segments and strong audio-visual synchronization under autoregressive generation, accelerating autoregressive inference without pipeline-specific retraining.
Qijun Gan, Chen-Wei Zhang, Mei-Guang Jin et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.