Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the do...
Yu-Xi Liu, Hao-Yu Li, Ze-Kun Zhang et al.· 0 citations
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and o...
Lianghua Huang, Zhigang Wu, Yupeng Shi et al.· arXiv.org· 2 citations
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse low-rank hybrids alleviate this cost by combining a local sparse branch with a global compressed br...
Ze-Kun Zhang, Yi-Xiang Cai, Yu-Xi Liu et al.· 1 citation
Wan-Streamer v0.2 is presented, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model, which raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-side signal-to-signal latency at 25 FPS.
Lianghua Huang, Zhigang Wu, Yupeng Shi et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.