Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...
Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al.· 1 citation
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the do...
Yu-Xi Liu, Hao-Yu Li, Ze-Kun Zhang et al.· 0 citations
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse low-rank hybrids alleviate this cost by combining a local sparse branch with a global compressed br...
Ze-Kun Zhang, Yi-Xiang Cai, Yu-Xi Liu et al.· 1 citation
This paper presents a proactive defense framework for securing LLMs against evolving multi-turn adversarial attacks that combines disruption, misdirection, and adaptation across successive interaction turns and employs a cooperative multi-agent architecture.
Si-Yuan Li, Zehao Liu, Hao-Yu Li et al.· 0 citations
ZenGen, an integrated framework for measuring, internalizing, and grounding social intelligence, and Actio, a harness-controlled inference architecture that routes four typed supports into reasoning demonstrate the effectiveness of typed runtime support.
ZenGen Team, Xiang Ao, Jingping Bi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.