Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...
Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al.· 0 citations
G-Frame, an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training and synthesizes a specialized corpus of 363,045 chains-of-thought and 199,589 question-answer pairs.
Runzhe Liu, Biquan Bie, Zi-Hao Wang et al.· arXiv.org· 0 citations
This work introduces principled criteria for desirable VQ behavior and demonstrates that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse, and instantiate this framework using a Wasserstein-based objective with an efficient closed-for...