Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the...
Kai-Chen Zhang, Yuzhong Hong, Jun-Wei Bao et al.· 0 citations
A memory-free recent-window protocol is fixed and a transparent and reproducible reference for open-source streaming-video research is provided, combining verifiable streaming-video data, thinking-mode OPD, and instruct-mode deployment.
Keming Wu, Baoyi Wang, Kai-Chen Zhang et al.· 0 citations
Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.
This work reformulates the video diffusion sampling as a frame-indexed stochastic process over noise levels, and constructs a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling.
Yue-Ting Zhu, Yue-Hao Song, Kai-Chen Zhang et al.· 1 citation
The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.