We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The...
Hong-Yuan Tao, Xing-Gang Wang, Liang-Hui Zhu et al.· 0 citations
Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocabulary performance through exhaustive vocabulary traversal, yet suffers from prohibitive computational overhead as target categories scale. I...
Hao Peng, Yong-Kang Li, Zhao-Xiang Liu et al.· 0 citations
DreamWAM is introduced, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics, showing that robust world-action learning depends not only on predicting the future, but on representing it in a fo...
Shanglin Yuan, Weiheng Zhao, Xin Shi et al.· 4 citations
Faster-WAM introduces a sparse future-conditioning framework that computes future representations once and selectively reuses them throughout action denoising, and proposes SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fu...
Weiheng Zhao, Haoyi Jiang, Xin Shi et al.· 10 citations
This work reformulates the video diffusion sampling as a frame-indexed stochastic process over noise levels, and constructs a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling.
Yueting Zhu, Yuehao Song, Kaichen Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.