Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...
Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al.· 0 citations
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat''when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reaso...
Yi-Jun Yang, Shenghe Zheng, Wenbo Li et al.· 0 citations
This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.
Xiong-Hao Wu, Yi-Jun Yang, Shi-Long Zhou et al.· 2 citations
The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.
Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.