Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear...
Qiwen Gu, Jifan Li, Bing-Jie Gao et al.· 0 citations
High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \...
Qiwen Gu, Bingjie Gao, Rui Chen et al.· 0 citations
Inspired by render-based compression, this work renders textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning, and concludes that discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.
Shuochen Chang, Qingyang Liu, Shaobo Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.