A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations natu...
Weiliang Chen, Haowen Sun, Jun Gao et al.· 2 citations
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at th...
Yi-Xin Zheng, Jiangran Lyu, Yun-Tian Deng et al.· 0 citations
GIF, an agentic Generation framework for Interactive and Functional object compositions, recast this problem as disentangled reconstruction followed by relative pose recovery, revealing diversity scaling in both simulation and real-world deployment.
Long-Ji Xu, Zhi-Qi Zhang, Mi Yan et al.· 1 citation· ⚡1
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and gove...
Han-Yang Wang, Yiyang Cai, Weiliang Chen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.