AffordanceWAM is introduced, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World, and supports affordance as an effective interface for both vision-language-action learning and huma...
Jia-Di You, Qi-Ze Yu, Yue Chen et al.· 0 citations
Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instructi...
Xin Shen, Chengyou Jia, Ke Xing et al.· 1 citation
The predictive divergence mask is proposed, which asks whether the next policy-gradient step will increase or decrease the same divergence used by the trust region, and the resulting masks improve RL training across model scales and precision settings.