Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-sp...
Zhen-Yu Liu, Zhang-Quan Chen, Ke-Yi Chen et al.· 0 citations
ViP-Rig is a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones into a frozen pretrained autoregressive generator.
Zihan Qin, Ming-Ze Sun, Yifan Mao et al.· arXiv.org· 1 citation
D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-po...
Junhao Chen, Mingjin Chen, Henghaofan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.