While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maxi...
Peng Ling, Yingda Yin, Lingting Zhu et al.· 0 citations
A framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement is presented, and observation-consistent supervision is introduced that aligns each target scene with the visual evidence available in its inp...
Kai Li, Lu-Tao Jiang, Zhenyang Li et al.· 1 citation
The key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling and replacing global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the...