While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primarily rely on diversity-based selection, discarding similar tokens to maxi...
Peng Ling, Yingda Yin, Lingting Zhu et al.· 0 citations
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In pra...
A framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement is presented, and observation-consistent supervision is introduced that aligns each target scene with the visual evidence available in its inp...
Kai Li, Lu-Tao Jiang, Zhenyang Li et al.· 1 citation
The key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling and replacing global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the...