Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We at...
Yuqun Wu, Yao Xiao, Chu-Hang Zou et al.· 0 citations
Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance segmentation is limi...
Jiang-Shan Gong, Yuqun Wu, Qiqian Fu et al.· 1 citation
We present FIRE3D, a unified framework that takes a single RGB image or casual RGB video and transforms it into simulation-ready 3D scene assets for games and interactive applications in under a minute. At the core of FIRE3D is a feed-forward, end-to-end network that predicts a compositional scene representation from p...
Hongchi Xia, Tianhang Cheng, Wei-Chiu Ma et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.