Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok...
Xiang-Qi Li, Li-Bo Huang, Jia-Rui Zhao et al.· 0 citations
Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragi...
Wei-Lun Feng, Chuan-Guang Yang, Haotong Qin et al.· IEEE Transactions on Pattern...· 3 citations
3D scene generation has rapidly evolved, significantly promoting the innovation of content creation. In this context, interaction techniques serve as a pivotal bridge connecting user intent with the generative models, thereby enabling precise control, real-time feedback and personalized customization of complex 3D scen...
Yuqi Li, Si-Wei Meng, Chuan-Guang Yang et al.· Proceedings of the Thirty-Fi...· 30 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.