Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge...
Gui-Biao Liao, M. Xiang, Heng Li et al.· 0 citations
Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous text...
Meng-Huan Zhang, Qing Cai, Fan Zhang et al.· IEEE Transactions on Image P...· 0 citations
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In pra...
This paper introduces a novel conjunctive multi-attribute verification mechanism to explicitly combat attribute under-representation and significantly outperforms existing methods and establishes a new state-of-the-art on challenging FG-OVD benchmarks, demonstrating a more robust approach to compositional visual reason...
Jiaming Li, Zhijia Liang, Shuangyin Liu et al.· International Journal of Com...· 0 citations
MoRoute is introduced, a unified multimodal video generation framework that formulates a frozen VLM and a pretrained video DiT with different architectures as heterogeneous experts connected through dynamic layer routing.
Chong Gao, Jie Ma, Zhan Peng et al.· arXiv.org· 0 citations
UPS-GRPO is developed, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency and introduces a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improve...
Keyang Zhong, Kuo Wang, Peng Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.