Skip to content

Author

Guanbin Li

We have 6 of 201 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images

Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge...

Gui-Biao Liao, M. Xiang, Heng Li et al. · 0 citations
Sep 2026

Mitigating Textual Noise in Multimodal FGVC via Hierarchical Semantic Purification and Multi-Stage Alignment

Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous text...

Meng-Huan Zhang, Qing Cai, Fan Zhang et al. · 0 citations
Jul 2026

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In pra...

Zijie Wang, Wei Zhang, Wei-Ming Zhang et al. · 0 citations
Aug 2026

GUIDED++: Enhancing Discrimination with Conjunctive Verification for Fine-Grained Open-Vocabulary Object Detection

This paper introduces a novel conjunctive multi-attribute verification mechanism to explicitly combat attribute under-representation and significantly outperforms existing methods and establishes a new state-of-the-art on challenging FG-OVD benchmarks, demonstrating a more robust approach to compositional visual reason...

Jiaming Li, Zhijia Liang, Shuangyin Liu et al. · 0 citations
Jul 2026

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

MoRoute is introduced, a unified multimodal video generation framework that formulates a frozen VLM and a pretrained video DiT with different architectures as heterogeneous experts connected through dynamic layer routing.

Chong Gao, Jie Ma, Zhan Peng et al. · 0 citations
Preprint Aug 2026

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

UPS-GRPO is developed, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency and introduces a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improve...

Keyang Zhong, Kuo Wang, Peng Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.