Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-in...
Tuo Liang, Di-Sheng Liu, Neng-Bo Wang et al.· 0 citations
This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier, and synthesizes benchmark design, evaluation protocols, and modeling paradigms based on multimodal alignment, evidence-grounded reasoning, and controlled gener...
Tuo Liang, Zhe Hu, Disheng Liu et al.· arXiv.org· 0 citations
This survey provides a comprehensive and unified overview of recent advances in spatial intelligence for VLMs, summarize core concepts behind spatial reasoning in VLMs, analyze why spatial failures occur, and organize existing solutions into a clear framework spanning prompting-based techniques, model improvements, exp...
Disheng Liu, Tuo Liang, Zhe Hu et al.· Artificial Intelligence Revi...· 7 citations
Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data, is proposed, a training-free framework that recovers a frozen VLA at inference time without policy fine-tuning or failure-specific recovery training.
Yan-Yan Zhang, Disheng Liu, Kai Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.