Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and phys...
Chao-Qian Mu, Wen-Hao Wu, Zi-Chen Liang et al.· 0 citations
Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of gene...
Jun-Lan Xiao, Jun-Wei Jiang, Zai-Bin Zhang et al.· 0 citations
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedur...
Zhong-Bo Zhang, Jia-Yi Jin, Yi-Fan Wang et al.· 0 citations
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors....
Zaibin Zhang, Jun-Lan Xiao, Zhong-Bo Zhang et al.· 3 citations
GROVE is introduced, a training-free framework that supports both behaviors with one memory grown causally from a continuous video stream, and achieves the best results among the compared methods.
Sitong Gong, Caixin Kang, Tianyu Yan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.