Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evi...
Hao-Xuan Ma, Yi-Hao Liu, Yu-Tao Sun et al.· 0 citations
DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.
Hao-Xuan Ma, Jin-Fei Qi, Yi-Cheng Xiao et al.· 0 citations
Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from whole images to specific objects and regions. However, these methods face two fundamental challenges. First, the scarcity of high-quality mask...
Yicheng Xiao, Haoxuan Ma, Caorui Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.