Skip to content

Author

Haoxuan Ma

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning

Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evi...

Hao-Xuan Ma, Yi-Hao Liu, Yu-Tao Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Latent Memory Routing for Vision-Language Reasoning

DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.

Hao-Xuan Ma, Jin-Fei Qi, Yi-Cheng Xiao et al. · 0 citations
Preprint Aug 2026

PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle

Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from whole images to specific objects and regions. However, these methods face two fundamental challenges. First, the scarcity of high-quality mask...

Yicheng Xiao, Haoxuan Ma, Caorui Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.