Skip to content

Author

Yicheng Xiao

We have 5 of 18 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning

Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evi...

Hao-Xuan Ma, Yi-Hao Liu, Yu-Tao Sun et al. · 0 citations
#machine learning Preprint Sep 2026

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from"Rollout Silencing"and low-quality gradient signals in standard sampling procedures. In this work, we propose...

Yi-Meng Ye, Shuang Chen, Wen-Xuan Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Latent Memory Routing for Vision-Language Reasoning

DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.

Hao-Xuan Ma, Jin-Fei Qi, Yi-Cheng Xiao et al. · 0 citations
Preprint Aug 2026

PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle

Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from whole images to specific objects and regions. However, these methods face two fundamental challenges. First, the scarcity of high-quality mask...

Yicheng Xiao, Haoxuan Ma, Caorui Li et al. · 0 citations
Preprint Aug 2026

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.

Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al. · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.