Skip to content

Author

Xin-Lei Yu

We have 6 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer caus...

Haojie Huang, Xin-Lei Yu, Cheng-Ming Xu et al. · 2 citations
#machine learning Preprint Sep 2026

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students...

Shuai Dong, Yong-Fu Zhu, Yu-Qi Xu et al. · 0 citations
Preprint Jul 2026

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propos...

Tengfei Liu, Yang Shi, Yuran Wang et al. · 0 citations
Jul 2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...

Qixun Wang, Yang Shi, Le-Tian Cheng et al. · 0 citations
Preprint Jul 2026

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.

Hongyu Qu, Jianzhe Gao, Xiaobin Hu et al. · 1 citation
Preprint Aug 2026

Latent On-Policy Self-Distillation

This work introduces Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience.

Gui-Min Zhang, Jiayang Lyu, Ran Sun et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.