Skip to content

Author

Xiangnan He

We have 4 of 59 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an...

Hou-Cheng Jiang, Bo-Xuan Zhang, Qi-Yong Zhong et al. · 0 citations
Open access Aug 2026

RosePO: Customized Preference Alignment in LLM-Based Recommendation

This work proposes RosePO, a framework to refine LLM-based recommendation through pairwise preference optimization with personalized smoothing, and incorporates a personalized smoothing factor predicted by a user oracle into the optimization objective.

Jia-Yi Liao, Xiang-Nan He, Ruo-Bing Xie et al. · 0 citations
Jul 2026

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

This work proposes ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization, enabling sustained performance improvements over extended training horizons.

Kexin Huang, Junkang Wu, Jinda Lu et al. · 0 citations
Preprint Aug 2026

PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity is proposed, which demonstrates that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base mo...

Jiawei Feng, Jiancan Wu, Xingyu Zhu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.