Skip to content

Author

Haiyun Guo

We have 3 of 71 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

This work proposes PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition, which alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition.

Ruiming Liang, Yinjie Zhong, Yizhen Yuan et al. · 1 citation
Preprint Aug 2026

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation and improves over matched vanilla OPSD reruns on every benchmark at all three model scales.

Zhi-Yan Hou, Xinyu Tang, Hongyan An et al. · 1 citation
Review Aug 2026

Continual Learning in Transition

Anchored by this tri-axial framework, representative methods are systematically surveyed, the ongoing transition of continual learning is traced, and the key challenges, broader implications, and future directions arising from this paradigm shift are discussed.

Zhi-Yan Hou, Dan Zhang, Tao Feng et al. · 0 citations