Skip to content

Author

Peng Cheng

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

This work proposes MSign, a new optimizer that periodically applies matrix sign operations to restore stable rank in a 5M-parameter NanoGPT model, and proves theoretically that these two conditions jointly cause exponential gradient norm growth with network depth.

Lianhai Ren, Yucheng Ding, Xiao Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.