Preprint
Jul 2026
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
This work proposes PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition, which alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition.
Ruiming Liang, Yinjie Zhong, Yizhen Yuan et al.
· 1 citation