ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update, achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods.
Shi-Cheng Fang, Yi-Wen Zhao, Wen-Bo Tian et al.
· 0 citations