Preprint
Jul 2026
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function across auxiliary tasks before RLHF training, is introduced, providing theoretical guarantees for policy invariance, analyze representation drift sensitivity, and formally address incentive misalignment from entropy maximization.
Yu-An Chu
· 0 citations