Preprint
Jul 2026
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
LatentRM is a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards through on-policy optimization of the latent reasoning space end-to-end.
Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang et al.
· 0 citations