Preprint
Aug 2026
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
Self-Reflective Policy Optimization (SRPO) enables LLMs to analyze their own completed trajectories, synthesize errors into concise"reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals.
Jialong Liu, Yuling Shi, Ning Yang et al.
· 1 citation