When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment
UECR-GRPO is introduced, which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels and uses the signed teacher--old-policy token gap to redistribute the verifier-derived component.