ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
The prior requires no target-policy rollout before the first selection, but the posterior thereafter uses target-policy outcomes, and the measured ThinkPrior+DAPO composition reduces generated rollouts by 10.6% while both arms retain the same 3840-rollout update budget.
Tommy Sha, Skylar Zhai, Si-Qi Zhao
· 1 citation
· ⚡1