Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
A Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor is proposed.