EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning
Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching.
Nikita Khomich, L. Hermansson, Ido Hakimi
· 0 citations