Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning
It is argued that the right resolution is state-dependent, and GACA, a critic-free estimator whose granularity follows an uncertainty-based criticality proxy is proposed, improves task success over GRPO and GiGPO at both 1.5B and 7B scales.