ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills...