AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Results show that long-horizon reflective data is an effective route toward self-improving agents, and synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration.