Results show that long-horizon reflective data is an effective route toward self-improving agents, and synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration.
Hong-Jin Qian, Chao-Fan Li, Kun Luo et al.· 0 citations
DisCo is presented, a skill-powered research agent that creates skills and uses them during research, and yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families.
Jianlyu Chen, Yuyang Hu, Hong-Jin Qian et al.· 1 citation
This work introduces AREX, a family of Recursively Self-Improving (RSI) deep research agents that substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
Shuqi Lu, Chaofan Li, Kun Luo et al.· arXiv.org· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.