AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each accepted rewrite beco...
Dhruv Srikanth, Bingchen Zhao, Di-Xing Xu et al.· 1 citation
The introduction of AutoData, an agent that searches directly over executable selection algorithms, suggests that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
Yan Meng, Dhruv Srikanth, Bingchen Zhao et al.· 2 citations
SpecBench is introduced, a benchmark comprising 30 systems-level programming tasks ranging from short horizon tasks like building a JSON parser to ultra long horizon tasks like building an entire OS kernel from scratch, which offers a principled testbed for measuring whether coding agents build genuine working systems...