The introduction of AutoData, an agent that searches directly over executable selection algorithms, suggests that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
Abstract
LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collabo...
Haotian Luo, Hao-Yu Wang, Ze-Yu Qin et al.· 0 citations
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...
Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al.· 0 citations
This study suggests that LLM agents can provide practical value for tabular ML by expanding the design space, and suggests that the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models.
Renat Sergazinov, Artem Chistyakov, Sergey Pankevich et al.· 0 citations
TuiML is a self-contained machine-learning library built for AI agents, with native algorithms across supervised, unsupervised, time-series, data handling, tuning, and evaluation tasks, and Benchmarks show TuiML remains predictively competitive with scikit-learn and Weka.
Nilesh Verma, N. Lim, Albert Bifet et al.· 0 citations
Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate...
Rui-Feng Yuan, Yi-Zhi Li, Ya-Xin Du et al.· 0 citations
Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...
Hao-Ran Zhao, Wei Du, Dingwen Yang et al.· 0 citations