Skip to content

AutoData: Agentic Search for Pre-training Data Selection

Sep 2026 · 2 citations · 31 references
Computer Science

TL;DR

The introduction of AutoData, an agent that searches directly over executable selection algorithms, suggests that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

Abstract

LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collabo...

Haotian Luo, Hao-Yu Wang, Ze-Yu Qin et al. · 0 citations
#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations
#machine learning Preprint Sep 2026

Agentic Search Spaces for Tabular Machine Learning

This study suggests that LLM agents can provide practical value for tabular ML by expanding the design space, and suggests that the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models.

Renat Sergazinov, Artem Chistyakov, Sergey Pankevich et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TuiML: Machine Learning for AI Agents

TuiML is a self-contained machine-learning library built for AI agents, with native algorithms across supervised, unsupervised, time-series, data handling, tuning, and evaluation tasks, and Benchmarks show TuiML remains predictively competitive with scikit-learn and Weka.

Nilesh Verma, N. Lim, Albert Bifet et al. · 0 citations
#natural language process... Preprint Sep 2026

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate...

Rui-Feng Yuan, Yi-Zhi Li, Ya-Xin Du et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...

Hao-Ran Zhao, Wei Du, Dingwen Yang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.