Skip to content

Agentic Search Spaces for Tabular Machine Learning

Sep 2026 · 0 citations · 82 references
Computer Science

TL;DR

This study suggests that LLM agents can provide practical value for tabular ML by expanding the design space, and suggests that the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models.

Abstract

Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a concrete use case: whether state-of-the-art agentic AI systems can design extended HPO search spaces for established tabular models that outperform the standard search spaces provided by the model authors. Specifically, we represent each tabular model as a modular pipeline covering preprocessing, embeddings, architecture, training, and inference. We then task the agent to propose candidate code implementations for each module and use a classical HPO algorithm to jointly optimize over these candidates and the model's default hyperparameters. Compared with the base HPO spaces, the expanded search spaces improve the performance of nearly every model family across a suite of 45 datasets, with average relative gains of 0.6%, rising to 2.0% on small-to-medium regression datasets. Notably, these gains come at no extra tuning cost: the enlarged spaces outperform the base under the same tuning and ensembling budgets. The gains transfer to the recent TabArena benchmark, where the agentic spaces improve the official Elo scores of four of the five model families and the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models. Overall, our study suggests that LLM agents can provide practical value for tabular ML by expanding the design space.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Across several reasoning tasks and model families, the proposed frontier learning approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.

Robin Faro, S. Ramesh, Ilija Bogunovic et al. · 0 citations
Preprint Aug 2026

Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations, achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage.

Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer et al. · 2 citations
#artificial intelligence Preprint Sep 2026

AutoData: Agentic Search for Pre-training Data Selection

The introduction of AutoData, an agent that searches directly over executable selection algorithms, suggests that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

Yan Meng, Dhruv Srikanth, Bingchen Zhao et al. · 2 citations
#machine learning Preprint Sep 2026

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Hao Liang, Ming-Rui Chen, Hengyi Feng et al. · 1 citation
#artificial intelligence Preprint Sep 2026

BOReFT: Manifold Steering of Language Models for Black-box Optimization

Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how completely and efficiently the model's search space is explore...

Dhruv Agarwal, Rico Angell, Kavitha Srinivas et al. · 0 citations
Preprint Aug 2026

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories

DeltaML-Bench is introduced, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories that indicates that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experim...

Josias Moukpe, Priyanka Aryal, M. Kenney · 2 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.