This study suggests that LLM agents can provide practical value for tabular ML by expanding the design space, and suggests that the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models.
Abstract
Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a concrete use case: whether state-of-the-art agentic AI systems can design extended HPO search spaces for established tabular models that outperform the standard search spaces provided by the model authors. Specifically, we represent each tabular model as a modular pipeline covering preprocessing, embeddings, architecture, training, and inference. We then task the agent to propose candidate code implementations for each module and use a classical HPO algorithm to jointly optimize over these candidates and the model's default hyperparameters. Compared with the base HPO spaces, the expanded search spaces improve the performance of nearly every model family across a suite of 45 datasets, with average relative gains of 0.6%, rising to 2.0% on small-to-medium regression datasets. Notably, these gains come at no extra tuning cost: the enlarged spaces outperform the base under the same tuning and ensembling budgets. The gains transfer to the recent TabArena benchmark, where the agentic spaces improve the official Elo scores of four of the five model families and the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models. Overall, our study suggests that LLM agents can provide practical value for tabular ML by expanding the design space.
Across several reasoning tasks and model families, the proposed frontier learning approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Robin Faro, S. Ramesh, Ilija Bogunovic et al.· 0 citations
Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations, achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage.
Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer et al.· 2 citations
The introduction of AutoData, an agent that searches directly over executable selection algorithms, suggests that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
Yan Meng, Dhruv Srikanth, Bingchen Zhao et al.· 2 citations
DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Hao Liang, Ming-Rui Chen, Hengyi Feng et al.· 1 citation
Language models are increasingly used as proposal models for black-box search, from program optimization to molecular design. Existing approaches typically improve proposals through iterative prompting or parameter updates, offering limited control over how completely and efficiently the model's search space is explore...
Dhruv Agarwal, Rico Angell, Kavitha Srinivas et al.· 0 citations
DeltaML-Bench is introduced, a benchmark comprising 48 tasks sourced from research papers that require agents to improve published baselines within imperfect, open-source repositories that indicates that scaffolding design and integrity checks are important considerations when deploying agents for autonomous ML experim...
Josias Moukpe, Priyanka Aryal, M. Kenney· 2 citations· ⚡1
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026