Skip to content
Book Open access

PACE: Unleashing the Power of Code Embeddings to Boost AutoML Agents

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 6724-6735 · 0 citations · 25 references

Abstract

Large Language Model (LLM)-driven AutoML agents have shown strong capabilities in constructing end-to-end machine learning pipelines. However, their effectiveness is limited by costly execution-based feedback, which can make the search for high-quality solutions inefficient under restricted computational budgets. We propose PACE (Pre-execution Admission via Code Embeddings), an online-adaptive admission control framework that improves budgeted sample efficiency by estimating candidate utility prior to execution from within-run execution history, without training a separate offline predictor. The core idea of PACE is to leverage latent structure in the solution space as a within-run admission signal. It projects candidate solutions into multi-view semantic embeddings, dynamically organizes executed candidates into clusters, and estimates new candidates by their proximity to historical elite regions. Moreover, PACE aggregates multi-view embeddings via an adaptive reweighting strategy that prioritizes views with higher discriminative power. This enables PACE to bias the agent toward high-potential regions under a limited computational budget while retaining exploration when local structure is weak. We demonstrate that PACE operates as a plug-and-play admission layer for AutoML agents, serving as either an execution gate or a search prior. In the tested settings, it improves the density of elite solutions found within fixed budgets without modifying the underlying agent architecture. Code, configurations, and prompt templates are publicly available at https://github.com/fendss/PACE.

Read PDF

Similar papers

#small language model Book Open access Sep 2026

PROMISE: Process Reward Models for Unlocking Test-Time Scaling Laws in Generative Recommendations

This work proposes Promise, a novel framework that integrates dense, step-by-step verification into generative models, and unlocks Test-Time Scaling Laws in recommender systems, demonstrating that by increasing inference compute, smaller models can match or surpass larger models.

Cheng-Cheng Guo, Kuo Cai, Yu Zhou et al. · 0 citations
#artificial intelligence Review Sep 2026

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leav...

Zheng-Yu Chen, Lin-Feng Liu, Hong Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UNBIND: UNlearning By INference-time Directional Steering for Code LLMs

Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patt...

Zhengyang Shan, Jia-Yu Xin, Yan-Jun Lin et al. · 0 citations
#machine learning Preprint Sep 2026

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

This work uses empirical studies to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment, and proposes a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models.

Qi-Yao Ma, Jun-Shan Zhang, Zhe Zhao · 0 citations
#machine learning Preprint Sep 2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...

Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.