Skip to content
Preprint

Validation-Frontier Representation Selection under Constrained Observation

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

Adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.

Abstract

AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state representation when raw accuracy is not the only operational criterion. We propose a validation-frontier selector that combines balanced accuracy with penalties for feature cost, overfit gap, and validation-test instability. In a focused public-tabular benchmark using three scikit-learn datasets, five observation regimes, 45 matched task cells, 720 candidate actions, and 405 representation rows, the adaptive selector improves frontier score over full trace features by 0.025801 while reducing mean feature count by 22.733. Balanced-accuracy difference is small and not statistically significant. A broader offline stress test gives mixed results. The supported claim is therefore bounded: adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.

View source

Similar papers

#machine learning Preprint Aug 2026

PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision, is introduced, showing why predictive fit, decision reliability, and pruning method quality require separate evidence.

Hao Ye, Gao-Peng Zhang · 0 citations
#artificial intelligence Review Aug 2026

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.

Madhusudan Srinivasan, Namith Nishal Raphae · 0 citations
Preprint Aug 2026

ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning

Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable. However, existing methods often infer environments from a separate reference model and select representations before fitting the classifier used at deployment, leaving both decisions misaligned w...

Qianqian Wang, Yun-Shan Li, Dawei Huang et al. · 0 citations
Preprint Aug 2026

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, ev...

Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki · 3 citations
Open access Aug 2026

Accurate and Interpretable Prediction of Exploration Input–Output Matching Under Data Scarcity: An Ensemble Learning Framework

Accurate prediction of input–output relationships in natural gas exploration is essential for improving exploration efficiency and optimizing investment allocation. However, this task is severely hindered by data sparsity and strong nonlinear characteristics inherent in oil and gas exploration systems, rendering conven...

Xiao Chen, Hui Liu, Weiyun Zhan et al. · 0 citations
Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Kartik Jhawar, Li-Po Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.