Skip to content
Preprint

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

Sep 2026 · 1 citation · 31 references
Computer Science

TL;DR

The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Abstract

Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain enough information to determine the best action at all? We show that this is not guaranteed, even with a perfect selector. If an observation channel makes two deployments look the same while their TTA rankings differ, reliable selection is impossible from that channel; richer evidence can restore the decision only when it resolves the relevant ambiguity. We make this boundary exact in a finite-batch Gaussian TTA model, where doing nothing beats mean recentering for small shifts, recentering wins beyond a unique critical shift, and the boundary shrinks as $1/\sqrt n$. Public benchmark studies on CIFAR-100-C and DomainNet-126 show the same failure mode with modern TTA methods: changing only deployment structure can reverse the oracle action while global order-blind evidence remains unchanged. The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation

Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for every incoming test batch, which can incur substantial annotat...

M. Huzaifa, Lea Schönherr, Thorsten Eisenhofer · 0 citations
Preprint Aug 2026

When Test-Time Adaptation Helps, Harms, or Becomes Inactive: A Condition-Level Study on CIFAR-10-C

Results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive, and suggest that reliability filtering can substantially restrict effective adaptation.

Sreeja Guha Majumdar, Aratrika Saha · 0 citations
Preprint Sep 2026

Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation

Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target ob...

You-Jia Zhang, Hui-Ling Liu, Soyun Choi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

"System One"decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya...

Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay · 1 citation
#artificial intelligence Preprint Sep 2026

Useful to Whom? Sample Value Is Defined Only Relative to the Learner

What kind of data does a model need in order to learn? Coreset selection makes this question concrete: under a budget, keep the samples most useful for training. Easy-first and geometric coverage criteria can win in different budget regimes, separated by a crossover boundary. We ask whether this boundary is fixed by th...

Yang-Ze Liu, Xiao-Long Yin, Zhong-Yi Han · 0 citations
#artificial intelligence Preprint Sep 2026

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the o...

Naveen Vakada, Ming-Yuan Li, Shao-Xiong Ji · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.