Skip to content
Preprint

When Test-Time Adaptation Helps, Harms, or Becomes Inactive: A Condition-Level Study on CIFAR-10-C

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

Results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive, and suggest that reliability filtering can substantially restrict effective adaptation.

Abstract

Test-time adaptation (TTA) aims to improve model robustness under distribution shift by adapting a source model using unlabeled test data. Although methods such as TENT and EATA have demonstrated gains on corrupted data, aggregate accuracy can obscure the conditions under which adaptation fails or provides little benefit. We present a controlled comparison of three TTA strategies---BatchNorm-statistics adaptation (BN-Adapt), entropy-minimization adaptation (TENT), and reliability-filtered adaptation (a scoped re-implementation of EATA)---against an unadapted source model on the full CIFAR-10-C benchmark, covering 15 corruption types and 5 severity levels. All three methods improve mean accuracy over the source model by 12.2--13.3 percentage points (Wilcoxon signed-rank $p<10^{-12}$). However, each method underperforms the source model on 8.0--9.3\% of conditions, with failures concentrated in low-severity corruptions where the source model already performs near ceiling, particularly brightness, fog, contrast, and defocus blur. We further find that EATA closely tracks the gradient-free BN-Adapt baseline, with a mean absolute difference of 0.09 percentage points, compared with 1.08 percentage points relative to TENT. This suggests that reliability filtering can substantially restrict effective adaptation, causing EATA to behave more like a BatchNorm-statistics baseline than an entropy-minimization method. These results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive.

View source

Similar papers

Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Kartik Jhawar, Li-Po Wang · 1 citation
#machine learning Preprint Aug 2026

Continual Test-Time Adaptation via Entropy Sensitivity-Guidance in Strict Online Setting

Sensitivity-Guided Erasing Adaptation (SEGA) is introduced, a method for strict online continual TTA (CTTA) on corruption-style streams that yields consistent robustness and stability gains over strong CTTA baselines while reducing backward passes through sensitivity-based gating.

Chandler Timm C. Doloriel, Yun-Bei Zhang, M. Siddiqui et al. · 0 citations
Preprint Sep 2026

Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation

Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target ob...

You-Jia Zhang, Hui-Ling Liu, Soyun Choi et al. · 0 citations
Open access Sep 2026

AdaTent++: reliability-guided hybrid test-time adaptation for traffic sign detection under distribution shift

Test-time adaptation (TTA) can improve detector robustness under distribution shift, but indiscriminate online updating may amplify errors from unreliable frames. We present AdaTent++, a reliability-guided hybrid TTA framework for traffic sign detection. For each unlabeled frame, the framework combines predictive entro...

Xiao-Mei Gai, Mei-Chun Wang, Yan-Tong Guo et al. · 0 citations
#machine learning Preprint Sep 2026

Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs

This work systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length, finding that sentiment steering survives quantization intact.

Saurav Bhandari, Benjamin Wade · 0 citations
Preprint Sep 2026

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-ven...

Li-Li Wang, Jing Li, Xiao-Wen Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.