Results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive, and suggest that reliability filtering can substantially restrict effective adaptation.
Abstract
Test-time adaptation (TTA) aims to improve model robustness under distribution shift by adapting a source model using unlabeled test data. Although methods such as TENT and EATA have demonstrated gains on corrupted data, aggregate accuracy can obscure the conditions under which adaptation fails or provides little benefit. We present a controlled comparison of three TTA strategies---BatchNorm-statistics adaptation (BN-Adapt), entropy-minimization adaptation (TENT), and reliability-filtered adaptation (a scoped re-implementation of EATA)---against an unadapted source model on the full CIFAR-10-C benchmark, covering 15 corruption types and 5 severity levels. All three methods improve mean accuracy over the source model by 12.2--13.3 percentage points (Wilcoxon signed-rank $p<10^{-12}$). However, each method underperforms the source model on 8.0--9.3\% of conditions, with failures concentrated in low-severity corruptions where the source model already performs near ceiling, particularly brightness, fog, contrast, and defocus blur. We further find that EATA closely tracks the gradient-free BN-Adapt baseline, with a mean absolute difference of 0.09 percentage points, compared with 1.08 percentage points relative to TENT. This suggests that reliability filtering can substantially restrict effective adaptation, causing EATA to behave more like a BatchNorm-statistics baseline than an entropy-minimization method. These results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive.
The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.
Sensitivity-Guided Erasing Adaptation (SEGA) is introduced, a method for strict online continual TTA (CTTA) on corruption-style streams that yields consistent robustness and stability gains over strong CTTA baselines while reducing backward passes through sensitivity-based gating.
Chandler Timm C. Doloriel, Yun-Bei Zhang, M. Siddiqui et al.· 0 citations
Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target ob...
You-Jia Zhang, Hui-Ling Liu, Soyun Choi et al.· 0 citations
Test-time adaptation (TTA) can improve detector robustness under distribution shift, but indiscriminate online updating may amplify errors from unreliable frames. We present AdaTent++, a reliability-guided hybrid TTA framework for traffic sign detection. For each unlabeled frame, the framework combines predictive entro...
Xiao-Mei Gai, Mei-Chun Wang, Yan-Tong Guo et al.· Frontiers in Signal Processi...· 0 citations
This work systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length, finding that sentiment steering survives quantization intact.
Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-ven...
Li-Li Wang, Jing Li, Xiao-Wen Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.