Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedG...
Janet Wang, Yun-Bei Zhang, Xiao Wang et al.· 0 citations
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu...
Yun-Bei Zhang, Zi-Jian Jin, Yuan-Zhe Liu et al.· 0 citations
Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorl...
Yuqi Li, Xi Xiao, Yunbei Zhang et al.· arXiv.org· 4 citations
The ability of AI systems to improve their behavior during deployment is becoming increasingly important. As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by exploiting test-time information and additional computation...
Shuai-Cheng Niu, Guo-Hao Chen, Yaofo Chen et al.· 2 citations
Dynamic Hub-and-Spoke Memory is proposed, a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception in streaming video understanding.
Xinru Jiang, Lin Zhao, Xi Xiao et al.· 4 citations
Diversify, Anchor, and Filter (DAF), a stabilization framework that augments entropy-based adaptation with a marginal diversity loss that resists collapse, a cross-modal anchor consistency loss that constrains feature drift relative to a frozen source model, and feature salience filtering that skips low-value backward...
Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana et al.· 0 citations
Sensitivity-Guided Erasing Adaptation (SEGA) is introduced, a method for strict online continual TTA (CTTA) on corruption-style streams that yields consistent robustness and stability gains over strong CTTA baselines while reducing backward passes through sensitivity-based gating.
Chandler Timm C. Doloriel, Yunbei Zhang, M. Siddiqui et al.· 0 citations
This work injects two complementary semantic priors into Visual prompt tuning, a cascaded scheme that integrates both priors throughout ViT adaptation, and proposes a cascaded scheme that integrates both priors throughout ViT adaptation.
Xi Xiao, Xing-Jian Li, Cheng Han et al.· Trans. Mach. Learn. Res.· 0 citations
This comprehensive survey formally defines the CTTA problem, analyzes the diverse continual domain shift patterns that characterize different evaluation protocols, and proposes a hierarchical taxonomy that categorizes existing methods into three families: optimization-based strategies (entropy minimization, pseudo-labe...