Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question,...
Jin-Tao Tong, Yujing Lou, Zhan-Ming Shen et al.· 0 citations
STD is proposed, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases.
Shuo Zhang, Jin-Tao Tong, Yi-Xiong Zou et al.· 0 citations
CoReFuse-Med is proposed, a Corruption-aware Rebalanced Fusion framework that suppresses corruption during feature transmission and rebalances modality contributions during high-level fusion, demonstrating improved accuracy and robustness under modality-quality discrepancies.
Yu-Chen Pei, Xiao-Yu Hu, Yi-Xiong Zou et al.· 0 citations
Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes with only a few samples while avoiding forgetting base classes. However, current methods show a tendency to misclassify novel-class samples into base classes, which we find to be caused by the excessive focus on base-class-discriminat...
Haichen Zhou, Y. Lyu, Yixiong Zou et al.· IEEE transactions on multime...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.