Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...
Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al.· 0 citations
LaP-Forensics is presented, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence that supports the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.
Can Wang, Yuhao Wang, Yu-She Cao et al.· arXiv.org· 1 citation
These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
Siyuan Ma, Yiqin Luo, Zhangji et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.