Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...
Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al.· 0 citations
This work introduces TempFinRAG, a point-in-time evaluation protocol built from public filings and XBRL facts, and introduces TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluates the framework on complementary evidence-grounded, numerical, conversational, and multi-tabl...
Lanju Tao, Zheng-Ji Li, Ying-Rui Ji et al.· Symmetry· 0 citations
This work injects two complementary semantic priors into Visual prompt tuning, a cascaded scheme that integrates both priors throughout ViT adaptation, and proposes a cascaded scheme that integrates both priors throughout ViT adaptation.
Xi Xiao, Xing-Jian Li, Cheng Han et al.· Trans. Mach. Learn. Res.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.