A new multimodal ICL framework is proposed that combines contrastive demonstration modeling with the self-refinement capability of MLLMs and consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).
Ming-Bo Yang, Wen-Qiang Wang, Zhaolu Kang et al.· 1 citation
This work introduces a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment.
Peng Chen, Kai-Ge Li, Wei Wang et al.· arXiv.org· 0 citations
COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.
Cheng-Hua Zhu, Zhaolu Kang, Qifan Shi et al.· 1 citation
Performance-Driven Demonstration Selection (PDDS), which directly aligns demonstration selection with ICL performance, is proposed, which formulates selection as predicting the target LLM’s downstream task performance for a given query–in-context pair, replacing proxy heuristics with a performance-aware objec-tive.
Wen-Qiang Wang, Ming-Bo Yang, Ai-Ping Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.