Video world models achieve long-range temporal consistency by storing KV cache during generation, but the growing cache makes KV cache memory a major deployment bottleneck, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on...
Jia-Qi Zhao, Xiao-Bin Hu, Bo Yin et al.· 0 citations
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as a...
Hongyu Qu, Guang-Ming Yao, Ling Xing et al.· 0 citations
Experiments on OmniMedVQA subsets show that Med-OPD consistently outperforms SFT and standard OPD across CT, MRI, Disease Diagnosis, and Lesion Grading, and demonstrate that evidence-aware distillation can better strengthen medical VLMs'reliance on key visual evidence and improve reliable multimodal medical reasoning.
Yun-Hang Qian, Jiaquan Yu, Jiawei Liu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.