Preprint
Aug 2026
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
VisCache is proposed, a plug-and-play framework for coarse-to-fine KV pruning without training that consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference.
Lyuke Wang, Zhuo Li, Guangxu Zhu
· 0 citations