MVP: A Mobile 3D-Stacked VLM Accelerator for Efficient Video Understanding by Leveraging Dynamic Sparse Attention Patterns
The emergence of Vision-Language Models (VLMs) has enabled multimodal reasoning, e.g. video understanding, yet their extension to long-context inference remains bottlenecked by the “token explosion”. This surge in sequence length leads to prohibitive attention computation overhead and memory-bound KV cache access. Whil...