Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units,...
Xu-Dong Wang, Hao Wu, Hao-Zhe Hu et al.· 0 citations
Results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss.
Yichu Fang, Si-Tong Wei, Hao-Zhe Hu et al.· 3 citations
WIDE is presented, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios, and a pruning--kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent...