This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.
Abstract
Extreme compression techniques, including vector quantization (VQ) and input-dependent sparsity, can significantly reduce the memory footprint of large language models (LLMs). However, a key challenge remains in translating such compression into practical efficiency. On conventional systolic-array-based accelerators, VQ incurs high dequantization overhead, while the irregular patterns of input-dependent sparsity are difficult to exploit. In this study, we address these challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference. Vortex adopts a bi-flow execution strategy that efficiently supports vector-quantized models across both prefill and decoding workloads, and we further optimize it through systematic design space exploration. On the algorithm side, we propose codebook-wise contextual sparsity to align with VQ execution. Across end-to-end workloads, Vortex achieves $8.03\times$--$23.7\times$ speedup and $5.68\times$--$12.5\times$ energy reduction over state-of-the-art accelerators.
Large language models (LLMs) exhibit strong performance across applications, but their inference is computationally intensive, posing significant challenges for edge deployment. Quantization is among the most effective and widely used optimizations. In particular, ternary-weight quantization further lowers compute cost...
Chang-Xu Liu, Yi-Fan Song, Yi-Feng Yang et al.· ACM Transactions on Design A...· 0 citations
Design rules and a reproducible evaluation protocol are contributed that jointly report quality, memory, and end-to-end speed, and a foundation for automated pipeline search under realistic single-GPU constraints is provided.
This work investigates lightweight BF16 weight compression schemes tailored for PIM-based LLM inference by focusing on exponent-oriented compression methods that exploit the locality and redundancy present in BF16 exponent fields.
Sabiha Tajdari, Akhil Shekar, Kevin Skadron et al.· 0 citations
Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.
Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al.· 3 citations
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trini...