Skip to content
Preprint

Vortex: Bridging Extreme Compression and Efficient LLM Inference

Sep 2026 · 0 citations · 67 references
Computer Science

TL;DR

This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.

Abstract

Extreme compression techniques, including vector quantization (VQ) and input-dependent sparsity, can significantly reduce the memory footprint of large language models (LLMs). However, a key challenge remains in translating such compression into practical efficiency. On conventional systolic-array-based accelerators, VQ incurs high dequantization overhead, while the irregular patterns of input-dependent sparsity are difficult to exploit. In this study, we address these challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference. Vortex adopts a bi-flow execution strategy that efficiently supports vector-quantized models across both prefill and decoding workloads, and we further optimize it through systematic design space exploration. On the algorithm side, we propose codebook-wise contextual sparsity to align with VQ execution. Across end-to-end workloads, Vortex achieves $8.03\times$--$23.7\times$ speedup and $5.68\times$--$12.5\times$ energy reduction over state-of-the-art accelerators.

View source

Similar papers

#edge computing Open access Sep 2026

SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models

Large language models (LLMs) exhibit strong performance across applications, but their inference is computationally intensive, posing significant challenges for edge deployment. Quantization is among the most effective and widely used optimizations. In particular, ternary-weight quantization further lowers compute cost...

Chang-Xu Liu, Yi-Fan Song, Yi-Feng Yang et al. · 0 citations

Tailoring LLM Weight Compression For PIM Architectures

This work investigates lightweight BF16 weight compression schemes tailored for PIM-based LLM inference by focusing on exponent-oriented compression methods that exploit the locality and redundancy present in BF16 exponent fields.

Sabiha Tajdari, Akhil Shekar, Kevin Skadron et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 3 citations
Preprint Aug 2026

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trini...

Mohammad Mozaffari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.