SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models
Large language models (LLMs) are often memory-bandwidth bound during autoregressive decoding, so reducing weight storage does not automatically produce parallel speedup when the compressed representation is irregular. We present SSQT, a post-training framework that jointly applies hardware-aligned structured sparsifica...