SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models
Large language models (LLMs) exhibit strong performance across applications, but their inference is computationally intensive, posing significant challenges for edge deployment. Quantization is among the most effective and widely used optimizations. In particular, ternary-weight quantization further lowers compute cost...