TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models
Speculative decoding accelerates large language model inference through collaboration between a lightweight draft model and a target verifier. Existing methods mainly improve the draft side, while the target model is typically kept dense and unchanged. We show that, under domain-specific inference, full-depth target ve...