Tiered Speculative Decoding with Uplink Scheduling for End-Edge LLM Inference
With the large-scale deployment of LLM-driven services, centralized cloud inference faces increasing serving loads and latency bottlenecks. End-edge collaborative speculative decoding has emerged as a promising paradigm because its draftand-verify mechanism enables efficient collaboration between local generation and e...