Preprint
Aug 2026
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference
This work presents a distributed inference framework that integrates speculative decoding across edge and cloud, and shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement.
D. J. Bajpai, K. Upadhyay, M. Hanawal
· 0 citations