A Cloud--Edge Collaborative Large Language Model Inference Framework Based on Historical Context Matching
Cloud–edge collaborative inference has emerged as a promising paradigm to address the latency, energy, and privacy challenges of large language models (LLMs). However, current offloading mechanisms often struggle to efficiently capture the dynamic semantic dependencies between historical context and ongoing queries. Th...