Skip to content
#small language model Book Open access

Edge-Cloud Collaborative RAG with Multi-State Semantic Caching

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 37 references

TL;DR

A Multi-State Semantic Cache (MSSC) framework that dynamically maps knowledge chunks into three mutually exclusive states: zero cache, index cache, and full cache is proposed that reduces average service latency by over 40% in dynamic scenarios and suppresses backhaul reconfiguration traffic by up to 38% compared to conventional binary caching baselines.

Abstract

Enterprise applications are increasingly adopting collaborative Retrieval Augmented Generation (RAG) frameworks that synergize Small Language Models (SLMs) at the network edge with Large Language Models (LLMs) in the cloud for joint reasoning. However, this paradigm is bottlenecked by the capacity asymmetry between massive cloud-based knowledge repositories and resource-constrained edge servers. Existing edge-cloud collaborative frameworks rely on monolithic storage strategies that fail to exploit the structural asymmetry between the vector index of knowledge chunks and raw documents. This limitation induces a severe conflict between service latency and cache reconfiguration traffic. This paper proposes a Multi-State Semantic Cache (MSSC) framework that dynamically maps knowledge chunks into three mutually exclusive states: zero cache, index cache, and full cache. By leveraging semantic graph diffusion, our framework captures conceptual dependencies to predict query trajectories, enabling predictive state transitions among three states. Furthermore, we formulate the cache scheduling process as a Model Predictive Control (MPC) problem to optimize the trade-off between service latency and reconfiguration traffic. Experimental evaluations on real-world datasets demonstrate that, while strictly preserving generation quality, our framework reduces average service latency by over 40% in dynamic scenarios and suppresses backhaul reconfiguration traffic by up to 38% compared to conventional binary caching baselines.

Read PDF

Similar papers

#edge computing Oct 2026

A Cloud--Edge Collaborative Large Language Model Inference Framework Based on Historical Context Matching

Cloud–edge collaborative inference has emerged as a promising paradigm to address the latency, energy, and privacy challenges of large language models (LLMs). However, current offloading mechanisms often struggle to efficiently capture the dynamic semantic dependencies between historical context and ongoing queries. Th...

Xianzhong Tian, Yu Wang, Guanpeng Zhu · 0 citations
#edge computing Preprint Sep 2026

AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference

AceSpec, an asymmetric edge-cloud collaborative framework that employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink and introduces a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local ca...

Yi-Da Zhang, Zhi-Yong Gao, Shuai-Bing Yue et al. · 0 citations
2026

Workflow-Aware Expert Routing for Distributed LLM Serving Over the Edge-Cloud Continuum

Deploying Large Language Models (LLMs) over the edge-cloud continuum faces severe stability challenges due to the conflict between stochastic network topology and complex workflow dependencies. Existing schedulers, relying either on computationally prohibitive Graph Neural Networks (GNNs) or topology-agnostic heuristic...

Yan Gao, Shaoyuan Huang, Yonghui Ye et al. · 0 citations
#graph neural networks Open access Sep 2026

Synergizing large and small models in cloud-edge continuum: a spatiotemporal hypergraph approach for dynamic offloading

This paper proposes a spatiotemporal hypergraph-driven framework integrating high-order topological feature extraction with dynamic resource modeling, and introduces dynamic hypergraph sequences to naturally encompass local conflict domains, mitigating the topological blind spots and “over-smoothing” issues inherent in...

Kun Ding, Xi-Wen Qiu, Nian-Feng Weng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization

Amid the rapid advancement of physical-world intelligence, cloud-edge collaborative large language models (LLMs) have emerged as a promising roadmap for practical LLM deployment. However, existing cloud-edge paradigms struggle to balance global consensus with local personalization, which fails to satisfy the need for a...

Victor H. Chen, Hai-Rui Yu, S. K. Chung et al. · 0 citations
Book Open access Aug 2026

BRIDGE: Block-Wise Speculative Coordination for Cloud-Edge Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) improves factuality by conditioning LLMs on retrieved evidence, yet real-world knowledge is often split across tiers: cloud-based RAG can exploit large public corpora, whereas edge-based RAG is the natural place to access private, user-specific stores. This raises a key question: ho...

Yuting Li, Shaoyuan Huang, Xiangqi Liu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.