Pretrained Foundation Models (PFMs) enable highaccuracy inference services but are typically deployed in remote datacenters, resulting in prohibitively high inference delay. Mobile Edge Computing (MEC) can mitigate such high delays by caching PFMs or their fine-tuned variants on cloudlets located close to end users. Ho...
Li-Zhe Zhou, Qiu-Fen Xia, Zi-Chuan Xu et al.· IEEE Transactions on Paralle...· 0 citations
MetaKV is proposed, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.
Rui-Kun Luo, C. Gu, Jing Yang et al.· Proceedings of the 32nd ACM...· 0 citations
MetaKV is proposed, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.
Ruikun Luo, C. Gu, Jing Yang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.