On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily opt...
Zheng-Xiang Huang, Sheng-Heng Chen, Chao-Yue Niu et al.· 0 citations
C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.
Chuheng Du, Jun-Yi Chen, Hanlin Tang et al.· Proceedings of the 32nd ACM...· 2 citations
Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification, makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden.
Yijie Ma, Chao-Yue Niu, Fan Wu et al.· 0 citations
Gleam is a novel and network-efficient framework for task-generic GPU sharing across local-area CUDA devices, with three key contributions that reduce bandwidth overhead in CUDA API remoting through automatic model weight caching, and mitigate accumulated latency from frequent API calls by asynchronous execution.
This work proposes a novel topology named the Balanced Sparse Tree (BST), which is a topology characterized by symmetric design and sparse connections, motivated by hypergraph theory and Steiner Systems, and demonstrates the superiority of BST over the state-of-the-art in network scale, latency, bandwidth, and cost.
Shaoteng Liu, Dejun Kong, Huitian Wang et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.