Long-running LLM agent workflows often revise interior context spans while retaining long suffixes. Although suffix tokens remain unchanged, altered causal histories and rotary positions prevent exact reuse of their offloaded key-value (KV) states. Full suffix recomputation wastes prefill work, while indiscriminate reu...
Guo-Tao Yang, Rui Guo, Si-Wei He et al.· 0 citations
Serverless edge computing, despite its flexibility and efficiency, is hindered by high startup latency during peak load. Remote fork, employing either Checkpoint/Restore (C/R) or Remote Direct Memory Access (RDMA), offers a potential solution for function scaling acceleration. Although RDMA fork is faster, the opportun...
Zhe-Xiong Li, De-Ze Zeng, Lin Gu et al.· IEEE Transactions on Mobile...· 0 citations
System is presented, a reference-aware codec that retrieves exact-token historical spans for prefill uplinks, reuses the reconstructed uplink state for same-round prefill downlinks, and generates boundary-specific decode references with lightweight causal predictors.
Guo-Tao Yang, Ming-Ze Zhao, Hao-Peng Li et al.· 0 citations
This work presents AsymSpec, which addresses uplink-gated verification and invalid dependent work with two corresponding mechanisms, and shows that AsymSpec delivers 2.82-28.03$\times the output-token throughput of the strongest baseline.
Guotao Yang, Hao Chen, Rui Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.