Long-running LLM agent workflows often revise interior context spans while retaining long suffixes. Although suffix tokens remain unchanged, altered causal histories and rotary positions prevent exact reuse of their offloaded key-value (KV) states. Full suffix recomputation wastes prefill work, while indiscriminate reu...
Guo-Tao Yang, Rui Guo, Si-Wei He et al.· 0 citations
System is presented, a reference-aware codec that retrieves exact-token historical spans for prefill uplinks, reuses the reconstructed uplink state for same-round prefill downlinks, and generates boundary-specific decode references with lightweight causal predictors.
Guo-Tao Yang, Ming-Ze Zhao, Hao-Peng Li et al.· 0 citations
This work presents AsymSpec, which addresses uplink-gated verification and invalid dependent work with two corresponding mechanisms, and shows that AsymSpec delivers 2.82-28.03$\times the output-token throughput of the strongest baseline.
Guotao Yang, Hao Chen, Rui Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.