Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing me...
Ruo-Ling Qi, Yi-Rui Liu, Xuan'er Wu et al.· 0 citations
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundar...
Yi-Rui Liu, Ruo-Ling Qi, Xuan'er Wu et al.· 0 citations
This work presents \textsc{Revise}, a validity-guided runtime for fine-grained recovery in structured agent workflows, a validity-guided runtime for fine-grained recovery in structured agent workflows that matches a latest-version oracle with no stale outputs or effects.
Ruoling Qi, Xuan'er Wu, Peng-Hang Liu et al.· 0 citations
Tail-Replay is presented, a prefix caching mechanism that enables unconstrained token-level prefix reuse in hybrid large language models and is evaluated on three Gated DeltaNet-based hybrid models using the LongBench and RULER benchmarks.
Yi-Rui Liu, Ruo-Ling Qi, Xuan'er Wu et al.· 1 citation
LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to res...
Yi-Rui Liu, Ruoling Qi, Long-Wen Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.