Open access
Jul 2026
AdaptiveKV: Accelerating KV Cache Offloading with a Bandwidth-Adaptive Memory Allocation Mechanism
A bandwidth-oriented memory allocation mechanism, named AdaptiveKV, which is self-adaptive to CXL-enabled memory pools and KV cache offloading scales for LLM inference acceleration, and achieves a maximum speedup in LLM inference throughput compared to the state-of-the-art strategies.
Yibo Tang, Lizhou Wu, Yang Ou et al.
· ACM Transactions on Architec... · 0 citations