Skip to content

Author

Roozbeh Bostandoost

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single f...

Michael Wang, Keith Li, Roozbeh Bostandoost · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.