Skip to content

Author

Yi Kang

We have 1 of 4 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while als...

Qi-Hu Xie, Zi-Wei Li, Yi Kang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.