Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs
On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing designs were primarily opt...