Empowering Hybrid Attention Models on NPUs
HA-NPU is presented, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms, and enhances execution efficiency by reorganizing the dataflow of the LA components across three levels.