Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
Flux Attention is introduced, a context-aware framework that dynamically optimizes attention computation at the layer level by integrating a lightweight Layer Router into frozen pretrained LLMs, which adaptively routes each layer to FA or SA based on the input context.