#machine learning
Apr 2026
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
Flux Attention is introduced, a context-aware framework that dynamically optimizes attention computation at the layer level by integrating a lightweight Layer Router into frozen pretrained LLMs, which adaptively routes each layer to FA or SA based on the input context.
Quantong Qiu, Zhiyi Hong, Yi Yang et al.
· arXiv.org · 0 citations