Skip to content

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Semantic-Geometric Decoupled Routing is proposed, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances and yields an explicit closed-form block routing score without token-level search or post-hoc calibration.

Abstract

Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult. Existing routers often pool post-RoPE token representations, which entangles semantic aggregation with RoPE-induced geometry and attenuates local positional cues through high-frequency phase cancellation. To resolve this mismatch, we propose \textbf{Semantic-Geometric Decoupled Routing}, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances. This decomposition yields an explicit closed-form block routing score without token-level search or post-hoc calibration. Experiments on long-context text and video tasks show that our method approaches full-attention accuracy across 4K--128K contexts, keeps routing overhead below 3.4 ms, and achieves a 5.03$\times$ speedup over FlashAttn at a 128K context length.

View source

Similar papers

#small language model Preprint Aug 2026

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

The results establish BF1 as a reproducible sparse operator and selective retrofit primitive with real long-context systems value, and evaluates numerical correctness, selected-interaction scaling, kernel performance, partial-model inference, and matched next-token language modeling.

Hina Dixit · 0 citations
#machine learning Preprint Sep 2026

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely...

Joe McKenna, Anastasios Alexandridis, Nathan Susanj et al. · 0 citations
#natural language process... Preprint Sep 2026

CEDAR: Error-Bounded Residual Routing for Efficient Long-Context Attention

This work introduces Coarse-to-fine Error-aware Dynamic Attention Routing (CEDAR), a coarse-to-fine method that keeps the language model frozen while preserving global coverage and derives an output-error bound governed by within-chunk key/value dispersion and uses it to allocate a variable refinement budget.

Si-Yu Li, Dong Wang, Jie Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SMat-Attention: Structured Long-Context Sequence Modeling

Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes...

E. Anand, Abdullah Ateyeh, Archer Wang et al. · 1 citation
#machine learning Preprint Sep 2026

Block Sparse Attention with Log-Linear Complexity

PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.

Bo-Hao Tang, Zhen Qin, Yu-Qi Pan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries

The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, wit...

M. Paolicelli, Alessandro Petruzzelli, Alessandro Francesco Maria Martina et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.