Skip to content

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

Sep 2026 · 0 citations · 57 references
Computer Science

TL;DR

Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.

Abstract

Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually use a lightweight selector to score context units, followed by hard Top-K selection that blocks gradients from the language modeling loss. Consequently, these methods commonly distill layer-wise dense attention distributions. Although this encourages the selector to rank context units by dense attention weights in the original model, the ranking is not directly aligned with their impact on predictions under a fixed attention budget (i.e., the number of attended context units per query), potentially wasting the limited budget on less useful units. To address this misalignment, we propose Simple Attention Sparsification (SAS), a gated sparse attention mechanism that optimizes context ranking end-to-end with the language modeling loss. The key idea is to inject the selector's continuous scores into attention logits during training, allowing the loss to update the selector through standard backpropagation. We identify several choices crucial for this simple design to work well in practice: placing the gate inside the attention softmax in log form, using normalized softmax gates to calibrate historical context against the always-retained current block, and preserving continuous selector scores so the model learns relative priorities rather than only hard selections. To support long-sequence training, we implement a memory-efficient Triton kernel that integrates SAS into FlashAttention-style computation. Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.

View source

Similar papers

#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation With Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations
#machine learning Preprint Sep 2026

You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likel...

Jia-Rong Wen, Qi Wang, Yun Qu et al. · 0 citations
Preprint Aug 2026

TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation

This work proposes a novel Token Selection approach for Preference Optimization in LLM-based sequential Recommendation, i.e., TSPORec, which accurately pinpoints informative tokens throughout the entire textual content to improve recommendation performance.

Wenqiao Zhu, Chao Xu, Haipang Wu et al. · 0 citations
#machine learning Preprint Sep 2026

DimPO: Dimensionality Reduction for Attention using Preference Optimization

A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full...

Vojtěch Lanz, Yu-Fei Cui, Prasanna Parthasarathi · 0 citations
#artificial intelligence Preprint Sep 2026

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean d...

Chu-Xu Song, Jiu-Qi Wei, Zhen-Can Peng · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing...

Yury Nahshan, Nati Daniel, Jacob Goldberger et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.