SinkProbe is built, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and applies it to four small models that differ only in how they mix tokens and depth, and three results follow.
Abstract
Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea with Kimi Delta Attention and Attention Residuals behind a one million token window, eight times past the range where these diagnostics have been reported. This paper asks whether the fix survives that jump. We build SinkProbe, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and apply it to four small models that differ only in how they mix tokens and depth. Three results follow. The training objective produces the sink, not the architecture. Gating did not reproduce its published effect at our scale. Sink mass, activations and position bias moved independently. Code, data and the measurement protocol are released at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
Declarative Attention is introduced, a protocol that elicits the model to declare where it needs to attend within its chain-of-thought, partitioning generation into three modes:(full context),(a specific region), and(recent output only).
Namgyu Ho, Huzama Ahmad, Woosung Koh et al.· 1 citation
LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.
Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al.· 1 citation
How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the s...
Timur Mudarisov, M. Burtsev, Radu State· 0 citations
Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, although the benefit of global access varies across prediction positions. We find that, before global attention is computed for the current step, the decoding states...
Hai-Bo Feng, Ruiqi Liang, Dong-Yang Jin et al.· 0 citations
Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which laye...
Mu-Yu He, Yu-Chen Liu, Ran Tao et al.· 0 citations
Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bou...
Meysam Ghaffari, Nina Fatehi, Bhaskar Sen et al.· 0 citations