Long-context "context rot" (accuracy degradation as input length grows, even when task difficulty is held fixed) is localized here to a specific, causal mechanism inside a small Mixture-of-Experts language model (OLMoE-1B-7B-0924), then tested for generality across two further architectures. Forced-choice retrieval accuracy falls from 0.938 (256 tokens) to 0.688 (3,840 tokens) on a controlled needle-in-a-haystack substrate. A deconfounded linear probe shows the target fact remains ~99.5% decodable at its source position on every failing prompt, ruling out storage loss; decodability specifically at the readout position degrades instead (0.714 vs. 0.960 on model-right prompts). Sixteen attention heads, identified from short prompts alone, carry that content to the readout; on long failing prompts their attention to the fact collapses (0.432 to 0.187), localized to those heads well beyond a random-head null (specificity p<0.0005). A pre-softmax attention boost restricted to exactly those 16 heads, over the fact's span, repairs all 14 failing prompts against strength-matched random-head and wrong-span controls (answer probability 0.238 to 0.986, dz=5.55). Three router-level and readout-level interventions were tried first and failed: two that verifiably restored MoE specialist routing did not move accuracy, and a residual-stream content injection at the readout was content-independent — showing router "starvation" is a downstream correlate of the transport failure, not its cause. The full causal chain of head identification, localized collapse, and causal repair replicates on a second, architecturally distinct MoE (Granite-3.0-3B-A800M: 5/5 failures repaired against matched controls) and on a dense transformer with no MoE component at all (Pythia-2.8B: 11/11 repaired, dz=1.33-1.44; the collapse measurement itself on this substrate is significant and specific but falls under the project's own effect-size floor, and is reported as suggestive rather than confirmatory). This indicates the attention-transport failure and its repair do not depend on Mixture-of-Experts routing. A training-free, IDF-weighted lexical detector, requiring zero forward passes of the model under study, locates the failing span and recovers 100.8% and 99.4% of the oracle repair on the two primary OLMoE substrates, 99.1% on Granite, and 51% on a harder Pythia substrate built specifically to weaken lexical anchoring. Its boundary was tested, not assumed. It is unaffected by paraphrase (100% hit, 100.8% of oracle) but fails completely on multi-hop composition as a single pass (0% hit); a training-free two-stage chain recovers 45.0% of the oracle effect there with no labels, and a registered labeled fallback recovers 61.4%. A follow-up track replaces lexical overlap with sentence embeddings and a graph walk specifically to test coreference, where the relevant content shares zero tokens with the question by construction. The result is a registered split, not a single verdict: a discourse-adjacency edge fully solves coreference where the referring expression is immediately adjacent to its antecedent (100% of oracle, 22 of 22 repaired), but on a harder substrate with antecedent distance drawn independently per prompt, both that mechanism and a distance-tolerant version built specifically to extend it fail identically past distance one (dz=0.48). Both are reported as registered negative results, not discarded. Every claim above is preregistered before evaluation, with effect-size floors (|d| or |dz| >= 0.8) alongside significance, sign-flip and label-shuffle permutation tests (2,000 draws), matched controls, and Benjamini-Hochberg FDR correction. Results that miss a registered bar are reported as such rather than dropped — including three failed intervention families, an initially wrong lexical-detector prediction (logged and corrected), and both coreference negative results above.
Manjunath Bhaskar· Zenodo (CERN European Organi...· 0 citations
Long-context "context rot" (accuracy degradation as input length grows, even when task difficulty is held fixed) is localized here to a specific, causal mechanism inside a small Mixture-of-Experts language model (OLMoE-1B-7B-0924), then tested for generality across two further architectures. Forced-choice retrieval accuracy falls from 0.938 (256 tokens) to 0.688 (3,840 tokens) on a controlled needle-in-a-haystack substrate. A deconfounded linear probe shows the target fact remains ~99.5% decodable at its source position on every failing prompt, ruling out storage loss; decodability specifically at the readout position degrades instead (0.714 vs. 0.960 on model-right prompts). Sixteen attention heads, identified from short prompts alone, carry that content to the readout; on long failing prompts their attention to the fact collapses (0.432 to 0.187), localized to those heads well beyond a random-head null (specificity p<0.0005). A pre-softmax attention boost restricted to exactly those 16 heads, over the fact's span, repairs all 14 failing prompts against strength-matched random-head and wrong-span controls (answer probability 0.238 to 0.986, dz=5.55). Three router-level and readout-level interventions were tried first and failed: two that verifiably restored MoE specialist routing did not move accuracy, and a residual-stream content injection at the readout was content-independent — showing router "starvation" is a downstream correlate of the transport failure, not its cause. The full causal chain of head identification, localized collapse, and causal repair replicates on a second, architecturally distinct MoE (Granite-3.0-3B-A800M: 5/5 failures repaired against matched controls) and on a dense transformer with no MoE component at all (Pythia-2.8B: 11/11 repaired, dz=1.33-1.44; the collapse measurement itself on this substrate is significant and specific but falls under the project's own effect-size floor, and is reported as suggestive rather than confirmatory). This indicates the attention-transport failure and its repair do not depend on Mixture-of-Experts routing. A training-free, IDF-weighted lexical detector, requiring zero forward passes of the model under study, locates the failing span and recovers 100.8% and 99.4% of the oracle repair on the two primary OLMoE substrates, 99.1% on Granite, and 51% on a harder Pythia substrate built specifically to weaken lexical anchoring. Its boundary was tested, not assumed. It is unaffected by paraphrase (100% hit, 100.8% of oracle) but fails completely on multi-hop composition as a single pass (0% hit); a training-free two-stage chain recovers 45.0% of the oracle effect there with no labels, and a registered labeled fallback recovers 61.4%. A follow-up track replaces lexical overlap with sentence embeddings and a graph walk specifically to test coreference, where the relevant content shares zero tokens with the question by construction. The result is a registered split, not a single verdict: a discourse-adjacency edge fully solves coreference where the referring expression is immediately adjacent to its antecedent (100% of oracle, 22 of 22 repaired), but on a harder substrate with antecedent distance drawn independently per prompt, both that mechanism and a distance-tolerant version built specifically to extend it fail identically past distance one (dz=0.48). Both are reported as registered negative results, not discarded. Every claim above is preregistered before evaluation, with effect-size floors (|d| or |dz| >= 0.8) alongside significance, sign-flip and label-shuffle permutation tests (2,000 draws), matched controls, and Benjamini-Hochberg FDR correction. Results that miss a registered bar are reported as such rather than dropped — including three failed intervention families, an initially wrong lexical-detector prediction (logged and corrected), and both coreference negative results above.
Manjunath Bhaskar· Zenodo (CERN European Organi...· 0 citations