Language models can over-condition on irrelevant preceding text: predictions already supported by local context may still change when distant, unrelated prefix tokens are perturbed. This interference is especially consequential in long, packed, or distractor-heavy contexts, where useful evidence and irrelevant spans co...
Jin-Chang Zhu, Hao-Lan He, Yi-Cheng Ding et al.· 0 citations
Latent chain-of-thought models move intermediate reasoning from emitted text into continuous states, improving compactness but hiding the causal object. We introduce SCIT, the Suffix Cache Interchange Test, a causal protocol that constructs exact source-recipient counterfactuals, patches declared cache segments, and id...
GapSight is proposed, a framework for learning visual re-reading: a VLM first takes a global glance, then selectively returns to a free-form region when the question calls for local evidence, and Mechanism analyses show that the router rescues concrete wrong answers, adapts its action rate by task, and forms a favorabl...
Jinchang Zhu, Rong Fu, Yicheng Ding et al.· 1 citation
Results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance, and show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.