Skip to content

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

This work provides the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, and believes that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.

Abstract

In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We provide the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, across three families (Llama-3.2-3B, Qwen-2.5-3B, Gemma-3-4B). Five independent methods, node and edge attribution, superposition role analysis, causal ablation, and path patching, converge on gating, with the same heads, in the same late-layers, are found to be reweighted rather than replaced with a high node overlap (0.60-0.82). Behaviorally, tuning shifts models toward parametric memory, making instruct models reject a terse counterfactual context far more than base ones, the opposite of a naive user-following expectation. Yet this added skepticism is a factor of framing since it disappears when the same false claim is delivered as a coherent, evidential passage. The robustness that instruction tuning buys against terse injection is therefore real but narrow. More broadly, we believe that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

It is shown that user-preferring conflict resolution can coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone, while directions selected mainly for pooled separability steer poorly.

Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi et al. · 0 citations
Preprint Aug 2026

Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution

The authors train 26M-parameter transformers on a synthetic language where recency and rarity are exactly coextensive, and separate them with a minimal causal edit that inverts one cue while holding the truth, token count and answer position fixed.

Yijun Liao, Fan-Wei Liang · 0 citations
#machine learning Preprint Sep 2026

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA)...

Muhammad Usama, Dong Eui Chang · 0 citations
#artificial intelligence Preprint Sep 2026

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

MemCalib-RL is proposed, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation and achieves the best overall performance while better balancing over-use and under-use.

Rui-Ke Cao, Fan-Yu Zhao, Fu-Gen Yao et al. · 0 citations
Preprint Aug 2026

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

The complete pipeline -- detect, localize, and release -- is submitted to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented.

Xi-Ning Xun · 0 citations
Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.