Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs
This work provides the first mechanistic base-vs-instruct comparison of conflict-resolution circuits, and believes that because the conflict circuit is preserved rather than rebuilt, interpretability and control tools calibrated on base models should transfer directly to their deployed instruct siblings.