Right Answer, Wrong Mechanism: Detecting Pernicious Divergence in Causal Interventions
Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless a...