Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally drive model behavior. We demonstrate that this gap cannot be closed from the geometry of probe weights alone: the featur...