Preprint
Jul 2026
The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity
It is argued that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge, and a cheap instrument for measuring it is given.
Chih-Ting Liao, Xinzhuo Cao
· 0 citations