Skip to content
Preprint

The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity

Jul 2026 · 0 citations · 48 references
Computer Science

TL;DR

It is argued that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge, and a cheap instrument for measuring it is given.

Abstract

A linear probe that recovers a conserved quantity from a learned dynamics model's activations is routinely read as evidence that the model uses that quantity. We show this inference is unsound. Across mechanical, circuit, and partial-differential-equation (PDE) systems, and on a 158M-parameter pretrained PDE foundation model, energy and other invariants are linearly decodable at $R^2 \approx 1$ yet causally inert on next-state prediction: overwriting the decoded direction with a donor state's value (single-step activation interchange) leaves the forward pass essentially unchanged (transfer-corr $\tau \approx 0$). The same direction in the same representation becomes causally load-bearing ($\tau \to +1$) the moment the training objective rewards the invariant, so deployment is a property of the objective, not of the representation or the probe. We further show that when an invariant is deployed is governed by a precise algebraic predicate (its relation to the prediction output), by flipping a single invariant from inert to load-bearing by changing only the output's algebra. Finally, the gap has teeth: across models that all decode the target at $R^2 = 1.00$, the deployment gap forecasts out-of-distribution (OOD) accuracy ($r = +0.97$) where decodability is blind. We argue that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge, and we give a cheap instrument for measuring it.

View source

Similar papers

Preprint Jul 2026

How are linear representations learned? Exact solutions to the dynamics of abstraction

In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call"abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.

William Yang, Andrew M. Saxe, Peter E. Latham · 0 citations
Preprint Aug 2026

Learned proposals in trans-dimensional inference are optimal at equilibrium, not during assembly

Inferring the dimension of a model - the number of components needed to explain data - jointly with the parameters is a pervasive problem, from counting sources in an image to mixture modeling, and reversible-jump Markov chain Monte Carlo solves it exactly but mixes slowly. Learned proposals are well established at fixed dimension, but whether they can accelerate the dimension-changing moves themselves has remained largely untested. We show that the answer has a structural origin: the optimal proposal for the dimension-changing birth move is a different object in different phases of the run. While the fit is being assembled it must match the current residual - a state-dependent quantity no state-independent network can represent - but at equilibrium it degenerates to the posterior's single-component marginal, which is exactly the distribution an adaptive normalizing flow learns from the sampler's own history. A learned state-independent birth proposal is therefore useless in one phase and optimal in the other. Controlled experiments confirm the attribution: applied with an exact Metropolis--Hastings correction that leaves the target invariant for any network, the learned births leave acceptance rates unchanged yet accelerate model-order mixing - in a ten-seed benchmark they meet a pre-specified stopping rule in six of ten runs, typically several times sooner, where a strong hand-tuned baseline meets it in one (one-sided p=0.03) - and an isolation experiment shows the same flow deployed within-model buys nothing. Making no domain-specific assumptions, the same sampler counts sources in a noisy image and reconstructs signals across scientific domains, including gravitational waves from ground- and space-based detectors and a scalp EEG recording. We release the method as HyperWave, an open-source package.

A. Sasli, N. Karnesis, M. Karamanis et al. · 0 citations
Preprint Aug 2026

Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control

An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics and certified stability - that modulates cognition without performing it? Ours is a field on the module graph governed by a family of PDEs on the graph Laplacian, advancing with an adaptive-depth reasoner. We certify the stability of the integrator of the whole family - an integrator certificate, not a closed-loop one. New, and proved here: a discrete Schur-Cohn criterion for Verlet with velocity coupling, necessary and sufficient per latent root, with no commutation hypothesis. The answer is threefold: substance no, structure only in part, certifiability yes. The type of the field's physics is irrelevant for accuracy: wave, diffusion, gated mixtures and a 2D Navier-Stokes substrate tie. A twenty-seed preregistered deconfounding campaign bounds the structural claim: at equalized caps the second-order effect is strong in one family (+0.087 [+0.042, +0.132], t=4.0) but is not detected in the other (+0.014 [-0.013, +0.040], n.s.), so part of the original contrast was capacity, not order; and a matched-interface GRU is indistinguishable in the first and nominally exceeds the field in the second (-0.035 [-0.067, -0.002]). What distinguishes the field is not capability but that its one-step operator admits an exact runtime stability check - a difference of kind, not of existence: learned recurrences carry certificates too, sufficient and conservative ones. A kill-gate with a positive control finds no evidence for the field as evidence accumulator (Delta AUC +0.0007 [-0.0065, +0.0079] vs a 0.03 threshold). A dynamic internal field is a viable, certifiable compute governor, but not an enhancer of cognition: it modulates, it does not think.

F. Arrabal-Campos, Ignacio Fernández, Francisco G. Montoya et al. · 1 citation
Preprint Aug 2026

Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW

This work forms AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates, and derives an exact multistep error decomposition and establishes first-order finite-horizon accuracy under local smoothness and controlled activation switching.

Kangning Liu, Suyan Li · 0 citations
Preprint Jul 2026

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contact stiffness. A certificate-gated protocol first certifies each parameter as recoverable from raw observations, then measures whether it enters the latent, so a null result can be attributed to the objective rather than to the environment. The resulting identifiability map has two organizing mechanisms and one frontier. Inputs limit what can be known, while prediction targets decide what is retained. Stiffness enters the latent only when touch is forecast ($R^2=0.50$, compared with $-0.02$ when the same signal is merely fused into the input), and under single-step prediction a vision-only latent discards even perfectly visible object state. Drag marks the frontier. It carries a recoverability certificate of 0.89 yet plateaus near 0.13 under every deterministic prediction objective we test, while a supervised head on the same trunk reaches 0.45. Parameters whose readout is slow and ratio-type under the sensed coordinates fall outside what these objectives acquire. On RH20T, an input-target factorial across scaling curves reproduces both mechanisms across two robots and 4,258 episodes. Every arm missing information or prediction pressure stays flat over a fivefold data range, and only the full multimodal objective forecasts force beyond a persistence baseline, with held-out gains that grow with scale. Objective structure determines which physical parameters a latent acquires, and additional data improves only the parameters it already acquires.

Kaizhen Tan, Xin Xu, Siru Tao et al. · 2 citations
Preprint Aug 2026

Hidden Gauge Controls Feature Specialization in ReLU Networks

Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire a teacher feature while the others become redundant. We call the identity of that neuron feature ownership and ask whether it can be controlled by a parameter choice invisible to the initial predictor. In a tractable Gaussian teacher--student model, we fix the complete initial function and vary only a positive-homogeneous scaling gauge. Opposite gauges produce distinct feature trajectories and a sharp $\Theta(D^2)$ separation in specialization time that no global change of clock can explain. Among any fixed number of initially duplicate students, assigning the favorable gauge to one neuron deterministically selects it as the owner and drives the remaining functional contribution to zero. An exact reaction--transport decomposition attributes the effect to different mobilities for changing a feature's coefficient and direction. We prove global selection and functional pruning, extend finite-time selection to visible perturbations and small-step full-batch gradient descent, and verify the predicted loss, alignment, pruning, and dissipation trajectories in population and finite-sample training. The initial predictor therefore determines neither when the feature is learned nor which neuron learns it.

Tongxi Wang · 1 citation