This work makes the fraction $\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance.
Abstract
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance. Granted a channel rich enough to echo any action the human could take, higher $\varepsilon_t$ weakly lowers every monotone measure of the power of a human with a message-independent fallback. An oracle rewarded by per-round approval cultivates reliance beyond a closed-form patience threshold, so the same reward weights leave the optimal oracle answering in episodic deployments and cultivating in long-memory ones. An influence bound certified once at deployment is blind to that horizon and bounds the loss no lower than its trivial ceiling. An exogenous cap on influence bounds the guarantee the human loses, and a short enough memory reset removes the incentive to cultivate, while neither recovers the value already steered away. In a closed-form example the optimal oracle never cultivates in fifteen-round sessions and does in sixteen.
The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing, and what works is reconstructing the whole steered activation with those dimensions pinned to their natural values.
The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead.
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equ...
The myopic escalation threshold is derived in closed form, characterise the optimal policy via dynamic programming, and it is proved that the optimal policy is a time-varying threshold with no shape assumption on the raw signal.
A four-stage audit for frozen proximal policy optimization policies without retraining examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives.
Xing-Fei Zeng, Xin Zhong, Nan-Ting Li et al.· 0 citations
Can a restricted computational model predict well while omitting distinctions required by its explanatory task? We define boundary sufficiency relative to an outcome, a representation, and a declared family of input interventions. Exact sufficiency requires a common response law on every fiber of the retained represent...
Guillaume Vimeney· Persistence· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.