Skip to content

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

Sep 2026 · 0 citations · 58 references
Computer Science

TL;DR

To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to sustain the induced preference along subsequent trajectories after the trigger disappears.

Abstract

World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. Backdoor attacks offer a targeted and stealthy means of exploiting such supply chains, yet their threat to interactive world-model agents remains largely unexplored. To fill this gap, we present TrojanWorld, a backdoor framework for world-model agents that induces attacker-specified behavior by steering internal imagination. A physical object placed in the scene acts as the trigger, enabling deployment-time activation through the agent's native observation pipeline without digitally manipulating the observation stream. To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to sustain the induced preference along subsequent trajectories after the trigger disappears. Together, these mechanisms establish an end-to-end attack chain from physical perception through corrupted imagination to malicious action selection. Experiments with the TD-MPC2, DreamerV3, and R2-Dreamer systems across the DeepMind Control, MetaWorld, MyoSuite, and RoboDesk benchmarks show that under trigger activation, TrojanWorld achieves a target-action deviation as low as 0.026 while retaining at least 98.8% of the corresponding clean performance. Even after trigger removal, the compromised agent can remain trapped in the induced behavioral trajectory, continuing to execute attacker-specified actions.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control

Pretrained world models, learned simulators that encode an observation into a latent state and predict how it evolves under actions, are beginning to be reused as off-the-shelf dynamics backbones for control, like pretrained encoders and language models are reused today. We show that this reuse opens a supply-chain bac...

Roberto Riaño, Gorka Abad, S. Picek et al. · 0 citations
Preprint Aug 2026

EnvHarness: Awakening Static Worlds for Agent Learning

Environment Harness is proposed, a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic, enabling continuous, targeted co-evolution of the policy and its environment.

Chengsong Huang, Zifeng Wang, Rujun Han et al. · 7 citations
#artificial intelligence Preprint Sep 2026

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yi...

Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan et al. · 0 citations
Preprint Sep 2026

Backdoors in Learning-Based Industrial Robotic Arm Manipulation: An Empirical Security Study

Learning-based models (e.g., visuomotor and Vision-Language-Action (VLA)) are increasingly explored for industrial robotic manipulation, where model predictions are directly translated into physical actions. This tight coupling between model behavior and physical execution makes hidden security vulnerabilities particul...

Zi-Jian Zhang, Zhen Zeng, Zhong-Shu Gu et al. · 0 citations
2026

Action-Level Backdoor Attacks Against Deep Reinforcement Learning Systems via Adaptive Reward Exploration

Deep Reinforcement Learning (DRL) has demonstrated remarkable capabilities in domains such as robotics, finance, and autonomous systems. With the increasing cost of training, DRL models are increasingly shared and reused via model marketplaces, cloud platforms, and open-source repositories. This trend exposes DRL syste...

Ou-Bo Ma, L. Du, Yang Dai et al. · 0 citations
Preprint Aug 2026

BehaviorWorldGen: Closing the Loop between Action Models and World Simulators via Controllable Behavior-Aware Structured World Generation

Modern driving action models are increasingly improved in a self-improvement loop, where a learned world simulator imagines future observations and the resulting data is fed back to refine the action model. However, the bottleneck of this loop lies in the simulators'inability to generate behaviorally plausible response...

Jiaqi Wang, Zhuo Zhang, Hai-Ning Guan et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.