What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructions during LLM inference, thereby subverting the agent's decision. Exi...