Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks
Two universal tool-based defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores the agent's original toolset prior to planning.
Xiao-Yan Li, Yun-Li Wang
· 0 citations