Skip to content

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

Two universal tool-based defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores the agent's original toolset prior to planning.

Abstract

Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion. However, they are increasingly vulnerable to adversarial attacks, including direct prompt injection, indirect prompt injection, memory poisoning, and backdoor attacks, which exploit the model's openness to prompt injection and tool manipulation. In this work, we explore practical and generalizable defense strategies within a unified framework across these four attack types. We introduce two universal tool-based defenses: Attacker Tool Filtering, which uses anomaly detection (e.g., Isolation Forest) to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores the agent's original toolset prior to planning. Additionally, we incorporate prompt-based defenses: Chain-of-Thought prompting and self-reflection techniques to enhance reasoning and task paraphrasing to mitigate attacks. Experimental results across both four open-source LLMs (Gemma2-9B, Qwen2-7B, LLaMA3-8B, and LLaMA3.1-8B) and three proprietary LLMs (GPT-3.5, GPT-4, and GPT-5) show that our methods significantly reduce the Attack Success Rates (ASR), achieving 0% ASR in many settings, while preserving or even improving the original task success rate. These findings highlight the promise of simple, modular, multi-layered defenses for strengthening the security and robustness of tool-integrated LLM agents. The code is available at https://github.com/Xiaoyan-Lisa/Defenses-for-Tool-Integrated-LLM-Agents-Against-Adversarial-Attacks.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.

Zi Liang, XiaoYu Xu, Yanyun Wang et al. · 1 citation
#artificial intelligence Preprint Oct 2026

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's o...

Mohamed Dhouib, Clément Elliker, Alexi Canesse et al. · 0 citations
Review Open access Aug 2026

Securing the Prompt Pipeline: A Systematic Review of Defense Mechanisms Against Prompt-Based Attacks in LLM Agents

A systematic review and structured descriptive synthesis of research on defenses against prompt-based attacks in language model and agent systems reveals trade-offs between security effectiveness, performance, and system complexity as well as major gaps in benchmarks, indirect attack coverage, and multi-agent evaluatio...

Sana Mourad, E. E. Abdallah, Mohammad Ababneh · 0 citations
Preprint Aug 2026

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness, and ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.

Yutao Mou, Pengfei Yang, Zhenfei Yin et al. · 2 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.