Skip to content
Review

On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

Aug 2026 · 0 citations · 136 references
Computer Science

TL;DR

A four-layer taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers is contributed, and seven open problems centered on containment are identified.

Abstract

Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-world privileges---calling APIs, modifying files, and querying databases---a compromised reasoning step can trigger unauthorized data access, irreversible state changes, or cascading failures, yet the security research community has not kept pace. To quantify the state of the field, we conducted a systematic literature review under PRISMA 2020 guidelines across six databases, screening 743 records and retaining 85 papers (2023--2025) on agentic LLM security. Attack research outpaces defense work by 3.9:1. Perception-layer vulnerabilities (prompt injection, jailbreaking, adversarial perturbations) dominate, accounting for 66\% of papers, while action-layer vulnerabilities (tool misuse, code injection, sandbox escape) appear in only 4.7\%, misaligned with real-world risk. Code execution security accounts for 3.5\%, and tool-augmented agents 12\%. We contribute a four-layer taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers, and identify seven open problems centered on containment. Agentic LLM insecurity stems from architectural coupling, where weak isolation allows vulnerabilities to propagate across layers.

View source

Similar papers

Review Aug 2026

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or cascading compromise. We argue that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns; existing safeguards are typically local to one lifecycle boundary or one call. Guided by this threat model, we present ClawSentry, an open-source, framework-agnostic security supervision gateway for agent runtimes. Before a skill package is ever executed, First-use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read-only agentic review (locus A). At runtime, a three-tier progressive decision engine--a deterministic L1 layer, a rule-anchored L2 semantic reviewer, and a read-only L3 evidence-seeking agent--spends contextual review only on the residual ambiguity, while a session-level anti-bypass mechanism recognizes tool-switching and rephrased retries (loci B--C); a post-action path feeds high-severity evidence non-retroactively into later review (locus D). An Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals. On SkillInject with Codex/GPT-5.4, contextual ASR falls from 39.55% to 2.61% while contextual TSR moves only from 83.78% to 83.05%. Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines ASR to 9.09--15.03% from 33.5--49.7% unprotected, and aggregate TSR on clean skills remains 98.7%.

Kai Wang, Zeming Wei, Biaojie Zeng et al. · 0 citations
Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 0 citations
Preprint Aug 2026

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

This work proposes a self-evolving test-time defense built around a persistent, cross-interaction rule memory that substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.

Tongshen Hu, Bryan Hooi · 0 citations
Conference Jul 2026

A Survey of Fuzzing Techniques for Large Language Model Agents: Challenges, Methods, and Future Directions

With the rapid development of large language models (LLMs), LLM-based agents are evolving from single-turn text generators into autonomous decision-making systems capable of task planning, tool invocation, memory retrieval, and environmental interaction. However, the openness of agent environments and the nondeterminism of model-driven decision logic introduce security and reliability risks that are difficult to capture with conventional software testing. Fuzzing, as an automated and scalable vulnerability discovery technique, provides a promising approach for assessing the robustness of LLM agent systems. This paper presents a systematic survey of fuzzing techniques for LLM agents. First, we analyze the distinctive architecture of LLM agents and summarize three representative threat categories: indirect prompt injection, planning hallucination, and tool abuse. Second, we review existing fuzzing methods from three perspectives: test-case generation and mutation, environment simulation and interaction monitoring, and test-oracle construction. Particular attention is paid to LLM-enhanced fuzzers that exploit the semantic understanding and generation capabilities of LLMs to improve mutation quality and attack effectiveness. Finally, we connect the main technical challenges with an actionable research roadmap, including semantic coverage, cost-aware execution, reproducible oracles, multi-agent fuzzing, gray-box guidance, and automated repair.

Yang Cao · 0 citations
Book Open access Aug 2026

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models brings together researchers and practitioners from data mining, machine learning, security, and responsible AI to address issues from a data-centric, system-level perspective.

Lu Lin, Jinghui Chen, Ting Wang et al. · 0 citations