The Security Duality of Large Language Model Agents: A Systematic Review of Self-Security Protection and Cybersecurity Empowerment
Abstract
Large Language Model (LLM) agents integrate memory, tool invocation, environment interaction, and multi-agent collaboration, evolving from passive text-generation tools into autonomous systems that execute complex tasks. This capability leap creates a pronounced security duality. On the one hand, LLM agents face emerging threats such as jailbreak attacks, prompt injection, memory poisoning, and multi-agent collusion, requiring systematic self-security protection. On the other hand, they empower cybersecurity in vulnerability detection, penetration testing, threat intelligence, and malicious code analysis, while their autonomous attack capabilities, e.g., zero-day exploitation, raise concerns about malicious abuse. Existing reviews often treat these two lines separately. This review systematically examines key studies from 2023 to 2026, covering attacks on agents and agent-enabled cyber operations. We propose a unified analytical framework with two dimensions: Dimension 1, self-security protection, and Dimension 2, cybersecurity empowerment. We organize self-security threats into model-layer safety alignment failure and fine-tuning-induced safety weakening; prompt-layer jailbreak and indirect prompt injection; memory/knowledge-layer retrieval-augmented generation (RAG) poisoning and long-term memory contamination; and multi-agent-layer prompt infection and secret collusion. For cybersecurity empowerment, we review vulnerability detection and remediation, automated penetration testing, cyber threat detection, autonomous vulnerability exploitation, and information ecosystem security. We further identify four coupling mechanisms: security weaknesses as attack entry points, tool integration expanding the attack surface, bidirectional capability enhancement, and double-edged multi-agent architectures. Finally, we distill five challenges—security-capability parity, native security design, multi-agent governance, ecosystem-oriented evaluation, and responsible capability release—and outline future directions for responsible development and deployment of LLM agents.