LLM-Powered Automated Attacks
Abstract
The increasing integration of large language models (LLMs) into systems introduces new attack surfaces that extend beyond traditional software vulnerabilities. While LLMs are commonly protected by prompt-level security mechanisms, recent researches show that these controls can be bypassed through carefully crafted inputs. In this paper, we propose an interaction model based on a Generative Adversarial Network (GAN) conceptual analogy and a dual-LLM framework for systematically examining and testing the security boundaries of LLMs through malicious code generation. The framework consists of two LLMs that iteratively produce attack-oriented prompts, interpret and implement them. Experimental results demonstrate that the proposed approach can generate outputs that are similar to realworld attack patterns, such as SQL injection and cross-site scripting. This research highlights the importance of LLM security controls and emphasizes the need for proactive, automated evaluation methods to improve the robustness and governance of LLM-based systems.