Skip to content
Preprint

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

Konur Tholl F. Rivest Mariam El Mezouar Adrian Taylor Ranwa Al Mallah
Jul 2026 · 0 citations · 28 references
Computer Science

TL;DR

The use of a Large Language Model (LLM) to improve autonomous defensive decision-making within an ACO environment is investigated and an online policy distillation framework is proposed that transfers the LLM's defensive policy into a lightweight RL agent containing only 64,910 parameters, reducing model size by several orders of magnitude while maintaining effective defensive capabilities.

Abstract

Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication. ACO applications commonly employ Reinforcement Learning (RL) agents to learn defensive behaviors through interaction with environments. However, RL agents typically require extensive exploration during training, often resulting in unstable behavior and poor initial decision-making before converging toward effective defense strategies. In this work, we investigate the use of a Large Language Model (LLM) to improve autonomous defensive decision-making within an ACO environment. Through prompt engineering rather than fine-tuning, we demonstrate that an 8-billion parameter LLM pretrained on cybersecurity data can outperform a baseline RL agent in a modified CybORG CAGE Challenge 2 environment. We then propose an online policy distillation framework that transfers the LLM's defensive policy into a lightweight RL agent containing only 64,910 parameters, reducing model size by several orders of magnitude while maintaining effective defensive capabilities. This provides a pathway toward operationalizing frontier cybersecurity models within lightweight, deployable agents. To evaluate transferability, we construct CybORG scenarios ranging from 4 to 12 hosts and assess the approach across varying network configurations. We also evaluate teacher-guided RL stabilization strategies and observe that none consistently surpass the optimized teacher policy, suggesting policy-alignment limitations between reward-driven RL optimization and teacher-guided defense strategies. Our results demonstrate the potential of cybersecurity-focused LLMs as sources of expertise for autonomous cyber defense, while policy distillation provides a practical path toward operationalizing frontier cybersecurity models within efficient, scalable agents.

View source

Similar papers

Preprint Aug 2026

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.

Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi et al. · 0 citations
2026

Cyber Task Automation With Knowledge-Infused Reinforcement Learning and LLM-Guided Policies

As cyber threats continue to evolve, there is a need for Autonomous Cyber Defense (ACD) strategies capable of fast and context-aware responses. Reinforcement learning (RL) has shown promise in automating cyber defense by exploring and learning effective countermeasures. However, RL often struggles with sparse reward signals and insufficient context to handle diverse attack scenarios. Furthermore, the convergence time of an RL agent is often high, making it difficult to train the agent in online settings. To address these challenges, we propose a large language model (LLM)-enhanced RL method that builds and queries a knowledge base (KB) derived from agent–environment interactions. We leverage the pre-trained knowledge of an LLM on different cybersecurity frameworks and use the LLM to analyze parts of the KB to generate appropriate actions for the RL agent. The LLM-generated output is infused into the RL training process to improve performance and reduce convergence time. To validate our approach, we formulate two RL problems: a contextual bandit problem, which accounts for possible misclassifications of network flows by the detection module, and a multi-step RL problem, which considers that adversarial actions may be missed by monitoring or detection tools. For the contextual bandit problem, we develop a custom environment guided by the MITRE ATT&CK framework, while for the multi-step RL problem, we use a prominent Cybersecurity simulation platform named CybORG. Experimental results show that our proposed approach outperforms the baseline RL by over 75% and 65% in the contextual bandit and multi-step RL settings, respectively, in terms of selecting more effective actions.

Md. Shamim Towhid, Shahrear Iqbal, E. P. Neto et al. · 0 citations
Preprint Aug 2026

Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination

Cyber threats continue to escalate in both frequency and sophistication, necessitating more adaptive and scalable defense strategies. This paper explores how Large Language Models (LLMs) can bolster cybersecurity simulations by automating the creation of synthetic environments and identifying latent vulnerabilities. We employ YAML as a structured representation format for simulating complex network configurations, thereby enabling Large Language Model-driven pipelines to support and improve reinforcement learning (RL) agent training. Comparative studies examine the advantages of LLM-based techniques over classical approaches such as Double Q-learning with Prioritized Experience Replay (PER), emphasizing increased efficiency, higher adaptability, and enhanced realism in cyberattack simulations. In empirical benchmarks across multiple synthetic topologies, LLM-instantiated Python agents achieved up to a 94.5% compromise rate while executing in 0.02-0.06 seconds per assessment---a ~25,000x to 50,000x speedup over traditional RL training cycles. Our findings underscore the transformative potential of integrating LLMs into cybersecurity research, ultimately paving the way for more intelligent and robust cyber-defense systems.

S. Kampakis, Fabio Rovai, Marcos Charalambides et al. · 0 citations
Conference Jul 2026

Can LLM Agents Replace Reinforcement Learning Agents in Cyber Defence Automation: A Case Study Using the DARPA CAGE-2 Challenge

As cyber attacks grow more sophisticated, defenders need autonomous systems that are fast, adaptable, and explainable. Over the last decade, various strategies have been proposed for automated cyber defence (as opposed to static rule-based or signature-based), including those based on reinforcement learning (RL). Researchers have proposed various algorithms to improve RL-based defenders and evaluated them using simulation-based frameworks like the DARPA CAGE-2. Although RL showed promise, it has many limitations, for example, the lack of a realworld training environment and the need for extensive training, which is time-consuming. In this case study, we investigate whether LLM agents can be used instead of RL agents to automate cyber defence. Large Language Models (LLMs) can reason over natural language and generalize from extensive pretraining. They are attractive for cyber defence because they can read textbased system states and make human-like, explainable decisions. We propose a unique way to convert CAGE-2 states to natural language and a domain-specific fine-tuning method that improve the average reward and reduce hallucination significantly, beating existing RL-based agents and state-of-the-art LLM agents.

Arijit Diganto, S. Lohrasbi, Euclides Carlos Pinto Neto et al. · 0 citations
Open access Aug 2026

L-ARLPT: An LLM-Augmented Reinforcement Learning Framework for Autonomous Penetration Testing

A Large Language Model-enhanced Autonomous Reinforcement Learning Penetration Testing framework that leverages the domain knowledge embedded in a Large Language Model to perform tactical planning, thereby pruning the original action space into a compact set of candidate actions.

Rufeng Zhan, Junyi Zhu, Yinghui Xu et al. · 0 citations
Jul 2026

Agentic AI for Offensive Security: LLM-guided Autonomous Red Teaming in a Limited Cyber-range Environment

Agentic artificial intelligence (AI) is increasingly being explored for automating offensive security and red teaming tasks, enabling systems that can coordinate multi-step cyber operations through structured decision-making. While prior research has investigated reinforcement learning (RL) agents and large language models (LLMs) for penetration testing, most studies are evaluated in simulated or abstract environments, with limited empirical validation in real cyber-range settings. This paper presents a controlled experimental evaluation of an LLM-guided offensive security pipeline against deterministic scripted baselines in a cyber-range environment. Using a vulnerable Kioptrix virtual machine and a Kali Linux attacker, we implement three deterministic pipelines, fixed-path, sequential and rule-based, alongside two configurations of an LLM-guided agent: an initial version (LLM V1) and a refined constrained controller (LLM V2). All approaches operate within a restricted and auditable action space executed through predefined tools. Across repeated trials, the initial LLM configuration exhibits reduced reliability and increased execution cost due to exploratory behaviour. In contrast, the refined controller achieves a 100% success rate, reduces execution steps, eliminates wasted actions and matches the efficiency of rule-based automation. These results show that, within a controlled cyber-range environment, LLM-guided agents can approximate deterministic performance when appropriate constraints are applied. This suggests that agent-based approaches to offensive security can support semi-autonomous red teaming workflows, provided that decision-making is governed by structured control policies.

Atif Chowhan, Yasir Hamid · 0 citations