Cloud-Deployed Adaptive SSH Honeypot with Reinforcement Learning and LLM-based Dynamic Response Generation
Abstract
A honeypot is a decoy computer system that is intentionally deployed to attract cyber attackers and gather threat intelligence in a controlled environment. However, honeypot frameworks like the popularly used Cowrie SSH honeypot completely rely on static responses that do not change based on the attacker’s behavior. This is because experienced attackers easily recognize these unrealistic responses within a few seconds and disconnect their session right away. This inability to change and improve is the biggest drawback of honeypot technology. In this paper, we propose the design of the Adaptive SSH Honeypot framework, which incorporates Q-learning, Reinforcement Learning, and the Llama 3.1 8B Large Language Model via the Groq API with the Cowrie framework. The system is implemented on a three-VM AWS EC2 cloud computing infrastructure setup, wherein incoming SSH connections to port 22 are silently redirected via iptables to Cowrie at port 2222. The state of each attacker’s command is categorized into a four-dimensional tuple state (C, H, B, R), of 96 possible states. The Q-learning agent uses six deception actions via the ε-greedy policy to update its Q-table via the Bellman equation after each command. The LLM will dynamically produce three candidate responses for a Linux shell command and return the top-ranked result. Honeytoken bait files with fake credentials, API keys, and DB passwords are strategically placed throughout the imitation filesystem to detect attacker intent. The Paramiko sandbox simulation and offline JSON log-based training approach eliminates the problem of cold start attackers prior to actual attacker connections. The experimental results on the real-world AWS EC2 deployment show a substantial performance gain of +187.1% in average session duration and +154.8% in commands per session over a static Cowrie baseline, with an average of 0.5 honeytoken hits per session.