Preprint
Jul 2026
Cybersecurity Detection Classification with Reasoning-enabled Language Models
This work trains a chain-of-thought reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards, and shows that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.
Amol Khanna, Manu Nandan, Cristian Viorel Popa et al.
· 0 citations