With the increasingly aggressive cyber threat landscape for governments, businesses, and institutions, as information and/or cybersecurity implementations are increasingly under scrutiny by regulators, it has been pointed out that governance failure is one of the major reasons for a weakened cybersecurity posture. A major component of Cyber/information security governance is the development, adoption, and implementation of a comprehensive information and/or cyber security policy document. The policy document must be in compliance with international or national standards and, if possible, with regulatory guidelines. However, it is often observed that policy documents are often incomplete with respect to industry standards or regulations and require revision when subjected to a thorough audit. Identifying the gaps between the controls and processes documented in the policy and those required in the regulations or standards necessitates extensive manual effort. The advent of Generative AI tools such as Large Language Models (LLMs) led to use of LLMs and Agentic AI tools to automate such compliance checks, as seen in a few research publications in recent times. However, such reported use of LLMs are experimented with high resource environments such as expensive GPUs and memory based servers. For smaller organizations such expensive compute platform may not be easily available. In this article, we benchmark the compliance checking tasks on LLMs that do not require GPU and high memory usage and the effectiveness of such resource constrained LLMs in compliance checking. Our experiments demonstrated that the low resource LLMs can provide good agreement/accuracy in compliance checking of policy documents against standards by experimenting with ISO 27002:2022 controls against multiple policy documents.
: As cybersecurity regulations such as ISO/IEC 27001 and the NIS2 Directive continue to expand in scope and complexity, organizations face growing challenges in translating regulatory obligations into actionable security policies and audit-ready evidence. Conventional compliance approaches rely on manual interpretation of regulatory texts, fragmented documentation repositories, and ad hoc audit preparation, introducing operational bottlenecks and exposing organizations to non-compliance risks. This paper presents a compliance management platform that operationalizes regulatory requirements through structured, expert-guided control implementation. It combines NLP extraction with human-supervised annotation to convert regulatory texts into machine-readable frameworks, enabling multi-framework management (ISO/IEC 27001:2022 and NIS2), control mapping, evidence tracking, and role-based audit workflows. In a task-based usability study with twelve participants, the platform scored 83.3 on the System Usability Scale (SUS), rated “excellent,” indicating that embedded guidance can reduce expertise barriers in cybersecurity compliance management.
M. Andrade, J. Almeida, José L. Oliveira· Proceedings of the 23rd Inte...· 0 citations
The findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly, which has important implications for the design and evaluation of AI-assisted security decision-support systems.
The growing complexity and frequency of cyberattacks make cybersecurity risk assessment an increasingly demanding task for organisations, requiring substantial expertise, resources, and adherence to established standards. This work explores the applicability of Large Language Model (LLM) to cybersecurity risk assessment, with a focus on threat identification and risk scoring. The paper presents a standalone consistency analysis across five models, measuring accuracy and stability under lexical, structural, and noisy prompt perturbations using an OWASP-oriented rubric. Building on the analysis results, we present a modular LLM-based system that combines Retrieval-Augmented Generation, MITRE ATT&CK-Aligned threat evaluation, rubric-constrained risk scoring, and a Judge Reviewer, orchestrated through a Beliefs–Desires–Intentions control loop. The validation against incidents from the VERIS and EuRepoC datasets highlights limitations and weaknesses, and allows identifying the architectural and structural mitigations that can reduce prompt sensitivity in LLM-based risk assessment.
The progressive convergence of Information Technology (IT) and Operational Technology (OT) environments has introduced new cybersecurity challenges for industrial and critical infrastructure systems. This work presents a generalized OT cybersecurity architecture that combines the Purdue reference model with Zero Trust principles to enforce strict segmentation, continuous verification, and controlled information flows across IT/OT boundaries. The architecture incorporates Artificial Intelligence (AI)-driven monitoring to support anomaly detection, contextual risk assessment, and automated response mechanisms under operational constraints. Additionally, a governance and assurance layer is discussed, aligning AI-enabled security functions with recognized risk management frameworks and auditable controls to ensure trustworthiness, resilience, and operational sustainability in high-impact industrial deployments. The proposal is further contextualized with prior AI-RMFgoverned IoT-as-a-Service and OWASP ML05 middleware contributions that address AI-enabled IoT security, model protection, and governance requirements.
Yair Rivera Julio, Ángel D. Pinto-Mangones, Frank A. Ibarra et al.· 2026 6th International Confe...· 0 citations
To improve cybersecurity across industries, Cyber Threat Intelligence (CTI) is becoming increasingly crucial. This systematic review explores how CTI practices are evolving in response to advancements in Artificial Intelligence (AI), particularly in the context of Large Language Models (LLMs). We examined 61 peer-reviewed studies using the PRISMA methodology, which demonstrates a strict selection procedure founded on specified inclusion, exclusion, and quality standards. This approach aligns with the scope of similar systematic reviews in the field of cyber threat intelligence. The review provides a comparative synthesis of CTI research capabilities across threat detection and prediction, attribution, forecasting, and automated reporting. We classify these approaches into three categories: conventional methods, those enhanced by AI and Machine Learning, and those based on LLMs. Our findings indicate that LLMs offer significant advantages in contextual reasoning, processing unstructured threat intelligence, and generating actionable mitigation plans. However, challenges such as model explainability, data privacy, system interoperability, and standardization impede their integration into operational environments. In addition to highlighting the potential and practical limitations of LLMs in CTI, this study identifies research gaps and proposes methods to create scalable, secure, and flexible CTI systems that support real-time cyber defense.
Hilalah Alturkistani, Abdul Ghafar Jaafar, S. Chuprat et al.· International journal of res...· 0 citations
Most vulnerability pipelines remain predictioncentric: they output scores or labels and defer decisions to engineers, even when outputs are compressed, imbalanced, or unreliable under Continuous Integration and Continuous Deployment (CI/CD) shift. We introduce AEGIS (Autonomous Enhanced Guardian for Intelligent Security), a policy-governed framework that treats vulnerability management as a constrained CI/CD decision process in which learned signals serve as evidence and are translated into admissible actions under explicit constraints. AEGIS combines graph-based risk estimation, epistemic uncertainty via stochastic inference, and a symbolic policy guard that maps evidence to auditable decisions: Block, Warn, and Pass. A key design principle is separation of concerns: perception estimates risk while governance determines admissible actions. Irreversible automation is permitted only when risk is high and uncertainty is low, while uncertain cases are deferred to controlled review. This separation makes it possible to revise policy thresholds, cost assumptions, and review budgets without retraining the perception model. We evaluate AEGIS on an extreme-imbalance patch stream used as a stress-test setting and an expanded multi-project dataset that enables more stable estimation of decision outcomes. The evaluation reports the policy thresholds, model settings, symbolic predicates, ablations, and sensitivity settings used in policy replay. Results provide preliminary evidence that policygoverned control supports more interpretable decision behavior under uncertainty, while enabling controlled trade-offs between automation, safety, and review load.
Imad Abdallah, Lunjin Lu· Annual International Compute...· 0 citations