Skip to content
Conference Open access

A Hierarchical Evaluation Framework for LLM-driven Threat Modelling Tools

2026 · Proceedings of the 23rd International Conference on Security and Cryptography · pp. 689-700 · 0 citations · 38 references

TL;DR

A systematic evaluation framework for LLM-driven threat modelling tools to support tool selection, observing the general LLM-integration, governance risks, and allowing for comparison of tool output is introduced.

Abstract

: AI adoption has accelerated with the rise of LLMs, and people within the field of security are increasingly exploring their practical value. Threat modelling is central to secure system development, yet it remains largely manual and the value of LLM-driven tools is unclear. Even when LLMs prove useful, selecting the right one can be more challenging than using it. This paper introduces a systematic evaluation framework for LLM-driven threat modelling tools to support tool selection, observing the general LLM-integration, governance risks, and allowing for comparison of tool output. Using Goal-Question-Metric, we derive evaluation metrics and show the value of the framework on a set of state-of-the-art LLM-driven threat modelling tools. The results show our metrics distinguish both performance and governance risks, providing a basis for organisations to ensure automation strengthens rather than burdens their threat modelling process.

Read PDF

Similar papers

Aug 2026

ARAMIS: A unified and scalable methodology for industrial cyber security risk assessment

The rapid digitisation of critical infrastructure has made traditional fragmented risk assessment practices increasingly challenging to scale. For global industrial leaders managing hundreds of diverse projects, there is a real need for a unified methodology that ensures technical rigour, cross-project reproducibility, and scalability. This paper introduces the Advanced Risk Assessment Methodology for Industrial Systems (ARAMIS), an innovative framework developed through a strategic partnership between Airbus Protect and Alstom. ARAMIS merges the structured, requirement-driven security levels of ISA/IEC 62443 with the scenario-based approach of Expression des Besoins et Identification des Objectifs de Sécurité Risk Manager (EBIOS RM). The paper details the five-module structure of ARAMIS, its unique multilayered modelling of operational scenarios and its algorithmic approach to calculating security levels target (SL-T). Finally, it discusses the implementation of the methodology within the Fence risk management tool to ensure seamless reproducibility and knowledge capitalisation across global project portfolios. This article is also included in The Business & Management Collection which can be accessed at https://hstalks.com/business/.

Serge Benoliel, Florence Foudrain · 0 citations
Conference Jul 2026

Limitations of Large Language Models for Cybersecurity Risk Assessment

The growing complexity and frequency of cyberattacks make cybersecurity risk assessment an increasingly demanding task for organisations, requiring substantial expertise, resources, and adherence to established standards. This work explores the applicability of Large Language Model (LLM) to cybersecurity risk assessment, with a focus on threat identification and risk scoring. The paper presents a standalone consistency analysis across five models, measuring accuracy and stability under lexical, structural, and noisy prompt perturbations using an OWASP-oriented rubric. Building on the analysis results, we present a modular LLM-based system that combines Retrieval-Augmented Generation, MITRE ATT&CK-Aligned threat evaluation, rubric-constrained risk scoring, and a Judge Reviewer, orchestrated through a Beliefs–Desires–Intentions control loop. The validation against incidents from the VERIS and EuRepoC datasets highlights limitations and weaknesses, and allows identifying the architectural and structural mitigations that can reduce prompt sensitivity in LLM-based risk assessment.

Monica Chingate, Gabriele Gatti, Cataldo Basile · 0 citations
Review Aug 2026

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.

Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj et al. · 0 citations
Jul 2026

Risk Management Framework for LLM-Enabled Identity Tools: Mapping OWASP LLM Top 10 to NIST CSF 2.0

Large language models (LLMs) are increasingly embedded in identity and access management (IAM) tools that support workflows such as account recovery, access request triage, provisioning, policy interpretation, and privileged access handling. In these settings, security risk is often dominated not by the model in isolation but by workflow exposure: who can trigger the system, what identity data and systems it can access, what actions it can execute, and which governance safeguards constrain behavior. We present a workflow-centric risk assessment method for LLM-enabled identity tools that uses the OWASP Top 10 for LLM Applications as a threat taxonomy and NIST Cybersecurity Framework (CSF) 2.0 as a governance outcomes layer. We instantiate OWASP categories as a compact library of 20 IAM-relevant, workflow-anchored threat scenarios and assign inherent risk scores per scenario. For each scenario, we map relevant CSF 2.0 Categories/Subcategories and score outcome coverage across tool/architecture archetypes. Residual risk is estimated by scaling inherent risk by uncovered outcome coverage, yielding an auditable signal to prioritize risk treatment and determine when workflows require mandatory human escalation versus safe automation. We additionally report backend sensitivity results under a fixed scenario suite and scoring rubric to quantify variance across LLM backends.

Sanaa S. Mironov, Shahmir Rizvi, B. Shariati · 0 citations
Preprint Jul 2026

A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms

As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority. While a myriad of high-level ethical guidelines have emerged, criticism persists that these frameworks remain abstract and lack concrete mechanisms for implementation. This paper conducts a critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset from the OECD. Through empirical mapping and descriptive comparative analysis, we identify significant asymmetries in ethical focus, lifecycle coverage, stakeholder targeting, and tool typology. Our findings show a strong emphasis on fairness, transparency, and robustness, with comparatively little attention paid to explainability, digital security, and environmental sustainability. Moreover, most tools and certifications concentrate on post-development stages, with limited guidance for early design or data collection phases. Educational initiatives and policy engagement are notably underdeveloped, suggesting that current TAI efforts are dominated by technical and procedural measures within industry contexts. We argue that bridging the persistent chasm between AI principles and practice requires expanding ethical objectives, embedding ethics across the AI lifecycle, and fostering broader multi-stakeholder participation. This study provides both a diagnosis of existing implementation gaps and actionable recommendations for advancing more holistic, inclusive, and enforceable AI governance

Michael Papademas, Xenia Ziouvelou, K. Karpouzis et al. · 0 citations
Review Aug 2026

Demystifying cyber threat intelligence: A first-principles approach to capability development and vendor evaluation

The case is made for a first-principles approach that CTI teams can adopt as an unbiased anchor to guide their decisions around establishing an adequate CTI capability, and pragmatic recommendations to assist CTI teams with qualifying their prospective vendors to ensure good fit are offered.

Aaron Aubrey Ng · 0 citations