This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance and illustrates how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications.
Abstract
Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, the European Union Aviation Safety Agency (EASA) emphasizes the need to demonstrate the representativeness and completeness of the Operational Design Domain (ODD) and the associated data distributions used during development and verification. Despite this requirement, a structured engineering process for defining target distributions and evaluating representativeness within ODDs remains largely unexplored. This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance. Starting from the methodical identification of suitable target distributions, a process flow is proposed that guides developers from ODD definition and parameter distribution modeling to the quantitative assessment and interpretation of coverage results with respect to EASA's learning assurance objectives. As quantitative measures, the chi-squared goodness-of-fit test is examined and found unsuitable for the large data sets arising in this setting, leading to the adoption of the Kullback--Leibler divergence and Cram\'er's $V$ for the representativeness assessment. The method is demonstrated using the example of AI-based airborne collision avoidance, employing experimental data from previous Horizontal Collision Avoidance System (HCAS) and Vertical Collision Avoidance System (VCAS) simulations. The results illustrate how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications and contribute toward a systematic Safety-by-Design AI engineering process aligned with emerging EASA guidance.
The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies on numerous safety-critical systems. For future safety-critical AI-based systems, EASA requires a Safety-by-Design approach, which can be achieved by using Safety Nets that combine neural network compression with lookup tables to ensure 100 % correct runtime behavior across the discretized operational design domain. Although Safety Nets have been studied, no comprehensive study of their performance characteristics and system design trade-offs has been conducted. This work presents the first systematic analysis of the trade-off between neural network and lookup table size in Safety Nets. By systematically comparing neural networks with diverse architectures, this study identifies optimal design parameters that minimize overall storage and memory requirements while maintaining certification compliance. Results demonstrate that architectures with 3 to 5 hidden layers, each with approximately 50 to 100 nodes, combined with one-hot encoding, achieve the best balance. In these configurations, neural networks accurately represent at least 97 % of the data, while compact lookup tables handle the remaining errors. The resulting Safety Nets reduce the system size by almost three orders of magnitude, fitting within the memory budget of current avionics hardware while guaranteeing 100 % correct outputs across the entire discretized input space, as required by EASA guidelines. This work provides the first-ever open-source implementation of Safety Nets for HCAS and VCAS with replicable results, demonstrating a practical pathway toward certifiable AI-based systems in aviation and establishing Safety Nets as a viable Safety-by-Design solution for safety-critical applications.
Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann et al.· 0 citations
Safety of Automated Driving Systems (ADSs) is arguably one of the main remaining
barriers before widespread market deployment. While there exists a plethora of
methods for planning a trajectory that fulfils certain constraints, what those
constraints should look like, to enable effective planning of safe trajectories,
is still being discussed. In this article, we generalize the concept of
Precautionary Safety (PCS) and present a framework providing constraints on the
tactical and operational decisions of the ADS. Such constraints consider the
ADS’ capabilities, the external conditions, knowledge of statistically relevant
events and behaviors of other traffic actors, as well as the controllability of
these events. The proposed framework enables assessment of the statistical
fulfilment of quantitative risk acceptance criteria (QRACs), including
requirements on accident, injury, and fatality rates. The framework further
provides a means to dynamically adapt the constraints used for trajectory
planning, i.e., to adapt the driving to the situation at hand. A case study,
considering a possible collision scenario with a jaywalking pedestrian and a
rear-end collision with a trailing vehicle, is provided to showcase the
applicability and usefulness of the presented framework. The simulation-based
case study displays the safety benefits from considering QRACs with multiple
injury risk levels and further shows how the proposed PCS framework can be
applied in practice.
Magnus Gyllenhammar, Gabriel Rodrigues de Campos, Fredrik Sandblom et al.· SAE International Journal of...· 0 citations
The use of Artificial Intelligence (AI) is becoming established as a component of safety-critical engineering systems such as the autonomous transportation systems, energy infrastructure, industrial automation, and structural monitoring. As much as AI models and especially deep neural networks prove to be more accurate in prediction, their black box nature create serious issues when it comes to interpretability and safety guarantees. In safety critical areas explainability is not only a good idea, but a precursor to regulatory compliance, accountability and mitigation of hazards. This paper creates a comprehensive mathematical conceptualization of using interpretability metrics and formal safety verification protocols of AI-based engineering systems. Based on the recent developments in explainable artificial intelligence (XAI), mechanistic interpretability, and formal verification theory, we operationalize interpretability as a measurable functional characterization of model structures and safety verification as meet-in-the-middle of reachable state spaces. Our abstraction of causal circuits and verification of time logics constraints using Shapley-based attribution functions make our integrated optimisation model offer explanation fidelity and safety invariants. The framework illustrates how the measures of interpretability can be integrated in the form of the side-constraints in formal verification pipelines, such that the model decision-making can be transparent and safely proven. Findings suggest that the metrics of coupling explanation coupled with reachability analysis can minimize unsafe decision regions and enhance calibration of the trust. The paper adds a mathematically based architecture that can be able to fill in the gap between heuristic XAI methods and serious engineering safety requirements. It is discussed in implications to aerospace, autonomous systems and industrial control environment as well as computational trade-offs and scalability issues. The model suggested creates a direction towards the certifiable AI systems in which interpretability and safety are not set at cross purposes.
Moses Adeolu Agoi, Emmanuel Taiwo Agoi, Oluwanifemi Opeyemi Agoi et al.· International Journal of App...· 0 citations
Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.
T. Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal et al.· 0 citations
The reality of deploying artificial intelligence in safety-critical systems, such as autonomous vehicles, medical diagnoses and weather forecasting, is considered, including how an AI's mathematical properties relate to its benefit and risk profile.
T. Kolda· Philosophical transactions....· 1 citation