Jul 2026· Philosophical transactions. Series A, Mathematical, physical, and engineering sciences· Vol 384 2324· 1 citation· 32 references
Medicine
TL;DR
The reality of deploying artificial intelligence in safety-critical systems, such as autonomous vehicles, medical diagnoses and weather forecasting, is considered, including how an AI's mathematical properties relate to its benefit and risk profile.
Abstract
We consider the reality of deploying artificial intelligence (AI) in safety-critical systems, such as autonomous vehicles, medical diagnoses and weather forecasting. Our discussion is grounded in the mathematical nature of AI systems, including how an AI's mathematical properties relate to its benefit and risk profile. Benefits include the ability to learn models from data even when no physical model exists, increased automation and enhanced speed compared with traditional approaches. Risks of AI include its opaque (mis-)understanding of the world, failures on out-of-distribution inputs, its insatiable appetite for data and computation and the ongoing challenge of aligning the AI's objectives with human values. Such risks are potentially manageable with clear-eyed expectations, and our hope in this work is to clarify what can be expected. This article is part of the theme issue 'Safe, secure and robust AI for safety-critical systems'.
The rapid integrationof artificial intelligence (AI) into everyday life has brought transformative benefits, but it has also sharpened concerns surrounding safety, security and robustness, particularly in safety-critical domains. This special issue seeks to address these interconnected challenges in a holistic manner, recognizing that dependable AI systems must be not only performant, but also trustworthy across a wide range of operational conditions. The articles bring together contributions in the fields of computer science and mathematics, while also addressing concerns about ethics, accountability and regulation.
This article is part of the theme issue ‘Safe, secure and robust AI for safety-critical systems’.
Ajitha Rajan, D. Higham· Philosophical Transactions o...· 0 citations
Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.
T. Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal et al.· 0 citations
AI alignment is generally associated with ethics and social aspects. For high-risk AI systems, principles such as safety, stability and performance are crucial for their adoption in real-world application, to build trust and increase efficiency of a process. By safety it is meant both safety of the system itself but also safety of the process and environment. Any decision taken by a high-risk AI system should preserve safety, ensure continuous operation, real-time functioning, smooth control, and improve process efficiency. Human oversight needs to be continuously ensured, both to preserve safety of the process in case of transition from autonomous to manual mode in case of failures, but also to allow for contextual knowledge of the decisions. All these objectives are included in the AI operational alignment. High-risk AI systems designed to automatize complex processes in critical environments are continuously adapting to dynamic contexts, thus a static alignment might not be sufficient to properly assess their behaviour. The paper describes an approach to automated dynamic operational alignment in a complex high-risk process, exemplified by a case study on autonomous drilling. Possible sources of AI misalignments in this case are discussed and their potential implications.
Rodica Mihai, B. Daireaux, E. Cayeux· Scientific Reports· 0 citations
This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance and illustrates how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications.
Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann et al.· 0 citations
Background: The safety of Autonomous Driving (AD) remains a barrier to its widespread adoption, as evidenced by recent incidents. Factors such as a complex environment, evolving technologies, and shifting regulatory and customer requirements necessitate continuous monitoring and improvement of AD software. This is a process that may favor software and system engineering supported by DevOps. The iterative nature of the DevOps process is crucial, serving two purposes: satisfying customer demands through continuous im- provement of the function and providing a framework for timely responses to unknown bugs or incidents. However, any update to the software must follow rigorous safety processes prescribed by standards, regulations, and the state of the art in industry. Incorporating these safety activities into the DevOps forms an iterative process called DevSafeOps. These necessary activities are vital for safety assurance, and may inherently lead to a compromise in rapidity.Research Goal: In this work, we identify the challenges of rapid DevSafeOps in AD development and explore existing solutions. Subsequently, we propose multiple approaches for accelerating safety analysis, requirements engineering, code generation, and synthetic data generation in DevSafeOps cycles.Methods: Diverse research methods are utilized to address each research objective. Interview studies and a systematic literature review are conducted to identify the challenges, research gaps, and existing approaches. Then, design science, interview study, case study, and experimentation are employed to design and evaluate new approaches to address our research goal.Results: Initially, the challenges and research gaps related to each essential activity for the safety of automated driving are identified (Papers A and B), together with the proposed solutions presented in the literature (Paper B). Two approaches are proposed to accelerate the design phase (i.e., analysis and requirements engineering) as an initial step in DevSafeOps. We adapt System Theoretic Process Analysis (STPA) to enable distributed development within automotive system engineering (Paper C). As an alternative approach, a Large Language Model (LLM)-based multi-agent Hazard Analysis and Risk Assessment (HARA) prototype is proposed and evaluated to enable automation (Papers D and E). The rule-based software-implementation phase is accelerated through LLM-based code generation conducted through a conversation in a simulation environment (Papers F and G). To connect the design phase to operation, a vision-language model (VLM) is employed to enable rapid closed-loop DevSafeOps (Paper H). In parallel, a complementary solution is introduced to address the specific needs of Machine Learning (ML)-based software development. As data act as requirements for ML, it is crucial to generate data in a controlled manner to obtain a su!ciently sized population of critical scenarios for training the expected behavior. Hence, through synthetic data generation using three-dimensional Gaussian Splatting (3DGS), ML-based software development in the DevSafeOps cycle is covered (Paper I).Conclusions: This thesis first identifies multiple challenges in achieving rapid DevSafeOps in AD development and then proposes several approaches for addressing these challenges across different phases of the DevSafeOps cycle. To accelerate the design phase, we introduce an adaptation of STPA for multiparty distributed development and employ multi-agent LLMs as a parallel approach for HARA. We further examine how LLMs and VLMs can support safety concept design, code generation, and monitoring activities with reduced engineer involvement, while defining necessary safeguarding measures. Finally, we investigate 3DGS as an effective and rapid DataOps technique within DevSafeOps, enabling improved data generation and augmentation for ML-based software development.