Development of a Framework for Evaluating Large Language Model Safety and Reliability: a Proof-of-Concept Evaluation
Large language models (LLMs) are entering clinical decision support faster than methodology can characterise their safety. Aggregate accuracy treats all errors as interchangeable and cannot support safe deployment under Software as a Medical Device (SaMD) and EU AI Act frameworks. To develop and demonstrate a framework...