A lifecycle-based framework utilizing the collaborative “red-blue-purple” teaming model is introduced, utilizing the collaborative “red-blue-purple” teaming model to ensure that clinical guardrails remain robust against evolving adversarial tactics.
Abstract
The emergence of large language models (LLMs) provides new avenues for clinical support in healthcare and dentistry. However, these models often exhibit unpredictable behaviours when challenged by adversarial or misleading inputs. Recent data indicate that nearly 20% of LLM outputs contain safety risks or biases, necessitating rigorous evaluation prior to clinical use. This review examines AI red teaming, a systematic approach for identifying system vulnerabilities through simulated attacks. It details methodological approaches and outcome measures while proposing a structured framework to integrate these safety evaluations into the clinical AI lifecycle. This review focuses on prompt-based attacks, such as prompt injection and jailbreaking, which are highly relevant in medical settings. It evaluates various testing strategies, including manual expert reviews, automated “attacker” models, and hybrid human-in-the-loop systems. A lifecycle-based framework is introduced, utilizing the collaborative “red-blue-purple” teaming model. This approach spans pre-deployment testing, live deployment monitoring, and iterative review audits to ensure that clinical guardrails remain robust against evolving adversarial tactics. Safe implementation of LLMs in dentistry and healthcare requires continuous, iterative adversarial testing rather than static assessments. Success depends on standardized protocols, multidisciplinary collaboration between clinicians and AI researchers, and the development of domain-specific benchmarks. Bridging existing regulatory gaps through these structured frameworks is vital for ensuring LLMs are safe, reliable, and clinically fit for patient care.
The Review examines rapid LLM adoption in clinical care, outlining emerging security and safety risks across development stages, key protective layers, clinically relevant threats and current mitigation responsibilities in a single integrated framework.
J. Clusmann, O. Freyer, Max Ostermann et al.· Nature· 0 citations
A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.
Jiazhen Pan, Bailiang Jian, Paul Hager et al.· Nature Health· 0 citations
A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.
Noah Wright· International Journal of Mod...· 0 citations
It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
K. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
OBJECTIVES
Evaluate whether general-purpose large language models (LLMs) demonstrate competencies suitable for antimicrobial stewardship (AMS) support and characterize their failure modes.
METHODS
Cross-sectional evaluation of seven LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, Llama-3.3-70b-instruct, Qwen 2.5-72b-instruct, DeepSeek-chat-v3.1) using 30 clinical scenarios mapped to ESCMID AMS competency frameworks. Scenarios included deliberate traps for fabrication and dangerous recommendations. Six AMS experts from the Netherlands and Spain performed blinded dual evaluation using content scores (0-5 scale) and binary safety flags for fabrication and danger. Standard and incentivizing prompt framings were compared.
RESULTS
Four commercial models achieved mean content scores above 3.9/5.0: Claude Sonnet 4.5 (4.06), Gemini 2.5 Pro (3.96), Grok 4 (3.96), and GPT-5 (3.94). Open-weight models scored significantly lower (2.94-3.57). No model achieved more than 63% responses free of fabrication or danger flags. However, fabrication did not impair clinical utility in non-trap scenarios (all within-category comparisons p>0.20). Danger flags ranged from 6.7% to 16.7% across models, with no significant difference between commercial and open-weight models. Incentivizing prompts were associated with a consistent 0.48-point-content score improvement (p=0.006), though significance attenuated after accounting for scenario-level clustering. Evaluators endorsed LLMs as useful AMS support tools with moderate supervision (5/6), identifying documentation preparation and trainee education as promising applications.
CONCLUSIONS
Medically untrained LLMs demonstrate competencies suitable for supervised AMS support. Fabrication remains the central safety challenge and requires verification workflows; danger, though less frequent (6.7-16.7%), concentrated in identifiable and therefore mitigable failure modes. Non-clinical stewardship tasks (education, documentation, communication) can benefit now, whereas clinical recommendations require expert oversight. Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk.
Ángela Abejez-Arrizabalaga, Galadriel Pellejero-Sagastizabal, Rocío Aznar-Gimeno et al.· Clinical Microbiology and In...· 0 citations
The rapid deployment of large language models (LLMs) in healthcare settings makes the reliability of their built-in guardrails against malicious queries a question of urgent practical consequence. Yet the robustness of these mechanisms against deliberate misuse (in the healthcare context) remains poorly understood. In this paper, we investigate this question empirically, using AI-assisted medical note manipulation as a concrete case study. We make four novel contributions. First, we develop a reproducible manipulation pipeline that takes publicly available seed medical note templates and use commercial LLMs to produce customized manipulated notes by substituting patient names, provider identities, dates, and medical conditions across multiple model families, input formats, and prompt phrasings. Second, we conduct a systematic empirical evaluation of LLM guardrail robustness for medical note manipulation. Our experimental results reveal substantial weaknesses and inconsistencies in contemporary commercial LLM guardrails, including low refusal rates for several model families. Third, we utilize a combination of automated metrics and human annotation-based metrics to assess the correctness of requested manipulations. Fourth, we conduct a user-study to assess the believability of manipulated medical notes, finding that the best manipulations are visually indistinguishable from original documents to human raters. Finally, we discuss implications for responsible guardrail design in LLMs, AI safety policies, and the broader ethics of deploying LLMs in healthcare settings.