2026· International Journal of Advanced Computer Science and Applications· 0 citations· 56 references
TL;DR
A conceptual Clinical Co-pilot Framework is proposed to position GenAI as a collaborative partner that supports clinicians rather than replaces them, which provides a conceptual basis for future empirical validation and may help inform the responsible implementation of GenAI in healthcare.
Abstract
Generative artificial intelligence (GenAI), large language models (LLMs), and multimodal foundation models are rapidly transforming healthcare by extending artificial intelligence beyond traditional predictive analytics toward collaborative clinical intelligence. Recent advances have enabled applications in clinical decision support, medical documentation, workflow optimization, patient communication, and personalized care planning. However, the existing literature remains fragmented across technical evaluations, specialty-specific applications, and governance discussions, resulting in a limited understanding of how these technologies collectively function within clinical environments. This study conducted a systematic review and thematic synthesis to examine the emerging role of GenAI as a clinical co-pilot in healthcare. Following the PRISMA 2020 framework, literature was retrieved from Web of Science, PubMed, and Scopus databases. A total of 3,124 records were identified, and 41 studies published between 2023 and March 2026 met the eligibility criteria and were included in the final qualitative synthesis. Quality assessment indicated that 87.8% of the included studies were classified as moderate or high quality, providing a robust methodological basis for the thematic synthesis. Thematic analysis revealed three interrelated domains underlying the evolution of GenAI-enabled clinical intelligence: 1) Perception and Fact Anchoring, involving multimodal data integration, retrieval-augmented generation (RAG), and domain-specific medical intelligence; 2) Clinical Agency and Collaboration, encompassing ambient digital scribing, agentic clinical decision support, workflow augmentation, and personalized care; and 3) Governance and Responsibility, including human-in-the-loop oversight, privacy protection, regulatory compliance, fairness, and trustworthy AI practices. Based on these findings, a conceptual Clinical Co-pilot Framework is proposed to position GenAI as a collaborative partner that supports clinicians rather than replaces them. The framework provides a conceptual basis for future empirical validation and may help inform the responsible implementation of GenAI in healthcare.
Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· 0 citations
Rare diseases create distinctive challenges for scientific communication and health-system organization because of small patient populations, heterogeneous phenotypes, limited natural history data, fragmented referral pathways, and persistent diagnostic delays. This narrative review aimed to synthesize the literature on Medical Science Liaison (MSL) practice and rare-disease ecosystems and to propose a pragmatic framework for Field Medical excellence. A targeted search of PubMed and publicly available professional guidance repositories was conducted through July 2026 using terms related to Medical Affairs, MSL competencies, rare diseases, diagnostic delay, real-world evidence, center development, scientific insights, and impact measurement. Sources were selected for conceptual relevance rather than exhaustive coverage. The synthesis identified six interdependent domains through which Field Medical can create scientific and system-level value: Scientific Mastery, Insight Intelligence, Center Development, System Navigation, Strategic Partnership, and Impact Measurement. The framework emphasizes scientific independence, contextual interpretation of uncertain evidence, longitudinal stakeholder engagement, translation of field insights into action, and evaluation through outcome-oriented indicators rather than activity counts alone. It is intended as an organizing model, not as a validated competency instrument. Future studies should test its content validity, feasibility, and association with measurable improvements in diagnostic readiness, research feasibility, stakeholder capability, and patient-centered care pathways.
Ilma Nascimento· Research, Society and Develo...· 0 citations
This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.
Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al.· International Journal of Sur...· 0 citations
INTRODUCTION
Objective Structured Clinical Examinations (OSCEs) are widely used to assess clinical competence, but face challenges related to examiner workload, scoring variability, delayed feedback, and resource demands. Although AI may address these constraints and support precision medical education, the evidence remains fragmented. This scoping review maps AI applications in OSCEs.
METHODS
We followed PRISMA-ScR. We searched MEDLINE, Scopus, Embase, Web of Science, ERIC, LILACS, and IEEE Xplore from inception to June 2025. Eligible studies examined AI within any phase of an OSCE in health professions education. Data charting captured AI form, technology, OSCE phase, competencies, outcomes, faculty and resource implications, and ethical/governance issues. Synthesis used a hybrid approach: deductive coding with FACETS, SAMR, and operationalized P4 properties, plus inductive coding for emergent themes. Findings were stratified by evidence maturity rather than formal quality scoring.
RESULTS
Of 421 records screened, 22 studies were included. AI was used for learner preparation, station/material construction, scoring/evaluation, and operational delivery. Benefits were strongest for grading, feedback speed, and consistency in structured tasks, but weaker for relational competencies. Most applications reflected SAMR Augmentation or Modification. Personalization dominated P4 alignment. Ethical concerns centered on privacy, bias, accuracy, transparency, access, and human oversight.
DISCUSSION AND CONCLUSION
AI currently augments rather than transforms OSCE assessment, performing best in structured, observable tasks and least well in relational, situated, and culturally mediated competencies. Claims that AI delivers precision medical education through OSCEs are not yet supported by evidence; alignment with P4 is partial and conditional. Realizing the potential of AI in OSCEs will require human-in-the-loop governance with concrete safeguards and equity-focused implementation in resource-limited settings.
Sergio Andrés León-Ariza, María Camila Orobio-Pinzón, Héctor Miguel Ibáñez-Gutiérrez et al.· Medical Teacher· 0 citations
It is suggested that AI systems are commonly performance-validated but insufficiently outcome-verified, which contributes to limited adoption and must be bridged through workflow-native design, prospective clinical validation, interoperable digital infrastructure, and alignment of technological innovation with healthcare system readiness.
Imran Khan, A. Mehreen, Noor Fatima et al.· 0 citations
Artificial intelligence has been proposed as a corrective to three persistent problems in mental healthcare: diagnostic imprecision, the trial-and-error character of treatment selection, and the episodic nature of clinical monitoring. The volume of primary research has expanded rapidly, yet few tools have altered routine practice. This critical narrative review examines evidence across the three domains in which artificial intelligence has been most extensively applied to mental health, namely diagnostic classification and risk detection, treatment personalisation, and continuous patient monitoring, and asks why demonstrated technical performance has so rarely converted into demonstrated clinical benefit. Literature was identified through a bibliographic metadata registry, a biomedical citation index, targeted searching of scholarly and institutional sources, and backward and forward citation tracking, covering January 2015 to 11 June 2026, with earlier work retained where conceptually necessary. Evidence was appraised for design adequacy, validation strategy, sample representativeness, outcome definition and reporting transparency, then synthesised thematically rather than study by study. Three findings recur. Apparent accuracy is systematically inflated by internal validation, small and selected samples, and reference standards of limited reliability; where external validation has been attempted, discrimination frequently falls towards chance. The three domains differ markedly in evidential maturity, since monitoring and conversational intervention now rest on randomised evidence and pooled effect estimates, whereas diagnostic classification and treatment-response prediction remain largely at the model-development stage. The binding constraints on translation are infrastructural and epistemic rather than algorithmic, encompassing narrow training populations, unreliable outcome labels, absent prospective evaluation and immature governance. Unresolved questions include whether any model confers benefit over routine care in prospective use, how algorithmic outputs should enter clinical judgement, and how safety should be established for generative systems operating outside professional supervision. Progress will depend less on model refinement than on representative longitudinal datasets, standardised outcome definitions, prospective impact evaluation and governance capable of distinguishing wellness products from clinical instruments.
Oyebode Mary Oluwabunmi, Anyebe Daniel Ameh, Jacob Miracle Godswill et al.· Journal of medicine and heal...· 0 citations