Artificial intelligence in medical education: a narrative review across four functional domains.
Abstract
Background
Artificial intelligence (AI) is increasingly reshaping medical and health-professions education through adaptive tutoring, generative content creation, simulation analytics, automated assessment, and diagnostic-reasoning support. Since 2023, large language models and multimodal AI systems have expanded AI from relatively narrow analytic tools into interactive educational agents capable of dialogue, feedback, and content generation.
Objective
This narrative review asks: How does AI contribute to medical education across four core educational functions, and what methodological, ethical, and implementation constraints should guide responsible integration? To answer this question, we synthesize evidence across four functional domains: AI-supported instruction and content generation, AI-augmented simulation and procedural training, AI-supported assessment and learning analytics, and AI-assisted diagnostic reasoning and clinical cognition.
Methods
A structured search of PubMed, Scopus, and Web of Science (January 2018-January 2025), supplemented by citation tracking, identified 540 records, of which 90 met the eligibility criteria. Evidence was analyzed using an iterative qualitative synthesis approach that combined inductive coding of educational mechanisms with deductive organization into the four predefined functional domains. Findings were summarized descriptively because of substantial heterogeneity in interventions, outcomes, and study designs.
Results
In instruction and content generation, AI-supported tools were associated with preliminary gains in knowledge acquisition, learner engagement, and self-regulated study, although accuracy depended on supervision and prompt quality. In simulation and procedural training, computer-vision, motion-analytics, and VR/AR systems supported immediate feedback and were associated with improved procedural efficiency in early studies. In assessment and learning analytics, AI tools reduced feedback latency and improved scoring consistency, but validity, fairness, and explainability remained central concerns. In diagnostic reasoning, AI-supported case platforms and LLM-based dialogue improved short-term reasoning performance and metacognitive calibration, while evidence for long-term clinical transfer remained limited.
Conclusion
AI appears most defensible as an augmentative educational partner that strengthens feedback, personalization, and competency-based progression, rather than as an autonomous substitute for educators. The current evidence supports cautious, supervised implementation, but remains early-phase and methodologically uneven. Multicenter validation, transparent reporting, equity-focused implementation, and structured AI-literacy training are essential before AI can be integrated more broadly into high-stakes educational workflows.