Artificial intelligence in clinical decision-making: a comparison of ChatGPT 5.0 and Gemini 3.0 in otologic cases
Large Lingual Models (LLMs) can suggest treatment options based on a patient’s history and symptoms. However, they may lack the ability to deliver fully patient-specific recommendations, which may lead to misdiagnosis or inappropriate treatment decisions. This study compared the diagnostic accuracy, investigation recommendations, and treatment planning consistency of two LLMs, ChatGPT 5.0 and Gemini 3.0, using real-world otologic cases. Fifty retrospective cases with diverse clinical characteristics (symptoms, examination findings, audiometric tests, and radiological sections) were selected. Three experienced otorhinolaryngologists established the final diagnoses as a gold standard. The cases were processed by both LLMs, which were instructed to act as ENT specialists. Two blinded reviewers independently evaluated the responses using the Artificial Intelligence Performance Instrument (AIPI) across four domains: summarization, differential diagnosis, additional examination, and therapy options. Inter-rater reliability was assessed via Cohen’s kappa. Both models demonstrated high performance in managing otologic cases. However, Gemini 3.0 significantly outperformed ChatGPT 5.0 in diagnostic accuracy, investigation/therapy planning, and overall AIPI scores ( p < 0.05). Subgroup analysis revealed that Gemini 3.0 performed higher AIPI scores in ‘vertigo’ cases, while similar performances were found in ‘otits media’, ‘hearing loss’ and ‘facial paralysis’ groups. Inter-rater agreement for AIPI scores was excellent. To our knowledge, this is the first study evaluating LLMs using real-world otologic data. While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases. Future research should include larger populations and endoscopic imaging to further validate these findings.