Aug 2026· Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie· 0 citations· 21 references
Medicine
TL;DR
Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.
OBJECTIVE
To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs) - ChatGPT, Microsoft Copilot, and Google Gemini - when applied to real-world clinical scenarios.
METHODS
Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into "primary" and "complex" glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0-100 scale.
RESULTS
Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p = .367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p = .009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p = .014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity.
CONCLUSIONS
LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.
Müge Toprak, B. Yılmaz Tuğan, N. Yüksel· Seminars in Ophthalmology· 0 citations
OBJECTIVE
Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator.
MATERIALS AND METHODS
We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications.
RESULTS
The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training.
DISCUSSION
VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends.
CONCLUSION
The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.
I. Strechen, P. Krishnan, O. Kilickaya et al.· International Journal of Med...· 0 citations
Glaucoma is a leading cause of irreversible blindness worldwide. Ophthalmologists diagnose glaucoma through a structured reasoning process by sequentially evaluating optic nerve head characteristics before reaching a final diagnosis, whereas existing AI systems typically perform direct image classification without providing clinically meaningful reasoning. We present the first clinically annotated fundus reasoning dataset, comprising 1,077 fundus photographs paired with expert-authored six-step diagnostic reports. Building on this dataset, we develop a reasoning-driven vision-language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis. The generated reports are clinically validated, achieving the best performance across all evaluated clinical findings, including a cup-to-disc ratio mean absolute error of 0.070, an ISNT Kendall distance of 1.73, and the highest semantic agreement with expert reports (BERTScore-F1 = 0.874). The resulting framework also improves glaucoma diagnosis, achieving a balanced accuracy of $94.7%$ and precision of $94.8%$, demonstrating that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance. Code and data are available at url{https://glaucoma-cot.github.io/}.
Kaichen Zhou, Yuzhen Chen, E. Yildiz et al.· medRxiv· 0 citations
Artificial intelligence in ophthalmology encounters a continual challenge: Systems proficient in picture classification seldom yield quantifiable enhancements in patient outcomes. The primary concern is the disparity between pixel-level performance metrics and their clinical significance. Primary obstacles encompass data bias, domain shift, and label noise, exacerbated by the lack of prospective, randomized deployment trials. The frequent disregard for patient-centered objectives, cost-effectiveness, and equity evaluations is significant. Rectifying these deficiencies necessitates stringent external validation, established decision criteria, and ongoing surveillance within actual clinical practices. Transparent reporting criteria and the deliberate incorporation of human-factors engineering are essential. Only by bridging this gap can algorithmic accuracy be converted into significant diagnostic precision for glaucoma, diabetic retinopathy, and macular conditions (specifically diabetic macular edema and age-related macular degeneration). This paper aims to assess the limits of using high-performing artificial intelligence systems in ocular image processing, which seldom lead to enhanced patient outcomes, and to delineate the scientific, clinical, and practical techniques required to close this gap.
Marco Zeppieri, Matteo Capobianco, Federico Visalli et al.· World Journal of Methodology· 0 citations