Aug 2026· Ophthalmology and Therapy· 0 citations· 28 references
Medicine
TL;DR
TabPFN achieved the lowest mean absolute error (MAE) on all Eyerobo scenarios and was competitive with Ridge on auto-refractor scenarios; on the headline Eyerobo Pre + IOL scenario, the per-eye paired difference in MAE between TabPFN and the next-best gradient-boosting learner was 0.117 D in favour of TabPFN.
Abstract
INTRODUCTION
A cross-sectional comparative device study to compare the Eyerobo Vision Screener (VS), a portable handheld photorefractor, against a conventional auto-refractor for machine-learning-based prediction of cycloplegic spherical equivalent (SE) and power vector components (J0, J45) in a pediatric myopia cohort, to evaluate whether supplementary IOL Master biometry narrows the performance gap between devices, and to report screening-oriented classification metrics (sensitivity, specificity, positive and negative predictive values, area under the receiver operating characteristic curve [AUC]) at clinically meaningful referral thresholds that quantify the device-model combination's triage performance.
Methods
Data from 1129 eyes of 574 patients aged 6-18 years were collected using three ophthalmic devices before and after pharmacological dilation. All refractive measurements were decomposed into power vectors (SE, J0, J45) following Thibos et al. Twelve clinically motivated feature scenarios were constructed from a symmetric framework comparing the Eyerobo VS and auto-refractor (each under predilation, post-dilation, and combined conditions), with and without IOL Master biometry. TabPFN, a tabular foundation model requiring no hyperparameter tuning, was evaluated alongside seven conventional machine-learning algorithms using patient-level five-fold grouped cross-validation repeated over three random seeds. Agreement was assessed using Bland-Altman analysis, Pearson correlation, and clinical threshold analysis. In addition to regression metrics, a screening-classification analysis was performed at three prespecified clinical thresholds (cycloplegic SE ≤ - 0.50 D, ≤ - 3.00 D, and ≤ - 6.00 D), reporting sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and AUC with 95% paired bootstrap confidence intervals. Formal paired statistical comparison between TabPFN and the next-best algorithm and a one-eye-per-patient sensitivity analysis are also reported.
Results
TabPFN achieved the lowest mean absolute error (MAE) on all Eyerobo scenarios and was competitive with Ridge on auto-refractor scenarios; on the headline Eyerobo Pre + IOL scenario, the per-eye paired difference in MAE between TabPFN (0.425 D) and the next-best gradient-boosting learner (XGBoost, 0.542 D) was 0.117 D in favour of TabPFN, with paired Wilcoxon signed-rank p < 0.001 . For predilation SE prediction, the auto-refractor achieved an MAE of 0.295 D (87.4% within ± 0.50 D), compared with 0.657 D (61.6%) for the Eyerobo VS. Adding IOL Master biometry to the Eyerobo reduced this gap by 64%, yielding an MAE of 0.425 D (72.0%). For astigmatic power vector components, the gap narrowed further: J0 MAE was 0.106 D (auto-refractor) versus 0.143 D (Eyerobo + IOL), and J45 MAE was 0.060 D versus 0.088 D. Auto-refractor post-dilation predictions approached cycloplegic accuracy (MAE = 0.203 D, 96.0% within ± 0.50 D). Subgroup analysis showed the Eyerobo performed best for mild myopia (MAE = 0.398 D, 75.6% within ± 0.50 D) and degraded for high myopia and the small nonmyopic stratum. In the screening-classification analysis, the Eyerobo and the auto-refractor were operationally equivalent at the any-myopia threshold (Eyerobo AUC 0.964 [95% CI 0.939-0.982], sensitivity 0.946, specificity 0.894; auto-refractor AUC 0.969 [0.948-0.986], sensitivity 0.960, specificity 0.886; paired AUC difference 0.005 [95% CI - 0.022 to + 0.034 ], formally noninferior under a prespecified δ = 0.05 margin).
Conclusions
Within this single-center pediatric referral cohort, the Eyerobo VS combined with machine learning produced cycloplegic SE estimates of a precision consistent with use as a triage adjunct to identify children who should proceed to a full cycloplegic examination, with near-equivalent performance to the auto-refractor for astigmatic components. The addition of IOL Master biometry narrowed the device gap by 64%. TabPFN achieved the best performance among the evaluated algorithms within this cohort and requires no hyperparameter tuning, but the modest margin over linear and gradient-boosting baselines and the absence of external validation preclude any recommendation as the definitive algorithm of choice. The proposed approach is not a substitute for cycloplegic refraction; prospective external validation in a general pediatric screening cohort is required before clinical deployment. Video abstract. See video abstract at https://youtu.be/wy7_Fw5Dnw4 or in the online or HTML version of the manuscript. (MP4 4546 KB).
Purpose: To determine the accuracy of the fundus red reflex test using a smartphone and autorefractor to detect refractive errors in children.Methods: This cross-sectional study included 1302 eyes of 651 school-age children aged 6- 13 years with visual acuity worse than 6/9 in at least one eye, examined using the Snellen chart. Study participants were selected using purposive sampling, after which non-cycloplegic refractive measurement was performed using two instruments: a smartphone-based application (MD EyeCare) and an autorefractor. The detection of refractive errors using the online platform was compared with autorefraction as the reference standard.Results: Our study found that 46.10% of participants with abnormal red reflex had significant myopia, 41.20% had nonsignificant refractive errors, and 12.70% had significant hypermetropia. The mean difference between the smartphone-based application and the autorefractor was ≤−3.00 D (range: −2.88 to −2.63), with a sensitivity of 0.89 to 0.90 and a specificity of 0.86 to 0.87 for significant myopia, and the area under the curve was between 0.94 and 0.97. The mean difference between the smartphone-based application and the autorefractor was ≥+1.00 D (range: +0.13 to +0.38), with a sensitivity of 0.76 and a specificity of 0.97 for significant hypermetropia, and the area under the curve was between 0.88 and 0.96.Conclusion: A smartphone-based application could be used as a screening tool for refractive errors in children, especially in resource-limited settings.
Muhammad Syauqie, Lutfah Rif’ati, W. Diarsvitri et al.· Journal of Ophthalmic & Visi...· 0 citations
Smartphone-based fundus imaging (SBFI) is an emerging approach with potential relevance for global ophthalmic care, including in low- and middle-income countries and other resource-constrained settings. This scoping review, based on a structured literature search, synthesises current SBFI technology, clinical and teaching applications, implementation challenges and future directions. Analysing 30 hardware solutions (based on three principal technical designs) and 274 scientific publications and other relevant sources, we found that some SBFI devices can provide fields of view and image quality sufficient to support screening or triage for referable diabetic retinopathy, glaucoma-related optic nerve head changes, and selected retinopathy of prematurity applications. Reported sensitivity, specificity and gradability varied by disease, device, protocol, operator, and reference standard; no universal performance threshold could be inferred. After brief training, allied healthcare workers can acquire usable images in selected settings, suggesting that SBFI may support task-shifted screening and referral pathways. Artificial intelligence may further support scalability by assisting image-quality and protocol compliance assessment, frame selection, montage generation and disease classification. However, costs, maintenance requirements, data governance and domain shift remain important implementation considerations. Key challenges include variable image quality and fields of view, heterogeneous reporting, smartphone compatibility, regulatory compliance, photobiological safety verification, data security, and patient privacy. We therefore suggest minimum reporting standards to increase comparability across studies. Large-scale implementation and cost-effectiveness studies are needed to determine whether improved access to ophthalmic imaging can translate into sustainable and equitable eye-care services.
R. Lechtenboehmer, C. Cleland, F. Malerbi et al.· Progress in retinal and eye...· 0 citations
PURPOSE
To determine whether objective biometry can guide intraocular lens formula selection in post-laser vision correction (LVC) eyes with indeterminate classification of actual LVC type (total keratometry [TK] TKave - keratometry [K] Kave = 0.00 to 0.06), and to compare refractive accuracy among contemporary LVC formula strategies.
METHODS
A total of 507 post-LVC eyes (inclusive of both myopic [M-LVC] and hyperopic [H-LVC] eyes) with TKave - Kave between 0.00 and 0.06 measured by swept-source optical coherence tomography biometry (IOLMaster 700; Carl Zeiss Meditec AG) were analyzed. Refractive prediction errors were calculated for 22 formulas, including eight H-LVC, eight M-LVC, and six non-LVC formulas. Formula performance was assessed using root mean square error rankings and heteroscedasticity testing for dependent data. The main outcome measure was refractive prediction error and proportion of eyes within ±0.50 diopters (D) of target refraction.
RESULTS
H-LVC formulas demonstrated superior refractive accuracy across all eyes, regardless of reported LVC type. Pearl DGS and Barrett True K No-History (H-LVC) achieved the highest proportion of eyes (approximately two-thirds) within ±0.50 D of target. K-based H-LVC formulas outperformed variants incorporating posterior K or TK. Standard formulas showed intermediate performance, whereas M-LVC formulas ranked lowest. Differences between top-performing formulas were statistically significant.
CONCLUSIONS
In post-LVC eyes with TKave - Kave between 0.00 and 0.06, H-LVC formulas provide the most accurate refractive outcomes, independent of reported treatment type. These findings challenge reliance on historical LVC classification for formula selection and support a measurement-driven approach in which biometry-guided formula selection improves outcomes when history is unreliable in ambiguous post-LVC eyes.
David L. Cooke, J. Wendelstein, Kamran M. Riaz· Journal of refractive surger...· 0 citations
To develop a predictive model based on school-based vision screening programs for evaluating hyperopic reserve status in children aged 5–10 years.
This cross-sectional study included 8,035 students aged 5–10 years from senior kindergarten and primary school in Guangdong Province, China. Participants underwent ocular biometric measurements, autorefraction, and questionnaire assessment before cycloplegia, after which autorefraction was repeated. An ensemble learning model was constructed to predict low hyperopic reserve (LHR), a refractive status below age-appropriate norms but not yet myopic, using NCR ocular parameters and questionnaire data. Model performance was evaluated and compared through nested 5-fold cross-validation. Sensitivity and robustness analyses were conducted to verify the comprehensive model performance. SHapley Additive exPlanations (SHAP) was utilized to interpret feature contributions and model logic.
On the test set, the ensemble learning model achieved an area under the curve (AUC) of 0.807, an accuracy of 0.737, and a precision of 0.747, outperforming all the single machine learning models. Compared to the direct NCR-based assessment and the +0.50 D / +0.75 D fixed-offset benchmarks (accuracy = 0.614–0.669, precision = 0.593–0.683), the ensemble learning model improved accuracy by 0.068–0.123 and precision by 0.064–0.154. SHAP analysis identified axial length to corneal radius ratio (SHAP value = 0.652), spherical equivalent (SHAP value = 0.514), and school type (SHAP value = 0.312) as the most three important predictors.
The proposed ensemble learning model offers a noncycloplegic, convenient approach for school-based vision programs to predict LHR in children aged 5–10 years. This model and its risk-assessment tool may support existing school-based vision programs in preserving children's hyperopic reserve.
Jingwei Jiang, Yu Lu, Meng Li et al.· Frontiers in Public Health· 0 citations
Objectives: The purpose of this study was to evaluate the ability of computer vision models to detect myopia from standard colour fundus photographs and to compare their diagnostic performance with that of experienced ophthalmologists. Methods: A previously published dataset of 324 retinal fundus images labelled as myopic or non-myopic based on cycloplegic refraction as used for model training and internal validation. Images were acquired using a non-mydriatic 45° fundus camera. Final model evaluation was performed on an independent test set of 50 images from different patients who were not included in the original dataset. YOLOv8 and YOLOv11 variants were trained for binary classification. Internal validation used patient-level cluster bootstrap confidence intervals, whereas image-level bootstrap confidence intervals were estimated for the independent test set. Pairwise model comparisons were adjusted using the Holm–Bonferroni correction. Five experienced ophthalmologists independently classified the test set, and their consensus was compared with the selected YOLO models using DeLong’s and exact McNemar tests. Results: Internal validation identified YOLOv8-m and YOLOv11-n as the best-performing models according to a predefined composite score used exclusively for model selection. On the independent test set, YOLOv11-n achieved the highest area under the curve (AUC = 0.889), followed by YOLOv8-m (0.806), although the difference was not statistically significant (DeLong test, p > 0.05). The clinical consensus achieved an AUC of 0.832, with no significant difference compared with either model. Exact McNemar testing likewise revealed no statistically significant differences in paired classification outcomes between either AI model and the clinical consensus. Limitations include the small, single-centre, class- and age-imbalanced dataset and the limited number of expert observers. Conclusions: Although neither YOLOv8-m nor YOLOv11-n showed statistically significant differences from the clinical consensus on this independent test set, these findings should be interpreted cautiously given the relatively small, single-centre study population. Larger multicentre studies with independent external validation are warranted to confirm the generalisability, robustness, and potential role of clinician-driven computer vision models as decision support tools for myopia screening.
Nicola Rizzieri, Luca Dall'Asta, Maris Ozoliņš· Journal of Clinical Medicine· 0 citations
Structured preoperative variables can partially reproduce single-center clinician-selected refractive procedure patterns but do not establish optimal surgical recommendation or external generalizability.
Yinhao Li, Gang Li, Chuanyun Xu et al.· BMC Ophthalmology· 0 citations