Machine learning can support diagnosticians in this effort, as demonstrated here utilizing multiple rating scales, the TASI, and the TAP, but there is a risk for bias when using machine learning and as such, no algorithm should replace expert clinical judgment.
Abstract
Abstract Background Best practices in diagnosing autism spectrum disorder require an expert diagnostician to integrate multiple sources of information, including direct observations, interviews, and rating scales completed by knowledgeable informants. Machine learning methods are well‐positioned to integrate disparate information sources to improve classification. This study sought to evaluate whether a machine learning algorithm could support appropriately weighting multiple sources of diagnostically relevant information. Methods Telehealth diagnostic assessments were conducted for 639 toddlers (age mean = 30.4, SD = 4.3 months; 29.7% female) already receiving early intervention (EI). The Toddler Autism Symptom Interview (TASI) and the TELE‐ASD‐PEDS (TAP) were completed during separate visits. Each child's caregivers and usual EI providers asynchronously completed rating scales of autistic characteristics developed for this study. Calibration and validation data subsets were created using a 2:1 split. The optimal algorithm was developed using elastic net regularized regression in the calibration dataset, which was then evaluated in the validation dataset. Statistical fairness criteria evaluated whether the proposed algorithm functioned similarly across multiple protected group statuses. Results The prevalence of autism in the sample was 80.4%. The algorithm developed in the calibration dataset included both the caregiver‐ and EI provider‐completed rating scales, four items from the TASI, and six items from the TAP. Model performance was high in both the calibration and validation samples (sensitivity >0.90; specificity >0.75; kappa >0.60), but statistical fairness criteria varied. False positives were relatively rare, but were more common in advantaged groups, which may suggest systematic under‐classification by the algorithm in minoritized groups. Conclusion Diagnostic practices for autism require integrating multiple sources of information. Machine learning can support diagnosticians in this effort, as demonstrated here utilizing multiple rating scales, the TASI, and the TAP. However, there is a risk for bias when using machine learning and as such, no algorithm should replace expert clinical judgment.
A toddler‑centered Case‑Based Reasoning (CBR) framework that emulates clinicians’ decision‑making by retrieving and adapting similar historical cases and delivers transparent, interpretable recommendations via comparable cases, supporting clinician trust is introduced.
Hachemi Yamina· ITEGAM- Journal of Engineeri...· 0 citations
It is demonstrated that lay behavioral descriptions can provide diagnostically valuable information comparable to clinical observations, although they are not readily summarized by AI.
Gondy Leroy, Himanshu Nimbarte, Madhuri Sai Kandula et al.· Frontiers in Digital Health· 0 citations
Background: Digital behavioral phenotyping of autism spectrum disorder (ASD) offers a promising approach for developing more scalable diagnostic frameworks across diverse global contexts. Machine learning (ML) models show promise for ASD diagnosis using behavioral videos, but critical questions remain regarding whether models trained on data from one country work in another, and how the background of the raters affects the accuracy. Our work addresses these questions by testing whether ML models can accurately diagnose ASD across different populations and rater groups. Methods: This work evaluates the performance of a supervised ML framework for binary classification of ASD versus non-ASD [speech, language and communication disorders (SLC) + neurotypical (NT)] in a cohort of 227 children in Bangladesh. We first assessed the cross-domain model transferability of a clinical-instrument-trained logistic regression model (LR-9) on behavioral ratings that were based on videos of Bangladeshi children interacting with caregivers and toys at two major child development centers in Dhaka, Bangladesh. We then trained five diverse classifiers (Logistic Regression, Random Forest, XGBoost, SVM, and RuleFit) on the full annotated Bangladeshi dataset. Using SHAP-based consensus elbow feature selection, we identified a compact set of features that maintained the performance. Finally, we developed ensemble models to improve predictive stability. Results: The LR-9 model, originally trained on U.S. clinical instrument data, was evaluated on video-based behavioral ratings from 214 Bangladeshi children. When tested on Bangladeshi clinician ratings, the LR-9 model achieved a sensitivity of 86.1% (95% CI: [0.78–0.93]) and AUC of 0.79 (95% CI: [0.73–0.86]). The distinction across rater groups was between trained raters (clinicians and students) and crowd workers, who showed lower sensitivity 28.5% (95% CI: [0.21, 0.39]). When tested on the aggregated ratings from all groups, the model achieved an AUC of 0.78 (95% CI: [0.72–0.84]). Inter-rater reliability followed the same pattern: individual agreement was fair (Krippendorff’s α = 0.26), but the multi-rater consensus was reliable (ICC(1,k) = 0.84), with Bangladeshi clinicians showing the highest agreement (α = 0.34) and crowd workers the lowest (α = 0.20). We then trained new models directly on the Bangladeshi ratings. All model types achieved similar AUC values (0.86–0.89), with overlapping confidence intervals. Using just 8–11 key behaviors kept the similar performance while cutting the features by 66–75%. Combining ensembles gave similar results (e.g., Bayesian averaging: AUC 0.88 [0.78, 0.95]) but with more stable predictions. Conclusion: This study provides evidence that mobile video-based ASD diagnosis can achieve comparable performance (AUC: 0.89 [0.76, 0.96]) to models trained on clinical instrument data. This work contributes to the development of broader adaptable autism detection tools, bypassing the dependence on traditional clinical instrument data.
Saimourya Surabhi, K. Dunlap, Parnian Azizian et al.· BioMedInformatics· 0 citations
This Attention Deficit Hyperactivity Disorder (ADHD) remains substantially under-diagnosed among university students despite affecting 2–8% of this population. Campus health services, facing persistent resource constraints, frequently accumulate assessment backlogs of 6–12 months. This paper presents a machine learning framework for automated ADHD pre-screening that combines structured psychometric assessments with natural language processing (NLP)-derived features extracted from free-text clinical self-reports. Drawing on 506 university student responses, we engineer 124 multimodal features spanning four validated instruments, the Adult ADHD Self-Report Scale (ASRS), Beck Anxiety Inventory (BAI), Beck Depression Inventory (BDI-II), and Adult Attachment Scale (AAS), together with unstructured diagnostic text. Mutual Information-based feature selection reduces dimensionality to 20 features, yielding a 2% accuracy gain. A comparative evaluation across five classifiers reveals Logistic Regression as the top performer, achieving 81.4% accuracy and an AUC of 0.881. SHAP (SHapley Additive exPlanations) analysis confirms clinical meaningfulness by identifying BAI Item 8 (somatic anxiety), ASRS inattention items, and prior mental health history as the principal risk factors. The system is deployed as an interactive web application that delivers calibrated risk assessments suited to clinical triage in resource limited settings.
Abstract Objectives Building on innovations for autism detection—where artificial intelligence (AI)-based models monitor clinical data within electronic health records—this study evaluates the context for clinical decision support (CDS) deployment and identifies design preferences. Materials and Methods This observational study utilized contextual inquiry to elicit perspectives from 8 clinicians and twenty caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics. Data were analyzed using rapid qualitative analysis techniques. Results Workflow analysis identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points. Technologies that streamlined screening included patient portals, digital tablets, and note templates. Clinicians identified 2 major barriers—limited screening tool accuracy and challenges in implementing follow-up steps—and 3 facilitators: electronic screening, early intervention provider input, and staff referral coordination support. For design, CDS should include clear, actionable outputs, with explanations of prediction data, visual summaries linked to next steps, and educational resources. Embedding CDS within the EHR, with outputs delivered at key points during the clinical encounter, along with caregiver-facing materials, would improve workflow efficiency. Discussion Findings highlight key integration points for an autism detection AI-based CDS tool and stress the need for clinical utility and caregiver-centered communication. Effective design requires alignment with clinical workflow, including the timing of outputs, meaningful explanations, and integration with caregiver communication. Conclusion Findings will inform the design of an AI-based CDS tool for autism detection, providing workflow-informed integration points and user preferences. Future work should refine explainability and optimize delivery of outputs within clinical encounters to support decision-making and caregiver engagement.
Adesuwa Emovon, Lauren P. Driggers-Jones, Matthew M Engelhard et al.· JAMIA Open· 0 citations