Skip to content
Review Open access

Validation of an AI-powered mobile application for personalizing medical note explanations: a mixed-methods evaluation

Aug 2026 · Frontiers in Digital Health · 0 citations · 31 references

TL;DR

This mixed-methods evaluation suggests that a deliberately constrained, language-focused AI system can improve the accessibility of medical notes while preserving clinical accuracy and safety while extending into clinical interpretation.

Abstract

Nearly half of adults struggle to understand written health information, making medical communication a persistent barrier to effective care. While artificial intelligence has potential to improve health communication, few patient-facing tools have undergone systematic validation for personalized medical explanation. Patiently AI is a mobile application designed to clarify clinician-authored medical notes using large language models with audience-specific adaptations (child, teenager, adult, carer) and tone variations (friendly, informative, reassuring). A three-phase mixed-methods evaluation was conducted: (1) computational readability analysis of 210 AI-generated explanations using established metrics; (2) expert review by 15 healthcare professionals assessing medical accuracy, safety, and communication quality; and (3) a patient survey of 54 participants evaluating preferences, comprehension, and acceptance. AI-generated explanations demonstrated consistent improvements in readability, with mean Flesch–Kincaid Grade Level decreasing by 2.96 levels (10.57–7.61), Flesch Reading Ease increasing by 31.9 points (37.7–69.6), and Gunning Fog Index decreasing by 4.09 points (14.5–10.4); all improvements were statistically significant (all P  ≤ 0.002). Readability gains were greatest for younger audiences (child: 4.25 grade-level reduction; adult: 1.80). Expert reviewers rated outputs highly for medical accuracy (4.49 ± 0.83/5), clarity (4.53 ± 0.77/5), and trustworthiness (4.37 ± 0.90/5), with 87.3% assessed as clinically safe. Inter-rater agreement across the 15 reviewers was substantial (Gwet's AC1 = 0.72 for safety assessments). Among patients, 70.0% of responses preferred AI-generated explanations ( P  < 0.001), with 98.1% comprehension accuracy and high ratings for clarity (4.58 ± 0.65/5) and confidence in care (4.19 ± 0.85/5). Overall, 70.4% indicated a likelihood of using the application. This mixed-methods evaluation suggests that a deliberately constrained, language-focused AI system can improve the accessibility of medical notes while preserving clinical accuracy and safety. Patiently AI demonstrates a scalable approach to supporting health literacy and patient engagement without extending into clinical interpretation.

Read PDF

Similar papers

Open access Jul 2026

Quality of AI-generated exercise awareness messages for older adults aligned with the ICFSR consensus: a comparative study of three LLMs.

BACKGROUND Large Language Models (LLMs) are emerging as potential tools for health communication and patient education. However, their ability to translate complex medical guidelines into accessible, safe, and accurate messages for older adults remains insufficiently evaluated. To assess the capacity of ChatGPT-4, Claude AI, and Deepseek to generate exercise awareness messages for older adults, aligned with international expert consensus. MATERIALS AND METHODS This was a cross-sectional observational study conducted between April and August 2025. Using the ICFSR 2021 Global Consensus on optimal exercise recommendations as the gold standard, we evaluated messages generated by three LLMs for 13 common chronic conditions in older adults. A standardized prompt was used to generate messages addressing exercise prescription considerations, disease progression, and recommended modalities. Two independent expert evaluators; a geriatrician and a sports medicine physician assessed five dimensions using a 5-point Likert scale; accuracy, clarity, safety, behavioral relevance, and absence of fabrication. Readability was measured using Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG). Inter-rater agreement was assessed using intraclass correlation coefficient (ICC). RESULTS Claude achieved the overall scores from both evaluators at 4.63 ± 0.55 and 4.60 ± 0.55, followed by ChatGPT-4 at 4.55 ± 0.50 and 4.48 ± 0.53 and Deepseek at 4.15 ± 0.87 and 4.17 ± 0.86. Inter-rater agreement was moderate was 0.668. Claude demonstrated accuracy scores at 4.58 ± 0.64, while ChatGPT-4 excelled in clarity with 5 ± 0. All models achieved perfect scores for absence of fabrication (5 ± 0). Readability indices revealed high complexity across all LLMs, with median FKGL values ranging from 10.84 to 10.93, corresponding to 10th-11th grade reading level, exceeding recommended levels for older populations. A significant correlation was found between Claude's accuracy scores and FKGL (r = 0.599, p = 0.030). CONCLUSION LLMs, particularly Claude, ChatGPT-4 and Deepseek demonstrate strong potential for generating accurate, safe, and hallucination-free exercise awareness messages for older adults. However, readability remains above recommended levels, requiring optimization.

Y. Bayala, L. Ouedraogo, A. Cissé et al. · 0 citations
Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability, usefulness, overall quality, and readability of AI-generated patient education materials. Evidence specifically evaluating AI-generated patient education for hallux valgus, a condition strongly influenced by patient expectations and treatment preferences, remains limited. This cross-sectional comparative study evaluated the performance of three large language models—ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3—in responding to 20 patient-centered questions related to hallux valgus. Questions were developed using AI-assisted question generation and publicly available Google Trends search patterns and categorized into four clinical domains. AI-generated responses were anonymized and independently assessed by three orthopaedic surgeons for reliability, usefulness, and overall quality using 7-point Likert-based reliability and usefulness scales and the Global Quality Scale (GQS). Readability was analyzed using six standardized indices. Inter-rater agreement and between-model comparisons were statistically evaluated. Gemini-2.5-Flash demonstrated modestly higher overall reliability, particularly in questions related to etiology and clinical presentation. DeepSeek-V3 achieved higher usefulness scores in the long-term outcomes and quality-of-life domain and produced significantly more readable content, as reflected by higher Flesch Reading Ease scores and lower grade-level indices. In contrast, Gemini-2.5-Flash generated linguistically more complex responses requiring higher educational levels for comprehension. Overall usefulness and global quality scores did not differ significantly among models. Qualitative review also identified occasional examples of oversimplified or potentially misleading information. Despite these differences, several readability metrics exceeded recommended patient health-literacy thresholds. Contemporary AI-based conversational agents can provide patient-oriented information regarding hallux valgus with variable reliability and readability characteristics, although statistically significant differences were observed across models. A trade-off between factual accuracy and linguistic accessibility was observed. AI tools should therefore be regarded as adjuncts to, rather than replacements for, clinician-led patient education. Awareness of AI limitations and appropriate clinical guidance remain essential to ensure safe, accurate, and patient-centered information delivery. Not applicable.

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Review Open access Aug 2026

When AI Sounds More Helpful: Users’ Perceptions of AI-Generated and Physician-Provided Health Information

AI-powered conversational agents are becoming part of the everyday Internet information ecosystem, reshaping how users seek, interpret, and act on health-related information outside clinical encounters. As large language model (LLM)-based chatbots are increasingly used as on-demand digital health information tools, understanding how users perceive their credibility, usefulness, and limitations is essential for the responsible design of future Internet-based health services. This mixed-method survey study examined how general adults evaluated healthcare-related question–answer pairs provided by physicians and generated by AI chatbots. A sample of U.S.-based adults recruited through Prolific (N = 62) rated each answer on clarity, usefulness, appropriateness of detail, trustworthiness, and perceived evidence, and provided open-ended explanations of their judgments. Primary mixed-effects analyses showed that both ChatGPT- and Claude-generated responses received higher overall participant ratings than physician-provided responses, although the estimated difference was substantially larger for Claude (ChatGPT–physician estimate = 0.250, 95% CI [0.135, 0.364]; Claude–physician estimate = 0.825, 95% CI [0.710, 0.939]). ChatGPT received higher ratings on four of the five dimensions but not on clarity, whereas Claude received higher ratings across all five dimensions. However, physician, ChatGPT, and Claude responses were always presented first, second, and third, respectively. Response source was therefore confounded with presentation position, and the observed differences cannot be attributed exclusively to source. The responses were also not matched for length or format. Qualitative findings showed that participants valued detailed, specific, and evidence-like explanations. Participants also expressed concerns about hallucination, privacy, over-reliance, and the need for clinician verification. These findings suggest that LLM-based chatbots may be perceived as useful supplemental information tools within future Internet health ecosystems, but their deployment should include safeguards that support transparency, verification, and appropriate reliance.

Tianyi Wang, Masooda N. Bashir · 0 citations
Review Open access Jul 2026

“But it sounded confident”: the role of accuracy, tone, and disclaimers in users' medical decision-making

Introduction Artificial intelligence (AI)-powered chatbots are increasingly used in healthcare for applications ranging from symptom triage to lifestyle guidance. Their effectiveness depends not only on their ability to provide reliable information but also on users engaging with their advice while remaining aware of potential inaccuracies. This study investigated how users perceive AI-generated medical advice, with a particular focus on the roles of accuracy, conversational tone, and disclaimers. Methods A survey was conducted with 115 participants who evaluated 15 chatbot scenarios representing five common health problems. Participants assessed chatbot responses varying in accuracy, tone, and the presence of disclaimers. Associations between these factors, user trust, willingness to seek second opinions, and participant characteristics (including age, health literacy, and prior chatbot use) were examined. Results Accuracy emerged as the strongest predictor of user trust, although participants did not consistently identify inaccurate advice. Individual characteristics, including age, health literacy, and prior chatbot experience, showed no significant associations with trust. Differences in chatbot tone had little influence on user perceptions, with chatbots generally being viewed as confident regardless of phrasing. The effects of disclaimers on willingness to seek second opinions were mixed and varied across scenarios, suggesting that disclaimers alone may not effectively prevent over-reliance on chatbot advice. Open-ended responses highlighted trust in healthcare professionals and the importance of clearly communicating AI limitations. Discussion This exploratory study, based on self-reported intentions in hypothetical scenarios, suggests that the accuracy of AI-generated medical advice is the primary determinant of user trust. Conversational tone appears to have limited influence, while disclaimers may not consistently promote appropriate caution. These findings indicate that healthcare chatbots should prioritize accuracy, clearly communicate their limitations, and encourage critical evaluation to achieve an appropriate balance between user trust and scepticism.

Kerstin Denecke, E. Gabarron · 0 citations
Review Open access Aug 2026

The Scale for AI Literacy in Health Care Workers: Development and Validation

Abstract Background AI is increasingly embedded in health care systems; yet, validated instruments for assessing AI literacy among health care workers remain limited. Existing measures are often designed for students or general populations and may not adequately reflect competencies required in health care practice. Objective This study aimed to develop and validate the Scale for AI Literacy in Health Care Workers (SAIL-HCW), a new instrument designed to assess AI literacy across domains relevant to health care practice. Methods A 3-phase instrument development study was conducted. In Phase 1, conceptual domains were identified through a literature review, and an initial item pool was generated. In Phase 2, content validity was assessed by 4 subject-matter experts, and face validity was evaluated with 26 health care workers. Feedback from both groups informed item refinement. In Phase 3, psychometric testing was conducted using survey data from health care workers in a single health care organization. A total of 425 participants completed the survey. The dataset was randomly split into 2 subsamples for exploratory factor analysis (n=212) and confirmatory factor analysis (n=213). Model fit was evaluated using unidimensional, correlated-factor, higher-order, and bifactor models. Reliability was assessed using Cronbach alpha and McDonald omega. Item performance was examined using corrected item-total correlations (CITC), item discrimination analysis, and inter-item correlations. Construct validity was assessed using prior AI training, frequency of AI use, and self-rated AI literacy. Results Phase 2 feedback from experts and health care workers supported the proposed domain structure and informed item refinement, including revision of wording and removal of redundant items. The final SAIL-HCW consists of 14 items across 7 domains, including AI concept, data fluency, AI evaluation, AI in practice, ethics and regulation, AI in system, and continuous learning. In Phase 3, the bifactor model showed the best fit compared with alternative models (comparative fit index and Tucker-Lewis index>0.93; root-mean-square error of approximation<0.06; standardized root-mean-square residual<0.05), indicating a general AI literacy factor alongside domain-specific factors. Internal consistency for the total scale was high (Cronbach α=0.937; ω=0.938). Domain-level reliability ranged from 0.635 to 0.797. All items significantly discriminated between high- and low-scoring groups (P<.001), with CITC values ranging from 0.570 to 0.785. Construct validity was supported, with higher SAIL-HCW scores observed among participants with prior AI training, higher frequency of AI use, and higher self-rated AI literacy (all P<.001). Conclusions The SAIL-HCW provides initial evidence of validity and reliability for assessing AI literacy among health care workers. Findings suggest that AI literacy may be represented as a general construct with additional domain-level components. The scale may be useful for research and educational evaluations, although further validation in other settings is required.

Chin-Siang Ang, Sakura Ito, Saumya Bajaj et al. · 0 citations
Open access Jul 2026

Augmenting medical data interpretation with Large Language Models (LLMs): a comparative analysis of patient empowerment, information processing, and technology acceptance.

LLM-augmented interpretation of medical data compares with healthcare professional-led interpretation across different data modalities, excelling in enhancing comprehension, control, and efficiency while healthcare professionals provide superior relational value through trust, confidence, and emotional support.

Pouyan Esmaeilzadeh · 0 citations