Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.
Abstract
Language barriers hinder healthcare, particularly during case history-taking, a key part of diagnosis. While multilingual artificial intelligence (AI) chatbots offer solutions, there is fragmented evidence of their effectiveness and impact. This systematic review followed PRISMA 2020 guidelines, examining studies published between 2015 and 2025 on multilingual AI chatbots in healthcare across four databases (Google Scholar, Scopus, Web of Science, and PubMed), using a two-stage screening process. Data extraction focused on applications, supported languages, underlying technologies, target populations, and clinical outcomes. From 503 records, 49 studies, covering primary care, telemedicine, oncology, mental health, and other areas, met the criteria. Supported languages included English, Spanish, Arabic, Chinese, Hindi, and other underrepresented languages. In individual system evaluations using heterogeneous methodologies and evaluation settings, AI chatbots achieved a diagnostic accuracy ranging from 72–92%. Core technologies included large language models (LLMs), bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), retrieval-augmented generation (RAG), speech recognition, and distillation. The findings show that these improve clinical workflow (30–70% time savings) and patient engagement, reduce language barriers, and promote health equity. However, the overall evidence certainty was low to moderate, reflecting the predominance of prototype and proof-of-concept studies. Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias. Future research should include real-world studies, diverse populations, standardized outcome measures, and long-term equity assessments.
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
Background Scotland faces a severe public health crisis with drug-related deaths reaching 267 per million people, ranking second globally after the United States. Medication-Assisted Treatment (MAT) represents a proven intervention for heroin addiction. However, healthcare professionals struggle with accessing and interpreting current MAT standards through fragmented information systems and time-consuming manual searches across multiple websites. Despite advances in healthcare chatbots leveraging Large Language Models (LLMs), no specialized systems exist to support MAT delivery or integrate advanced technologies like Retrieval-Augmented Generation (RAG) and Knowledge Graphs for addiction treatment. Objective To develop and evaluate an AI-driven chatbot prototype that integrates LLMs, RAG, and Knowledge Graphs to enhance healthcare professionals’ access to MAT standards in Scotland, addressing current barriers in information delivery and clinical decision-making. Methods We employed a mixed-methods approach combining a survey of 39 MAT healthcare professionals (31% response rate) and systematic literature review following PRISMA guidelines. The chatbot prototype was developed using Llama2 language model, Neo4j knowledge graphs, and custom RAG implementation. Data was ethically collected from Public Health Scotland and Healthcare Improvement Scotland websites. Performance was evaluated using BLEU and ROUGE metrics, with prototype deployment via Streamlit interface. Results Survey findings revealed significant challenges with current communication methods: only 5 of 39 respondents rated existing systems as “exceptional,” while 17 rated them as “average” or below. Primary challenges included decentralized information (n=13) and time-consuming access processes (n=8). Literature review of 14 healthcare chatbot studies identified a critical gap in MAT-specific applications. The developed prototype demonstrated moderate performance with BLEU score of 36.64, ROUGE-1 score of 0.48, and ROUGE-L score of 0.42. The knowledge graph successfully integrated 227 nodes, 136 relationships, and 8 characteristics representing comprehensive MAT standards. The system successfully retrieved relevant MAT standards information in response to queries about specific MAT standards, medication protocols, and implementation guidance Conclusions To our knowledge, this study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment. While performance metrics indicate potential for enhancing MAT information delivery, further development is needed to improve semantic understanding and response naturalness. The prototype establishes a foundation for future integration with electronic health records and broader healthcare systems, with the potential to support improved treatment outcomes for individuals with heroin addiction in Scotland, subject to longitudinal clinical validation.
Sandra C. Nwobi, Zainab Loukil, Abbas Jawahar· Frontiers in Digital Health· 0 citations
Artificial intelligence, particularly AI-based chatbots, has gained increasing attention as a tool to support pharmacists in providing drug information services, especially in settings with pharmacist shortages. This study compared the competencies of ChatGPT, ChatGPT Pro, Perplexity, and Perplexity Pro in responding to drug-related clinical questions, and examined the effect of an Enhanced Task Translation (ETT) technique on chatbot performance, after English technical terms were embedded with Thai-language prompts. An experimental comparative design was employed. Two sets of Thai-language multiple-choice questions (MCQs), each comprising 120 items based on the Thai Pharmacy Licensing Examination, were administered: one standard set and one ETT set. Cardiology-focused clinical case scenarios were additionally presented using a SOAP-note format. Performance was assessed using a rubric adapted from the American Society of Health-System Pharmacists (ASHP) guidelines across four domains: question classification, source citation, evidence application, and communication. All four chatbots surpassed the 60% passing threshold: 72.08–83.33% for the standard set and 73.75–80.83% for the ETT set. Perplexity Pro demonstrated the highest MCQ performance, while the inclusion of English technical terms did not consistently improve results across models. In the clinical case assessment, ChatGPT Pro achieved the highest rubric score, showing strong evidence use and clinical reasoning. These findings suggest that AI chatbot competency in drug-related queries is broadly comparable to that of pharmacists. The ETT technique did not produce notable performance differences between Thai-only and mixed-language prompts. AI chatbots may support pharmacists by generating initial structured responses to drug information requests. The proposed evaluation framework provides a potential model for assessing AI competency in non-English settings and underscores the importance of jointly evaluating knowledge accuracy and clinical judgment to ensure safe integration into drug information services.
Inthira Kanchanaphibool, Panyanat Aonpong, Thanaphat Dabngoen et al.· Thai Bulletin of Pharmaceuti...· 0 citations
Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic.
Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale.
Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness (
p
= 0.068) and a language effect on safety (
p
= 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20,
p
= 0.009) and completeness (κ = 0.29,
p
= 0.001) and showed a very low κ in the safety scale (κ = 0.02,
p
= 0.81).
AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
Sufyan Alomair, Maryam Alsuwayq, Walla Alabbad et al.· Frontiers in Medicine· 0 citations
Hospitals frequently face challenges in delivering timely and accessible information to patients due to high inquiry volumes, language barriers, and limited staff availability. This paper proposes a multilingual, voice-enabled hospital chatbot that provides real-time assistance through both text and speech interfaces. The proposed system leverages Large Language Model (LLM)-based sentence embeddings using Sentence-BERT for semantic similarity-driven question answering, along with Google Translate API for multilingual support and Google Text-toSpeech for voice responses. The chatbot supports multiple Indian languages, including English, Hindi, Telugu, Tamil, Kannada, and Marathi, enabling inclusive communication across diverse user groups. Designed as a Flask-based web application with a responsive Bootstrap interface, the system aims to achieve effective contextual understanding, reduced response time, and improved accessibility when compared to traditional rule-based hospital inquiry systems. The proposed approach highlights the potential of LLM-driven semantic retrieval-based conversational agents in enhancing patient engagement and improving operational efficiency in healthcare environments.
Kamisetty Mythri Sridevi, Kallagunta Srividhya· International Journal of Eng...· 0 citations
INTRODUCTION
Telehealth is a strategic component of primary health care and has advanced in Brazil through the National Telehealth Program. Its benefits can be enhanced by artificial intelligence (AI), which has emerged as a promising tool. This study aims to compare the performance of real human and AI-generated responses to queries submitted to the teleconsultation services of the Telehealth Center of the UFMG Faculty of Medicine (NUTEL FM-UFMG), a member of the Telehealth Brazil Program.
METHODS
This is a comparative cross-sectional study of 180 real human and AI-generated responses, evaluated in a blinded manner according to quality criteria (medical adequacy, conciseness, coherence, and comprehensibility), risk potential, authorship identification accuracy, and inquiry resolution. Data from NUTEL FM-UFMG (January 2020 to May 2024) were utilized, covering cardiology, endocrinology, and obstetrics/gynecology (OB-GYN). Statistical analysis included the Shapiro-Wilk test, Kruskal-Wallis test, Nemenyi multiple comparison test, chi-square test, and Fisher's exact test.
RESULTS
Across all specialties, a significant difference was observed in comprehensibility, with AI mean scores surpassing those of humans. For the remaining quality criteria, as well as for risk potential and inquiry resolution, no significant differences were found, despite AI scoring higher than humans. Within specific specialties, significant differences was observed in endocrinology (except conciseness) and cardiology (in conciseness); AI showed superior means. Across all specialties, as well as individually within endocrinology and OB-GYN, the accuracy of authorship identification (human vs. AI) was statistically significant.
CONCLUSION
Despite existing limitations, AI demonstrates substantial potential as a support tool for teleconsultation services.
Gabriela Dário Mendes Barros, Carlos Eduardo Menezes Amaral, César Macieira et al.· Telemedicine journal and e-h...· 0 citations