Skip to content
Open access

Design and Evaluation of a Source-Grounded Medical LLM for Clinical Decision Support and Patient Care in Trustworthy Diagnostic Systems

Sep 2026 · Diagnostics · 0 citations · 31 references

Abstract

Background/Objectives: Large language models (LLMs) are increasingly explored for clinical decision support and digital health applications. However, reliable diagnostic assistance remains challenging for low-resource medical languages such as Turkish due to limited source grounding, transparency, clinical safety, and localized medical knowledge. This study presents TurkishMedLLM, a source-grounded and safety-aware Turkish medical LLM designed for clinician-supervised diagnostic decision support, symptom interpretation, and patient care. Methods: The methodology integrates multi-source Turkish medical data ingestion, schema standardization, duplicate removal, quality filtering, supervised fine-tuning, embedding generation, vector indexing, Qwen3-8B fine-tuning, retrieval-augmented generation (RAG), and multi-layer evaluation. The system was evaluated using retrieval metrics, ROUGE, RAGAS, DeepEval, and clinical safety assessments. As a use case, TurkishMedLLM was integrated into the AI-based Diabetes Care (AIDCare) mHealth platform, which supports patient queries related to lifestyle management, symptoms, diagnosis, and treatment of diabetes. The system generates safety-aware responses with clinician-in-the-loop validation before delivery through the mobile application. Results: After pre-processing, the final dataset comprised 232,926 unique documents, including 210,791 Turkish medical question–answer pairs and 22,135 hospital medical articles. The retrieval module achieved Hit Rate@1, Hit Rate@3, and Hit Rate@5 of 94.67%, 98.67%, and 100.00%, respectively, indicating consistent retrieval of clinically relevant evidence. QLoRA fine-tuning achieved a validation loss of 0.9373 and ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum scores of 0.8338, 0.5476, 0.7836, and 0.7372, respectively. The fine-tuned RAG system achieved RAGAS faithfulness and response relevancy scores of 0.91 and 0.88, respectively, while DeepEval achieved an answer relevancy score of 0.90. Clinical safety assessment achieved a caution score of 0.93, indicating generally evidence-grounded and clinically cautious responses for symptom interpretation and diagnosis-related patient support. Conclusions: Combining retrieval grounding, parameter-efficient fine-tuning, and multi-layer safety evaluation provides a promising approach for clinician-supervised medical AI in Turkish. TurkishMedLLM demonstrates potential for symptom interpretation, differential diagnostic support, and trustworthy digital health applications.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.