When Stable Answers are Not Enough: Evaluating Response Consistency and Medical Alignment in Consumer-Facing LLMs
Abstract
Large Language Models (LLMs) are increasingly used by consumers as sources of health information. Evaluating the quality of such systems requires considering multiple quality dimensions rather than relying on a single indicator. Response consistency is often interpreted as a sign of reliability, but it remains unclear whether consistent responses also align with medically grounded reference information. This paper examines two quality dimensions of consumer-facing health LLMs: response consistency and reference-based medical alignment. Four general-purpose LLMs, namely ChatGPT, Claude, Gemini, and Mistral, were evaluated on 60 real patient questions from the K-QA dataset. Each model answered each question five times. Consistency was measured using lexical and semantic similarity metrics, while medical alignment was assessed against physicianauthored answers and Must-Have criteria through self- and crossevaluation. The results show high semantic consistency across all models, while reference-based medical alignment varied substantially across models, question types, and evaluator conditions. Correlation analysis revealed only weak associations between consistency and alignment, indicating that stable responses do not necessarily contain the medically relevant information expected by the reference criteria. These findings suggest that response consistency and reference-based medical alignment should be evaluated independently when assessing consumer-facing health LLMs, rather than treating response stability as evidence of medically adequate behaviour.