Skip to content
Conference

When Stable Answers are Not Enough: Evaluating Response Consistency and Medical Alignment in Consumer-Facing LLMs

Aug 2026 · 2026 IEEE 34th International Requirements Engineering Conference Workshops (REW) · pp. 487-496 · 0 citations · 30 references

Abstract

Large Language Models (LLMs) are increasingly used by consumers as sources of health information. Evaluating the quality of such systems requires considering multiple quality dimensions rather than relying on a single indicator. Response consistency is often interpreted as a sign of reliability, but it remains unclear whether consistent responses also align with medically grounded reference information. This paper examines two quality dimensions of consumer-facing health LLMs: response consistency and reference-based medical alignment. Four general-purpose LLMs, namely ChatGPT, Claude, Gemini, and Mistral, were evaluated on 60 real patient questions from the K-QA dataset. Each model answered each question five times. Consistency was measured using lexical and semantic similarity metrics, while medical alignment was assessed against physicianauthored answers and Must-Have criteria through self- and crossevaluation. The results show high semantic consistency across all models, while reference-based medical alignment varied substantially across models, question types, and evaluator conditions. Correlation analysis revealed only weak associations between consistency and alignment, indicating that stable responses do not necessarily contain the medically relevant information expected by the reference criteria. These findings suggest that response consistency and reference-based medical alignment should be evaluated independently when assessing consumer-facing health LLMs, rather than treating response stability as evidence of medically adequate behaviour.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.