Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology
BackgroundGeneral-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance, pathology-relevant errors and hallucinations, their burden, and potenti...