Large Language Model Chatbots Cannot Reliably Calculate Clinical Risk Scores—A Comparative Accuracy Study
Background: Large language models (LLMs) are increasingly accessible to healthcare providers and patients for clinical decision support, yet their ability to perform precise mathematical calculations required for validated risk scores remains unexplored, and errors could compromise patient safety. The EuroSCORE II requ...