When 99.98% is too Good to be True: Preventing Rule-Induced Overfitting in Embodied Clinical AI for Surgical Readmission Prediction
Embodied Artificial Intelligence (AI) systems are increasingly used to support clinical decision-making, telemedicine follow-up, and resource allocation, particularly in remote and resource-constrained healthcare settings. In these deployments, predictive models are embedded within clinical workflows and operate under human oversight, making safety, transparency, and reliability essential for regulatory-compliant use. A key but underexplored factor affecting trustworthiness is how supervision signals are derived from unstructured clinical text. This paper analyses the impact of text-derived label construction on embodied clinical AI for postoperative risk monitoring. Using surgical readmission prediction from free-text clinical notes, we compare two weak supervision strategies: a myopic keyword-based heuristic and a context-aware labeling framework that accounts for negation, temporal scope, and clinical severity. Although the naive approach achieves near-perfect accuracy (up to 99.98%), we show that this performance is driven by rule-induced target leakage, where models learn documentation artifacts rather than clinically meaningful deterioration signals. We propose an auditable, context-aware labeling protocol aligned with the requirements of deployable clinical decision support systems. While trading inflated accuracy for more realistic performance, the proposed approach improves discrimination and patient-level calibration-properties essential for safe human-in-the-loop operation in telemedicine and remote care. These findings highlight that trustworthy embodied AI depends not only on model sophistication, but also on clinically grounded and transparent supervision mechanisms.