Skip to content
Open access

Beyond Keyword Filters: Calibrated Monte-Carlo Risk Gating for Safe Multilingual Colorectal-Cancer LLM Dialogue

Sep 2026 · Big Data and Cognitive Computing · 0 citations · 20 references

Abstract

Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating the risk of a user turn from a bootstrap ensemble over multilingual sentence representations, calibrating it with Platt scaling and deciding at a single threshold between an informative answer and referral to a clinician. Evaluated on 450 oncologist-approved prompts in English, Turkish and Spanish under a scenario-level split, the gate reaches a guardrail F1 of 0.961, blocks 98.2 percent of harmful prompts and refuses 10.8 percent of legitimate questions, while the keyword, regular-expression and fuzzy layers that dominate current practice reach at most 0.034 and fire on five of 450 prompts, a separation that holds at a corrected q of 0.0003 and survives Bonferroni correction. Ablation locates the mechanism, with the semantic representation carrying the discriminative signal, Platt scaling lowering the expected calibration error from 0.216 to 0.069, and the ensemble predicting its own errors at an AUROC of 0.856 and removing them entirely at 50 percent coverage. Calibrated selective prediction over semantic representations makes multilingual medical-dialogue safety measurable, tunable to an explicit operating point, and consistent across the three languages tested.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.