SENTRYLINE is presented, a living guideline-aware clinical question answering system that retrieves guideline passages through a vectorless hierarchical RAG pipeline and returns a role-specific answer with inline citations, factual and temporal verification reports, and drift detection notes that surface when a guideline has been updated.
Abstract
Oncology care operates at constant pressure of absorbing rapidly evolving evidence base in biomedicine. The American Society of Clinical Oncology (ASCO) addresses this through living guidelines, but the format introduces a new burden: any recommendation can change at any point, across multiple versioned documents. We present SENTRYLINE, a living guideline-aware clinical question answering system. SENTRYLINE retrieves guideline passages through a vectorless hierarchical RAG pipeline and returns a role-specific answer with inline citations, factual and temporal verification reports, and drift detection notes that surface when a guideline has been updated. We construct ASCOBENCH, a benchmark of 405 three-turn conversations across four question categories with gold answers from expert annotators(clinicians), and use test set to evaluate SENTRYLINE against five baselines under an LLM-as-judge framework. Experiments across three generation backbones show consistent improvements over four retrieval baselines and ASCO's guideline assistant, with particularly strong gains on Reasoning and Role-Specific questions where multi-hop synthesis and register adaptation are required
Case2Flow is introduced, a task designed to retrieve the most relevant guideline flowchart for a given patient case from a collection of guideline documents and CRISP, a training-free scoring method that sharpens late-interaction retrieval by suppressing uninformative patches, discounting ambiguous token matches, and i...
Jia-Le Wei, Yu-Fan Chen, Alexander Jaus et al.· 0 citations
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated,...
J. Cahoon, C. Stanwyck, Sulaiman Somani et al.· 0 citations
Results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision, and suggest that EMQ gives the most stable external transfer and retention among same-budget SFT variants.
Yangmin Huang, Shu Quan, He Geng et al.· 0 citations
Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to newly available evidence, but retrieved information may be irrelevant, incomplete, or conf...
Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating the...
Abdurrahim Kızılay, Kerem Gencer· Big Data and Cognitive Compu...· 0 citations
Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical rera...
Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.