Skip to content

SentryLine: Evidence-Grounded Question Answering over Evolving Documents in Oncology Care

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

SENTRYLINE is presented, a living guideline-aware clinical question answering system that retrieves guideline passages through a vectorless hierarchical RAG pipeline and returns a role-specific answer with inline citations, factual and temporal verification reports, and drift detection notes that surface when a guideline has been updated.

Abstract

Oncology care operates at constant pressure of absorbing rapidly evolving evidence base in biomedicine. The American Society of Clinical Oncology (ASCO) addresses this through living guidelines, but the format introduces a new burden: any recommendation can change at any point, across multiple versioned documents. We present SENTRYLINE, a living guideline-aware clinical question answering system. SENTRYLINE retrieves guideline passages through a vectorless hierarchical RAG pipeline and returns a role-specific answer with inline citations, factual and temporal verification reports, and drift detection notes that surface when a guideline has been updated. We construct ASCOBENCH, a benchmark of 405 three-turn conversations across four question categories with gold answers from expert annotators(clinicians), and use test set to evaluate SENTRYLINE against five baselines under an LLM-as-judge framework. Experiments across three generation backbones show consistent improvements over four retrieval baselines and ASCO's guideline assistant, with particularly strong gains on Reasoning and Role-Specific questions where multi-hop synthesis and register adaptation are required

View source

Similar papers

Preprint Aug 2026

Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval

Case2Flow is introduced, a task designed to retrieve the most relevant guideline flowchart for a given patient case from a collection of guideline documents and CRISP, a training-free scoring method that sharpens late-interaction retrieval by suppressing uninformative patches, discounting ambiguous token matches, and i...

Jia-Le Wei, Yu-Fan Chen, Alexander Jaus et al. · 0 citations
Preprint Sep 2026

A Living Benchmark for Information Retrieval from Electronic Health Records

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated,...

J. Cahoon, C. Stanwyck, Sulaiman Somani et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models

Results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision, and suggest that EMQ gives the most stable external transfer and retention among same-budget SFT variants.

Yangmin Huang, Shu Quan, He Geng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to newly available evidence, but retrieved information may be irrelevant, incomplete, or conf...

Shuai Wang, Yi-Ze Zhao, Qing-Yu Chen · 0 citations
Open access Sep 2026

Beyond Keyword Filters: Calibrated Monte-Carlo Risk Gating for Safe Multilingual Colorectal-Cancer LLM Dialogue

Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating the...

Abdurrahim Kızılay, Kerem Gencer · 0 citations
#natural language process... Preprint Oct 2026

Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes

Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical rera...

Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.