Leveraging dynamic prompting for outcome prediction of cancer patients using large language models and electronic health record notes
Abstract
Outcome prediction from unstructured EHR notes remains challenging, especially for rare cancers with limited representation in large language model (LLM) pre-training data. We present a dynamic prompting framework that retrieves semantically similar patient examples, constructs tailored few-shot prompts and integrates note summaries to enhance outcome predictions. We evaluated this approach on a retrospective cohort of 503 breast cancer and 475 glioma patients using EHR notes from the first 180 days post-diagnosis. Dynamic prompting substantially improved glioma prediction while producing modest improvements for breast cancer, consistent with our hypothesis that dynamic prompting benefits rare cancers more. The combined summarization-driven dynamic-prompting approach achieved the highest performance while maintaining stability. Embedding visualizations suggested alignment with established prognostic markers in both cohorts. In this proof-of-concept, dynamic prompting improved outcome prediction for a lower-incidence cancer (glioma) in a single cohort while maintaining performance for a common cancer (breast), establishing it as a promising approach for LLM-based outcome prediction from unstructured EHR notes. Validation across additional low-incidence tumor types and institutions is required before these findings can be generalized.