Skip to content
Review Open access

Overview of #SMM4H-HeaRD 2026 – Task 6: Predicting TNM staging from pathology reports

2026 · Proceedings of the 11th Social Media Mining for Health Research and Applications (SMM4H-HeaRD 2026) Workshop and Shared Tasks · pp. 332-337 · 1 citation · 17 references

TL;DR

These results suggest that while fine-tuned domain-specific encoders excel at surface-level extraction, larger general-purpose LLMs may be more robust when staging must be inferred from contextual clinical findings.

Abstract

This paper provides an overview of Task 6 from the Social Media Mining for Health/Health Real-World Data shared task (#SMM4H-HeaRD 2026), which focused on predicting TNM staging from pathology reports from TCGA. Seven teams submitted systems spanning fine-tuned clinical encoders, open-source generative LLMs, and closed-source API models. On a straightforward test set, most teams achieved near-perfect F1 scores (average 0 . 993 , 0 . 972 , and 0 . 957 for T, N, and M). However, on a harder tiebreak set where explicit TNM notation was removed and staging had to be inferred from clinical descriptions, performance dropped substantially (average 0 . 725 , 0 . 783 , and 0 . 846 ). Notably, the two teams using large closed-source API models generalized best to the harder set, achieving the highest T and N scores despite not leading on the easy set. These results suggest that while fine-tuned domain-specific encoders excel at surface-level extraction, larger general-purpose LLMs may be more robust when staging must be inferred from contextual clinical findings. All teams surpassed baseline overall performance on both test sets.

Read PDF

Similar papers

Open access Jul 2026

Development of a benchmarking dataset for symptom detection using large language models

Abstract Objectives To develop a pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters. Materials and Methods We created a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts (264 encounters; 16 symptoms; double-coded and adjudicated). Nine different LLMs from 4 vendors (OpenAI, Meta, DeepSeek, Moonshot AI) were used as examples to test our evaluation pipeline; outputs were assessed for correct structure and symptom information. Results Of 3085 excerpts, 2087 (68%) contained symptoms. Pain, cough, and shortness of breath were most common; LLMs achieved F1 scores ranging 0.66-0.88 for these symptoms with minimal prompt engineering. Of tested models, GPT-4.1 demonstrated the best overall performance. Discussion Our evaluation pipeline and benchmarking dataset are publicly available and applicable to various LLMs, including open-source models. Conclusion This work supports the development and optimization of models that seek to improve patient symptom understanding.

Joshua Davis, B. Durieux, C. V. van Dongen et al. · 0 citations
Open access Jul 2026

Large language models for interpretation of health checkup results

Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Jiwon You, Hangsik Shin · 0 citations
Open access Jul 2026

Fine-Tuned Large Language Models for Detecting Social Isolation from Unstructured Clinical Notes

Objectives: This study aimed to leverage FLAN-T5-Large, BERT, RoBERTa, and Gemma-2-2B, with fine-tuning, to identify instances of social isolation and social support within unstructured clinical notes. Materials and Methods: Annotated clinical note spans containing social context cues were used to fine-tune each model. Performance was evaluated using Accuracy, Precision, Recall, and Macro-F1 score. A structured prompt was used to instruct the model to perform classification task and mitigate overgeneralization. Performance comparisons across the models assessed sensitivity, robustness, and false positive reduction. Results: FLAN-T5-Large achieved highest performance, with Macro-F1 of 0.92{+/-}0.04, demonstrating balanced results across classes: social isolation (F1 = 0.91{+/-}0.03), no social isolation (F1 = 0.94{+/-}0.05), and social support (F1 = 0.90{+/-}0.04). Gemma-2-2B produced comparable results, with Macro-F1 score of 0.89{+/-}0.10. BERT and RoBERTa achieved lower Macro-F1 scores of 0.77{+/-}0.17 and 0.80{+/-}0.21 respectively, with variability across categories. Discussion: A major contribution of this work is precise identification of multiple concepts related to social connectedness. By integrating annotated examples of both true and false positives, including negations and contextually ambiguous terms, the model better distinguished relevant social context cues from noise. Training on both social isolation and support provided a dual framework for comparative analyses and patient stratification. Conclusion: Transformer-based NLP models, particularly FLAN-T5-Large, demonstrated potential for identifying social isolation and social support in clinical text. These findings support the use of generative AI techniques to enhance detection of social isolation from EHRs, advancing context-aware healthcare analytics.

L. Chinthala, C. Lemon, A. Shaban-Nejad et al. · 0 citations
Jul 2026

Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer.

PURPOSE To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes. METHODS Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139. RESULTS This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions). CONCLUSION Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.

Syed Arsalan Ahmed Naqvi, I. Riaz, Amir Saeidi et al. · 0 citations
Preprint Jul 2026

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized''Honest Threshold Tuning''procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.5790$ and a secondary $F_1$ of $0.9657$. In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary $F_1$ of $0.5780$ and a secondary $F_1$ of $0.9599$-essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall $0.3571$, ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging ($0.3564$), and a zero-shot MedGemma-4B run with a PubMed-style prompt ($0.3186$), spanning a wide range of model scales and training costs. Code: https://github.com/dsgt-arc/imageclef-caption-2026.

Bowen Wang, Youwen Zhang, Ritesh Mehta · 0 citations