Skip to content

Author

Daniel Smolyak

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

Scalable extraction of social determinants of health from clinical notes in a sepsis cohort using instruction-tuned language models.

Objectives Social determinants of health (SDOH) are incompletely captured in structured electronic health records (EHRs) but are frequently documented in unstructured clinical notes. We evaluated large language models (LLMs) for extracting SDOH from clinical text. Materials and Methods We constructed an adult sepsis cohort from the Medical Information Mart for Intensive Care-IV (Sequential Organ Failure Assessment ≥2) and analyzed clinical notes from 1 year prior to 30 days following suspected infection. Three instruction‑tuned, decoder‑only LLMs (Mistral‑Instruct‑7B-v0.2, DeepSeek‑R1‑Distill‑Qwen‑14B, and GPT‑oss‑20B) were evaluated using structured prompts with predefined label schemas and few‑shot examples. Performance was benchmarked against a clinically validated annotated dataset and compared with a fine‑tuned encoder‑decoder baseline. Macro‑F1 scores were reported. A gold‑standard Intensive Care Unit (ICU) sepsis subset was independently annotated by 3 reviewers to assess domain‑level performance and ensemble strategies. Results Decoder‑only models outperformed the fine‑tuned encoder‑decoder baseline across SDOH domains. GPT‑oss achieved the highest macro‑F1 score (0.79) compared with Flan‑T5‑XXL (0.57). Prompt refinement substantially improved extraction accuracy. Ensemble majority voting increased robustness across domains, while unanimous agreement yielded high precision but limited coverage. In a subsequent mortality analysis, extracted SDOH did not independently predict 30-day mortality, which was instead associated with established clinical and demographic risk factors. Discussion Instruction‑tuned decoder‑only LLMs can reliably extract multiclass SDOH from unstructured clinical notes without task‑specific fine‑tuning. Ensemble and agreement‑based strategies provide practical operating points for high‑precision clinical deployment. Conclusion These findings support the feasibility of leveraging LLMs to enrich EHRs with structured SDOH data, providing a scalable approach for incorporating social context into downstream risk stratification and health outcome prediction.

D. Salazar, Pankaj Dipankar, Daniel Smolyak et al. · 0 citations