LLM-Augmented Hybrid Representations for Disease Category Classification from Clinical Notes
Abstract
Clinical notes contain rich diagnostic evidence, but their long, noisy, and heterogeneous form makes automatic disease category classification difficult. We study a DiReCT-derived closed-set disease category classification task and propose an LLM-Augmented Hybrid Representation for lightweight clinical note classification. The central intuition is that an LLM can act as a clinical text processor, reformulating noisy notes into compact, clinically focused feature text that emphasizes diagnostically relevant cues and facilitates downstream representation learning and classification. Instead of using the LLM as the final predictor, we use it to generate clinically focused feature text from each note. This generated text is encoded with TF-IDF and MPNet, then fused with lexical and semantic representations of the original note before XGBoost classification. Across three random seeds on a fixed train/validation/test split, TF-IDF achieves 0.735 mean test accuracy, and combining TF-IDF with MPNet improves performance to 0.804. The proposed LLM-Augmented Hybrid reaches 0.850 mean test accuracy, achieving the best observed performance over the strongest original-note baseline and a reproduced same-model, same-task, closed-set direct LLM classifier baseline at 0.824. These results show that LLM-generated feature text provides useful representation augmentation for clinical note classification and offers a practical way to use LLMs beyond direct prediction.