Abstract Background Adverse drug events (ADEs) pose significant public health challenges and economic burdens. While substantial ADE information is documented in unstructured clinical notes, its extraction remains difficult due to semantic complexity. Large language models (LLMs) offer promising text comprehension capabilities but are often hindered by domain-specific hallucinations. Objective This study aims to evaluate the effectiveness of retrieval-augmented generation (RAG) in improving the identification of ADEs using LLMs from Chinese clinical narratives and to establish a paradigm for this task. Methods We collected and preprocessed 19,983 Chinese clinical notes, retaining 18,432 high-quality records. Following a rigorous annotation and deduplication process, we established a gold-standard reference dataset (n=2510) and an ADE knowledge base (n=5144) using a standardized JSON schema. We evaluated 3 state-of-the-art LLMs (DeepSeek-V3 [DeepSeek], ERNIE 3.5-8K [Baidu], and GPT-4o [OpenAI]) under 3 prompt strategies: nonaugmented generation (NAG), static-augmented generation (SAG), and RAG. Performance was comprehensively assessed using precision, recall, and F1-score across 3 recognition matching levels (L1 exact, L2 sentence, and L3 overlap) via 1000 bootstrap resamples. Model robustness was further validated from real-world clinical progress notes, reflecting real-world ADE prevalence. Results We successfully constructed and publicly released the first Chinese ADE corpus derived from clinical notes. Across the tested LLMs, RAG yielded higher F1-scores than NAG and SAG at the L3 level. The optimal configuration, DeepSeek-V3 with RAG, achieved an overall L3-level F1-score of 0.9638 (95% CI 0.9541‐0.9727). Notably, the RAG approach increased the recall of GPT-4o from 0.6419 under NAG to 0.9241 under RAG (FDR P=.003). Evaluation on real-world datasets demonstrated clinical utility, with the RAG prompt maintaining high discriminatory capability (specificity: 0.9821; F2-score: 0.8885). Error analysis revealed that RAG successfully resolved common identification errors, both omissions and commissions, that were intractable for nonaugmented models. Conclusions Synergizing a curated domain-specific knowledge base with LLMs via a RAG architecture is an effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes. This approach can mitigate hallucinations in LLMs, providing a foundational open-source benchmark and a robust technical framework to advance pharmacovigilance, drug safety research, and clinical decision support.
Junlong Ma, Xuehong Wu, Zeying Feng et al.· Journal of Medical Internet...· 0 citations
Membrane molecular recognition features (MemMoRFs) are lipid-binding intrinsically disordered regions (IDRs) that undergo disorder-to-order transitions to mediate critical membrane dynamics. Consequently, their dysregulation is closely linked to severe human pathologies, including neurodegenerative diseases and viral infections. Despite their biological significance, annotations for MemMoRFs are scarce, limiting the accuracy of computational predictors. We introduce PreMemMoRF, a deep learning framework that leverages transfer learning to alleviate data scarcity. The model is pre-trained on linear interacting peptides (LIPs) with similar conformational transitions and fine-tuned on MemMoRF datasets, capturing generalizable binding-related sequence features. PreMemMoRF outperforms existing predictors across multiple metrics and demonstrates robust performance on transmembrane and membrane-associated proteins. It also performs consistently in short linear motif prediction, highlighting cross-task generalizability. Proteome-wide analysis in yeast shows that predicted scores exhibit systematic differences across distinct transmembrane topological regions and are consistent with established physicochemical constraints of membrane proteins. Collectively, these results validate PreMemMoRF as a robust and reliable computational framework for the large-scale identification of MemMoRFs.
Chenxi Xia, Jiayi Hao, Hao Liu et al.· IEEE journal of biomedical a...· 0 citations