Skip to content

Diagnosis classification in EMR data using latent representations and SNOMED-CT mapping for improved medical data integration.

Aug 2026 · Medical and Biological Engineering and Computing · 0 citations · 18 references
Medicine

TL;DR

A diagnosis classification model that automatically maps diagnosis spans in EMR data to the standardized clinical ontology Systematized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) is developed, demonstrating strong potential for scalable and privacy-preserving medical concept normalization in real-world clinical environments.

View source

Similar papers

Open access Aug 2026

A text mining and ontology-based approach using phenotypes to obtain relevant literature for rare diseases

Summary Diagnosing rare diseases remains a major challenge due to limited clinical knowledge and the frequent absence of diagnostic criteria. We present a digital framework that leverages large language models and biomedical text embeddings to bridge this gap. By mapping Human Phenotype Ontology terms to a shared vector space with millions of PubMed abstracts and full-text articles, our method enables phenotype-driven semantic search and ranks literature relevant to patient symptoms, even without explicit disease mentions. Validated on OMIM-derived benchmarks and applied to RASopathies, including NF1, Noonan, and Costello syndromes, our approach retrieved expected findings, supporting differential diagnosis and research. The framework is implemented in an open-source Python package, py-semtools, and it can be integrated into clinical decision support systems or adapted to other ontologies and corpora. This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.

Jesús Pérez-García, Federico García-Criado, F. Pazos et al. · 0 citations
Conference Open access 2026

Natural Language Processing for Prediction of Chronic Diseases from Electronic Health Records

This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.

U. Luke, P. Asuquo, Victor Anaga et al. · 0 citations
Open access Jul 2026

Medical code embeddings from claims-based co-occurrences: a unified semantic space for ICD-10 diagnoses and ATC medications.

OBJECTIVE The analysis of care trajectories derived from electronic health records and claims data has become increasingly common in biomedical informatics. This has enabled large-scale studies of care processes, yet widely used binary code representations result in high-dimensional, sparse data that fail to capture semantic relationships between medical concepts. Learning dense vector representations (embeddings) has emerged as a promising approach to address these limitations. We aimed to construct and share joint embeddings for the International Classification of Diseases (ICD-10) and the Anatomical Therapeutic Chemical (ATC) classification system, providing reusable semantic representations of diagnoses and treatments from real-world claims data. MATERIALS AND METHODS Using claims records from 1.5 million patients, we defined code co-occurrences within temporal windows and constructed a Positive Pointwise Mutual Information (PPMI) matrix spanning ICD-10 and ATC codes. Singular Value Decomposition (SVD) was applied to derive a low-dimensional embedding space. Evaluation combined UMAP visualization, nearest-neighbor retrieval, and a code-level classification task based on ICD chapters and ATC classes. RESULTS The embeddings reflected the hierarchical organization of ICD-10 and ATC and revealed associations across coding systems, including clinically relevant diagnosis-treatment relationships. The classification task achieved mean AUCs of 0.93 for ICD-10 and 0.90 for ATC, indicating strong grouping of semantically related codes. DISCUSSION The embeddings provide a reusable, code-level semantic representation that can support code retrieval, reduce manual code grouping, and be aggregated into patient-level features without training a task-specific model. CONCLUSION We release the first openly available joint ICD-10-ATC embedding space derived from real-world claims data, providing a reusable resource for biomedical informatics research.

C. Faujour, S. Bouée, C. Emery et al. · 0 citations
Review Open access Jul 2026

The Orphanet Nomenclature and Classification of Rare Diseases for Improved Patient Recognition and Data Interoperability: Qualitative and Quantitative Analysis

Background Although individually uncommon, rare diseases (RDs) collectively affect an estimated 329-624 million people worldwide. There are over 6500 known RDs, 85% of which affect fewer than 1 person per million. Consequently, the critical amount of data necessary to improve knowledge, care, and treatment can only be achieved through cumulative data collection across countries. However, RDs remain underrepresented in medical terminologies and classification systems, hindering data sharing, interoperability, and public health monitoring. Objective This paper presents the Orphanet Nomenclature and Classification of RDs detailing its content, production and update methodology, and mappings to other semantic resources. It also provides an up-to-date count of RDs based on the consensus operational definition describing their distribution by medical domain. Methods The Orphanet Nomenclature of RDs is a multilingual standardized system composed of clinical entities, each defined by a unique and time-stable ORPHAcode, a preferred term, synonyms, a classification level, and a textual definition. This nomenclature is structured into 3 classification levels organized within a multihierarchical and multiparental classification system by medical domain. Its production, updates, and mappings to major biomedical resources rely on standardized and published procedures, continuous literature review, manual curation, and expert validation, reflecting advancements in RDs knowledge and clinical practice. Presented data metrics were computed using the Orphanet July 2025 release to quantitatively characterize the content, structure, classification, and semantic alignments of the Orphanet Nomenclature and Classification system. Results As of July 2025, the Orphanet Nomenclature of RDs includes a total of 9784 active clinical entities, including 6527 disorders (corresponding to the RDs definition), 1084 subtypes of disorders, and 2173 groups of disorders. Disorders are multiclassified into 29 classification hierarchies, each corresponding to a distinct medical domain, accurately representing the complex multisystemic nature of RDs. Extensive qualified mappings ensure semantic interoperability: 97.4% (6355/6527) of disorders are mapped to at least 1 ICD-10 (International Statistical Classification of Diseases, Tenth Revision) code (415/6527, 6.4% with an exact proximity relationship), 71.8% (4683/6527) are mapped to at least 1 ICD-11 (International Classification of Diseases, Eleventh Revision) Mortality and Morbidity Statistics code (958/6527, 14.7% with an exact relationship) and 94.8% (6191/6527) are mapped to Systematized Nomenclature of Medicine Clinical Terms (all with an exact relationship). Genetic disorders represent 72.2% (4715/6527) of all RDs, and 63.4% (4141/6527) are mapped to at least 1 phenotypic Online Mendelian Inheritance in Man number. Conclusions The Orphanet Nomenclature and Classification of RDs is the only RDs-specific interoperable medical terminology meeting the needs of health care, research, and public health systems. By addressing the underrepresentation of RDs in medical terminologies, it enables accurate RDs identification, coding, and monitoring, supporting cross-border data interoperability, and contributing to improved knowledge, policymaking, and ultimately better care for people living with an RD.

C. Lucano, D. Lagorce, A. Olry et al. · 0 citations
Preprint Aug 2026

xMICD: Explainable Representation of Multiple ICD Codes

Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing them effectively remains challenging. Existing approaches often face a trade-off between predictive performance and interpretability: grouping-based representations are interpretable but may lose information, while embedding-based representations achieve strong predictive performance but are difficult to interpret. We propose Explainable Representation of Multiple ICD Codes (xMICD), a method for constructing low-dimensional patient representations from sets of ICD codes. xMICD combines clinically meaningful diagnostic groupings with similarity in a pre-trained ICD embedding space. Instead of using binary group membership, the method assigns codes to groups via similarity-based relative assignments, yielding features that reflect how closely a patient's diagnoses align with each clinical group. Experiments on large-scale EHR datasets demonstrate that xMICD achieves predictive performance comparable to embedding-based representations such as ICD2Vec across multiple clinical prediction tasks. At the same time, the resulting features remain clinically interpretable because each dimension corresponds to a recognizable diagnostic group. xMICD therefore provides a practical way to integrate embedding-based semantic relationships into interpretable clinical feature spaces for machine learning models.

P. Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong et al. · 0 citations
Open access Aug 2026

Precision pharmacology: deep learning infused ontological framework with E-GRU enhancement for tailored medicine prescriptions

Advanced Clinical Decision Support Systems significantly influence patient care, with medicine prescriptions being a vital area of research. Ontology, a growing discipline in the semantic web, enables hierarchical domain representation, thereby allowing finer data access to be achieved. Deep Learning (DL) supports pattern recognition in Electronic Health Records (EHR), which include patient demographics and diagnosis histories. Prescribing medications with minimal adverse effects is crucial, particularly for patients who require multiple drugs, as drug interactions can result in more complex conditions. This study introduces an integrated approach that combines Ontology with DL neural networks to improve prescription accuracy. This study proposes NexusOpti, a model featuring an Enhanced Gated Recurrent Unit (E-GRU) layer. To understand drug–disease interactions, hierarchical data were extracted from the International Classification of Diseases (ICD) and Anatomical Therapeutic Chemical (ATC) ontologies. These structured data were processed using a self-attention mechanism to enhance the recommendation precision. This integration not only addresses data security concerns but also improves the accuracy of the medicine recommendations. The model was evaluated using key metrics such as the hit ratio and normalised discounted cumulative gain (NDCG). The NexusOpti model, incorporating the Enhanced Gated Recurrent Unit (E-GRU) layer, outperforms the existing GRU model in terms of NDCG and Hit Ratio metrics. 13% of improvement in performance was oberved to the comparison between NexusOpti with the E-GRU and the GRAM baseline model. These findings highlight the effectiveness of the model in advancing personalised, safer, and data-driven medication prescriptions.

Harichandra Khalingarajah, A. Vasudevan, P. Abinaya et al. · 0 citations