Skip to content

Bioinformatics and Machine-Learning Identification of Pulmonary Hypertension Biomarkers and Candidate Therapeutic Compounds.

Aug 2026 · Journal of Visualized Experiments · Vol 234 · 0 citations
Medicine

TL;DR

Findings support the five genes as candidate PH biomarkers and VX-745 as a computational drug-repositioning hypothesis requiring experimental validation and VX-745 as a computational drug-repositioning hypothesis requiring experimental validation.

Abstract

This study aimed to identify pulmonary hypertension (PH)-associated molecular biomarkers and candidate small-molecule compounds using public transcriptomic data and independent validation resources. Three Gene Expression Omnibus datasets (GSE22356, GSE33463, and GSE48149) were integrated following normalization, probe annotation, and ComBat batch-effect correction. Differential expression analysis, weighted gene co-expression network analysis, functional enrichment analysis, protein-protein interaction network analysis, and three machine-learning algorithms were used to identify core feature genes. Diagnostic performance was evaluated using receiver operating characteristic curves. External validation included an independent lung-tissue cohort (GSE117261), a pulmonary artery single-cell RNA-sequencing dataset (GSE210248), and quantitative reverse transcription PCR validation in independent lung-tissue samples. Connectivity Map-based drug repositioning and molecular docking were used to screen candidate compounds. Seventy-eight differentially expressed genes were identified, and CXCL10, JUN, IFIH1, MX1, and TLR7 were selected as core feature genes. In the independent GSE117261 lung-tissue cohort, JUN showed the strongest external support, whereas replication of the other genes was variable. Quantitative reverse transcription PCR in 20 biologically independent pulmonary arterial hypertension samples and 20 control samples confirmed upregulation of all five genes. The apparent five-gene qRT-PCR model and 100 repeated stratified five-fold cross-validation analyses both yielded an area under the curve of 1.000, although the small cohort requires cautious interpretation and independent prospective validation. Single-cell analysis of GSE210248 supported altered communication between immune and structural cells and smooth muscle cell phenotypic switching. BRD-K91900765/VX-745 ranked highest in Connectivity Map screening. MAPK14/p38α, its established pharmacological target, was included as a positive-reference docking protein, whereas docking against the five biomarker-associated proteins was treated as exploratory. These findings support the five genes as candidate PH biomarkers and VX-745 as a computational drug-repositioning hypothesis requiring experimental validation.

View source

Similar papers

Open access Aug 2026

Integrative transcriptomic and machine learning analyses identify NOL7 and TXLNA as candidate blood biomarkers for Parkinson’s disease within an RNA modification-associated co-expression network

Parkinson’s disease (PD) pathogenesis involves complex molecular mechanisms, with emerging evidence implicating RNA modifications (RM). This study sought to explore RM-associated key genes and their roles in PD progression. Transcriptomic datasets GSE6613 (training) and GSE72267 (validation) were analyzed. Differentially expressed genes (DEGs) between PD and controls were recognized. WGCNA was performed to identify RM-associated module genes. Machine learning algorithms, Receiver Operating Characteristic (ROC) analysis, and gene expression validation were applied to screen key genes. Nomogram construction, functional enrichment, immune infiltration analysis were conducted to investigate the biological mechanisms and therapeutic potential of the key genes. The bioinformatics findings were further supported by RT-qPCR experiments in a small clinical cohort ( n  = 5 per group), though these preliminary results require validation in larger independent samples. Through intersection of 507 DEGs and 2,091 RM-associated module genes, 63 candidate genes related to RM in PD were identified. Machine learning, ROC analysis, gene expression validation, and clinical experiments were further employed to identify two key genes (NOL7 and TXLNA). A nomogram constructed based on key genes demonstrated moderate diagnostic efficacy for PD, with an area under the curve of 0.792. Enrichment analyses revealed associations of the key genes with neuroactive pathways, such as spliceosome and ribosome. Immune infiltration analysis suggested a negative correlation between NOL7 and NKT (cor = -0.30, P  < 0.05). Furthermore, danazol was predicted to be a compound associated with both NOL7 and TXLNA. Molecular docking analysis revealed that danazol exhibited a relatively favorable binding affinity for NOL7 (-6.3 kcal/mol), whereas its binding affinity for TXLNA was weaker (-4.7 kcal/mol). NOL7 and TXLNA were validated as blood biomarkers for PD derived from an RNA modification-associated transcriptional module, offering insights into epigenetic dysregulation and immune interactions. The nomogram provided a preliminary framework for PD risk assessment, and the drug prediction offered potential candidates for future therapeutic exploration.

Meiling Chen, Peng Chen, Liya Suo et al. · 0 citations
Open access Aug 2026

Identification and Validation of Mannose Metabolism-Related Biomarkers in COPD Through Integrated Bioinformatics and Machine Learning Analysis: A Pilot Study

Background Chronic obstructive pulmonary disease (COPD) represents a progressive respiratory disorder marked by sustained airflow restriction and ongoing inflammatory processes. Recently, mannose metabolism has emerged as a significant factor in chronic disease development. This investigation explored how mannose metabolism-related genes (MMRGs) contribute to COPD pathogenesis and evaluated their utility as candidate diagnostic and therapeutic targets. Methods We obtained blood sample gene expression data from COPD patients and healthy controls via the GEO database. Differentially expressed genes (DEGs) were identified and intersected with MMRGs to obtain candidate genes. Three machine learning algorithms combined with expression validation across independent datasets were applied to identify biomarkers. A nomogram prediction model was constructed and its diagnostic performance was assessed using receiver operating characteristic (ROC) curve analysis. Subsequently, gene set enrichment analysis (GSEA), immune infiltration analysis, drug prediction, and molecular docking were performed. Results Nineteen candidate genes were identified from 1685 DEGs and subsequently screened for two biomarkers: MAN1C1 and MAN2B2. A nomogram model constructed on the basis of the two showed moderate discriminatory efficacy (area under the curve (AUC) = 0.701). In addition, GSEA analysis showed that both were co-enriched in pathways such as TNF’s target up-regulated gene sets. The immune infiltration results revealed significant differences (p < 0.05) between COPD and controls in a total of 12 categories of immune cells, such as activated B cells. Finally, drug prediction revealed 12 and 3 potential drugs for MAN1C1 and MAN2B2, respectively, with trichostatin A showing a potential binding conformation. Conclusion This study revealed the potential roles of MMRGs in COPD and identified novel biomarkers. These findings provided new insights and research foundations for the early diagnosis, personalized treatment, and drug development of COPD.

Xiaodan Li, Jin Wang, Zhongbai Hu et al. · 0 citations
Open access Jul 2026

Integrative Bioinformatics and Machine Learning Analysis Identifies Novel Molecular Biomarkers in Prostate Adenocarcinoma

Prostate adenocarcinoma is characterized by substantial inter-patient heterogeneity, limiting the clinical reliability of conventional diagnostic tools, including prostate-specific antigen testing. This limitation underscores the need for robust molecular biomarkers that may complement conventional diagnostic tools, highlighting the urgent need for biomarkers capable of enhancing diagnostic accuracy and enabling more precise risk stratification. In the present study, transcriptomic data from The Cancer Genome Atlas (TCGA) were analyzed using an integrative bioinformatics and machine learning pipeline., The proposed workflow was designed as a stepwise and reproducible biomarker prioritization framework in which differential expression analysis, functional enrichment, protein–protein interaction (PPI) based network interpretation, graph-convolutional feature selection, and hybrid ensemble machine learning were sequentially integrated. Differential gene expression analysis was combined with pathway enrichment (Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Reactome), protein–protein interaction network construction, and graph-convolutional feature selection. Multiple machine learning algorithms, including Random Forest, Gradient Boosting Machine, Support Vector Classifier, Artificial Neural Network, and AdaBoost, were systematically evaluated. A hybrid ensemble model integrating Gradient Boosting Machine and Random Forest (GBM+RF) was subsequently developed. Model performance was assessed using accuracy, sensitivity, specificity, and area under the Receiver Operating Characteristic (ROC) and externally validated using the independent GSE14206 dataset. The analysis revealed a coordinated molecular pattern characterized by dysregulated cell cycle activity and enhanced interferon-mediated immune signaling. Protein–protein interaction analysis identified STAT1 and PLK1 as highly connected network hub genes within immune-related and cell-cycle-associated modules. Among the evaluated models, the hybrid GBM+RF framework achieved the highest predictive performance on the TCGA dataset, with AUC: 0.9526; Accuracy: 97.49%. External validation using the GSE14206 dataset confirmed the robustness of this model (AUC: 0.9156; Accuracy: 91.53%). These findings support a broader multi-gene candidate signature in prostate adenocarcinoma, in which machine learning prioritized genes such as XAF1, APP, RPA3, IFIH1, UBE2D2, RSAD2, KIF2C, and PLK1, while STAT1 and PLK1 provided complementary network-level biological relevance. The proposed framework provides a robust and transferable strategy for biomarker discovery and precision oncology.

H. Kurt, Sabire Kılıçarslan, M. M. Çiçekliyurt et al. · 0 citations
Open access Jul 2026

Bioinformatics identification of candidate biomarkers associated with T cell proliferation in atherosclerosis via WGCNA and machine learning.

BACKGROUND Atherosclerosis (AS) is a complex chronic disease caused by the development of atherosclerotic plaques. T cell proliferation exerts a vital influence on development of the AS. This research aimed to conduct a comprehensive analysis to computationally screen candidate T cell proliferation-related biomarkers in AS and to explore their potential molecular mechanisms. METHODS The transcriptional datasets of AS patients were obtained from the Gene Expression Omnibus (GEO) repository. To identify differentially expressed-T cell proliferation-related genes (DE-TPRGs), we integrated differential gene expression analysis with weighted gene co-expression network analysis (WGCNA), taking the intersection of DEGs and WGCNA-derived T cell proliferation-related module genes (TRMGs). Biomarkers were selected and validated through machine learning algorithms and expression levels. Moreover, a nomogram for predicting AS risk was developed based on the biomarkers. Enrichment analysis was employed to examine relevant pathways, while immune cell infiltration analysis was conducted to investigate the connection between immune cells and biomarkers. Finally, the construction of molecular regulatory and compound prediction networks, as well as molecular docking and molecular dynamic simulations, further validated the key regulatory roles of biomarkers in AS. RESULTS Overall, two candidate genes (CLU and GUCY1B3) were determined to be potentially associated with the progression of AS. The nomogram constructed from these biomarkers showed good performance in predicting AS risk. Reactome rRNA processing and reactome translation were notably enriched in the pathways related to two biomarkers. CLU and GUCY1B3 showed positive computational corrections with activated dendritic cells, suggesting a possible involvement of these immune processes in AS. Moreover, 50 miRNAs (like hsa-miR-4504) and 68 transcription factors (TFs) (like MYB and TAL1) were found to have relationships with biomarkers. Importantly, the results and molecular dynamics simulation analyses predicted potential binding interactions between these candidate genes and bisphenol A, with GUCY1B3 demonstrating relatively greater binding stability. CONCLUSIONS The bioinformatics findings suggested that CLU and GUCY1B3 may serve as candidate biomarkers warranting further experimental investigation in AS associated with T cell proliferation.

Wen Xiong, Xiang Long, Feng Lu et al. · 0 citations
Open access Aug 2026

Machine learning–driven identification and experimental validation of key biomarkers in the bile acid metabolic pathway associated with ulcerative colitis

Bile acids are shown to participate in inflammatory responses. This study was designed to investigate the functions of bile acid metabolism-associated genes (BAMGs) in ulcerative colitis (UC), identify the potential biomarkers based on eleven machine learning algorithms. Seven independent UC transcriptomic datasets were retrieved from the GEO database. Differentially expressed genes, weighted gene co-expression network analysis (WGCNA), and multiple machine learning algorithms were integrated to identify key BAMGs. Subsequently, enrichment analysis, immune cell analysis and single cell analysis were performed to explore the biological functions and immunological characteristics. The dextran sulfate sodium (DSS) induced colitis model in mice was then established and validated the results through western blot and immunohistochemical (IHC) analysis. In addition, peripheral blood samples were collected from UC patients for the detection of feature gene expression by quantitative real-time PCR (RT-qPCR). Through integrative analysis, three feature BAMGs ( CH25H , SLC23A1 and PHYH ) were identified. Unsupervised clustering based on the three-gene signature stratified UC patients into two distinct subgroups exhibiting divergent immune status. In DSS-treated mice, western blot and IHC confirmed significantly reduced SLC23A1 and PHYH protein levels and elevated CH25H protein expression in colonic tissues. RT-qPCR analysis of PBMCs from UC patients showed consistent gene expression. Immune cell analysis showed obvious association between the key BAMGs and inflammatory cells including naïve B cells, neutrophils, monocytes, CD8 T cells, and macrophages. Single-cell analysis revealed that the three feature genes were differentially expressed across T- and B-cell subsets, indicating their potential involvement in UC. This study identified a novel of BAMGs and preliminary revealed their interaction with immune cells in the development of UC. Downregulation of SLC23A1 and PHYH and upregulation of CH25H may contribute to UC pathogenesis and represent potential biomarkers.

Yuqing Wu, Danyan Gu, Jin Liu et al. · 0 citations
Open access Aug 2026

Identification and Experimental Validation of Key Biomarkers for Rheumatoid Arthritis Based on Bioinformatics Analysis and Machine Learning

Objective Rheumatoid arthritis (RA) is a chronic autoimmune joint disease driven by dysregulated immune cells and transcription factors. Despite known molecular alterations, systematic screening of key biomarkers and their link to the immune microenvironment remains lacking, particularly regarding extensive multi-algorithm cross-validation across multiple independent cohorts. This study employs bioinformatics and machine learning to identify potential RA biomarkers, aiming to support diagnosis and targeted therapy. Methods Multiple RA-related transcriptomic datasets derived from synovial tissue were integrated from the GEO database to screen differentially expressed genes (DEGs). Weighted gene co-expression network analysis (WGCNA) was performed to identify RA-associated modules. A total of 107 parameter and algorithm permutations from 11 distinct machine learning approaches were employed to screen key feature genes. The optimal model was selected based on average AUC values across training and validation sets, and the final three genes were identified by integrating individual diagnostic performance, biological relevance, and experimental validation. Diagnostic performance was evaluated using receiver operating characteristic (ROC) curves, while decision curve analysis (DCA) and confusion matrices were applied to validate the clinical net benefit and classification performance of the model. Immune infiltration analysis was used to assess alterations in immune cell composition within the RA microenvironment. Collagen-induced arthritis (CIA) was used to establish rat models of RA in Sprague-Dawley (SD) rats with a modest sample size (control n = 4, CIA n = 6). Ankle joint tissues were harvested for pathological examination, and the key targets were further validated by immunohistochemistry, serving as a preliminary biological corroboration of the computational findings. Results Through differential expression analysis and WGCNA, a set of RA-related candidate genes was identified. After combined screening using 107 parameter and algorithm permutations and ROC curve evaluation, FOSL2, JUN, and EGR1 were ultimately determined as potential biomarkers for RA. These genes demonstrated good individual diagnostic accuracy (AUC > 0.8). Immune infiltration analysis consistently revealed significant enrichment of mast cells in the RA microenvironment. The CIA model rats were successfully established, and immunohistochemistry results showed significantly high expression of FOSL2, JUN, and EGR1 in the synovial tissue. Conclusion This study identifies FOSL2, JUN, and EGR1 as potential markers for RA, supporting their potential roles in RA pathogenesis and clinical application.

Yuxin Han, Pengrui Wang, Yifei Wang et al. · 0 citations