Gene Swap Randomization (GSR) is introduced, an empirical framework that preserves pathway size and multi- pathway gene membership in the null model, enabling explicit adjustment for pathway overlap and improves biological insight by distinguishing pathway-specific genetic signal from enrichment driven by pathway overlap.
Abstract
Genome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, and extensive efforts are underway to translate these variant-level signals into biological mechanism. A widely applied approach is pathway enrichment analysis, which tests whether genetic associations concentrate within biological pathways beyond background polygenic expectations. However, pathway databases contain extensive sharing of genes across pathways (pathway overlap), an underappreciated source of bias that creates structural dependencies in enrichment statistics and obscures pathway-specific genetic signal. Moreover, the degree of pathway overlap is increasing as pathway resources expand. Here, we introduce Gene Swap Randomization (GSR), an empirical framework that preserves pathway size and multi- pathway gene membership in the null model, enabling explicit adjustment for pathway overlap. Applying GSR to enrichment results from the Molecular Signatures Database (MSigDB) across twelve complex traits and four pathway analysis approaches (MAGMA, PascalX, GSA-MiXeR, and PRSet), we show that pathway overlap can produce enrichment under polygenicity even in the absence of pathway-specific biology. GSR improves prioritization of biologically relevant pathways supported by independent gene-disease associations (Open Targets, Malacards), regulatory interactions (DoRothEA), and tissue-specific expression patterns (GTEx). GSR improves concordance with external benchmarks in 60.8% of comparisons overall and 79.3% disease- association benchmarks, corresponding to improvement in 10 of 16 aggregated method-validation framework comparisons. We demonstrate that pathway overlap is a key source of bias in GWAS pathway enrichment, that pathway-specific disease enrichment persists after conditioning on overlap, and that GSR improves biological insight by distinguishing pathway-specific genetic signal from enrichment driven by pathway overlap.
This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets.
Christina Y. Feng, P. Sugier, Nan Zou et al.· Statistics in Medicine· 0 citations
The SNPannotator package is introduced, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings and provide a practical framework for efficiently deriving biologically meaningful insights from GWAS data.
Alireza Ani, I. Nolte, Zoha Kamali et al.· Bioinformatics· 1 citation
An approach comprising locus-specific stratification (LSS) and gene regulatory prioritisation score (GRPS), which uniquely considers multi-signals during fine-mapping and target gene identification, to address issues arising from multi-signals in complex diseases.
Jing Zhang, Qiao-Qiao Liu, Ye Zhu et al.· Nature Communications· 0 citations
To support reproducible best practice, a simple command set is provided for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.
Brian Timoney, P. Guasoni, Komal Zade et al.· bioRxiv· 0 citations
The Multi-Omics Causal Resource Database (MOCR-DB) is an interactive platform that integrates large-scale GWAS summary statistics from UK Biobank, FinnGen, and the COVID-19 Host Genetics Initiative with molecular quantitative trait locus (QTL) datasets to provide a unified framework for genetic correlation, causal infe...
Using summary statistics from genome-wide association studies for 37 cancer types, extensive genome-wide and local genetic correlations among cancers are identified and 33 novel functional genes harboring previously unreported cancer risk variants are identified.
Xiao-Hong Wu, Yu-Qing Yan, Wen Cao et al.· Briefings in Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.