Skip to content
Open access

A biobank-scale method for learning modulators of gene-environment interaction underlying human complex traits from multiple environmental exposures.

Aug 2026 · Genome Research · pp. gr.282105.126 · 0 citations
Medicine

TL;DR

Efficient multi-eNvironmental Gene-environment Interaction iNference Estimator (ENGINE), a supervised variance-component framework that learns an embedding that combines multiple environmental exposures while jointly estimating additive, G×E, and heteroskedastic noise components, making polygenic G×E analysis tractable when both the number of individuals and the number of SNPs reach the millions.

Abstract

It is increasingly recognized that genetic effects on complex traits and diseases are shaped by environmental context. Biobanks that measure diverse environmental exposures alongside genotypes and phenotypes at scale enable systematic study of gene-environment (G×E) interactions. Existing approaches, however, are limited in their ability to accurately model polygenic G×E involving many exposures across genome-wide genetic variants. It is unclear which exposure combinations are relevant for a given trait while distinguishing true interactions from environment-dependent heteroskedastic noise. To address these challenges, we develop Efficient multi-eNvironmental Gene-environment Interaction iNference Estimator (ENGINE), a supervised variance-component framework that learns an embedding that combines multiple environmental exposures while jointly estimating additive, G×E, and heteroskedastic noise components. To enable biobank-scale inference, ENGINE makes a single pass over the genotype matrix to cache genotype-dependent summaries, then assembles normal-equation components and gradients at each iteration. In simulations, ENGINE controls type I error rates, achieves high power, and accurately recovers the environmental embedding while remaining efficient at biobank-scale. It is roughly five-fold faster than the state-of-the-art method at biobank scale, making polygenic G×E analysis tractable when both the number of individuals and the number of SNPs reach the millions. Applied to five complex traits paired with lifestyle exposures in N = 291,273 unrelated white British individuals and M = 454,207 common SNPs (MAF>0.01) from the UK Biobank, ENGINE recovered G×E variance that was on average 1.4-fold larger than that captured by a single exposure and 5.5-fold larger than that captured by the first principal component of the exposures.

Read PDF

Similar papers

Open access Sep 2026

GeneSIS: enhancing transferability of polygenic scores with variant-level gene-by-sex interaction effects

Advancing precision medicine requires accurate prediction of disease liability across populations and contexts. A major challenge is the limited transferability of polygenic scores (PGS) across genetic ancestry groups. We present GeneSIS (GENE and Sex Interaction Score), a supervised statistical learning framework for...

Yosuke Tanigawa, M. Kellis · 0 citations
Open access Aug 2026

Locus-specific gene-context interactions improve polygenic prediction

Polygenic scores (PGS) are a primary output of large-scale genetic studies and are being deployed in clinical and non-clinical settings. However, current PGS assume simple additive models that ignore context-specific genetic effects, which likely reduce their accuracy and robustness. To address this, we developed PGSC,...

Renée Fonseca, C. Caggiano, Manuela Costantino et al. · 1 citation
Open access Aug 2026

Deviations from genetic additivity driven by rare variants at biobank scale

Additive genetic models are the default for genome-wide association studies, but deviations from additivity are crucial for understanding disease mechanisms and therapeutic responses. Yet existing methods for testing nonadditivity are computationally infeasible for large-scale analysis or rely on Hardy-Weinberg assumpt...

Frederik H. Lassen, S. S. Venkatesh, N. Baya et al. · 0 citations
Aug 2026

Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

A novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables that significantly improves classification accuracy and model interpretability compared to using the original feature space alone is propose...

J. Byun, Dheeman Saha, Younghun Han et al. · 0 citations
Open access Sep 2026

Multivariate GWAS framework for family trios with parental phenotypes to control dynastic effects

Effect size estimates in genome-wide association studies (GWAS) and Mendelian randomization (MR) based on unrelated individuals are often confounded by dynastic effects (DE). Existing family-trio-based methods can mitigate this bias, but they typically discard parental phenotypes and cannot jointly model multiple corre...

Shun Zhang, Jia-Hao Mai, Qiao-Feng Zhong et al. · 0 citations
Open access Aug 2026

Additive Multilocus Burden and Epistatic Interactions Improves Genetic Risk Predictions for Complex Diseases

An extended PRS (ePRS) framework for type 2 diabetes (T2D) that incorporates locus-by-locus non-additive effects beyond those captured by additive single-locus PRS or linkage disequilibrium (LD) tagging is developed, with potential for clinical use pending prospective validation.

K. Multerer, P. Atkinson, L. Woods et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.