Interpretable multimodal learning for integrating neuroimaging and genetic data in Alzheimer’s disease
Early detection of Alzheimer's disease (AD) requires models that combine brain structure changes with genetic risk, but existing methods struggle to align these different data types. We present R-GenIMA, an interpretable multimodal large language model that pairs a region-of-interest vision transformer with genetic prompting to jointly analyze structural MRI and single nucleotide polymorphisms (SNPs). Each brain region becomes a visual token and SNP profiles are encoded as structured text, letting the model link regional atrophy to genetic factors through cross-modal attention. Tested on the ADNI cohort, R-GenIMA performs well in classifying four groups: normal cognition, subjective memory concerns, mild cognitive impairment, and AD. Beyond accuracy, it produces biologically meaningful explanations, identifying stage-specific brain regions and genes. The model consistently highlighted known AD risk genes (APOE, BIN1, CLU, RBFOX1) and revealed stage-specific patterns: striatal involvement in subjective decline, frontotemporal changes in early impairment, and broad network disruption in AD. These results show that interpretable multimodal AI can integrate imaging and genetics to reveal disease mechanisms, providing a foundation for clinical tools that enable earlier risk assessment and inform precision treatment in Alzheimer's disease.