Disease-associated variants reside frequently in noncoding cis-regulatory elements (CREs), yet their functional consequences remain poorly understood. We performed a large-scale lentiMPRA in human excitatory neurons, quantifying the impact of >46,000 naturally occurring variants across >27,000 candidate CREs near 524 disease-associated genes. These data improved regulatory variant effect predictions beyond state-of-the-art models. Significant allelic effects occurred at comparable rates across common, rare, and singleton variants, demonstrating that, within MPRA-measurable effects, population frequency carries limited information about per-variant regulatory impact. Variant effect detectability and magnitude were governed primarily by baseline activity of the enclosing regulatory element and local sequence context. Regulatory effects were distributed across numerous transcription factors rather than concentrated in master regulators, consistent with a combinatorial enhancer architecture. We establish a large-scale functional variant catalog and provide a complementary benchmark and resource for developing and evaluating models of noncoding regulatory variation.
Kilian Salomon, Chengyu Deng, P. Dash et al.· bioRxiv· 0 citations
Structural variants are a major source of genomic variation and contribute to human disease and evolution through diverse mechanisms, yet their functional interpretation remains challenging. We present CADD-SV v2.0, an improved machine learning framework for scoring SV deleteriousness that expands on the original CADD-SV implementation. This version introduces a unified Random Forest model trained on an expanded set of proxy-neutral and proxy-deleterious variants drawn from human and non-human primate genomes. The model integrates updated genomic annotations, including constraint metrics, regulatory elements, and chromatin architecture features. It scores Deletions, Insertions, Duplications and Inversions based on a single scoring framework that uses both the variant and its flanking regions. To complement this framework, we also explore sequence-based annotations derived from SegmentNT, a deep learning model that provides functional predictions from DNA sequence at nucleotide resolution. Our analysis evaluated whether sequence-derived functional signals can provide additional information for SV prioritization and whether additional models with these features alone or in combination with previous coordinate-based annotations can be used.\ CADD-SV v2.0 outperforms its previous version and other tools in prioritizing deleterious variants across major SV types, including some previously unsupported, and substantially improves the computational workflow, increasing predictive power for genome-wide SV interpretation.