Over 99.9% of known protein sequences lack experimentally validated functional annotations. We present ProtJEPA, a multimodal Joint-Embedding Predictive Architecture that trains a sequence-only student encoder to predict joint embeddings spanning ten biological modalities—sequence, structure, knowledge graph, protein interactions, literature, localization, tissue expression, GO function, anatomy, and disorder—requiring only sequence at inference. The key innovation is target whitening, which eliminates severe anisotropy in joint targets (mean cosine 0.984 to 0.086) and prevents representation collapse without covariance regularization. On 1,828 held-out dark proteins with zero primary Pfam family overlap with training, ProtJEPA achieves 58.07% Hit@10 on zero-shot GO retrieval (+2.80 pp, p = 0.020), 69.99% enzyme class accuracy (+9.64 pp, p < 0.001), and +11.87 pp subcellular localization at 1% labels (p < 0.001). Under realistic dark-protein deployment conditions where relational modalities are unavailable, ProtJEPA significantly outperforms naive concatenation of remaining modalities. Cross-domain evaluations on drug–target interaction and disorder prediction confirm transfer beyond training modalities, with the T1-only < ESMC < ProtJEPA ordering replicated across six independent tasks. Ablations establish that Phase 1 aggregator pretraining and target whitening are each independently load-bearing.
Vaibhava Lakshmi Ravideshik, Jinha Kim, M. Kellis· bioRxiv· 0 citations
Deciphering the regulatory consequences of sequence divergence across human evolution is essential to understanding the molecular basis of human-specific traits and disease. Although millions of derived alleles distinguish humans from great apes, only a small fraction are likely to influence human-specific traits. Previous studies have focused on regions of elevated sequence divergence, assuming that rapid evolution reflects functional adaptation, yet individual high-impact regulatory mutations evade such scans. Here, we apply sequence-to-function deep learning to predict chromatin accessibility across modern human, archaic hominin, and great ape personalized genomes, identifying lineage-specific cis-regulatory elements (linCREs) across diverse cellular contexts. Compared to conserved elements, linCREs are shorter, less pleiotropic, less conserved, and enriched in neurodevelopmental pathways. Many linCREs occur in regions with limited sequence divergence that acceleration-based approaches would overlook. We validate lineage-specific enhancer activity through luciferase reporter assays and demonstrate that a single motif-generating derived allele nominated by model interpretability tools drives a hominin-specific neurodevelopmental enhancer.
Riley J. Mangan, Nikitha Thoduguli, Dimitar Ivanov et al.· bioRxiv· 0 citations
An encyclopedia of enhancer–gene regulatory interactions in the human genome is built, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes, and improving analyses linking noncoding variants to target genes and cell types for common, complex diseases.
A. Gschwind, Kristy S. Mualim, Alireza Karbalayghareh et al.· Nature· 5 citations
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
Manoj M Wagle, Yongheng Wang, Soham Samanta et al.· bioRxiv· 0 citations