Skip to content

Author

Ethan Pickering

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Biology‐informed neural networks learn nonlinear representations from omics data to improve genomic prediction and biological discovery

SUMMARY Traditional genotype‐to‐phenotype models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes can achieve higher predictive fit, but remain impractical since such data are unavailable at deployment or design time. Biology‐informed neural networks (BINNs) overcome this limitation by encoding pathway‐level inductive biases and leveraging multi‐omics data only during training, while using genotype data alone during inference. Here, we extend BINNs for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge. By directly embedding omics‐derived priors, BINN outperforms conventional models in low‐data (n < p) regimes and enables sensitivity analyses that expose biologically meaningful traits. Applied to maize gene expression and multi‐environment field trial data, BINN improves rank correlation accuracy within and across most subpopulations under sparse data conditions and nonlinearly identifies genes that GWAS/transcriptome‐wide association studies may fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show that highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests that BINNs learn biologically relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework for improved genomic prediction accuracy and biological discovery that can guide genomic selection, candidate gene selection, pathway enrichment, and gene‐editing prioritization.

Katiana Kontolati, R. J. Gladstone, Ian Davis et al. · 0 citations
Open access Jul 2026

CASCADE recovers promoter-associated regulatory motifs from cell-type-resolved DNA language-model attributions

Gene expression is governed by regulatory DNA and their associated trans factors acting in specific cell types, yet the sequences underlying this control remain poorly mapped in plants. Genome-pretrained DNA language models provide a route to interrogate regulatory sequence directly, but their attributions have largely been interpreted using bulk or whole-tissue data, and standard attribution pipelines can preferentially highlight sequences downstream of the transcription start (TSS) site rather than promoter-associated signals. Here, we train a celltype-resolved sequence-to-expression model from a single-cell soybean (Glycine max) atlas by coupling a soybean-adapted Genomic Pre-trained Network (GPN) to a shared sequence encoder with 66 cell-type-specific output heads. Across 38,339 protein-coding genes, the model achieves a mean per-cell-type, across-gene Pearson correlation of 0.683 and, recast as a highversus-low expression classification, reaches an area under the ROC curve of 0.92 to 0.97 across tissues, at or above dedicated plant sequence models. We then introduce ContextAware Significance of Cross-gene Attribution for Discovering Elements (CASCADE), a positionspecific statistical framework for identifying model-derived candidate regulatory elements from in silico saturation mutagenesis. Relative to the pooled null used by TF-MoDISco, CASCADE shifts motif recovery from downstream of the transcription start site toward promoter sequence, with 77% of CASCADE-exclusive motifs, compared with 12% of TF-MoDISco-exclusive motifs, falling within the promoter. Applied across the atlas, CASCADE identifies approximately 1.39 million candidate elements spanning broadly active, tissue-restricted and cell-type-restricted classes. Together, these analyses establish a position-aware approach for extracting promoterassociated regulatory hypotheses from sequence models and generate a cell-type-resolved map of candidate cis-regulatory elements.

Ali Farghadan, Robert J. Schmitz, Scott A. Jackson et al. · 0 citations