Skip to content
Review Open access

From linear models to deep learning: statistical advances in genomic selection for animal breeding

Aug 2026 · Journal of Animal Science and Biotechnology · Vol 17 · 0 citations · 160 references
Medicine

TL;DR

A comprehensive overview of GS methodologies is provided, first covering the statistical foundations of linear mixed and Bayesian models, and then modern ML and DL approaches, to provide practical guidance for optimizing genomic evaluation strategies in the era of big data breeding.

Abstract

Genomic selection (GS) has revolutionized animal breeding by accelerating genetic gain through genome-wide marker data. As genotyping technologies advance and data dimensionality grows, the statistical foundations of GS are shifting from classical linear frameworks, which assume additive genetic effects, toward advanced computational models that capture complex nonlinear relationships in genomic data. The commercialization of genotyping arrays for livestock and poultry, coupled with steadily declining sequencing costs, has led to an exponential increase in the availability of high-density genomic data. However, challenges persist, including scenarios where the number of genetic markers far exceeds the number of samples with phenotypic data, and the growing complexity of relationships within genomic data. These issues significantly limit the applicability of traditional evaluation models. In parallel, computational power has increased significantly over the last few decades, providing the capacity necessary for highly complex analyses. While traditional linear mixed models provide a robust framework for incorporating biological priors and modeling additive genetic effects, they often rely on simplified assumptions. In contrast, machine learning (ML) and deep learning (DL) algorithms, which do not rely on predefined parametric models, are well-suited to capturing complex nonlinear relationships and offer effective solutions to the aforementioned challenges. This review provides a comprehensive overview of GS methodologies. We first cover the statistical foundations of linear mixed and Bayesian models, and then survey modern ML and DL approaches. We discuss the assumptions, advantages, and limitations of each method and, by comparing the computational efficiency and predictive accuracy of these diverse approaches, aim to provide practical guidance for optimizing genomic evaluation strategies in the era of big data breeding.

Read PDF

Similar papers

Review Open access Jul 2026

Genomic Selection in Animal Breeding: Principles, Applications, and Future Perspectives

Genomic selection (GS) has transformed modern animal breeding by enabling the prediction of genetic merit using dense genome-wide molecular markers rather than relying solely on pedigree and phenotypic information. Since its conceptual introduction in 2001, GS has become a cornerstone of genetic improvement programs in livestock species, particularly dairy cattle, and has subsequently expanded to beef cattle, sheep, goats, swine, poultry, and aquaculture. The integration of high-density single nucleotide polymorphism (SNP) genotyping with advanced statistical prediction models has substantially increased the accuracy of breeding value estimation, shortened generation intervals, and accelerated rates of genetic gain. Compared with conventional best linear unbiased prediction (BLUP) and marker-assisted selection (MAS), genomic selection captures the combined effects of thousands of loci distributed throughout the genome, making it highly effective for complex quantitative traits governed by many genes of small effect. Recent advances, including single-step genomic BLUP, Bayesian prediction methods, whole-genome sequence analysis, functional genomics, multi-omics integration, artificial intelligence, and precision livestock farming technologies, have further enhanced the scope and efficiency of genomic prediction. These innovations are facilitating simultaneous improvement in productivity, fertility, feed efficiency, disease resistance, animal welfare, and environmental sustainability. Moreover, genomic information is increasingly being integrated with genome editing technologies such as CRISPR to support precision breeding strategies. This review summarizes the historical evolution, fundamental principles, methodological developments, and practical applications of genomic selection in livestock breeding while highlighting emerging innovations and future research directions that are expected to shape next-generation animal improvement programs.

S. Pathak, Amit Kumar, Vaishali Sah · 0 citations
Review Open access Jul 2026

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

This review describes the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS and presents 30 methods designed to leverage AI in GWAS.

S. D’Antona, Mawada Elmagboul Abdalla Abakar, Daniele Ramazzotti et al. · 0 citations
Open access Aug 2026

Biology‐informed neural networks learn nonlinear representations from omics data to improve genomic prediction and biological discovery

SUMMARY Traditional genotype‐to‐phenotype models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes can achieve higher predictive fit, but remain impractical since such data are unavailable at deployment or design time. Biology‐informed neural networks (BINNs) overcome this limitation by encoding pathway‐level inductive biases and leveraging multi‐omics data only during training, while using genotype data alone during inference. Here, we extend BINNs for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge. By directly embedding omics‐derived priors, BINN outperforms conventional models in low‐data (n < p) regimes and enables sensitivity analyses that expose biologically meaningful traits. Applied to maize gene expression and multi‐environment field trial data, BINN improves rank correlation accuracy within and across most subpopulations under sparse data conditions and nonlinearly identifies genes that GWAS/transcriptome‐wide association studies may fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show that highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests that BINNs learn biologically relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework for improved genomic prediction accuracy and biological discovery that can guide genomic selection, candidate gene selection, pathway enrichment, and gene‐editing prioritization.

Katiana Kontolati, R. J. Gladstone, Ian Davis et al. · 0 citations
Open access Jul 2026

Using Deep Learning Models as a Genetic Architecture for the Simulation of Breeding Schemes.

In several simulation studies, long-term selection led to the rapid depletion of genetic variance. These outcomes differ from real-life observations that we aim to replicate, thereby highlighting a fundamental limitation of current classical quantitative genetic simulation models. Deep learning (DL) models have demonstrated promising results in capturing complex interactions essential for maintaining genetic variance; thus, we hypothesize that DL-based genetic simulation models may preserve more genetic variance than classical models, because the biological pathways underlying complex traits exhibit interactions that classical models ignore. The primary objective of this study was to introduce alternative DL-based genetic simulation models and compare them with classical genetic simulation models in terms of their retention of additive genetic variance under truncation selection in a simulated full-sib pig breeding scheme using real haplotypes as founders. After 20 generations of directional truncation selection, the classical models (A, ADAA, and ADAAADDD) retained between 55% and 64% of their initial additive genetic variance. In contrast, while the DL_simple model lost all its additive variance, the DL medium retained 92-98% of its additive variance, and the DL_complex model's initial additive variance increased by 296-314%. This paper introduces DL-based genetic simulation models and concludes that their ability to retain additive genetic variance depends on the models' architectural complexity. When sufficiently complex, DL-based models exhibit greater retention of additive genetic variance because they intrinsically capture epistatic interactions that are converted into additive variance, as selection progresses. Thus, affirming the role of non-additive genetic effects in maintaining long-term genetic variation.

Olumide Onabanjo, Theo H. E. Meuwissen, H. M. Gjøen et al. · 0 citations
Open access Jul 2026

Reversible jumps for the joint modeling of the dimension of pseudo-functional models in chromosomal windows in genome-wide selection.

BACKGROUND With the advance of large-scale genotyping and the high availability of single nucleotide polymorphism (SNP) markers, genomic selection (GS) has become fundamental for plant and animal breeding. However, this technique faces three classical problems: multicollinearity, high marker dimensionality, and high computational cost. To address these problems, mixed models or Bayesian inference methods are generally used. More recently, bin-based genomic window approaches have been widely used, as they can group redundant markers, reducing dimensionality and mitigating the effects of multicollinearity. However, the performance of these methods depends, among other factors, on the number, size, and location of the bins. In general, existing approaches group markers based on linkage disequilibrium (LD) patterns or use fixed-size genome partitions. RESULTS A transdimensional method was proposed to jointly infer the dimension, location, and composition of genomic compartments, as well as the parameters associated with the model, through reversible jump Markov chain Monte Carlo (RJ-MCMC) sampling. The method was evaluated in simulated F2 and F10 populations under oligogenic and polygenic architectures, as well as in human SNP data from the HapMap project with simulated phenotypes and real eucalyptus data. In the simulated datasets, the method achieved the best results under oligogenic architectures (F2 and F10). In more complex polygenic scenarios and real data, performance was satisfactory and comparable to that of rr-BLUP, Bayes B, and Hu, considering mean squared error (MSE), mean absolute error (MAE), and coefficient of determination (R2). CONCLUSIONS The proposed method incorporates model complexity as an inferential component, allowing the automatic estimation of the number, location, and composition of genomic compartments jointly with the model effects. The method demonstrated robust predictive performance, in addition to improving the ability to identify causal regions, which may represent a promising methodological alternative for both genomic selection and QTL mapping.

E. G. Moura, Carlos Pereira da Silva, Luciano Antonio de Oliveira et al. · 0 citations
Open access Jul 2026

Assessing the Role of Marker Density and Minor Allele Frequency on Machine Learning–Driven Genomic Selection Accuracy in Grapevine

Although grapevine (Vitis spp.) is among the oldest and most economically significant fruit species globally, its genetic improvement faces major bottlenecks due to long juvenile periods and extended cycles for phenotypic evaluation. In this context, genomic selection (GS) has emerged as an effective alternative to traditional selection, offering a robust framework to optimize breeding programs by significantly reducing generation intervals while enhancing predictive accuracy (PA) in early generations and expected genetic gains (EGGs). Nevertheless, factors such as minor allele frequency (MAF) and population size can significantly affect predictive models, even to the point of making their use unfeasible in breeding programs. In this context, this study evaluated the effect of data dimensionality reduction on GS accuracy by selecting single-nucleotide polymorphisms (SNPs) based on MAF thresholds. The experimental design tested the predictive capacities of four machine learning (ML) algorithms (ElasticNet, K-Neighbors, Support Vector Machine Regression, and XGBoost) alongside the conventional Genomic Best Linear Unbiased Prediction (gBLUP) model. These were validated using three SNP datasets (11,115, 9,494, and 6,100 markers) filtered by MAF levels of 0.05, 0.1, and 0.2 across six genetic traits, and EGGs were compared between conventional breeding and GS via the breeders’ equation. The results revealed that the ML models exhibited remarkable stability, with no significant differences in PA across the different MAF-based SNP densities, except for berry length, which showed a substantial difference with XGBoost at an MAF of 0.2. Conversely, gBLUP demonstrated high sensitivity to dimensionality reduction, with its performance significantly impacted by MAF filtering across all the traits. These results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction. Additionally, compared with chemical traits, morphological traits generally had greater predictive ability. Furthermore, every GS model provided estimated genetic gains superior to traditional breeding, with improvements ranging from an 8.90-fold increase in berry length to a 2.86-fold increase in total soluble solids, confirming that GS integration is promising for enhancing breeding efficiency in grapevines.

F. R. Francisco, Geovani Luciano de Oliveira, Guilherme Francio Niederauer et al. · 0 citations