Skip to main content
Journal of Animal Science and Biotechnology logoLink to Journal of Animal Science and Biotechnology
. 2026 Aug 7;17:159. doi: 10.1186/s40104-026-01477-w

From linear models to deep learning: statistical advances in genomic selection for animal breeding

Lifei Zhang 1, Mingzhu Zhang 1, Xinle Wang 1, Cancan Chen 1, Shunzhe Wang 1, Yujie Zhao 1, Yanjun Zhang 1, Ruijun Wang 1,✉, Yongbin Liu 1,✉
PMCID: PMC13449407  PMID: 42563180

Abstract

Genomic selection (GS) has revolutionized animal breeding by accelerating genetic gain through genome-wide marker data. As genotyping technologies advance and data dimensionality grows, the statistical foundations of GS are shifting from classical linear frameworks, which assume additive genetic effects, toward advanced computational models that capture complex nonlinear relationships in genomic data. The commercialization of genotyping arrays for livestock and poultry, coupled with steadily declining sequencing costs, has led to an exponential increase in the availability of high-density genomic data. However, challenges persist, including scenarios where the number of genetic markers far exceeds the number of samples with phenotypic data, and the growing complexity of relationships within genomic data. These issues significantly limit the applicability of traditional evaluation models. In parallel, computational power has increased significantly over the last few decades, providing the capacity necessary for highly complex analyses. While traditional linear mixed models provide a robust framework for incorporating biological priors and modeling additive genetic effects, they often rely on simplified assumptions. In contrast, machine learning (ML) and deep learning (DL) algorithms, which do not rely on predefined parametric models, are well-suited to capturing complex nonlinear relationships and offer effective solutions to the aforementioned challenges. This review provides a comprehensive overview of GS methodologies. We first cover the statistical foundations of linear mixed and Bayesian models, and then survey modern ML and DL approaches. We discuss the assumptions, advantages, and limitations of each method and, by comparing the computational efficiency and predictive accuracy of these diverse approaches, aim to provide practical guidance for optimizing genomic evaluation strategies in the era of big data breeding.

Keywords: Animal breeding, Bayesian, Deep learning, GBLUP, GEBV, Genomic selection, Machine learning, Prediction accuracy, ssGBLUP

Introduction

The fundamental objective of animal breeding is to increase the frequency of favorable alleles in populations, thereby enhancing economically important traits. Traditional methods have relied on estimating breeding values from phenotypic records and genetic relationships, primarily through pedigree-based statistical approaches. For decades, the best linear unbiased prediction (BLUP) proposed by Henderson [1], which utilizes the numerator relationship matrix, has served as the cornerstone of genetic evaluation. Advances in sequencing technologies, bioinformatics, and molecular breeding have progressively elucidated the genetic basis of complex traits, enabling the identification of underlying genes or markers. This shift facilitated the integration of marker-assisted selection (MAS), which incorporates trait-associated markers into breeding programs [2]. The advent of high-throughput genotyping technologies and the seminal proposal by Meuwissen et al. [3] in 2001 marked the formal emergence of genomic selection (GS). By replacing expected pedigree relationships with realized genomic relationships, GS directly estimates genomic estimated breeding values (GEBV), substantially reducing generation intervals and improving prediction accuracy, especially for low-heritability or hard-to-measure traits [4, 5]. A crucial factor in the success of GS is the quality and quantity of phenotypic data, particularly within the reference population used to train prediction models. Because phenotypes serve as the essential training labels that allow models to associate genotypes with trait expression, establishing a robust reference population with high-quality records is fundamental. This dependency becomes particularly critical for new traits, where accurate genomic predictions are often bottlenecked by the lack of extensive phenotypic records. Consequently, optimizing phenotyping strategies remains a pivotal challenge in maximizing GS efficiency. Another critical factor influencing the accuracy of genomic predictions is heritability, which quantifies the proportion of phenotypic variation attributable to genetic differences. Traits with higher heritability generally yield more accurate GEBVs, as the genetic signal is stronger relative to environmental noise. Conversely, low-heritability traits require larger and more precisely phenotyped reference populations to achieve comparable prediction accuracy. Understanding heritability also informs breeding program design, guiding decisions regarding reference population size and composition, phenotyping intensity, and trait selection for genomic evaluation. The workflow of GS in animal breeding is shown in Fig. 1.

Fig. 1.

Fig. 1

The workflow of GS in animal breeding. The reference population refers to individuals possessing both high-density genotypes and accurate phenotypes. This population serves as the training set for estimating marker effects. The candidate population consists of young individuals that are genotyped but lack phenotypic records. Their GEBVs are predicted using the trained model. The selected population comprises the top-ranking candidates derived from the candidate population based on GEBVs, which are chosen for breeding and subsequently incorporated into the reference population for model updating. Workflow: following quality control of the genotype data, the reference population is used to train statistical or ML models. These models are applied to the candidate population to calculate GEBVs. Once selected individuals have been phenotyped, the phenotypic data are integrated back into the reference population, creating a closed-loop update mechanism that improves prediction accuracy for the next generation

As GS has become the standard in livestock breeding, its underlying statistical models have diversified to address biological complexities. The evolution began with linear mixed models, such as genomic BLUP (GBLUP) [5] and its extension, single-step GBLUP (ssGBLUP) [6, 7], which efficiently integrate genotyped and non-genotyped individuals. While robust for traits governed primarily by additive genetic effects, these models typically assume normally distributed marker effects, thereby limiting their suitability for traits influenced by major genes. Bayesian regression methods subsequently emerged to accommodate heterogeneous genetic architectures through variable shrinkage of marker effects [3, 8].

More recently, the exponential growth of phenotypic and genomic big data has pushed parametric models to their limits, driving the adoption of machine learning (ML) and deep learning (DL) in quantitative genetics [9, 10]. Unlike traditional approaches requiring explicit genetic assumptions, ML and DL learn complex patterns directly from data, offering the potential to capture non-linear interactions (e.g., dominance and epistasis), though explicitly disentangling these effects remains a challenge in DL applications. Nevertheless, these methods introduce challenges related to interpretability, overfitting, and computational demands. While DL and ML methods often demonstrate superior accuracy, their high computational complexity and extensive training times pose significant barriers to routine application. In contrast, traditional linear models like GBLUP remain the industry standard due to their computational efficiency and robustness, even if they may underutilize complex genetic architectures. Consequently, assessing the cost-effectiveness of advanced algorithms is a pivotal challenge in modern genomic evaluation. A substantial improvement in genetic gain is required to justify the increased computational resources. Beyond computational cost, adopting ML/DL approaches requires careful consideration of the trade-off between predictive performance and the risk of overfitting, as well as the need for model interpretability. Furthermore, DL models pose additional challenges related to parameter tuning and stability. Specifically, DL model hyperparameters are often unknown a priori or change dynamically during training, introducing variability in results. To achieve stable and reproducible predictions, strategies such as cross-validation, hyperparameter optimization, early stopping, dropout, and ensemble methods are commonly employed. Careful model validation is therefore essential to ensure that the added complexity of DL translates into reliable GEBVs rather than overfitting or unstable predictions.

This review systematically explores the statistical foundations of GS, tracing its evolution from classical linear and Bayesian models to emerging ML and DL techniques. By comparing their underlying assumptions, performance characteristics, and practical applicability, we aim to provide guidance for optimizing genomic evaluation in modern animal breeding.

Linear statistical model in genomic selection

BLUP model

The BLUP method was first proposed by Henderson [1]. However, its widespread application in breeding was delayed until the mid-1970s when advancements in computing power rendered the required complex calculations feasible. As the cornerstone of traditional genetic evaluation, BLUP operates within the framework of mixed linear models and has proven highly effective for traits with moderate to high heritability. However, for traits with low heritability, the phenotypic variance is largely dominated by environmental noise, making it difficult for BLUP to accurately separate the genetic signal from the residual effects based solely on phenotypic records. Furthermore, for traits that are difficult or expensive to measure, the scarcity of phenotypic data limits the model's statistical power, leading to reduced selection efficiency and accuracy [11]. The fundamental mixed linear model is expressed as:

y=Xb+Za+e, 1

where y is the vector of observations; b and a represent fixed and random effects, respectively; X and Z are their incidence matrices; and e is the vector of random residuals. To estimate these effects, Henderson [1] derived the mixed model equations (MME):

X′R-1XX′R-1ZZ′R-1XZ′R-1Z+G-1b^a^=X′R-1yZ′R-1y, 2

when assuming R=Iσe2 and G=Aσa2 (where A is the additive genetic relationship matrix), the MME can be simplified to:

X′XX′ZZ′XZ′Z+A-1kb^a^=X′yZ′y,k=σe2σa2. 3

GBLUP model

VanRaden [5] proposed the GBLUP method, offering a robust extension of the traditional BLUP framework. This approach utilizes a mixed linear model in which the pedigree relationship matrix is replaced by a genomic relationship matrix (G) constructed from high-density single nucleotide polymorphism (SNP) markers. The GBLUP linear model is expressed as:

y=Xg+Za+e, 4

where y denotes the vector of observed values; X represents the matrix of observed values for the genotype (g) at a single locus; Z is the association matrix for the additive effect of genes, where a denotes the additive effect of genes; e denotes the residual vector. The corresponding MME is as follows:

X′XX′ZZ′XZ′Z+G-1kg^a^=X′yZ′y,k=σe2σa2. 5

The G matrix is constructed from all SNP markers. Assuming a=Zg, then Var(a)=ZZ′σg2, σa2=σg22∑pi(1-pi). The ZZ′ matrix can be normalized to yield G=ZZ′2∑pi(1-pi) and Var(a)=Gσa2, where pi denotes the minimum allele frequency at a locus i, and Z represents the individual genotype matrix. In addition to the method proposed by VanRaden [5], researchers have explored other approaches for constructing matrices. For instance, Goddard et al. [12] proposed a blended G matrix, G=A+b(Gm-A), to incorporate pedigree information, thereby enhancing the flexibility of GBLUP for varied genetic architectures. Yang et al. [13] proposed calculating the G matrix based on weights:

Ajk=1N∑iAijk=1N∑i(xij-2pi)(xik-2pi)2p(1-pi),j≠k1+1N∑ixij2-(1+2pi)xij+2pi22pi(1-pi),j=k. 6

These approaches differ primarily in their assumptions regarding the distribution of marker effects. The standard method proposed by VanRaden [5] assumes that all markers contribute equally to the genetic variance, which is computationally efficient and robust for highly polygenic traits. However, it may yield suboptimal results for traits influenced by major genes. To address data limitations, Goddard et al. [12] introduced a blended matrix incorporating pedigree information, which enhances the stability of variance component estimation, particularly in populations with limited sample sizes or weak genetic connections. In contrast, the weighting strategy suggested by Yang et al. [13] allows for heterogeneous variances across markers. This approach is advantageous for livestock populations where complex traits are often governed by a few loci with large effects, as it assigns higher weights to influential SNPs, thereby improving prediction accuracy compared to the equal-weight assumption. In GS, multi-trait GBLUP models often provide more accurate GEBVs than single-trait models by leveraging genetic correlations [14]. Wang et al. [15] further noted that modifying kinship derivation enables GBLUP to maintain significant computational advantages, especially for complex traits. However, a fundamental limitation of standard GBLUP is its assumption that all markers contribute equally to genetic variance. In reality, genomic landscapes are often dominated by a few large-effect loci, while most markers have minor effects. This discrepancy was addressed by Zhang et al. [16] through the trait-specific marker derived relationship matrix BLUP (TABLUP), which integrates marker effect estimation with matrix construction by assigning trait-specific weights to individual SNPs. Haque et al. [17] demonstrated the substantial potential of iterative weighted GBLUP in predicting complex carcass traits within a cohort of nearly 20,000 Hanwoo cattle. Using real phenotypic and genotypic data, this study highlights the practical applicability of the method in livestock breeding programs. By iteratively reweighting SNP effects to enhance the contribution of influential loci, the weighted GBLUP model achieved a predictive accuracy improvement of up to 8.97% relative to conventional GBLUP. This performance closely approached that of more computationally demanding Bayesian approaches, offering a favorable balance between prediction accuracy and computational efficiency. However, a limitation of this iterative approach is its dependence on the quality of initial variance estimates, and it may require careful monitoring to ensure convergence.

ssGBLUP model

In livestock breeding practice, constraints such as high genotyping costs often result in a significant number of individuals having pedigree and phenotypic records but lacking genotype data. To utilize this diverse information, Misztal et al. [18] proposed a theoretical framework to integrate pedigree, genotype, and phenotype data. Subsequently, Legarra et al. [6] and Christensen et al. [7] independently derived methods to construct a unified kinship matrix (H) by merging the pedigree relationship matrix (A) and the genomic relationship matrix (G). This integration is particularly effective in livestock populations, which are typically characterized by extensive linkage disequilibrium (LD) blocks due to historical bottlenecks and selection. Such an extensive LD structure allows markers to effectively capture the effects of quantitative trait loci (QTLs) even at a distance, ensuring that the measurable genetic effects are well-distributed and captured across the genome within the H matrix framework. This ssGBLUP method enables the simultaneous estimation of breeding values for both genotyped and non-genotyped individuals, significantly expanding the applicability of GBLUP methods [19]. The structural evolution from the A and G matrices to the combined H matrix is systematically illustrated in Fig. 2. The H matrix is constructed as follows [18]:

H=A11-A12A22-1A21+A12A22GA22-1A21A12A22-1GGA22-1A21G, 7

where A11 and A22 represent the pedigree sub-matrices for non-genotyped and genotyped individuals, respectively. To address potential singularity issues that would otherwise make the G matrix non-invertible, a blending approach is typically employed: Gw=(1-w)G+wA22, where w is the weighting factor representing the proportion of polygenic inheritance effects. Given the computational burden of directly inverting the H matrix, Christensen et al. [7] derived a simplified form for its inverse: H-1=A-1+000G-1-A22-1. Following the construction of H-1, the process of solving the MME remains consistent with traditional BLUP:

X′XX′ZZ′XZ′Z+H-1kb^a^=X′yZ′y,k=σe2σa2. 8

Fig. 2.

Fig. 2

Schematic of relationship matrix construction in linear mixed models. The diagram illustrates the integration of independently derived relationship matrices. It depicts the transition from the traditional pedigree-based BLUP (A matrix) and genomic-based GBLUP (G matrix) to the integrated ssGBLUP. By constructing the H matrix, which combines both pedigree and genomic information, the model enables a more comprehensive and accurate estimation of breeding values for both genotyped and non-genotyped individuals

Despite its widespread adoption, ssGBLUP faces challenges regarding the compatibility between the G and A matrices. This incompatibility stems primarily from incomplete pedigree records and discrepancies in allele frequency definitions [20]. Specifically, ssGBLUP assumes that the allele frequencies used to construct the G matrix represent the unselected base population, which is often difficult to define in practical settings. Failure to properly account for the stratification of allele frequencies across different generations or sub-populations results in a fundamental incompatibility between the G and A matrices. This discrepancy introduces substantial bias into the MME, typically manifesting as the inflation or deflation of GEBVs and a biased estimation of genetic variance, thereby compromising the prediction accuracy for both genotyped and non-genotyped individuals. To mitigate this, Christensen [21] introduced parameters γ and s to correct for population bias. Furthermore, to address the lack of pedigree connections in multi-population evaluations, Legarra et al. [22] proposed the metafounders theory. Studies on cross-population genomic prediction have demonstrated that the metafounder approach effectively captures ancestral correlations, yielding predictive accuracies at least as high as standard ssGBLUP models that assume unrelated populations [23, 24]. Beyond the frequentist framework, Fernando et al. [25] introduced single-step Bayesian regression (ssBR). Recent evaluations in buffalo populations suggest that ssBR variants can outperform ssGBLUP in predicting polygenic traits, especially when genotyping resources are limited [26].

To further overcome the limitations of standard ssGBLUP, which assumes equal variance for all markers and may fail to capture major genes embedded in a polygenic background, recent research has shifted toward integrating genome-wide association study (GWAS) results into the single-step framework. Pang et al. [27] introduced the single-step genome-wide association assisted BLUP model, which relaxes the infinitesimal assumption by assigning differential weights to SNPs based on GWAS-derived effect sizes. Their empirical findings, derived from both simulation scenarios and real pig data, demonstrated that the single-step genome-wide association assisted BLUP with pseudo-quantitative trait nucleotides as covariates significantly outperforms both standard ssGBLUP and iterative weighted ssGBLUP. These markers are identified through GWAS and LD pruning. Crucially, this method enhances prediction accuracy not only for traits with large-effect loci but also remains effective for complex, polygenic traits by capturing a broader spectrum of genetic variance through weighted markers. Despite the development of these advanced models, standard ssGBLUP continues to demonstrate robust practical utility. For instance, Mancisidor et al. [28] observed that ssGBLUP improved the accuracy of breeding value predictions for fiber diameter standard deviation in Huacaya alpacas by up to 6.44% compared to traditional BLUP, even with limited genotypic records. Addressing the persistent challenges of pedigree incompleteness and base population definition, Himmelbauer et al. [29] used a simulated cattle population to theoretically demonstrate that the metafounder approach significantly reduces evaluation bias and dispersion compared to traditional unknown-parent groups in scenarios with fragmented pedigrees. By accurately aligning pedigree and genomic relationships, the metafounder framework, enhanced by genotype-based grouping, offers a superior strategy for managing complex population structures in modern genomic evaluations.

RRBLUP model

Traditional least squares methods typically treat marker effects as fixed. However, in genomic prediction, the number of SNP markers (p) often far exceeds the sample size (n), leading to multicollinearity and making the simultaneous estimation of all marker effects via ordinary least squares mathematically intractable. To circumvent this, early approaches relied on selecting only a small subset of significant SNPs, which often led to overfitting and failed to capture the small effects of rare variants. To address these limitations, Whittaker et al. [30] proposed a ridge regression approach, now commonly known as ridge regression BLUP (RRBLUP). Unlike least squares, RRBLUP treats all SNP effects as random variables following a normal distribution with a constant variance. By employing a linear mixed model, RRBLUP estimates the effect of every marker simultaneously and calculates the GEBVs as the sum of these effects. In GS, RRBLUP is categorized as an indirect approach (estimating marker effects first), while GBLUP is a direct approach (estimating individual relationships directly). Although the two methods are mathematically equivalent and yield comparable accuracy under Gaussian assumptions, their computational efficiency differs based on data dimensions. Since p≫n in most modern genomic datasets [31], RRBLUP is generally more computationally demanding and time-consuming than GBLUP.

Recent empirical studies have continued to highlight the utility and limitations of RRBLUP in animal breeding programs. A study on Rhode Island Red chickens genotyped with a 50 K SNP array integrated ML approaches with classical methods and compared their performance with RRBLUP and BayesA across 10 economic traits. While ML models generally outperformed RRBLUP for traits such as body weight and eggshell strength, RRBLUP and BayesA exhibited 2%–58% higher predictive accuracy for egg number traits, highlighting the robustness of RRBLUP for traits with a predominantly additive genetic architecture. Notably, the incorporation of GWAS significant SNPs marked a ‘sweet spot' where the advantage of RRBLUP was most pronounced (up to 58%), whereas standard predictions showed more marginal differences [32]. Similarly, Sahebalam and Gholizadeh [33] evaluated different approaches for estimating the shrinkage parameter in RRBLUP using pig genomic datasets, finding that the allelic frequencies RRBLUP variant delivered comparable or superior predictive accuracy while substantially reducing computational burden. However, such accuracy gains can sometimes be attributed to overfitting to the specific dataset rather than representing a general methodological advance, underscoring the importance of validation in independent populations.

Bayesian model

Bayesian methods typically assume a sparse genetic architecture where only a small proportion of SNPs are linked to QTLs with large effects, while the majority have negligible or no effect. The linear model is defined as:

y=Xb+∑j=1MZjαj+e, 9

where y is the vector of phenotypic observations, b is vector of fixed effects, M is the number of markers, αj is the effect of the j-th marker, X and Zj are the incidence matrices of b and αj, respectively, and e is the vector of random residuals. The residuals follow a normal distribution N(0,Iσe2), where I is the identity matrix and σe2 is the residual variance. Solving the model yields an estimate of the marker effect α, subsequently, the GEBV for individual i is calculated as gEBVi=∑Zijαj, where Zij is the genotype of individual i at locus j, and αj is the estimated effect of locus j.

To enhance estimation efficiency, Mathew et al. [34] applied Markov chain Monte Carlo (MCMC) theory, followed by the introduction of Gibbs sampling to address computational complexity. These developments led to the "Bayesian Alphabet" family of models, widely used for GEBV prediction, including BayesA [3], BayesB [3], BayesC [8], BayesCπ [8], BayesDπ [8], and Bayesian Lasso [35]. To better capture complex genetic architectures and incorporate prior biological information, the "Bayesian Alphabet" has continuously evolved, with many models derived directly from one another. Further advancements led to BayesR [36], which utilizes a mixture of four normal distributions to categorize SNP effects into null, small, medium, and large classes, offering enhanced predictive accuracy for traits with varying genetic architectures. Building upon BayesR, the BayesRC [37] model incorporates prior biological knowledge by classifying markers into distinct functional categories. Most recently, models such as BayesRCO [38] have extended this framework to handle overlapping functional annotations, allowing multi-annotated SNPs to be modeled more accurately [39]. This evolution highlights a clear transition from uniform genomic priors to highly flexible, biologically informed architectures.

BayesA and BayesB models

BayesA, proposed by Meuwissen et al. [3], utilizes a multi-level prior distribution where SNP effects gi follow a normal distribution N(0,σgi2), and the variance σgi2 follows an inverse Chi-square distribution χ-2(ν,S). Meuwissen et al. [3] set the degrees of freedom ν and scale parameter S at 4.012 and 0.002, respectively, while others have explored alternative priors [40, 41]. The resulting posterior distribution of the effect variance σgi2 is also an inverse Chi-square distribution χ-2(ν+ni,gi′g′) [42].

BayesB was developed to better reflect genomic reality by assuming most segments have no effect. It introduces a parameter π representing the probability that a marker has zero effect and uses a mixture distribution for marker variances. While BayesA assumes all SNPs have effects (π = 0), BayesB is more accurate for traits controlled by a few large-effect QTLs [3, 43]. Due to its complex posterior, BayesB requires a combination of Gibbs and Metropolis–Hastings (MH) sampling [3]. The prior distribution assumptions in BayesB are as follows:

σgi2=0withprobabilityπ, 10
σgi2∼χ-2(ν,S)withprobability(1-π). 11

Habier et al. [8] proposed deriving S when ν=4.2. Since Eσak2=νasa2νa-2=σ~a2, it follows that Sa2=σ~a2νa-2νa. Here, σ~a2 is the variance of the additive effect for a randomly sampled locus, which can be related to the additive-genetic variance explained by the SNP, σ~s2 as:

σ~a2=σ~s21-π∑k=1K2pk1-pk, 12

where pk denotes the allele frequency of SNP k, while the three parameters ν, S, and π are 4.2339, 0.0429, and 0.947, respectively [3]. The choice of these prior parameters directly impacts the model's output. By changing the priors from fixed, arbitrary values to empirical values derived from actual genetic variance, several key differences can be expected in the results. First, it alters the estimated variance of individual markers, allowing the model to more accurately distinguish between loci with major effects and background polygenic noise. Second, it allows the model to adapt flexibly to different genetic architectures across various traits. Ultimately, these differences lead to improved prediction accuracy for the GEBVs.

BayesC, BayesCπ, and BayesDπ models

To reduce the subjectivity associated with manually setting π, Habier et al. [8] proposed several optimized methods, including BayesC, BayesCπ, and BayesDπ. The BayesC method utilizes a fixed π and assumes a common variance for all SNPs that possess an effect. In contrast, BayesCπ treats π as an unknown parameter following a uniform prior distribution. Further extending this flexibility, BayesDπ estimates both π and the scale parameter S from the data. Notably, the performance of BayesCπ is highly sensitive to factors such as heritability and marker density, which makes it particularly suitable for analyzing traits with medium-to-high heritability [44].

Bayesian Lasso model

The Bayesian Lasso (BL), proposed by Park and Casella [35], obtains linear regression coefficients by normalizing predictor variables xij and centering the response values yi. It addresses the l1 penalty regression problem by minimizing the following objective function with respect to β={βj}:

∑i=1Nyi-∑jxijβj2+λ∑j=1pβj. 13

This is equivalent to a constrained sum of squares minimization: ∑|βj|≤s. A distinguishing feature of BL is the assumption that marker effect variances follow an exponential distribution, resulting in a Laplace prior for the marker effects. This Laplace distribution is particularly effective for genomic prediction as it assigns higher probabilities to extreme values, whether large or small [45]. To systematically compare the differences among these methods, their prior distribution assumptions regarding the key parameters π, SNP effects, and their variances are summarized in Table 1.

Table 1.

Comparison of prior distribution assumptions in different Bayesian models

Bayesian models Prior distribution of π Prior distribution ofgi Prior distribution ofσgi2
BayesA 0 gi∼N(0,σgi2) σgi2∼χ-2(ν,S)
BayesB Fixed value gi∼N(0,σgi2) σgi2∼χ-2(ν,S)
BayesC Fixed value gi∼N(0,σg2) σg2∼χ-2(ν,S)
BayesCπ U(0, 1) gi∼N(0,σg2) σg2∼χ-2(ν,S)
BayesDπ U(0, 1) gi∼N(0,σg2) σg2∼χ-2(ν,Gamma(1,1))
Bayesian LASSO 0 gi∼N(0,σgi2) σgi2∼Exp(λ2/2)

The primary challenge in applying these Bayesian methods lies in the formulation of appropriate prior distributions for hyperparameters. While Bayesian models typically offer improved prediction accuracy over BLUP-based methods by estimating a larger set of parameters, they impose a significantly higher computational burden [46]. The MCMC process is particularly intensive, requiring tens of thousands of iterations where all marker effects must be re-estimated in each cycle. This sequential nature consumes substantial time, often conflicting with the stringent timeliness requirements of practical animal and plant breeding [47].

To enhance computational efficiency, various optimized versions have been developed, such as fastBayesA [48], fBayesB [49], wBSR [50], emBayesR [51], EBL [52], BayesRS [53], and BayesTA [54]. Despite these advancements, the classical "Bayesian Alphabet" remains the most prevalent in current GS practices. Ultimately, the predictive success of these models depends on the degree to which their underlying prior assumptions align with the actual genetic architecture of the target trait [55]. Recent empirical studies in diverse livestock populations have further substantiated these theoretical expectations. In a comprehensive 2025 study of over 16,000 Chinese Holstein cattle, BayesR emerged as the most robust model for production traits, achieving a peak predictive accuracy of 0.625—outperforming both GBLUP and several SNP-weighted models [56]. This superiority is particularly pronounced for traits with known major-effect QTLs, such as milk fat and protein yield, where Bayesian models more effectively capture the non-infinitesimal genetic architecture. Furthermore, in the genomic evaluation of Alpine Merino sheep, research revealed that while GBLUP is sufficient for low-heritability traits, BayesA and BayesB provide a significant accuracy advantage as trait heritability and marker density increase, particularly when utilizing high-density SNP arrays [57]. Additionally, for complex reproductive traits in Hanwoo cattle, recent applications of BL and BayesR recorded a 2% to 6% improvement in prediction reliability over traditional linear methods [58]. These findings collectively indicate that the Bayesian framework remains a cornerstone of GS, offering an optimal balance of biological interpretability and predictive power across different species and genetic architectures. However, as genomic datasets grow exponentially in size and increasingly incorporate highly non-linear multi-omics layers, traditional parametric models face significant computational and structural bottlenecks. This escalating data complexity has catalyzed the exploration of non-parametric ML approaches, which offer novel solutions to these emerging challenges.

Machine learning approaches in genomic selection

Besides parametric models based on BLUP and Bayesian theory, GS increasingly incorporates non-parametric ML approaches. Unlike traditional parametric and Bayesian linear regression, ML models do not necessitate explicit assumptions regarding the underlying data distribution. These models learn complex patterns directly from training populations to predict phenotypes in target populations with similar genetic architectures [59]. ML methods are primarily classified into supervised and unsupervised learning [60], with mainstream prediction techniques in GS encompassing neural networks, ensemble methods, and kernel-based approaches [61]. As summarized in Fig. 3, these non-parametric models offer distinct advantages in handling high-dimensional and non-linear biological data. However, their application is not without challenges. ML models may struggle with the adequate representation of rare alleles due to data sparsity, and they carry an inherent risk of overfitting during the learning process if regularization is not strictly applied. Furthermore, the "black-box" nature of some ML algorithms can obscure the biological interpretability of the genotype-to-phenotype prediction mechanisms. These aspects will be explored in detail in this chapter.

Fig. 3.

Fig. 3

Classification of the most popular parametric and non-parametric ML regression models. This diagram provides a comprehensive taxonomy of genomic prediction models utilized in animal biotechnology. At the center, models are bifurcated into parametric and non-parametric approaches. The upper hemisphere details established statistical frameworks, including classical methods and Bayesian methods. The lower hemisphere illustrates the expanding landscape of ML, featuring kernel methods, ensemble methods, and DL architectures. This categorization serves as a roadmap for selecting optimal computational tools to enhance the accuracy of GEBV across diverse livestock populations

Decision tree

The decision tree (DT) is a widely used supervised learning algorithm primarily employed for classification and regression tasks based on recursive partitioning [62]. First introduced by Breiman et al. [63], this method partitions large heterogeneous datasets into multiple smaller homogeneous subsets to form a hierarchical structure. In this tree structure, each intermediate node corresponds to a feature, each branch represents a value of that feature, and each terminal node stores a specific category or a regression function [64]. Following the development of the ID3 and C4.5 algorithms by Quinlan [65], the classification and regression tree (CART) method emerged as a significant advancement, utilizing the Gini impurity to identify optimal segmentation points [64, 66]. A recent study by Liu et al. [67] in yellow-feathered broilers revealed that tree-based ensemble methods markedly outperformed traditional GBLUP and Bayesian approaches, achieving over 60% improvement in predictive accuracy for carcass traits. This suggests that while individual trees provide a robust foundation for modeling nonlinear relationships, they are often prone to high variance and overfitting. Therefore, their integration into ensemble systems is essential. By aggregating predictions from multiple diverse trees, ensemble methods significantly reduce model variance and improve generalization ability, thereby optimizing performance in complex GS scenarios characterized by intricate genetic architectures.

Ensemble learning

Ensemble learning is a supervised learning approach that constructs and combines multiple weak learners to form a comprehensive strong learner, also known as a multi-classifier system. Common strategies include bagging, boosting, and stacking [68–70]. In GS, ensemble methods often outperform single ML models by mitigating individual model biases. However, they may require more extensive tuning of structural and training hyperparameters (e.g., learning rate, number of estimators, tree depth) to optimize model architecture and prevent overfitting, which differs from the hyperparameter prior specification in Bayesian frameworks [71].

Bagging

Bagging involves training m base learners independently on data subsets generated via random sampling with replacement from the original dataset D, a process known as bootstrap sampling that resamples the observations (rows) of the dataset [68]. The final prediction is obtained through voting or averaging (Fig. 4a). This parallelized approach significantly reduces prediction variance and is robust against data noise [72]. However, it is important to note that when dealing with unbalanced datasets, standard bootstrap sampling can inadvertently generate subsets with skewed class distributions (e.g., lacking minority class samples), potentially introducing bias into the model. Therefore, specialized sampling strategies or weighting schemes are often required in such scenarios to ensure robust performance. Nevertheless, bagging remains sensitive to the randomness inherent in bootstrap sampling, and its overall stability can still be influenced by fluctuations in the training data [73].

Fig. 4.

Fig. 4

Schematic architectures of ensemble learning and neural networks frameworks in genomic prediction. This diagram illustrates the implementation of advanced computational frameworks for GS. a Bagging utilizes parallel weak learners to ensure model robustness through random sampling. b Boosting optimizes prediction accuracy via sequential weight adjustment based on learning errors. c Stacking achieves superior performance by consolidating outputs from multiple base models into a meta-learner. d MLP processes high-dimensional SNP data through hierarchical layers and nonlinear activation functions (sigmoid and ReLU) to capture complex genetic effects

Boosting

Boosting iteratively trains a series of weak learners, where each subsequent model focuses on the samples misclassified or poorly predicted by its predecessor by adjusting instance weights [69, 74] (Fig. 4b). One classic implementation is adaptive boosting (Adaboost) [75], which assigns higher weights to hard-to-classify samples and determines the final output through a weighted decision process. Adaboost.RT is an enhanced version specifically optimized for regression tasks in biological data [76]. Recent studies demonstrate Adaboost's advantages in genomic prediction, particularly in handling complex non-linear interactions within high-dimensional data. It has been shown to effectively improve prediction accuracy for various complex traits compared to traditional linear models [77, 78]. The Adaboost.RT regression model can be written as:

fx=∑a=1Mlog1εafax/∑a=1Mlog1εa, 14

where fx represents the final predicted value; fax denotes the prediction value of the a-th weak learner; εa is the error rate of fax.

Gradient boosting machines (GBM) employ gradient descent error minimization with boosting techniques to improve the overall goal of more accurately estimating the target response by iteratively fitting new models [79]. Boosting is commonly used in ML to combine weak predictor models, or weak learners, into a stronger model. The final predictive model is expressed by the following formula:

y=μ+∑i=1Mv×hiy;X+e, 15

where y represents the phenotypes and X the genotypes. The model starts with an initial prediction μ (the population mean). At each iteration i, rather than fitting y directly, GBM trains a new weak learner hi to fit the negative gradient of the loss function with respect to the prediction of the previous model. The contribution of hi is scaled by a shrinkage factor v (learning rate) and added to the cumulative model. This process is repeated M times, with each iteration incrementally reducing the overall loss, until the final residual e is minimized.

In addition to Adaboost and GBM, currently popular algorithms such as gradient boosting decision trees (GBDT) [80] and extreme gradient boosting (XGBoost) [81] also fall under the category of boosting algorithms. These advanced boosting frameworks typically incorporate regularization terms and second-order gradient information to enhance generalization. However, their performance is highly dependent on the configuration of hyperparameters, making rigorous hyperparameter tuning essential for achieving optimal predictive accuracy across different traits.

Stacking

Stacking employs a hierarchical framework where multiple heterogeneous base learners are trained in the first layer, and their outputs serve as training data for a second-layer meta-model (Fig. 4c) [70]. Unlike bagging or boosting, stacking aims to leverage the diverse strengths of different algorithms to optimize predictive performance.

Random forest

The random forest (RF), an ensemble learning algorithm proposed by Breiman [82], constructs a "forest" by combining the bagging method with multiple DTs. The RF regression model aggregates the predictions from individual base learners to form a robust final output, expressed by the following equation:

fx=1B∑b=1BtbΨby;x, 16

where fx represents the predicted value from the RF regression; the predictor variable tbΨby;x is constructed using the bootstrap sample of data Ψby;x during the b-th iteration of the DT; and B denotes the number of DTs included in the RF. By aggregating multiple weak learners through a voting or averaging mechanism, RF effectively forms a strong learner capable of processing high-dimensional biological datasets [83].

In GS, RF exhibits exceptional robustness and effectively mitigates overfitting, particularly when hyperparameters such as tree depth are appropriately tuned. However, it is important to acknowledge that RF shares the limitations of bagging regarding unbalanced data. The random subsetting process can result in bootstrap samples that lack sufficient representation of minority classes, potentially biasing the model. Therefore, applying class weights or stratified bootstrapping is often recommended to ensure fair learning across all classes. This robustness is further enhanced by the stochastic nature of its construction process, which involves both random sampling and random feature selection. Its predictive performance frequently surpasses other ML algorithms, particularly in handling high-dimensional data and modeling non-linear patterns inherent in genomic datasets [84, 85]. For instance, Wang et al. [77] demonstrated that RF and GBDT provided significant accuracy gains over traditional models for reproductive traits such as litter weight and piglets born alive, noting that precise hyperparameter tuning further boosted performance. Despite its advantages, the computational cost of RF training grows significantly with the increasing depth and number of nodes; however, this challenge is effectively mitigated by the algorithm's high potential for parallelization. Modern RF implementations, particularly when combined with feature selection tools like the Boruta algorithm, can significantly boost computational efficiency. Moreover, substantial training datasets are required to achieve stable and satisfactory predictive performance in complex animal breeding populations [86].

Kernel-based algorithms

Support vector machines

The support vector machines (SVM), a non-parametric supervised learning method proposed by Cortes and Vapnik [87], demonstrate powerful adaptability for complex biological datasets [88, 89]. While SVM performs linear classification on separable data, its true strength lies in processing non-linearly separable data through kernel functions. These functions map the original input space into a high-dimensional feature space where data becomes linearly separable. This "kernel trick" enables the efficient computation of inner products in high-dimensional space directly within the original input space, significantly reducing computational complexity and memory usage [90, 91]. The SVM model is expressed as follows:

fx=β0+hxT, 17

where fx denotes the prediction function, β represents the weight vector, β0 is the intercept term, and hxT denotes the kernel function. The most commonly used kernel functions include: the linear kernel: Kxi,xj=xi′xj, the polynomial kernel: Kxi,xj=γxi′xj+rd, the Gaussian kernel: Kxi,xj=exp-γ‖xi-xj‖2, and the sigmoid kernel: Kxi,xj=tanhγxi′xj+r [92].

In the context of GS, this mechanism is crucial as it transforms genomic prediction from linear to nonlinear models. Compared to traditional linear methods, this approach offers greater flexibility in capturing non-additive genetic effects across the entire genome, thereby enhancing the prediction accuracy of GEBVs. While SVM shares principles with kernel ridge regression (KRR), they differ in their assumptions; specifically, unlike SVM, KRR typically assumes that only certain feature data significantly influence breeding value estimation [93]. Existing research has demonstrated the potential of SVM methods in GS [94]. Despite its potential, SVM faces challenges with large datasets, where solving the quadratic programming problem becomes computationally intensive [95]. Furthermore, direct genotype input can be challenging in populations with low-density SNP data. To address these issues, Long et al. [96] proposed replacing raw genotype data with a genomic relationship matrix as the input kernel. This optimization transforms the feature-based SVM into a relationship-based model, which not only enhances computational efficiency but also helps prevent model overfitting, especially when dealing with low-density SNP data. Recent empirical evidence further underscores the robustness of SVM in livestock breeding. Cross-breed evaluations in beef and dairy cattle have demonstrated that SVM-based regression can achieve superior accuracy compared to GBLUP by effectively capturing complex genetic architectures [56]. In addition, in the context of poultry breeding, a recent study on yellow-feathered broilers has shown that support vector regression (SVR) models can significantly outperform traditional linear methods like GBLUP in predicting specific carcass traits, including half-eviscerated and eviscerated weights [67]. Regarding computational scalability, recent large-scale studies in swine populations validate that utilizing a genomic relationship matrix directly as the kernel input serves as an effective computational strategy. This approach significantly mitigates the computational intensity typically associated with SVM training while maintaining high predictive stability even with low-density genotype data [97].

Reproducing kernel Hilbert space

Reproducing kernel Hilbert space (RKHS) theory offers a powerful framework for statistical learning, primarily addressing three categories of problems. First, the subspace problem in function spaces [98], where mapping the problem space to a higher-dimensional space can make the problem easier to solve when defined precisely on an RKHS subspace. Second, the kernel-based embedding of positive semi-definite functions [99], which allows SVMs and other methods to transform nonlinear problems into linear ones via the kernel trick. A classic example is Parzen's [100] work linking RKHS to stochastic processes via covariance kernels. Third, geometric embeddings [101], where RKHS preserves distance-preserving transformations. These mathematical properties provide the foundation for mapping complex biological data into high-dimensional feature spaces. Building on these theoretical strengths, Gianola et al. [102] introduced RKHS to genomic prediction as a typical semi-parametric method. RKHS offers a more general covariance structure than GBLUP, with pedigree-based BLUP and MAS models being special cases of the RKHS framework [103]. Numerous studies have demonstrated that RKHS outperforms traditional parametric models in genome-wide selection [104–106]. A critical factor in RKHS specification is the choice of the reproducing kernel; multi-kernel models are particularly beneficial for capturing complex genetic architectures and simplifying kernel selection [107].

Other machine learning approaches

Beyond mainstream ML methods, several other approaches have been explored for genomic prediction, albeit less frequently. The k-nearest neighbors (KNN), a non-parametric method often used for classification and regression, estimates phenotypic values (yj^) based on the genomic proximity of the K nearest neighbors:

yj^=1K∑k=1Kyk, 18

where yj^ is the predicted phenotype of individual j; K is the number of nearest neighbors; and yk is the observed phenotype of the k-th nearest neighbor found in the training set. The proximity of neighbors is defined by the genomic distance between individuals. The Euclidean distance dxi,xj between two individuals, i and j, is calculated using the formula shown below:

dxi,xj=∑p=1Pxpi-xpj2, 19

where P is the total number of SNP markers; xpi and xpj represent the genotype codes for the p-th SNP marker of individuals i and j, respectively. The prediction is derived from the K nearest individuals with the smallest distance values. When determining the K value, a guided search method (such as cross-validation) can be employed to test different K values, ultimately selecting the K value that minimizes the mean squared error (MSE) between the actual and predicted values.

In some studies, KNN-based genomic prediction accuracy in dairy cows [108] was lower than that of other classical and ML models. Similarly, naive Bayes (NB), a probabilistic classification algorithm, has also been employed in ML applications. It exhibits high predictive capability, though this varies depending on conditions and environments. The NB classification model has been successfully utilized to identify SNPs associated with mortality in chickens [109, 110].

NB demonstrated superior performance in predicting survival traits in Holstein cattle compared to RF [111], yet showed limitations in horse breed prediction relative to specialized Bayesian methods [112]. Other instance-based algorithms, such as IB1 and IB5, have also been tested but generally fall short of the predictive benchmarks set by Bayesian linear procedures [113]. Beyond these conventional approaches, advanced non-linear models have also been developed to address environmental complexity. The kernelized Bayesian matrix factorization (KBMF) algorithm has proven effective in capturing genotype-by-environment interaction (G × E), outperforming standard BLUP by integrating weather forecasts into breeding value estimations [114]. Originally demonstrated in crops, these kernel-based strategies are gaining traction in animal breeding for traits sensitive to management and climate, such as those in porcine populations [115].

Integration of multi-omics data using machine learning in genomic selection

In addition to the aforementioned ML methods for genomic prediction, the current frontier of genomic prediction lies in the integration of multi-omics data to move beyond purely additive genetic models. It is crucial to acknowledge the "no free lunch" theorem in ML, which posits that no single algorithm universally outperforms others across all possible problems. The efficacy of a method is inherently tied to the specific data structure, the underlying genetic architecture of the trait, and the modeling assumptions. Early breakthroughs like multilayered LASSO pioneered the interactive learning of genetic, transcriptional, and metabolomic data in plants [116]. Similar multi-omics-guided frameworks have since been extended to animal and model organism breeding. These examples, while demonstrating the potential of integrating diverse data layers, also highlight that performance gains are context-dependent. For instance, Ye et al. [117] utilized whole-genome sequencing combined with gene expression data in the Drosophila Genetic Reference Panel. They preselected SNPs via expression QTL (eQTL) mapping of significant genes identified from transcriptome-wide association studies and incorporated these functionally informed SNPs into both GBLUP and genomic feature BLUP models. This approach resulted in substantially improved prediction accuracy for specific complex traits, underscoring the importance of tailoring the analytical approach to the biological question at hand.

Additionally, innovative kernel-based models such as transcriptome BLUP, multi-omics BLUP, multi-omics single-step BLUP, weighted multi-omics single-step BLUP, and fusion similarity BLUP have been introduced [118, 119]. These methods, particularly weighted multi-omics single-step BLUP, optimally integrate intermediate omics layers (e.g., transcriptomics) into the genomic relationship matrix. By capturing regulatory and functional variants that standard GBLUP misses, these fusion similarity approaches have yielded accuracy gains of 3.37%–4.18% in beef cattle across diverse genetic architectures [119]. Collectively, these studies underscore that the strategic incorporation of functional omics data, rather than simple feature expansion, is pivotal for the next generation of GS programs.

Parallel to these matrix-based approaches, Bayesian frameworks have also been actively developed to directly or indirectly integrate other omics layers. Instead of constructing multiple relationship matrices, these methods typically utilize multi-omics data to inform prior distributions. For instance, advanced models like BayesRC and BayesRCO indirectly integrate these layers by classifying SNPs into distinct functional categories based on omics annotations and applying differential shrinkage to biologically active regions. This Bayesian multi-omics integration effectively prioritizes causal variants, thereby enhancing both prediction accuracy and biological interpretability compared to BLUP-based methods.

Deep learning methods in genomic selection

In genomic prediction, it is a universally recognized principle across all statistical and computational frameworks that model performance is highly task-dependent, with no single algorithm offering a universal solution. While classical ML algorithms have been extensively applied to genomic prediction, their capacity to improve accuracy is often constrained by the high dimensionality of genomic data and the computational intensity inherent in training. As a sophisticated evolution of ML, DL has emerged as a powerful alternative. While its success, like all genomic prediction methods, fundamentally relies on the availability of massive datasets, DL is uniquely characterized by its ability to construct multi-layered architectures that autonomously extract complex, non-linear features [120]. At the core of DL are artificial neural networks (ANN), which draw structural inspiration from the human brain's interconnected biological neurons [121]. An ANN typically comprises an input layer, multiple hidden layers, and an output layer. Within this hierarchy, each node or "neuron" receives input through learnable weights and activation thresholds [122]. The refinement of these parameters is governed by the Backpropagation algorithm, which minimizes prediction error by propagating error backward from the output to the preceding layers—a mechanism essential for training deep architectures [122].

In the context of GS, DL models represent the current state-of-the-art for functional prediction, particularly for complex quantitative traits. Their superior performance stems from an innate capacity for self-learning and the modeling of intricate, non-linear biological relationships [123, 124]. Recent advances have also seen the rise of explainable AI within regulatory genomics, aiming to decode the "black box" of DL models [125]. The following sections provide a detailed review of the DL architectures most extensively utilized in the current GS landscape.

Multilayer perceptron

The multilayer perceptron (MLP), a fundamental class of feedforward neural networks, was first introduced to the field of genomic prediction by Okut et al. [126] to predict body weight index in mice. Structurally, an MLP consists of an input layer, one or more hidden layers, and an output layer. Each layer comprises processing units termed neurons, which are interconnected through a system of weights (w) that represent the strength of the connections [127, 128]. In a typical MLP, each neuron in a hidden layer receives inputs from all neurons in the preceding layer, computes a weighted sum, and then applies a non-linear activation function (σ) to generate an output. From a biological perspective, the step function (either 0 or 1) most closely mirrors the excitation-inhibition mechanism of biological neurons. However, due to its non-differentiability, modern practical applications predominantly utilize the sigmoid or ReLU functions to facilitate gradient-based optimization. Figure 4d presents a diagram of a three-layer MLP with four input layer units, five hidden layer units, and one output layer unit. Formally, an MLP can be defined as a universal function approximator. The general mapping from genomic markers to phenotypes is expressed as:

y=fX;θ, 20

where X is the n×p input matrix representing p genomic markers for n individuals, y is the n×1 output vector of phenotypes, and θ denotes the set of all learnable parameters (weights and biases). The goal of f∙ is to learn an optimal non-linear mapping from X to y. The final model output is a linear combination of the non-linear transformations from the terminal hidden layer:

y=fX;θ,W,b=ϕX;θTW+b, 21

where ϕX;θ represents the final hidden representation after X has passed through all hidden layers (parameterized by θ and their non-linear transformations. W and b are the weights and biases of the output layer, respectively. For a single-layer MLP, this process is simplified to:

y^=σXW1+bW2, 22

where W1 and b connect the input layer to the hidden units, and W2 connects the hidden layer to the output. σ is the non-linear activation function applied to the hidden layer. To optimize the learnable parameters, the network minimizes a loss function representing the discrepancy between predicted and observed values [129]. For regression tasks in GS, the MSE is the standard objective function:

Ly,y^=12n∑i=1n‖yi-yi^‖22, 23

where Ly,y^ is the total loss. n is the total number of individuals (observations) in the training set. yi is the true observed phenotype for individual i; and yi^ is the predicted phenotype for individual i. ‖∙‖22 denotes the Euclidean squared norm, which calculates the squared difference between yi and yi^. The weights are iteratively adjusted to minimize L, though the objective function may exhibit multiple local minima, potentially leading to convergence at sub-optimal solutions.

Despite their theoretical promise, empirical results for ANNs in genomic prediction remain mixed; for instance, while some authors [130] reported improved prediction accuracy for complex traits in Nellore cattle, whereas others found that gradient boosting and SVM algorithms outperformed ANNs [71, 131]. Still other researchers [132, 133] noted that although ANN performance varied across traits, it generally did not surpass that of classical or other non-parametric models. Despite its theoretical promise, the ANN approach faces several practical challenges: it typically requires very large datasets for effective training, is prone to overfitting, and depends on numerous hyperparameters whose optimal values can be difficult to identify [134, 135]. Consequently, ongoing research focuses on improving hyperparameter optimization [136, 137]. For instance, applying "differential evolution" as a population-based evolutionary heuristic can serve as an effective search method within the arbitrarily complex hyperparameter space of DL models [136]. Furthermore, the objective function may exhibit multiple local optima, and some algorithms may converge to suboptimal solutions where parameter values yield poor predictive capability. To mitigate the issues of regularization and uncertainty, Bayesian neural networks (BNNs) have emerged as a robust variant. By assuming a joint prior distribution over the weights and inferring a posterior distribution from the data, BNNs provides a probabilistic framework for training. This approach allows for weight sampling during forward passes and establishes a theoretical link to Gaussian processes, showing significant potential in animal breeding applications [128, 138–140].

Convolutional neural networks

Convolutional neural networks (CNNs) are a class of feedforward neural networks inspired by cognitive neuroscience. Unlike MLPs, where hidden layers are exclusively fully connected, CNNs architectures integrate convolutional layers, pooling layers, and fully connected layers [71, 134, 141]. This specialized structure allows CNN to capture local spatial correlations and hierarchical patterns within input data.

The convolutional layer serves as the primary feature extractor. By sliding a trainable convolution kernel (or filter) over the input, which in GS studies typically consists of one-dimensional SNP sequences, the model performs multiply-accumulate operations to generate feature maps [142]. The core operation of the convolutional layer is defined as follows:

glx=ϕ∑l′=1pfl′∗Kl′,lx. 24

In this formula, we define the mapping from a p-dimensional input f(x)=f1(x),⋯,fp(x) to a Q-dimensional output g(x)=(g1(x),⋯,gq(x)). First, the l′-th input feature map fl′ is convolved (denoted by ∗) with a trainable kernel Kl′,l. The kernel Kl′,l represents the specific filter connecting the l′-th input channel to the l-th output channel. Subsequently, the results from all p input channels are aggregated using the summation operator ∑. This step accumulates information from all input dimensions to generate the pre-activation value for the l-th output channel. Finally, this aggregated value is passed through a non-linear activation function ϕ to obtain the final l-th output feature map, glx.

To further condense the extracted features and enhance computational efficiency, a pooling layer is typically introduced. Common methods include max pooling and average pooling, which reduce the spatial dimensions of feature maps while maintaining robustness to positional variations:

glx=Pflx′:x′∈Vx. 25

This formula describes a feature aggregation process: V(x) defines a non-overlapping or overlapping pooling window centered at position x. P is the pooling function, which aggregates all input values flx′ within the Vx neighborhood into a single output value. glx is the resulting summarized feature map obtained after pooling the l-th feature map fl. Finally, one or more fully connected layers integrate these high-level features to perform phenotypic prediction tasks, including regression or classification [143].

In the context of GS, CNN-based models such as DeepGS [144] and ResGS [145] have demonstrated significant potential. While still largely exploratory, CNNs have shown superior performance in predicting traits like litter size and accommodating large non-additive genetic variances [71, 141]. This advantage is largely attributed to their ability to automatically extract local features and capture high-order epistatic interactions among markers through convolutional filters, which are crucial for complex traits governed by non-additive effects. Beyond single-omic prediction, CNNs are increasingly utilized to integrate multi-omics data. For instance, Fu et al. [146] developed a CNN model to prioritize candidate genes by merging diverse omics information, supported by the ISwine knowledge base. Another prominent DL algorithm is the recurrent neural network (RNN) [135]. Due to their inherent architecture, RNNs are highly suitable for modeling spatio-temporal structures and processing sequential information, finding wide application in fields such as time series prediction, natural language processing, and speech recognition [147]; however, their use in GS remains limited. Specifically, the training process for RNNs, particularly due to their complex feedback mechanisms, typically demands significant computational resources. In practical applications, the use of RNNs still faces several challenges [148]. To clarify the hyperparameter settings and optimization strategies for the models used in this study, we summarize them in Table 2.

Table 2.

Summary of key hyperparameters, optimization methods, and tuning tools for major ML and DL models used in GS

Category Representative methods Key hyperparameters Common optimization methods Hyperparameter tuning tools
Tree-based ensemble GBM (e.g., XGBoost) Learning rate, number of trees (n_estimators), max depth of trees, subsample ratio, min child weight Grid search, random search, Bayesian optimization Common hyperparameter optimization libraries: Optuna, Hyperopt, and scikit-optimize; hyperparameter tuning using AutoML libraries: Auto-Sklearn, H2O AutoML, and Auto-PyTorch; Algorithm-based hyperparameter tuning: Bayesian optimization, HyperBand, and genetic algorithms
RF Number of trees (n_estimators), max depth of trees, number of features (mtry), min samples per leaf Grid search, random search, out-of-bag error estimation
Kernel-based SVM/SVR Regularization parameter (C), kernel coefficient (γ), kernel type, epsilon (ϵ) Grid search, random search, cross-validation
RKHS/KRR Kernel function, bandwidth (h), regularization parameter (λ) Grid search, cross-validation
KNN Number of neighbors (k), distance metric (Euclidean, Manhattan, etc.), weight function Grid search, cross-validation
DL ANN (MLP) Learning rate, number of hidden layers/neurons, activation function, batch size, dropout rate, regularization (L1/L2) Grid search, random search, Bayesian optimization, differential evolution
CNN Number of convolutional layers/filters, kernel size, stride, pooling type, padding, learning rate Grid search, random search, Bayesian optimization

Recent empirical applications and comparative studies

This section synthesizes recent empirical studies comparing classical parametric models with non-parametric ML and DL methods in GS. To ensure a comprehensive review, we conducted a literature search using keywords such as "genomic selection", "machine learning", "deep learning", "genomic prediction", and "animal breeding" across databases including PubMed, Web of Science, and Google Scholar. The inclusion criteria focused on peer-reviewed studies published in English that explicitly compared ML/DL methods with traditional statistical methods in animal breeding contexts. We prioritized recent studies (mainly within the last 5–10 years) to capture the latest advancements, while also including seminal works essential for the theoretical context. By evaluating the performance of these selected studies across major livestock species, we focus on the prediction accuracy of complex traits. Table 3 summarizes these key investigations from 2018 to 2026, illustrating the relative strengths and current trends in genomic prediction for animal breeding programs.

Table 3.

Summary of empirical comparisons (2018–2026) between parametric models and ML methods in GS across major livestock species

Species Year References No. of animals No. of SNPs Traits Traditional/ML methods Key findings
Pig (real and simulated data) 2018 Waldmann [139] 3,226 (simulated), 3,141 (real) 9,723 (simulated), 50,276 (real) Quantitative trait GBLUP, BL, ABNN ABNN outperformed GBLUP and BL in prediction accuracy
Pig (real and simulated data) 2020 Waldmann et al. [141] 3,226 (simulated), 808 (real) 9,723 (simulated), 50,174 (real) Quantitative trait GBLUP, LASSO, CNN CNN reduced prediction error by > 25% in simulated data and ~3% in real pig data compared to GBLUP and LASSO
Chinese Yorkshire pigs 2022 Wang et al. [86] 2,566 44,922 Total number of piglets born and number of piglets born alive GBLUP, ssGBLUP, BayesHE, SVR, KRR, RF, Adaboost.R2 ML methods significantly outperformed traditional methods, and hyperparameter tuning substantially improved ML performance
Pigs (PIC commercial dataset + real dataset from Chifeng national pig nucleus herd, North China) 2023 Xiang et al. [85] 2,691–3,161 (PIC), 4,307–6,211 (Chifeng) 41,670 (PIC), 34,604 (Chifeng) T1, T2, T3, T4, T5, average daily gain, and total number of piglets born GBLUP, RF, SVM, XGBoost, CNN Non-linear ML approaches demonstrated feasibility and reliability for genomic prediction and genetic site screening in pigs, especially capturing complex/non-linear relationships
Rongchang pigs (Chinese local breed) 2025 Wu et al. [149] 485 22,686 (chip), 3,472,412 (imputed) Backfat thickness, loin and thoracic height, and girth circumference GBLUP, Bayesian, KRR, SVR, RF, GBDT, LightGBM, Adaboost ML methods outperformed traditional methods by 6.6–8.1% in prediction accuracy
Pigs (multiple) 2025 Su et al. [150] 1,187–4,146 (5 datasets) 32,909–52,843 (5 datasets) Reproductive and growth traits; PIC traits are hidden GBLUP, LASSO, SVR, KRR-rbf, KRR-cos, KRR-sig, stacking, Adaboost, RF, ENET, CNN, MLP GBLUP showed strong performance; stacking, SVR, and KRR-rbf also achieved high prediction accuracy and stability
Holstein dairy cattle (real and simulated data) 2020 Abdollahi et al. [71] 3,226 (simulated), 3,141 (real) 9,723 (simulated), 50,276 (real) Sire conception rate GBLUP, BayesB, RF, GBM, MLP, CNN GBM outperformed DL and traditional methods in the real bull fertility dataset, while DL showed limited advantages
Hanwoo cattle 2021 Srivastava et al. [151] 7,324 45,624 Carcass weight, marbling score, backfat thickness, and eye muscle area GBLUP, RF, XGBoost, SVM XGBoost outperformed for carcass weight and marbling score, whereas GBLUP was best for backfat thickness and eye muscle area; however, GBLUP showed the lowest MSE for all traits
Cattle (Holstein Friesian) 2022 Nazzicari et al. [152] 1,033 26,503 Simulated complex traits: additivity, dominance, and epistasis Stacked kinship CNN, GBLUP-A, GBLUP-optim GBLUP outperformed CNN: CNN showed lower root MSE but also lower Pearson correlation and normalized discounted cumulative gain compared to GBLUP-optim. GBLUP remains a strong benchmark for additive traits
Chinese Holstein dairy cattle 2024 Wang et al. [153] 5,962–6,558 45,944 Seven traits: milk production traits, type traits, and one health trait

GBLUP, BayesB, BayesCπ

RF, BANNs

BANNs framework (especially BANN_100kb) outperformed GBLUP, RF, BayesB, and BayesCπ across all seven traits
North American Holstein dairy cattle 2024 Pedrosa et al. [154] 4,511 57,600 Behavioral traits from automatic milking systems GBLUP, BL, MLP, CNN DL methods (MLP/CNN) showed slightly higher accuracies than GBLUP, but the results may not be sufficient to justify their use over traditional methods due to higher computational demand
Nellore beef cattle 2024 Mota et al. [155] 1,024 305,128 Feed efficiency-related traits STGBLUP, MTGBLUP, BayesA, BayesB, BayesC, BRR, BL, MLNN, SVR SVR and MTGBLUP outperformed MLNN, STGBLUP, and all Bayesian regression methods
Composite beef cattle 2024 Hay [156] 4,680 40,533 Birth weight, weaning weight, and yearling weight GBLUP, BayesA, BayesB, SVM, CNN, MLP, RF GBLUP showed robust performance, though RF outperformed it for weaning weight
Holstein dairy cattle 2026 Schwarz et al. [157] 7,413 45,161 Endometritis, retained placenta, stillbirth maternal, and ovary cycle disturbances GWAS (mixed linear model), RF RF complements GWAS: identified non-additive genetic architectures and epistatic effects missed by GWAS; top RF markers showed chromosomal clustering different from GWAS signals
Inner Mongolian cashmere goats 2025 Liu et al. [158] 2,256 50,728 Fiber length, cashmere diameter, and cashmere production GBLUP, RF, XGBoost, GBDT, LightGBM LightGBM, RF, and GBDT provided higher genomic prediction accuracy than traditional GBLUP for cashmere traits in cashmere goats
Yellow-feathered broilers 2025 Liu et al. [67] 989–1,347 16,962–31,934 Eight traits: laying, growth, and carcass traits (highlighted half-eviscerated weight and eviscerated weight) GBLUP, BayesA, BayesB, BayesC, BL, BRR, SVR, RF, GBDT, XGBoost, LightGBM, KRR, MLP Traditional methods (GBLUP and Bayesian) outperformed ML methods in 5 out of 8 traits, whereas ML methods (especially SVR and RF) showed significantly higher accuracy for carcass traits (highlighted half-eviscerated weight and eviscerated weight)
Yellow-feathered broilers 2026 Chai et al. [159] 834 500–100,000 Eight traits: egg production and quality traits GBLUP, MLP, RF RF outperformed for low-heritability traits (e.g., egg number, shell traits), while GBLUP was better for high-heritability egg weight traits
Simulated livestock population 2025 Yang et al. [160] 12,500 (training), 500 (testing) 55,043 Quantitative trait GBLUP, 2GBLUP, weighted 2GBLUP, RF, SVR QTL integration benefits GBLUP/SVR but not RF, possibly due to data structure or RF's unsuitability for this genomic prediction scenario

All abbreviations used in this table are defined in the 'Abbreviations' section

As summarized in Table 3, empirical studies from 2018 to 2026 have increasingly explored ML and DL methods in GS, frequently benchmarking them against traditional linear models like GBLUP, ssGBLUP, and Bayesian regressions. Results illustrate that ML approaches often outperform traditional parametric models, particularly for traits characterized by complex or non-additive genetic architectures. However, performance varies depending on trait heritability, population size, and validation design. Most investigations focus on pigs and dairy or beef cattle. However, emerging applications can also be seen in goats and poultry. The evaluated traits cover a broad range: specifically, for cattle, these include milk production and behavioral traits. In pigs, the focus is on reproductive and growth traits. Other examples include cashmere production in goats, as well as egg production and carcass traits in broilers. Despite these advancements, the landscape of genomic prediction faces new hurdles. With the exponential surge in the number of whole-genome markers, achieving accurate predictions under high-dimensional and small-sample conditions remains a formidable challenge. While traditional methods like KNN and SVM have been widely adopted, limitations exist. Notably, recent studies by Ali et al. [161] on human genomic data demonstrate that the improved double-weighted k-nearest neighbor algorithm can more effectively filter out noise. By incorporating feature weighting and exponential distance weighting, it enhances classification accuracy, suggesting that integrating such weighted distance algorithms into animal GS models represents a promising direction for future breeding value predictions.

Overall, non-parametric ML and DL approaches frequently outperform traditional parametric models, particularly for complex or low-heritability traits involving non-additive effects or heterogeneous genetic architectures. In empirical studies, reported accuracy gains range from 3% to 10%, whereas these gains can reach up to 25%. Ensemble methods and kernel-based techniques demonstrate particular robustness and stability across scenarios. DL architectures excel when non-linear interactions or large datasets are present; however, they are sensitive to hyperparameter tuning and risk overfitting. In contrast, linear models like GBLUP remain competitive or superior for additive, high-heritability traits and smaller datasets, while also offering advantages in computational efficiency and interpretability. These trends underscore the potential of ML and DL to address limitations of classical frameworks in the big data era, yet they also highlight the critical need for larger reference populations and rigorous model validation.

Conclusions and future perspectives

The statistical foundations of GS have undergone a remarkable evolution, beginning with pedigree-based linear mixed models that excel in predicting additive, high-heritability traits. The field then progressed to Bayesian regressions, which better accommodate heterogeneous marker effects and oligogenic architectures. Today, in the era of big data breeding, non-parametric ML and DL approaches are capable of modeling complex nonlinear interactions. Traditional parametric models remain computationally efficient and robust; consequently, they are widely implemented in routine livestock breeding programs, providing reliable predictions with interpretable assumptions. In contrast, ML and DL methods demonstrate increasing promise for low-heritability or non-additive traits, and empirical studies in cattle, pigs, poultry, and goats have shown clear accuracy gains with these advanced methods.

Nevertheless, no single method universally dominates; performance depends on trait genetic architecture, dataset size, population structure, and computational resources. Ensemble techniques and kernel-based methods often provide stable improvements with moderate demands, whereas DL architectures shine in large-scale, high-dimensional scenarios but require careful hyperparameter optimization to mitigate overfitting and enhance interpretability. Looking ahead, hybrid models combining parametric efficiency with non-parametric flexibility, alongside multi-omics integration and explainable AI, are poised to accelerate genetic progress in animal breeding further. These advances will support sustainable livestock production by enabling more precise selection for complex traits such as disease resistance, feed efficiency, and environmental resilience. Ultimately, judicious model selection, guided by empirical validation and breeding objectives, will maximize genomic gains in practical programs worldwide.

Acknowledgements

The authors would like to acknowledge the use of DeepL (web-based), Qwen-3.7, and Grammarly (v1.2.279.1925) for assistance with English language editing and grammar checking to improve the readability of this manuscript. The authors take full responsibility for the content of this paper.

Abbreviations

ABNN

Approximate Bayesian neural networks

Adaboost

Adaptive boosting

ANN

Artificial neural networks

BANNs

Biologically annotated neural networks

BayesHE

Bayesian horseshoe

BL

Bayesian Lasso

BLUP

Best linear unbiased prediction

BNN

Bayesian neural networks

BRR

Bayesian ridge regression

CART

Classification and regression tree

CNN

Convolutional neural networks

DL

Deep learning

DT

Decision tree

ENET

Elastic net

eQTL

Expression QTL

GBDT

Gradient boosting decision trees

GBLUP

Genomic BLUP

GBM

Gradient boosting machines

GEBV

Genomic estimated breeding values

G × E

Genotype-by-environment interaction

GS

Genomic selection

GWAS

Genome-wide association study

KBMF

Kernelized Bayesian matrix factorization

KNN

K-nearest neighbors

KRR

Kernel ridge regression

LD

Linkage disequilibrium

LightGBM

Light gradient boosting machine

MAS

Marker-assisted selection

MCMC

Markov chain Monte Carlo

MH

Metropolis-Hastings

ML

Machine learning

MLNN

Multi-layer neural network

MLP

Multilayer perceptron

MME

Mixed model equations

MSE

Mean squared error

MTGBLUP

Multi-trait GBLUP

NB

Naive Bayes

QTL

Quantitative trait locus

RF

Random forest

RKHS

Reproducing kernel Hilbert space

RNN

Recurrent neural network

RRBLUP

Ridge regression BLUP

SNP

Single nucleotide polymorphism

ssBR

Single-step Bayesian regression

ssGBLUP

Single-step GBLUP

STGBLUP

Single-trait GBLUP

SVM

Support vector machines

SVR

Support vector regression

TABLUP

Trait-specific marker derived relationship matrix BLUP

XGBoost

Extreme gradient boosting

Authors’ contributions

LZ: Conceptualization, methodology, writing – original draft, and visualization. MZ, XW, CC, SW, YZ, and YZ: writing – review & editing. YL and RW: conceptualization, supervision, funding acquisition, and writing – review & editing. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Inner Mongolia Autonomous Region Breeding Joint Research Project (YZ2023011), the Science and Technology Program of the Inner Mongolia Autonomous Region (2025YFHH0226), the Fundamental Research Funds of Directly Subordinate Universities of Inner Mongolia Autonomous Region (BR251201), the Construction and Demonstration of a Genomic Information and Smart Breeding Platform for Meat-Producing Animals (2025KJHZ005), the Inner Mongolia Autonomous Region "Leading the Charge with Open Competition" topic (2022JBGS0024), the Program for Innovative Research Team in Universities of Inner Mongolia Autonomous Region (NMGIRT2322), the Earmarked Fund for China Agriculture Re-search System of Mutton Sheep (CARS-38), the Science and Technology Plan of Inner Mongolia Autonomous Region (2023KYPT0021), the Research on Key Technologies for Breeding New Strains of Low-Fat and High-Yield Grassland Short-Tailed Sheep (NC2024005).

Data availability

No datasets were generated or analysed during the current study.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Contributor Information

Ruijun Wang, Email: nmgwrj@126.com.

Yongbin Liu, Email: Liuyongbin@imau.edu.cn.

References

  • 1.Henderson CR. Best linear unbiased estimation and prediction under a selection model. Biometrics. 1975;31:423–47. [PubMed] [Google Scholar]
  • 2.Soller M, Beckmann JS. Genetic polymorphism in varietal identification and genetic improvement. Theor Appl Genet. 1983;67:25–33. 10.1007/BF00303917. [DOI] [PubMed] [Google Scholar]
  • 3.Meuwissen TH, Hayes BJ, Goddard ME. Prediction of total genetic value using genome-wide dense marker maps. Genetics. 2001;157:1819–29. 10.1093/genetics/157.4.1819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Goddard ME, Hayes BJ. Genomic selection. J Anim Breed Genet. 2007;124:323–30. 10.1111/j.1439-0388.2007.00702.x. [DOI] [PubMed] [Google Scholar]
  • 5.VanRaden PM. Efficient methods to compute genomic predictions. J Dairy Sci. 2008;91:4414–23. 10.3168/jds.2007-0980. [DOI] [PubMed] [Google Scholar]
  • 6.Legarra A, Aguilar I, Misztal I. A relationship matrix including full pedigree and genomic information. J Dairy Sci. 2009;92:4656–63. 10.3168/jds.2009-2061. [DOI] [PubMed] [Google Scholar]
  • 7.Christensen OF, Lund MS. Genomic prediction when some animals are not genotyped. Genet Sel Evol. 2010;42:2. 10.1186/1297-9686-42-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Habier D, Fernando RL, Kizilkaya K, Garrick DJ. Extension of the bayesian alphabet for genomic selection. BMC Bioinf. 2011;12:186. 10.1186/1471-2105-12-186. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Crossa J, Montesinos-Lopez OA, Costa-Neto G, Vitale P, Martini JWR, Runcie D, et al. Machine learning algorithms translate big data into predictive breeding accuracy. Trends Plant Sci. 2025;30:167–84. 10.1016/j.tplants.2024.09.011. [DOI] [PubMed] [Google Scholar]
  • 10.Ma W, Zheng W, Qin S, Wang C, Lei B, Liu Y. DeepAnnotation: a novel interpretable deep learning–based genomic selection model that integrates comprehensive functional annotations. Gigascience. 2025;14:giaf083. 10.1093/gigascience/giaf083. [DOI] [PMC free article] [PubMed]
  • 11.Dekkers JCM. Application of genomics tools to animal breeding. Curr Genomics. 2012;13:207–12. 10.2174/138920212800543057. [DOI] [PMC free article] [PubMed]
  • 12.Goddard ME, Hayes BJ, Meuwissen THE. Using the genomic relationship matrix to predict the accuracy of genomic selection. J Anim Breed Genet. 2011;128:409–21. 10.1111/j.1439-0388.2011.00964.x. [DOI] [PubMed] [Google Scholar]
  • 13.Yang J, Benyamin B, McEvoy BP, Gordon S, Henders AK, Nyholt DR, et al. Common SNPs explain a large proportion of the heritability for human height. Nat Genet. 2010;42:565–9. 10.1038/ng.608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Guo X, Christensen OF, Ostersen T, Wang Y, Lund MS, Su G. Improving genetic evaluation of litter size and piglet mortality for both genotyped and nongenotyped individuals using a single-step method. J Anim Sci. 2015;93:503–12. 10.2527/jas.2014-8331. [DOI] [PubMed] [Google Scholar]
  • 15.Wang J, Zhou Z, Zhang Z, Li H, Liu D, Zhang Q, et al. Expanding the BLUP alphabet for genomic prediction adaptable to the genetic architectures of complex traits. Heredity. 2018;121:648–62. 10.1038/s41437-018-0075-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhang Z, Liu J, Ding X, Bijma P, de Koning D-J, Zhang Q. Best linear unbiased prediction of genomic breeding values using a trait-specific marker-derived relationship matrix. PLoS ONE. 2010;5:e12648. 10.1371/journal.pone.0012648. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Haque MA, Jang EB, Lee HD, Shin DH, Jang JH, Kim JJ. Performance of weighted genomic BLUP and bayesian methods for hanwoo carcass traits. Trop Anim Health Prod. 2025;57:38. 10.1007/s11250-025-04293-y. [DOI] [PubMed] [Google Scholar]
  • 18.Misztal I, Legarra A, Aguilar I. Computing procedures for genetic evaluation including phenotypic, full pedigree, and genomic information. J Dairy Sci. 2009;92:4648–55. 10.3168/jds.2009-2064. [DOI] [PubMed] [Google Scholar]
  • 19.Song H, Zhang J, Zhang Q, Ding X. Using different single-step strategies to improve the efficiency of genomic prediction on body measurement traits in pig. Front Genet. 2018;9:730. 10.3389/fgene.2018.00730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Kudinov AA, Koivula M, Aamand GP, Strandén I, Mäntysaari EA. Single-step genomic BLUP with many metafounders. Front Genet. 2022;13:1012205. 10.3389/fgene.2022.1012205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Christensen OF. Compatibility of pedigree-based and marker-based relationship matrices for single-step genetic evaluation. Genet Sel Evol. 2012;44:37. 10.1186/1297-9686-44-37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Legarra A, Christensen OF, Vitezica ZG, Aguilar I, Misztal I. Ancestral relationships using metafounders: finite ancestral populations and across population relationships. Genetics. 2015;200:455–68. 10.1534/genetics.115.177014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Xiang T, Christensen OF, Legarra A. Technical note: Genomic evaluation for crossbred performance in a single-step approach with metafounders. J Anim Sci. 2017;95:1472–80. 10.2527/jas.2016.1155. [DOI] [PubMed] [Google Scholar]
  • 24.Xiang T, Nielsen B, Su G, Legarra A, Christensen OF. Application of single-step genomic evaluation for crossbred performance in pig. J Anim Sci. 2016;94:936–48. 10.2527/jas.2015-9930. [DOI] [PubMed] [Google Scholar]
  • 25.Fernando RL, Dekkers JC, Garrick DJ. A class of Bayesian methods to combine large numbers of genotyped and non-genotyped animals for whole-genome analyses. Genet Sel Evol. 2014;46:50. 10.1186/1297-9686-46-50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Paiva TB, Macedo J, Alves JJ, Santos D, Aspilcueta-Borquis RR, Tonhati H, et al. Single-step Bayesian regression methods for genomic evaluation of milk yield of Murrah buffaloes. J Dairy Res. 2025;92:66–8. 10.1017/S0022029925000317. [DOI] [PubMed] [Google Scholar]
  • 27.Pang Z, Wang W, Huang P, Zhang H, Zhang S, Yang P, et al. Enhancing genomic prediction accuracy with a single-step genomic best linear unbiased prediction model integrating genome-wide association study results. Animals. 2025;15:1268. 10.3390/ani15091268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Mancisidor B, Cruz A, Gutiérrez G, Burgos A, Morón JA, Wurzinger M, et al. ssGBLUP method improves the accuracy of breeding value prediction in huacaya alpaca. Animals. 2021;11:3052. 10.3390/ani11113052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Himmelbauer J, Schwarzenbacher H, Fuerst C, Fuerst-Waltl B. Exploring unknown parent groups and metafounders in single-step genomic best linear unbiased prediction: Insights from a simulated cattle population. J Dairy Sci. 2024;107:8170–92. 10.3168/jds.2024-24891. [DOI] [PubMed] [Google Scholar]
  • 30.Whittaker JC, Thompson R, Denham MC. Marker-assisted selection using ridge regression. Genet Res. 2000;75:249–52. 10.1017/s0016672399004462. [DOI] [PubMed] [Google Scholar]
  • 31.Wang C, Zhang Q, Jiang L, Qian R, Ding X, Zhao Y. Comparative study of estimation methods for genomic breeding values. Sci Bull. 2016;61:353–6. 10.1007/s11434-016-1014-1. [Google Scholar]
  • 32.Li X, Chen X, Wang Q, Yang N, Sun C. Integrating bioinformatics and machine learning for genomic prediction in chickens. Genes. 2024;15:690. 10.3390/genes15060690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sahebalam H, Gholizadeh M. Different approaches for estimating the shrinkage factor in ridge regression BLUP for genomic selection. Sci Rep. 2025;15:42142. 10.1038/s41598-025-26193-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Mathew B, Bauer AM, Koistinen P, Reetz TC, Léon J, Sillanpää MJ. Bayesian adaptive Markov chain Monte Carlo estimation of genetic parameters. Heredity. 2012;109:235–45. 10.1038/hdy.2012.35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Park T, Casella G. The Bayesian Lasso. J Am Stat Assoc. 2008;103:681–6. 10.1198/016214508000000337.
  • 36.Erbe M, Hayes BJ, Matukumalli LK, Goswami S, Bowman PJ, Reich CM, et al. Improving accuracy of genomic predictions within and between dairy cattle breeds with imputed high-density single nucleotide polymorphism panels. J Dairy Sci. 2012;95:4114–29. 10.3168/jds.2011-5019. [DOI] [PubMed] [Google Scholar]
  • 37.MacLeod IM, Bowman PJ, Vander Jagt CJ, Haile-Mariam M, Kemper KE, Chamberlain AJ, et al. Exploiting biological priors and sequence variants enhances QTL discovery and genomic prediction of complex traits. BMC Genomics. 2016;17:144. 10.1186/s12864-016-2443-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Mollandin F, Gilbert H, Croiseau P, Rau A. Accounting for overlapping annotations in genomic prediction models of complex traits. BMC Bioinf. 2022;23:365. 10.1186/s12859-022-04914-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Cai W, Cole JB, Goddard ME, Li J, Zhang S, Song J. Mammary gland multi-omics data reveals new genetic insights into milk production traits in dairy cattle. PLoS Genet. 2025;21:e1011675. 10.1371/journal.pgen.1011675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.ter Braak CJF, Boer MP, Bink MCAM. Extending Xu’s Bayesian model for estimating polygenic effects using markers of the entire genome. Genetics. 2005;170:1435–8. 10.1534/genetics.105.040469. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Xu S. Estimating polygenic effects using markers of the entire genome. Genetics. 2003;163:789–801. 10.1093/genetics/163.2.789. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Wang C, Rutledge J, Gianola D. Marginal inferences about variance components in a mixed linear model using Gibbs sampling. Genet Sel Evol. 1993;25:41. 10.1186/1297-9686-25-1-41. [Google Scholar]
  • 43.Gianola D, De Los CG, Hill WG, Manfredi E, Fernando R. Additive genetic variability and the Bayesian alphabet. Genetics. 2009;183:347–63. 10.1534/genetics.109.103952. [DOI] [PMC free article] [PubMed]
  • 44.van den Berg I, Fritz S, Boichard D. QTL fine mapping with bayes C(π): a simulation study. Genet Sel Evol. 2013;45:19. 10.1186/1297-9686-45-19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Simeão RM, Resende MDV, Alves RS, Pessoa-Filho M, Azevedo ALS, Jones CS, et al. Genomic selection in tropical forage grasses: current status and future applications. Front Plant Sci. 2021;12:665195. 10.3389/fpls.2021.665195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Gu J, Guo J, Zhang Z, Xu Y, Qadri QR, Zhang Z, et al. Molecular design-based breeding: a kinship index-based selection method for complex traits in small livestock populations. Genes. 2023;14:807. 10.3390/genes14040807. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Yin L, Zhang H, Zhou X, Yuan X, Zhao S, Li X, et al. KAML: Improving genomic prediction accuracy of complex traits using machine learning determined parameters. Genome Biol. 2020;21:146. 10.1186/s13059-020-02052-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Sun X, Qu L, Garrick DJ, Dekkers JCM, Fernando RL. A fast EM algorithm for BayesA-like prediction of genomic breeding values. PLoS ONE. 2012;7:e49157. 10.1371/journal.pone.0049157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Meuwissen THE, Solberg TR, Shepherd R, Woolliams JA. A fast algorithm for BayesB type of prediction of genome-wide estimates of genetic value. Genet Sel Evol. 2009;41:2. 10.1186/1297-9686-41-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Hayashi T, Iwata H. EM algorithm for Bayesian estimation of genomic breeding values. BMC Genet. 2010;11:3. 10.1186/1471-2156-11-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Wang T, Chen YPP, Goddard ME, Meuwissen THE, Kemper KE, Hayes BJ. A computationally efficient algorithm for genomic prediction using a Bayesian model. Genet Sel Evol. 2015;47:34. 10.1186/s12711-014-0082-4. [DOI] [PMC free article] [PubMed]
  • 52.Mutshinda CM, Sillanpää MJ. Extended Bayesian LASSO for multiple quantitative trait loci mapping and unobserved phenotype prediction. Genetics. 2010;186:1067–75. 10.1534/genetics.110.119586. [DOI] [PMC free article] [PubMed]
  • 53.Brøndum RF, Su G, Lund MS, Bowman PJ, Goddard ME, Hayes BJ. Genome position specific priors for genomic prediction. BMC Genomics. 2012;13:543. 10.1186/1471-2164-13-543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Wang CL, Ding XD, Wang JY, Liu JF, Fu WX, Zhang Z, et al. Bayesian methods for estimating GEBVs of threshold traits. Heredity. 2013;110:213–9. 10.1038/hdy.2012.65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Song H, Hu H. Strategies to improve the accuracy and reduce costs of genomic prediction in aquaculture species. Evol Appl. 2021;15:578–90. 10.1111/eva.13262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Zheng W, Zhang Q, He J, Han B, Zhang Q, Sun D. Comparative evaluation of SNP-weighted, Bayesian, and machine learning models for genomic prediction in Holstein cattle. BMC Genomics. 2025;26:1037. 10.1186/s12864-025-12218-0. [DOI] [PMC free article] [PubMed]
  • 57.Zhu S, Guo T, Yuan C, Liu J, Li J, Han M, et al. Evaluation of Bayesian alphabet and GBLUP based on different marker density for genomic prediction in Alpine Merino sheep. G3 (Bethesda). 2021;11:jkab206. 10.1093/g3journal/jkab206. [DOI] [PMC free article] [PubMed]
  • 58.Haque MA, Lee YM, Ha JJ, Jin S, Park B, Kim NY, et al. Genomic predictions in Korean Hanwoo cows: A comparative analysis of genomic BLUP and Bayesian methods for reproductive traits. Animals. 2023;14:27. 10.3390/ani14010027. [DOI] [PMC free article] [PubMed]
  • 59.Pirompud P, Sivapirunthep P, Punyapornwithaya V, Chaosap C. Application of machine learning algorithms to predict dead on arrival of broiler chickens raised without antibiotic program. Poult Sci. 2024;103:103504. 10.1016/j.psj.2024.103504. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Navaz AN, T. El-Kassabi H, Serhani MA, Oulhaj A, Khalil K. A novel patient similarity network (PSN) framework based on multi-model deep learning for precision medicine. J Pers Med. 2022;12:768. 10.3390/jpm12050768. [DOI] [PMC free article] [PubMed]
  • 61.Liu Z, Moroz YS, Isayev O. The challenge of balancing model sensitivity and robustness in predicting yields: a benchmarking study of amide coupling reactions. Chem Sci. 2023;14:10835–46. 10.1039/d3sc03902a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Jiang T, Gradus JL, Rosellini AJ. Supervised machine learning: a brief primer. Behav Ther. 2020;51:675–87. 10.1016/j.beth.2020.05.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Breiman L, Friedman JH, Olshen RA, Stone CJ. Classification and regression trees (CART). Biometrics. 1984;40:358. 10.2307/2530946. [Google Scholar]
  • 64.Lu J, Lu X, Wang Y, Zhang H, Han L, Zhu B, et al. Comparison between logistic regression and machine learning algorithms on prediction of noise-induced hearing loss and investigation of SNP loci. Sci Rep. 2025;15:15361. 10.1038/s41598-025-00050-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Quinlan JR. Induction of decision trees. Mach Learn. 1986;1:81–106. 10.1007/BF00116251. [Google Scholar]
  • 66.Dou Y, Liu J, Meng W, Zhang Y. Comparative analysis of supervised learning algorithms for prediction of cardiovascular diseases. Technol Health Care. 2024;32:241–51. 10.3233/THC-248021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Liu B, Liu H, Tu J, Xiao J, Yang J, He X, et al. An investigation of machine learning methods applied to genomic prediction in yellow-feathered broilers. Poult Sci. 2025;104:104489. 10.1016/j.psj.2024.104489. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Breiman L. Bagging predictors. Mach Learn. 1996;24:123–40. 10.1023/A:1018054314350. [Google Scholar]
  • 69.Schapire RE. The boosting approach to machine learning: an overview. In: Denison DD, Hansen MH, Holmes CC, Mallick B, Yu B, editors. Nonlinear estimation and classification. Lecture Notes in Statistics, vol 171. New York: Springer. 2003;171:149–71. 10.1007/978-0-387-21579-2_9.
  • 70.Sigletos G, Paliouras G, Spyropoulos CD, Hatzopoulos M, Cohen W. Combining information extraction systems using voting and stacked generalization. J Mach Learn Res. 2005;6:1751–82. [Google Scholar]
  • 71.Abdollahi-Arpanahi R, Gianola D, Peñagaricano F. Deep learning versus parametric and ensemble methods for genomic prediction of complex phenotypes. Genet Sel Evol. 2020;52:12. 10.1186/s12711-020-00531-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Wang W, Wu Y, Liu W, Fu T, Qiu R, Wu S. Tensile performance mechanism for bamboo fiber-reinforced, palm oil-based resin bio-composites using finite element simulation and machine learning. Polymers. 2023;15:2633. 10.3390/polym15122633. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Crisci C, Ghattas B, Perera G. A review of supervised machine learning algorithms and their applications to ecological data. Ecol Modell. 2012;240:113–22. 10.1016/j.ecolmodel.2012.03.001. [Google Scholar]
  • 74.Fu S, Chen H. Machine learning-based intelligent scoring system for English essays under the background of modern information technology. Bhardwaj A, editor. Comput Intell Neurosci. 2022;2022:6912018. 10.1155/2022/6912018. [DOI] [PMC free article] [PubMed]
  • 75.Freund Y. Experiment with a new boosting algorithm. Morgan Kaufmann. 1996;148–56.
  • 76.Shrestha DL, Solomatine DP. Experiments with AdaBoost.RT, an improved boosting scheme for regression. Neural Comput. 2006;18:1678–710. 10.1162/neco.2006.18.7.1678. [DOI] [PubMed]
  • 77.Wang J, Chai J, Chen L, Zhang T, Long X, Diao S, et al. Enhancing genomic prediction accuracy of reproduction traits in rongchang pigs through machine learning. Animals. 2025;15:525. 10.3390/ani15040525. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Liang M, Miao J, Wang X, Chang T, An B, Duan X, et al. Application of ensemble learning to genomic selection in Chinese simmental beef cattle. J Anim Breed Genet. 2021;138:291–9. 10.1111/jbg.12514. [DOI] [PubMed] [Google Scholar]
  • 79.Natekin A, Knoll A. Gradient boosting machines, a tutorial. Front Neurorob. 2013;7:21. 10.3389/fnbot.2013.00021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Stat. 2001;29 (5):1189–232. 10.1214/aos/1013203451.
  • 81.Chen T, Guestrin C. XGBoost: a scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016;785–94. 10.1145/2939672.2939785.
  • 82.Breiman L. Random forests. Mach Learn. 2001;45:5–32. 10.1023/A:1010933404324. [Google Scholar]
  • 83.Ghosh D, Cabrera J. Enriched random forest for high dimensional genomic data. IEEE/ACM Trans Comput Biol Bioinform. 2022;19:2817–28. 10.1109/TCBB.2021.3089417. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Wang K, Yang B, Li Q, Liu S. Systematic evaluation of genomic prediction algorithms for genomic prediction and breeding of aquatic animals. Genes. 2022;13:2247. 10.3390/genes13122247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Xiang T, Li T, Li J, Li X, Wang J. Using machine learning to realize genetic site screening and genomic prediction of productive traits in pigs. FASEB J. 2023;37:e22961. 10.1096/fj.202300245R. [DOI] [PubMed] [Google Scholar]
  • 86.Wang X, Shi S, Wang G, Luo W, Wei X, Qiu A, et al. Using machine learning to improve the accuracy of genomic prediction of reproduction traits in pigs. J Anim Sci Biotechnol. 2022;13:60. 10.1186/s40104-022-00708-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Cortes C, Vapnik V. Support-vector networks. Mach Learn. 1995;20:273–97. 10.1007/BF00994018. [Google Scholar]
  • 88.Huang S, Cai N, Pacheco PP, Narrandes S, Wang Y, Xu W. Applications of support vector machine (SVM) learning in cancer genomics. Cancer Genom Proteom. 2018;15:41–51. 10.21873/cgp.20063. [DOI] [PMC free article] [PubMed]
  • 89.Zhang S, Zhou Z, Chen X, Hu Y, Yang L. pDHS-SVM: A prediction method for plant DNase I hypersensitive sites based on support vector machine. J Theor Biol. 2017;426:126–33. 10.1016/j.jtbi.2017.05.030. [DOI] [PubMed] [Google Scholar]
  • 90.Kupp SD, VanGordon IA, Gönen M, Esener S, Eksi SE, Ak Ç. Interpretable and integrative analysis of single-cell multiomics with scMKL. Commun Biol. 2025;8:1160. 10.1038/s42003-025-08533-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Cichonska A, Ravikumar B, Parri E, Timonen S, Pahikkala T, Airola A, et al. Computational-experimental approach to drug-target interaction mapping: a case study on kinase inhibitors. Plos Comput Biol. 2017;13:e1005678. 10.1371/journal.pcbi.1005678. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Galal A, Talal M, Moustafa A. Applications of machine learning in metabolomics: disease modeling and classification. Front Genet. 2022;13:1017340. 10.3389/fgene.2022.1017340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Liang M, An B, Li K, Du L, Deng T, Cao S, et al. Improving genomic prediction with machine learning incorporating TPE for hyperparameters optimization. Biology. 2022;11:1647. 10.3390/biology11111647. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Zhao W, Lai X, Liu D, Zhang Z, Ma P, Wang Q, et al. Applications of support vector machine in genomic prediction in pig and maize populations. Front Genet. 2020;11:598318. 10.3389/fgene.2020.598318. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Caragea C, Sinapov J, Silvescu A, Dobbs D, Honavar V. Glycosylation site prediction using ensembles of support vector machine classifiers. BMC Bioinf. 2007;8:438. 10.1186/1471-2105-8-438. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Long N, Gianola D, Rosa GJM, Weigel KA. Application of support vector regression to genome-assisted prediction of quantitative traits. Theor Appl Genet. 2011;123:1065–74. 10.1007/s00122-011-1648-y. [DOI] [PubMed] [Google Scholar]
  • 97.Li M, Hall T, MacHugh DE, Chen L, Garrick D, Wang L, et al. KPRR: A novel machine learning approach for effectively capturing nonadditive effects in genomic prediction. Briefings Bioinf. 2025;26:bbae683. 10.1093/bib/bbae683. [DOI] [PMC free article] [PubMed]
  • 98.Manton JH, Amblard P-O. A primer on reproducing kernel Hilbert spaces. Found Trends Signal Process. 2015;8:1–126. 10.1561/2000000050. [Google Scholar]
  • 99.Ljoncheva M, Stepišnik T, Kosjek T, Džeroski S. Machine learning for identification of silylated derivatives from mass spectra. J Cheminf. 2022;14:62. 10.1186/s13321-022-00636-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Parzen E. Extraction and detection problems and reproducing kernel Hilbert spaces. J Soc Ind Appl Math Ser A Control. 1962;1:35–62. 10.1137/0301004. [Google Scholar]
  • 101.Gaudelet T, Day B, Jamasb AR, Soman J, Regep C, Liu G, et al. Utilizing graph machine learning within drug discovery and development. Briefings Bioinf. 2021;22:bbab159. 10.1093/bib/bbab159. [DOI] [PMC free article] [PubMed]
  • 102.Gianola D, Fernando RL, Stella A. Genomic-assisted prediction of genetic value with semiparametric procedures. Genetics. 2006;173:1761–76. 10.1534/genetics.105.049510. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.de Los CG, Gianola D, Rosa GJM. Reproducing kernel Hilbert spaces regression: a general framework for genetic evaluation. J Anim Sci. 2009;87:1883–7. 10.2527/jas.2008-1259. [DOI] [PubMed] [Google Scholar]
  • 104.Blondel M, Onogi A, Iwata H, Ueda N. A ranking approach to genomic selection. PLoS ONE. 2015;10:e0128570. 10.1371/journal.pone.0128570. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.González-Recio O, Rosa GJM, Gianola D. Machine learning methods and predictive ability metrics for genome-wide prediction of complex traits. Livest Sci. 2014;166:217–31. 10.1016/j.livsci.2014.05.036. [Google Scholar]
  • 106.Xu Y, Xu C, Xu S. Prediction and association mapping of agronomic traits in maize using multiple omic data. Heredity. 2017;119:174–84. 10.1038/hdy.2017.27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Pérez P, De Los CG. Genome-wide regression and prediction with the BGLR statistical package. Genetics. 2014;198:483–95. 10.1534/genetics.114.164442. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108. Karacaören B. An evaluation of machine learning for genomic prediction of hairy syndrome in dairy cattle. Anim Sci Pap Rep. 2022;40:45–58.
  • 109.Long N, Gianola D, Rosa GJM, Weigel KA, Avendaño S. Machine learning classification procedure for selecting SNPs in genomic selection: application to early mortality in broilers. J Anim Breed Genet. 2007;124:377–89. 10.1111/j.1439-0388.2007.00694.x. [DOI] [PubMed] [Google Scholar]
  • 110.Long N, Gianola D, Rosa GJM, Weigel KA, Avendaño S. Marker-assisted assessment of genotype by environment interaction: a case study of single nucleotide polymorphism-mortality association in broilers in two hygiene environments. J Anim Sci. 2008;86:3358–66. 10.2527/jas.2008-1021. [DOI] [PubMed]
  • 111.van der Heide EMM, Veerkamp RF, van Pelt ML, Kamphuis C, Athanasiadis I, Ducro BJ. Comparing regression, naive bayes, and random forest methods in the prediction of individual survival to second lactation in Holstein cattle. J Dairy Sci. 2019;102:9409–21. 10.3168/jds.2019-16295. [DOI] [PubMed]
  • 112.Putnová L, Štohl R. Comparing assignment-based approaches to breed identification within a large set of horses. J Appl Genet. 2019;60:187–98. 10.1007/s13353-019-00495-x. [DOI] [PubMed] [Google Scholar]
  • 113.Burocziova M, Riha J. Horse breed discrimination using machine learning methods. J Appl Genet. 2009;50:375–7. 10.1007/BF03195696. [DOI] [PubMed] [Google Scholar]
  • 114.Gillberg J, Marttinen P, Mamitsuka H, Kaski S. Modelling G×E with historical weather information improves genomic prediction in new environments. Bioinformatics. 2019;35:4045–52. 10.1093/bioinformatics/btz197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Chen SY, Freitas PHF, Oliveira HR, Lázaro SF, Huang YJ, Howard JT, et al. Genotype-by-environment interactions for reproduction, body composition, and growth traits in maternal-line pigs based on single-step genomic reaction norms. Genet Sel Evol. 2021;53:51. 10.1186/s12711-021-00645-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Hu X, Xie W, Wu C, Xu S. A directed learning strategy integrating multiple omic data improves genomic prediction. Plant Biotechnol J. 2019;17:2011–20. 10.1111/pbi.13117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Ye S, Li J, Zhang Z. Multi-omics-data-assisted genomic feature markers preselection improves the accuracy of genomic prediction. J Anim Sci Biotechnol. 2020;11:109. 10.1186/s40104-020-00515-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Xue Y, Zhou L, Zhuo Y, Li W, Ma S, Du H, et al. FSBLUP: a novel strategy of fusion similarity matrix construction via optimally integrating intermediate omics data to enhance genomic prediction. Genome Biol. 2026;27:27. 10.1186/s13059-026-03931-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Liang M, An B, Chang T, Deng T, Du L, Li K, et al. Incorporating kernelized multi-omics data improves the accuracy of genomic prediction. J Anim Sci Biotechnol. 2022;13:103. 10.1186/s40104-022-00756-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Li D, Bai L, Wang R, Ying S. Research progress of machine learning in extending and regulating the shelf life of fruits and vegetables. Foods. 2024;13:3025. 10.3390/foods13193025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.ZareBidaki M, Allahyari E, Zeinali T, Asgharzadeh M. Occurrence and risk factors of brucellosis among domestic animals: an artificial neural network approach. Trop Anim Health Prod. 2022;54:62. 10.1007/s11250-022-03076-z. [DOI] [PubMed] [Google Scholar]
  • 122.Rumelhart DE, Hinton GE, Williams RJ. Learning representations by back-propagating errors. Nature. 1986;323:533–6. 10.1038/323533a0. [Google Scholar]
  • 123.Bellot P, de Los CG, Pérez-Enciso M. Can deep learning improve genomic prediction of complex human traits? Genetics. 2018;210:809–19. 10.1534/genetics.118.301298. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Xu L, Guo Z, Liu X. Prediction of essential genes in prokaryote based on artificial neural network. Genes Genomics. 2020;42:97–106. 10.1007/s13258-019-00884-w. [DOI] [PubMed] [Google Scholar]
  • 125.Novakovsky G, Dexter N, Libbrecht MW, Wasserman WW, Mostafavi S. Obtaining genetics insights from deep learning via explainable artificial intelligence. Nat Rev Genet. 2023;24:125–37. 10.1038/s41576-022-00532-2. [DOI] [PubMed] [Google Scholar]
  • 126.Okut H, Gianola D, Rosa GJM, Weigel KA. Prediction of body mass index in mice using dense molecular markers and a regularized neural network. Genet Res. 2011;93:189–201. 10.1017/S0016672310000662. [DOI] [PubMed] [Google Scholar]
  • 127.Angermueller C, Pärnamaa T, Parts L, Stegle O. Deep learning for computational biology. Mol Syst Biol. 2016;12:878. 10.15252/msb.20156651. [DOI] [PMC free article] [PubMed]
  • 128.Gianola D, Okut H, Weigel KA, Rosa GJ. Predicting complex quantitative traits with Bayesian neural networks: a case study with Jersey cows and wheat. BMC Genet. 2011;12:87. 10.1186/1471-2156-12-87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Meng Y, Zhan J, Li K, Yan F, Zhang L. A rapid and precise algorithm for maize leaf disease detection based on YOLO MSM. Sci Rep. 2025;15:6016. 10.1038/s41598-025-88399-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Brito Lopes F, Magnabosco CU, Passafaro TL, Brunes LC, Costa MFO, Eifert EC, et al. Improving genomic prediction accuracy for meat tenderness in Nellore cattle using artificial neural networks. J Anim Breed Genet. 2020;137:438–48. 10.1111/jbg.12468. [DOI] [PubMed] [Google Scholar]
  • 131.Shahinfar S, Al-Mamun HA, Park B, Kim S, Gondro C. Prediction of marbling score and carcass traits in Korean Hanwoo beef cattle using machine learning methods and synthetic minority oversampling technique. Meat Sci. 2020;161:107997. 10.1016/j.meatsci.2019.107997. [DOI] [PubMed] [Google Scholar]
  • 132.Azodi CB, Bolger E, McCarren A, Roantree M, De Los Campos G, Shiu SH. Benchmarking parametric and machine learning models for genomic prediction of complex traits. G3 (Bethesda). 2019;9:3691–702. 10.1534/g3.119.400498. [DOI] [PMC free article] [PubMed]
  • 133.Chateigner A, Lesage-Descauses MC, Rogier O, Jorge V, Leplé JC, Brunaud V, et al. Gene expression predictions and networks in natural populations supports the omnigenic theory. BMC Genomics. 2020;21:416. 10.1186/s12864-020-06809-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.Pérez-Enciso M, Zingaretti LM. A guide for using deep learning for complex trait genomic prediction. Genes. 2019;10:553. 10.3390/genes10070553. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Montesinos-López OA, Montesinos-López A, Pérez-Rodríguez P, Barrón-López JA, Martini JWR, Fajardo-Flores SB, et al. A review of deep learning applications for genomic selection. BMC Genomics. 2021;22:19. 10.1186/s12864-020-07319-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Han J, Gondro C, Reid K, Steibel JP. Heuristic hyperparameter optimization of deep learning models for genomic prediction. G3 (Bethesda). 2021;11:jkab032. 10.1093/g3journal/jkab032. [DOI] [PMC free article] [PubMed]
  • 137.Peters SO, Sinecen M, Kizilkaya K, Thomas MG. Genomic prediction with different heritability, QTL, and SNP panel scenarios using artificial neural network. IEEE Access. 2020;8:147995–8006. 10.1109/ACCESS.2020.3015814. [Google Scholar]
  • 138.Pérez-Rodríguez P, Gianola D, Weigel KA, Rosa GJM, Crossa J. Technical note: an R package for fitting Bayesian regularized neural networks with applications in animal breeding. J Anim Sci. 2013;91:3522–31. 10.2527/jas.2012-6162. [DOI] [PubMed]
  • 139.Waldmann P. Approximate Bayesian neural networks in genomic prediction. Genet Sel Evol. 2018;50:70. 10.1186/s12711-018-0439-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Van Bergen GHH, Duenk P, Albers CA, Bijma P, Calus MPL, Wientjes YCJ, et al. Bayesian neural networks with variable selection for prediction of genotypic values. Genet Sel Evol. 2020;52:26. 10.1186/s12711-020-00544-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Waldmann P, Pfeiffer C, Mészáros G. Sparse convolutional neural networks for genome-wide prediction. Front Genet. 2020;11:25. 10.3389/fgene.2020.00025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015;521:436–44. 10.1038/nature14539. [DOI] [PubMed] [Google Scholar]
  • 143.Ali M, Benfante V, Basirinia G, Alongi P, Sperandeo A, Quattrocchi A, et al. Applications of artificial intelligence, deep learning, and machine learning to support the analysis of microscopic images of cells and tissues. J Imaging. 2025;11:59. 10.3390/jimaging11020059. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Ma W, Qiu Z, Song J, Cheng Q, Ma C. DeepGS: predicting phenotypes from genotypes using deep learning. bioRxiv. 2017. 10.1101/241414.
  • 145.Xie Z, Xu X, Li L, Wu C, Ma Y, He J, et al. Residual networks without pooling layers improve the accuracy of genomic predictions. Theor Appl Genet. 2024;137:138. 10.1007/s00122-024-04649-2. [DOI] [PubMed] [Google Scholar]
  • 146.Fu Y, Xu J, Tang Z, Wang L, Yin D, Fan Y, et al. A gene prioritization method based on a swine multi-omics knowledgebase and a deep learning model. Commun Biol. 2020;3:502. 10.1038/s42003-020-01233-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Yang B, Xu Y. Applications of deep-learning approaches in horticultural research: a review. Hortic Res. 2021;8:123. 10.1038/s41438-021-00560-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Wang Z, Liao W, Jin Y, Wang Z. Performance guarantees of recurrent neural networks for the subset sum problem. Biomimetics. 2025;10:231. 10.3390/biomimetics10040231. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Wu P, Wang J, Chen X, Wang T, Guo Z, Diao S, et al. Optimization of genomic breeding value prediction for growth traits in Rongchang pigs through machine learning techniques. Mach Learn Appl. 2025;22:100747. 10.1016/j.mlwa.2025.100747. [Google Scholar]
  • 150.Su R, Lv J, Xue Y, Jiang S, Zhou L, Jiang L, et al. Genomic selection in pig breeding: comparative analysis of machine learning algorithms. Genet Sel Evol. 2025;57:13. 10.1186/s12711-025-00957-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Srivastava S, Lopez BI, Kumar H, Jang M, Chai HH, Park W, et al. Prediction of hanwoo cattle phenotypes from genotypes using machine learning methods. Animals. 2021;11:2066. 10.3390/ani11072066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Nazzicari N, Biscarini F. Stacked kinship CNN vs. GBLUP for genomic predictions of additive and complex continuous phenotypes. Sci Rep. 2022;12:19889. 10.1038/s41598-022-24405-0. [DOI] [PMC free article] [PubMed]
  • 153.Wang X, Shi S, Ali Khan MdY, Zhang Z, Zhang Y. Improving the accuracy of genomic prediction in dairy cattle using the biologically annotated neural networks framework. J Anim Sci Biotechnol. 2024;15:87. 10.1186/s40104-024-01044-1. [DOI] [PMC free article] [PubMed]
  • 154.Pedrosa VB, Chen SY, Gloria LS, Doucette JS, Boerman JP, Rosa GJM, et al. Machine learning methods for genomic prediction of cow behavioral traits measured by automatic milking systems in North American Holstein cattle. J Dairy Sci. 2024;107:4758–71. 10.3168/jds.2023-24082. [DOI] [PubMed] [Google Scholar]
  • 155.Mota LFM, Arikawa LM, Santos SWB, Fernandes Júnior GA, Alves AAC, Rosa GJM, et al. Benchmarking machine learning and parametric methods for genomic prediction of feed efficiency-related traits in Nellore cattle. Sci Rep. 2024;14:6404. 10.1038/s41598-024-57234-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.Hay EH. Machine learning for the genomic prediction of growth traits in a composite beef cattle population. Animals. 2024;14:3014. 10.3390/ani14203014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157.Schwarz L, Heise J, Bennewitz J, Thaller G, Tetens J. Genomic assessment of reproduction traits in Holstein dairy cattle across 3 lactations using additive genetic models and post hoc random forest analysis. J Dairy Sci. 2026;109:1647–64. 10.3168/jds.2025-26432. [DOI] [PubMed]
  • 158.Liu J, Yan X, Li W, Xue SH, Wang Z, Su R. Genomic selection for cashmere traits in inner mongolian cashmere goats using random forest, gradient boosting decision tree, extreme gradient boosting and light gradient boosting machine methods. Animals. 2025;15:2940. 10.3390/ani15202940. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 159.Chai H, Yang Y, Wang D, Ning C, Zhang X, Wang W, et al. Comparative accuracy of machine learning and GBLUP for predicting genomic estimated breeding values in chickens. Genes. 2026;17:315. 10.3390/genes17030315. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160.Yang J, Calus MPL, Wientjes YCJ, Meuwissen THE, Duenk P. Incorporating information of causal variants in genomic prediction using GBLUP or machine learning models in a simulated livestock population. J Anim Sci Biotechnol. 2025;16:118. 10.1186/s40104-025-01250-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 161.Ali A, Khan Z, Du H, Aldahmani S. Double weighted k nearest neighbours for binary classification of high dimensional genomic data. Sci Rep. 2025;15:12681. 10.1038/s41598-025-97505-2. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

No datasets were generated or analysed during the current study.


Articles from Journal of Animal Science and Biotechnology are provided here courtesy of BMC

RESOURCES