Skip to main content
Translational Animal Science logoLink to Translational Animal Science
. 2025 Jul 12;9:txaf095. doi: 10.1093/tas/txaf095

Improved genomic prediction accuracy by genetic relatedness using a crossbred pig population

Euiseo Hong 1,#, Yoonji Chung 2,#, Suyeon Maeng 3, In-Cheol Cho 4, Seung Hwan Lee 5,
PMCID: PMC12342970  PMID: 40799624

Abstract

Genomic prediction is crucial in animal breeding because it facilitates the selection of superior individuals based on genotype data. The success of genomic prediction is determined by its accuracy, which depends on the size of the reference population and relatedness between the reference and test populations. However, not all populations have large, highly genetically related reference populations. In this study, we evaluated the genomic prediction accuracy of three crossbreds and seven purebred populations using crossbred animals as a reference population and determined whether crossbred could be used as a reference population for small purebred populations. Genomic prediction accuracy was assessed using the genomic best linear unbiased prediction (GBLUP) for backfat thickness and carcass weight traits. Data from 29 Bisaro, 91 Duroc, 50 Duroc × Korean Native Pig (DK), 36 Iberian, 34 Korean Native Pig (KNP), 85 Landrace, 50 Landrace × Korean Native Pig (LK), 50 Landrace × Yorkshire × Duroc (LYD), 37 Meishan, and 49 Yorkshire pigs were used as test populations, whereas data from 245 DK, 964 LK, and 967 LYD crossbreds were used as the reference population. The findings indicated that the prediction accuracy of purebreds was higher when they were genetically related to the crossbred population, with accuracies ranging from 0.36 to 0.53 for backfat thickness and from 0.26 to 0.46 for carcass weight. In contrast, unrelated breeds showed lower accuracies, ranging from 0.16 to 0.48 for backfat thickness and from 0.13 to 0.40 for carcass weight. These results suggest that using crossbred populations related to the purebred population being predicted can improve prediction accuracy, especially for breeds with limited data. The prediction accuracy increased as the size of the reference population increased, regardless of genetic relatedness. Notably, small reference populations yielded higher accuracy when they were genetically related to the target animals, underscoring the importance of genetic similarity in addition to population size. These results highlight that using crossbred animals for reference populations is advantageous for genomic predictions because large populations can be rapidly established.

Keywords: crossbreed reference population, genomic prediction, genomic relatedness, reference size


Using crossbred as a reference population for genomic prediction is a valuable approach for small purebred populations, with prediction accuracy increasing proportionally with genetic similarity between the reference and test populations.

INTRODUCTION

Genomic prediction is a transformative tool in animal and plant breeding because it can identify genetically superior individuals as genotype data become available (Meuwissen et al., 2001). The success of genomic prediction is determined by its accuracy, which indicates the reliability of predicting the future phenotype of a test individual. The reference population size and relationship between the reference and test populations play crucial roles in accuracy (Goddard, 2009; Lee et al., 2017).

However, obtaining a sufficiently large reference population with both phenotypic and genotypic data is challenging for many traits and breeds (VanRaden et al., 2009). To address this limitation, alternative strategies such as across-breed prediction, which uses single nucleotide polymorphism (SNP) effects from other breeds with large reference populations, and multi-breed prediction, which combines data from multiple breeds, have been explored. Van den Berg et al. (2019) found that across-breed predictions are less accurate than within-breed predictions because of the lower genetic relatedness between populations. This result is consistent with those of other studies showing that greater genomic relationships between the reference and test populations improve genomic prediction accuracy (Lee et al., 2008). Olson et al. (2012) also found that the prediction accuracy obtained when using SNP effects from other breeds was low and adding data from other breeds did not significantly improve accuracy for large populations such as Holstein and Jersey cows. However, they also observed that information from other breeds could compensate for small population sizes, such as in the Brown Swiss breed, where across-breed predictions showed improved accuracy.

In the swine industry, production animals are generally three-way or four-way crosses designed to capture heterosis and breed complementarity (Sellier, 1976; Wakchaure et al., 2015). These breeding strategies generate large populations of crossbred animals, which integrate the genetic diversity of multiple purebred lines. In such scenarios, using large-scale crossbreed data as a reference population can significantly enhance the accuracy of genomic predictions. Moreover, crossbred reference populations offer a cost-effective alternative because genotyping purebreds from multiple lines is expensive and resource intensive. By genotyping crossbred animals, a single reference population can be used to simultaneously estimate breeding values across multiple purebred lines. For example, instead of genotyping 2,000 animals from each of the two purebred lines, 4,000 crossbred animals can be genotyped as the reference for both lines. This approach reduces costs and improves prediction accuracy by utilizing a larger reference population and minimizing the bias in genetic variance caused by small sample sizes (Wientjes et al., 2020). In addition to genotyping, the availability of phenotypic data is a critical factor influencing the accuracy of genomic prediction. The decision to collect phenotypes from purebred or crossbred animals can affect the predictive power, especially when the genetic correlations between purebred and crossbred populations are imperfect. While this limitation may reduce the accuracy of predictions based on SNP effects derived from genetically different populations, crossbred animals remain a valuable source of data in cases where phenotyping purebred animals is challenging.

However, the use of crossbred reference populations has some limitations. A key challenge is the low genetic relatedness between crossbred reference populations and purebred selection candidates. In purebred reference populations, reference animals are often from the same generation as the selection candidates, which results in strong genetic relationships. By contrast, crossbred reference populations typically share their closest genetic relationships with purebred reference animals at the grandparent level, which weakens their genetic relatedness (Rothschild and Ruvinsky, 2011). This issue is particularly pronounced in production systems with multiple hierarchical levels where genetic relatedness between the reference and test populations is further diminished (Wientjes et al., 2020). Despite these limitations, crossbred populations remain a practical alternative for constructing large reference populations, particularly for breeds with small purebred populations.

The genetic relatedness between reference and test populations is pivotal for accurate genomic prediction (Clark et al., 2012; Pszczola et al., 2012). Higher genetic relatedness ensures a better alignment of linkage disequilibrium (LD) patterns between populations, leading to more accurate predictions (De Roos et al., 2009). However, prediction accuracy may decline when genetic relatedness is limited, as in the case of genetically distant reference populations. In such cases, increasing the size of the reference population becomes essential. Although crossbred animals often exhibit lower genetic relatedness than purebred animals used as selection candidates, they can compensate for this limitation by significantly expanding the reference population and capturing a broader spectrum of genetic diversity (VanRaden et al., 2011). This strategy helps reduce the loss of accuracy associated with low genetic relatedness, making crossbred a practical resource for genomic prediction in breeds with limited data.

In this study, we compared the accuracy of genomic predictions using three different populations with different genetic relatedness to evaluate the application scope of crossbred populations as reference populations and their genomic prediction performance. Furthermore, the relationship between genetic relatedness and prediction accuracy was examined, and the effects of genetic relatedness and reference population size were confirmed.

MATERIALS AND METHODS

Approval from the Animal Care and Use Committee was not needed for this study because it used previously collected data obtained through standard farm management practices.

Data

Genotyped crossbred animals included 295 Duroc × Korean Native Pig (DK), 1,014 Landrace × Korean Native Pig (LK), and 1,017 Landrace × Yorkshire × Duroc (LYD) pigs. Genotyped purebred animals included 29 Bisaro, 91 Duroc, 36 Iberian, 34 Korean native pigs (KNP), 85 Landrace, 37 Meishan, and 49 Yorkshire pigs. Genotypic data were obtained using the Illumina PorcineSNP60 Genotyping BeadChip. Genotype data for the Bisaro, Iberian, and Meishan breeds were obtained from previous studies (Burgos-Paz et al., 2013; Yang et al., 2017). Among the purebreds, data for KNP were provided by the National Institute of Animal Science, Korea.

The reference population consisted of three crossbred groups (DK, LK, and LYD), and the test population was divided into three categories based on their genetic relatedness to the reference population: 1) breeds identical to the reference population (DK, LK, and LYD), 2) breeds related to the reference population (Duroc, KNP, Landrace, and Yorkshire), and 3) breeds unrelated to the reference population (Bisaro, Iberian, and Meishan). For the test population, 50 animals were randomly selected from each of the DK, LK, and LYD breeds (n = 150 total), and these individuals were not included in the reference population. The related group included 91 Duroc, 34 KNP, 85 Landrace, and 49 Yorkshire, and the unrelated group included 29 Bisaro, 36 Iberian, and 37 Meishan pigs.

The number of genotyped animals in each group is summarized in Table 1. To assess the prediction accuracy based on the reference population size, reference populations were constructed by randomly selecting 500, 1,000, 1,500, and 2,000 animals from the entire reference population.

Table 1.

Number of genotyped animals

Population Breed Number of animals
Reference population Duroc × Korean Native Pig (DK) 245
Landrace × Korean Native Pig (LK) 964
Landrace × Yorkshire × Duroc (LYD) 967
Population Group Breed Number of animals
Test population Same Breed DK 50
LK 50
LYD 50
Related Breed Duroc 91
Korean Native Pig (KNP) 34
Landrace 85
Yorkshire 49
Unrelated Breed Bisaro 29
Iberian 36
Meishan 37

In the quality control process, the following exclusion criteria were applied: SNPs located on sex chromosomes, genotype call rates less than 90%, and minor allele frequencies below 1%. After merging all datasets and performing quality control, a common set of 22,783 SNPs remained. The phenotypic data used in this study included backfat thickness and carcass weight. Phenotypic statistics are presented in Supplementary Table 1.

Principal Component Analysis

The study utilized principal component analysis (PCA) to investigate the genetic distances between breeds. This method simplifies the data while preserving the relationships between the breeds. Applying PCA to biallelic genotype data identifies the eigenvalues and eigenvectors of the allele frequency covariance matrix, which reduces the data to a small number of dimensions termed principal components (PCs). Each PC captures the proportion of genomic variation. The data are then projected onto the space defined by these PC axes, allowing the visualization of samples and their distances from each other in a scatter plot. In the visualizations, overlapping samples indicate a shared genetic identity because of common ancestry or origin (Patterson et al., 2006).

Genomic Prediction

The most straightforward approach for genomic prediction is applying the genomic best linear unbiased prediction (GBLUP) (VanRaden, 2008). The following GBLUP model was fitted for genomic prediction using MTG2 v2.22 software (Lee and Van der Werf, 2016):

y=Xb+Zu+e,

where y is a vector of phenotypes; X and Z are design matrices allocating phenotypes to vectors b and u, which include fixed effects (sex and breed) and genomic breeding value, respectively; and e is a vector of residual errors distributed as N(0,Iσe2), with identity matrix I and error variance σe2. Additional fixed effects, such as farm or batch, and random effects, such as litter or pen, were excluded from the model because the necessary information was unavailable. The absence of additional fixed effects may have reduced the model’s ability to account for environmental variance, which could result in lower prediction accuracy due to the confounding of environmental effects with genetic values (Van Bebber et al., 1997). Genomic breeding values were distributed as N(0,Gσg2), where G is a genomic relationship matrix and σg2 is the genetic variance explained by genomic variants. Variance components were estimated using all animals with available phenotypes and these estimates were then fixed across all predictions. The heritability of the traits was estimated as h2=σg2^/(σg2^+σe2^). The genomic relationship matrix was constructed using GCTA v1.94.1 software with the following equation (Yang et al., 2011):

Gjk=1Ni=1N(xij2pi)(xik2pi)2pi(1pi),

where xij and xik are the genotypes (coded as 0, 1, or 2) of individuals j and k at SNP i; pi is the allele frequency of SNP i; and N is the total number of SNPs.

Realized accuracy, which is commonly used to evaluate the performance of genomic predictions, was not assessed in this study owing to the absence of phenotype records for the purebred animals in the test population. Therefore, prediction accuracy was evaluated as the theoretical accuracy, which was calculated using the prediction error variance (PEV) as follows (Misztal and Wiggans, 1988):

1PEVσg2.

The variance components derived from all phenotyped animals were also used in this calculation. While theoretical accuracy is a valid and widely used metric, it represents an expected correlation between the genomic estimated breeding value (GEBV) and true breeding value under model assumptions such as the correct specification of fixed and random effects, accurate estimation of variance components and does not account for real-world factors such as phenotypic measurement error or genotype-by-environment interactions. Thus, it may slightly overestimate actual predictive performance (Dekkers et al., 2021).

Effective Number of Chromosome Segments

The accuracy of genomic prediction depends on several factors, including the genetic variance captured relative to the total variance, sample size of the reference population, and effective number of chromosome segments (Me). Me is a critical factor for estimating independent parameters. As the number of independent chromosome segments decreases, the number of independent parameters that must be estimated also decreases.

Me can be empirically obtained using a genomic relationship matrix expressed as follows:

Me=1var(Gij),

where Gij is the genomic relationship between individual i from the test population and j from the reference population. In this study, Me was estimated using MTG2 v2.22 software (Lee and Van der Werf, 2016).

Phylogenetic Analysis

A phylogenetic tree was constructed to visualize genetic relationships among breeds based on SNP data. Pairwise genetic distances were calculated using Euclidean distance, and a neighbor-joining tree was constructed using the “ape” package (v5.8) in R (Paradis and Schliep, 2019). In addition, TreeMix (v1.13) was used to infer phylogenetic structure among breeds based on allele frequency data. (Pickrell and Pritchard, 2012).

RESULTS

Principal Component Analysis

PCA was performed to investigate the genetic distance among the investigated pig breeds. The analysis indicated that the first, second, and third PCs (PC1, PC2, and PC3, respectively) accounted for 37.5%, 13.9%, and 7.7% of the total genetic variance, respectively (Figure 1). The PCA plot highlighted distinct genetic clustering among the different populations, with LYD displaying a broader spread, reflecting the wide genetic variation within this breed, which spanned from the maternal breeds (Landrace and Yorkshire) to the paternal breed (Duroc). By contrast, the LK breed was positioned between the Landrace and KNP breeds, whereas the DK breed was placed between the Duroc and KNP breeds (Figure 1-A). Notably, the Bisaro and Iberian breeds were slightly separated, and the Meishan breed was distinct from all other populations (Figure 1-B).

Figure 1.

Figure 1.

Principal component analysis of breeds. A) The first and second principal components of each population, B) the second and third principal components of each population.

Genetic Parameter Estimates

Heritability was estimated using only animals with both genotype and phenotype records, as phenotype records were not available for all genotyped individuals. The results are presented in Table 2. The heritability estimates were moderate, with backfat thickness having a heritability of 0.42 ± 0.03 and carcass weight having a heritability of 0.31 ± 0.03. These estimates suggest that both traits have a substantial genetic component, with backfat thickness exhibiting a slightly higher heritability.

Table 2.

Variance components and heritability estimates

Trait Genetic variance Residual variance Heritability
Backfat thickness 13.69 ± 1.36 18.91 ± 0.82 0.42 ± 0.03
Carcass weight 31.39 ± 3.59 69.13 ± 2.71 0.31 ± 0.03

Genomic Prediction Accuracy by Breed

The accuracy of the GEBV was evaluated across different test populations (Figure 2 and Supplementary Table 2). The analysis showed that the accuracy of GEBV was the highest when the test population consisted of individuals from the same breed as the reference population, followed by genetically related breeds, and the lowest for genetically unrelated breeds.

Figure 2.

Figure 2.

Genomic prediction accuracy by breed using a crossbred reference population combining DK, LK, and LYD. A) Accuracy for backfat thickness, and B) accuracy for carcass weight. The red dotted line indicates the mean prediction accuracy of the genetic relatedness group.

Specifically, the Same Breed group exhibited a mean accuracy of 0.47 for backfat thickness and 0.42 for carcass weight, highlighting the importance of using the same breed for reference and test populations to achieve the best prediction accuracy. By contrast, the accuracy decreased when the test population was a related breed but not the same as the reference population (Related Breed). In this group, the mean accuracy was 0.42 for backfat thickness and 0.37 for carcass weight, reflecting a decline in prediction accuracy owing to genetic differences between the reference and test populations. The lowest accuracy was observed when the breeds of the reference and test populations were completely unrelated (Unrelated Breed). In this group, the mean accuracy was 0.28 for backfat thickness and 0.23 for carcass weight, showing that genetic differences between populations can significantly reduce the accuracy of GEBV predictions. These results highlight the critical role of genetic similarity between the reference and test populations in achieving high GEBV accuracy.

While the overall accuracy followed the genetic similarity between the reference and test populations, notable differences were observed in prediction accuracy across breeds within each group. In the Same Breed group, LK showed the highest mean accuracy for both backfat thickness (0.52) and carcass weight (0.47), followed by DK (0.49 and 0.43) and LYD (0.41 and 0.36), which showed the largest variability. In the Related Breed group, KNP and Duroc had relatively high accuracies, with backfat thickness accuracies of 0.53 and 0.51 and carcass weight accuracies of 0.45 and 0.46, respectively. By contrast, Landrace and Yorkshire showed lower accuracies, with backfat thickness accuracies of 0.30 and 0.36 and carcass weight accuracies of 0.26 and 0.31, respectively. In the Unrelated Breed group, Meishan showed relatively higher accuracies than the other breeds in this group, with a backfat thickness accuracy of 0.48 and carcass weight accuracy of 0.40. By contrast, Bisaro and Iberian exhibited the lowest accuracies, with backfat thickness accuracy of 0.17 and 0.16 and carcass weight accuracies of 0.14 and 0.13, respectively.

Effect of Genetic Relatedness on Prediction Accuracy

The relationship between the GEBV accuracy and genetic relatedness between the reference and test populations was analyzed using the effective number of chromosome segments (Me) (Figure 3 and Figure 4; Supplementary Table 3). A lower Me indicates a closer genetic relationship with the reference population, and the results showed that a lower Me corresponded to a higher GEBV accuracy.

Figure 3.

Figure 3.

Relationship between accuracy and the effective number of chromosome segments ( Me ) by genetic relatedness group. A) Accuracy for backfat thickness, and B) Accuracy for carcass weight. The dotted lines represent the mean of each value. Each dot represents an individual.

Figure 4.

Figure 4.

Relationship between accuracy and the effective number of chromosome segments ( Me ) by breed. A) Accuracy for backfat thickness, and B) Accuracy for carcass weight. The dotted lines represent the mean of each value. Each dot represents an individual.

Specifically, the Same Breed group, which had the lowest mean Me of 79.71, demonstrated the highest mean accuracy of 0.57 for backfat thickness and 0.51 for carcass weight. Conversely, the Related Breed group exhibited a higher mean Me of 131.49 and correspondingly lower accuracy of 0.46 for backfat thickness and 0.42 for carcass weight, whereas the Unrelated Breed group had the highest mean Me of 552.9 and lowest accuracy of 0.3 for backfat thickness and 0.25 for carcass weight (Figure 3). These findings underscore the critical role of genetic relatedness, as measured by Me, in determining the accuracy of GEBV predictions.

Within the Same Breed group, DK had the lowest mean Me (68.78), followed by LK (74.57) and LYD (95.80). Despite having a slightly higher Me than DK, LK showed the highest accuracy among all breeds, with 0.52 for backfat thickness and 0.47 for carcass weight. DK also showed high accuracies of 0.49 and 0.43, whereas LYD exhibited lower accuracies of 0.41 and 0.36. Among Related Breeds, Duroc and KNP had relatively low Me values (48.27 and 86.85) and showed high prediction accuracies. By contrast, Landrace and Yorkshire had higher Me values (235.71 and 136.25) and lower accuracies, with 0.30 and 0.36 for backfat thickness and 0.26 and 0.31 for carcass weight. Within the Unrelated Breed group, Meishan exhibited a high mean Me (597.70) but achieved relatively high prediction accuracies, with 0.50 for backfat thickness and 0.40 for carcass weight. By contrast, Bisaro and Iberian, which had mean Me values of 378.82 and 647.08, respectively, showed substantially lower accuracies, with 0.17 and 0.16 for backfat thickness and 0.14 and 0.13 for carcass weight (Figure 4).

Effect of Reference Population Size on Prediction Accuracy

To determine the effect of reference population size on genomic prediction, we analyzed the relationship between the reference population size and GEBV accuracy based on the genetic relatedness between the reference and test populations (Figure 5). The analysis demonstrated that genomic prediction accuracy increased as the reference population grew. This trend was consistently observed across all genetically related groups.

Figure 5.

Figure 5.

Mean Accuracy by genetic relatedness group according to reference size. A) Accuracy for backfat thickness, and B) Accuracy for carcass weight

Notably, the results also indicated that genetic relatedness plays a more critical role than the reference population size alone. For example, for both backfat thickness and carcass weight, using 1,000 reference animals from genetically related breeds yielded higher prediction accuracies than those when using 2,000 animals from unrelated breeds. Specifically, for backfat thickness, the accuracy with 1,000 animals from the Related Breed group was 0.31, whereas it was 0.26 with 2,000 animals from the Unrelated Breed group. A similar trend was observed for carcass weight, where the accuracy was 0.27 for the Related Breed group (1,000 animals), compared to 0.21 for the Unrelated Breed group (2,000 animals) (Supplementary Table 4).

For backfat thickness, the Same Breed group showed a steady increase in accuracy from 0.26 with 500 reference animals to 0.36 with 1,000, 0.42 with 1,500, 0.46 with 2,000, and 0.47 with the full set of 2,176 animals. The Related Breed group followed a similar trend, with accuracies increasing from 0.22 to 0.31, 0.37, 0.41, and finally 0.42. In the Unrelated Breed group, the initial accuracy was 0.10 with 500 animals, increasing to 0.16, 0.22, 0.26, and 0.28 as the reference population expanded. A similar pattern was observed for carcass weight. In the Same Breed group, accuracy increased from 0.22 with 500 animals to 0.31, 0.36, 0.41, and 0.42 with the full reference set. The Related Breed group showed improvement from 0.19 to 0.27, 0.32, 0.36, and 0.37, while the Unrelated Breed group started at 0.08 and increased to 0.13, 0.18, 0.21, and 0.23 across the same reference size increments. These results demonstrate that increasing the size of the reference population leads to high prediction accuracies for both traits. Moreover, the magnitude of improvement was high in genetically related groups, underscoring the combined importance of reference population size and genetic relatedness in enhancing genomic prediction accuracy.

DISCUSSION

In this study, we explored the potential use of crossbred as a reference population to improve the accuracy of genomic prediction of small purebred populations. We compared the prediction accuracy at different levels of genetic relatedness using crossbred as a reference population and explored the relationship between genetic relatedness and prediction accuracy. We also analyzed how the size of the reference population affected the accuracy of genomic prediction, depending on the genetic relatedness. Prediction accuracy increased with a higher genetic relatedness between the reference and test populations. Genetic relatedness was higher when purebreds were included among the crossbred pairs, indicating the potential use of crossbred reference populations in genomic prediction. Moreover, the accuracy increased as the size of the reference population increased, regardless of genetic relatedness, and the magnitude of improvement was greater in groups with higher genetic relatedness to the reference population. These findings highlight that using a large number of crossbred animals for genomic prediction can improve prediction accuracy compared with that when using a small number of purebreds as the reference population in cases when phenotypic and genotype data from purebreds are relatively difficult to obtain.

Using the SNP effects of crossbred to predict the breeding values of animals unrelated to the crossbred population resulted in a much lower accuracy than that when using the breeding values of related animals. This result is similar to the findings of Steyn et al. (2019), who simulated multiple breeds assuming the same quantitative trait locus (QTL) effects and found poor prediction accuracy for across-breed predictions. Other studies on cattle have also shown that using data from one breed to predict breeding values for another breed decreases accuracy (Olson et al., 2012; Kachman et al., 2013). This may be attributed to the low genetic relatedness between the reference and test populations. As shown in Figure 3, accuracy decreased as the genetic relatedness with the reference population decreased. Other studies have reported that low genetic relatedness between populations decreases accuracy and increases bias (Pszczola et al., 2012; Fraslin et al., 2022). For genomic predictions, both the relationship between the reference and test populations and LD between SNPs and QTLs affect prediction accuracy (Habier et al., 2010; Wientjes et al., 2013). Owing to differences in LD patterns among breeds, the accuracy of across-breed predictions is inevitably lower than that of within-breed predictions (Veroneze et al., 2014; Raymond et al., 2018).

Differences were found in prediction accuracy within groups. In the Same Breed group, the accuracy for LYD was lower than that for DK and LK. This could be attributed to the genetic diversity of LYD, as shown in Figure 1. The high genetic diversity of the population may reduce the genetic relatedness between the reference and test populations, increasing the likelihood that information from the reference population does not accurately reflect genetic variation in the test population. In other words, the accuracy was likely lower because of low genetic relatedness. (Lee et al., 2008). Within the Related Breed group, the accuracy for the Duroc and KNP breeds was higher than that for the Landrace and Yorkshire breeds. This difference was because of the different Me values, which represent the genetic relatedness between the reference and test populations across these breeds. Duroc had a mean Me of 48.27, and KNP had a mean Me of 86.85. By contrast, Landrace had a mean Me of 235.71, and Yorkshire had a mean Me of 136.25. Genomic prediction accuracy was thus highly correlated with Me, with a correlation coefficient of −0.56. When Meishan, which showed high genomic prediction accuracy despite an unusually high Me, was excluded, the correlation coefficient increased to −0.78. This indicates that Me is directly negatively correlated with genomic prediction accuracy.

Although the Meishan breed had a low genetic relatedness with the reference population (Figure 4) and was genetically distant from the reference population (Figure 1), it exhibited abnormally high prediction accuracy compared with that of the other breeds with a low genetic relatedness, such as the Bisaro and Iberian. In this study, Me was used as an indicator of genetic relatedness, calculated from the variance of the genomic relationship matrix. This variance corresponds to the average r2 value of LD across genome-wide SNPs, providing insight into the degree of relatedness among individuals (Goddard et al., 2011). In other words, Me indirectly represents the LD structure within the population. By contrast, GBLUP estimates breeding values based on the genetic similarity between individuals, which depends on the presence of shared genetic variants rather than LD patterns. This suggests that although compared to the reference population, the Meishan breed has a distinct LD structure, it shares a greater number of common genetic variants with the reference population than with the Bisaro or Iberian breeds, highlighting that SNP genetic variants, rather than LD patterns, play a significant role in GBLUP predictions. This observation aligns with the phylogenetic analysis results (Supplementary Fig. 1 and 2), which showed that Meishan shares a common ancestor with KNP. As KNP is a parental breed of the DK and LK populations in the reference dataset, the genetic architecture of Meishan is likely to be more similar to that of these reference populations. The presence of shared genetic variants between the reference population and Meishan breed could account for its high prediction accuracy despite the low genetic relatedness. In addition, breeds such as Meishan, Iberian and Bisaro had relatively small sample sizes like 29 to 37 animals in the test set, which may have contributed to their high accuracy variability, resulting in large accuracy differences between them. Moreover, the limited sample size can cause sampling error, meaning that the observed accuracy could be influenced by the particular subset of animals selected for testing, and that accuracy may depending on the set.

Increasing the size of the reference population is thus critical for improving the accuracy of genomic prediction. Numerous studies have consistently demonstrated this (Lee et al., 2017; Song et al., 2019; Takeda et al., 2021). A positive correlation was found between genetic relatedness and prediction accuracy, with the accuracy increasing as the reference population size increased (Figure 4). Wientjes et al. (2020) performed simulation studies and reported that accuracy is maximized as the reference population size increases, thus, showing that a large population can compensate for the shortcomings associated with low genetic relatedness within the reference population. Lee et al. (2017) and Takeda et al. (2021) demonstrated that the accuracy of genomic predictions increased sharply with the reference population size before gradually stabilizing. Lee et al. (2017) observed a steep increase in accuracy up to a reference population size of 10,000 animals, followed by a gradual improvement up to 25,000 animals. Similarly, Takeda et al. (2021) reported a sharp increase in accuracy for up to 12,000 animals in cattle. These findings underscore the differences in the impact of the reference population size, depending on the study context and breed characteristics. In this study, the maximum reference population size was limited to 2,176 animals. As a result, only the initial phase of rapid improvement in prediction accuracy could be observed, which is consistent with the early increasing trends reported in previous studies. However, we could not determine the reference population size threshold beyond which gains in prediction accuracy begin to diminish. Therefore, further studies utilizing larger reference populations are warranted to explore how genetic relatedness influences the rate of accuracy improvement and identify the population size threshold beyond which the marginal gains in accuracy decline. Thus, when securing a large reference population is challenging, crossbreed data represents a practical and effective solution for enhancing the accuracy of genomic predictions in small purebred populations.

Although genomic prediction accuracies exceeding 0.7 have been reported (Wiggans et al., 2011), such values are typically achieved under favorable conditions, including large-scale reference populations of more than 30,000 animals (Wiggans et al., 2017), high-quality phenotypic records, and strong genetic relatedness between reference and test animals (Meuwissen et al., 2016). In contrast, this study used relatively small reference populations ranging from 245 to 967 animals and limited genetic connectedness between the reference and test populations. These limitations may have contributed to the lower genomic prediction accuracies, which reached 0.52 for backfat thickness and 0.47 for carcass weight within the Same Breed group.

CONCLUSION

Genomic prediction accuracy increases with higher genetic relatedness between the reference and test populations as well as with larger reference populations. Furthermore, the results showed that genetically related reference populations yielded higher accuracy relative to genetically unrelated populations, even when the related group was smaller in size. Therefore, using crossbred animals with higher genetic relatedness as a reference population can result in more accurate predictions and enable the rapid construction of large reference populations. Using crossbred populations helps address the challenges of genomic prediction in small or underrepresented breeds, thereby supporting more robust and accurate breeding programs. However, this study was limited by the relatively small sample size, which may have introduced variability in the prediction accuracy estimates. Further research should investigate additional factors such as genotype-by-environment interactions and trait-specific heritability. Integrating advanced genomic methodologies, such as Bayesian approaches or machine learning models, may further enhance prediction accuracy and improve the effectiveness of genomic selection in diverse breeding programs.

Supplementary Material

txaf095_suppl_Supplementary_Tables_S1-S4_Figures_S1-S2

Contributor Information

Euiseo Hong, Department of Bio-Big Data and Precision Agriculture, Chungnam National University, Daejeon, 34134, Republic of Korea.

Yoonji Chung, Institute of Agricultural Science, Chungnam National University, Daejeon, 34134, Republic of Korea.

Suyeon Maeng, Division of Animal & Dairy Science, Chungnam National University, Daejeon, 34134, Republic of Korea.

In-Cheol Cho, Subtropical Livestock Research Institute, National Institute of Animal Science, Rural Development Administration, Jeju, 63242, Republic of Korea.

Seung Hwan Lee, Division of Animal & Dairy Science, Chungnam National University, Daejeon, 34134, Republic of Korea.

Author Contributions

Euiseo Hong (Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Visualization, Writing - original draft, Writing - review & editing), Yoonji Chung (Formal analysis, Methodology, Project administration, Supervision, Validation, Writing - review & editing), Suyeon Maeng (Methodology), In-Cheol Cho (Resources), and Seung Hwan Lee (Conceptualization, Project administration, Resources, Supervision, Validation, Writing - review & editing)

Funding

This research was supported by Chungnam National University (This work was supported by the National Research Foundation of Korea’s Brain Korea 21 FOUR Program).

Conflict of Interest statement

There are no conflicts of interest to declare.

Literature Cited

  1. Burgos-Paz, W., Souza C. A., Megens H. J., Ramayo-Caldas Y., Melo M., Lemús-Flores C., Caal E., Soto H. W., Martínez R., Alvarez L. A.,. et al. 2013. Porcine colonization of the Americas: a 60k SNP story. Heredity 110:321–330. doi: https://doi.org/ 10.1038/hdy.2012.109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Clark, S. A., Hickey J. M., Daetwyler H. D., and van der Werf J. H... 2012. The importance of information on relatives for the prediction of genomic breeding values and the implications for the makeup of reference data sets in livestock breeding schemes. Genet. Sel. Evol. 44:4. doi: https://doi.org/ 10.1186/1297-9686-44-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Dekkers, J. C., Su H., and Cheng J... 2021. Predicting the accuracy of genomic predictions. Genet. Sel. Evol. 53:1–23. doi: https://doi.org/ 10.1186/s12711-021-00675-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. De Roos, A., Hayes B. J., and Goddard M... 2009. Reliability of genomic predictions across multiple populations. Genetics 183:1545–1553. doi: https://doi.org/ 10.1534/genetics.109.104935 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Fraslin, C., Yáñez J. M., Robledo D., and Houston R. D... 2022. The impact of genetic relationship between training and validation populations on genomic prediction accuracy in Atlantic salmon. Aquacult. Rep. 23:101033. doi: https://doi.org/ 10.1016/j.aqrep.2022.101033 [DOI] [Google Scholar]
  6. Goddard, M. 2009. Genomic selection: prediction of accuracy and maximisation of long term response. Genetica 136:245–257. doi: https://doi.org/ 10.1007/s10709-008-9308-0 [DOI] [PubMed] [Google Scholar]
  7. Goddard, M. E., Hayes B. J., and Meuwissen T. H... 2011. Using the genomic relationship matrix to predict the accuracy of genomic selection. J Anim. Breed. Genet. 128:409–421. doi: https://doi.org/ 10.1111/j.1439-0388.2011.00964.x [DOI] [PubMed] [Google Scholar]
  8. Habier, D., Tetens J., Seefried F. -R., Lichtner P., and Thaller G... 2010. The impact of genetic relationship information on genomic breeding values in German Holstein cattle. Genet. Sel. Evol. 42:1–12. doi: https://doi.org/ 10.1186/1297-9686-42-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Kachman, S. D., Spangler M. L., Bennett G. L., Hanford K. J., Kuehn L. A., Snelling W. M., Thallman R. M., Saatchi M., Garrick D. J., Schnabel R. D.,. et al. 2013. Comparison of molecular breeding values based on within-and across-breed training in beef cattle. Genet. Sel. Evol. 45:1–9. doi: https://doi.org/ 10.1186/1297-9686-45-30 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Lee, S. H., Clark S., and Van Der Werf J. H... 2017. Estimation of genomic prediction accuracy from reference populations with varying degrees of relationship. PLoS One 12:e0189775. doi: https://doi.org/ 10.1371/journal.pone.0189775 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Lee, S. H., and Van der Werf J. H... 2016. MTG2: an efficient algorithm for multivariate linear mixed model analysis based on genomic information. Bioinformatics 32:1420–1422. doi: https://doi.org/ 10.1093/bioinformatics/btw012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Lee, S. H., Van Der Werf J. H., Hayes B. J., Goddard M. E., and Visscher P. M... 2008. Predicting unobserved phenotypes for complex traits from whole-genome SNP data. PLoS Genet. 4:e1000231. doi: https://doi.org/ 10.1371/journal.pgen.1000231 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Meuwissen, T., Hayes B., and Goddard M... 2016. Genomic selection: a paradigm shift in animal breeding. Anim Front 6:6–14. doi: https://doi.org/ 10.2527/af.2016-0002 [DOI] [Google Scholar]
  14. Meuwissen, T. H., Hayes B. J., and Goddard M... 2001. Prediction of total genetic value using genome-wide dense marker maps. Genetics 157:1819–1829. doi: https://doi.org/ 10.1093/genetics/157.4.1819 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Misztal, I., and Wiggans G... 1988. Approximation of prediction error variance in large-scale animal models. J. Dairy Sci. 71:27–32. doi: https://doi.org/ 10.1016/s0022-0302(88)79976-2 [DOI] [Google Scholar]
  16. Olson, K. M., VanRaden P. M., and Tooker M. E... 2012. Multibreed genomic evaluations using purebred Holsteins, Jerseys, and Brown Swiss. J. Dairy Sci. 95:5378–5383. doi: https://doi.org/ 10.3168/jds.2011-5006 [DOI] [PubMed] [Google Scholar]
  17. Paradis, E., and Schliep K... 2019. ape 5.0: an environment for modern phylogenetics and evolutionary analyses in R. Bioinformatics. 35:526–528. doi: https://doi.org/ 10.1093/bioinformatics/bty633 [DOI] [PubMed] [Google Scholar]
  18. Patterson, N., Price A. L., and Reich D... 2006. Population structure and eigenanalysis. PLoS Genet. 2:e190. doi: https://doi.org/ 10.1371/journal.pgen.0020190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Pickrell, J. K., and Pritchard J. K... 2012. Inference of population splits and mixtures from genome-wide allele frequency data. PLoS Genet. 8:e1002967. doi: https://doi.org/ 10.1371/journal.pgen.1002967 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Pszczola, M., Strabel T., Mulder H., and Calus M... 2012. Reliability of direct genomic values for animals with different relationships within and to the reference population. J. Dairy Sci. 95:389–400. doi: https://doi.org/ 10.3168/jds.2011-4338 [DOI] [PubMed] [Google Scholar]
  21. Raymond, B., Bouwman A. C., Schrooten C., Houwing-Duistermaat J., and Veerkamp R. F... 2018. Utility of whole-genome sequence data for across-breed genomic prediction. Genet. Sel. Evol. 50:1–12. doi: https://doi.org/ 10.1186/s12711-018-0396-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Rothschild, M. F., and Ruvinsky A... 2011. The genetics of the pig. CABI.
  23. Sellier, P. 1976. The basis of crossbreeding in pigs; a review. Livest. Prod. Sci. 3:203–226. doi: https://doi.org/ 10.1016/0301-6226(76)90016-6 [DOI] [Google Scholar]
  24. Song, H., Zhang J., Zhang Q., and Ding X... 2019. Using different single-step strategies to improve the efficiency of genomic prediction on body measurement traits in pig. Front. Genet. 9:730. doi: https://doi.org/ 10.3389/fgene.2018.00730 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Steyn, Y., Lourenco D. A., and Misztal I... 2019. Genomic predictions in purebreds with a multibreed genomic relationship matrix. J. Anim. Sci. 97:4418–4427. doi: https://doi.org/ 10.1093/jas/skz296 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Takeda, M., Inoue K., Oyama H., Uchiyama K., Yoshinari K., Sasago N., Kojima T., Kashima M., Suzuki H., Kamata T.,. et al. 2021. Exploring the size of reference population for expected accuracy of genomic prediction using simulated and real data in Japanese Black cattle. BMC Genomics 22:1–11. doi: https://doi.org/ 10.1186/s12864-021-08121-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Van Bebber, J., Reinsch N., Junge W., and Kalm E... 1997. Accounting for herd, year and season effects in genetic evaluations of dairy cattle: a review. Livest. Prod. Sci. 51:191–203. doi: https://doi.org/ 10.1016/s0301-6226(97)00058-4 [DOI] [Google Scholar]
  28. Van den Berg, I., Meuwissen T., MacLeod I., and Goddard M... 2019. Predicting the effect of reference population on the accuracy of within, across, and multibreed genomic prediction. J. Dairy Sci. 102:3155–3174. doi: https://doi.org/ 10.3168/jds.2018-15231 [DOI] [PubMed] [Google Scholar]
  29. VanRaden, P., Van Tassell C., Wiggans G., Sonstegard T., Schnabel R., Taylor J., and Schenkel F... 2009. Invited review: reliability of genomic predictions for North American Holstein bulls. J. Dairy Sci. 92:16–24. doi: https://doi.org/ 10.3168/jds.2008-1514 [DOI] [PubMed] [Google Scholar]
  30. VanRaden, P. M. 2008. Efficient methods to compute genomic predictions. J. Dairy Sci. 91:4414–4423. doi: https://doi.org/ 10.3168/jds.2007-0980 [DOI] [PubMed] [Google Scholar]
  31. VanRaden, P. M., O’Connell J. R., Wiggans G. R., and Weigel K. A... 2011. Genomic evaluations with many more genotypes. Genet. Sel. Evol. 43:1–11. doi: https://doi.org/ 10.1186/1297-9686-43-10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Veroneze, R., Bastiaansen J. W., Knol E. F., Guimarães S. E., Silva F. F., Harlizius B., Lopes M. S., and Lopes P. S... 2014. Linkage disequilibrium patterns and persistence of phase in purebred and crossbred pig (Sus scrofa) populations. BMC Genet. 15:1–9. doi: https://doi.org/ 10.1186/s12863-014-0126-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Wakchaure, R., Ganguly S., Praveen P. K., Sharma S., Kumar A., Mahajan T., and Qadri K... 2015. Importance of heterosis in animals: a review. Int. J. Adv. Eng. Technol. Innov. Sci. 1:1–5. www.ijaetis.org [Google Scholar]
  34. Wientjes, Y. C., Bijma P., and Calus M. P... 2020. Optimizing genomic reference populations to improve crossbred performance. Genet. Sel. Evol. 52:1–18. doi: https://doi.org/ 10.1186/s12711-020-00573-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Wientjes, Y. C., Veerkamp R. F., and Calus M. P... 2013. The effect of linkage disequilibrium and family relationships on the reliability of genomic prediction. Genetics. 193:621–631. doi: https://doi.org/ 10.1534/genetics.112.146290 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Wiggans, G., VanRaden P., and Cooper T... 2011. The genomic evaluation system in the United States: past, present, future. J. Dairy Sci. 94:3202–3211. doi: https://doi.org/ 10.3168/jds.2010-3866 [DOI] [PubMed] [Google Scholar]
  37. Wiggans, G. R., Cole J. B., Hubbard S. M., and Sonstegard T. S... 2017. Genomic selection in dairy cattle: the USDA experience. Annu. Rev. Anim. Biosci. 5:309–327. doi: https://doi.org/ 10.1146/annurev-animal-021815-111422 [DOI] [PubMed] [Google Scholar]
  38. Yang, B., Cui L. L., Perez-Enciso M., Traspov A., Crooijmans R. P. M. A., Zinovieva N., Schook L. B., Archibald A., Gatphayak K., Knorr C.,. et al. 2017. Genome-wide SNP data unveils the globalization of domesticated pigs. Genet. Sel. Evol. 49:71. doi: https://doi.org/ 10.1186/s12711-017-0345-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Yang, J. A., Lee S. H., Goddard M. E., and Visscher P. M... 2011. GCTA: a tool for genome-wide complex trait analysis. Am. J. Hum. Genet. 88:76–82. doi: https://doi.org/ 10.1016/j.ajhg.2010.11.011 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

txaf095_suppl_Supplementary_Tables_S1-S4_Figures_S1-S2

Articles from Translational Animal Science are provided here courtesy of Oxford University Press

RESOURCES