Skip to main content
Journal of Animal Science logoLink to Journal of Animal Science
. 2020 Jun 4;98(6):skaa184. doi: 10.1093/jas/skaa184

Genomic prediction using pooled data in a single-step genomic best linear unbiased prediction framework

Johnna L Baller 1,, Stephen D Kachman 2, Larry A Kuehn 3, Matthew L Spangler 1
PMCID: PMC7314383  PMID: 32497209

Abstract

Economically relevant traits are routinely collected within the commercial segments of the beef industry but are rarely included in genetic evaluations because of unknown pedigrees. Individual relationships could be resurrected with genomics, but this would be costly; therefore, pooling DNA and phenotypic data provide a cost-effective solution. Pedigree, phenotypic, and genomic data were simulated for a beef cattle population consisting of 15 generations. Genotypes mimicked a 50k marker panel (841 quantitative trait loci were located across the genome, approximately once per 3 Mb) and the phenotype was moderately heritable. Individuals from generation 15 were included in pools (observed genotype and phenotype were mean values of a group). Estimated breeding values (EBV) were generated from a single-step genomic best linear unbiased prediction model. The effects of pooling strategy (random and minimizing or uniformly maximizing phenotypic variation within pools), pool size (1, 2, 10, 20, 50, 100, or no data from generation 15), and generational gaps of genotyping on EBV accuracy (correlation of EBV with true breeding values) were quantified. Greatest EBV accuracies of sires and dams were observed when there was no gap between genotyped parents and pooled offspring. The EBV accuracies resulting from pools were usually greater than no data from generation 15 regardless of sire or dam genotyping. Minimizing phenotypic variation increased EBV accuracy by 8% and 9% over random pooling and uniformly maximizing phenotypic variation, respectively. A pool size of 2 was the only scenario that did not significantly decrease EBV accuracy compared with individual data when pools were formed randomly or by uniformly maximizing phenotypic variation (P > 0.05). Pool sizes of 2, 10, 20, or 50 did not generally lead to statistical differences in EBV accuracy than individual data when pools were constructed to minimize phenotypic variation (P > 0.05). Largest numerical increases in EBV accuracy resulting from pooling compared with no data from generation 15 were seen with sires with prior low EBV accuracy (those born in generation 14). Pooling of any size led to larger EBV accuracies of the pools than individual data when minimizing phenotypic variation. Resulting EBV for the pools could be used to inform management decisions of those pools. Pooled genotyping to garner commercial-level phenotypes for genetic evaluations seems plausible although differences exist depending on pool size and pool formation strategy.

Keywords: beef cattle, DNA pooling, genomic prediction

Introduction

Millions of phenotypic records are collected annually within commercial sectors of livestock industries, including commercial herds, feedlots, and abattoirs. However, most of these records are not included in genetic evaluations because of the lack of available pedigree ties between the commercial and nucleus (seedstock) animals. Examples of traits routinely recorded in commercial settings include carcass merit, disease incidence, female fertility, and growth traits. Many of these trait complexes represent economically relevant traits, those that have a direct source of revenue or cost at the commercial level. Pedigree ties inherently exist between commercial and seedstock animals, but they are often unknown due to lack of recording, group mating, or the pedigree simply does not follow an animal through the entire production system (Bell et al., 2017). Kinship ties can be resurrected through genomic relationships; but even with the decreasing cost of genotyping, it is still not economically feasible to genotype all commercial animals.

Pooling DNA for genome-wide association studies (GWAS) has been shown to reduce the cost of genotyping (Sham et al., 2002) by selectively grouping animals based on phenotype and then genotyping a combined pool of DNA (Darvasi and Soller, 1994). Many studies have identified candidate quantitative trait loci (QTL) for traits using this approach—e.g., general cognitive ability in children (Fisher et al., 1999), fertility in Holstein cattle (Huang et al., 2010), low reproductive cattle with the presence of single nucleotide polymorphism (SNP) mapped to the Y chromosome (McDaneld et al., 2012), colorectal and prostate cancer in a Polish population (Gaj et al., 2012), and somatic cell score in Valdostana Red Pied cattle (Strillacci et al., 2014). Recently, pooled data have also been used for genetic prediction within a simulated aquaculture population (Sonesson et al., 2010), Brahman and Tropical composite cattle (Henshall et al., 2012; Reverter et al., 2016), Merino sheep (Bell et al., 2017), and a simulated cattle data set (Alexandre et al., 2019). Pooling data, genotypes and thus phenotypes, not only reduce the cost of genotyping but also allow the inclusion of phenotypes that are typically only observed at the commercial level in genetic evaluations.

The aims of large-scale genetic evaluations should be to improve commercial-level phenotypes that directly impact the profitability of commercial enterprises. However, the majority of, if not all, phenotypes recorded in nucleus (seedstock) settings are indicator traits. A comprehensive genetic evaluation would combine indicator traits from nucleus animals with the target phenotypes from commercial animals, and to do so would require the use of individual and pooled data simultaneously. However, in some species (e.g., beef cattle), not all parent animals are genotyped thus necessitating the use of both pedigree and genomic kinship as in single-step genomic best linear unbiased prediction (ssGBLUP). Moreover, the estimated breeding values (EBV) of pools could themselves be used to inform management-level decisions. To our knowledge, previous literature has not investigated the accuracy of EBV of the pools themselves. Consequently, the objectives of this paper were to quantify the impact of pool size, method of assigning animals to pools, and generational gaps between the genotyped nucleus (seedstock) and commercial animals on the resulting accuracy of EBV of parents and grandparents and of the pools in a ssGBLUP framework utilizing simulation.

Materials and Methods

Animal care and use committee approval was not obtained for this study as all data were simulated.

Simulation

The simulated data used for the analysis were previously described by Baller et al. (2019). Briefly, a purebred beef cattle population was simulated using Geno-Diver (Howard et al., 2017). Five replicates were simulated, each with a different founder genome. Individuals contained 29 chromosomes, with 29 QTL per chromosome. Markers mimicked those from a 50k SNP panel and were randomly distributed across the genome. Locations of the markers and QTL were randomly drawn from separate uniform distributions. A phenotype with a heritability of 0.4 was simulated. The Markovian Coalescence Simulator (MaCS) program (Chen et al., 2009) generated a founder genome in which a large amount of short-range linkage disequilibrium was created. The founder population was assumed to have an effective population size of 70. Founder animals were randomly selected and mated for five generations in order to establish a pedigree. For an additional 10 generations, individuals were randomly mated with the caveat that individuals with an additive relationship of 0.125 or greater were not mated together. These last 10 generations were selectively replaced based on the highest EBV determined by pedigree-based BLUP with replacement rates of 0.2 and 0.4 for dams and sires, respectively. Animals remained within the breeding population until they were culled for low EBV or until they had been a parent for 12 generations.

Pooling

Individuals born in generation 15 (n = 2,000) were assigned to pools, where each individual was included in only one pool per scenario. Pool sizes included 2, 10, 20, 50, or 100 individuals, resulting in 1,000, 200, 100, 40, or 20 pools, respectively. The pool size was consistent within each scenario. Pool assignments were determined in three ways: randomly, minimize phenotypic variation within a pool, and uniformly maximize phenotypic variation within a pool. In order to construct random pools, individuals were randomly assigned a pool number where the only caveat was consistent pool sizes. To minimize phenotypic variation within pools, individuals were ranked based on phenotype and then grouped together dependent on pool size such that for pool size 2, for example, the first two ranked animals were grouped together. This resulted in individuals with the smallest phenotype in one pool while individuals with the largest phenotype were in another pool. To uniformly maximize phenotypic variation within pools, individuals were again ranked based on phenotype. Individuals with ranks i, i+r, …, i+r(q-1) were assigned to pool i, where r is the number of pools and q is the pool size. For example, when pool size was 2 and thus 1,000 pools were constructed, individuals with ranks one and 1,001 were assigned to pool one, and individuals with ranks 1,000 and 2,000 were assigned to pool 1,000. Minimizing and uniformly maximizing phenotypic variation within pools were chosen to demonstrate extreme cases of pooling strategies. Minimizing variation within pools increases variation between pools. Alternatively, maximizing variation within pools decreases variation between pools.

Once the individuals were assigned to pools, the phenotypic record for a given pool was determined as the average of the individuals contributing to the pool. Pooling allele frequency (PAF) for each SNP is based on the normalized intensity of red and green signals from the genotyping assay and is an estimate of the proportion of alternate alleles at every SNP locus (McDaneld et al., 2012). These PAF can be used instead of traditional genotype calls of “0,” “1,” or “2” of individual animals. Pooled genotypes were constructed by averaging the genotype calls across the SNP for all individuals in a pool, resulting in quantitative PAF ranging from 0 to 2. In the current study, all genotypes were assumed to be known without error. Additionally, error associated with the formation of pooled genotypes was also ignored, for example, no over- or under-representation of one individual’s DNA in a pool. Thus, it was assumed that no additional residual variation was introduced through the process of generating pooled genotypes or genotyping. In real populations, PAF can only range from 0 to 1. Within real data, a minor adjustment can be made to genotype calls and PAF so that they are on the same scale (Bell et al., 2017).

To mimic a commercial setting where pedigree ties are known to exist between the commercial and seedstock individuals but are not often recorded, the animals in generation 15 were not included in the pedigree. Therefore, the only ties between the pools and individuals in the rest of the population were quantified through genomic relationships.

As a means of comparison, pool sizes of 1 in generation 15 were also considered, which is equivalent to individuals having their own phenotypes and genotypes included in the analysis. In this case, PAF was not needed; genotypes entered the evaluation as the typical calls of “0,” “1,” or “2.” Scenarios in which no information was included from generation 15 were also considered to serve as the alternate extreme comparison. This set of scenarios enables the illustration of the EBV accuracy gained with individual or pooled data compared with no data being utilized, which represents the current situation for many livestock industries, particularly those that are nonintegrated.

Missing generations of genotypes

All individuals (n = 32,000) from the 15 generations had a genotype retained. However, in real livestock populations, genotypes of founder individuals are usually missing and there can be a generational gap between genotyped seedstock animals (e.g., natural service sires in beef cattle, an initial reference population) and commercial animals due to the cost of genotyping. Additionally, genotyped ancestors might be sparse because animals selected for genotyping may be superior or may have an associated phenotype of particular interest or importance (Boligon et al., 2012). Thus, generational gaps of genotyping were induced. For all scenarios, genotypes were retained once selection began (individuals born in generation 6 or after). Four scenarios were considered: individuals up to and including those born in generation 11 were genotyped (Gen11); up to and including those born in generation 12 were genotyped (Gen12); up to and including those born in generation 13 were genotyped (Gen13); and up to and including those born in generation 14 were genotyped (Gen14). All individuals in generations 6 through 14, no matter what scenario was considered, had a recorded phenotype as well as known pedigree relationships. Individuals born in generations 12, 13, or 14 that were not genotyped were included in the pedigree and were phenotyped. Individuals born in generations 0 through 5 that appeared in a three-generation pedigree of the individuals born in generation 15 were included in the pedigree and phenotyped, whereas all others were excluded from the analysis.

Analysis

Single-step GBLUP, which combines genomic and pedigree information in a kinship matrix typically known as H (Aguilar et al., 2010; Christensen and Lund, 2010), was used in order to calculate EBV. The model used when only individual data were included in the analysis was y=Xb+Zu+e, where y is a vector of individual phenotypic observations, X is a known incidence matrix relating observations to fixed effects, b is a vector of fixed effects, Z is a known incidence matrix relating observations to random additive genetic effects, u is a vector of random additive genetic effects, and e is a vector of random residuals. It was assumed var[u]=G=Hσu2 and var[e]=R=Iσe2. The only fixed effect considered was the intercept because no other systematic effects were simulated. The inverse of H (H1) was constructed as:

H1=A1+[000G1A221]

where A1 is the inverse of the numerator relationship matrix constructed using all animals in the pedigree using the principles derived by Henderson (1976). Matrix A22 is the pedigree-based relationship matrix of only the genotyped animals and is constructed according to Colleau (2002). The genomic relationship matrix, G, was calculated in the following way. First, a genomic relationship matrix (Graw) was computed as MM2 Σ pi(1pi), where MM is the centered genotype incidence matrix for individuals and pi is the allelic frequency of the second allele of the ith SNP (VanRaden, 2008). Christensen et al. (2012) formulated a matrix (Gscale) in order to make Graw and A22  compatible by forcing the mean off-diagonal and diagonal elements of Graw to equal the mean off-diagonal and diagonal elements of A22. This was done by setting Gscale=βGraw+α, where β and α are found by solving the following system of linear equations:

diag(Graw)¯β+α=diag(A22)¯
Graw¯β+α=A22¯

Lastly, the matrix Gscale was blended with A22 with coefficients of 0.95 and 0.05, respectively, as suggested by VanRaden (2008) to produce the final genomic relationship matrix (G).

When pooled data were added to the analysis and following the notation established by Su et al. (2018), the underlying model was T[y=Xb+(ZS)(Wu)+e] where vectors y, u, and e and matrices X and Z are defined the same as above. Let m equal the number of individuals that were not pooled, and again q equals the number of individuals in a pool and r equal the number of pools. Matrix T has dimensions (m + r) × (m + rq) and is a design matrix that linked individual observations to the individuals in the pools they were contained in. Matrix S has dimensions (m + rq) × (m + r) and is an indicator matrix that linked individual genotypes to pooled genotypes. Matrix W has dimensions (m + r) × (m + rq) and is also a design matrix that linked individual breeding values to the breeding values of individuals in the pools they were contained in. Let j denote an animal and k denote a pool. Elements Tkj, Wkj, and Sjk were 1 when j = k for individuals in generations 0 through 14, 1q if the jth animal in generation 15 belonged to the kth pool, and 0 otherwise. The matrices T and W average phenotypes and breeding values within pools. Elements Sjk were 1 if the jth animal in generation 15 belonged to the kth pool and 0 otherwise.

Given the assumptions that individual data (genotypes and phenotypes) were unknown for individuals contained in pools, as could be the case in practice, the final prediction model was y=Xb+Zu+e, where y is a vector of individual observations of animals in generations 0 through 14 and pooled phenotypic observations of animals in generation 15, X is a known incidence matrix relating individual and pooled observations to fixed effects, b is a vector of fixed effects, Z is a known incidence matrix relating individual or pooled observations to random additive genetic effects, u is a vector of random additive genetic effects of the individual animals in generations 0 through 14 and pooled animals in generation 15, and e is a vector of random residuals. It was assumed var[u]=G=Hσu2 and var[e]=R=diag(1q)σe2 because the observations in y are heterogeneous in information content given some phenotypes are individuals and others are means of groups of individuals. The inverse of H was constructed in the same fashion above except that the allelic frequencies, pi, were estimated from individuals in generations 0 through 14 as well as the pools. The inverse of H and H was constructed within R (R Core Team, 2017) and then used within ASReml v4.1 software (Gilmour et al., 2015) for the estimation of breeding values.

Accuracy of EBV for sires and dams was estimated as the correlation between true breeding value (TBV) and predicted EBV. The EBV accuracies were estimated for each sex and the generation in which they were born. Accuracy of EBV for pools was estimated as the correlation between the average TBV of the individuals within the pool and the predicted EBV of the pool. To determine the significance of effects on the EBV accuracy, analysis of variance tests were performed with the following model:

yijklm=μ+αi+βj+γk+αβij+αγik+βγjk+αβγijk+bl+eijklm

where yijklm is the EBV accuracy of sires/dams born in generations 11, 12, 13, or 14 or pools; μ is the overall mean; α is the effect of generational gap; β is the effect of pooling strategy; γ is the effect of pool size; b is the random effect of replicate; and e is the random residual. It was assumed b and e were distributed normally with a mean of zero and variance of σb2 and σe2, respectively. Significance was determined at the 0.05 level.

Expectations of pooled genomic relationships

Let G0 represent a genomic relationship matrix with no pooling. Let Gp  represent the expectation of the genomic relationship matrix when considering pooled and non-pooled individuals. The expected genomic relationship matrix is a function of G0 and can be partitioned into four distinct submatrices such that Gp=[G11pG12pG21pG22p] where G11p is the submatrix of relationships between individuals in generations 1 through 14, G12p and G21pare the submatrices of relationships between individuals in generations 1 through 14 and the pools, and G22p is the submatrix of relationships between the pools. Similarly, the genomic relationship matrix can be partitioned into four distinct submatrices such that G0=[G110G120G210G220] Again, let q equal the pool size. The expectations of Gp are as follows:

  1. G11p=G110.

  2. {G22p}kk=(1q1){G220}kk(1q1) where {G22p}kk is the kk′ element of G22p corresponding to pools k and k′ and {G220}kk is the kk′ submatrix of G220 corresponding to individuals in pools k and k′.

  3. {G12p}jk={G120}jk(1q1) where {G12p}jk is the jk′ element of G12p corresponding to individual j and pool k and {G120}jk is the jk′ submatrix of G120 corresponding individual j and to individuals in pool k.

From the expectations above, it can be seen that for a pool of unrelated individuals, the diagonal elements of G22 are equal to 1q, the off-diagonals of G22 are proportional to 1q2, and the elements of G12   and   G21 are proportional to 1q. However, as individuals in pools become more related, the diagonal of G22p is expected to be greater than 1q.

Results and Discussion

Pooling

The number of dams contributing to a pool was equal to the pool size because dams have one progeny per generation. However, sires have 20 progeny per generation and so the number of contributing sires to a pool depended on pool size. The average number of contributing sires to a pool across pooling scenarios was 1, 1.99, 9.57, 18.22, 39.76, and 63.96 for pools of 1, 2, 10, 20, 50, and 100, respectively. On average, random assignment led to the most sires contributing to a pool, whereas minimizing phenotypic variation led to the smallest. However, these differences were small. The largest discrepancy was seen with a pool size of 100; random assignment led to an average of 0.96 more contributing sires than when minimizing phenotypic variation within pools.

The correlations of the average phenotype and the average TBV within pools are depicted in Figure 1. Three distinct patterns emerge when considering pool formation. Randomly assigning individuals to pools led to approximately the same correlation between the average phenotype and average TBV regardless of pool size. When minimizing phenotypic variation within pools, the smallest correlation between average phenotype and average TBV was observed with a pool size of 1 and increased as pool size increased. A large increase was observed between pool sizes of 1 and 2, and again between pool sizes of 2 and 10. After pools of size 10, the gain in the correlation between average phenotype and average TBV plateaued with increasing pool size and approached 1. When considering uniformly maximizing phenotypic variation within pools, the largest correlation was observed with a pool size of 1 and the smallest with a pool size of 100.

Figure 1.

Figure 1.

Correlation of average phenotype and average TBV in pools resulting from different pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools) and pool sizes.

Figures 2 and 3 represent the average relationships of individuals across pools and within pools, respectively. Regardless of the pooling strategy or pool size, the average relationship of individuals across different pools was approximately equal. Relative to relationships of individuals within pools, random assignment led to approximately equal relationships regardless of pool size with the exception of pool sizes of 2, due to random chance. When minimizing phenotypic variation within pools, relationships were the lowest for pool sizes of 2, the highest for pool sizes of 10, and intermediate for pools of 20, 50, and 100. Grouping individuals together with the same sire based on similar phenotypes was unlikely, especially with groups of two. Grouping some half-sibs together was more likely with pool sizes of 10, which led to the increase in average relationship within pools. The average relationships declined again with pools sizes of 20, 50, and 100 because of the large number of individuals in the pools. When uniformly maximizing phenotypic variation within pools, average relationships within pools were approximately equal with the exception with pools of 2, which led to the lowest relationships. This was because individuals with differing phenotypic values were grouped together and given a moderate heritability it was expected that they would not be highly related.

Figure 2.

Figure 2.

Average relationships of individuals across pools resulting from different pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools) and pool sizes.

Figure 3.

Figure 3.

Average relationships of individuals within pools resulting from different pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools) and pool sizes.

If individuals in pools were unrelated, expected values of the diagonal of G22p were 1qk, where qk is the size of the pool. The average realized values of the diagonal elements of G22p were 0.99, 0.50, 0.12, 0.07, 0.04, and 0.03 for pool sizes of 1, 2, 10, 20, 50, and 100, respectively. Slight deviations of realized values are due to the fact that some related individuals were pooled together.

EBV accuracies of sires and dams

Figures 4 and 5 depict the EBV accuracies of sires and dams, respectively, by generation of birth that resulted from different generational gaps in genotyping, pooling strategies, and pool sizes. Results of grandsires/dams are not shown as they follow the same patterns as sires/dams except delayed by one generation. Across all scenarios, the only significant effect was the generational gap in genotyping with the exception of sires and dams born in generation 11. The EBV accuracies of sires born in generation 11 were not significantly impacted by any effects while EBV accuracies of dams born in generation 11 were significantly impacted by both genotyping gaps and pool sizes.

Figure 4.

Figure 4.

EBV accuracies of sires (estimated as the correlation between TBV and predicted EBV) by generation of birth resulting from different generational gaps in genotyping (Gen11 = individuals up to and including those born in generation 11 were genotyped; Gen12 = individuals up to and including those born in generation 12 were genotyped; Gen13 = individuals up to and including those born in generation 13 were genotyped; Gen14 = individuals up to and including those born in generation 14 were genotyped), pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools), and pool sizes with error bars along x-axis.

Figure 5.

Figure 5.

EBV accuracies of dams (estimated as the correlation between TBV and predicted EBV) by generation of birth resulting from different generational gaps in genotyping (Gen11 = individuals up to and including those born in generation 11 were genotyped; Gen12 = individuals up to and including those born in generation 12 were genotyped; Gen13 = individuals up to and including those born in generation 13 were genotyped; Gen14 = individuals up to and including those born in generation 14 were genotyped), pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools), and pool sizes with error bars along x-axis.

Generational gaps of genotyping

Across all scenarios, the lowest EBV accuracies were observed when genotyping occurred only through generation 11, and the largest were observed when genotyping occurred through generation 14. Increases in EBV accuracy due to larger reference populations have been well documented in the literature (e.g., Hayes et al., 2009; Daetwyler et al., 2010). Additionally, in a simulated data set, Lourenco et al. (2017) found that the accuracy of genomic EBV (GEBV) when using ssGBLUP increased as more genotyped individuals were used. Note that when genotyping occurred through generation 14, this represented a situation where all information was used. Accuracies of EBV by year of birth for sires and dams were impacted by the generation in which genotyping stopped and EBV accuracies were highest when the genotyping occurred through or past the generation considered. Table 1 provides the least-squares means of EBV accuracies when different generational gaps in genotyping were considered. All differences of least-squares means were significant.

Table 1.

Least-squares means estimates of EBV accuracies due to generational gaps of genotyping

Sires2 Dams3
Generation genotyping stops1 14 13 12 11 14 13 12 11
Gen11 0.38 0.82 0.83 0.76 0.48 0.53 0.60 0.83
Gen12 0.41 0.83 0.87 0.72 0.50 0.54 0.83 0.85
Gen13 0.46 0.90 0.90 0.79 0.53 0.82 0.84 0.85
Gen14 0.79 0.91 0.90 0.83 0.82 0.83 0.84 0.86
Standard Error 0.064 0.013 0.022 0.090 0.020 0.016 0.005 0.008

1Gen11, individuals up to and including those born in generation 11 were genotyped; Gen12 = individuals up to and including those born in generation 12 were genotyped; Gen13 = individuals up to and including those born in generation 13 were genotyped; Gen14 = individuals up to and including those born in generation 14 were genotyped.

2Sires born in generations 14, 13, 12, or 11.

3Dams born in generations 14, 13, 12, or 11.

The increase in EBV accuracy from when the sires and dams in a generation were genotyped vs. when they were not was dependent on sex and the total number of progeny they had contributing to the evaluation. The largest increase in EBV accuracy resulting from additional genotypes was observed with sires and dams born in generation 14. Accuracy of EBV increased by 70% and 54% for sires and dams, respectively, from when genotyping stopped at generation 13 to 14. Accuracy of EBV increased by 9% and 47% for sires and dams born in generation 13, respectively, from when genotyping stopped at generation 12 to 13. Sires born in generation 14 only had progeny that were born in generation 15, which were those that were pooled. Sires born in generation 13 had 20 individually genotyped/pedigreed progeny in addition to the progeny that were pooled in generation 15. The increase in EBV accuracy from when sires were and were not genotyped was not as large for sires born in generation 13 as those born in generation 14 because EBV accuracy of the sires was already relatively high due to the 20 individual progeny born in generation 14 that were at least in the pedigree. The same concept applied to sires born in generations 11 and 12. Dams, on the other hand, had large increases in EBV accuracy from when they were and were not genotyped compared with sires born in the same generation because they had only one progeny per generation. Predictive ability of young animals for growth traits, measured as the correlation between corrected phenotypes and GEBV, increased from when reference populations included only top bulls with accuracy for birth weight greater than 0.85 (n = 1,628) to when all genotyped animals were included (n = 33,162) for an Angus population (Lourenco et al., 2015). The gaps in genotyping in the current research could reflect a similar situation in which the top accuracy animals (accuracy accumulated because of more progeny) were included in the evaluation. From this result, it can be concluded that the quantity and quality of the information used for evaluation matters.

Connectedness between individuals—deduced from pedigrees or genotypes—impacted EBV accuracies, with the latter giving rise to higher EBV accuracies. Additionally, the number of pedigreed progeny also impacted the EBV accuracies. With more pedigreed progeny already in the evaluation, EBV accuracies of sires did not increase as substantially from when individuals themselves were genotyped and when they were not genotyped. The EBV accuracies of sires and dams as a result of pooling were generally higher than if no data from generation 15 entered the evaluation. This was consistent whether the sires or dams in question were genotyped or were not.

Pooling strategy

Although not significant overall, significant differences were found when looking at pairwise differences in least-squares means of different pooling strategies. Differences were not significant between random assignment and uniformly maximizing phenotypic variation but were significant for the other pairwise comparisons. Minimizing phenotypic variation within pools led to larger EBV accuracies than the other two scenarios. The largest differences in least-squares means were found in sires born in generation 14 where minimizing phenotypic variation resulted in an increase of EBV accuracy of 8% and 9% compared with random assignment and uniformly maximizing variation, respectively. Although other comparisons between these pooling scenarios were statistically significant when sires/dams were born in other generations, the difference may not be practically different. The average increase across generations born and sires/dams was approximately 1% (results not shown).

Henshall et al. (2012) concluded that pooling by the rank of phenotype within contemporary groups led to results more correlated with individual genotyping than pooling based on ranked, pre-adjusted phenotypes across contemporary groups. The current study did not include designed systematic effects; therefore, contemporary groups were not considered when constructing pools. Within simulation, Alexandre et al. (2019) pooled individuals based on two traits, one with a heritability of 0.1 (trait 1) and the other of 0.4 (trait 2). The pools were constructed based on trait 1, trait 2, a combination of both, or randomly. Relationships between pools and 200 sires were estimated by genomic relationships alone. Construction of pools based on a single trait was similar to minimizing phenotypic variation within pools in the current study. Accuracies of GEBV, estimated as the correlation of GEBV and TBV, for a single trait were greatest when pools were constructed based on the trait itself and lowest when pools were constructed randomly. Therefore, the ways in which pools are constructed do impact the EBV accuracies of prediction.

Pooling size

Again, while the effect of pool size was not significant overall, some pairwise comparisons of least-squares means did show significant differences. Least-squares means of sire EBV accuracies are presented in Table 2. The EBV accuracies of sires resulting from pool sizes of 10, 20, 50, or 100 were not significantly different from those when no information from generation 15 was included in the evaluation when pools were constructed randomly or by maximizing phenotypic variation. Exceptions to this were for pool sizes of 10 and 20 using either pooling strategy (sires born in generation 14 had significantly increased accuracy) and when pools of size 20 uniformly maximized variation (sires born in generation 13 had significantly increased accuracy). Estimated BV accuracies resulting from pool sizes of 2 were intermediate to situations in which progeny in generation 15 were individually genotyped and when no information from generation 15 was used. Additionally, the only differences in EBV accuracies resulting from pooling and individual data that were not significantly different were with pool sizes of 2. The gain in additional information when pooling randomly or by uniformly maximizing phenotypic variation within pools was not significant when progeny were grouped in pool sizes greater than 10 compared with when data from generation 15 were not used at all, often a numerical gain in accuracy was not even observed. A pooling size of 2 was the only scenario that did not decrease the EBV accuracy significantly when pools were formed randomly or by uniformly maximizing phenotypic variation within pools.

Table 2.

Least-squares means estimates of EBV accuracies of sires due to pooling strategy, pool size, and generational gaps in genotyping

Born in Generation3
14 13 12 11
Pooling Strategy1 Pool Size2 Gen 114 Gen 125 Gen 136 Gen 147 Gen 11 Gen 12 Gen 13 Gen 14 Gen 11 Gen 12 Gen 13 Gen 14 Gen 11 Gen 12 Gen 13 Gen 14
Random 1 0.40 0.45 0.52b 0.87b 0.83 0.83 0.92b 0.93b 0.84 0.91b 0.92b 0.92b 0.82 0.80 0.85 0.88
2 0.38 0.42 0.48 0.82 0.82 0.83 0.91 0.91 0.83 0.89b 0.91 0.90 0.78 0.72 0.80 0.84
10 0.37 0.39 0.44 0.77a 0.82 0.83 0.90a 0.90a 0.83 0.86a 0.89a 0.89 0.72 0.69 0.77 0.80
20 0.37 0.39 0.43 0.75a 0.82 0.82 0.89a 0.90a 0.83 0.86a 0.89a 0.89a 0.73 0.68 0.76 0.79
50 0.36 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.90a 0.83 0.85a 0.88a 0.88a 0.74 0.69 0.76 0.80
100 0.37 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.88a 0.88a 0.74 0.69 0.76 0.80
0 0.37 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.88a 0.88a 0.77 0.70 0.77 0.80
Minimize 1 0.40 0.45 0.52b 0.87b 0.83 0.83 0.92b 0.93b 0.84 0.91b 0.92b 0.92b 0.82 0.80 0.85 0.88
2 0.40 0.46 0.54b 0.86b 0.83 0.83 0.92b 0.93b 0.84 0.90b 0.92b 0.91b 0.80 0.78 0.84 0.88
10 0.41 0.47 0.54b 0.85b 0.83 0.84 0.92b 0.92b 0.84 0.88a 0.91b 0.90 0.77 0.76 0.82 0.87
20 0.40 0.46 0.53b 0.84b 0.83 0.84 0.92b 0.92b 0.84 0.87a 0.90 0.90 0.77 0.76 0.82 0.87
50 0.39 0.44 0.50 0.82 0.83 0.83 0.91 0.91 0.83 0.87a 0.90 0.90 0.77 0.74 0.81 0.85
100 0.38 0.42 0.47 0.80 0.83 0.83 0.91 0.91 0.83 0.87a 0.90a 0.90a 0.76 0.72 0.80 0.84
0 0.37 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.88a 0.88a 0.77 0.70 0.77 0.80
Uniformly 1 0.40 0.45 0.52b 0.87b 0.83 0.83 0.92b 0.93b 0.84 0.91b 0.92b 0.92b 0.82 0.80 0.85 0.88
Maximize 2 0.38 0.41 0.46 0.81 0.82 0.82 0.90a 0.90 0.83 0.89b 0.90 0.90 0.74 0.71 0.77 0.81
10 0.36 0.38 0.43 0.75a 0.82 0.82 0.89a 0.89a 0.83 0.86a 0.89a 0.89a 0.73 0.68 0.75 0.79
20 0.37 0.38 0.42a 0.74a 0.82 0.82 0.89a 0.89a 0.83 0.86a 0.89a 0.88a 0.74 0.69 0.76 0.79
50 0.36 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.89a 0.88a 0.74 0.69 0.76 0.79
100 0.37 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.88a 0.88a 0.75 0.69 0.76 0.80
0 0.37 0.38 0.42a 0.73a 0.82 0.82 0.89a 0.89a 0.83 0.85a 0.88a 0.88a 0.77 0.70 0.77 0.80
Standard Error 0.073 0.015 0.023 0.100

1Random = individuals were randomly assigned to pools; Minimize = individuals were pooled so that phenotypic variation within pools was minimized; Uniformly maximize = individuals were pooled so that phenotypic variation within pools was uniformly maximized.

21 = individually genotyped and phenotyped; 2 = pool size of 2; 10 = pool size of 10; 20 = pool size of 20; 50 = pool size of 50; 100 = pool size of 100; 0 = data from generation 15 did not enter the evaluation.

3Sires born in generations 14, 13, 12, or 11.

4Gen11 = individuals up to and including those born in generation 11 were genotyped.

5Gen12 = individuals up to and including those born in generation 12 were genotyped.

6Gen13 = individuals up to and including those born in generation 13 were genotyped.

7Gen14 = individuals up to and including those born in generation 14 were genotyped.

aWithin a column and pooling strategy, the least-squares means difference with a pool size of one is significant.

bWithin a column and pooling strategy, the least-squares means difference with when no information from generation 15 is included is significant.

When minimizing phenotypic variation within pools, EBV accuracies of sires resulting from pool sizes of 50 or 100 were not significantly different than those when no information from generation 15 was included. Additionally, EBV accuracies from all pool sizes were not significantly different than individual information from generation 15 with the exception of pool sizes of 10, 20, 50, and 100 when sires were born in generation 12 and genotyping stopped at generation 12 or with pool sizes of 100 when genotyping stopped at either generation 13 or 14. These results also show that EBV accuracies from large pool sizes (50 or 100) show no improvement compared with when data from generation 15 were excluded completely. It also shows that overall, even though there is a reduction in EBV accuracy resulting from pooling compared with individual data, the reduction is not statistically significant. These results are consistent with Alexandre et al. (2019) who suggested pool sizes of 10 in order to not only retain EBV accuracy but also save on genotyping costs. However, Kuehn et al. (2018) suggested pool sizes of at least 20. In a study investigating the efficiency of estimated genomic relationships of pools to the animals that make up the pools and to other potentially related individuals, Kuehn et al. (2018) found that technical error (error due to the genotyping of the intensity of the fluorescent dye) was a minimal contribution to the total pooled error. It was also suggested the use of large pools because they are less prone to pool construction error—the planned representation of individual DNA to the pool. Thus, the impact of errors associated with PAF and pool construction decreases with large pool sizes.

Although some statistically significant differences were found for pairwise comparisons of least-squares means of EBV accuracy of dams, differences in EBV accuracy did not exceed 0.02, and thus results are not presented.

When comparing the decrease in EBV accuracy due to pooling compared with individual data, Alexandre et al. (2019) reported larger decreases compared with those presented herein and were dependent on the heritability of the trait. Alexandre et al. (2019) reported large drops in GEBV accuracy from individual data to pool sizes of 2 and 10, but began to plateau with pool sizes of 20, 25, 50, and 100 for the trait with a heritability of 0.4, when pools were constructed based on the trait itself. The same authors reported that when pools were constructed randomly, GEBV accuracy of the trait with a heritability of 0.4 resulting from pool sizes of 10 was comparable to the GEBV accuracy of the lowly heritable trait. The more dramatic decreases in GEBV accuracy observed by Alexandre et al. (2019) may be caused by the fact that only a sire’s own phenotype and the pools’ phenotypes were entered into the evaluation. In the current study, other relatives’ information also entered into the evaluation, so that the decrease in information as pool sizes became larger were not as detrimental, justifying the use of single-step evaluation.

Presumably, results from when no information from generation 15 was included in the evaluation would serve as a lower boundary for EBV accuracy and the upper boundary would be defined by the case when progeny born in generation 15 were genotyped individually. However, when sires/dams were not genotyped and pools were constructed to minimize phenotypic variation within pools, EBV accuracies resulting from pooling were actually higher than if generation 15 had individual data. The EBV accuracies were maximized at pool sizes of 10. This phenomenon was likely a result of both the increased relationship within pools and the confidence in the average phenotype representing the pooled phenotype, determined by the correlation of average phenotype and average TBV in pools. These differences in EBV accuracy from individual data from generation 15 to any pool size were not significant except pools of 10, 20, 50, and 100 for sires born in generation 12 and genotyping stopped at generation 12, as already noted previously.

EBV accuracy of pools

The EBV accuracy of pools is given in Figure 6. An analysis of variance showed that the effects of pool size and the interaction between pool size and pooling strategy to be significant. Pools sizes of 100 had the lowest EBV accuracy and pool size of 1 had the largest EBV accuracy using random assignment and uniformly maximizing phenotypic variation within pools. However, the effect of pooling when uniformly maximizing variation had larger effects on the EBV accuracy compared with random assignment to pools, seen by larger decreases in EBV accuracy as pool sizes increased. When pools were formed by minimizing phenotypic variation, pool sizes of 100 led to the largest EBV accuracies for the pools while individual data led to the lowest EBV accuracy. Accuracies of EBV resulting from pool sizes of 10 were significantly different compared with pool sizes of 2 and 1. However, EBV accuracies resulting from pools of 10 compared with pool sizes of 20, 50, or 100 were not significantly different.

Figure 6.

Figure 6.

EBV accuracies of pools (estimated as the correlation between the average TBV of the individuals within the pool and predicted EBV of the pool) resulting from different generational gaps in genotyping (Gen11 = individuals up to and including those born in generation 11 were genotyped; Gen12 = individuals up to and including those born in generation 12 were genotyped; Gen13 = individuals up to and including those born in generation 13 were genotyped; Gen14 = individuals up to and including those born in generation 14 were genotyped), pooling strategies (Random = randomly allocated to pools; Minimize = minimize phenotypic variation within pools; Uniformly Maximize = uniformly maximize phenotypic variation within pools), and pool sizes with error bars along x-axis.

Practical applications of pooling phenotypes and genotypes have been used before. Bell et al. (2017) used dag scores in Merino sheep to pool individuals in commercial flocks, resulting in categorical phenotypes and PAF for each of the pools. These PAF were combined with individual sire genotypes into a hybrid genomic relationship matrix (h-GRM) for the use in GBLUP estimations of GEBV of the sires. Pregnancy and lactation status, a categorical phenotype, in Brahman cows were used to pool cattle (Reverter et al., 2016). The resulting PAF from the pools were combined with individual genotypes of herd and stud bulls into an h-GRM for use in GBLUP estimations of GEBV for the fertility of bulls. The bulls were not the sires of the cows in the pools. Within both studies, pedigrees were unknown for the animals used for pooling. These studies showed the potential use of pooling to estimate GEBV of direct parents (Bell et al., 2017) or of seedstock individuals (Reverter et al., 2016). The work of Bell et al. (2017) and Reverter et al. (2016) represent the practical applications of the current study. However, because individual genotypes were not available, the loss of GEBV accuracy was unknown, warranting further research in this area. Additionally, both Bell et al. (2017) and Reverter et al. (2016) pooled individuals based on similar categorical phenotypes, which would be similar to minimizing phenotypic variation within pools using a quantitative phenotype. The current research demonstrates the validity of work such as Bell et al. (2017) and Reverter et al. (2016), especially when pools are constructed in order to minimize phenotypic variation within the pools and pool size is less than 50. Results from such studies should lead to EBV accuracy that is not significantly different than when individual data are included. Further research with ssGBLUP and pooling DNA and phenotypic data are needed within real populations.

Conclusions

Accuracies of EBV from this simulation represent theoretical maximum EBV accuracies; realized EBV accuracies resulting from pooling could be less due to lab and genotyping errors. However, the results presented in this paper show the potential use of pooling data in order to economically make use of commercial data in genetic evaluations. The use of pooled phenotypes and genotypes in combination with a ssGBLUP evaluation can be a potential way to economically leverage the plethora of phenotypes from commercial sectors in combination with the individual-level data (genotypes and phenotypes) from nucleus (seedstock) animals. When pools were constructed in such a way that minimized the phenotypic variation within pools, pool sizes of 2, 10, 20, or 50 did not generally lead to differences in EBV accuracy that are statistically different than when individual progeny data were used. Sires with prior low EBV accuracy benefited the most from pooled observations. Additionally, the resulting EBV for the pools could be used to inform management decisions. Such examples would be using the EBV for marketing purposes or specialized feeding programs.

Acknowledgments

This project is based on research that was partially supported by the Nebraska Agricultural Experiment Station with funding from the Hatch Act (Accession Number 1011203) through the USDA National Institute of Food and Agriculture.

Glossary

Abbreviations

EBV

estimated breeding value

GEBV

genomic estimated breeding value

GWAS

genome-wide association study

h-GRM

a hybrid genomic relationship matrix

PAF

pooling allele frequency

QTL

quantitative trait loci

SNP

single nucleotide polymorphism

ssGBLUP

single-step genomic best linear unbiased prediction

TBV

true breeding value

Conflict of interest statement

The authors declare no conflicts of interest to objectively present this research. Mention of a trade name, proprietary product, or specific equipment does not constitute a guarantee or warranty by the USDA and does not imply approval to the exclusion of other products that may be suitable. USDA is an equal opportunity provider and employer.

Literature Cited

  1. Aguilar I., Misztal I., Johnson D. L., Legarra A., Tsuruta S., and Lawlor T. J.. . 2010. Hot Topic: A unified approach to utilize phenotypic, full pedigree, and genomic information for genetic evaluation of Holstein final score. J. Dairy Sci. 93:743–752. doi: 10.3168/jds.2009-2730 [DOI] [PubMed] [Google Scholar]
  2. Alexandre P. A., Porto-Neto L. R., Karaman E., Lehnert S. A., and Reverter A.. . 2019. Pooled genotyping strategies for the rapid construction of genomic reference populations. J. Anim. Sci. 97:4761–4769. doi: 10.1093/jas/skz344 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Baller J. L., Howard J. T., Kachman S. D., and Spangler M. L.. . 2019. The impact of clustering methods for cross-validation, choice of phenotypes, and genotyping strategies on the accuracy of genomic predictions. J. Anim. Sci. 97:1534–1549. doi: 10.1093/jas/skz055 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Bell A. M., Henshall J. M., Porto-Neto L. R., Dominik S., McCulloch R., Kijas J., and Lehnert S. A.. . 2017. Estimating the genetic merit of sires by using pooled DNA from progeny of undetermined pedigree. Genet. Sel. Evol. 49:28. doi: 10.1186/s12711-017-0303-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Boligon A. A., Long N., Albuquerque L. G., Weigel K. A., Gianola D., and Rosa G. J.. . 2012. Comparison of selective genotyping strategies for prediction of breeding values in a population undergoing selection. J. Anim. Sci. 90:4716–4722. doi: 10.2527/jas.2012-4857 [DOI] [PubMed] [Google Scholar]
  6. Chen G. K., Marjoram P., and Wall J. D.. . 2009. Fast and flexible simulation of DNA sequence data. Genome Res. 19:136–142. doi: 10.1101/gr.083634.108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Christensen O. F., and Lund M. S.. . 2010. Genomic prediction when some animals are not genotyped. Genet. Sel. Evol. 42:2. doi: 10.1186/1297-9686-42-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Christensen O. F., Madsen P., Nielsen B., Ostersen T., and Su G.. . 2012. Single-step methods for genomic evaluation in pigs. Animal 6:1565–1571. doi: 10.1017/S1751731112000742 [DOI] [PubMed] [Google Scholar]
  9. Colleau J. J. 2002. An indirect approach to the extensive calculation of relationship coefficients. Genet. Sel. Evol. 34:409–421. doi: 10.1186/1297-9686-34-4-409 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Daetwyler H. D., Pong-Wong R., Villanueva B., and Woolliams J. A.. . 2010. The impact of genetic architecture on genome-wide evaluation methods. Genetics 185:1021–1031. doi: 10.1534/genetics.110.116855 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Darvasi A., and Soller M.. . 1994. Selective DNA pooling for determination of linkage between a molecular marker and a quantitative trait locus. Genetics 138:1365–1373. doi: 10.1007/bf00222881 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Fisher P. J., Turic D., Williams N. M., McGuffin P., Asherson P., Ball D., Craig I., Eley T., Hill L., Chorney K., . et al. 1999. DNA pooling identifies QTLs on chromosome 4 for general cognitive ability in children. Hum. Mol. Genet. 8:915–922. doi: 10.1093/hmg/8.5.915 [DOI] [PubMed] [Google Scholar]
  13. Gaj P., Maryan N., Hennig E. E., Ledwon J. K., Paziewska A., Majewska A., Karczmarski J., Nesteruk M., Wolski J., Antoniewicz A. A., . et al. 2012. Pooled sample-based GWAS: a cost-effective alternative for identifying colorectal and prostate cancer risk variants in the Polish population. PLoS One 7:e35307. doi: 10.1371/journal.pone.0035307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Gilmour A. R., Gogel B. J., Cullis B. R., Welham S. J., and Thompson R.. 2015. ASReml User Guide Release 4.1 Functional Specification. Hemel Hempstead (UK):VSN International. Available from https://asreml.kb.vsni.co.uk/wp-content/uploads/sites/3/2018/02/ASReml-4.1-Functional-Specification.pdf [Google Scholar]
  15. Hayes B. J., Bowman P. J., Chamberlain A. J., and Goddard M. E.. . 2009. Invited Review: Genomic selection in dairy cattle: progress and challenges. J. Dairy Sci. 92:433–443. doi: 10.3168/jds.2008-1646 [DOI] [PubMed] [Google Scholar]
  16. Henderson C. R. 1976. A simple method for computing the inverse of a numerator relationship matrix used in prediction of breeding values. Biometrics 32:69–83. doi: 10.2307/2529339 [DOI] [Google Scholar]
  17. Henshall J. M., Hawken R. J., Dominik S., and Barendse W.. . 2012. Estimating the effect of SNP genotype on quantitative traits from pooled DNA samples. Genet. Sel. Evol. 44:12. doi: 10.1186/1297-9686-44-12 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Howard J. T., Tiezzi F., Pryce J. E., and Maltecca C.. . 2017. Geno-diver: a combined coalescence and forward-in-time simulator for populations undergoing selection for complex traits. J. Anim. Breed. Genet. 134:553–563. doi: 10.1111/jbg.12277 [DOI] [PubMed] [Google Scholar]
  19. Huang W., Kirkpatrick B. W., Rosa G. J., and Khatib H.. . 2010. A genome-wide association study using selective DNA pooling identifies candidate markers for fertility in Holstein cattle. Anim. Genet. 41:570–578. doi: 10.1111/j.1365-2052.2010.02046.x [DOI] [PubMed] [Google Scholar]
  20. Kuehn L. A., McDaneld T. G., Keele J. W.. 2018Quantification of genomic relationship from DNA pooled samples. In: Proceedings of the World Congress on Genetics Applied to Livestock Production; February 12 to 16; Auckland, New Zealand. http://www.wcgalp.org/proceedings/2018/quantification-genomic-relationship-dna-pooled-samples. Accessed 11 June 2020.
  21. Lourenco D. A. L., Fragomeni B. O., Bradford H. L., Menezes I. R., Ferraz J. B. S., Aguilar I., Tsuruta S., and Misztal I.. . 2017. Implications of SNP weighting on single-step genomic predictions for different reference population sizes. J. Anim. Breed. Genet. 134:463–471. doi: 10.1111/jbg.12288 [DOI] [PubMed] [Google Scholar]
  22. Lourenco D. A., Tsuruta S., Fragomeni B. O., Masuda Y., Aguilar I., Legarra A., Bertrand J. K., Amen T. S., Wang L., Moser D. W., . et al. 2015. Genetic evaluation using single-step genomic best linear unbiased predictor in American Angus. J. Anim. Sci. 93:2653–2662. doi: 10.2527/jas.2014-8836 [DOI] [PubMed] [Google Scholar]
  23. McDaneld T. G., Kuehn L. A., Thomas M. G., Snelling W. M., Sonstegard T. S., Matukumalli L. K., Smith T. P., Pollak E. J., and Keele J. W.. . 2012. Y are you not pregnant: identification of Y chromosome segments in female cattle with decreased reproductive efficiency. J. Anim. Sci. 90:2142–2151. doi: 10.2527/jas.2011-4536 [DOI] [PubMed] [Google Scholar]
  24. R Core Team. 2017. R: A language and environment for statistical computing. Vienna (Austria):R Foundation for Statistical Computing; Available from https://www.R-project.org/. [Google Scholar]
  25. Reverter A., Porto-Neto L. R., Fortes M. R., McCulloch R., Lyons R. E., Moore S., Nicol D., Henshall J., and Lehnert S. A.. . 2016. Genomic analyses of tropical beef cattle fertility based on genotyping pools of Brahman cows with unknown pedigree. J. Anim. Sci. 94:4096–4108. doi: 10.2527/jas.2016-0675 [DOI] [PubMed] [Google Scholar]
  26. Sham P., Bader J. S., Craig I., O’Donovan M., and Owen M.. . 2002. DNA pooling: a tool for large-scale association studies. Nat. Rev. Genet. 3:862–871. doi: 10.1038/nrg930 [DOI] [PubMed] [Google Scholar]
  27. Sonesson A. K., Meuwissen T. H., and Goddard M. E.. . 2010. The use of communal rearing of families and DNA pooling in aquaculture genomic selection schemes. Genet. Sel. Evol. 42:41. doi: 10.1186/1297-9686-42-41 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Strillacci M. G., Frigo E., Schiavini F., Samoré A. B., Canavesi F., Vevey M., Cozzi M. C., Soller M., Lipkin E., and Bagnato A.. . 2014. Genome-wide association study for somatic cell score in Valdostana Red Pied cattle breed using pooled DNA. BMC Genet. 15:106. doi: 10.1186/s12863-014-0106-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Su G., Madsen P., Nielsen B., Ostersen T., Shirali M., Jensen J., and Christensen O. F.. . 2018. Estimation of variance components and prediction of breeding values based on group records from varying group sizes. Genet. Sel. Evol. 50:42. doi: 10.1186/s12711-018-0413-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. VanRaden P. M. 2008. Efficient methods to compute genomic predictions. J. Dairy Sci. 91:4414–4423. doi: 10.3168/jds.2007-0980 [DOI] [PubMed] [Google Scholar]

Articles from Journal of Animal Science are provided here courtesy of Oxford University Press

RESOURCES