Skip to main content
Molecular Biology and Evolution logoLink to Molecular Biology and Evolution
. 2026 Apr 28;43(5):msag107. doi: 10.1093/molbev/msag107

GAPIT version 4: integration of GWAS into genomic prediction

Jiabo Wang 1,2,✉, Zhiwu Zhang 3,✉
Editor: Carlos Schrago
PMCID: PMC13183175  PMID: 42047095

Abstract

Genomic prediction leverages all available markers, irrespective of their statistical significance in genome-wide association studies (GWAS). Recent advancements in marker density, sample sizes, and sophisticated statistical GWAS methods have demonstrated that integrating GWAS results can potentially boost the accuracy of genomic predictions. The Genomic Association and Prediction Tool (GAPIT) has recently begun incorporating GWAS findings into its prediction framework, streamlining this approach, referred to as GWAS-Assisted Genomic Best Linear Unbiased Prediction (GAGBLUP). A sufficient simulation study revealed that the benefits of GAGBLUP depend on the GWAS model used. Multiple-locus models, such as Bayesian information and linkage-disequilibrium iteratively nested keyway (BLINK), outperformed single-locus models, such as the mixed linear model. Specifically, when BLINK GWAS results in a real trait were incorporated into Genomic Best Linear Unbiased Prediction (GBLUP), prediction accuracy improved by over 20% compared to GBLUP alone. This approach integrates the trait-specific insights from GWAS with the polygenic modeling capacity of GBLUP, resulting in more stable prediction across varying genetic backgrounds. This broader applicability enhances the utility of genomic selection in breeding programs, enabling its deployment across a wider range of crops and trait architectures.

Keywords: genomic selection, genetic markers, breeding, accuracy, linkage disequilibrium

Graphical Abstract

Graphical Abstract.

For image description, please refer to the figure legend and surrounding text.

Introduction

Discovery and prediction represent two fundamental and complementary goals of scientific research. In genomic research, these goals are primarily pursued through genome-wide association studies (GWAS) for discovery and genomic prediction (GP) for prediction, respectively (Fernando and Garrick 2013). GWAS aim to identify genetic loci associated with specific traits and estimate their effects, often by pinpointing regions near causal genes. In contrast, GP focuses on directly predicting breeding values by estimating the overall genetic merit underlying the traits (Wang et al. 2018). In GP, incorporating known major genes can improve the accuracy of phenotype predictions (Reyes-Valdés 2000; Nishio and Satoh 2015; Spindel et al. 2016; Wolc et al. 2016).

However, the outcomes of incorporating associated markers identified through GWAS have been inconsistent. Some findings have been unreliable due to contamination, which arises from using phenotypes of individuals in the testing population when GWAS is initially performed on all individuals, including both the training and testing sets (Chiwina et al. 2023). When contamination is eliminated, only one-third of the 200 simulated traits showed improved outcomes from integrating GWAS performed using a population structure (often referred to as a Q matrix) and kinship matrix (K) model into GP conducted via ridge regression (Rice and Lipka 2019). The Q + K GWAS model was introduced in 2006 by Yu et al. (Yu et al. 2006) to fully control the inflated P values in addition to the Q model introduced by Pritchard in 2000 (Pritchard et al. 2000). Since then, multiple methods have been developed to control both false positives and false negatives to enhance statistical power.

To minimize confounding between the kinship and the markers being tested in the Q + K mixed linear model (MLM), individuals were substituted with their respective groups in the compressed MLM (Zhang et al. 2010). Additional improvements include the enriched compressed MLM (Li et al. 2014), SUPER (Wang et al. 2014), multiple loci mixed linear model (MLMM) (Segura et al. 2012), FarmCPU (Liu et al. 2016), and Bayesian information and linkage-disequilibrium iteratively nested keyway (BLINK) (Huang et al. 2019).

Genomic Association and Prediction Tool (GAPIT) is a widely used software package in R that implemented all of these GWAS methods (Lipka et al. 2012; Tang et al. 2016; Wang and Zhang 2021a). Furthermore, GAPIT also implemented multiple GP models, including the Genomic Best Linear Unbiased Prediction (GBLUP), compressed BLUP, and SUPER BLUP. The implementation of the GWAS and GP methods has been documented in the first three versions.

Although GWAS and GP have traditionally been modeled independently, their integration has recently gained attention. This integration differs from marker-assisted selection (MAS), which is also conducted. GWAS-identified markers are typically subjected to rigorous validation before being applied in selection processes. In contrast, integrating GWAS into GP seeks to enhance prediction accuracy without requiring extensive validation of associated markers. However, because multiple associated markers may be in high linkage disequilibrium (LD) and reflect the same causal gene, careful selection among these markers is necessary before incorporating them as covariates in GP models.

Existing integration strategies, such as SUPER BLUP, often rely on binning significant markers to reduce LD before constructing kinship matrices for GP. While such approaches can perform well for traits governed by a few moderate- to large-effect loci, their accuracy tends to decline substantially for highly polygenic traits with numerous small-effect variants. Therefore, their performance is highly dependent on prior knowledge of the trait's genetic architecture, which limits their general applicability in real-world breeding scenarios where such information is often unavailable.

To overcome this constraint, we have enhanced the architectural flexibility and adaptive capacity of the prediction framework implemented in the GAPIT package. This improvement allows the method to integrate GWAS-derived markers as fixed effects while constructing a refined kinship matrix from the remaining markers as random effects, thereby maintaining competitive performance across both oligogenic and polygenic trait architectures.

Results

GAPIT consists of five sequential modules (Wang and Zhang 2021b). The first module prepares the data by handling tasks, such as data formatting and genotype imputation. The second module focuses on quality control, including the filtering of markers. The third module interprets the data, calculating metrics, such as minor allele frequency, principal component analysis, kinship, LD, and marker density. The fourth module performs GWAS and GP, integrating GWAS with GP. The fifth module generates tables and graphs to interpret results. The integration of GWAS into GP was conducted exclusively within the fourth module.

GAPIT has implemented several GWAS methods, including GLM, MLM, CMLM, SUPER, MLMM, FarmCPU, and BLINK. Their performance was evaluated using a trait simulated from a maize dataset comprising 282 inbred lines genotyped with a 55K Single Nucleotide Polymorphism markers (SNP) array. The trait was governed by 20 quantitative trait nucleotides (QTNs) and exhibited 75% heritability. Each of the seven GWAS methods was applied to the full set of 282 inbred lines, with the simulation replicated 40 times. The GWAS results for the first replicate are depicted in Manhattan plots (Fig. 1), while the average performance across all 40 replicates is summarized in Table 1. At a 5% type I error rate after Bonferroni multiple-test correction, BLINK demonstrated the highest statistical power (0.26) with a false discovery rate (FDR) of 0.32. In contrast, GLM exhibited the highest FDR (0.77) with a power of 0.13. Generally, multilocus models showed significantly higher statistical power and lower FDR compared to single-locus models.

Figure 1.

For image description, please refer to the figure legend and surrounding text.

Advances of GWAS models. The evaluation was performed on a trait simulated from the genotypes of 55K SNPs genotyped on 282 maize inbred lines. The trait was controlled by 20 QTNs with 75% heritability. The evaluation was replicated 40 times. The first replicate is illustrated with Manhattan plot for GWAS performance using all individuals. The threshold of 5% of type I error is indicated by the horizontal green line after Bonferroni multiple-test correction. The average performances are summarized in Table 1 across the 40 replicates.

Table 1.

Performances of seven GWAS methods implemented in GAPITa.

Method Positives QTNs QTN ratio False positives FDR Power
GLM 11.48 2.61 0.23 8.87 0.77 0.13
MLM 5.42 2.12 0.39 3.30 0.61 0.11
CMLM 5.27 2.09 0.40 3.18 0.60 0.10
SUPER 17.36 3.70 0.21 13.66 0.79 0.19
MLMM 6.94 4.64 0.67 2.30 0.33 0.23
FarmCPU 6.70 4.82 0.72 1.88 0.28 0.24
BLINK 7.61 5.18 0.68 2.43 0.32 0.26

aThe performances were examined on a trait simulated from the genotypes of 55K SNPs genotyped on 282 maize inbred lines. The trait was controlled by 20 QTNs with 75% heritability.

FDR, false discovery rate.

In multilocus models, such as MLMM, FarmCPU, and BLINK, the correlation among associated markers is minimal, enabling these markers to be directly used as covariates in GBLUP for GWAS-assisted GP. In contrast, single-locus models, such as GLM and MLM, often detect associated markers with strong correlations, especially within regions of high LD (see Fig. 1). For example, on maize chromosome 4, multilocus models identify a single QTN above the significance threshold, while single-locus models detect multiple SNPs: two with GLM, MLM, and CMLM and three with SUPER, of which only one is a true QTN.

The prediction accuracy of GBLUP, incorporating these filtered markers, was evaluated using both simulated data (maize) and real data (peach). All seven GWAS methods were evaluated using simulated data to assess prediction accuracy through 5-fold cross-validation. The best-performing method was then applied to real trait to evaluate prediction accuracy in peach data. The simulated trait matched the one used in the study of the seven GWAS methods, which assessed their performance in terms of power and FDR. The entire population was divided into five folds, with one fold designated as the testing population and the remaining four as the training population. GWAS was performed on the phenotypes of the training population to identify associated markers. These markers were then filtered using QR decomposition, and the remaining markers were incorporated as covariates in GBLUP using phenotype of the training population to predict the breeding values of the testing population. Prediction accuracy was estimated by Pearson's correlation between the predicted breeding values and the true breeding values of the testing population. This process was repeated, iterating through each of the five folds as the testing population, and the entire procedure was replicated 40 times. The resulting prediction accuracies are presented in Fig. 2 and Table 2.

Figure 2.

For image description, please refer to the figure legend and surrounding text.

Prediction accuracy of GBLUP with and without results by seven GWAS methods. The evaluation was performed on a trait simulated from the genotypes of 55K SNPs genotyped on 282 maize inbred lines. The trait was controlled by 20 QTNs with 75% heritability. The evaluation was replicated 40 times. For each replicate, 5-fold cross-validation was conducted. GWAS was conducted on the training population only to select the associated SNPs as covariates for GBLUP to evaluate prediction accuracy. The selection of associated SNPs was determined by the Bonferroni threshold at 5% of type I error after the multiple-test correction.

Table 2.

Prediction accuracy of GBLUP without and with incorporating results of using seven GWAS methodsa.

Method Mean Maximum Minimum Median SE
GBLUP 0.5020 0.6338 0.3212 0.5100 0.0124
GLM 0.7071 0.8455 0.4452 0.7072 0.0143
MLM 0.7078 0.8511 0.4419 0.7083 0.0138
CMLM 0.7033 0.8511 0.4419 0.7083 0.0137
SUPER 0.7616 0.8919 0.5402 0.7688 0.0115
MLMM 0.7666 0.8867 0.5069 0.7913 0.0139
FarmCPU 0.7873 0.9111 0.5414 0.8061 0.0131
BLINK 0.7879 0.8977 0.5889 0.7925 0.0123

aPrediction accuracies were examined using 5-fold cross-validation on a trait simulated from the genotypes of 55K SNPs genotyped on 282 maize inbred lines. The trait was controlled by 20 QTNs with 75% heritability. Prediction accuracy is displayed for mean, median, minimum, and maximum among 40 replicates of cross-validation with five folds.

SE: standard error.

Prediction accuracies are higher than GBLUP when GWAS results are incorporated using any of the seven methods compared to GBLUP without such incorporation. The magnitude of improvement in prediction accuracy was closely aligned with the statistical power rankings of the respective GWAS methods, with BLINK demonstrating the highest performance. Generally, multilocus GWAS methods outperform single-locus methods. Notably, while SUPER GWAS delivers accuracy comparable to multilocus models, it also reports the highest number of positives (17.36). The average number of QTNs identified by SUPER (3.7) falls between that of other single-locus methods (2.09 to 2.61) and multilocus methods (4.64 to 5.18) (Table 1). These findings indicate that filtering associated markers using QR decomposition is an effective strategy.

Next, the optimal integration approach (GBLUP with BLINK) was compared to two standard practices in molecular breeding—MAS and GBLUP—using real trait data from a peach dataset. Both MAS and GWAS-Assisted Genomic Best Linear Unbiased Prediction (GAGBLUP) utilized the BLINK GWAS method. Five-fold cross-validation with 40 replicates was conducted. For MAS and GAGBLUP, GWAS analyses were performed solely on the training population, excluding the testing population's phenotypes from both GWAS and GPs with GBLUP or GAGBLUP. Since true breeding values were unavailable, prediction accuracy was assessed using Pearson's correlation coefficients between the predicted and observed phenotypes of the testing population. The results indicated that GAGBLUP outperformed both MAS and standard GBLUP. Specifically, GAGBLUP improved prediction accuracy by 30% compared to MAS with BLINK GWAS and by 21% compared to GBLUP without GWAS integration (Table 3).

Table 3.

Prediction accuracy using three methods on real trait: peach pollen sterilitya.

Method Mean Maximum Minimum Median SE
MAS 0.5289 0.6889 0.3944 0.5312 0.0070
GBLUP 0.5658 0.7028 0.4452 0.5851 0.0067
GAGBLUP 0.6855 0.8453 0.5080 0.6879 0.0083

aPrediction accuracies were examined using 5-fold cross-validation. Prediction accuracy is displayed for mean, median, minimum, and maximum among 40 replicates of cross-validation with five folds.

SE, standard error; MAS, marker-assisted selection; GAGBLUP: GWAS-Assisted Genomic Best Linear Unbiased Prediction.

To evaluate the performance of GAGBLUP across diverse genetic architectures, we compared the prediction accuracy of MAS, GBLUP, and GAGBLUP (Fig. 3) on traits simulated under varying heritability levels (75%, 50%, 25%) and numbers of QTNs (2, 10, 100, 150, 500, 1,000). The advantage of the GWAS + GS approach (GAGBLUP) diminishes when the number of causal genes is either very small (eg 5) or very large (eg >100). This advantage is also contingent on marker density. In both limiting scenarios, GAGBLUP performs as good as the better one between MAS and gBLUP. In all cases, GAGBLUP ensures the highest accuracy. It prevents losing the benefit of MAS when using GBLUP or reversely prevents losing the benefit of GBLUP when using MAS. With lower marker density, the probability of causal genes being tagged is reduced, further minimizing any potential advantage. Overall, GAGBLUP was the only method that consistently delivered the best or equally best prediction accuracy across all tested genetic architectures.

Figure 3.

For image description, please refer to the figure legend and surrounding text.

Prediction accuracy of GAGBLUP, MAS, and GBLUP under different genetic architectures. Heritability was set to 75%, 50%, and 25% to represent high, median, and low levels, respectively. The number of QTNs was set to 2, 10, 100, 150, 500, and 1,000. Forty repeated runs were performed for each combination of genetic architecture parameters. The difference value (rounded to two decimal places) between GAGBLUP and GBLUP is indicated next to the corresponding data points (or as marked in the figure).

Materials and methods

Two datasets were used in this study. One was peach data with 242 accessions phenotyped on pollen sterility and genotyped on 145,608 SNPs (Li et al. 2022, 2023). The peach dataset, with real phenotype data, was used to evaluate the performance of three prediction methods: MAS based on associated markers identified by GWAS using BLINK, GBLUP, and GAGBLUP, which incorporated the BLINK-identified associated markers as covariates.

The other dataset was from 282 maize inbred lines genotyped with the 55K SNP array (Cook et al. 2012). The dataset can be downloaded at Panzea website (https://cbsusrv04.tc.cornell.edu/users/panzea/download.aspx?filegroupid=7). After quality control, 51,182 SNPs remained in the maize dataset. Phenotype simulation was performed by two approaches. The one is with the 20 QTNs randomly selected from SNPs on the first five maize chromosomes, accounting for 75% heritability. The other is with the multiple combination of the number of the QTN (2, 10, 100, 150, 500, 1,000) and the levels of heritability (25%, 50%, and 75%). QTN effects were drawn from a standard normal distribution. Breeding values were computed for all individuals by multiplying their QTN genotypes by the corresponding QTN effects. Simulated phenotype values were then generated by adding the breeding values to residual values, which were sampled from a normal distribution with a mean of zero and a residual variance. The residual variance was calculated based on the variance of the breeding values and the specified heritability. The maize dataset was used to assess the performance of seven GWAS models and to investigate the prediction accuracy of integrating their GWAS results into GBLUP.

The redundancy among the associated markers identified by GWAS must be limited before their incorporation as covariates in GP using GBLUP. For the genotype matrix (X) with n (individuals) rows and m (markers sorted on their significance) columns, we applied QR decomposition (X = Q × R) to derive an orthogonal matrix Q (n by m) and an upper triangular matrix R (m by m). The markers that correspond to near-zero diagonal elements of R were removed. This filtering aims to select core markers with high contributions to the genotype matrix (Abdollahi-Arpanahi et al. 2022). The zero elements are determined by the product between the absolute maximum of the diagonal elements and alpha (which was 1e-10 as default). An element with absolute value below the threshold is considered as zero element. The GAGBLUP model constructs the kinship matrix using a subset of markers selected based on their estimated effect sizes derived from GWAS. Specifically, markers are filtered according to a threshold ratio applied to the distribution of absolute GWAS effect estimates. This threshold ratio—defined by the pv.cut parameter in the GAPIT package—can range from 0 to 1, with a default value of 0.2. This setting excludes the 20% of markers with the smallest estimated effects from the kinship matrix construction.

The performance of seven GWAS methods was evaluated using 40 replicates and all 282 individuals. The methods assessed were GLM, MLM, CMLM, SUPER, MLMM, FarmCPU, and BLINK. For each replicate, positives were categorized into true positives (SNPs among the 20 QTNs) and false positives (non-QTN SNPs). Power was calculated as the proportion of true-positive QTNs out of the 20 QTNs. The FDR was determined as the proportion of false-positive SNPs among the total positives, which included both false positives and true-positive QTNs.

Prediction accuracy was assessed by dividing the entire population from the peach and maize datasets into training and testing sets. Five-fold cross-validation was used for the evaluation. Accuracy (instant) was calculated for each fold and then averaged across the five folds (Zhou et al. 2017). GWAS was performed on the training population, which consisted of four folds. The associated markers identified in the training population were used as covariates to predict the testing population using GBLUP, with the phenotypes of the testing population masked. In the peach dataset, prediction accuracy was calculated as Pearson's correlation between the predicted breeding values and the actual phenotypes in the testing population. In the maize dataset, prediction accuracy was determined as Pearson's correlation between the predicted breeding values and the simulated breeding values in the testing population. The standard error based on the results of 40 repeated runs by computing the standard deviation and then dividing it by the square root of N-1.

GWAS models can be selected by specifying the “model” parameter, such as model = “BLINK” or model = c(“FarmCPU”, “BLINK”). When multiple GWAS models are specified, GAPIT executes each model independently. The input genotype data may include individuals with phenotypes for training or those without phenotypes for testing. Phenotypic data should include only individuals with phenotypes for training; if testing individuals are included in the phenotypic data, their phenotypes must be coded as NA. Predictions are generated for individuals both with and without phenotypes. The associated markers identified by GWAS are used as covariates for prediction. By default, a type I error threshold of cutOff = 0.05 is applied to determine the associated markers following Bonferroni multiple-test correction. QR was implemented in GAPIT. Users can use GAGBLUP method with two parameters (buspred = TRUE and lmpred = TRUE or FALSE). The “buspred” means whether use those significant markers to do GS. The “lmpred” means whether use general linear model or MLM to predict.

GAPIT outputs GBLUP, which predicts random individual effects based on a kinship structure derived from all markers. When a GWAS is incorporated, the breeding values include the sum of the effects of selected associated markers in addition to the GBLUP component. Additionally, the covariate variable (CV) data may contain inheritable factors—such as known genes or principal components derived from all markers—that should contribute to the breeding values. GAPIT requires that these additional genetic factors be positioned at the beginning of the CV data, enabling GAPIT to determine how many of these factors are inheritable and should be included in the breeding values. The remaining factors are treated as noninheritable.

GAPIT produces four output files that enable users to derive various components of the estimates and predictions. Phenotype y is conceptually decomposed as y = CVI + CVN + QTNs + GBLUP + e, where CVI are covariates of inheritable factors, CVN denotes covariates of noninheritable factors, QTNs are associated markers identified, GBLUP is the prediction of random individual effects, and the e represents residuals. The four standard outputs are BLUE = CVI + CVN, GBLUP, BreedingValue = CVI + QTNs + GBLUP, and Prediction = CVI + CVN + QTNs + GBLUP. Users can derive specific elements from these outputs. For example, e = y—Prediction, CVN = Prediction—BreedingValue, CVI = BLUE—CVN, and QTNs = BreedingValue—GBLUP—CVI.

Discussion

GP leverages all markers, irrespective of their statistical significance in GWAS. Integrating GWAS results into GP can improve accuracy, particularly by increasing marker density and sample size and, most notably, by advancing statistical GWAS methods (Bernardo 2014; Odilbekov et al. 2019; Merrick et al. 2021; Degen and Müller 2023). The updated GAPIT integrates GBLUP with GWAS results, enabling users to apply one or more GWAS methods without requiring manual coding. Previously, users had to perform GWAS separately, extract associated SNPs, filter them based on correlation, format a covariate input file, calculate the contributions of associated SNPs by extracting their effects and multiplying them by the corresponding SNP genotypes, and then compute total breeding values. Now, GAPIT streamlines this process, requiring only phenotype and genotype data from the user to complete all steps seamlessly.

Various strategies exist for integrating GWAS into GP. Some GP methods inherently incorporate GWAS methodologies, enabling a seamless integration. GP approaches can be broadly categorized into bottom-up and top-down methods. In the bottom-up approach, the effects of individual genetic markers are first estimated, and these marker effects are subsequently aggregated to calculate the breeding values of individuals. This category includes methods such as ridge regression, Least Absolute Shrinkage and Selection Operator, and Bayesian models, such as BayesA, BayesB, and BayesC (Pérez and De Los Campos 2014). Notably, the estimated effects of genetic markers or their statistical significance, tested against the null hypothesis of zero effects, can be directly applied in GWAS. The integration of GWAS into the GP framework has already been implemented for models in the bottom-up category. In contrast, the top-down approach, such as GBLUP, constructs kinship among individuals using genetic markers and predicts individual breeding values with a MLM, without initially incorporating GWAS results. However, the accumulation of known genes and the increasing power of GWAS have driven the need for GBLUP to integrate GWAS results to enhance prediction accuracy. The relatedness among associated markers diminishes their effectiveness as covariates in GBLUP. When dense marker sets are used, the number of associated markers can become excessive—potentially surpassing the number of individuals—which may lead to model overfitting or linear dependence. To address this, we applied the QR decomposition algorithm (Abdollahi-Arpanahi et al. 2022) to select relevant markers (see Materials and methods for details).

This study used simulated phenotypes in maize and real phenotypes in peach to compare the performance of GAGBLUP and GBLUP. An approach of simulated maize phenotype was controlled by 20 QTNs randomly selected from SNPs on the first five chromosomes. This design allowed us to examine how GAGBLUP integrates GWAS-derived insights with genome-wide polygenic breeding values. Since the last five chromosomes lacked QTNs, we could evaluate how false signals from these regions diminished GAGBLUP's prediction accuracy. This setup also disadvantaged GBLUP, as these chromosomes contributed no breeding value, diluting the kinship matrix instead. Although GAGBLUP relied on the same kinship matrix, it leveraged additional GWAS signals, mitigating this issue. This partially explains GAGBLUP's superior performance, improving prediction accuracy by over 40% compared to GBLUP in the simulated maize data, versus a 21% improvement in the real peach phenotypes. Other contributing factors may include heritability, the number of genes, and their effect distributions.

In this study, we also conducted systematic simulations specifically designed to evaluate performance across a spectrum of genetic architectures—from oligogenic to highly polygenic scenarios. The comparison of GAGBLUP, MAS, and GBLUP confirms that the advantage of GAGBLUP is not universal but is highly contingent on genetic architecture. Specifically, the results fully support the observation that GAGBLUP's relative advantage diminishes when the number of causal genes is either very small (eg 5) or very large (eg >100), and it performs as good as the better one MAS and GBLUP in these limiting scenarios. Most importantly, across all tested architectures, GAGBLUP ensures robust, high accuracy. It effectively integrates the strengths of MAS and GBLUP: It prevents losing the benefit of MAS when using GBLUP and, conversely, prevents losing the benefit of GBLUP when using MAS. Therefore, while it does not universally surpass both, it provides a reliable and safe integrative strategy. This finding provides direct, though preliminary, evidence that our integrated approach remains stable and applicable even under complex genetic backgrounds.

Conclusions

This study demonstrated that integrating GWAS into GP via GAPIT outperforms GBLUP in accuracy. Among seven GWAS methods, BLINK excelled, enhancing accuracy in simulated maize data by over 50% and in real peach data by 21%, due to higher statistical power and lower FDR. Multilocus models (eg BLINK and FarmCPU) outdid single-locus models by minimizing marker relatedness, optimized by QR decomposition to prevent overfitting. In peach, GAGBLUP with BLINK surpassed MAS and GBLUP. These findings affirm the efficacy of GWAS–GP integration, with GAPIT advancing breeding precision by incorporating multiple GWAS methods into GBLUP.

Contributor Information

Jiabo Wang, Key Laboratory of Qinghai-Tibetan Plateau Animal Genetic Resource Reservation and Utilization, Sichuan Province and Ministry of Education, Southwest Minzu University, Chengdu 610041, China; Department of Crop and Soil Sciences, Washington State University, Pullman, WA 99164, USA.

Zhiwu Zhang, Department of Crop and Soil Sciences, Washington State University, Pullman, WA 99164, USA.

Author contributions

J.W.: software, data curation, writing—original draft, visualization, testing, and validation. Z.Z.: conceptualization, methodology, supervision, and writing—review & editing. Both authors read and approved the final manuscript.

Funding

This work was partly supported by the Science and Technology Plan Project of the Xizang Autonomous Region, China (XZ202502ZY0001); the Endowment and Research Project (No. 126593) from the Washington Grain Commission and USDA award (2020-67021-32460); the National Key Research and Development Project of China, China (2022YFD1601601); the Heilongjiang Province Key Research and Development Project, China (2022ZX02B09); and the Fundamental Research Funds for the Central Universities, China (Southwest Minzu University, ZYN2024050).

Conflicts of interests

The authors have declared no competing interests.

Data availability

The GAPIT source code, demo script, and demo data are freely available on the GAPIT website (www.zzlab.net/GAPIT).

References

  1. Abdollahi-Arpanahi  R, Lourenco  D, Misztal  I. A comprehensive study on size and definition of the core group in the proven and young algorithm for single-step GBLUP. Genet Sel Evol.  2022:54:1–14. 10.1186/s12711-022-00726-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Bernardo  R. Genomewide selection when major genes are known. Crop Sci. 2014:54:68–75. 10.2135/cropsci2013.05.0315. [DOI] [Google Scholar]
  3. Chiwina  K  et al.  Genome-wide association study and genomic prediction of fusarium wilt resistance in common bean core collection. Int J Mol Sci.  2023:24:15300. 10.3390/ijms242015300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Cook  JP  et al.  Genetic architecture of maize kernel composition in the nested association mapping and inbred association panels. Plant Physiol. 2012:158:824–834. 10.1104/pp.111.185033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Degen  B, Müller  NA. A simulation study comparing advanced marker-assisted selection with genomic selection in tree breeding programs. G3: Genes, Genomes, Genetics. 2023:13:1–9. 10.1093/g3journal/jkad164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Fernando  RL, Garrick  D. Bayesian methods applied to GWAS BT - genome-wide association studies and genomic prediction. In: Gondro  C, van der Werf  J, Hayes  B  Humana Press; 2013. p. 237–274. [DOI] [PubMed] [Google Scholar]
  7. Huang  M, Liu  X, Zhou  Y, Summers  RM, Zhang  Z. BLINK: a package for the next level of genome-wide association studies with both individuals and markers in the millions. Gigascience. 2019:8:giy154. 10.1093/gigascience/giy154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Li  M  et al.  Enrichment of statistical power for genome-wide association studies. BMC Biol. 2014:12:73. 10.1186/s12915-014-0073-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Li  X  et al.  Single nucleotide polymorphism detection for peach gummosis disease resistance by genome-wide association study. Front Plant Sci.  2022:12:763618. 10.3389/fpls.2021.763618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Li  X  et al.  Multiple-statistical genome-wide association analysis and genomic prediction of fruit aroma and agronomic traits in peaches. Hortic Res.  2023:10:uhad117. 10.1093/hr/uhad117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Lipka  AE  et al.  GAPIT: genome association and prediction integrated tool. Bioinformatics. 2012:28:2397–2399. 10.1093/bioinformatics/bts444. [DOI] [PubMed] [Google Scholar]
  12. Liu  X, Huang  M, Fan  B, Buckler  ES, Zhang  Z. Iterative usage of fixed and random effect models for powerful and efficient genome-wide association studies. PLoS Geneticsenetics. 2016:12:e1005767. 10.1371/journal.pgen.1005767. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Merrick  LF, Burke  AB, Chen  X, Carter  AH. Breeding with major and minor genes: genomic selection for quantitative disease resistance. Front Plant Sci.  2021:12:1–22. 10.3389/fpls.2021.713667. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Nishio  M, Satoh  M. Genomic best linear unbiased prediction method including imprinting effects for genomic evaluation. Genet Sel Evol.  2015:47:1–10. 10.1186/s12711-015-0091-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Odilbekov  F, Armoniené  R, Koc  A, Svensson  J, Chawade  A. GWAS-Assisted Genomic prediction to predict resistance to Septoria tritici blotch in Nordic winter wheat at seedling stage. Front Genet.  2019:10:1224. 10.3389/fgene.2019.01224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Pérez  P, De Los Campos  G. Genome-wide regression and prediction with the BGLR statistical package. Genetics. 2014:198:483–495. 10.1534/genetics.114.164442. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Pritchard  JK, Stephens  M, Donnelly  P. Inference of population structure using multilocus genotype data. Genetics. 2000:155:945–959. 10.1093/genetics/155.2.945. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Reyes-Valdés  MH. A model for marker-based selection in gene introgression breeding programs. Crop Sci. 2000:40:91–98. 10.2135/cropsci2000.40191x. [DOI] [Google Scholar]
  19. Rice  B, Lipka  AE. Evaluation of RR-BLUP genomic selection models that incorporate peak genome-wide association study. Signals in maize and sorghum. Plant Genome.  2019:12:1–14. 10.3835/plantgenome2018.07.0052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Segura  V  et al.  An efficient multi-locus mixed-model approach for genome-wide association studies in structured populations. Nat Genet.  2012:44:825–830. 10.1038/ng.2314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Spindel  JE  et al.  Genome-wide prediction models that incorporate de novo GWAS are a powerful new tool for tropical rice improvement. Heredity (Edinb).  2016:116:395–408. 10.1038/hdy.2015.113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Tang  Y  et al.  GAPIT version 2: an enhanced integrated tool for genomic association and prediction. Plant Genome. 2016:9:1–9. 10.3835/plantgenome2015.11.0120. [DOI] [PubMed] [Google Scholar]
  23. Wang  J, Zhang  Z. GAPIT version 3: boosting power and accuracy for genomic association and prediction. Genomics Proteomics Bioinformatics. 2021:19:629–640. 10.1016/j.gpb.2021.08.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Wang  J  et al.  Expanding the BLUP alphabet for genomic prediction adaptable to the genetic architectures of complex traits. Heredity (Edinb).  2018:121:648–662. 10.1038/s41437-018-0075-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Wang  Q, Tian  F, Pan  Y, Buckler  ES, Zhang  Z. A SUPER powerful method for genome wide association study. PLoS One. 2014:9:e107684. 10.1371/journal.pone.0107684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Wolc  A  et al.  Mixture models detect large effect QTL better than GBLUP and result in more accurate and persistent predictions. J Anim Sci Biotechnol.  2016:7:7. 10.1186/s40104-016-0066-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Yu  J  et al.  A unified mixed-model method for association mapping that accounts for multiple levels of relatedness. Nat Genet.  2006:38:203–208. 10.1038/ng1702. [DOI] [PubMed] [Google Scholar]
  28. Zhang  Z  et al.  Mixed linear model approach adapted for genome-wide association studies. Nat Genet.  2010:42:355–360. 10.1038/ng.546. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Zhou  Y, Isabel Vales  M, Wang  A, Zhang  Z. Systematic bias of correlation coefficient may explain negative accuracy of genomic prediction. Brief Bioinform.  2017:18:744–753. 10.1093/bib/bbw064. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The GAPIT source code, demo script, and demo data are freely available on the GAPIT website (www.zzlab.net/GAPIT).


Articles from Molecular Biology and Evolution are provided here courtesy of Oxford University Press

RESOURCES