Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2026 Feb 23;27:321. doi: 10.1186/s12864-026-12685-z

Integration and imputation of GBS-derived and DArTseq-derived SNP markers in assessing genetic diversity of bread wheat genotypes

Hossein Abdi 1, Hadi Alipour 1,, Iraj Bernousi 1, Reza Darvishzadeh 1, Sima Fatanatvash 1, Aras Türkoğlu 2
PMCID: PMC13036897  PMID: 41731365

Abstract

Background

Wheat (Triticum aestivum L.) is a globally paramount crop. Iranian landraces serve as a vital resource for enriching wheat gene banks worldwide, and deciphering the diversity in its genotypes is crucial for breeders. Genotyping-by-Sequencing (GBS) and Diversity Array Technology (DArT) are two important platforms for generating single nucleotide polymorphisms (SNP) markers. The integration of molecular marker data from different genotyping platforms is crucial for a comprehensive analysis of genetic variation in wheat germplasm. The aim of this study was to integrate and impute SNP markers derived from GBS and DArTseq platforms, and to employ the dataset for assessing the genetic diversity of Iranian bread wheat genotypes and for detecting selection signatures.

Results

This study integrated molecular marker data from two genotyping platforms (GBS and DArTseq) through imputation to enable a unified analysis of genetic diversity in bread wheat germplasm. We first imputed missing data for 357 Iranian bread wheat accessions genotyped via GBS. This process more than tripled the number of usable SNP markers obtained through GBS. Subsequently, we imputed markers for the remaining genotypes using a reference set of 90 accessions genotyped with DArTseq technology. These sequential imputation steps yielded a consolidated dataset of 46,876 high-quality GBS-derived SNP and 3,417 high-quality DArTseq-derived SNP markers. The results obtained from the two marker systems demonstrated a high degree of complementarity, effectively distinguishing cultivars from landraces. Furthermore, cluster analysis delineated the genotypes into three distinct groups. Furthermore, these markers were used to identify signatures of natural and artificial selection by detecting high Fst values. Our results showed that the genomic regions under selection, identified by SNPs contain genes involved in regulatory processes related to DNA transcription, cell wall organization, protein phosphorylation, and defense response to biotic stresses. These pathways are particularly significant in the differentiation of populations in response to environmental pressures. In contrast, genes associated with DArTseq-derived SNP markers were mainly involved in more general pathways such as transcription regulation and cell structure processes, which may indicate the lower sensitivity of this system in detecting directional selection. Nevertheless, the identification of distinct selection signatures by DArTseq-derived SNP markers underscores their complementary role in genomic studies.

Conclusions

The presented framework enables effective integration of multi-platform marker data, enhancing genetic diversity assessment and revealing new selection signatures in wheat. The resulting imputed dataset forms a foundational resource for subsequent genome-wide association and genomic selection studies.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-026-12685-z.

Keywords: Genotyping-by-Sequencing, Iranian wheat landraces and cultivars, Missing data, Selection signature

Introduction

The nutritional and economic significance of bread wheat is well-established today, and efforts to enhance its yield continue relentlessly. Iran, being one of the centers of wheat genetic diversity, possesses a wide range of accessions that can play a significant role in breeding programs [1]. The success of breeding programs depends largely on the presence and identification of genetic diversity in wheat genotypes. For the past hundred years, diversity has been identified by recording phenotypic traits, and more recently, with modern molecular tools. Molecular markers serve as powerful tools in various domains of genetics and genomics, including genome sequencing, DNA fingerprinting, and genetic mapping. Furthermore, they are instrumental in identifying genes that control specific traits, quantitative trait locus (QTL) mapping, determining heterotic groups, investigating kinship relationships, and facilitating both marker-assisted selection and genomic selection [25]. The use of molecular markers, in addition to being stable and reliable, also saves time and increases the efficiency of breeding programs. Single nucleotide polymorphisms (SNPs) are the most abundant type of marker in plant genomes and offer several advantages over other molecular markers. SNP markers result from a single nucleotide difference between two DNA sequences, reflecting the true nature of allelic variation [6].

So far, a variety of high-throughput platforms have been introduced for genotyping plant samples. Two common examples are Genotyping-by-Sequencing (GBS) and Diversity Array Technology (DArTseq). GBS facilitates the discovery of novel SNPs in plant species with complex genomes, without requiring prior sequence information [7]. Its cost-effectiveness and high SNP density make it particularly well-suited for studying species with large genomes and high biodiversity [8]. A recognized challenge of GBS is its high missing data rate; however, this can be effectively mitigated using various imputation strategies [9]. On the other hand, DArTseq is recognized for its high capacity to detect single-nucleotide differences at a genome-wide scale and performs effectively in species with large and complex genomes like wheat [10, 11]. Several studies have confirmed the effectiveness of GBS-derived SNP [1215] and DArTseq-derived SNP [1, 1619] markers in assessing wheat genetic diversity. Research further suggests that integrating these two marker systems can provide a more comprehensive understanding of genetic variation in wheat [2023].

Overall, both DArTseq and genotyping-by-sequencing are considered high-throughput single nucleotide polymorphisms discovery techniques, but they differ in terms of their data generation mechanism, genomic coverage rate, raw data volume, and bioinformatics resource requirements. The choice between them depends on the study objectives, the complexity of the plant genome, the level of financial resources, and the availability of sequencing equipment. In polyploid and non-model species such as wheat, combining data from both methods can provide richer complementary genetic information for breeders. On the other hand, in some countries, limited budgets and infrastructure often constrain researchers’ ability to perform large-scale genotyping. Therefore, effectively integrating the expanding genetic and genomic resources from diverse platforms and germplasm panels is crucial for leveraging their full potential in breeding programs [24]. In general, all genome-wide genotyping technologies are somewhat complementary and can be integrated through imputation [25]. Torkamaneh and Belzile [26] conceptually defined two types of missing data: (i) missing genotype, where certain individuals lack a genotype call at a locus that was successfully genotyped in other individuals within the population; and (ii) untyped locus, where data for a specific SNP is absent for all individuals in the population, with the possible exception of a few individuals common to different datasets. The imputation of these missing data types, based on the available genetic information, is crucial for ensuring the integrity of downstream analyses. Mainly, missing data imputation serves various purposes, including facilitating high-throughput sequencing mapping, combining different genotyping arrays, transitioning from low-density to high-density chips, detecting genotyping errors, reducing costs, increasing statistical power in association studies, and enabling meta-analyses.

This study had three primary objectives. First, we sought to increase the density of GBS-derived SNP markers through imputation. Second, we utilized available DArTseq-derived SNP marker data from a subset of genotypes to impute this data across the entire panel. Finally, by applying both imputed SNP datasets to a collection of 357 wheat genotypes, we aimed to assess genetic diversity and identify genomic regions under natural or artificial selection. This integrated approach enabled a direct and robust comparison of the two marker systems within an identical genetic background.

Materials and methods

Plant materials

Two large collections of Iranian wheat accessions, previously genotyped using different platforms, were considered in this study. One collection consisted of 357 accessions (87 cultivars and 270 landraces) genotyped by GBS technology [12]. The cultivars, which represent diverse breeding origins, have been released for cultivation in various climatic regions of Iran over the past century. The landraces were collected to be representative of all geographical regions of Iran. All landraces were obtained as accessions from the USDA National Plant Germplasm System. Further details for these accessions are provided in (Supplementary Table 1). In another collection, 2,403 Iranian wheat landraces from the genebank of the International Maize and Wheat Improvement Center (CIMMYT) were evaluated using DArTseq technology [27]. These landraces were selected from a collection of 6,800 accessions, the details of which are available in (Supplementary Table 2). A comparison of the two sets revealed 90 common accessions (Supplementary Table 3). Using these shared accessions, we completed the genotyping information for a total of 357 Iranian wheat accessions and assessed their genetic diversity using both GBS-derived and DArTseq-derived SNP markers.

Genotyping

GBS-based genotyping

SNP markers were identified in a set of 357 accessions using GBS, as previously described by Alipour et al. [12]. The Poland et al. [28] approach was employed, utilizing two restriction enzymes, PstI (CTGCAG) and MspI (CCGG). After digestion by enzymes, barcoded adapters were ligated to each DNA sample using T4 ligase (New England Bio-Labs Inc.). The size-selected library was sequenced on an Ion Proton system (Life Technologies Inc.). After trimming the sequencing reads to 64 bp, identical reads were collapsed into sequence tags. The unique tags were then internally aligned with a tolerance of up to 3 bp mismatches to detect single nucleotide polymorphisms. SNP calling was performed using the Universal Network Enabled Analysis Kit (UNEAK) GBS pipeline [29] within the TASSEL 3.0 bioinformatics package [30]. The filtering steps were applied as follows: reads with a quality score below 15 were discarded, and only SNPs with a heterozygosity rate of less than 10%, a minor allele frequency (MAF) above 1%, and less than 20% missing data were retained for subsequent analysis.

DArTseq-based genotyping

A diverse panel of 2,403 accessions was genotyped using the DArTseq platform developed by Singh et al. [27] at CIMMYT (https://hdl.handle.net/11529/10045). The methodology employed two complexity reduction methods, as described by Sansaloni et al. [31], utilizing the restriction enzymes PstI and HpaII. Following digestion, barcode adapters were ligated to the samples to enable multiplexing of 96 samples in a single lane of an Illumina HiSeq 2500 system (Illumina Inc., San Diego, CA). The initial marker filtering was performed using a proprietary analytical pipeline, which applied thresholds for reproducibility, call rate, and average read depth. For details on the germplasm and genotyping, refer to previous studies [1, 32]. Following initial data acquisition, we applied the same filtering criteria to the DArTseq dataset.

Imputation and integration of genotyping data

A major limitation of the GBS approach is the prevalence of missing data. To address this, we accurately imputed the missing SNPs using BEAGLE [33] software (V3.3.2) and the w7984 reference genome [34]. In the BEAGLE, missing data are first imputed for haplotypes based on allele frequencies within the population. The initial imputed data are then used to construct haplotype-cluster models, which represent a specific class of hidden Markov models (HMMs). Each model is applied to a single chromosome, with the number of levels in the model corresponding to the number of markers. This approach enables rapid and accurate estimation of missing alleles. The imputation method used in this study was adapted from the work of Alipour et al. [9]. Briefly, sequence tags were aligned to the reference genome using BLASTn (nucleotide BLAST). Following haplotype phasing for all individuals, SNPs were ranked based on their allele frequency. For SNPs that mapped to multiple chromosomal positions, the position with the lowest E-value was selected. Following the application of various quality filters, this process yielded a final set of 46,876 high-quality GBS-derived SNPs. On the other hand, the DArTseq-derived SNP markers for the remaining samples (87 cultivars and 180 landraces) were imputed using a reference set of 90 accessions for which both GBS and DArTseq data were available. A total of 3,417 high-quality DArTseq-derived SNP markers were successfully generated for all 357 accessions. Finally, these two marker systems were compared in assessing genetic diversity.

Statistical analysis

The rates of transition (Ts; A↔G and T↔C) and transversion (Tv; A↔T, A↔C, T↔G, and C↔G) mutations, along with their ratios, were calculated using Microsoft Excel. Marker density was calculated as the number of SNPs per megabase (Mb) for each chromosome, using the physical positions of markers on the reference genome. Density plots were generated with a sliding window of 1 Mb to visualise genome‑wide SNP distribution. A suite of population genetic parameters was estimated using the adegenet package in R-4.4.2 program [35]. These parameters included: observed heterozygosity (Ho), expected heterozygosity (Hs), expected heterozygosity under random mating (Ht), gene diversity among samples (Dst), corrected Ht (Htp), corrected Dst (Dstp), the fixation index (Fst), corrected Fst (Fstp), the inbreeding coefficient (Fis), and Jost’s D (Dest). Moreover, to partition the calculated genetic variation between landrace and cultivar groups, an analysis of molecular variance (AMOVA) along with the genetic differentiation index (PhiPT) was performed using the pegas package in R software [36]. The generation of the kinship matrix and its heatmap visualization were performed using the GAPIT R-package [37]. This matrix represents the genetic relationships between individuals based on their proportion of shared alleles. Principal component analysis (PCA) and hierarchical cluster analysis with Euclidean distance and Ward’s minimum variance method (specifically Ward.D2) were performed using the factoextra R package [38]. Subsequently, the PCA biplot was generated using the component scores in Excel.

Gene annotation and KEGG enrichment analysis

The Fst statistic measures the reduction in heterozygosity within subpopulations relative to the total population. Loci with extreme Fst values are potential targets of divergent selection, either natural or artificial [39]. Consequently, in this study, sequences surrounding high-Fst markers were retrieved from the EnsemblPlants database (http://plants.ensembl.org) and annotated using the BLAST tool against the International Wheat Genome Sequencing Consortium (IWGSC) RefSeq v1.0 reference genome [40]. The BLASTn results were filtered using an expect value (E-value) threshold of < 1e− 10 and a percent identity threshold of > 95%. Gene Ontology (GO) terms, encompassing molecular function, biological process, and cellular component, were obtained for the selected genes from the Ensembl-Gramene database (http://ensembl.gramene.org). These genes were subsequently analyzed using the DAVID bioinformatics database (https://david.ncifcrf.gov/) for GO enrichment analysis. Furthermore, the Kyoto Encyclopedia of Genes and Genomes (KEGG; https://www.genome.jp/kegg/) was employed to identify potential biological pathways [41, 42]. This tool maps genes to predefined pathways based on their functional annotations. Significantly enriched GO terms and KEGG pathways were reported based on a significance threshold of P < 0.05.

Results

Nucleotide substitution rates and genetic diversity and differentiation measures

A total of 45,838 GBS-derived SNP and 3,417 DArTseq-derived SNP markers were mapped across the various chromosomes and sub-genomes of wheat. For both marker types, the B sub-genome contained the highest marker density, followed by the A and D sub-genomes. Conversely, chromosomes 4D, 5D, and 3D had the lowest marker coverage. Specifically, chromosomes 3B (4,213), 2B (3,967), and 6B (3,934) possessed the highest number of GBS-derived SNP markers, while chromosomes 1B (287), 2B (282), and 7 A (274) had the highest number of DArTseq-derived SNP markers. Across the entire genome, the rate of transition mutations was higher than that of transversion mutations. The transition/transversion (Ts/Tv) ratio was 2.25 for GBS-derived SNP markers and 1.75 for DArTseq-derived SNP markers. This trend was also consistent within the individual sub-genomes. Interestingly, sub-genome A exhibited the highest Ts/Tv ratio among the three sub-genomes. At the chromosomal level, this ratio ranged from 2.66 (chromosome 2 A) to 1.66 (chromosome 1D) based on GBS-derived SNP markers, and from 2.23 (chromosome 5D) to 1.22 (chromosome 6D) based on DArTseq-derived SNP markers (Table 1).

Table 1.

Statistical summary of studied wheat genotypes based on 46,877 GBS-derived SNP and 3,417 DArTseq-derived SNP markers

Chromosome SNP$ DArT
NM Ts (%) Tv (%) Ts/Tv NM Ts (%) Tv (%) Ts/Tv
1 A 2410 1734 (72.0) 676 (28.0) 2.57 213 139 (65.3) 74 (34.7) 1.88
1B 3141 2171 (69.1) 970 (30.9) 2.24 287 185 (64.5) 102 (35.5) 1.81
1D 1002 626 (62.5) 376 (37.5) 1.66 80 49 (61.3) 31 (38.8) 1.58
2 A 2884 2096 (72.7) 788 (27.3) 2.66 203 130 (64.0) 73 (36.0) 1.78
2B 3967 2708 (68.3) 1259 (31.7) 2.15 282 169 (59.9) 113 (40.1) 1.50
2D 1438 922 (64.1) 516 (35.9) 1.79 86 50 (58.1) 36 (41.9) 1.39
3 A 2033 1403 (69.0) 630 (31.0) 2.23 177 116 (65.5) 61 (34.5) 1.90
3B 4213 2988 (70.9) 1225 (29.1) 2.44 260 178 (68.5) 82 (31.5) 2.17
3D 780 498 (63.8) 282 (36.2) 1.77 47 29 (61.7) 18 (38.3) 1.61
4 A 2741 1928 (70.3) 813 (29.7) 2.37 183 115 (62.8) 68 (37.2) 1.69
4B 1503 1039 (69.1) 464 (30.9) 2.24 112 69 (61.6) 43 (38.4) 1.60
4D 284 197 (69.4) 87 (30.6) 2.26 27 15 (55.6) 12 (44.4) 1.25
5 A 1505 1011 (67.2) 494 (32.8) 2.05 171 111 (64.9) 60 (35.1) 1.85
5B 3161 2177 (68.9) 984 (31.1) 2.21 213 136 (63.8) 77 (36.2) 1.77
5D 646 406 (62.8) 240 (37.2) 1.69 42 29 (69.0) 13 (31.0) 2.23
6 A 2053 1424 (69.4) 629 (30.6) 2.26 142 94 (66.2) 48 (33.8) 1.96
6B 3934 2764 (70.3) 1170 (29.7) 2.36 237 152 (64.1) 85 (35.9) 1.79
6D 797 508 (63.7) 289 (36.3) 1.76 60 33 (55.0) 27 (45.0) 1.22
7 A 3170 2224 (70.2) 946 (29.8) 2.35 274 180 (65.7) 94 (34.3) 1.91
7B 3190 2278 (71.4) 912 (28.6) 2.50 219 138 (63.0) 81 (37.0) 1.70
7D 986 626 (63.5) 360 (36.5) 1.74 102 60 (58.8) 42 (41.2) 1.43
Genome A 16,796 11,820 (70.4) 4976 (29.6) 2.38 1363 885 (64.9) 478 (35.1) 1.85
Genome B 23,109 16,125 (69.8) 6984 (30.2) 2.31 1610 1027 (63.8) 583 (36.2) 1.76
Genome D 5933 3783 (63.8) 2150 (36.2) 1.76 444 265 (59.7) 179 (40.3) 1.48
Total 45,838 31,728 (69.2) 14,110 (30.8) 2.25 3417 2177 (63.7) 1240 (36.3) 1.75

NM Number of markers, Ts Transition, and Tv Transversion

$1,038 SNP markers were located on the unknown chromosome

Analysis of population genetic parameters revealed that the observed heterozygosity (Ho) based on GBS-derived SNP markers was highest for chromosomes 2B, 3 A, 5 A, 5B, and 7B, and lowest for chromosomes 4B, 3D, and 4 A. In contrast, based on DArTseq-derived SNP markers, chromosome 2B exhibited the highest Ho, while chromosome 5D showed the lowest. A comparison of the two marker systems indicated that the D sub-genome had a higher Ho value when assessed with GBS-derived SNP than with DArTseq-derived SNP markers. This pattern was reversed in sub-genomes A and B. The genome-wide expected heterozygosity (He) for GBS-derived SNP markers (0.295) compared to DArTseq-derived SNP markers (0.212). Among chromosomes, this statistic for GBS-derived SNP markers ranged from 0.238 (2D) to 0.34 (2 A), whereas for DArTseq-derived SNP markers, it varied from 0.111 (4B) to 0.246 (5B). The estimated values for the expected heterozygosity under random mating (Ht) and the corrected Ht (Htp) were similar to the He values (Table 2). The frequency distribution of individuals and markers across different heterozygosity values was compared (Fig. 1). Although both marker systems showed the highest frequencies at low heterozygosity, their distribution patterns differed. With increasing heterozygosity, the frequency of individuals based on DArTseq-derived SNP markers declined uniformly. Conversely, GBS-derived SNP data showed a minor peak in individual frequency at 0.05 heterozygosity. Furthermore, the heterozygosity range differed markedly between the two marker systems. GBS-derived SNP markers displayed a very low frequency at a heterozygosity of 0.25, whereas no DArTseq-derived SNP marker exceeded a heterozygosity of 0.12.

Table 2.

Statistical summary of studied wheat genotypes based on 46,877 GBS-derived SNP and 3,417 DArTseq-derived SNP markers

Chromosome Ho Hs Ht Htp Dst Dstp Fst Fstp Fis Dest
SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT SNP DArT
1 A 0.027 0.032 0.270 0.222 0.303 0.239 0.336 0.256 0.033 0.017 0.067 0.034 0.085 0.047 0.143 0.082 0.900 0.840 0.100 0.052
1B 0.027 0.029 0.308 0.218 0.336 0.233 0.364 0.248 0.028 0.015 0.056 0.030 0.068 0.041 0.119 0.073 0.908 0.828 0.089 0.047
1D 0.025 0.021 0.276 0.216 0.299 0.227 0.321 0.238 0.022 0.011 0.045 0.021 0.063 0.037 0.112 0.068 0.909 0.881 0.067 0.032
2 A 0.027 0.031 0.340 0.239 0.368 0.258 0.396 0.276 0.028 0.018 0.057 0.037 0.067 0.045 0.115 0.079 0.912 0.820 0.089 0.058
2B 0.029 0.035 0.310 0.232 0.341 0.247 0.372 0.262 0.031 0.015 0.062 0.030 0.078 0.040 0.136 0.073 0.903 0.791 0.097 0.048
2D 0.024 0.018 0.238 0.156 0.260 0.168 0.282 0.180 0.022 0.012 0.044 0.024 0.066 0.040 0.116 0.071 0.901 0.843 0.064 0.035
3 A 0.029 0.031 0.289 0.175 0.322 0.186 0.355 0.197 0.033 0.011 0.066 0.023 0.084 0.031 0.144 0.055 0.897 0.709 0.099 0.036
3B 0.027 0.031 0.301 0.225 0.350 0.249 0.399 0.274 0.049 0.025 0.097 0.050 0.112 0.060 0.183 0.099 0.908 0.809 0.147 0.076
3D 0.023 0.021 0.239 0.184 0.261 0.195 0.283 0.207 0.022 0.012 0.044 0.023 0.063 0.042 0.110 0.075 0.904 0.862 0.066 0.033
4 A 0.023 0.028 0.307 0.210 0.343 0.226 0.378 0.242 0.035 0.016 0.071 0.032 0.088 0.041 0.154 0.072 0.920 0.799 0.111 0.051
4B 0.022 0.016 0.243 0.111 0.265 0.117 0.287 0.123 0.022 0.006 0.044 0.012 0.065 0.018 0.114 0.034 0.908 0.723 0.066 0.020
4D 0.025 0.017 0.272 0.152 0.305 0.174 0.337 0.196 0.033 0.022 0.066 0.044 0.085 0.066 0.148 0.117 0.909 0.864 0.101 0.067
5 A 0.029 0.028 0.294 0.193 0.314 0.201 0.334 0.209 0.020 0.008 0.040 0.016 0.057 0.025 0.101 0.046 0.900 0.753 0.060 0.025
5B 0.029 0.032 0.314 0.246 0.353 0.263 0.391 0.279 0.038 0.017 0.077 0.033 0.091 0.044 0.153 0.078 0.906 0.844 0.118 0.053
5D 0.024 0.014 0.258 0.128 0.288 0.143 0.319 0.158 0.030 0.015 0.061 0.030 0.082 0.046 0.139 0.079 0.906 0.866 0.089 0.044
6 A 0.027 0.024 0.300 0.234 0.327 0.255 0.354 0.275 0.027 0.020 0.053 0.040 0.067 0.052 0.116 0.090 0.909 0.878 0.083 0.062
6B 0.025 0.027 0.303 0.242 0.338 0.260 0.373 0.278 0.035 0.018 0.070 0.036 0.086 0.048 0.144 0.087 0.914 0.876 0.106 0.056
6D 0.028 0.022 0.287 0.216 0.305 0.225 0.323 0.234 0.018 0.009 0.037 0.018 0.052 0.033 0.093 0.062 0.902 0.884 0.056 0.027
7 A 0.027 0.032 0.293 0.213 0.326 0.232 0.360 0.252 0.033 0.020 0.066 0.039 0.082 0.050 0.139 0.088 0.902 0.790 0.101 0.060
7B 0.029 0.032 0.287 0.198 0.310 0.206 0.333 0.213 0.023 0.008 0.046 0.015 0.060 0.024 0.104 0.045 0.898 0.802 0.069 0.024
7D 0.026 0.023 0.268 0.204 0.289 0.222 0.309 0.239 0.021 0.017 0.041 0.035 0.057 0.048 0.101 0.087 0.906 0.859 0.063 0.056
A Genome 0.027 0.030 0.301 0.213 0.331 0.229 0.362 0.245 0.031 0.016 0.061 0.032 0.077 0.042 0.132 0.074 0.906 0.797 0.094 0.050
B Genome 0.027 0.030 0.300 0.219 0.334 0.234 0.368 0.250 0.034 0.016 0.068 0.031 0.083 0.041 0.140 0.073 0.907 0.817 0.104 0.049
D Genome 0.025 0.020 0.260 0.186 0.283 0.200 0.305 0.213 0.023 0.013 0.045 0.027 0.064 0.043 0.113 0.077 0.905 0.864 0.068 0.041
Whole Genome 0.027 0.029 0.295 0.212 0.326 0.227 0.358 0.243 0.031 0.015 0.063 0.031 0.078 0.042 0.134 0.074 0.906 0.815 0.096 0.048

Ho Observed heterozygosity, Hs Expected heterozygosity, Ht Expected heterozygosity in the random-mating, Htp Corrected Ht, Dst Gene diversity among samples, Dstp Corrected Dst, Fst Fixation index, Fstp Corrected Fst, Fis Inbreeding coefficient, and Dest Jost’s D

Fig. 1.

Fig. 1

Comparison of heterozygosity in markers and individuals using GBS-derived SNP (A) and DArTseq-derived SNP (B) data

The genetic diversity (Dst) and corrected Dst (Dstp) estimated with DArTseq-derived SNP markers was approximately half that obtained with GBS-derived SNPs. The distribution of Dst and Dstp across chromosomes also differed between the two marker systems. Based on GBS-derived SNP data, the highest Dst values were observed on chromosomes 6B (0.035), 5B (0.038), and 3B (0.049), while the lowest were on chromosomes 6D (0.018), 5 A (0.020), and 7D (0.021). In contrast, for DArTseq-derived SNP markers, the maximum Dst value was found on chromosome 3B (0.025), and the minimum on chromosome 4B (0.006). We also analyzed the fixation index (Fst) and the corrected Fst (Fstp). The values for both statistics were higher when estimated with GBS-derived SNP markers compared to DArTseq-derived SNP markers. Based on DArTseq-derived SNP markers, the Fst did not differ significantly among sub-genomes. In contrast, estimates from GBS-derived SNP markers revealed a significantly higher Fst in sub-genome B (0.083) than in sub-genomes A (0.077) and D (0.064). A more detailed examination of the GBS-derived SNP marker data showed that chromosome 3B (0.112) had the maximum Fst value, while chromosomes 6D (0.052), 5 A (0.057), and 7D (0.057) had the minimum. Conversely, with DArTseq-derived SNP markers, chromosome 4B (0.018) had the lowest Fst and chromosome 4D (0.066) the highest. The application of the Fstp correction led to an overall increase in the fixation index values. The inbreeding coefficient (Fis) was 0.906 based on GBS-derived SNP markers and 0.815 based on DArTseq-derived SNP markers at the whole-genome level. This index showed little variation among chromosomes, and for both marker types, chromosome 3 A exhibited the lowest Fis value.

Analysis using Jost’s D (Dest) revealed higher genetic diversity with GBS-derived SNP markers compared to DArTseq-derived SNP markers. Based on GBS-derived SNP data, the sub-genomes were ranked as B (0.104) > A (0.094) > D (0.068). In contrast, DArTseq-derived SNP markers showed similar, yet lower, values for sub-genomes A (0.050), B (0.049), and D (0.049). Furthermore, the range of Dest values across individual chromosomes was broader for GBS-derived SNPs (0.056 on 6D to 0.147 on 3B) than for DArTseq-derived SNP markers (0.020 on 4B to 0.076 on 3B) (Table 2). The frequency distribution of markers across different density categories is shown in Fig. 2. As anticipated, the frequency was very high at low densities and decreased sharply as density increased. The cumulative frequency of DArTseq-derived SNP markers at low densities increased with a gentler slope compared to that of GBS-derived SNP markers.

Fig. 2.

Fig. 2

Frequency of GBS-derived SNP (A) and DArTseq-derived SNP (B) markers at different densities

Analysis of molecular variance

MANOVA results revealed a significant difference between Iranian wheat cultivars and landraces populations using both GBS-derived SNP and DArTseq-derived SNP marker systems (P < 0.0001). Notably, the PhiPT value for GBS-derived SNP markers (0.182) was slightly higher than that of DArTseq-derived SNP markers (0.128). Considering spring and winter wheat cultivars and landraces collected 1,000 km apart, within-population diversity was substantially higher than between-population diversity for both marker systems (Table 3).

Table 3.

Analysis of molecular variance (AMOVA) between wheat cultivars and landraces based on GBS-derived SNP and DArTseq-derived SNP markers

Source df SS MS Variance components PhiPT p-value
SNP Among Pops 1 1585975.5 1585975.5 11,654 0.182 < 0.0001
Within Pops 355 18,572,335 52316.4 52,316
Total 356 20158310  56624.5
DArT Among Pops 1 58359.9 58359.9 421.6 0.128 < 0.0001
Within Pops 355 1022652.4 2880.7 2880.7
Total 356 1081012.3 3036.5

Df degrees of freedom, SS sum of squares, MS mean squared, PhiPT genetic differentiation index among population, and Pops populations

PCA, cluster, and kinship analysis

The first two principal components collectively explained 23% and 19.5% of the total genetic variation for the GBS-derived SNP and DArTseq-derived SNP markers, respectively. The wide distribution of genotypes in these biplots indicates a high degree of genetic diversity among them. The PCA biplot revealed a clear separation between cultivars and landraces for both molecular marker types, a distinction that was more pronounced with GBS-derived SNP markers. However, some cultivars were in a similar position to the landraces (Fig. 3). Cluster analysis using both SNP markers segregated the genotypes into three major groups, with the second group containing the majority of cultivars in both cases. A comparison of the two dendrograms showed that 38 genotypes were grouped differently when using DArTseq-derived SNP markers compared to GBS-derived SNP markers (Fig. 4). Consistent with the cluster analysis, the kinship heatmap revealed that the studied germplasm was divided into three main groups, highlighting both the genetic relationships and distinctions between them. The intensity of the red color along the heatmap’s diagonal represents the degree of genetic relatedness, as estimated by the GBS-derived SNP and DArTseq-derived SNP markers (Fig. 5).

Fig. 3.

Fig. 3

PCA biplot showing the separation between wheat cultivars (green) and landraces (red) based on GBS-derived SNP (A) and DArTseq-derived SNP (B) marker data

Fig. 4.

Fig. 4

Dendrogram from cluster analysis of 357 wheat genotypes based on GBS-derived SNP (A) and DArTseq-derived SNP (B) markers

Fig. 5.

Fig. 5

Kinship analysis of 357 wheat genotypes visualized by a heatmap using GBS-derived SNP (A) and DArTseq-derived SNP (B) markers

Functional assessment of genomic regions under selection

In this study, high-Fst markers were used to identify genomic regions in the bread wheat population. The genome-wide distribution of outlier Fst data from both platforms is presented in Supplementary Fig. 1, which visually illustrates the location and density of the selective signatures identified. A total of 886 high-Fst GBS-derived SNP markers were identified, and after performing BLAST and comparing with the IWGSC RefSeq v1.0 reference genome, 176 overlapping genes within 100 kb of flanking regions were detected around these markers (Supplementary Table 4). These genes were involved in various biological processes, including regulation of DNA-templated transcription, cell wall organization, protein phosphorylation, and defense response to others organism. Additionally, the genes were linked to important metabolic processes such as diterpenoid metabolic process, fatty acid biosynthetic process, and glucuronoxylan biosynthetic process. In contrast, the analysis of DArTseq-derived SNP markers revealed that out of 9 identified markers, 3 overlapping genes were detected, which were mainly associated with biological processes such as regulation of DNA-templated transcription and cell wall organization (Supplementary Table 4). These results highlight significant differences between the two marker systems in terms of the distribution of high-Fst genomic regions. Moreover, KEGG analysis revealed that the genes identified with GBS-derived SNP markers are associated with important biological pathways, including plant hormone signal transduction, plant-pathogen interaction, valine, leucine and isoleucine degradation, butanoate metabolism, metabolic pathways, peroxisome, phenylalanine, tyrosine and tryptophan biosynthesis, carbon fixation by the Calvin cycle, glycerolipid metabolism, and ubiquitin mediated proteolysis (Fig. 6a). In contrast, DArTseq-derived SNP marker analysis identified the ubiquitin mediated proteolysis pathway (Fig. 6b). This pathway highlights the role of DArT-related genes in protein degradation processes, which are crucial for regulating protein levels and cellular processes (Supplementary Table 5).

Fig. 6.

Fig. 6

KEGG pathways associated with selection signatures detected by GBS-derived SNP (A) and DArTseq-derived SNP (B) markers

Discussion

Genetic diversity is a cornerstone of plant breeding and is now routinely assessed using molecular markers. In this study, both GBS-derived and DArTseq-derived SNP markers revealed substantial genetic diversity within the evaluated bread wheat germplasm. Our findings align with those of Sansoloni et al. [19], who conducted a comprehensive genetic diversity analysis of a global panel comprising 56,342 hexaploid wheat accessions—8.1% of which were of Iranian origin. That study not only documented genome- and chromosome-level patterns of polymorphism but also highlighted that a considerable fraction of the genetic diversity preserved in landraces remains untapped in modern breeding programs. These observations carry direct implications for wheat improvement: it underscores the presence of specific, underutilized genetic reservoirs including potentially adaptive alleles from regions such as Iran that can be deliberately targeted by breeders. By incorporating such diversity through strategic crosses, breeders can develop new cultivars with improved resilience, yield potential, and adaptation to marginal environments, thereby broadening the increasingly narrow genetic base of elite wheat germplasm. Enhancing the power of these markers provides a more comprehensive characterization of diversity within plant species. One important strategy is the imputation of SNP markers within the GBS framework, as the high frequency of missing data is a key characteristic of GBS [43]. In the present study, the application of imputation led to an almost threefold increase in the number of usable GBS-derived SNP markers. The success of this missing data imputation approach is consistent with findings reported in several previous studies [4447]. Marker imputation has become a standard method in modern genetics for enhancing genome coverage. It enables researchers to genotype large sample sets using low-density, cost-effective arrays and subsequently impute them to a higher density or even to the sequence level, utilizing information from a small, reference panel [48]. Several algorithms and software packages have been developed for genetic marker imputation, among which BEAGLE is recognized as one of the most efficient and suitable tools—and was therefore employed in our study. Further details regarding imputation methods and their comparative performance have been discussed extensively in previous literature [43, 4851].

In essence, imputation serves as a powerful tool for inferring missing genotypes and predicting unknown genetic loci with high accuracy, thereby enabling more comprehensive genetic analyses [26]. Furthermore, genotype imputation allows for the integration of germplasm populations genotyped with different platforms, facilitating downstream analyses and maximizing the utility of accumulated genetic resources [24]. The choice of genotyping platform significantly influences the accuracy of genotype imputation. Specifically, hybridization-based SNP arrays generally yield more reliable imputation results than GBS approaches [52, 53]. There is optimism about integrating data from various sources, including GBS [54]. The large-scale and cost-effective application of GBS requires a combination of sample pooling and genotype imputation [55, 56]. However, selecting an appropriate genotyping method and effectively integrating datasets from diverse sources remain challenging [25]. In this study, we successfully applied this approach to expand the DArTseq-derived SNP marker dataset from 90 to 357 genotypes, imputing data for 267 additional genotypes. In contrast to conventional genotype imputation from low- to high-density marker sets, the pooling strategy introduced by Clouard and Nettelblad [55] reduces the number of samples for microarray testing without decreasing the marker density. In a related report, the Practical Haplotype Graph (PHG) tool was used to impute missing genotypes in wheat inference panels with high accuracy. These panels, which included wheat cultivars and recombinant inbred lines, were genotyped using various sequencing approaches such as exome capture, GBS, and whole-genome skim-seq sequencing [57]. Nyine et al. [58] employed two distinct scenarios for genotype imputation in winter wheat. In the first scenario, missing data from sparse markers were imputed using a combination of the iSelect 90 K SNP array and GBS SNPs. In the second scenario, the shared 90 K and/or GBS markers between a target panel of 307 accessions and a reference panel of 62 wheat lines were used to impute ungenotyped SNP sites. Similar imputation strategies have been applied in other plant species [24, 26, 59, 60]. In maize, a set of 35 million SNPs identified through whole-genome resequencing (WGR) of 1,268 inbred lines was imputed onto a larger panel of over 10,000 lines, which had previously been genotyped using 500,000 GBS-derived SNPs. This process also achieved a high imputation accuracy of 98% [61]. For empirical imputation, Zhao et al. [24] utilized a canola population comprising 160 accessions and lines common to both GBS via transcriptome (GBS-t) and WGS references, along with 24 doubled haploid (DH) lines that overlapped between the WGS and skim-WGS datasets. Imputed genotypes can be useful in identifying new candidate genes controlling quantitative traits [62, 63]. Despite these advantages, several caveats regarding imputation should be considered. Firstly, it is an optional procedure, and its application should be guided solely by the research aims [64]. Secondly, the accuracy of imputation is contingent on multiple factors [9], and it carries the inherent risk of introducing genotyping errors [53]. On the other hand, given the recent development of functional markers like competitive allele-specific PCR (KASP) in wheat, these platforms can be readily deployed for pre-breeding purposes within the characterized germplasm.

The summary of population genetic statistics confirms the robust efficiency of both the GBS-derived SNP and DArTseq-derived SNP marker systems. The difference in Ht and Ho indicated high population differentiation. The observed pattern for the number of markers on different chromosomes was similar to previous studies [47, 65, 66]. This pattern can be attributed to variations in chromosome size and the differential distribution of genetic diversity across the genome. The patterns of genetic diversity and allele frequency in the D genome, compared to the A and B genomes, align with the known population bottleneck induced by polyploidization [57]. In both marker systems, the D genome comprised 13% of the markers. Although the wheat D genome is often underrepresented in genotyping platforms due to its lower polymorphism rate, GBS provided better coverage of it than SNP chip genotyping [67]. Furthermore, GBS-scored SNPs represent a promising marker platform for genetic diversity and genomic selection studies in winter wheat compared to array-scored SNPs [53]. A previous study noted a high frequency of heterozygous genotypes on chromosomes 2 A and 4D, and found no correlation between pre- and post-imputation heterozygosity rates [47]. In a study of a recombinant inbred lines (RIL) population, Bajgain et al. [67] found that while the effect of missing data on the heterozygosity rate was not statistically significant, a general decreasing trend in heterozygosity was observed as the proportion of missing values increased. Discrepancies in diversity statistics between genotyping systems have also been reported in species other than wheat. For instance, in maize, the average expected heterozygosity was lower for Genotyping-by-Sequencing (GBS) than for SNP arrays [46]. In our study, imputation increased the number of available markers in the GBS dataset and the number of individuals in the DArTseq dataset, ultimately enabling a more comprehensive assessment of genetic diversity. The heterozygosity estimates and MANOVA statistics obtained in this study differed from those reported in our earlier assessments of the same germplasm [12, 68]. This discrepancy likely reflects the enhanced data completeness and reduced missing data bias achieved through marker imputation, underscoring its positive impact on the accuracy and reliability of genetic diversity parameters. This process, however, should be applied cautiously. Previous studies have reported that estimates of key diversity parameters (He, Ho, and Fis) can be significantly biased when derived from imputed data compared to non-imputed, complete data. The extent of this bias is dependent on the initial proportion of missing data within a dataset. Specifically, as the rate of missing observations increases, imputation tends to produce upwardly biased estimates of heterozygosity and downwardly biased estimates of the inbreeding coefficient [69]. Consequently, while diversity studies may benefit only marginally from imputing missing data, the power of association mapping is substantially enhanced by this process [44].

The higher percentage of variation explained by the GBS-derived biplots, compared to those from DArTseq, is likely attributable to the difference in the number of markers between the two systems. However, a similar study that also utilized a much higher number of GBS markers than DArT markers reported that biplots from both systems explained only 28% of the variation and yielded similar results [70]. This discrepancy suggests that factors beyond mere marker count, such as marker distribution and the specific genetic diversity of the population under study, may also influence the outcomes. Elbasyoni et al. [53] noted that while GBS-generated SNPs provide a high volume of markers, they often contain substantial missing data. In contrast, array-based SNPs are of higher quality but come at a significantly greater cost per sample. Crucially, despite these differences, both platforms identified similar genetic patterns within their panel, with 90% of the lines clustering into common genetic groups. This finding is consistent with our results, as evidenced by the cluster analysis dendrogram and PCA biplot. Furthermore, in a related study, a comparison between clusters obtained from SilicoDArT and SNP markers revealed a significant positive correlation between the two marker systems [66]. The separation of modern cultivars from landraces was clearly reflected in both marker systems, suggesting that a substantial portion of the allelic diversity present in landraces has remained underutilized in modern breeding programs. This finding aligns with a global study by Sansoloni et al. [19], in which DArT markers and multidimensional scaling revealed that approximately 70% of landrace accessions diverged significantly from the mean position of elite breeding lines.

The estimated genetic relatedness among genotypes within the same group varied slightly between the two marker systems. These findings justify our rationale for implementing imputation across both marker systems for a more accurate assessment of genetic diversity. The close relationships observed between cultivars, particularly with GBS-derived SNP data, are likely due to shared parentage in their pedigrees. Negro et al. [46] reported that despite differences in allele frequency spectra between GBS and DNA array technologies, both showed similar trends in organizing population structure and genetic relatedness. However, it is important to note that array- or chip-based SNP markers are subject to ascertainment bias in downstream applications, which can influence the assessment of genetic relationships among individuals [67]. On the other hand, shared population history and genetic linkage resulting from familial relationships can influence the imputation of missing genotypes [56]. All populations studied were of Iranian origin and, consequently, were genetically closely related. This relatively close relationships meant that a relatively limited number of markers were sufficient for their complete characterization [70]. As a result, the population structures inferred from the two marker systems were largely congruent. Therefore, since most of our markers were imputed, the consistency observed in genetic analyses such as the cluster analysis and kinship matrix between both marker systems can be largely justified. Nevertheless, the high-resolution data generated by GBS can overcome the challenges associated with resolving complex phylogenetic relationships [71]. Accurate estimation of the kinship matrix is ​​crucial for modern plant breeding programs [72]. Such reliable estimation of relatedness between individuals requires a large number of polymorphic markers [73].

In this study, the comparison of two marker systems, GBS-derived SNP and DArTseq-derived SNP, in identifying genomic regions under selection (selection signatures) in bread wheat populations revealed significant differences in the genetic differentiation patterns between populations. The Fst index, as one of the most important indicators for identifying highly differentiated genomic regions between populations, was used, and the results showed that GBS-derived SNP markers, compared to DArTseq-derived SNP, have a broader distribution and higher resolution in detecting genomic regions under selection. While a large number of GBS-derived SNP markers with high Fst were identified, leading to the discovery of 176 overlapping genes in target genomic regions (1 A, 1B, 1D, 2 A, 2B, 3 A, 3B, 4 A, 5B, 5D, 6B, 7 A, 7B, and 7D), the DArT-seq system identified only 9 markers in these regions, with just three overlapping genes on chromosomes 3B and 7 A. These findings not only facilitate the identification of genomic regions associated with key agronomic traits and clarify the recent selection history in modern wheat breeding, but also provide candidate target al.leles for future improvement programs [19]. The differential selection signatures observed between the two marker systems likely arise because SNPs from different platforms cover distinct, yet proximate, genomic regions around genes. This propensity for SNPs from different platforms to capture variation in closely linked but distinct segments has previously been shown to lead to the identification of different quantitative trait loci (QTLs) [46]. This significant difference can be attributed to the technical nature of the two systems; GBS-derived SNPs, being single-nucleotide markers, are widely distributed across the genome and are capable of detecting subtle changes that have been under natural selection [12]. In contrast, DArTseq-derived SNP markers, which are based on the presence or absence of specific DNA fragments, have a more limited genomic coverage and lower sensitivity in detecting subtle genetic changes [74, 75]. Our results showed that the genomic regions under selection, identified by GBS-derived SNPs contain genes involved in regulatory processes related to DNA transcription, cell wall organization, protein phosphorylation, and defense response to biotic stresses. These pathways are particularly significant in the differentiation of populations in response to environmental pressures. In contrast, genes associated with DArTseq-derived SNP markers were mainly involved in more general pathways such as transcription regulation and cell structure processes, which may indicate the lower sensitivity of this system in detecting directional selection. Nevertheless, the identification of distinct selection signatures by DArTseq-derived SNP markers underscores their complementary role in genomic studies.

The genomic regions under selection detected in our study show both congruence and complementarity with those described in earlier wheat selection scans. For instance, Ayalew et al. [76] reported the strongest selection signal on chromosome 2 A in a panel of bread wheat cultivars, which is consistent with our identification of multiple high‑Fst GBS‑derived SNPs on 2 A and the presence of overlapping genes in this region. Similarly, using a 9 K SNP array, identified selected regions predominantly associated with yield potential, vernalization response, plant height, and biotic/abiotic stress tolerance [39]. Zhou et al. [77], analyzing 717 Chinese wheat landraces, uncovered 148 genomic selection sweeps linked to grain yield and disease tolerance, with several sweeps located on chromosomes 1 A, 2 A, 3B, 5B, 6B, and 7D – all of which also harbor selection signatures in our GBS dataset. More recently, Shan et al. [78] dissected genetic networks underlying environmental stress tolerance in wheat and identified selective sweeps on chromosomes 4 A, 5B and 7 A, while Sertse et al. [79] pinpointed adaptive genes (e.g., CIPK2, RBR1, AGL30, SIZ2, CIPK6, CPK4) on chromosome 4 A. Our own GBS‑based analysis detected selection signatures on 4 A and likewise revealed genes involved in stress adaptation, supporting the view that this chromosome is a recurrent target of selection in wheat improvement. Notably, Sansaloni et al. [19] performed a large-scale DArTseq-based diversity analysis in a global wheat panel and reported selective sweeps. Although their study relied primarily on DArTseq markers, our combined use of the GBS and DArTseq platforms allowed for a direct comparison of their resolution.

Furthermore, KEGG analysis revealed clear functional differences between the two marker systems. The pathways identified by GBS-derived SNPs included key processes such as plant hormone signal transduction, plant-pathogen interaction, and amino acid metabolism, all of which are strongly related to environmental adaptation and genetic differentiation. In contrast, the only pathway identified by DArTseq-derived SNP was ubiquitin-mediated proteolysis. Iranian wheat landraces, primarily cultivated under rain-fed conditions, have undergone thousands of years of natural and artificial selection, leading to their adaptation to various environmental stresses [66, 80]. Plant-pathogen interaction is one of the pathways identified in this study, which has also been mentioned in previous research on natural selection in plants. In particular, in the study of Ayalew et al. [76], it was shown that genes involved in the response to pathogens in wheat are under natural selection pressure. In the present study, gene regions under disease-related selection were also identified, indicating the role of natural selection in the genetic differentiation of populations in response to pathogens and increased resistance to diseases. Genes related to metal transporters were identified, especially in the processes of domestication and improvement of stress resistance in wheat [81]. These genes play an important role in regulating the transport of metals inside cells and protecting the plant against nutrient deficiency and environmental stresses. In the present study, the metal transporter gene Nramp4-like (LOC123069893), which helps regulate metal balance in plants, was identified in lysosome-related pathways and may play a role in defense processes and metal regulation in response to environmental stresses and diseases.

Conclusion

Our study sought to provide a platform that would enable the transfer of genomic diversity data across multiple populations. This study establishes a framework for the imputation and integration of GBS-derived SNP and DArTseq-derived SNP markers, effectively leveraging combined datasets for wheat genotyping. Beyond confirming the genetic variation patterns extracted via GBS, the imputation of missing DArTseq data contributed additional value by helping to pinpoint new genomic regions under selection. The presented framework can serve as a paradigm for future research seeking to combine diverse genomic datasets. Furthermore, the imputed and integrated dataset (46,876 GBS-derived SNPs and 3,417 DArTseq-derived SNP markers) for 357 wheat genotypes represents a resource for genome-wide association studies (GWAS) and genomic selection (GS) on this germplasm collection.

Supplementary Information

Supplementary Material 1. (487.8KB, xlsx)

Acknowledgements

Not applicable.

Permission for land study

The authors declare that all land experiments and studies were performed according to authorized rules.

Authors’ contributions

H.Alipour, I.B., and R.D. conceived the research idea and designed the experiment. H.Alipour, H.Abdi and A.T. performed the data analysis. H.Abdi and S.F. wrote the original draft. H.Alipour reviewed and revised the manuscript. All authors approved the final version.

Funding

This research received no external funding.

Data availability

The data are available in the European Variation Archive (EVA) under accession number PRJEB105302. The data is publicly available and can be accessed at: https://www.ebi.ac.uk/eva/?eva-study=PRJEB105302.

Declarations

Ethics approval and consent to participate

The authors declare that all study complies with relevant institutional, national, and international guidelines and legislation for plant ethics in the methods section. The samples were provided by the Seed and Plant Improvement Institute (SPII) and the University of Tehran, Karaj, Iran and all the landraces are available at USDA with USDA PI numbers (Supplementary Table 1). The authors declare that all permissions or licenses were obtained to collect the wheat plant.

Consent for publication

Not applicable. No human participant.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Vikram P, Franco J, Burgueño J, Li H, Sehgal D, Saint-Pierre C, et al. Strategic use of Iranian bread wheat landrace accessions for genetic improvement: Core set formulation and validation. Plant Breed. 2021;140:87–99. [Google Scholar]
  • 2.AmiteyeS. Basic concepts and methodologies of DNA marker systems in plant molecular breeding. Heliyon. 2021;7(10):e08093. 10.1016/j.heliyon.2021.e08093. [DOI] [PMC free article] [PubMed]
  • 3.SongL, Wang R, Yang X, Zhang A, Liu D. Molecular markers and their applications in marker-assisted selection (MAS) in bread wheat (Triticum aestivum L). Agric. 2023;13(3):642. 10.3390/agriculture13030642.
  • 4.KumarR, Das SP, Choudhury BU, Kumar A, Prakash NR, Verma R, et al. Advances in genomic tools for plant breeding: harnessing DNA molecular markers, genomic selection, and genome editing. Biol Res. 2024;57(1):80. 10.1186/s40659-024-00562-6. [DOI] [PMC free article] [PubMed]
  • 5.Agarwal M, Shrivastava N, Padh H. Advances in molecular marker techniques and their applications in plant sciences. Plant Cell Rep. 2008;27:617–31. [DOI] [PubMed] [Google Scholar]
  • 6.Ganal MW, Altmann T, Röder MS. SNP identification in crop plants. Curr Opin Plant Biol. 2009;12:211–7. [DOI] [PubMed] [Google Scholar]
  • 7.GetachewSE, Bille NH, Bell JM, Gebreselassie W. Genotyping by sequencing for plant breeding- a review. Adv Biotechnol Microbiol. 2019;14(4):555891. 10.19080/AIBM.2019.14.555891.
  • 8.Favre F, Jourda C, Besse P, Charron C. Genotyping-by-Sequencing Technology in Plant Taxonomy and Phylogeny. Methods Mol Biol. 2021;2222:167–78. [DOI] [PubMed] [Google Scholar]
  • 9.AlipourH, Bai G, Zhang G, Bihamta MR, Mohammadi V, Peyghambari SA. Imputation accuracy of wheat genotyping-by-sequencing (GBS) data using barley and wheat genome references. PLoS ONE. 2019;14(1):e0208614. 10.1371/journal.pone.0208614. [DOI] [PMC free article] [PubMed]
  • 10.Akbari M, Wenzl P, Caig V, Carling J, Xia L, Yang S, et al. Diversity arrays technology (DArT) for high-throughput profiling of the hexaploid wheat genome. Theor Appl Genet. 2006;113:1409–20. [DOI] [PubMed] [Google Scholar]
  • 11.Marone D, Panio G, Ficco DBM, Russo MA, De Vita P, Papa R, et al. Characterization of wheat DArT markers: Genetic and functional features. Mol Genet Genomics. 2012;287:741–53. [DOI] [PubMed] [Google Scholar]
  • 12.AlipourH, Bihamta MR, Mohammadi V, Peyghambari SA, Bai G, Zhang G. Genotyping-by-sequencing (GBS) revealed molecular genetic diversity of Iranian wheat landraces and cultivars. Front Plant Sci. 2017;8:1293. 10.3389/fpls.2017.01293. [DOI] [PMC free article] [PubMed]
  • 13.TomarV, Dhillon GS, Singh D, Singh RP, Poland J, Joshi AK, et al. Elucidating SNP-based genetic diversity and population structure of advanced breeding lines of bread wheat (Triticum aestivum L). PeerJ. 2021;9:e11593. 10.7717/peerj.11593. [DOI] [PMC free article] [PubMed]
  • 14.HussainS, Habib M, Ahmed Z, Sadia B, Bernardo A, Amand PS, et al. Genotyping-by-Sequencing Based Molecular Genetic Diversity of Pakistani Bread Wheat (Triticum aestivum L.) Accessions. Front Genet. 2022;13:772517. 10.3389/fgene.2022.772517. [DOI] [PMC free article] [PubMed]
  • 15.AlemuA, Feyissa T, Letta T, Abeyo B. Genetic diversity and population structure analysis based on the high density SNP markers in Ethiopian durum wheat (Triticum turgidum ssp. durum). BMC Genet. 2020;21(1):18. 10.1186/s12863-020-0825-x. [DOI] [PMC free article] [PubMed]
  • 16.ZhangLY, Liu DC, Guo XL, Yang WL, Sun JZ, Wang DW, et al. Investigation of genetic diversity and population structure of common wheat cultivars in northern China using DArT markers. BMC Genet. 2011;12(1):42. 10.1186/1471-2156-12-42. [DOI] [PMC free article] [PubMed]
  • 17.El-EsawiMA, Witczak J, Abomohra AEF, Ali HM, Elshikh MS, Ahmad M. Analysis of the genetic diversity and population structure of Austrian and Belgian wheat germplasm within a regional context based on DArT markers. Genes (Basel). 2018;9(1):47. 10.3390/genes9010047. [DOI] [PMC free article] [PubMed]
  • 18.Ebrahimi P, Karami E, Etminan A, Talebi R, Mohammadi R. Genetic Diversity and Genome-Wide Association Study for Some Agronomic Traits in Durum Wheat (Triticum turgidum L.) Using Whole-Genome DArTseq Markers. Plant Mol Biol Rep. 2025;43:1479–95. [Google Scholar]
  • 19.SansaloniC, Franco J, Santos B, Percival-Alwyn L, Singh S, Petroli C, et al. Diversity analysis of 80,000 wheat accessions reveals consequences and opportunities of selection footprints. Nat Commun. 2020;11(1):4572. 10.1038/s41467-020-18404-w. [DOI] [PMC free article] [PubMed]
  • 20.OvendenB, Milgate A, Wade LJ, Rebetzke GJ, Holland JB. Genome-wide associations for water-soluble carbohydrate concentration and relative maturity in wheat using SNP and DArT marker arrays. G3 Genes, Genomes, Genet. 2017;7(8):2821–30. 10.1534/g3.117.039842. [DOI] [PMC free article] [PubMed]
  • 21.TyrkaM, Tyrka D, Wędzony M. Genetic map of triticale integrating microsatellite, DArT and SNP markers. PLoS ONE. 2015;10(12):e0145714. 10.1371/journal.pone.0145714. [DOI] [PMC free article] [PubMed]
  • 22.BalochFS, Alsaleh A, Shahid MQ, Çiftçi V, De Sáenz LE, Aasim M, et al. A whole genome DArTseq and SNP analysis for genetic diversity assessment in durum wheat from central fertile crescent. PLoS ONE. 2017;12(1):e0167821. 10.1371/journal.pone.0167821. [DOI] [PMC free article] [PubMed]
  • 23.Jighly A, Oyiga BC, Makdis F, Nazari K, Youssef O, Tadesse W, et al. Genome-wide DArT and SNP scan for QTL associated with resistance to stripe rust (Puccinia striiformis f. sp. tritici) in elite ICARDA wheat (Triticum aestivum L.) germplasm. Theor Appl Genet. 2015;128:1277–95. [DOI] [PubMed] [Google Scholar]
  • 24.ZhaoH, MacLeod IM, Keeble-Gagnere G, Barbulescu DM, Tibbits JF, Kaur S, et al. Using genotype imputation to integrate Canola populations for genome-wide association and genomic prediction of blackleg resistance. BMC Genomics. 2025;26(1):215. 10.1186/s12864-025-11250-4. [DOI] [PMC free article] [PubMed]
  • 25.Torkamaneh D, Boyle B, Belzile F. Efficient genome-wide genotyping strategies and data integration in crop plants. Theor Appl Genet. 2018;131:499–511. [DOI] [PubMed] [Google Scholar]
  • 26.TorkamanehD, Belzile F. Scanning and filling: Ultra-dense SNP genotyping combining genotyping-by-sequencing, SNP array and whole-genome resequencing data. PLoS ONE. 2015;10(7):e0131533. 10.1371/journal.pone.0131533. [DOI] [PMC free article] [PubMed]
  • 27.Singh S, Sansaloni C, Petroli C, Ellis M, Kilian A. DArTseq-derived SNPs for wheat Iranian landrace accessions. CIMMYT Res Data Softw Repos Netw. 2014;V4. https://hdl.handle.net/11529/10045.
  • 28.PolandJA, Brown PJ, Sorrells ME, Jannink JL. Development of high-density genetic maps for barley and wheat using a novel two-enzyme genotyping-by-sequencing approach. PLoS ONE. 2012;7(2):e32253. 10.1371/journal.pone.0032253. [DOI] [PMC free article] [PubMed]
  • 29.LuF, Lipka AE, Glaubitz J, Elshire R, Cherney JH, Casler MD, et al. Switchgrass genomic diversity, ploidy, and evolution: novel insights from a network-based SNP discovery protocol. PLoS Genet. 2013;9(1):e1003215. 10.1371/journal.pgen.1003215. [DOI] [PMC free article] [PubMed]
  • 30.Bradbury PJ, Zhang Z, Kroon DE, Casstevens TM, Ramdoss Y, Buckler ES. TASSEL: Software for association mapping of complex traits in diverse samples. Bioinformatics. 2007;23:2633–5. 10.1093/bioinformatics/btm308. [DOI] [PubMed] [Google Scholar]
  • 31.SansaloniC, Petroli C, Jaccoud D, Carling J, Detering F, Grattapaglia D, et al. Diversity arrays technology (DArT) and next-generation sequencing combined: genome-wide, high throughput, highly informative genotyping for molecular breeding of Eucalyptus. BMC Proc. 2011;5:P54. 10.1186/1753-6561-5-S7-P54.
  • 32.Crossa J, Jarquín D, Franco J, Pérez-Rodríguez P, Burgueño J, Saint-Pierre C, et al. Genomic prediction of gene bank wheat landraces. G3 Genes, Genomes. Genet. 2016;6:1819–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Browning BL, Browning SR. A unified approach to genotype imputation and haplotype-phase inference for large data sets of trios and unrelated individuals. Am J Hum Genet. 2008;84:210–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.ChapmanJA, Mascher M, Buluç A, Barry K, Georganas E, Session A, et al. A whole-genome shotgun approach for assembling and anchoring the hexaploid bread wheat genome. Genome Biol. 2015;16(1):26. 10.1186/s13059-015-0582-8. [DOI] [PMC free article] [PubMed]
  • 35.Jombart T, Adegenet. A R package for the multivariate analysis of genetic markers. Bioinformatics. 2008;24:1403–5. [DOI] [PubMed] [Google Scholar]
  • 36.Paradis E. pegas: an R package for population genetics with an integrated-modular approach. Bioinformatics. 2010;26:419–20. [DOI] [PubMed] [Google Scholar]
  • 37.Lipka AE, Tian F, Wang Q, Peiffer J, Li M, Bradbury PJ, et al. GAPIT: Genome association and prediction integrated tool. Bioinformatics. 2012;28:2397–9. [DOI] [PubMed] [Google Scholar]
  • 38.Kassambara A, Mundt F. Package ‘factoextra’: Extract and visualize the results of multivariate data analyses. CRAN- R Packag. 2016;:84. https://github.com/kassambara/factoextra/issues%0Ahttps://github.com/kassambara/factoextra/issues%0Ahttps://cran.r-project.org/package=factoextra.
  • 39.Cavanagh CR, Chao S, Wang S, Huang BE, Stephen S, Kiani S, et al. Genome-wide comparative diversity uncovers multiple targets of selection for improvement in hexaploid wheat landraces and cultivars. Proc Natl Acad Sci U S A. 2013;110:8057–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.AppelsR, Eversole K, Feuillet C, Keller B, Rogers J, Stein N, et al. Shifting the limits in wheat research and breeding using a fully annotated reference genome. Science. 2018;361(6403):eaar7191. 10.1126/science.aar7191. [DOI] [PubMed]
  • 41.Ogata H, Goto S, Sato K, Fujibuchi W, Bono H, Kanehisa M. KEGG: Kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 1999;27:29–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kanehisa M, Furumichi M, Sato Y, Kawashima M, Ishiguro-Watanabe M. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 2023;51:D587–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.de Oliveira AA, Guimarães LJM, Guimarães CT, Guimarães PE, de O, Pinto M, de Pastina O. Single nucleotide polymorphism calling and imputation strategies for cost-effective genotyping in a tropical maize breeding program. Crop Sci. 2020;60:3066–82. [Google Scholar]
  • 44.HeS, Zhao Y, Mette MF, Bothe R, Ebmeyer E, Sharbel TF, et al. Prospects and limits of marker imputation in quantitative genetic studies in European elite wheat (Triticum aestivum L). BMC Genomics. 2015;16(1):168. 10.1186/s12864-015-1366-y. [DOI] [PMC free article] [PubMed]
  • 45.Schmidt M, Kollers S, Maasberg-Prelle A, Großer J, Schinkel B, Tomerius A, et al. Prediction of malting quality traits in barley based on genome-wide marker data to assess the potential of genomic selection. Theor Appl Genet. 2016;129:203–13. [DOI] [PubMed] [Google Scholar]
  • 46.NegroSS, Millet EJ, Madur D, Bauland C, Combes V, Welcker C, et al. Genotyping-by-sequencing and SNP-arrays are complementary for detecting quantitative trait loci by tagging different haplotypes in association studies. BMC Plant Biol. 2019;19(1):318. 10.1186/s12870-019-1926-4. [DOI] [PMC free article] [PubMed]
  • 47.Edae EA, Bowden RL, Poland J. Application of population sequencing (POPSEQ) for ordering and imputing genotyping-by-sequencing markers in hexaploid wheat. G3 Genes, Genomes. Genet. 2015;5:2547–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Phocas F. Genotyping, the Usefulness of Imputation to Increase SNP Density, and Imputation Methods and Tools. Methods Mol Biol. 2022;2467:113–38. [DOI] [PubMed] [Google Scholar]
  • 49.PeiYF, Li J, Zhang L, Papasian CJ, Deng HW. Analyses and comparison of accuracy of different genotype imputation methods. PLoS ONE. 2008;3(10):e3551. 10.1371/journal.pone.0003551. [DOI] [PMC free article] [PubMed]
  • 50.Shi F, Tibbits J, Pasam RK, Kay P, Wong D, Petkowski J, et al. Exome sequence genotype imputation in globally diverse hexaploid wheat accessions. Theor Appl Genet. 2017;130:1393–404. [DOI] [PubMed] [Google Scholar]
  • 51.DeMarino A, Amr Mahmoud A, Bose M, Ozan Bircan K, Terpolovsky A, Bamunusinghe V, et al. A comparative analysis of current phasing and imputation software. PLoS ONE. 2022;17(10):e0260177. 10.1371/journal.pone.0260177. [DOI] [PMC free article] [PubMed]
  • 52.Rasheed A, Hao Y, Xia X, Khan A, Xu Y, Varshney RK, et al. Crop Breeding Chips and Genotyping Platforms: Progress, Challenges, and Perspectives. Mol Plant. 2017;10:1047–64. [DOI] [PubMed] [Google Scholar]
  • 53.Elbasyoni IS, Lorenz AJ, Guttieri M, Frels K, Baenziger PS, Poland J, et al. A comparison between genotyping-by-sequencing and array-based scoring of SNPs for genomic prediction accuracy in winter wheat. Plant Sci. 2018;270:123–30. [DOI] [PubMed] [Google Scholar]
  • 54.Lell M, Gogna A, Kloesgen V, Avenhaus U, Dörnte J, Eckhoff WM, et al. Breaking down data silos across companies to train genome-wide predictions: A feasibility study in wheat. Plant Biotechnol J. 2025;23:2704–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.ClouardC, Nettelblad C. Genotyping of SNPs in bread wheat at reduced cost from pooled experiments and imputation. Theor Appl Genet. 2024;137(1):26. 10.1007/s00122-023-04533-5. [DOI] [PMC free article] [PubMed]
  • 56.TechnowF, Gerke J. Parent-progeny imputation from pooled samples for cost-efficient genotyping in plant breeding. PLoS ONE. 2017;12(12):e0190271. 10.1371/journal.pone.0190271. [DOI] [PMC free article] [PubMed]
  • 57.JordanKW, Bradbury PJ, Miller ZR, Nyine M, He F, Fraser M, et al. Development of the Wheat Practical Haplotype Graph database as a resource for genotyping data storage and genotype imputation. G3 Genes, Genomes, Genet. 2022;12(2):jkab390. 10.1093/g3journal/jkab390. [DOI] [PMC free article] [PubMed]
  • 58.Nyine M, Wang S, Kiani K, Jordan K, Liu S, Byrne P, et al. Genotype imputation in winter wheat using first-generation haplotype map SNPs improves genome-wide association mapping and genomic prediction of traits. G3 Genes, Genomes. Genet. 2019;9:125–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Chen L, Yang S, Araya S, Quigley C, Taliercio E, Mian R, et al. Genotype imputation for soybean nested association mapping population to improve precision of QTL detection. Theor Appl Genet. 2022;135:1797–810. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.ChanAW, Hamblin MT, Jannink JL. Evaluating imputation algorithms for low-depth genotyping-by-sequencing (GBS) data. PLoS ONE. 2016;11(8):e0160733. 10.1371/journal.pone.0160733. [DOI] [PMC free article] [PubMed]
  • 61.SwartsK, Bauer E, Glaubitz J, Ho T, Johnson L, Li Y, et al. A large scale joint analysis of flowering time reveals independent temperate adaptations in maize. bioRxiv. 2016:086082. 10.1101/086082. [DOI] [PMC free article] [PubMed]
  • 62.SakhaleSA, Yadav S, Clark LV, Lipka AE, Kumar A, Sacks EJ. Genome-wide association analysis for emergence of deeply sown rice (Oryza sativa) reveals novel aus-specific phytohormone candidate genes for adaptation to dry-direct seeding in the field. Front Plant Sci. 2023;14:1172816. 10.3389/fpls.2023.1172816. [DOI] [PMC free article] [PubMed]
  • 63.RahimiY, Bihamta MR, Taleei A, Alipour H, Ingvarsson PK. Genome-wide association study of agronomic traits in bread wheat reveals novel putative alleles for future breeding programs. BMC Plant Biol. 2019;19(1):541. 10.1186/s12870-019-2165-4. [DOI] [PMC free article] [PubMed]
  • 64.RajendranNR, Qureshi N, Pourkheirandish M. Genotyping by sequencing advancements in Barley. Front Plant Sci. 2022;13:931423. 10.3389/fpls.2022.931423. [DOI] [PMC free article] [PubMed]
  • 65.TehseenMM, Tonk FA, Tosun M, Istipliler D, Amri A, Sansaloni CP, et al. Exploring the genetic diversity and population structure of wheat landrace population conserved at ICARDA Genebank. Front Genet. 2022;13:900572. 10.3389/fgene.2022.900572. [DOI] [PMC free article] [PubMed]
  • 66.MahboubiM, Mehrabi R, Naji AM, Talebi R. Whole-genome diversity, population structure and linkage disequilibrium analysis of globally diverse wheat genotypes using genotyping-by-sequencing DArTseq platform. 3 Biotech. 2020;10(2):48. 10.1007/s13205-019-2014-z. [DOI] [PMC free article] [PubMed]
  • 67.Bajgain P, Rouse MN, Anderson JA. Comparing genotyping-by-sequencing and single nucleotide polymorphism chip genotyping for quantitative trait loci mapping in wheat. Crop Sci. 2016;56:232–48. [Google Scholar]
  • 68.AlipourH, Abdi H, Rahimi Y, Bihamta MR. Dissection of the genetic basis of genotype-by-environment interactions for grain yield and main agronomic traits in Iranian bread wheat landraces and cultivars. Sci Rep. 2021;11(1):17742. 10.1038/s41598-021-96576-1. [DOI] [PMC free article] [PubMed]
  • 69.Fu YB. Genetic diversity analysis of highly incomplete snp genotype data with imputations: An empirical assessment. G3 Genes. Genomes Genet. 2014;4:891–900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.PolandJ, Endelman J, Dawson J, Rutkoski J, Wu S, Manes Y, et al. Genomic Selection in Wheat Breeding using Genotyping-by‐Sequencing. Plant Genome. 2012;5(3):103–13. 10.3835/plantgenome2012.06.0006.
  • 71.Escudero M, Eaton DAR, Hahn M, Hipp AL. Genotyping-by-sequencing as a tool to infer phylogeny and ancestral hybridization: A case study in Carex (Cyperaceae). Mol Phylogenet Evol. 2014;79:359–67. [DOI] [PubMed] [Google Scholar]
  • 72.SemagnK, Iqbal M, Alachiotis N, N’Diaye A, Pozniak C, Spaner D. Genetic diversity and selective sweeps in historical and modern Canadian spring wheat cultivars using the 90K SNP array. Sci Rep. 2021;11(1):23773. 10.1038/s41598-021-02666-5. [DOI] [PMC free article] [PubMed]
  • 73.Baumung R, Sölkner J. Pedigree and marker information requirements to monitor genetic variability. Genet Sel Evol. 2003;35:369–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.White J, Law JR, MacKay I, Chalmers KJ, Smith JSC, Kilian A, et al. The genetic diversity of UK, US and Australian cultivars of Triticum aestivum measured by DArT markers and considered by genome. Theor Appl Genet. 2008;116:439–53. [DOI] [PubMed] [Google Scholar]
  • 75.Bidyananda N, Jamir I, Nowakowska K, Varte V, Vendrame WA, Devi RS, et al. Plant Genetic Diversity Studies: Insights from DNA Marker Analyses. Int J Plant Biol. 2024;15:607–40. [Google Scholar]
  • 76.AyalewH, Sorrells ME, Carver BF, Baenziger PS, Ma XF. Selection signatures across seven decades of hard winter wheat breeding in the Great Plains of the United States. Plant Genome. 2020;13(3):e20032. 10.1002/tpg2.20032. [DOI] [PMC free article] [PubMed]
  • 77.Zhou Y, Chen Z, Cheng M, Chen J, Zhu T, Wang R, et al. Uncovering the dispersion history, adaptive evolution and selection of wheat in China. Plant Biotechnol J. 2018;16:280–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Shan D, Ali M, Shahid M, Arif A, Waheed MQ, Xia X, et al. Genetic networks underlying salinity tolerance in wheat uncovered with genome-wide analyses and selective sweeps. Theor Appl Genet. 2022;135:2925–41. [DOI] [PubMed] [Google Scholar]
  • 79.SertseD, You FM, Klymiuk V, Haile JK, N’Diaye A, Pozniak CJ, et al. Historical selection, adaptation signatures, and ambiguity of introgressions in wheat. Int J Mol Sci. 2023;24(9):8390. 10.3390/ijms24098390. [DOI] [PMC free article] [PubMed]
  • 80.Fayaz F, Aghaee Sarbarzeh M, Talebi R, Azadi A. Genetic Diversity and Molecular Characterization of Iranian Durum Wheat Landraces (Triticum turgidum durum (Desf.) Husn.) Using DArT Markers. Biochem Genet. 2019;57:98–116. [DOI] [PubMed] [Google Scholar]
  • 81.Maccaferri M, Harris NS, Twardziok SO, Pasam RK, Gundlach H, Spannagl M, et al. Durum wheat genome highlights past domestication signatures and future improvement targets. Nat Genet. 2019;51:885–95. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (487.8KB, xlsx)

Data Availability Statement

The data are available in the European Variation Archive (EVA) under accession number PRJEB105302. The data is publicly available and can be accessed at: https://www.ebi.ac.uk/eva/?eva-study=PRJEB105302.


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES