Abstract
Blanchability, the ease of seed coat removal after roasting, is a critical post-harvest trait in groundnut (Arachis hypogaea L.) that directly influences processing efficiency and product quality. Despite its economic value, limited genetic understanding restricts breeding efforts for customized blanchability in groundnut. Here, we integrate whole-genome resequencing of 184 diverse groundnut genotypes with multi-season phenotyping to dissect the haplotype-level genomic architecture of blanchability. Genome-wide association studies identify 26 significant single-nucleotide polymorphism-trait associations across multiple chromosomes, six of which are further validated using KASP markers, with two successfully validating the expected allelic effects across breeding lines and genotypes. Haplo-pheno analyses identify distinct subspecies-specific signatures for the major associations on chromosomes Ah01, Ah05, Ah06, and Ah17. Superior high-blanchability haplotypes (Ah01HapBL4, Ah05HapBL3, Ah06HapBL5, Ah06HapBL10, and Ah17HapBL6) are predominantly found in the fastigiata subspecies from South Asia and South America. In contrast, the low-blanchability haplotypes (Ah01HapBL2, Ah05HapBL6, Ah06HapBL3, Ah17HapBL2) are enriched in the hypogaea subspecies, mainly from Africa. These contrasting haplotypes offer the flexibility to achieve either high or low blanchability tailored to specific end-use applications. The availability of diagnostic markers and donor genotypes harboring multiple favorable haplotypes provides immediate tools for haplotype-based breeding. Collectively, this study introduces blanchability as a novel, customizable breeding target and establishes a translational framework to enhance the processing quality and industrial value of groundnut through haplotype-based breeding.
Subject terms: Plant genetics, Natural variation in plants
Genomic investigation of the ICRISAT mini-core collection revealed subspecies-specific favorable haplotypes and diagnostic markers associated with blanchability, the ease of seed coat removal, a key processing trait in groundnut.
Introduction
Groundnut (Arachis hypogaea L.; 2n = 4x = 40) is a globally important oilseed and economic crop of multi-purpose utility1. Groundnut genotypes can be broadly categorized into two subspecies: hypogaea and fastigiata, and six distinct botanical varieties: hypogaea, hirsuta, fastigiata, vulgaris, aequatoriana, and peruviana2,3. Additionally, groundnut accessions exhibit a remarkable array of morphological diversity, grouped into four agronomic types: Virginia runner, Virginia bunch, Valencia bunch, and Spanish bunch. Each agronomic type exhibits distinct growth habits, plant architecture, pod morphology, seed size, maturity duration, and yield performance. The Virginia runner and Virginia bunch types, belonging to subspecies hypogaea, are characterized by an absence of flowers on the main stem, an alternating pattern of vegetative and reproductive nodes on lateral branches, and a predominantly spreading or semi-erect growth habit. Virginia bunch types generally produce larger seeds and mature earlier than runners. In contrast, the Spanish and Valencia bunch types, under subspecies fastigiata, exhibit erect or compact growth habits, the presence of flowers on the main stem, sequential reproductive nodes, and smaller seed sizes. While the Spanish types are known for early maturity, the Valencia types typically display three-seeded pods and are suitable for direct consumption in the form of nuts4.
Groundnut products, such as salted, raw, or roasted nuts, oil, butter, candies, and groundnut flour, hold significant industrial importance. Here, ‘blanching’, the removal of the seed coat, is amongst the most critical steps to enhance the quality and shelf life of groundnut-based products5. As a result, ‘blanchability’, which refers to the ease with which the seed coat is removed, has become a key trait for groundnut industries in recent years. Blanching is a multistep process involving drying, roasting, mechanical rubbing between hard and soft surfaces, and air blowing to remove the seed coat. Several devices, such as laboratory blanchers (American Society of Agricultural and Biological Engineers, 2006), have been developed to facilitate this process5–7. Improper removal of seed coats and germ remnants imparts a bitter taste to groundnut products, for example, in groundnut butter, thus adversely impacting their industrial and economic value5. The genotypes and breeding lines possessing >60% blanchability are desirable for the commercial applications such as groundnut butter, snack food, snack bars, groundnut flour, and are considered processing-grade; however, commercial buyers often use 50% as a minimum threshold for acceptance8. The blanchability percentage requirement varies according to the manufacturers and product demand; for instance, genotypes with high split seeds are more suitable for candies, groundnut butter production, while low-blanchability is preferred for beer nuts and seed products with various confectionery coatings, where testa retention is desirable for added texture or adherence of coatings5. Hence, both high and low-blanchability are valued in different industrial contexts, underscoring the importance of aligning breeding objectives with specific market requirements. This dual demand highlights the need for a nuanced approach in trait selection and cultivar development that considers the intended processing application.
Blanchability is influenced by a complex interplay of genetic, physiological, and environmental factors5. The groundnut seed coat, derived from the ovular integuments, is associated with multiple traits, including days to maturity, seed size, weight, and blanchability. Its properties are determined by both physical (thickness, density, fissures, cavitation) and chemical (lignin, polyphenols, alkaloids, anthocyanins) attributes9,10. Blanchability, in particular, is influenced by various factors such as genotype, kernel grade, post-harvest storage temperature, moisture content, storage conditions, and the thermal and hygroscopic properties of the cotyledon-skin interface5,11,12. These factors collectively determine the ease with which the seed coat detaches from the kernel during processing, thereby affecting both industrial efficiency and product quality. Previous studies have highlighted substantial variability in blanchability across genotypes and environments, underscoring the need for precise and reproducible phenotyping strategies. While earlier work often relied on large sample quantities for trait assessment, recent evidence indicates that smaller sample sizes (50 g) can also provide reliable estimates of blanching percentage under controlled conditions13,14.
Groundnut whole genome resequencing (WGRS) has revolutionized genetic research and breeding, providing valuable insights into genetic variation, trait mapping, and marker development. Several genome sequencing projects led to genetic dissection of complex traits using whole genome resequencing of germplasm, including plant architecture and oil biosynthesis15, seed weight16,17, lowering pattern, inner tegument color, growth habit, pod/seed weight18. Quantitative trait loci (QTL) mapping and genome-wide association studies (GWAS) have played a central role in dissecting the genetic basis of complex traits across diverse crop species, enabling breeders to identify genomic regions underlying phenotypic variation and accelerate marker-assisted selection (MAS)1.
In the past, bi-parental genetic mapping using the QTL-Seq approach (Middleton × Sutherland cultivar) reported genomic regions for blanchability on chromosome A06 and B01 with R2 values of 10.3% and 5.6%, respectively19. Furthermore, studies using the US mini-core germplasm have investigated phenotypic variation and genotype × environment interactions to better understand and enhance the genetic improvement of blanchability14. Similarly, association mapping in the ICRISAT mini-core collection using SNP (Single Nucleotide Polymorphism) array has revealed genomic regions linked to blanchability; however, limited marker density and resolution have restricted the precision of these discoveries13. More recently, haplotype-based breeding has emerged as a powerful approach that leverages combinations of neighboring SNPs in linkage disequilibrium (LD), providing greater resolution and predictive power than single-marker analyses. Haplotype-informed selection has been successfully implemented in several crops, including rice20,21, and wheat22 for yield and quality traits, while in pigeonpea there has been identification of superior genomic combinations associated with flowering time and drought tolerance 23,24. However, despite its critical industrial importance, a comprehensive blanchability-related genetic framework and haplotype signatures underscoring its selection pattern in diverse groundnut agronomic types remain largely elusive, thus hampering informed genetic improvement in breeding programmes. While haplotype-based breeding has begun to impact trait improvement in crops like rice20,21 and wheat22, and superior haplotypes has been explored well in pigeonpea23,24, groundnut still remains unexplored in this regard.
Here, we present the haplotype-based breeding framework utilizing the ICRISAT groundnut mini-core collection25, comprising 184 groundnut genotypes (57 Spanish bunch, 35 Valencia bunch, 48 Virginia bunch, 33 Virginia runner, 11 Unknown). Although we developed kompetitive allele-specific PCR (KASP) assay, the most popular and adapted platform26. Nevertheless, our study suggested higher confidence in haplotype-based selection. In groundnut, the integration of QTL/GWAS discovery, haplotype-based dissection, and KASP marker validation offers a robust framework for accelerating genetic improvement of key quality traits such as blanchability. These genomics-assisted tools together enable the translation of molecular discoveries into practical breeding applications, improving the efficiency and accuracy of cultivar development27.
Specifically, we aimed to (i) phenotype this panel for blanchability across multiple crop seasons (ii) identify genomic regions associated with blanchability through multi-locus GWAS approaches; BLINK (Bayesian-information and Linkage-disequilibrium Iteratively Nested Keyway) and FarmCPU (Fixed and random model Circulating Probability Unification) models, (iii) develop diagnostic KASP markers for efficient selection, and (iv) uncover superior haplotype variants linked to blanchability and evaluate their selection pattern across subspecies and geographical origins. To our knowledge, this study represents developing framework for the first haplotype-based breeding effort in groundnut, marking a significant step toward precision breeding in this major crop.
Importantly, we focus on blanchability, a novel, high-value trait with direct implications for the processing industry, which has been historically overlooked in genetic improvement efforts. Collectively, these findings establish a translational framework for haplotype-informed varietal development, incorporating blanchability as a customizable selection target in groundnut improvement programs and enabling the design of next-generation cultivars tailored to industrial demands.
Results
Genomic variant profiling, principal component analysis, and LD decay in the mini-core collection
A total of 561,099 SNPs were identified using the WGRS approach. Minor allele frequency (MAF) was applied at a 0.05 level. A working subset of 255,144 filtered SNPs with 0.2% heterozygosity and 0.05% MAF was used for GWAS analysis with Bonferroni correction of 5%, across 167 accessions were subsequently employed for association mapping. SNP density across chromosomes, segmented into 1 Mb windows, shows the overall SNP counts (Supplementary Fig. S1).
Heatmap and dendrogram of the kinship matrix, constructed using 255,144 polymorphic SNPs derived from WGRS data, revealed clear and well-defined clustering patterns among the genotypes within the mini-core collection (Supplementary Fig. S2) (Supplementary Table S1). From WGRS data, the population structure based on groundnut subspecies, botanical variety, and agronomic type also revealed two distinct clusters: cluster1 (88 genotypes): fastigiata, Valencia bunch/ hypogaea, Virginia runner/ vulgaris, Spanish bunch, while for cluster2 (72 genotypes): hypogaea, Virginia bunch or Virginia runner/ hirsuta/ aequatoriana/ peruviana/ vulgaris, Spanish bunch/ fastigiata, Valencia bunch (Supplementary Fig. S2) (Supplementary Table S1). The graphical representation of the LD characteristics of the mini-core collection is presented in Supplementary Fig. S3.
Previous WGRS studies in cultivated groundnut collections have reported marked variation in linkage disequilibrium (LD) decay across genomes and botanical varieties15,18. The mean LD decay of the whole genome reported in the previous study was estimated at 92.3 kb (r² = 0.16)15. Additionally, half-maximum LD decay distances reported for different botanical varieties were 99.4 kb (var. hypogaea), 174.5 kb (var. hirsuta), 5.6 kb (var. fastigiata), and 15.8 kb (var. vulgaris)18. In the present study using the ICRISAT mini-core collection, LD decay reached an r² threshold of 0.2 at approximately 50 kb, defining the genome-wide critical distance for linkage detection. Accordingly, genes located within this interval of blanchability-associated STAs were considered putative candidate genes.
Diverse blanchability profiles detected in the ICRISAT groundnut mini-core collection
We have generated and used data from two seasons in two replications with different sample sizes: 50 g rainy season (50_R), 50 g post-rainy season (50_PR), 50g_rainy_post-rainy (50R_PR), 200g_post-rainy (200_PR), as phenotypes for GWAS analysis. The highest blanchability observed was upto 70%, while 3.98% was the lowest. Blanchability is significantly influenced by environment, making G×E stability assessment essential for reliable donor identification13. In this context, genotypes were ranked based on Additive Main effects and Multiplicative Interaction (AMMI) Stability Values (ASV), and the most stable and unstable subsets (top and bottom deciles) were identified. Stable genotypes provide consistent performance across conditions and are therefore ideal donors for the groundnut breeding program (Supplementary Table S2, Supplementary Fig. S4). The blanchability percentage observed was highest (50-70%) in the Spanish bunch, followed by the Valencia bunch, while the agronomic type Virginia bunch and Virginia runner account for the lowest blanchability (3.98–20%) (Supplementary Table S3).
Genetic architecture of blanchability revealed by GWAS
GWAS was performed using 255,144 high-quality SNPs with less than 20% missing data. These SNPs were distributed across the twenty chromosomes of groundnut. SNP genotyping data of 255,144 SNPs, along with information on population structure and kinship matrix, were used for GWAS for blanchability with the phenotypic data obtained from the 2022 rainy season and 2022-2023 post-rainy season. After the calculation, the Bonferroni correction threshold level of -log (10) ‘p’ value was set to 6.0 for analysis with WGRS data. Deviations in p-values at the initial stage indicated the presence of population stratification. The FarmCPU and BLINK models were found to significantly affect the detection of associations contributing to the traits of interest (Fig. 1; Supplementary Fig. S5).
Fig. 1. Genome-wide identification of STAs controlling blanchability in groundnut.
A–D are from the FarmCPU model, while E–H are from the BLINK model: A, E: Blanchability 50R (50 g sample size in rainy season), B, F: Blanchability 50PR (50 g sample size in post-rainy season), C, G: Blanchability 200PR (200 g sample size in post-rainy season), and D, H: Blanchability 50PR_R (50 g sample size in post-rainy and rainy season). In these Manhattan plots, the highlighted (pink) region shows the gene associated with the validated polymorphic markers on chromosome Ah01 and Ah05, respectively. The other dots represent the significant STAs and the genes found to be associated with blanchability. Source data for the GWAS Manhattan plots (Supplementary Data 1).
A total of 26 SNP-trait associations (STAs) were identified with the two different GWAS models. Among these, eight STAs were associated with 200PR, with a range of ‘p’ value; 2.53 × 10−16 to 1.78 × 10−7 and phenotypic variance explained (PVE); 0.86 to 54.03%. While for data set with 50PR_R mean (nine STAs), 50PR (one STA), 50R (eight STAs), the ‘p’ values ranged from 1.31 × 10−12 to 5.77 × 10−8 with PVE range from 0 to 38.51%, the ‘p’ values ranged from 1.34 × 10−8 with PVE value 43.12%, the ‘p’ values range from 3.91 × 10−13 to 1.67 × 10−7 with PVE range from 0 to 38.75%, respectively. Among the STAs identified, the STA (S17_133752226) on chromosome Ah17 exhibited the maximum PVE value of 54.03% (Table 1).
Table 1.
Significant SNP trait associations identified for groundnut blanchability
| Blanchability | SNP ID | Chr | Position | p value | Model | PVE (%) |
|---|---|---|---|---|---|---|
| 50 R | S1_62503335 | 1 | 62503335 | 3.39E−09 | FarmCPU | 2.7 |
| S7_66675581 | 7 | 66675581 | 3.91E−13 | FarmCPU | 26.0 | |
| S7_71977840 | 7 | 71977840 | 3.07E−08 | FarmCPU | 17.0 | |
| S13_19831068 | 13 | 19831068 | 4.14E−09 | FarmCPU | 1.3 | |
| S13_105766607 | 13 | 105766607 | 1.67E−07 | FarmCPU | 3.4 | |
| S8_20932855 | 8 | 20932855 | 4.19E−08 | BLINK | 39.0 | |
| S15_112025568 | 15 | 112025568 | 1.55E−09 | BLINK | 21.0 | |
| 200PR | S11_124801197 | 11 | 124801197 | 2.20E−08 | FarmCPU | 8.2 |
| S18_8958415 | 18 | 8958415 | 1.02E−08 | FarmCPU | 12.0 | |
| S18_17173755 | 18 | 17173755 | 6.58E−10 | FarmCPU | 50.0 | |
| S6_70882693 | 6 | 70882693 | 1.59E−08 | BLINK | 15.0 | |
| 200PR | S17_133752226 | 17 | 133752226 | 1.23E−07 | BLINK | 54.0 |
| 50PR | S17_133752226 | 17 | 133752226 | 5.07E−14 | BLINK | 43.1 |
| 50PR_R | S1_49247173 | 1 | 49247173 | 2.39E−08 | FarmCPU | 23.0 |
| S3_29782297 | 3 | 29782297 | 7.61E−10 | FarmCPU | 3.4 | |
| S11_45896378 | 11 | 45896378 | 1.31E−12 | FarmCPU | 1.8 | |
| S14_130371390 | 14 | 130371390 | 2.76E−11 | FarmCPU | 2.0 | |
| S5_107081867 | 5 | 107081867 | 7.84E−12 | BLINK | 38.5 |
Significant STAs identified through GWAS analysis: 50R (Blanchability for 50 g rainy sample), 50PR (Blanchability for 50 g post-rainy sample), 200PR (Blanchability for 200 g post-rainy sample), 50PR_R (Blanchability for the mean of 50 g sample, rainy and post-rainy).
Across the four datasets, six polymorphic markers were identified for blanchability and were distributed across chromosomes Ah01 (1), Ah05 (1), Ah06 (1), Ah07 (1), Ah16 (1), and Ah17 (1) (Supplementary Fig. S6). This includes S1_62503335 (Ah01; 50_R), S5_107081867 (Ah05; 50_PR_R), S6_70882693 (Ah06; 200_PR), S7_66675581 (Ah07; 50_R), S16_34347730 (Ah16; 50_PR_R) and S17_133752226 (Ah17;200PR and 50PR). Furthermore, we determined the allelic effects and allelic distribution of these stable STAs (Supplementary Fig. S7).
The SNPs obtained from the analysis were compared against the groundnut reference genome assembly “Tifrunner v.2.0” using physical positions. The function of the candidate genes associated with lead SNPs were investigated using Peanut base (https://Peanutbase.org/) (Supplementary Table S4). This includes galactoside 2-alpha-L-fucosyltransferase-like protein (Arahy.1B9W47), glycerophosphoryl diester phosphodiesterase3-like (Arahy.TWT7AM), Protein kinase superfamily protein (Arahy.Q1WR01). The gene Arahy.1B9W47 was identified from S11_124801197 (Ah11), explaining 8.2% PVE, while Arahy.TWT7AM was associated with marker S3_29782297 (Ah03), accounting for 3.4% PVE. The Arahy.Q1WR01 gene, linked to S5_107081867 (Ah05), explained 38.51% PVE and possesses a validated SNP (snpAH00569), highlighting its potential significance in controlling blanchability traits.
Haplo–pheno analysis identifies favorable blanchability haplotypes
The availability of the WGRS dataset permitted a comprehensive genomic investigation at haplotype resolution. Notably, to determine the impact of haplotype diversity on phenotypic performance, haplo-pheno analysis was conducted, where differences in the phenotypic means were correlated with the distinct haplotype signatures of all 26 blanchability-related STAs. Significant haplotype diversity Ah01HapBL, Ah05HapBL, Ah06HapBL, and Ah17HapBL across 4 STAs; S1_62503335 (Ah01) (50_R), S5_107081867 (Ah05) (50_PR_R), S6_70882693 (Ah06) (200_PR), and S17_133752226 (Ah17) (200PR, 50PR); were observed (Fig. 2, Supplementary Table S5–S17). The objective was to identify superior haplotypes associated with high and low blanchability by comparing mean phenotypic values across haplotype groups. For candidate loci harboring more than two haplotypes, statistical significance was evaluated using one-way ANOVA followed by Tukey’s HSD post-hoc test, enabling the robust identification of superior haplotypes (Fig. 3). Several neighboring SNPs (to the lead SNP) that were individually non-significant (based on p-value) became highly informative when evaluated within LD-defined haplotype blocks (Supplementary Table S17 and Supplementary Fig. S8). The phenotypic range and mean of biallelic SNP vs haplotype block were comprehensively represented (Supplementary Fig. S8), where the haplotype block represents more precise blanchability percentage. These haplotypes revealed strong and biologically coherent phenotypic differentiation, with narrower and more precise phenotypic ranges that were not apparent from single-SNP analyses alone (Supplementary Fig. S8).
Fig. 2. Superior haplotypes to tailor blanchability in groundnut.
A Significant STAs were identified based on GWAS for blanchability. B, E, H, K- LD-based haplotype (Ah01HapBL), (Ah05HapBL), (Ah06HapBL), and (Ah17HapBL) clustering of the genomic region (Ah01_62503335), (Ah05_107081867), (Ah06_70882693), and (Ah17_133752226); respectively, associated with blanchability. There were 24 haplotype groups in total, of which 6 are depicted as they have at least 3 genotypes in their haplotype group. FlapJack software was used to visualize the haplotype groups. Source data for the genotype in the haplotype group (Supplementary Data 2). C, F, I, L- Haplo-pheno analysis revealed that Ah01HapBL4, Ah05HapBL3, Ah06HapBL5, Ah06HapBL10, and Ah17HapBL6 confer high blanchability (> 50%), while Ah01HapBL2, Ah05HapBL5, Ah05HapBL6, Ah06HapBL3, Ah06HapBL7, and Ah17HapBL2 for low-blanchability (<20%). Violin plots represent the kernel density distribution of the data, with the width indicating the relative frequency of observations. Source data for the Violin plots (Supplementary Data 3–6). D, G. The superior high and low blanchability haplotypes identified here can be used to customize blanchability in groundnut. J High blanchability trait predominantly belongs to the Spanish bunch and Valencia bunch agronomic type, while the Virginia runner and bunch have low blanchability. Boxplots show the median (center line), interquartile range (box), and whiskers extending to 1.5× the IQR, with values beyond the whiskers shown as outliers. Source data for the Boxplots (Supplementary Data 7). Note: Statistical tests and effect sizes specified. We have performed multiple range comparisons of the phenotypic means among all the haplotype groups, which comprised at least three genotypes. One-way ANOVA, followed by Tukey’s HSD post-hoc test, was used to determine statistical significance. The haplo-group(s) with the most optimal/preferred phenotypic mean values were designated as superior haplotype(s), n = number of genotypes.
Fig. 3. Development and validation of KASP markers from potential candidate genes for blanchability.
A Validation panel includes 20 high-blanchability genotypes and groundnut varieties (50–70%-blanchability), and five low-blanchability genotypes (1-10%-blanchability), NA: Not available. Two single-nucleotide polymorphisms (SNPs) show clear homozygous clusters for (B) snpAH00567 (Ah01_62503335) for PGR5-like protein (Arahy.N69HTX) gene, (C) snpAH00569 (Ah05_107081867) for Protein kinase superfamily (Arahy.Q1WR01) gene.
Four superior high-blanchability haplotypes (Ah01HapBL4 (53.67%), Ah05HapBL3 (51%), Ah06HapBL5 (63%), Ah06HapBL10 (69%), Ah17HapBL6 (61%) were defined with blanchability >50%, while low-blanchability haplotypes are Ah01HapBL2, Ah05HapBL5/Ah05HapBL6, Ah06HapBL3 and Ah17HapBL2 with <20% blanchability. The identified superior high-blanchability haplotypes, exhibited the highest mean blanchability compared to other haplotype groups and vice-versa for low blanchability haplotypes (Fig. 2).
Diagnostic KASP markers for blanchability
To develop and validate diagnostic markers for blanchability, a total of six SNPs, one each on Ah01, Ah05, Ah06, Ah07, Ah16, and Ah17 were targeted for the development of KASP markers. Primers were successfully developed for all six SNPs; S1_62503335 (TT/CC), S5_107081867 (AA/GG), S6_70882693 (AA/GG), S7_66675581 (AA/GG), S16_34347730 (CC/TT), and S17_133752226 (AA/GG), and validated on a panel with contrasting blanchability (Supplementary Fig. S6). The validation panel comprised of low and high-blanchability genotypes and breeding lines ranging from 6% to 80%. Of the six KASP markers selected for validation, two KASPs (snpAH00567, and snpAH00569) showed the expected polymorphism in the designed validation panel. Interestingly, these two polymorphic KASP markers differentiated genotypes and breeding lines with high and low-blanchability, which belonged to the genomic region detected on chromosomes Ah01 and Ah05 (Fig. 3). Additionally, markers from chromosome Ah17 were also validated in a few genotypes, but majorly the calls were missing, hence lacked confidence. Overall, the validated KASP markers accurately distinguished high and low-blanchability genotypes and breeding lines, and can be used as potential diagnostic markers for selecting blanchability breeding material at the very early stages of the varietal development process.
Discussion
Blanchability, the ease of seed coat removal after roasting, is a critical trait shaping the processing quality and industrial suitability of groundnut5. Despite its substantial economic relevance, it has remained underexplored in breeding programs, partly due to its quantitative nature and genotype-by-environment interactions. Blanchability is an environment-responsive trait and emphasizes the necessity of integrating G×E stability metrics and season-specific haplotypes for informed genomic interpretation and customized groundnut breeding. The increasing availability of high-resolution genomic resources and advanced GWAS methodologies now offers powerful tools to dissect such complex traits28–30.
In this study, we capitalized on the availability of the WGRS dataset in a globally representative mini-core collection to identify 26 significant STAs for blanchability and, importantly, favorable haplotypes for selected STAs (Fig. 1, Supplementary Fig. S5, Table 1). Among these, the STAs on chromosomes Ah01: S1_62503335 (PVE 2.7%) (50_R); Ah05: S5_107081867 (PVE 38.51%) (50PR_R) and Ah17; S17_133752226 (PVE 54%) (50PR, 200PR) were validated successfully using KASP markers, highlighting their strong association with blanchability, and could serve as diagnostic markers. The WGRS-based analysis identified a key STA on chromosome Ah17, accounting for up to 54% of the phenotypic variance, which is higher than any previously reported QTL for blanchability19. While the three newly developed KASP markers worked on the mini-core collection and reference set, only two diagnostic KASP markers, viz., SnpAH00567 (Ah01) and SnpAH00569 (Ah05), were fully functional in breeding populations (Fig. 3). KASP markers, although efficient and cost-effective, have limitations including allele-specific primer sensitivity, assay failures due to flanking-region polymorphisms, challenges in polyploid genomes, limited ability to capture multi-SNP haplotypes, and reduced transferability across diverse genetic backgrounds. The associated SnpAH00565 marker on Ah17, in KASP format presented validation challenges in the breeding lines, highlighting the need for alternative validation platforms, such as gel-based or sequencing-based assays, for high-resolution genotyping of this key STA. Importantly, this study identified contrasting haplotypes at major loci- Ah01, Ah05, Ah06, and Ah17; enabling clear differentiation between high- and low-blanchability classes. Superior haplotypes, such as Ah01HapBL4, Ah05HapBL3, Ah06HapBL5, Ah06HapBL10, and Ah17HapBL6, associated with blanchability levels of 51–69%, offer immediate targets for haplotype-assisted introgression into elite low-blanchability backgrounds. Conversely, haplotypes such as Ah01HapBL2, Ah05HapBL6, Ah06HapBL3, and Ah17HapBL2, each associated with blanchability of <20%, may serve targeted processing needs where seed coat retention is desired (e.g., beer nuts, coated seed products). This dual utility highlights the translational power of haplotype-based selection in customizing groundnut cultivars to meet diverse industrial specifications.
Comparative assessment of biallelic SNPs versus haplotype blocks revealed that haplotypes captured blanchability variation with greater precision, exhibiting narrower phenotypic ranges and clearer mean differentiation than individual SNPs alone (Supplementary Fig. S8). For example, the low-blanchability “T” allele at S1_62503335 spanned a wide phenotypic range (2.06–52.37%), limiting its predictive utility despite statistical significance, whereas the corresponding 10-SNP haplotype Ah01HapBL2 showed a substantially narrower and more consistent blanchability range (4.66–28.30%). These findings demonstrate that haplotype-level analyses provide a more biologically realistic representation of the genomic complexity, effectively overcoming key limitations of single-SNP GWAS and offering greater robustness for translational breeding applications.
Our haplo-pheno analysis provided novel insights, revealing a strong subspecies-specific differentiation in blanchability-associated haplotypes. Specifically, fastigiata genotypes, belonging to Spanish and Valencia bunch types, harbored multiple high-blanchability haplotypes, whereas hypogaea genotypes, including Virginia runner and Virginia bunch types, predominantly carried low-blanchability haplotypes (Fig. 2). This pattern is consistent with recent genome-wide studies that have demonstrated significant genomic divergence between fastigiata and hypogaea, shaped by distinct domestication histories and selection pressures. The previous study reported selective sweeps contributing to key agronomic traits, and also observed that the two subspecies originated from independent domestication trajectories, with strong population structure and ecological adaptation18,31. Additionally, the subspecies-specific regions were also identified under artificial selection related to seed quality and stress response32.
The identified genomic signatures likely underpin the observed haplotypic patterns for blanchability, a trait of increasing importance in groundnut processing and industry. The geographical distribution of these haplotypes further supports an adaptive component: high-blanchability haplotypes were more prevalent in accessions originating from South Asia and South America, suggesting inadvertent selection in response to local processing preferences or post-harvest practices. Our findings integrate phenotypic and genomic evidence, highlighting the role of historical selection and ecological adaptation in shaping the distribution of trait-associated haplotypes across groundnut subspecies.
Although our primary focus was on developing necessary resources for haplotype-based breeding for groundnut blanchability, candidate genes that were in high LD with the STAs suggest potential links to pathways associated with cell wall composition and remodeling, processes known to influence testa adhesion and seed coat traits. Notably, these included genes encoding galactoside 2-alpha-L-fucosyltransferase-like protein and glycerophosphoryl diester phosphodiesterase. In Arabidopsis thaliana, galactoside 2-alpha-L-fucosyltransferase-like protein was found to fucosylate xyloglucan, influencing cell wall integrity and potentially affecting cell adhesion33. Additionally, glycerophosphoryl diester phosphodiesterase has been identified to be involved crucially in cell wall organization34. While these candidate gene associations remain speculative, they may offer leads for future mechanistic dissection of blanchability, without detracting from the immediate utility of the identified haplotypes themselves. In a recent study, a major stable locus on chromosome A05 (qSCPRA05) has been identified through QTL-seq and GWAS analysis. Besides, the study also provides the candidate genes, laccase-encoding gene (Arahy.0C6ZNN) and six diagnostic markers associated with the blanchability35. In our study, we also have found the same chromosome Ah05 to possess superior haplotypes for blanchability and have been successfully validated; however, these loci corresponded to different genomic regions compared with those reported previously35.
In any case, the identified haplotypes can readily be deployed to customize groundnut blanchability. Genotypes such as ICG10890 (rainy), ICG9507 (post-rainy), and ICG297 (post-rainy) possess multiple high-blanchability haplotypes; these genotypes could be strategically used in haplotype-based breeding pipelines to develop cultivars with 80–90% blanchability levels (Fig. 4). Notably, ICG297 has also shown promise for other agronomic traits13, offering a valuable multi-trait donor for the development of elite cultivars. Finally, it is essential to emphasize that the required blanchability levels are not universal; these vary depending on the processing demands. High-blanchability is preferred for confectionery and nut butter applications, while low-blanchability genotypes are ideal for products requiring testa retention5,8. This necessitates a flexible, haplotype-guided breeding strategy that aligns with market-driven end-use demands, making blanchability a tractable and customizable breeding target. In summary, our study provides a robust genomic and haplotype-level framework for the precision breeding of blanchability in groundnut. The identification of subspecies-specific haplotypes, diagnostic markers, and donor genotypes equips breeders with scalable tools to develop next-generation groundnut cultivars tailored for processing efficiency, market value, and industrial compatibility.
Fig. 4. Haplotype-based breeding strategy to customize blanchability into diverse groundnut cultivars.
High-blanchability-associated haplotypes (Ah01HapBL4 (50_R), Ah05HapBL3 (50PR_R), Ah06HapBL05/BL10 (200PR), Ah17HapBL6 (50PR, 200PR)) (yellow star), are predominantly present in Spanish bunch and Valencia bunch groundnut types, whereas Virginia runner and Virginia bunch types commonly possess low-blanchability haplotypes (Ah01HapBL2, Ah05HapBL6, Ah06HapBL3, Ah17HapBL2) (brown star). The cross between high-blanchability genotypes (ICG297 and ICG9507/ ICG10890) is expected to result in improved blanchability (80–90%). Overall, the identified superior high-blanchability haplotypes can be used to tailor high-blanchability in any genetic background, independent of the agronomic type, to match groundnut genotypes to the required industrial demand. This will enable the haplotype introgression from high-blanchability genetic backgrounds into low-blanchability types, enabling the development of Virginia-type cultivars with high-blanchability, and vice-versa. This image is partly generated using BioRender (https://www.biorender.com/).
Conclusion
This study presents a comprehensive dissection of the genetic architecture for groundnut blanchability, a trait of high industrial relevance yet historically neglected in breeding programs. Through the integration of multi-season phenotyping with WGRS-based GWAS, we identified 26 significant SNP-trait associations, with major associations on chromosomes Ah01, Ah05, Ah06, and Ah17. Notably, high-blanchability haplotypes were predominantly associated with the fastigiata subspecies, while low-blanchability haplotypes were confined to hypogaea, underscoring distinct subspecies-specific genetic differentiation. In addition, we optimized a scalable phenotyping method for blanchability, developed and validated diagnostic KASP markers from significant STAs on chromosome Ah01 and Ah05, and identified donor genotypes harboring multiple favorable haplotype signatures. Together, these components constitute a robust genomic toolkit for haplotype-based selection of blanchability in groundnut. Importantly, the strategy established here offers the flexibility to tailor blanchability levels across diverse genetic backgrounds and agronomic types, thereby aligning varietal development with the specific requirements of global processing industries. In summary, this work not only introduces blanchability as a novel, tractable breeding target but also delivers a translational and scalable roadmap for developing next-generation groundnut cultivars with enhanced processing quality and commercial value.
Methods
Plant material and experimental field design
The diversity panel used in this study comprised 184 accessions from the groundnut mini-core collection developed at ICRISAT, India25. The seed material of mini-core was procured from ICRISAT genebank and this panel represents the global genetic diversity of the groundnut germplasm. This mini-core collection of 184 accessions includes accessions from six botanical varieties: hypogaea, hirsuta, fastigiata, peruviana, aequatoriana, and vulgaris. Also, it represents variability for four agronomic types: Spanish bunch (57 genotypes), Valencia bunch (35 genotypes), Virginia bunch (48 genotypes), and Virginia runner (33 genotypes), and the remaining 11 genotypes are currently of unknown botanical varieties and agronomic types in the source dataset (Supplementary Table S3). Field trials were conducted using an alpha-lattice design across two consecutive seasons: the 2022 rainy season (R) and the 2022–2023 post-rainy season (PR) with two replications, at ICRISAT, Patancheru (17°31′48.00″N, 78°16′12.00″E). Each genotype was sown in a 4 m single row with spacing of 10 cm between plants and 60 cm between rows. The experimental field was characterized by red soil with a pH range of 6.0 to 6.3. Recommended standard agronomic practices for groundnut cultivation were implemented throughout both growing seasons to ensure optimal crop growth36. After harvest, blanchability-related phenotypic observations were recorded13 using well-graded mature seeds for each replication7 (Supplementary Table S3).
Phenotyping and data analysis
All the samples of the experiment were phenotyped for blanchability using the protocol by Janila et al.7. Based on the previous findings, both sample sizes (50 g and 200 g) were included to (i) validate reliability under our experimental conditions and ICRISAT customized blancher, and (ii) ensure comparability with previous studies that used larger quantities. In this study, blanching was performed using an ICRISAT-designed rotary blancher at 30 rpm, with blanching durations of 60 sec for 50 g samples and 120 sec for 200 g samples, as optimized by Shah et al.13. Moreover, groundnut grown during the rainy season typically produces lower pod yield and smaller, less uniform kernels due to high humidity, and foliar diseases. In contrast, post-rainy (rabi) crops have higher pod filling, better seed uniformity, and greater kernel availability8. Because of this greater availability, we were able to evaluate a larger 200-g batch only in the post-rainy season13. The weight of the blanched seeds and splits was recorded, and the blanching percentage was calculated using the formula below:
The phenotypic variation observed in the mini-core collection of blanchability was analysed with 50 g rainy season (50_R), 50 g post-rainy season (50_PR), 50g_post-rainy_rainy (50_PR_R), 200 g post-rainy (200_PR)13 (Supplementary Table S3) (Fig. 5).
Fig. 5. Schematic representation of the workflow for dissecting blanchability in groundnut mini-core.
The process includes phenotyping blanchability via heat-induced blanching, seed coat removal, and post-blanching weight assessment over two crop seasons. Genotypic analysis involves GWAS, marker identification, and haplotype cataloging. Heat treatment loosens cell wall adhesion, enabling testa removal. The integration of phenotypic and genotypic data facilitates the identification of superior haplotypes for improved blanchability. This image is partly generated using BioRender (https://www.biorender.com/).
Elucidating genotype-by-environment interaction
Genotype-by-environment (G × E) interaction for blanchability was evaluated across two contrasting environments. A two-way analysis of variance (ANOVA) was first performed to partition the effects attributable to genotype, environment, and their interaction. Interaction structure was examined using the Additive Main effects and Multiplicative Interaction (AMMI) model. The genotype × environment deviation matrix was constructed by removing genotype and environment marginal means, after which singular value decomposition was applied to extract Interaction Principal Component Axes (IPCAs). Genotypic scores from the first two IPCAs were used to compute the AMMI Stability Value (ASV), providing a quantitative measure of stability across environments. Adaptability patterns were visualized, and G×E response trajectories were examined separately for stable and unstable groups. Genotype × environment (G×E) analyses were implemented in R using a two-way ANOVA and AMMI (Additive Main effects and Multiplicative Interaction) modelling based on singular value decomposition. The analysis pipeline was developed using publicly available R packages (R version 4.3.3), including stats (base R), dplyr and tidyr for data processing, and ggplot2 for visualization.
DNA extraction, sequencing, and SNP calling from WGRS data
Genomic DNA for each genotype was extracted using the Nucleospin Plant II kit (Macherey-Nagel, Düren, Germany) following the manufacturer’s protocol, with 100 mg tender leaf tissue used per sample from a single plant for each genotype to extract high-quality genomic DNA. Whole-genome resequencing was performed using the Illumina sequencing platform, achieving an average depth of 10X coverage per accession to ensure reliable variant detection and robust genome-wide analyses. Post-sequencing, the adapter sequences were initially trimmed, followed by filtering out the low-quality reads, i.e., with > 20% low-quality bases (quality value ≤ 7) and > 5% “N” nucleotides using SOAP2 to obtain high-quality sequencing data for further analysis. The filtered clean sequence reads were then mapped onto the cultivated tetraploid reference genome of “Tifrunner v.2.0”37 using SOAP2 software with parameters “-m 300 -x 600 -s 35 -l 32 -v 5 -p 4” followed by calculation of the likelihood of all possible genotypes for each sample using SOAPsnp3 to get the maximum likelihood estimation of the allele frequency in the population37. The low-quality variants were then filtered-out using the stringent criteria for sequencing depth (>10000 and <400), mapping times (>1.5), and quality score (<20); and at the end, the SNP loci with estimated allele frequency ≠0 or 1. In addition, SNPs with ≥50% missing data across genotypes were then further filtered out, and the remaining high-quality SNPs set was considered for further analysis (Supplementary Fig. S1).
Principal component analysis and GWAS for identifying SNPs associated with blanchability
Population structure was estimated using Principal Component Analysis (PCA) based on a genetic distance matrix (Roger’s matrix) to identify genetic clusters with similar genetic backgrounds within the mini-core collection. PCA analysis was conducted using the SelectionTools package in R (Supplementary Fig. S2). LD decay in the groundnut mini-core diversity panel was analyzed using PopLDdecay v.3.2938 and the SelectionTools package in RStudio (Supplementary Fig. S3). To address population structure and relatedness in the GWAS, both the Q matrix (from population structure analysis) and the K matrix (pairwise kinship estimates) were incorporated into the multi-locus model (BLINK and FarmCPU). Together, these matrices effectively manage confounding effects, reducing false-positive associations that could arise from shared ancestry or familial relationships. Quantile-Quantile (Q-Q) plots were well visualized by plotting observed -log10 (p) values against expected -log10 (p) values for genome-wide SNPs, using the CMplot package in RStudio. Manhattan plots were also visualized using CMplot39. To correct spurious associations in the GWAS, a Bonferroni correction was applied at a 5% significance level. The Bonferroni threshold calculation for significant marker-trait associations was determined using the formula α/n, where α is the overall significance level, and n represents the total number of SNPs used in the GWAS analysis. The Bonferroni threshold calculation (p-value) is 6.7 × 10−7 40,41. The percentage of PVE by significant SNPs was obtained from the models used, with PV for each SNP calculated as the squared correlation between phenotype and genotype. The significance threshold for STAs in WGRS-based GWAS analysis was set at –log10 (p) >6, ensuring the selection of reliable SNPs (Fig. 2; Supplementary Fig. S5) (Table 1). To identify STAs, GWAS analysis was performed using the multilocus model, FarmCPU, and BLINK implemented in the Genome Association and Prediction Integrated Tool (GAPIT) package using RStudio42. Multi-locus models BLINK and FarmCPU incorporate both population structure (Q) and kinship (K) matrices during association mapping. Further, the effects of alleles underlying significant and stable SNP markers were analyzed using methods described by Su et al. and Alemu et al.43,44. Genotypes were grouped based on specific SNP alleles, and mean differences were evaluated using the Kruskal-Wallis test followed by Dunn’s post-hoc multiple comparisons with correction for multiple testing (Supplementary Fig. S7). Further, the identified STAs were validated; for this, the DNA was extracted from the l00 mg leaf tissue of 25 samples (10 genotypes and 15 breeding lines) and were included in the marker validation panel.
Haplo-pheno analysis and identification of superior haplotypes for blanchability
For each significant STA identified from the GWAS study, haplotype construction was performed using a region-based and LD-informed strategy to ensure that haplotypes reflected true quantitative trait loci (QTL) peaks rather than isolated SNP signals45–47 (Supplementary Fig. S8). Following the identification of a lead SNP, we examined a 1 Mb genomic window to assess its local LD structure. Representative SNPs were then selected from within the same LD block as the lead SNP, ensuring that the haplotype captured a coherent genomic interval supported by multiple SNPs with comparable statistical strength. Up to ten SNPs were chosen to span a range of LD levels with the lead SNP. These representative SNPs were chosen to span a spectrum of LD levels with the significant SNP: i. low LD (r2 > 0.30 and <= 0.50); ii. SNPs with moderate LD (r2 > 0.50 and <= 0.70), and iii. high LD (r2 > 0.70). This LD binning approach identified the SNPs with varying degrees of correlation, capturing distinct haplotype block patterns around the lead SNP. The genotypic combinations of these representative SNPs were then used to define haplotype groups, such that all available unique patterns within the panel. Haplotype analysis was implemented in R (4.4.1) using publicly available packages, primarily data.table, together with base R functions for genotype filtering, haplotype construction, and data integration. FlapJack software was used to visualize the identified haplotype groups48. Subsequently, to assess the phenotypic impact of each haplotype, ‘haplo-pheno’ analysis was performed. This involved multiple range comparisons of the phenotypic means among all the haplotype groups, which comprised at least three genotypes. One-way ANOVA, followed by Tukey’s HSD post-hoc test, was used to determine statistical significance. The haplo-group(s) with the most optimal/preferred phenotypic mean values were designated as superior haplotype(s).
Development of KASP-based markers
GWAS analysis identified 26 significant STAs, of which 6 highly significant markers were selected for the development of KASP markers (Supplementary Fig. S6). The selected polymorphic SNPs were in genomic regions proximal to candidate genes distributed across three different chromosomes. Further, KASP markers were designed based on 300-base pair upstream and downstream flanking sequences. For each SNP, two allele-specific forward primers and one common reverse primer were synthesized at Intertek Pvt. Ltd., Hyderabad, India. The designed KASP markers were validated using a panel of 25 samples, comprising five high and five low-blanchability genotypes as well as 15 breeding lines with varying blanchability from ICRISAT (Fig. 3).
Peanut base webtool (https://Peanutbase.org/) was used to identify genes within the LD region using GBrowse (cultivated peanut) version 1. The genes corresponding to the STAs were identified using the peanut base reference genome sequence annotations (www.peanutbase.org). The physical positions of the SNP were mapped onto the Arachis hypogaea reference genome v1.0 available in PeanutBase (www.peanutbase.org), and the corresponding genes were identified. The SNP subsiding start and ending position of the gene or exons were explored for the gene discovery (Table 1). Genes involved in regulating cell wall integrity and adhesion, previously reported in the literature, were described in the context of blanchability.
Statistics and Reproducibility
All statistical analyses were performed using biologically independent samples. No technical replicates were pooled for statistical inference. For blanchability phenotyping, each genotype was evaluated across two seasons with two field replications, and mean values were used for downstream analyses. Sample sizes (n) for each analysis are reported in the figure legends and correspond to the number of independent genotypes or breeding lines within each group. Genotype-by-environment interactions were assessed using two-way analysis of variance (ANOVA), followed by Additive Main Effects and Multiplicative Interaction (AMMI) modeling to quantify stability across environments. AMMI Stability Values (ASV) were calculated to identify stable and unstable genotypes.
For GWAS, population structure (Q matrix) and kinship (K matrix) were incorporated into the BLINK and FarmCPU multi-locus models implemented in the GAPIT package in R. Significance thresholds were determined using Bonferroni correction (α = 0.05). Quantile–quantile plots were examined to assess model performance and control of false positives. Allelic effects of significant SNPs were assessed using the Kruskal-Wallis test followed by Dunn’s post-hoc multiple comparisons with correction for multiple testing. For haplotype–phenotype analyses, only haplotype groups comprising at least three biologically independent genotypes were included. Differences among haplotype groups were evaluated using one-way ANOVA followed by Tukey’s honestly significant difference (HSD) post hoc test. Effect sizes were estimated as the proportion of phenotypic variance explained (PVE) for significant SNP-trait associations. No statistical estimates were calculated for groups with n < 3. All analyses were conducted in R (version 4.3.3) using GAPIT, SelectionTools, CMplot, and PopLDdecay packages.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Description of Additional Supplementary Materials
Acknowledgements
The authors are thankful to the Indian Council of Agricultural Research (ICAR) through the ICAR-ICRISAT collaborative project, MARS Inc. USA, and Bill & Melinda Gates Foundation (BMGF), USA, through Tropical Legumes III project, TL3: OPP1114827, High-Throughput Genotyping (HTPG) project: OPP1130244. Priya Shah acknowledges the Joint Council of Scientific and Industrial Research-University Grant Commission (CSIR-UGC), Government of India, for the award of a fellowship for a Ph.D., and ICRISAT for providing research facilities.
Author contributions
M.K.P.: Conceived the idea, supervised, and finalized the manuscript. K.S., R.S.: Contributed seed material and in seed multiplication field trials of the mini-core collection. O.H.P.: Performed the Genotype × Environment analysis in the revised manuscript. P.J.: Contributed the breeding lines for marker validation. P.S.: Phenotyped mini-core collection, performed the analysis, and wrote the manuscript. D.K.M. and S.S.G.: performed the analysis and reviewed, edited, and improved the manuscript. R.A.: performed haplotype analysis and reviewed, edited, and improved the manuscript. M.S., P.Si, C.Z., S.K.B., M.Y., X.W., and R.K.V.: Reviewed, contributed to data interpretation, edited, and improved the manuscript. All authors have read and agreed to the published version of the manuscript.
Peer review
Peer review information
Communications Biology thanks Admas Alemu, Madhusudhana R. Janga, and the other anonymous reviewer(s) for their contribution to the peer review of this work. Primary Handling Editor: Johannes Stortz. A peer review file is available.
Data availability
The phenotypic data used in this work are provided in Supplementary Table S3. The sequencing data generated in this study are deposited in NCBI with bio-project IDs PRJNA1002116, PRJNA490835, and PRJNA490832. Source data for graphs in the main figures is compiled in Supplementary Data 1-7. Supplementary Table S1-S17 are provided in Supplementary Data 8.
Code availability
No custom scripts were generated for this study. All genomic and phenotypic analyses were conducted using publicly available open-source R packages (R version 4.3.3), including GAPIT (v2.0), SelectionTools, CMplot, PopLDdecay (v3.29), and base R functions, together with the dplyr, tidyr and data.table packages. Full details of software tools and versions are provided in the Methods section.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s42003-026-09850-1.
References
- 1.Parmar, S. et al. Recent advances in genetics, genomics, and breeding for nutritional quality in groundnut. Accelerated Plant Breeding, Volume 4: Oil Crops 111–137 (2022).
- 2.Stalker, H. T. Utilizing wild species for peanut improvement. Crop Sci.57, 1102–1120 (2017). [Google Scholar]
- 3.Krapovickas, A., Gregory, W. C., Williams, D. E. & Simpson, C. E. Taxonomy of the genus Arachis (Leguminosae). Bonplandia16, 7–205 (2007). [Google Scholar]
- 4.Chu, Y., Clevenger, J. P., Holbrook, C. C., Isleib, T. G. & Ozias-Akins, P. Registration of two peanut recombinant inbred lines (TifGP-5 and TifGP-6) resistant to late leaf spot disease. J. Plant Regist.16, 635–640 (2022). [Google Scholar]
- 5.Shah, P. et al. Industry perspective, genetics and genomics of peanut blanchability. Plant Sci. 355, 112473 (2025). [DOI] [PubMed]
- 6.Hoover, M. W. A rotary air impact peanut blancher. Peanut Sci.6, 84–87 (1979). [Google Scholar]
- 7.Janila, P., Aruna, R., Jagadish Kumar, E. & Nigam, S. N. Variation in blanchability in Virginia groundnut (Arachis hypogaea L.). J. Oilseeds Res.29, 116–120 (2012). [Google Scholar]
- 8.Janila, P., Nigam, S. N., Pandey, M. K., Nagesh, P. & Varshney, R. K. Groundnut improvement: use of genetic and genomic tools. Front. Plant Sci.4, 23 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wan, L. et al. Transcriptome analysis of a new peanut seed coat mutant for the physiological regulatory mechanism involved in seed coat cracking and pigmentation. Front. Plant Sci.7, 1491 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Na, X. et al. Damage characteristics and regularity of peanut kernels. Trans. Chin. Soc. Agric. Eng.26, 117–121 (2010). [Google Scholar]
- 11.Farouk, S. M., Brusewitz, G. H. & Paulsen, M. R. Blanching of peanut kernels as affected by repeated rewetting-drying cycles. Peanut Sci.4, 63–66 (1977). [Google Scholar]
- 12.Mozingo, R. W. Effects of genotype, digging date and grade on the blanchability of Virginia-type peanuts. Proceedings-APREA. APRES (USA) (1979).
- 13.Shah, P. et al. Identification of high blanchability donors, candidate genes and markers in groundnut. BMC Plant Biol.25, 1409 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Wright, G. C. et al. Breeding for improved blanchability in peanut: phenotyping, genotype× environment interaction and selection. Crop Pasture Sci.69, 1237–1250 (2018). [Google Scholar]
- 15.Lu, Q. et al. A genomic variation map provides insights into peanut diversity in China and associations with 28 agronomic traits. Nat. Genet.56, 530–540 (2024). [DOI] [PubMed] [Google Scholar]
- 16.Liu, Y. et al. Multi-omics profiling identifies candidate genes controlling seed size in peanut. Plants11, 3276 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Zhao, K. et al. Pangenome analysis reveals structural variation associated with seed size and weight traits in peanut. Nat. Genet.57, 1250–1261 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Zheng, Z. et al. Chloroplast and whole-genome sequencing shed light on the evolutionary history and phenotypic diversification of peanuts. Nat. Genet.56, 1975–1984 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Korani, W. et al. De novo QTL-seq identifies loci linked to blanchability in peanut (Arachis hypogaea) and refines previously identified QTL with low coverage sequence. Agronomy11, 2201 (2021). [Google Scholar]
- 20.Abbai, R. et al. Haplotype analysis of key genes governing grain yield and quality traits across 3K RG panel reveals scope for the development of tailor-made rice with enhanced genetic gains. Plant Biotechnol. J.17, 1612–1622 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Singh, P. et al. Superior haplotypes of genes associated with higher grain yield under reproductive stage drought stress in rice. J. Exp. Bot. eraf150 (2025). [DOI] [PMC free article] [PubMed]
- 22.Brinton, J. et al. A haplotype-led approach to increase the precision of wheat breeding. Commun. Biol.3, 712 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Kumar, K. et al. Identification of superior haplotypes for flowering time in pigeonpea through candidate gene-based association study of a diverse minicore collection. Plant Cell Rep.43, 156 (2024). [DOI] [PubMed] [Google Scholar]
- 24.Sinha, P. et al. Superior haplotypes for haplotype-based breeding for drought tolerance in pigeonpea (Cajanus cajan L.). Plant Biotechnol. J.18, 2482–2490 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Upadhyaya, H. D., Bramel, P. J., Ortiz, R. & Singh, S. Developing a mini core of peanut for utilization of genetic resources. Crop Sci.42, 2150–2156 (2002). [Google Scholar]
- 26.Pandey, M. K. et al. High-throughput diagnostic markers for foliar fungal disease resistance and high oleic acid content in groundnut. BMC Plant Biol.24, 262 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Kumar, R. et al. Advances in genomic tools for plant breeding: harnessing DNA molecular markers, genomic selection, and genome editing. Biol. Res.57, 80 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Pandey, M. K. et al. Translational genomics for achieving higher genetic gains in groundnut. Theor. Appl. Genet.133, 1679–1702 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Pandey, M. K. et al. Emerging genomic tools for legume breeding: current status and future prospects. Front. Plant Sci.7, 455 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Varshney, R. K., Nayak, S. N., May, G. D. & Jackson, S. A. Next-generation sequencing technologies and their implications for crop genetics and breeding. Trends Biotechnol.27, 522–530 (2009). [DOI] [PubMed] [Google Scholar]
- 31.Zhang, X. et al. Genome-wide association study of major agronomic traits related to domestication in peanut. Front. Plant Sci.8, 1611 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Liu, Y. et al. Genomic insights into the genetic signatures of selection and seed trait loci in cultivated peanut. J. Adv. Res.42, 237–248 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Rocha, J., Cicéron, F., Lerouxel, O., Breton, C. & de Sanctis, D. The galactoside 2-α-L-fucosyltransferase FUT1 from Arabidopsis thaliana: crystallization and experimental MAD phasing. Acta Crystallogr F. Struct. Biol. Commun.72, 564–568 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Hayashi, S. et al. The glycerophosphoryl diester phosphodiesterase-like proteins SHV3 and its homologs play important roles in cell wall organization. Plant Cell Physiol.49, 1522–1535 (2008). [DOI] [PubMed] [Google Scholar]
- 35.Liu, H. et al. Candidate gene identification and marker development for seed coat peeling rate in peanut (Arachis hypogaea L.). BMC Plant Biol.25, 959 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Nigam, S. N., Jordan, D. L. & Janila, P. Improving Cultivation of Groundnuts. (Burleigh Dodds Science Publishing Limited, 2018).
- 37.Bertioli, D. J. et al. The genome sequence of segmental allotetraploid peanut Arachis hypogaea. Nat. Genet.51, 877–884 (2019). [DOI] [PubMed] [Google Scholar]
- 38.Zhang, C., Dong, S.-S., Xu, J.-Y., He, W.-M. & Yang, T.-L. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics35, 1786–1788 (2019). [DOI] [PubMed] [Google Scholar]
- 39.Yin, L. Package” CMplot”. 2019. URLhttps://github.com/YinLiLin/R-CMplot/blob/master/CMplot.r (2019).
- 40.Bland, J. M. & Altman, D. G. Multiple significance tests: the Bonferroni method. Bmj310, 170 (1995). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Bush, W. S. & Moore, J. H. Chapter 11: Genome-wide association studies. PLoS Comput. Biol.8, e1002822 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Tang, Y. et al. GAPIT version 2: an enhanced integrated tool for genomic association and prediction. Plant Genome9, 2015–11 (2016). [DOI] [PubMed] [Google Scholar]
- 43.Su, J. et al. Genome-wide association study identifies favorable SNP alleles and candidate genes for waterlogging tolerance in chrysanthemums. Hortic. Res. 6, 21 (2019). [DOI] [PMC free article] [PubMed]
- 44.Alemu, A. et al. Genome-wide association analysis unveils novel QTLs for seminal root system architecture traits in Ethiopian durum wheat. BMC Genom.22, 1–16 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zuk, O. et al. Searching for missing heritability: designing rare variant association studies. Proc. Natl. Acad. Sci. USA.111, E455–E464 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Singh, S. et al. Harnessing genetic potential of wheat germplasm banks through impact-oriented-prebreeding for future food and nutritional security. Sci. Rep.8, 12527 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Sehgal, D. & Dreisigacker, S. Haplotypes-based genetic analysis: benefits and challenges. Vav. J. Genet. Breed.23, 803–808 (2019). [Google Scholar]
- 48.Milne, I. et al. Flapjack—graphical genotype visualization. Bioinformatics26, 3133–3134 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Description of Additional Supplementary Materials
Data Availability Statement
The phenotypic data used in this work are provided in Supplementary Table S3. The sequencing data generated in this study are deposited in NCBI with bio-project IDs PRJNA1002116, PRJNA490835, and PRJNA490832. Source data for graphs in the main figures is compiled in Supplementary Data 1-7. Supplementary Table S1-S17 are provided in Supplementary Data 8.
No custom scripts were generated for this study. All genomic and phenotypic analyses were conducted using publicly available open-source R packages (R version 4.3.3), including GAPIT (v2.0), SelectionTools, CMplot, PopLDdecay (v3.29), and base R functions, together with the dplyr, tidyr and data.table packages. Full details of software tools and versions are provided in the Methods section.





