Abstract
Crop domestication has induced a severe genetic bottleneck that reduces the adaptive diversity present in modern cultivars. Standard intra-population genomic prediction models reliant on linear reference genomes and SNPs fail to capture the full spectrum of phenotypic variance hidden in wild relatives. Realizing this potential requires broadening the predictive paradigm from selection within narrow breeding populations toward evolutionary-scale inference across the entire wild-to-cultivated continuum. This missing heritability is largely sequestered within complex structural variations such as presence-absence and copy number variants. These variations drive environmental adaptations but remain obscured by reference bias. To recover these unmapped structural variations the field is evolving from linear coordinates to high-dimensional genomic data representations. We review this transition by contrasting explicit graph topologies that map reticulate evolution with implicit encodings like K-mers that capture sequence composition independent of alignment. Processing these complex and high-dimensional features necessitates advanced computational tools. We synthesize emerging deep learning frameworks and highlight how Graph Neural Networks resolve inheritance paths in topological data while Transformer-based foundation models extract functional syntax from sequence context. These architectures effectively integrate structural variations to resolve non-additive effects such as epistasis missed by traditional models. Computing these hidden structural variations facilitates the precise utilization of wild germplasm. We demonstrate how AI-driven strategies enable zero-shot prediction for uncharacterized wild alleles and optimize genotype-by-environment interactions. Ultimately these approaches pave the way for accelerated de novo domestication of climate-resilient crops.
Keywords: Pangenomics, Structural variation, Deep learning, Missing heritability, Wild germplasm, Genomic prediction
Missing heritability in crop domestication
Crop domestication and improvement follow a trajectory of artificial selection and progressive genetic erosion that typically induces a domestication bottleneck (Fig. 1a) [1, 2]. Prolonged breeding selection has imposed severe genetic cleansing on modern cultivars, eroding unique allelic variation present in wild relatives and ancient landraces [3]. Cultivated soybeans lost over 50% of genetic variation, and approximately 2% to 4% of maize genes experienced strong artificial selection, consistent with findings from large-scale comparative analyses of diverse wild and cultivated genomes [4, 5]. This bottleneck widened selective sweep regions and accelerated fixation or elimination of variants with environmental adaptive value [6]. Beyond eroding adaptive diversity, reduced effective population size and intense directional selection allow deleterious variants to accumulate and fix, constraining selection efficacy and limiting genetic gains in modern breeding programs [7]. Beyond these evolutionary costs, the methods used to survey crop genomes impose their own constraints. Genomic predictions and association studies that rely primarily on SNP arrays or short-read sequencing designed around singular reference genomes harbor biases obscuring variations unique to wild species. Pangenomic investigations in Arabidopsis thaliana demonstrate that a single genome fails to capture the full spectrum of polymorphism required for unbiased characterization [8]. This methodological limitation suggests the missing heritability depicted in Fig. 1c stems from the inability of low-dimensional linear markers to detect complex variation forms hidden within wild germplasm, such as structural and copy number variants. High-resolution assemblies in white clover substantiate this perspective by revealing that copy number variations drive rapid environmental adaptation and provide explanatory power inaccessible through traditional marker systems [9].
Fig. 1.

Genetic bottlenecks and linear genomic limitations in crop domestication. a The transition from wild progenitors to modern cultivars imposes a severe genetic bottleneck that erodes adaptive diversity. b Traditional resequencing approaches reliant on single linear reference genomes systematically overlook structural variations unique to wild species such as large insertions which leads to substantial information loss. c The missing heritability iceberg illustrates that while SNPs represent the visible tip utilized in standard models, structural variations, presence-absence variations and copy number variations comprise the submerged genomic dark matter underlying a substantial fraction of phenotypic variance
To overcome the biased representation of genomic differences imposed by a single reference genome and short-read sequencing, pangenomics has emerged to characterize total genetic variation within a species population [10]. Work on the wild soybean genome demonstrated that approximately 20% of the pangenome consists of dispensable sequences containing high-value agronomic alleles lost in cultivated lines [11]. Sequence-oriented graph pangenomes break the constraints of linear coordinates through topological network structures, effectively storing and displaying complete genomic information including complex structural variations. Embracing this complexity is essential for comprehensive population genomics [12], and empirical construction of crop graph pangenomes has provided quantitative evidence to address this challenge. A tomato graph pangenome study demonstrated that when prediction models used only SNPs from a linear reference, estimated average heritability for metabolic traits was 0.33; introducing graph-derived SV markers increased average heritability to 0.41, a rise of 24% [13]. This study indicated structural variations explained approximately 65.9% of heritability, and for most traits, the contribution of SVs exceeded that of SNPs. The root of this phenomenon extends beyond incomplete linkage disequilibrium between structural variations and neighboring reference markers. Standard single nucleotide polymorphism calling is predominantly restricted to genomic regions shared by all individuals and thereby neglects polymorphisms within dispensable genomic regions. Specific variants located within structural variations that may function as causal agents are systematically missed alongside the structural variations themselves. Across the broad genetic background spanning wild and cultivated populations, numerous rare or variety-specific variations constituting core drivers of phenotypic variation remain overlooked in traditional analyses [14].
From a molecular mechanism perspective, these omitted structural variations often carry critical biological functions, particularly regarding crop production traits and adaptability to environmental stress. The tomato super-pangenome revealed immense genetic diversity hidden in wild tomatoes [15]. A cytochrome P450 gene, Sgal12g015720, functions normally in wild tomatoes to increase branching and fruit yield substantially, whereas in all modern cultivated tomatoes, this gene has been pseudogenized by a 244bp deletion in the first exon. Similar structural variations are key factors contributing to the decline in quality of modern crops. The chemical genetic roadmap of tomato flavor mapped by Tieman et al. [16] suggested that modern commercial varieties, in their pursuit of yield and transportability, have lost numerous key alleles controlling flavor volatile synthesis, resulting in flavor quality inferior to traditional heirlooms. Research on tomato anthocyanin accumulation revealed that the inability of most cultivated tomato flesh to accumulate anthocyanins stems from a loss-of-function splicing variant in the key transcription factor SlAN2-like and structural inactivation of the repressor protein gene SlMYBATV, whereas functional versions remain intact in wild tomatoes such as Solanum chilense [17]. These loss-of-function structural variations caused by domestication selection have long remained undetected due to the limitations of single reference genomes. A transcriptome meta-analysis of rice, tomato, and soybean further supported this view, finding that wild relatives exhibit, on average, higher expression of genes related to osmotic stress, drought, and defense [18]. These high expression levels are often driven by structural variations in cis-regulatory regions. Conversely, domesticated varieties are primarily enriched in gene expression related to auxin and yield. This suggests that a substantial amount of variation with adaptive value is found in wild germplasm in the form of structural variations or gene presence-absence variations, which directly regulate the breadth and depth of transcriptional defense responses [19, 20]. Recent evidence demonstrates that domestication alters the chromatin landscape with changes in accessible regulatory regions closely linked to selective sweeps [21]. Findings in soybean highlight that such regulation is not limited to protein-coding genes; complex structural repeats generating long noncoding RNAs have been shown to underlie multiple domestication traits [22]. These non-coding elements are often ignored by standard SNP panels yet act as critical post-transcriptional regulators.
In summary, the missing heritability in crop domestication and artificially improved traits results from the combined effects of reduced genetic diversity due to domestication bottlenecks and the deterministic bias of existing detection technologies. A vast amount of unobserved structural variation serves not only as a carrier of genetic variation but also as a core element regulating gene expression and phenotypic plasticity. Findings by Wainschtein et al. [23] in human genetics corroborate this conclusion, showing that rare variants captured by whole-genome sequencing explain the vast majority of pedigree heritability. The authors further acknowledge that structural variations tagged by these rare variants may contribute to the still-missing heritability. As pangenomics expands research horizons from single linear sequences to graph structures covering cultivars and wild relatives, researchers have confirmed that structural variation is a critical component in filling the heritability gap [13, 14, 24].
It is therefore important to clarify the predictive context in which these tools are proposed. This review extends beyond conventional genomic selection confined to a single breeding population, where prediction accuracy is governed primarily by the relatedness between training and test sets. Rather, we consider a broader spectrum of predictive objectives spanning the wild-to-cultivated continuum: estimating breeding values within populations at one end, and, at the progressively more challenging end, predicting the molecular function of variants in uncharacterized wild germplasm, guiding precision introgression, and informing de novo domestication [25, 26]. Traditional within-population selection thus represents one anchor of this spectrum rather than its entirety, which clarifies why integrating wild relatives, the principal reservoir of untapped structural variation, is essential to the goals considered here.
However, while explicit graph pangenomes theoretically resolve reference bias, significant technical bottlenecks remain in their actual construction and downstream analysis. First, graph structures fundamentally disrupt traditional linear coordinate systems. The existence of multiple paths and varying sequence lengths within the graph makes continuous base numbering impossible, rendering the migration of gene annotations and the cross-path localization of variants highly complex. Second, graph pangenomes face severe computational scalability challenges, known as the graph blowup phenomenon. Introducing too many rare variants to ensure completeness results in overly complex graph topologies, which not only increases memory overhead but also triggers serious multi-mapping issues for reads, paradoxically reducing the accuracy of variant detection. Furthermore, there is a lack of standardized visualization schemes that can simultaneously account for large-scale topological structures and base-level resolution, limiting intuitive interpretation of complex rearrangement regions in large eukaryotic genomes especially when analyzing highly heterozygous or polyploid crops [27]. Consequently the complexity of these topological structures impedes the identification of orthologous genes and regions based on synteny which is essential for elucidating evolutionary history and linking structural variations to evolutionary genomics [28]. These mathematical modeling difficulties and computational efficiency bottlenecks posed by explicit graph topologies constitute core challenges in current computational biology. This necessitates moving beyond traditional explicit graph mapping logic to explore next-generation prediction architectures based on deep learning and implicit feature encoding, such as K-mer embeddings and implicit convolutions, to efficiently integrate high-dimensional, sparse structural variation information with complex non-additive effects.
Novel data representation: from explicit graphs to implicit encodings
The evolution of genomic prediction technology is at a critical turning point, moving from low-dimensional linear matrices to high-dimensional topological structures. The core driver of this transition is the need to break the homogenized constraints of single cultivar reference genomes and embrace heterogeneous information from wild relatives across the breadth of evolution. Traditional SNP arrays and linear resequencing analyses are fundamentally flawed when dealing with substantial genetic divergence between modern cultivars and their wild ancestors. This flaw stems from the limitations of coordinate systems based on single reference genomes. A review by Du et al. [29] explicitly pointed out that pangenomes constructed solely from cultivars still cannot explain the loss of genetic diversity caused by domestication bottlenecks. In contrast, super-pangenome graphs integrating crops and their wild relatives have become the primary pathway to capture genetic dark matter systematically omitted by linear views and by pangenomes of limited phylogenetic scope that include only cultivar genomes [30, 31]. Therefore, the innovation in data representation is not merely an algorithmic optimization but a reconstruction of complex genomic architectures shaped by multiple evolutionary processes within computational models. This requires shifting the data perspective from linear sequence alignment to graph-theory-based topological networks and implicit feature encodings independent of reference systems (Fig. 2) [32].
Fig. 2.

Evolution of genomic data representation for capturing structural diversity. a Explicit graph representation replaces linear coordinate systems to map reticulate evolution directly where divergent wild sequences manifest as alternative topological paths or bubbles rather than unmapped reads to preserve structural variation topology. b Implicit encoding utilizes alignment-free strategies such as K-mer frequency distributions or high-dimensional embeddings to conceptualize the genome as position-independent feature sets and circumvents alignment errors inherent in divergent wild germplasm. c Revaluation of genomic absence contrasts with traditional imputation models by encoding gene absence and presence-absence variation as distinct biological signals indicative of selection sweeps or pseudogenization rather than treating null genotypes as missing data
Explicit graph representations fundamentally resolve the difficulty of positioning large-scale structural variations (SVs) in wild germplasm on linear reference genomes by constructing the genome as a variation graph composed of sequence nodes and edges. Linear reference genomes force all variations to map to the same coordinate axis, causing gene sequences specific to wild species to be frequently mislabeled as unaligned fragments and discarded. Conversely, graph-based super-pangenomes allow new sequences to exist as branching paths or independent nodes, thereby preserving evolutionary structural traces completely. Research by Shang et al. [33] and subsequent updates involving the construction of a rice super-pangenome containing the wild ancestor Oryza rufipogon revealed numerous SVs and presence-absence variations (PAVs) invisible in reference systems containing only cultivars; these variations carry key agronomic traits lost during domestication. Similarly, Gui et al. [34] confirmed in a maize pangenome study that integrating wild relative data via graph structures allows for the precise localization of SVs conferring adaptability to biotic and abiotic stresses, whereas these variations remain invisible in traditional analyses based on the single B73 reference genome. Li et al. [15] further quantified this advantage in studies of tomato pan-genome, finding that unique SVs carried by wild relatives directly regulate gene expression networks. Although the establishment of high-quality reference genomes for wild species partially addresses the visibility of such variations, graph pangenomes provide a necessary unified framework to systematically capture these regulatory elements that remain unmapped in standard cultivar-based alignments. Additionally, Liu et al. [35] found during the construction of graphs for soybean and its wild ancestor Glycine soja that graph structures effectively resolve ultra-large fragment PAVs exceeding 10kb, a level of structural difference that linear models are unable to handle. These findings are further supported by recent integrative approaches in other species such as millet [36] and barley [37], and the methodology is increasingly applied to capture complex variations in livestock genomics to overcome reference bias [38]. These lines of evidence collectively indicate that explicit graphs are not simple collections of variations but are high-fidelity carriers capable of accommodating cross-species evolutionary differences and capturing actual wild-specific variants with demonstrated phenotypic effects.
Although explicit graphs provide precise topological coordinates, the high computational cost of constructing whole-genome variation graphs for thousands of wild germplasm sequencing datasets has prompted researchers to explore parallel strategies based on K-mer implicit representations. This strategy treats the genome as a set of fixed-length oligonucleotides, freeing it from dependence on any reference genome, thus offering unique robustness when dealing with extreme genetic divergence. Karikari et al. [39] noted that K-mer analysis can directly quantify differences in genomic content in an unbiased manner, whether sequences originate from modern varieties or ancient wild species. Early maize research by Lu et al. [40] demonstrated that K-mers effectively capture PAVs in the genome. In recent applications, this implicit encoding has been shown to identify 25% more resistance-related loci than traditional SNP methods, particularly disease resistance gene regions originating from wild species that are entirely absent in reference genomes [41]. K-mer frequency distributions themselves constitute natural high-dimensional feature vectors derived without sequence alignment, and they have served as direct inputs to conventional machine-learning classifiers. Deep learning architectures such as Transformers instead learn distributed sequence embeddings through training, rather than relying on precomputed K-mer frequency-count features, thereby capturing complex genetic patterns implicit in wild germplasm while requiring no complex sequence alignment or variant calling.
In the process of integrating data from wild relatives, handling missing values in traditional genotyping constitutes another major challenge in data representation. Novel representation strategies transform these missing values into biologically meaningful inputs. Across the broad genetic background spanning wild and cultivated species, the failure of many genotype calls is not technical error but a reflection of physical loss or drastic rearrangement of genomic sequences during evolution. Recent studies by Weber et al. [42] and Krusenbaum and Wissuwa [43] emphasize that using pangenome graphs confirms that these low-detection regions often correspond to genuine deletions in the genome, known as PAVs. However, the biological validity of this missing-as-information strategy depends critically on the genotyping technology employed. In low-coverage sequencing-based marker systems such as genotyping-by-sequencing (GBS) and DArTseq, marker absence predominantly reflects stochastic sampling dropout due to insufficient read depth rather than genuine genomic deletions, and therefore carries minimal biological information. In contrast, high-coverage whole-genome sequencing (typically >10× depth) combined with pangenome reference alignment, or long-read sequencing platforms such as PacBio HiFi and Oxford Nanopore Technologies, can reliably distinguish technical missingness from biological absence by directly resolving structural variants at base-pair resolution. For deep learning applications, models should ideally differentiate these two classes of missingness through distinct encoding schemes: treating low-coverage dropouts as uninformative masked values while representing high-confidence absence calls as explicit presence-absence features. Within deep learning frameworks, these missing data points should not be simply imputed but encoded as explicit presence-absence patterns or special tensor states. This processing method enables models to learn the genetic effects carried by gene loss itself, as these genomic absences can be due to negative selection or genetic drift during domestication bottlenecks. Therefore, whether through super-pangenome graphs that materialize structural variations of wild species or through implicit encodings of K-mers and deletion patterns that capture the essence of sequence composition, novel data representations are building a high-dimensional computational foundation [44]. These advancements, coupled with multimodal deep learning architectures [45] and sophisticated statistical machine learning software [25], are bridging the domestication gap and accommodating the full spectrum of genetic variation from wild ancestors to modern cultivars.
Deep learning architectures for SV integration
Capturing local non-additive effects of structural variations
Traditional statistical models such as Bayesian Ridge Regression and Genomic BLUP (GBLUP) are linear regression models that focus on additive variant effects [46, 47]. This linear assumption becomes problematic when processing structural variations. Elements like transposon insertions [48] and large copy number variants frequently induce local non-additive effects and epistatic interactions that are filtered out as statistical noise in conventional frameworks. Deep learning architectures provide a mathematical framework to capture non-linear patterns and facilitate the approximation of high-order features and local epistatic interactions (Fig. 3) thereby assisting the identification of adaptive traits at an evolutionary level [49, 50]. Convolutional Neural Networks (CNNs) and Multi-Layer Perceptrons (MLPs) provide an alternative computational structure to address this biological complexity [51–53]. Research by Gower et al. [54] suggested that CNNs could differentiate neutral evolution selective sweeps and adaptive introgression when processing unphased genomic data. The NovGMDeep model developed by Sehrawat et al. [55] demonstrated through one-dimensional convolutions that networks trained on structural variations and transposons often outperform traditional models in phenotype prediction. Their convolution mechanisms utilize local receptive fields to recognize sequence motifs and linkage patterns without requiring strict marker independence. Vourlaki et al. [56] evaluated the performance of MLPs and CNNs in cross-population prediction in rice noting that deep learning models exhibit distinct advantages over the BayesC and reproducing kernel Hilbert space (RKHS) regression models when dealing with germplasm resources of vastly different genetic backgrounds. As Sendrowski et al. [26] highlighted in a recent review the field is experiencing a methodological shift from individual locus association to sequence-to-function mapping. Deep learning architectures facilitate training unified functional models that approximate the non-linear relationship between genomic context and phenotype thereby mitigating the limitations of traditional methods in handling complex linkage disequilibrium and unobserved structural variations.
Fig. 3.

Deep learning architectures for integrating structural variations. a Graph Neural Networks operate on two distinct classes of genomic graph. In a pangenome variation graph nodes denote DNA sequence segments and edges denote their connections, whereas in a population relationship graph nodes denote individuals and edges denote kinship. Graph convolution and message passing aggregate information from topological neighbors of a target node. Specifically, in the GNN schematic (a, right),
denotes the central target node (representing either a DNA sequence segment or an individual, depending on the input graph class) with its associated feature vector
, while
to
represent its topological neighbors with their respective feature vectors
to
. b Transformer-based Genomic Foundation Models adopt computational strategies from natural language processing by treating DNA sequences as tokens in a biological language, utilizing self-attention mechanisms to capture long-range dependencies within the sequence context. c These architectures converge on sequence-level objectives rather than whole-plant traits, resolving local epistasis where the combined effect of variants departs from the sum of their individual effects and decoding cis-regulatory syntax such as DNA motifs and enhancer-promoter chromatin loops
Resolving positional shifts and epistasis in genomic sequences
Large structural variations such as inversions and major translocations severely disrupt physical coordinates and render alignment-based positional tracking highly unreliable. This positional shift complicates the detection of epistasis where spatially distant regulatory elements interact to modulate traits. Transformer architectures address this challenge by shifting the analytical focus from fixed physical coordinates to contextual sequence relationships [57]. Xu et al. [58] reviewed foundation models in plant molecular biology noting that Transformers adopt computational strategies from natural language processing by treating genomic sequences as tokens in a biological language to infer sequence grammar from massive unaligned datasets of wild and cultivated species. Early models such as DNABERT pioneered this approach by tokenizing sequences into overlapping K-mers before mapping each token to a learned contextual embedding [59], and tokenization strategies have since diversified across K-mer and byte-pair schemes that shape the biological knowledge a model captures [60]. Across these variants the model learns distributed representations during training rather than relying on precomputed K-mer frequency statistics. The core utility of this architecture stems from its self-attention mechanism which calculates dependencies between distant sequence elements regardless of their absolute positions in a reference genome. The Cropformer framework proposed by Wang et al. [61] exemplifies this strategy. By employing a Transformer architecture through large-scale pre-training the model identified critical genetic variations related to maize flowering time and plant height by capturing specific attention patterns associated with artificial selection. Dalla-Torre et al. [62] reported in their research on the Nucleotide Transformer that pre-training on multi-species genomes improves the zero-shot prediction capability for unseen variants suggesting the model internalizes broad evolutionary constraints. To mitigate the computational burden where standard Transformer complexity grows quadratically with sequence length Nguyen et al. [63] developed the HyenaDNA architecture. This approach utilizes implicit convolution to replace standard attention mechanisms, where filter kernels are generated dynamically by a small neural network rather than stored as explicit parameter arrays, facilitating the evaluation of million-base-long sequences at single-nucleotide resolution. The PlantCaduceus model developed by Zhai et al. [64] incorporated the Mamba architecture to further resolve computational efficiency bottlenecks caused by highly repetitive sequences in plant genomes. These emerging architectures suggest that predictive models are progressively moving beyond localized genomic windows to evaluate long-range epistatic interactions and complex structural variations within extended contextual frameworks. This paradigm is advancing rapidly, with genome-wide variant-effect models now extending from plant genomes to human (GPN-MSA) and pan-domain (Evo 2) foundation models [65, 66].
Graph neural networks for graph-structured genomic data
Crop evolution frequently involves gene introgression from wild relatives forming a complex reticulate network rather than a simple bifurcating tree. Linear mathematical models compress this multi-dimensional evolutionary history into one-dimensional matrices which systematically discards topological inheritance paths and complex genomic rearrangements. The pangenome variation graphs introduced in Novel data representation: from explicit graphs to implicit encodings section provide a natural substrate for preserving these non-Euclidean relationships. Importantly, algorithmic tools have already demonstrated the computational utility of these graph structures at population scale. Giraffe enables short-read mapping to pangenome graphs containing thousands of embedded haplotypes, facilitating structural variant genotyping across 5202 diverse genomes [67]. GraphTyper2 leverages pangenome graphs for population-scale structural variant genotyping in nearly 50,000 individuals [68]. PanGenie combines haplotype-resolved pangenome references with k-mer counts to perform genome inference across the full variant spectrum without read alignment [69]. While these tools validate pangenome graphs as practical computational platforms, they rely on hand-crafted graph algorithms rather than learnable architectures capable of extracting complex non-linear patterns from graph topology.
Deep learning has recently begun to operate directly on pangenome graph data through diverse architectural strategies. Pangenome-aware DeepVariant converts graph haplotypes into pileup images alongside sample-specific read alignments and applies a CNN to infer genotypes, reducing variant calling errors by up to 25.5% compared to linear-reference approaches [70]. Swave transforms structural variant signals from assembly-derived pangenome graphs into projection wave representations and employs a Recurrent Neural Network to distinguish true variants from repetitive-sequence noise at population scale [71]. Most notably, GNNome introduces geometric deep learning directly on assembly graph topology, leveraging the symmetries inherent to genome graphs to identify paths corresponding to reconstructed genomic sequences and achieving telomere-to-telomere assembly quality [72]. These three approaches represent a progression from indirect graph utilization (image encoding) through sequential modeling (projection waves) to native graph computation (message passing on nodes and edges), with GNNs uniquely preserving the full topological information of the underlying genomic graph.
Beyond direct operation on genomic sequence graphs, GNNs have demonstrated broad utility across other graph-structured representations in genomics. For variant effect prediction, GNN-MAP integrates multimodal annotations on variant knowledge graphs to predict pathogenicity across multiple variant types [73], while complementary approaches combine DNA language model embeddings with Graph Convolutional Networks (GCNs) for disease-specific variant classification [74]. AMR-GNN further demonstrates that multi-representation graph frameworks enable accurate phenotype prediction directly from whole-genome sequencing data [75]. For genomic prediction of quantitative traits, Kihlman et al [76] proposed GCN-RS, which performs graph convolution on population-level genomic relationship graphs constructed from SNP data, achieving up to 49.4% improvement over GBLUP by aggregating information from topological neighbors in the kinship network. Related architectures encode genotypes and environments as interacting graph nodes: Morshedian and Domaratzki [77] fused graph attention with LSTM networks to model genotype-by-environment interactions for maize yield prediction, while contrastive signed graph diffusion has been applied to predict positive and negative gene-phenotype regulatory associations in crops [78]. At the gene network level, causality-aware GNNs classify entire regulatory pathways under different conditions to translate genotype to phenotype [79]. Collectively, these studies demonstrate that the message-passing mechanism of GNNs is inherently suited to capturing the non-Euclidean relational structures pervasive in genomic data [29]. However, it is important to note that the studies reviewed above employ graph structures as alternative representations of population relationships, variant annotations, or gene-phenotype networks rather than operating directly on pangenome variation graphs. To date, no study has applied GNN architectures to perform convolution directly on crop pangenome variation graphs for genomic prediction or structural variant effect assessment. Given that all foundational components are now in place (pangenome graphs encoding wild-cultivated structural diversity in Novel data representation: from explicit graphs to implicit encodings section and GNNs proven on related graph-structured genomic problems), this convergence represents a promising frontier for resolving the complex genetic effects of wild structural variations through native graph computation. Table 1 summarizes representative studies across these data representations, architectures, and analytical objectives, spanning from variant and locus mapping through within-population genomic prediction to cross-species functional inference.
Table 1.
Representative studies integrating pangenomics, structural variations (SVs), and deep learning methodologies in crop domestication and breeding
| Study | Crop | Data representation | Methodology | Analytical objective | Key findings & contributions |
|---|---|---|---|---|---|
| I. Variant Discovery and Pangenome Mapping | |||||
| Zhou et al. [13] | Tomato | Graph Pangenome (838 accessions) | LMM/GWAS (PacBio HiFi) | Heritability dissection & causal-SV mapping | Benchmarked missing heritability: SVs increased heritability by∼24%; mapped causal variants for flavor. |
| Shang et al. [33] | Rice | Super-pangenome (251 accessions) | GWAS (ONT/Illumina) | SV/domestication-locus mapping | Identified SVs/PAVs in NLR genes and domestication loci (OsSh1, PROG1) missed by linear references. |
| He et al. [80] | Setaria | Graph-based Genome (110 accessions) | Multi-env GWAS | Multi-environment trait–locus mapping | Mapped 68 traits across 13 environments; identified promoter PAVs in SiGW3 regulating yield. |
| Gui et al. [34] | Maize | Variant Graph (721 accessions) | Graph Genotyping/GWAS | SV cataloguing & PAV–trait mapping | Constructed a comprehensive map of 255k SVs; linked gene PAVs to complex agronomic traits. |
| II. Within-Population Genomic Prediction and Association | |||||
| Vourlaki et al. [56] | Rice | Linear Marker Matrix (SNP+SV) | MLP & CNN | Genomic trait prediction (cross-population) | Demonstrated that DL models integrating SV markers outperform Bayesian methods for complex traits, especially when training and target sets are not closely related. |
| Sehrawat et al. [55] | Arabidopsis/Rice | Marker Matrix (SNP, SV, TE) | Deep CNN (NovGMDeep) | Genomic selection (SV/TE-based) | Showed that models trained on SVs/TEs often achieve higher accuracy than SNP-only models. |
| Weber et al. [42] | Canola, Maize | Implicit Encoding (Failed SNPs) | GBM/SVM/GBLUP | Genomic selection (PAV-based) | Demonstrated that missing data (failed SNP calls) serves as a biological proxy for deletions; proved these features achieve predictive accuracy competitive with standard SNP markers. |
| Kihlman et al. [76] | Wheat/Multi | Genomic Relationship Graph (Nearest Neighbor) | GCN with Sub-sampling (GCN-RS) | Genomic prediction (relationship-graph) | Proposed a graph sampling architecture to model non-Euclidean population structures; achieved up to 49.4% MSE improvement over GBLUP by aggregating neighbor information. |
| Wang et al. [61] | Multi | Encoded Marker Matrix (SNP+SV) | CNN-Transformer (Cropformer) | Genomic selection (interpretable) | Outperformed SOTA GS models by up to 7.5%; enables high-resolution gene mining via attention weights and SHAP values. |
| Pan et al. [78] | Cotton/Rapeseed/Wheat | Signed Bipartite Graph (Gene-Phenotype) | Contrastive Signed Graph Diffusion (CSGDN) | Gene–phenotype association mining | Utilized signed graph diffusion and contrastive learning to predict positive/negative gene regulation; outperformed baselines by up to 9.28% AUC in sparse datasets. |
| III. Cross-Species Inference and Foundation Models | |||||
| Benegas et al. [81] | Arabidopsis/Brassicales | Raw DNA Sequence (unaligned) | DNA LM (GPN) | Cross-species variant-effect prediction (unsupervised) | Learned genome-wide variant effects by unsupervised pre-training on eight Brassicales genomes; outperformed conservation scores (phyloP, phastCons). |
| Moore et al. [82] | Arabidopsis→Tomato | Sequence + functional features | Transfer Learning | Cross-species functional transfer | Transferred specialized-metabolism gene-function predictions from a model species to a crop, improving F-measure from 0.74 to 0.92. |
| Mendoza-Revilla et al. [83] | Multi | Raw DNA Sequence (48 genomes) | LLM (AgroNT) | Cross-species functional annotation | A foundation model trained on 48 plant genomes for cross-species functional prediction. |
| Zhai et al. [64] | Multi | Raw DNA Sequence (512 bp) | Mamba (PlantCaduceus) | Cross-species zero-shot variant-effect prediction | Demonstrated exceptional cross-species generalization (Arabidopsis to Maize) for gene annotation and zero-shot deleterious mutation identification. |
| Nguyen et al. [63] | General | Raw DNA Sequence (1M bp) | Implicit CNN (HyenaDNA) | Long-range sequence modeling (method) | Achieved sub-quadratic scaling to model million-bp contexts, enabling the capture of large SVs at single-nucleotide resolution. |
Abbreviations: LMM Linear Mixed Model, GWAS Genome-Wide Association Study, GBM Gradient Boosting Machine, SVM Support Vector Machine, CNN Convolutional Neural Network, MLP Multi-Layer Perceptron, GBLUP Genomic Best Linear Unbiased Prediction, GCN Graph Convolutional Network, LLM Large Language Model, LM Language Model, PAV Presence-Absence Variation, SOTA State-Of-The-Art, TE Transposable Element, MSE Mean Squared Error, AUC Area Under the Curve, SHAP SHapley Additive exPlanations, NLR Nucleotide-binding Leucine-rich Repeat, DL Deep Learning
Applications: bridging the wild-cultivated gap
The integration of deep learning and pangenomics is advancing crop improvement from standard genomic selection to computational precision breeding (Fig. 4) [84]. This transition extends beyond utilizing genetic variation within cultivars and aims to bridge the genetic gap between wild and cultivated species [85, 86]. By evaluating uncharacterized variations in wild genomes parsing complex genotype-environment interactions and accelerating the de novo domestication of wild plants artificial intelligence is optimizing the utilization of biological diversity to address global climate change and food security challenges [87]. Throughout this review, artificial intelligence (AI) refers specifically to advanced deep learning architectures that capture non-linear epistatic effects and structural variations, in contrast to traditional linear approaches such as genome-wide association studies (GWAS) and GBLUP. Critically, although some adaptive wild variation has already entered modern cultivars through historical introgression, a large fraction of the structurally complex, environmentally adaptive alleles in wild relatives remains sequestered outside cultivated gene pools by linkage drag and reproductive barriers [11]. A cultivar-only pangenome therefore systematically misses this un-introgressed reservoir, which is precisely the genomic dark matter that precision introgression and de novo domestication aim to access [29].
Fig. 4.

AI-driven strategies for bridging the wild-cultivated gap. a Overcoming phenotypic scarcity a genomic foundation model, namely a deep learning model pretrained on massive unlabeled genomic sequences such as AgroNT or PlantCaduceus to learn broad biological rules, scores variant effects across the wild germplasm bank, and these variant effect predictions in turn nominate high-potential rare alleles linked to target traits such as drought and disease resistance, rather than predicting whole-plant traits directly from sequence. b Resolving G × E interactions deep learning integrates pangenomic data with multidimensional environmental variables to model epistatic and genotype-by-environment effects and predict phenotype plasticity under future climate scenarios. c Targeting domestication loci a workflow where AI, referring here to deep learning architectures that capture non-linear epistatic effects rather than traditional linear approaches such as GWAS and GBLUP, identifies key domestication loci in resilient wild plants guiding precision gene editing to create new climate-ready crops
Overcoming phenotypic scarcity in wild germplasm via zero-shot learning
Evaluating the adaptive potential of wild relatives is critical to overcoming the yield stagnation in modern cultivars but this process is severely hindered by the scarcity of phenotypic records for uncharacterized wild accessions. Traditional GWAS rely on the statistical significance of allele frequencies making it difficult to detect structural variations (SVs) that are abundant in wild populations but rare or absent in cultivars [14, 56]. However, genomic foundation models based on deep learning provide an alternative approach through zero-shot learning. Such foundation models are deep learning architectures pretrained on large-scale unlabeled multi-species genomic sequences, such as AgroNT and PlantCaduceus, from which they learn broad biological rules that can subsequently be adapted to diverse downstream predictive tasks. MacNish et al. [88] noted that machine learning can transfer biological knowledge across species, including from well-characterized models to crop genomes [64, 82], a strategy that helps overcome the challenge of scarce phenotypic data for orphan crops and wild germplasm resources. Wójcik-Gront et al. [89] further emphasized that deep learning architectures such as CNNs can directly analyze large-scale genomic sequences to identify specific motifs and structural variations thereby predicting potential agronomic traits in the absence of specific phenotypic records [51, 55]. In this context DNA sequence is processed as sequential data with underlying syntax. Foundation models proposed by Zhu et al. [50] demonstrate how large language models based on Transformer architectures process massive genetic determinants through pre-trained encoders [61, 62, 83]. This approach allows researchers to infer the functional impact of specific structural variations in wild genomes by extracting features from sequence context without direct phenotypic supervision. This capability for direct sequence-to-function prediction enables breeders to locate rare variants that carry adaptive functions but are often filtered as noise in traditional statistical models effectively utilizing the genetic potential in neglected wild germplasm [90–94].
A critical practical question is whether such cross-species transfer has been empirically demonstrated and how phylogenetically close the source and target species must be to remain informative. Recent foundation models provide concrete evidence that transfer succeeds for tasks governed by conserved sequence grammar. PlantCaduceus, a DNA language model pretrained on sixteen angiosperm genomes, was fine-tuned on a small set of labeled Arabidopsis data yet transferred to maize, a lineage that diverged approximately 160 million years ago across the monocot–dicot boundary, improving splice-donor and translation-initiation-site prediction by 1.45-fold and 7.23-fold over the best prior DNA language model and prioritizing deleterious variants with threefold lower minor-allele frequencies than alignment-based methods [64]. At broader scale, AgroNT, trained on 48 plant genomes, attained state-of-the-art regulatory annotation and tissue-specific expression prediction across species and characterized the effects of over ten million mutations in cassava [83], while the multi-species Nucleotide Transformer achieved robust zero-shot variant prioritization even in low-data regimes [62]. Transferability is nonetheless bounded by phylogenetic distance. A systematic evaluation of cross-species histone-modification prediction reported that accuracy declined steadily with divergence, with a Poaceae family-level model jointly trained on rice and maize generalizing to unprofiled grasses whereas cross-family transfer remained inconsistent [95]. This boundary has a molecular basis: a comparative survey of 284 plant genomes spanning 300 million years found that cis-regulatory sequences turn over rapidly, with only a small fraction of conserved non-coding elements retained across deep evolutionary time [96]. The achievable transfer range is therefore task-dependent, as deeply conserved coding and splicing syntax supports cross-family generalization whereas the faster-evolving regulatory grammar degrades with phylogenetic distance. This limit is sharpest for complex quantitative traits, where genomic-prediction accuracy decays approximately linearly with the genetic distance between training and target populations even within a single species [97]; the substantially greater divergence separating distinct species therefore imposes a still tighter bound, and cross-species prediction of traits such as yield remains unproven and is expected to require comparatively close phylogenetic relationships, motivating integrative cross-domain prediction frameworks as the way forward [98].
Resolving linkage drag and genotype-by-environment interactions
Introducing adaptive traits from wild genomes into cultivated crops introduces significant challenges primarily overcoming complex genotype-by-environment (G×E) interactions and the linkage drag of deleterious alleles [25]. Recent genomic analyses at the wild-crop interface have revealed that permeable species boundaries facilitate significant adaptive introgression reshaping the genomic composition of populations in response to selective pressures [99, 100]. Evolutionary investigations further demonstrate that the colonization of arid habitats drives the rewiring of stress-responsive gene networks and leaves distinct footprints of selective sweeps at loci underpinning local adaptation [101, 102]. Consequently wild relatives growing in such extreme environments have accumulated abundant genes for abiotic stress tolerance and climate resilience [103, 104] yet the performance of these genes in agricultural settings is strongly modulated by environmental context. Jubair and Domaratzki comprehensively reviewed the advantages of deep learning models in multi-environment trials noting that deep learning integrates heterogeneous environmental covariates such as meteorological data and soil parameters to capture non-linear dynamic interactions between the genome and the environment [105, 106] including the complex genetic mechanisms underlying intergenerational stress memory for traits like salinity tolerance [107]. Leveraging such capacity facilitates moving beyond classical genomic selection paradigms to pinpoint wild adaptive alleles and forecast their phenotypic stability under future climate scenarios. Research by Wójcik-Gront et al. [89] confirmed that AI-driven predictive analysis identifies heat tolerance and disease resistance genes from wild wheat and barley while simulating crop performance under environmental stress to guide the selection of resilient varieties. Addressing the limitation that wild traits are often accompanied by linkage drag of deleterious traits [108] such as the low yield associated with heat tolerance in wild rice Oryza australiensis requires advanced methodological integration. By combining multimodal data with deep learning models [45] breeders can define target genomic regions and utilize AI-assisted recombination prediction to break such unfavorable linkages and achieve introgression of superior wild alleles. However, the effectiveness of such multimodal deep learning approaches depends on the availability of large population samples and proper representation of genomic variation as discussed in Deep learning architectures for SV integration section, and predictions may remain subject to biases when these conditions are not met. This strategy shifts the utilization of wild germplasm from conventional crossing to precision breeding based on environmental adaptability prediction ensuring that introduced wild genes enhance climate resilience without sacrificing crop agronomic traits.
Targeting domestication loci for precision breeding
Applying artificial intelligence to accelerate the de novo domestication of wild plants offers an alternative strategy for crop breeding to address climate change rather than relying solely on the prolonged improvement of existing cultivars. This approach is equally vital for harnessing the genetic diversity of underutilized crops and their wild relatives to build resilient agricultural systems [109]. The core of this approach involves endowing wild plants that already possess adaptability to extreme environments with agronomic traits required for modern agriculture. Recent advances in phased pangenomes also enable the computational design of ideal haplotypes by purging deleterious structural variants [110]. Phillips et al. [108] emphasized that modifying key domestication genes including those controlling shattering and dormancy allows breeders to develop new crops suitable for agricultural production while retaining the stress-resistant genetic background of wild plants. Navigating this process requires advanced computational frameworks for precision breeding. For instance Zhu et al. [50] demonstrated how AI models can predict optimal allele combinations at the whole-genome level to guide the selection of gene editing targets. Complementing this approach MacNish et al. [88] reviewed how machine learning could in principle transfer gene-function knowledge from well-characterized major crops to less-studied orphan crops or wild species. The individual capabilities underlying this vision have been demonstrated separately rather than as a single integrated pipeline. Transfer-learning models trained on Arabidopsis have predicted specialized-metabolism gene functions across species, including in tomato [82], establishing that learned representations can generalize phylogenetically, and machine-learning frameworks such as QTG-Finder have prioritized causal genes within quantitative trait loci in Arabidopsis and rice [111]. Independently, the translation of domestication knowledge into orphan crops has been achieved through homology-guided genome editing rather than machine learning, as editing of orthologues of known tomato domestication genes rapidly improved productivity traits in the orphan Solanaceae crop groundcherry (Physalis pruinosa) [112, 113], while comparative genomic analyses continue to catalogue the selection signatures that define such domestication loci, as demonstrated in cowpea [114]. A fully integrated workflow, in which machine learning transfers domestication-locus knowledge from a major crop to an orphan crop and the prediction is then experimentally validated, has yet to be reported and represents a key direction for future work. The complexity of these targets was highlighted by Li et al. [115] who revealed that soybean shattering resistance arose from epistatic interactions between adjacent genes Sh1 and Pdh1, which exemplifies the type of non-linear epistatic architecture among linked domestication loci that deep learning models are designed to resolve. This domestication model combining deep learning-based genomic prediction and functional annotation with gene editing shortens the breeding cycle and provides diverse crop options to support the global food system.
Conclusion
The missing heritability in crop domestication is a biological reality rather than a statistical artifact because complex structural variations are systematically overlooked by single linear reference genomes. To recover this lost genetic diversity the field requires transition to high-dimensional genomic representations. Pangenomic graphs and implicit sequence encodings provide the essential data foundation to capture structural variations unique to wild species. Extracting actionable insights from these complex data structures requires moving beyond traditional linear additive models. Deep learning architectures serve as the dedicated computational engines capable of resolving the non-linear epistasis positional shifts and topological relationships inherent to structural variations. The integration of pangenomics and deep learning is a practical necessity for utilizing uncharacterized wild germplasm. By accurately predicting genotype-by-environment interactions and breaking the linkage drag associated with wild alleles this framework enables the precise introgression of adaptive traits. This targeted recovery of wild diversity offers a concrete data-driven pathway to accelerate the de novo domestication of wild plants and breed climate-resilient crops.
Acknowledgements
Not applicable.
Authors' contributions
All authors (Y.W., M.C., Y.M., A.T., and K.W.) contributed to the study conceptualization and writing of the original draft. Y.W., M.C., and Y.M. performed data curation. Y.W. and M.C. prepared the visualizations. A.T. and K.W. supervised the research, and K.W. acquired the funding. All authors reviewed, edited, and approved the final manuscript.
Funding
This work was supported by the Key Research and Development Project of Xinjiang Autonomous Region, China (Grant No. 2025B02008-1), the National Natural Science Foundation of China (Grant No. 32500528), the Natural Science Foundation of Xinjiang Uygur Autonomous Region (Grant No. 2024D01C216), and the “Tianchi Talents” introduction plan.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare that they have no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Yiquan Wang and Minnuo Cai contributed equally to this work.
Contributor Information
Aurélien Tellier, Email: aurelien.tellier@tum.de.
Kai Wei, Email: kaiwei@xju.edu.cn.
References
- 1.Gaut BS, Seymour DK, Liu Q, Zhou Y. Demography and its effects on genomic variation in crop domestication. Nat Plants. 2018;4(8):512–20. [DOI] [PubMed] [Google Scholar]
- 2.Olsen KM, Wendel JF. A bountiful harvest: genomic insights into crop domestication phenotypes. Annu Rev Plant Biol. 2013;64(1):47–70. [DOI] [PubMed] [Google Scholar]
- 3.Dziurdziak J, Podyma W, Bujak H, Boczkowska M. Tracking changes in the spring barley gene pool in Poland during 120 years of breeding. Int J Mol Sci. 2022;23(9):4553. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Zhang H, Mittal N, Leamy LJ, Barazani O, Song BH. Back into the wild—Apply untapped genetic diversity of wild relatives for crop improvement. Evol Appl. 2017;10(1):5–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Zhang H, Jiang H, Hu Z, Song Q, An YqC. Development of a versatile resource for post-genomic research through consolidating and characterizing 1500 diverse wild and cultivated soybean genomes. BMC Genomics. 2022;23(1):250. [DOI] [PMC free article] [PubMed]
- 6.Huang X, Huang S, Han B, Li J. The integrated genomics of crop domestication and breeding. Cell. 2022;185(15):2828–39. [DOI] [PubMed] [Google Scholar]
- 7.Moyers BT, Morrell PL, McKay JK. Genetic costs of domestication and improvement. J Hered. 2018;109(2):103–16. [DOI] [PubMed] [Google Scholar]
- 8.Igolkina AA, Vorbrugg S, Rabanal FA, Liu HJ, Ashkenazy H, Kornienko AE, et al. A comparison of 27 Arabidopsis thaliana genomes and the path toward an unbiased characterization of genetic polymorphism. Nat Genet. 2025;57(9):2289–301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Kuo WH, Wright SJ, Small LL, Olsen KM. De novo genome assembly of white clover (Trifolium repens L.) reveals the role of copy number variation in rapid environmental adaptation. BMC Biol. 2024;22(1):165. [DOI] [PMC free article] [PubMed]
- 10.Matthews CA, Watson-Haigh NS, Burton RA, Sheppard AE. A gentle introduction to pangenomics. Brief Bioinform. 2024;25(6):bbae588. [DOI] [PMC free article] [PubMed]
- 11.Li YH, Zhou G, Ma J, Jiang W, Jin LG, Zhang Z, et al. De novo assembly of soybean wild relatives for pan-genome analysis of diversity and agronomic traits. Nat Biotechnol. 2014;32(10):1045–52. [DOI] [PubMed] [Google Scholar]
- 12.Bao Z, Weigel D. Complexity welcome: Pangenome graphs for comprehensive population genomics. Quant Plant Biol. 2025;6:e34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zhou Y, Zhang Z, Bao Z, Li H, Lyu Y, Zan Y, et al. Graph pangenome captures missing heritability and empowers tomato breeding. Nature. 2022;606(7914):527–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Zhang Z, Viana JPG, Zhang B, Walden KK, Paul HM, Moose SP, et al. Major impacts of widespread structural variation on sorghum. Genome Res. 2024;34(2):286–99. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Li N, He Q, Wang J, Wang B, Zhao J, Huang S, et al. Super-pangenome analyses highlight genomic diversity and structural variation across wild and cultivated tomato species. Nat Genet. 2023;55(5):852–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Tieman D, Zhu G, Resende MF Jr, Lin T, Nguyen C, Bies D, et al. A chemical genetic roadmap to improved tomato flavor. Science. 2017;355(6323):391–4. [DOI] [PubMed] [Google Scholar]
- 17.Sun C, Deng L, Du M, Zhao J, Chen Q, Huang T, et al. A transcriptional network promotes anthocyanin biosynthesis in tomato flesh. Mol Plant. 2020;13(1):42–58. [DOI] [PubMed] [Google Scholar]
- 18.Yumiya M, Bono H. Meta-Analysis of Wild Relatives and Domesticated Species of Rice, Tomato, and Soybean Using Publicly Available Transcriptome Data. Life. 2025;15(7):1088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Bali S, Vining K, Gleason C, Majtahedi H, Brown CR, Sathuvalli V. Transcriptome profiling of resistance response to Meloidogyne chitwoodi introgressed from wild species Solanum bulbocastanum into cultivated potato. BMC Genomics. 2019;20(1):907. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wei K, Stam R, Tellier A, Silva-Arias GA. Copy number variation shapes structural genomic diversity associated with ecological adaptation in the wild tomato Solanum chilense. Mol Biol Evol. 2025;42(8):msaf191. [DOI] [PMC free article] [PubMed]
- 21.Graf C, Winkler TS, Maughan PJ, Stetter MG. Domestication shaped the chromatin landscape of grain amaranth. Nat Commun. 2025;16:10407. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wang W, Duan J, Wang X, Feng X, Chen L, Clark CB, et al. Long noncoding RNAs underlie multiple domestication traits and leafhopper resistance in soybean. Nat Genet. 2024;56(6):1270–7. [DOI] [PubMed] [Google Scholar]
- 23.Wainschtein P, Zhang Y, Schwartzentruber J, Kassam I, Sidorenko J, Fiziev PP. Estimation and mapping of the missing heritability of human phenotypes. Nature. 2026;649:1219–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Liu Z, Wang N, Su Y, Long Q, Peng Y, Shangguan L, et al. Grapevine pangenome facilitates trait genetics and genomic breeding. Nat Genet. 2024;56(12):2804–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Crossa J, Martini JW, Vitale P, Pérez-Rodríguez P, Costa-Neto G, Fritsche-Neto R, et al. Expanding genomic prediction in plant breeding: harnessing big data, machine learning, and advanced software. Trends Plant Sci. 2025;30(7):756–74. [DOI] [PubMed] [Google Scholar]
- 26.Sendrowski J, Bataillon T, Ramstein GP. In silico prediction of variant effects: promises and limitations for precision plant breeding. Theor Appl Genet. 2025;138(8):193. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Kaur H, Shannon LM, Samac DA. A stepwise guide for pangenome development in crop plants: an alfalfa (Medicago sativa) case study. BMC Genomics. 2024;25(1):1022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Wingen LU, Crosbie D, Hu Y, Kemen E, Liu X, Müller MC, et al.. Towards a quantitative view of the NLR gene family 4evolution in the genome space. EcoEvoRxiv; 2025.
- 29.Du ZZ, He JB, Jiao WB. Plant graph-based pangenomics: techniques, applications, and challenges. aBIOTECH. 2025;6(2):361–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Jayakodi M, Shim H, Mascher M. What are we learning from plant pangenomes? Annu Rev Plant Biol. 2025;76(1):663–86. [DOI] [PubMed] [Google Scholar]
- 31.He W, Li X, Qian Q, Shang L. The developments and prospects of plant super-pangenomes: Demands, approaches, and applications. Plant Commun. 2025;6(2):101230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Jayakodi M, Schreiber M, Stein N, Mascher M. Building pan-genome infrastructures for crop plants and their use in association genetics. DNA Res. 2021;28(1):dsaa030. [DOI] [PMC free article] [PubMed]
- 33.Shang L, Li X, He H, Yuan Q, Song Y, Wei Z, et al. A super pan-genomic landscape of rice. Cell Res. 2022;32(10):878–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Gui S, Wei W, Jiang C, Luo J, Chen L, Wu S, et al. A pan-Zea genome map for enhancing maize improvement. Genome Biol. 2022;23(1):178. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Liu Y, Du H, Li P, Shen Y, Peng H, Liu S, et al. Pan-genome of wild and cultivated soybeans. Cell. 2020;182(1):162–76. [DOI] [PubMed] [Google Scholar]
- 36.Wang W, Wu T, Fan G, Zhang S, Liu S, Jiang S. Integrating Pan-genome, GWAS, and Interpretable Machine Learning to Prioritize Trait-Associated Structural Variations in Setaria italica. Plant Commun. 2026;7(3). [DOI] [PMC free article] [PubMed]
- 37.Jayakodi M, Lu Q, Pidon H, Rabanus-Wallace MT, Bayer M, Lux T, et al. Structural variation in the pangenome of wild and domesticated barley. Nature. 2024;636(8043):654–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Liu GE. Exploring cattle structural variation in the era of long reads, pangenome graphs, and near-complete assemblies. J Anim Sci Biotechnol. 2025;16(1):158. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Karikari B, Lemay MA, Belzile F. k-mer-based genome-wide association studies in plants: Advances, challenges, and perspectives. Genes. 2023;14(7):1439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Lu F, Romay MC, Glaubitz JC, Bradbury PJ, Elshire RJ, Wang T, et al. High-resolution genetic mapping of maize pan-genome sequence anchors. Nat Commun. 2015;6(1):6914. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Jaegle B, Voichek Y, Haupt M, Sotiropoulos AG, Gauthier K, Heuberger M, et al. k-mer-based GWAS in a wheat collection reveals novel and diverse sources of powdery mildew resistance. Genome Biol. 2025;26(1):172. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Weber SE, Chawla HS, Ehrig L, Hickey LT, Frisch M, Snowdon RJ. Accurate prediction of quantitative traits with failed SNP calls in canola and maize. Front Plant Sci. 2023;14:1221750. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Krusenbaum L, Wissuwa M. Low-call-rate SNPs and presence–absence variation identified in the rice pan-genome can improve genomic prediction of rice gene bank accessions. Theor Appl Genet. 2025;138(12):1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Alemu A, Åstrand J, Montesinos-Lopez OA, y Sanchez JI, Fernandez-Gonzalez J, Tadesse W, et al. Genomic selection in plant breeding: Key factors shaping two decades of progress. Mol Plant. 2024;17(4):552–578. [DOI] [PubMed]
- 45.Montesinos-Lopez OA, Chavira-Flores M, Kismiantini, Crespo-Herrera L, Saint Piere C, Li H, et al. A review of multimodal deep learning methods for genomic-enabled prediction in plant breeding. Genetics. 2024;228(4):iyae161. [DOI] [PMC free article] [PubMed]
- 46.Montesinos-López OA, Montesinos-López A, Pérez-Rodríguez P, Barrón-López JA, Martini JW, Fajardo-Flores SB, et al. A review of deep learning applications for genomic selection. BMC Genomics. 2021;22(1):19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Montesinos-Lopez A, Crespo-Herrera L, Dreisigacker S, Gerard G, Vitale P, Saint Pierre C, et al. Deep learning methods improve genomic prediction of wheat breeding. Front Plant Sci. 2024;15:1324090. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Bozan I, Achakkagari SR, Anglin NL, Ellis D, Tai HH, Strömvik MV. Pangenome analyses reveal impact of transposable elements and ploidy on the evolution of potato species. Proc Natl Acad Sci. 2023;120(31):e2211117120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Fan W, Guo Z, Wang X, Zhang L, Liu Y, Cai C, Zhang K, Cheng F. Deep learning applications advance plant genomics research. Hortic Plant J. 2025;11(5):1791–806. [Google Scholar]
- 50.Zhu W, Li W, Zhang H, Li L. Big data and artificial intelligence-aided crop breeding: Progress and prospects. J Integr Plant Biol. 2025;67(3):722–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Flagel L, Brandvain Y, Schrider DR. The unreasonable effectiveness of convolutional neural networks in population genetic inference. Mol Biol Evol. 2019;36(2):220–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Torada L, Lorenzon L, Beddis A, Isildak U, Pattini L, Mathieson S, et al. ImaGene: a convolutional neural network to quantify natural selection from genomic data. BMC Bioinformatics. 2019;20(Suppl 9):337. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Nguembang Fadja A, Riguzzi F, Bertorelle G, Trucchi E. Identification of natural selection in genomic data with deep convolutional neural network. BioData Min. 2021;14(1):51. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Gower G, Picazo PI, Fumagalli M, Racimo F. Detecting adaptive introgression in human evolution using convolutional neural networks. Elife. 2021;10:e64669. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Sehrawat S, Najafian K, Jin L. Predicting phenotypes from novel genomic markers using deep learning. Bioinforma Adv. 2023;3(1):vbad028. [DOI] [PMC free article] [PubMed]
- 56.Vourlaki IT, Ramos-Onsins SE, Pérez-Enciso M, Castanera R. Evaluation of deep learning for predicting rice traits using structural and single-nucleotide genomic variants. Plant Methods. 2024;20(1):121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Suzuki S, Horie K, Amagasa T, Fukuda N. Genomic language models with k-mer tokenization strategies for plant genome annotation and regulatory element strength prediction. Plant Mol Biol. 2025;115(4):100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Xu F, Wu T, Cheng Q, Wang X, Yan J. Foundation models in plant molecular biology: advances, challenges, and future directions. Front Plant Sci. 2025;16:1611992. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Ji Y, Zhou Z, Liu H, Davuluri RV. DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics. 2021;37(15):2112–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Zhou X, Wang Z, Shang J, Li YE. DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences. 2025. arXiv preprint arXiv:2512.17126.
- 61.Wang H, Yan S, Wang W, Chen Y, Hong J, He Q, et al. Cropformer: An interpretable deep learning framework for crop genomic prediction. Plant Commun. 2025;6(3):101223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Dalla-Torre H, Gonzalez L, Mendoza-Revilla J, Lopez Carranza N, Grzywaczewski AH, Oteri F, et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nat Methods. 2025;22(2):287–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Nguyen E, Poli M, Faizi M, Thomas A, Wornow M, Birch-Sykes C, et al. Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. Adv Neural Inf Process Syst. 2023;36:43177–201. [Google Scholar]
- 64.Zhai J, Gokaslan A, Schiff Y, Berthel A, Liu ZY, Lai WY, et al. Cross-species modeling of plant genomes at single-nucleotide resolution using a pretrained DNA language model. Proc Natl Acad Sci. 2025;122(24):e2421738122. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Benegas G, Albors C, Aw AJ, Ye C, Song YS. A DNA language model based on multispecies alignment predicts the effects of genome-wide variants. Nat Biotechnol. 2025;43(12):1960–5. [DOI] [PubMed] [Google Scholar]
- 66.Brixi G, Durrant MG, Ku J, Naghipourfar M, Poli M, Sun G, et al. Genome modelling and design across all domains of life with Evo 2. Nature. 2026;652(8112):1349–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Sirén J, Monlong J, Chang X, Novak AM, Eizenga JM, Markello C, et al. Pangenomics enables genotyping of known structural variants in 5202 diverse genomes. Science. 2021;374(6574):abg8871. [DOI] [PMC free article] [PubMed]
- 68.Eggertsson HP, Kristmundsdottir S, Beyter D, Jonsson H, Skuladottir A, Hardarson MT, et al. GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs. Nat Commun. 2019;10(1):5402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Ebler J, Ebert P, Clarke WE, Rausch T, Audano PA, Houwaart T, et al. Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes. Nat Genet. 2022;54(4):518–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Asri M, Chang PC, Mier JC, Sirén J, Eskandar P, Kolesnikov A, et al. Pangenome-aware DeepVariant. bioRxiv. 2025:2025.06.05.657102. 10.1101/2025.06.05.657102.
- 71.Wang S, Xu T, Zhang P, Ye K. Population-level structural variant characterization using pangenome graphs. Nat Genet. 2026;58(3):664–72. [DOI] [PubMed] [Google Scholar]
- 72.Vrček L, Bresson X, Laurent T, Schmitz M, Kawaguchi K, Šikić M. Geometric deep learning framework for de novo genome assembly. Genome Res. 2025;35(4):839. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Yu H, He G, Wang W, Qin S, Wang Y, Bai M, et al. A graph neural network approach for accurate prediction of pathogenicity in multi-type variants. Brief Bioinform. 2025;26(2):bbaf151. [DOI] [PMC free article] [PubMed]
- 74.Ghadie M, Sardaar S, Trakadis Y. Disease-Specific Prediction of Missense Variant Pathogenicity with DNA Language Models and Graph Neural Networks. Bioengineering. 2025;12(10):1098. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Nguyen HA, Peleg AY, Wisniewski JA, Wang X, Wang Z, Blakeway LV, et al. AMR-GNN: a multi-representation graph neural network framework to enable genomic antimicrobial resistance prediction. Nat Commun. 2026;17:3555. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Kihlman R, Launonen I, Sillanpää MJ, Waldmann P. Sub-sampling graph neural networks for genomic prediction of quantitative phenotypes. G3: Genes Genomes Genet. 2024;14(11):jkae216. [DOI] [PMC free article] [PubMed]
- 77.Morshedian A, Domaratzki M. LSTM-attention-guided graph neural networks for integrated genotype-environment modeling in maize yield prediction. PLoS Comput Biol. 2026;22(5):e1013729. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Pan Y, Ji X, You J, Li L, Liu Z, Zhang X, et al. CSGDN: contrastive signed graph diffusion network for predicting crop gene–phenotype associations. Brief Bioinform. 2025;26(1):bbaf062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Triantafyllidis CP, Aguas R. Causality-aware graph neural networks for functional stratification and phenotype prediction at scale. NPJ Syst Biol Appl. 2025;11(1):92. [DOI] [PMC free article] [PubMed]
- 80.He Q, Tang S, Zhi H, Chen J, Zhang J, Liang H, et al. A graph-based genome and pan-genome variation of the model plant Setaria. Nat Genet. 2023;55(7):1232–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Benegas G, Batra SS, Song YS. DNA language models are powerful predictors of genome-wide variant effects. Proc Natl Acad Sci. 2023;120(44):e2311219120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Moore BM, Wang P, Fan P, Lee A, Leong B, Lou YR, et al. Within- and cross-species predictions of plant specialized metabolism genes using transfer learning. In Silico Plants. 2020;2(1):diaa005. [DOI] [PMC free article] [PubMed]
- 83.Mendoza-Revilla J, Trop E, Gonzalez L, Roller M, Dalla-Torre H, de Almeida BP, et al. A foundational large language model for edible plant genomes. Commun Biol. 2024;7(1):835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Farooq MA, Gao S, Hassan MA, Huang Z, Rasheed A, Hearne S, et al. Artificial intelligence in plant breeding. Trends Genet. 2024;40(10):891–908. [DOI] [PubMed] [Google Scholar]
- 85.Bohra A, Kilian B, Sivasankar S, Caccamo M, Mba C, McCouch SR, et al. Reap the crop wild relatives for breeding future crops. Trends Biotechnol. 2022;40(4):412–31. [DOI] [PubMed] [Google Scholar]
- 86.Saad NSM, Neik TX, Thomas WJ, Amas JC, Cantila AY, Craig RJ, et al. Advancing designer crops for climate resilience through an integrated genomics approach. Curr Opin Plant Biol. 2022;67:102220. [DOI] [PubMed] [Google Scholar]
- 87.Cortés AJ, López-Hernández F. Harnessing crop wild diversity for climate change adaptation. Genes. 2021;12(5):783. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.MacNish TR, Danilevicz MF, Bayer PE, Bestry MS, Edwards D. Application of machine learning and genomics for orphan crop improvement. Nat Commun. 2025;16(1):982. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Wójcik-Gront E, Zieniuk B, Pawełkowicz M. Harnessing AI-powered genomic research for sustainable crop improvement. Agriculture. 2024;14(12):2299. [Google Scholar]
- 90.Novakovsky G, Dexter N, Libbrecht MW, Wasserman WW, Mostafavi S. Obtaining genetics insights from deep learning via explainable artificial intelligence. Nat Rev Genet. 2023;24(2):125–37. [DOI] [PubMed] [Google Scholar]
- 91.Novakovsky G, Fornes O, Saraswat M, Mostafavi S, Wasserman WW. ExplaiNN: interpretable and transparent neural networks for genomics. Genome Biol. 2023;24(1):154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Reynolds J, Pan C. Benchmarking interpretability of deep learning for predictive genomics: Recall, precision, and variability of feature attribution. PLoS Comput Biol. 2025;21(12):e1013784. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Majdandzic A, Rajesh C, Koo PK. Correcting gradient-based interpretations of deep neural networks for genomics. Genome Biol. 2023;24(1):109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.van Hilten A, Kushner SA, Kayser M, Ikram MA, Adams HH, Klaver CC, et al. GenNet framework: interpretable deep learning for predicting phenotypes from genetic data. Commun Biol. 2021;4(1):1094. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Lv T, Han Q, Li Y, Liang C, Ruan Z, Chao H, et al. Cross-species prediction of histone modifications in plants via deep learning. Genome Biol. 2026;27(1):20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Amundson KR, Hendelman A, Ciren D, Yang H, de Neve AE, Tal S, et al. A deep-time landscape of plant cis-regulatory sequence evolution. Science. 2026;392(6800):eadt8983. [DOI] [PubMed] [Google Scholar]
- 97.Scutari M, Mackay I, Balding D. Using genetic distance to infer the accuracy of genomic prediction. PLoS Genet. 2016;12(9):e1006288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Arirangan S, de Oliveira LF, Hasan MN, B Sherman A, Tuinstra M, Brito LF, et al. Sharing approaches in predictive genomics across animals, plants and humans. Nat Genet. 2026;58:503–16. [DOI] [PubMed] [Google Scholar]
- 99.Li LF, Pusadee T, Wedger MJ, Li YL, Li MR, Lau YL, et al. Porous borders at the wild-crop interface promote weed adaptation in Southeast Asia. Nat Commun. 2024;15(1):1182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Wedger MJ, Xiao E, Butts TR, Chlapecka JL, Webster LC, Olsen KM. Recent crop-to-weed adaptive introgression has reshaped the genomic composition and geographical structure of US weedy rice (Oryza spp). Mol Ecol. 2025;34(24):e17604. [DOI] [PubMed] [Google Scholar]
- 101.Wei K, Silva-Arias GA, Tellier A. Selective sweeps linked to the colonization of novel habitats and climatic changes in a wild tomato species. New Phytol. 2023;237(5):1908–21. [DOI] [PubMed] [Google Scholar]
- 102.Wei K, Sharifova S, Zhao X, Sinha N, Nakayama H, Tellier A, et al. Evolution of gene networks underlying adaptation to drought stress in the wild tomato Solanum chilense. Mol Ecol. 2024;33(21):e17536. [DOI] [PubMed] [Google Scholar]
- 103.Singhal R, Izquierdo P, Ranaweera T, Segura Abá K, Brown BN, Lehti-Shiu MD, et al. Using supervised machine-learning approaches to understand abiotic stress tolerance and design resilient crops. Phil Trans B. 1927;2025(380):20240252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Varshney RK, Barmukh R, Bentley A, Nguyen HT. Exploring the genomics of abiotic stress tolerance and crop resilience to climate change. Plant Genome. 2024;17(1). [DOI] [PMC free article] [PubMed]
- 105.Jubair S, Domaratzki M. Crop genomic selection with deep learning and environmental data: A survey. Front Artif Intell. 2023;5:1040295. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Cortés AJ, López-Hernández F, Blair MW. Genome–environment associations, an innovative tool for studying heritable evolutionary adaptation in orphan crops and wild relatives. Front Genet. 2022;13:910386. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Thabet SG, Safhi FA, Börner A, Alqudah AM. Genetic associations determine the effects of intergenerational and transgenerational stress memory for salinity exposure histories in barley. Plant Cell Rep. 2025;44(1):25. [DOI] [PubMed] [Google Scholar]
- 108.Phillips A, Schultz CJ, Burton RA. New crops on the block: effective strategies to broaden our food, fibre, and fuel repertoire in the face of increasingly volatile agricultural systems. J Exp Bot. 2025;76(8):2043–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Stetter MG, Joshi DC, Singh A. Assessing and mining grain amaranth diversity for sustainable cropping systems. Theor Appl Genet. 2025;138(7):171. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Cheng L, Wang N, Bao Z, Zhou Q, Guarracino A, Yang Y, et al. Leveraging a phased pangenome for haplotype design of hybrid potato. Nature. 2025;640(8058):408–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Lin F, Fan J, Rhee SY. QTG-Finder: a machine-learning based algorithm to prioritize causal genes of quantitative trait loci in Arabidopsis and rice. G3: Genes Genomes Genet. 2019;9(10):3129–3138. [DOI] [PMC free article] [PubMed]
- 112.Lemmon ZH, Reem NT, Dalrymple J, Soyk S, Swartwood KE, Rodriguez-Leal D, et al. Rapid improvement of domestication traits in an orphan crop by genome editing. Nat Plants. 2018;4(10):766–70. [DOI] [PubMed] [Google Scholar]
- 113.Yaqoob H, Tariq A, Bhat BA, Bhat KA, Nehvi IB, Raza A, et al. Integrating genomics and genome editing for orphan crop improvement: a bridge between orphan crops and modern agriculture system. GM Crops Food. 2023;14(1):1–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Wu X, Hu Z, Zhang Y, Li M, Liao N, Dong J, et al. Differential selection of yield and quality traits has shaped genomic signatures of cowpea domestication and improvement. Nat Genet. 2024;56(5):992–1005. [DOI] [PubMed] [Google Scholar]
- 115.Li S, Wang W, Sun L, Zhu H, Hou R, Zhang H, et al. Artificial selection of mutations in two nearby genes gave rise to shattering resistance in soybean. Nat Commun. 2024;15(1):7588. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.
