Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2024 Oct 28;67(3):722–739. doi: 10.1111/jipb.13791

Big data and artificial intelligence‐aided crop breeding: Progress and prospects

Wanchao Zhu 1,2, , Weifu Li 3,4, , Hongwei Zhang 5,, Lin Li 2,
PMCID: PMC11951406  PMID: 39467106

ABSTRACT

The past decade has witnessed rapid developments in gene discovery, biological big data (BBD), artificial intelligence (AI)‐aided technologies, and molecular breeding. These advancements are expected to accelerate crop breeding under the pressure of increasing demands for food. Here, we first summarize current breeding methods and discuss the need for new ways to support breeding efforts. Then, we review how to combine BBD and AI technologies for genetic dissection, exploring functional genes, predicting regulatory elements and functional domains, and phenotypic prediction. Finally, we propose the concept of intelligent precision design breeding (IPDB) driven by AI technology and offer ideas about how to implement IPDB. We hope that IPDB will enhance the predictability, efficiency, and cost of crop breeding compared with current technologies. As an example of IPDB, we explore the possibilities offered by CropGPT, which combines biological techniques, bioinformatics, and breeding art from breeders, and presents an open, shareable, and cooperative breeding system. IPDB provides integrated services and communication platforms for biologists, bioinformatics experts, germplasm resource specialists, breeders, dealers, and farmers, and should be well suited for future breeding.

Keywords: artificial intelligence, biological big data, breeding, precision design breeding, systems biology


Artificial intelligence technologies integrate biological big data to assist crop genetics and breeding. Intelligent precision design breeding combines biological techniques, bioinformatics, and breeding art from breeders to enhance crop breeding.

graphic file with name JIPB-67-722-g007.jpg

INTRODUCTION

Crop breeding plays an important role in achieving increased grain production and economic development (Senapati et al., 2022). However, observations spanning approximately 50 years revealed that yields from 24% to 39% of all growing areas have not increased for maize (Zea mays), rice (Oryza sativa), wheat (Triticum aestivum), and soybean (Glycine max) (Ray et al., 2012). With population growth, greater urbanization, and the threats imposed by an unstable climate, grain production for major crops may not provide sufficient food in the future (Ray et al., 20122013). It is imperative to accelerate the development of high‐yielding crop varieties with high nutrient quality and resistance to biotic and abiotic stresses (Lenaerts et al., 2019).

With the rapid accumulation of biological big data (BBD), such as sequenced genomes, transcriptome and proteome data for multiple tissues and conditions, and the development of methods incorporating artificial intelligence (AI), scientists endeavor to effectively harness and integrate vast BBD through AI for crop genetics and breeding, yielding notable progress in designing better crops (Wang et al., 2020; Montesinos‐López et al., 2021). Various new AI methods have been developed in recent years, along with software and online tools (van Dijk et al., 2021; Yan and Wang, 2023). However, up‐to‐date reviews on the applications of AI methods to crop genetics and breeding are lacking, especially on how to combine BBD resources and AI methods to serve crop breeding. In this article, we first review the development of breeding technologies. Then, we discuss the BBD and AI models that are useful in plant genetics and breeding, as well as their attractive performance for accurate predictions. Finally, we offer perspectives on how to use AI‐aided methods for intelligent precision design breeding (IPDB).

GENETICS ACCELERATES THE DEVELOPMENT OF BREEDING TECHNOLOGIES

Breeding refers to the selection of desirable materials from a segregating population, and is based on general genetics theories that describe the genetic basis of trait segregation (Jansen and Nap, 2001). Therefore, breakthroughs in variety development were accompanied by advancements in genetics. For example, maize hybrid varieties were developed based on the theory of heterosis (Crow, 1998), and the Green Revolution in rice was facilitated by the discovery and utilization of the semi‐dwarf1 (sd1) mutant (Sasaki et al., 2002).

Development of gene mining methods for targeted breeding

Gene mining is an important field of genetics, and breeding technologies (such as marker‐assisted selection (MAS) and the generation of transgenic plants) rely heavily on the identification and characterization of functional genes. The functions of several thousands of genes have been elucidated in crops with forward genetics approaches (from phenotype to gene) such as map‐based cloning, genome‐wide association studies (GWASs), quantitative trait locus (QTL) mapping, and bulked‐segregant analysis (BSA) (Liu and Yan, 2019; Wang et al., 2019; Chen et al., 2021; Liu et al., 2022; Zhang et al., 2024); reverse genetics strategies (gene to phenotype) offer a complementary approach to exploring gene function, through methods such as gene editing, overexpression, and vast indexed mutant libraries (Wu et al., 2021). Then, these characterized genes will become choices or tools for molecular design breeding. However, the number of cloned genes is far lower than the total number of genes present in the reference genomes of crops (Wu et al., 2021). Moreover, the breeding landscape and goals will be different when taking into account the changing climate, human pursuits and demands as well as cropping and harvesting systems (Harfouche et al., 2019). Therefore, we still need to develop and optimize gene mining methods to cope with the demand for more functional genes in the era of molecular design breeding. Various methods were developed to accelerate gene mining, including next‐generation BSA, quantitative trait gene (QTG)‐Miner, and machine learning (ML) methods (Han et al., 2023; Wang et al., 2023c2023d). Implementation of these methods will accelerate the identification of more genes underpinning a target trait for further improvement based on these genes.

Germplasm innovations

Germplasm resources are the essential genetic materials for crop breeding (Sahu et al., 2023). Germplasm resources can be used to develop bi‐parental populations for breeding practices, including second‐cycle breeding populations and doubled‐haploid (DH) populations (Darrah and Zuber, 1986; Maqbool et al., 2020). Compared with bi‐parental populations, multi‐parental populations tend to show richer genetic diversity and higher mapping resolution, they mainly include nested association mapping (NAM) populations derived from crossing a common inbred line to several inbred lines, multi‐parent advanced generation inter‐cross (MAGIC) populations constructed by intercrossing several different lines, and complete‐diallel design plus unbalanced breeding‐like inter‐cross (CUBIC) populations constructed following a complete‐diallel design (Kidane et al., 2019; Arrones et al., 2020; Liu et al., 2020; Scott et al., 2020). De novo genome assembly of the founder lines used in generating the segregation populations can reveal abundant genetic variation and improve the power of quantitative mapping studies (Hufford et al., 2021; Shang et al., 2022). Therefore, germplasm resources and their derived populations show great value for genetic mapping and breeding (Xiao et al., 2021; Jiao et al., 2023).

Genetics‐based technologies for crop improvement

Crop improvement involves the selection of plants with the desired phenotype (Figure 1). According to basic genetic principles, favorable traits can be combined into a selected plant after breeders' selection (Wallace et al., 2018). Technologies including MAS, marker‐assisted recurrent selection (MARS), genomic selection (GS), transgenic improvement, and genome editing increase the efficiency of selection and pyramiding of favorable traits. Advancement of these genetics‐based technologies was driven by the development of various biological, genetic, statistical, and bioinformatic techniques and methods (Gao, 2021; Guo et al., 2021; Xu et al., 2021; Alemu et al., 2024). Currently, most of these technologies can be applied to crop breeding, depending on the maturation and cheapness of the integrated techniques and methods (such as genotyping and phenotyping methods). However, we should continue to develop new methods and optimize current methods to increase their efficiency, practicality, and serviceability.

Figure 1.

Figure 1

Development of breeding technologies

Empirical breeding usually includes phenotypic selection, and breeding based on general genetic theories, it is a process of accumulating data. Gene‐based or QTL‐based breeding relies on molecular markers or manipulating genes, it is a process that generates critical information. Model‐based prediction mainly infers GS, on the basis of data and information, a preliminary breeding knowledge is formed. AI breeding embodies wisdom and represents precision design breeding.

BREEDING METHODS CURRENTLY IN USE

Crop improvement relies on the development of breeding methods. Reviewing the history of past breeding methods will help us understand where and how we can move forward. Breeding methods can be classified into empirical breeding, QTL‐based or gene‐based breeding, and model‐based predictions (Figure 1).

Empirical breeding

Empirical breeding was, and is still, carried out by phenotypic observations in the field and indoors to select those lines carrying desired traits (Figure 1). It is somehow akin to domestication, whereby ancient farmers selected plants with higher yields and greater adaptability to the local environment according to their personal experience and knowledge. With the formulation of genetics principles in the late 19th century, controlled crosses and planned experiments supported an increase in the reliability of selection from segregating breeding populations (Figure 1). Pedigree‐based selection is a classical empirical breeding method in which different varieties or inbred lines are crossed to form a segregating population, from which the desirable lines were selected by multi‐environment testing (Gorjanc et al., 2015). There are differences between the testing of crops cultivated as inbred varieties, such as soybean, and those cultivated as hybrid varieties, such as maize. For inbred varieties, the offspring obtained by crossing two parental varieties can be tested directly in the field. For hybrid varieties, two or more heterotic groups are defined based on their genetic relationships, and the progeny obtained by crossing different lines are crossed to a tester line from a distinct heterotic group, forming hybrid lines for trait assessment (Wegenast et al., 2008; Mebratu et al., 2024). For both inbred and hybrid cultivars, the agronomic traits of the tested lines are compared with those of popular varieties (as controls) in tens of locations across multiple growing seasons. Those tested lines showing better performance than the controls are considered candidate new varieties.

QTL‐based or gene‐based breeding technologies

The molecular markers developed in the late 20th century were first used to identify QTLs and describe the genetic diversity of breeding materials (Paran and Michelmore, 1993; Dwivedi et al., 2017). They can genotype plants based on polymorphisms in functional genes or within fine‐mapping QTL intervals. MAS, mostly used for large‐effect genes, facilitates the introgression of favorable alleles into breeding lines (Figure 1). Transgenic breeding and genome editing can manipulate functional genes for trait improvement (Figure 1). At present, transgenic breeding is mostly used to improve resistance to pests and herbicides (Tabashnik and Carrière, 2017; Dong et al., 2021). Genome editing can knock in desirable alleles or knock out unfavorable genes in natural or breeding materials, and is expected to play an important role in molecular design breeding (Li et al., 2024a). Due to disparities in genetic transformation systems, the application of transgenic and genome editing technologies varies across distinct crops. For instance, the transformation efficiency in some crops (such as wheat and millet) remains relatively low, posing a constraint on the advancement of transgenic breeding and genome editing.

Model‐based genomic selection breeding technologies

For traits controlled by many QTLs, a linear model based on the linear regression between the genotypic data and the phenotypic data can help select the best lines (Hadasch et al., 2016). In fact, MAS and MARS can both implement a linear regression model, because the number of QTLs as independent variables remains limited (Massman et al., 2013; Hadasch et al., 2016). Unlike MAS and MARS, GS bypasses QTL detection, and instead employs statistical models to estimate the effects of genome‐wide markers and calculate the genomic estimated breeding values (GEBVs) of genotyped plants (Wang et al., 2018b) (Figure 1). Conventional models currently in use mainly comprise ridge‐regression best linear unbiased prediction (rrBLUP), genomic linear unbiased prediction (GBLUP), and Bayesian models (Wang et al., 2018b). AI‐based models mainly consist of many ML models such as random forest (RF), support vector machine (SVM), gradient boosting machine, and deep learning (DL) (Alemu et al., 2024). Some models are integrated into packages or software often with a graphical user interface, facilitating the implementation of GS for users who are not familiar with program scripts (Chen et al., 2023; Li et al., 2024b).

Research on GS methods and their applications have quickly evolved in the past few years and have been moving in two distinct directions. First, various strategies have been proposed to enhance GS performance, such as incorporating functional genes and QTLs as fixed effects, accounting for G (genotype) × E (environment) interactions, considering the covariance of multiple traits, balancing the relationship between the training and testing populations (Wang et al., 2018b; Xu et al., 2021). Second, researchers are developing better models to integrate various types of BBD and factors (Wu et al., 2024b). For example, deep neural network genomic prediction (DNNGP) can integrate multiple BBD sets, and greatly increase GS performance (Wang et al., 2023b); LightGBM can integrate multiple effects into ML‐based algorithms (Yan et al., 2021). With the improvement of these methods and strategies, GS has now become a mature method for predictive breeding. However, the utilization of GS varies across crops, because the genetic effects of hybrid crops include additive, dominance, and epistasis effects, whereas those of inbred crops include additive and additive by additive epistasis effects. GS models can integrate different genetic effects depending on the constitution of the training population (Wang et al., 2018b).

Limitations of current breeding technologies

Empirical selection and breeding have made a major contribution to crop improvement since the advent of agriculture 10,000 years ago. However, empirical breeding has several inherent shortcomings. First, the varieties were selected based on their phenotypic performance. This was done without prior knowledge of the genetic composition and cannot rely solely on phenotypic observations to identify the best lines that carry all favorable alleles (Gesteiro et al., 2023). Second, the selection of breeding materials by empirical breeding has often resulted in declining genetic diversity compared with that seen in germplasm resources (Maccaferri et al., 2003), limiting the potential for yield increase. Third, empirical breeding is time‐consuming, expensive, and usually inefficient (Alemu et al., 2024).

Quantitative trait locus‐based or gene‐based breeding technologies rely on gene resources that are derived from breeding materials, germplasm resources, wild relatives of crops, or other organisms. This approach works very well for some immunity‐related traits or traits conferring resistance to pests (Zhou et al., 2003; Zeng et al., 2017). For specific quantitative traits, however, only a few genes with marginal effects may be available for gene‐based breeding (Lanubile et al., 2017). Because most agronomic traits are controlled by many minor‐effect QTLs or QTGs, with possible complex interactions with the genetic background and environmental factors (Kasela et al., 2024), QTL‐based or gene‐based breeding technologies usually fall short of the improvement of complex traits in crops.

Model‐based breeding technologies can overcome the limits of gene‐based or QTL‐based breeding. However, they still have some other drawbacks. First, the cost of collecting the amount of data necessary to build a reliable model is extremely high, as each model requires large populations that need to be characterized in genotype and phenotype under multiple environments (He and Li, 2020). Second, genotypic data consist of randomly distributed single‐nucleotide polymorphisms (SNPs) that may be located physically far from the causal genes underlying a trait of interest, when marker density is not high. Thus, using nearby markers rather than a marker at the gene of interest may introduce bias in the estimation of gene effects and breeding values (Anilkumar et al., 2023). Third, unlike large transnational corporations with years of accumulated genotypic, phenotypic, and environmental data, many breeding institutions or programs are isolated entities with limited resources for data collection (Fritsche‐Neto et al., 2021; Zebosi et al., 2024).

Under the increasing pressure of greater grain and food demand worldwide, we need to develop state‐of‐the‐art technologies to accelerate the development of high‐yielding varieties. Genetic mechanisms should be systematically dissected to discover functional genes at the genome scale, and these results can then be integrated into breeding strategies to allow precision breeding.

BBD AND AI MODELS FOR PLANT GENETICS AND BREEDING

Biological big data are derived from plants, tissues, or cells of organisms through various assays, encompassing genomic and RNA sequencing data, epigenomic data, proteomic data, metabolic data, and phenotypic data (Wu et al., 2021). AI is a broad field that includes ML and DL. Machine learning can learn from data and make predictions using algorithms, while DL (a paradigm of ML) uses layered neural networks to process big data and tackle intricate problems. Artificial intelligence models can analyze complex BBD and extract useful information or knowledge by leveraging ML and DL. By processing extensive training datasets, these models autonomously learn to recognize patterns, thereby facilitating informed decision‐making and predictive analytics. In plant genetics and breeding, AI models hold the promise of synthesizing different types of BBD with varying dimensions, styles, ranges, and variations, and can be used for genetic and genomic dissection, network construction and gene mining, prediction of cis‐elements and protein structure, and phenotypic prediction (Figure 2).

Figure 2.

Figure 2

Summary of BBD sources and their applications

(A) Multiple types of BBD were applied in genetic studies. (B) BBD and AI for cis‐elements prediction, protein structure prediction and phenotypic predictions. “Ex” indicates “Expression”; “Me” indicates “Metabolome”; “En” indicates “Environmental factors.”

BBD and AI models for genetic dissection

Population‐based BBD can identify genomic regions associated with expression traits, metabolic traits, protein translation, and agronomic traits (Figure 2A) (Luo, 2015; Zhu et al., 2021; Zhao et al., 2023). A comprehensive analysis of QTLs underlying tested traits and BBD variations can reveal regulatory hotspots, enabling prioritization of the candidate genes (Li et al., 2020b). Combined with an evolutionary analysis and selective sweep analysis, BBD can highlight those genes that are conserved or diversified through evolution or favored during crop domestication and improvement (Figure 2A) (Turner‐Hissong et al., 2020). Significantly, the genomes of the parental accessions used for the genetic dissection of a trait may diverge from the reference genome in population genetics studies. Therefore, incorporating the pan‐genome into the analysis increases the likelihood of uncovering causal variants responsible for phenotypic variation (Figure 2A) (Li et al., 2020a; Wang et al., 2023a). Moreover, BBD can help to identify differentially expressed genes, differentially abundant proteins, or metabolites, as well as differentially methylated regions in the genome (Abnizova et al., 2024), revealing the regulatory mechanism associated with candidate genes or target traits (Figure 2A). Above all, BBD has become an indispensable resource in various aspects of genetic dissection.

On the basis of BBD, AI models with prioritization algorithms can assist the selection of candidate genes for functional validation. QTG‐Finder uses a supervised learning algorithm to rank genes in QTL regions, and is designed for gene prioritization in post‐GWAS analysis (Lin et al., 2019). GPrior, a gene prioritization tool utilizing five positive‐unlabeled bagging classifiers, has notably enhanced the precision of identifying disease‐associated genes (Kolosov et al., 2021). The QTG‐Miner utilizes shortest distance (SD) and ML algorithms to analyze multiple BBD for efficient fine mapping and cloning of QTGs in maize (Wang et al., 2023d). DeepBSA, a BSA method driven by DL, was developed for QTG fine mapping and cloning (Li et al., 2022b). Therefore, combining AI models and BBD has become an efficient method for prioritizing QTGs underlying tested traits.

BBD and AI models for network construction and gene mining

Artificial intelligence models have been applied to network construction based on BBD (Figure 2A). The popular tool, GENIE3, uses tree‐based ensemble methods (RF or Extra‐Trees) to process gene expression matrix for constructing gene regulatory networks (GRNs) (Huynh‐Thu et al., 2010). Another tree‐based pipeline, the regression tree pipeline for spatial, temporal, and replicate data (RTP‐STAR), can be used to infer GRNs by integrating multi‐omics data, aiding the analysis of the regulatory pathways underlying responses to jasmonic acid (Zander et al., 2020). Spatiotemporal clustering and inference of 'omics networks (SC‐ION) adopt an adapted version of GENIE3 to infer transcription factor‐centered GRNs with multi‐omic datasets, assisting the identification of signaling events in response to brassinosteroid (Clark et al., 2021). GRNBoost2, a framework of GENIE3 with stochastic gradient boosting machine algorithms, was developed for inferring cell‐specific GRNs using single‐cell transcription deep sequencing (scRNA‐seq) data (Moerman et al., 2019). The single‐cell regulatory network inference and clustering (SCENIC) pipeline, incorporating GRNBoost2, cis‐Target, AUCell, and nonlinear projection techniques, shows enhanced efficiency in the analysis of high‐dimensional scRNA‐seq data (van de Sande et al., 2020). BTNET applies adaptive boosting and gradient boosting algorithms to infer GRNs with time‐series expression data, while Beacon uses an SVM algorithm to construct GRNs and is implemented to identify many specific regulators of embryonic development (Ni et al., 2016; Park et al., 2018). Generally speaking, AI models can extract the regulatory relationships among biomolecules (such as genes or proteins) from complex BBD resources.

Based on the network properties and interactions of the profiled genes, AI models can be established to predict new functional genes based on their relationships with known genes (Han et al., 2023). A meta GRN constructed by RF with large‐scale RNA‐seq data found 68 transcription factors spanning a variety of metabolic pathways (Zhou et al., 2020). Similarly, a translatome–transcriptome multi‐omics GRN identified a transcription factor related to plant architecture and another transcription factor regulating the drought stress response, which was both empirically validated (Zhu et al., 2023). Based on transcriptome data, the R package ML‐based differential network analysis (mlDNA) uses an RF model to prioritize stress‐responsive genes; in a case study, it identified two genes that regulate the response to salt stress in Arabidopsis (Ma et al., 2014). Therefore, combining AI models and BBD has become an efficient method for network construction and functional gene identification.

BBD and AI models for predicting cis‐elements

Endogenous molecules such as phytohormones, as well as external stimuli such as environmental stresses, can alter gene expression. The responsive biomolecules may bind to regulatory elements in the genes (cis‐elements, especially in promoter and untranslated regions (UTRs)), and stimulate their expression. AI models can learn the sequence features of genes and predict the potential elements affecting their gene expression. Multiple AI models and tools have been developed for cis‐element prediction using BBD (Figure 2B). For example, DeepBind, DL‐based sequence analyzer (DeepSEA), transcription factor impute (TFImpute), DL‐based functional impact of non‐coding variants evaluator (DeFine), and deep feature interaction maps (DFIMs) used DL‐related models to predict TF‐binding and cis‐elements by learning the sequence characteristics from large‐scale chromatin profiles or ChIP‐seq data (Alipanahi et al., 2015; Zhou and Troyanskaya, 2015; Qin and Feng, 2017; Greenside et al., 2018; Wang et al., 2018a). ExPecto, a DL‐based framework, integrates DNA sequences, epigenomic chromatin profiles, and expression data to unveil the regulatory mechanisms of gene expression and the transcriptional consequences of genomic variations (Zhou et al., 2018). In maize, ML models with a k‐mer grammar were used to analyze genome sequences, which can not only perform regional annotations, but can also evaluate the putative effect of sequence variations on regulatory function, becoming a powerful tool to annotate regulatory sequences (Mejía‐Guerra and Buckler, 2019). Based on maize genomic sequences, Saliency, DL important features (DeepLIFT), and Occlusion were used to identify putative cis‐elements, demonstrating that the sequences flanking coding regions harbor important regulatory elements that affect transcription (Washburn et al., 2019). Identification of regulatory elements assisted by AI models is helpful for the discovery of functional elements and genes, facilitating manipulation of these elements to precisely modify gene expression, which has great potential for crop improvement.

BBD and AI models for predicting protein structure

Protein sequences and structure data can help to predict the structure of unknown proteins. Comprehensive protein databases, such as the National Center for Biotechnology Information (NCBI), UniProt, and Protein Data Bank (PDB), have accumulated abundant protein data, including protein sequences, secondary structures, 3D structures, protein–protein interaction data, and protein functions. These data can be used to train the prediction models and establish the relationship between sequence variations and protein functions.

Some AI‐based models/tools have been used for predicting protein structures and functions (Figure 2B), including protein sequence activity relationships (ProSAR), deep‐learning de novo peptide sequencing (DeepNovo), Alphafold, raw multiple sequence alignments (rawMSA), DEEPre, DPPI, BioSeqVAE, protein generative adversarial network (ProteinGAN), MSA VAE (for aligned sequence input), and AR‐VAE (for raw sequence input) (Fox et al., 2007; Tran et al., 2017; Hashemifar et al., 2018; Li et al., 2018; Mirabello and Wallner, 2018; Costello and Garcia Martin, 2019; Hawkins‐Hooker et al., 2021; Jumper et al., 2021; Repecka et al., 2021). These tools use protein sequences or protein–protein interaction data to explore protein functions by predicting their structures and interaction sites. For example, rawMSA predicts secondary structures, DEEPre annotates enzyme functions, and DPPI deduces protein–protein interactions (PPIs) and homodimeric interactions. These methods have high predictive power through modeling the sequence information of proteins (Hashemifar et al., 2018; Li et al., 2018; Mirabello and Wallner, 2018). Notably, the prediction accuracy of Alphafold was close to 80% (Abramson et al., 2024). These AI models are not only very helpful for basic research, but also have great potential in design breeding through identifying functional domains.

BBD and AI models for phenotypic prediction

The premise of efficient phenotypic prediction is to obtain comprehensive and accurate phenotypic data from the training population. Replicated field experiments in multiple environments can estimate the reliability of the phenotypic data, and phenotypic data showing low repeatability or heritability should not be used for training. Phenomic platforms have the facilities to collect large‐scale phenotypic data of various plant traits and AI models (such as ML) enable feature identification from these large amounts of data (Farooq et al., 2024). GS integrates these phenotypic data and genotypic data from a training population to predict the breeding values of a test population (Meuwissen et al., 2001; Jannink et al., 2010). Apart from these data, other types of BBD can be used to predict trait performance, including transcriptomic, metabolic, and environmental data (Figure 2B) (Fu et al., 2012; Dan et al., 2019; Azodi et al., 2020; Gemmer et al., 2020; Li et al., 2022a; Sharma et al., 2022; Yang et al., 2022; Zhu et al., 2022; Bi et al., 2023). Moreover, models that integrate multiple BBD demonstrate superior performance over those based on a single type of BBD (Azodi et al., 2020; Sandhu et al., 2021b).

Both conventional models and AI‐aided models were recently proposed for integrating various BBD sources for phenotypic prediction (Alemu et al., 2024). As conventional statistical models may have difficulty or come up short when dealing with the complexity of genotype‐to‐phenotype (G2P) relationships, several AI models have been proposed, including SVM, RF, multilayer perceptron (MLP), convolutional neural networks (CNNs), deep belief networks (DBNs), genomic breeding machine, DNNGP, soybean deep neural genomic prediction (SoyDNGP), and DeepCCR (a DL method based on CNNs combined with bidirectional long short‐term memory) (Tong and Nikoloski, 2021; Yan et al., 2021; Yang et al., 2022; Wang et al., 2023b; Ma et al., 2024) (Table 1; Figure 2B). However, the superiority of AI models over conventional prediction models warrants further investigation (Montesinos‐López et al., 2021).

Table 1.

Tools that integrate AI models for phenotypic prediction

Method, software, platform AI models Crops Traits References
GS SVM Alfalfa, maize, rice Forage quality traits, yield traits Fu et al. (2012); Xu et al. (2016); Biazzi et al. (2017)
GS RKHS Banana, sugar cane, maize, strawberry, rice, Morphology‐related traits, fruit traits, yield‐related traits, sugar content, plant height, flowering date, nitrogen balance index Gezan et al. (2017); Lyra et al. (2017); Ben Hassen et al. (2018); Nyine et al. (2018); Deomano et al. (2020)
GS SVM, RF Maize Biomass Lima et al. (2018)
GS SVM, RKHS Rice Plant height, yield‐related traits Xu et al. (2018)
GS RKHS, RF Cassava, wheat, wheatgrass, rice Biomass, plant height, yield‐related traits Onogi et al. (2015); Spindel et al. (2015); Battenfield et al. (2016); Zhang et al. (2016); Wolfe et al. (2017)
GS MTDL Wheat Grain yield, days to heading, plant height Montesinos‐López et al. (2019b)
GS MLP Wheat, maize, Arabidopsis Yield‐related traits, flowering traits, dry biomass, pH, lodging, grain color, resistance to leaf rust, fusarium head blight, stripe rust, leaf spot diseases, gray leaf spot Gianola et al. (2011); González‐Camacho et al. (2012); Pérez‐Rodríguez et al. (2012); Montesinos‐López et al. (2018a2018b); Khaki and Wang (2019); Montesinos‐López et al. (2019a2019b); Pérez‐Rodríguez et al. (2020)
GS CNN, MLP Strawberry and blueberry

Soluble solid content,

percentage of culled fruit

Zingaretti et al. (2020)
GS DBN Maize Yield‐related traits, flowering traits Rachmatia et al. (2017)
GS CNN Soybean

Yield‐related traits, protein content,

oil content, moisture content, pH

Liu et al. (2019)
GS SVM, MLP, CNN Wheat Spectral reflectance indices Sandhu et al. (2021a)
DeepGS DL, CNN Wheat Yield‐related traits, SDS sedimentation, grain protein content, and pH Ma et al. (2018)
DLGWAS CNNs Soybean Yield, protein content, oil content, moisture content, height Liu et al. (2019)
LightGBM GBM Maize Days to tasseling, pH, ear weight Yan et al. (2021)
TOP A flexible ML algorithm Maize, rice 18 agronomic traits Yang et al. (2022)
DNNGP Deep neural network Maize, wheat, tomato Multiple traits Wang et al. (2023b)
SoyDNGP A deep and slim network structure Soybean Protein content, oil content, yield‐related traits, flowering traits, pH Gao et al. (2023)
GPformer Transformer‐based DL Soybean, maize, rice, wheat pH, oil content, protein content, yield‐related traits, node number Wu et al. (2024a)
Smart breeding platform Traditional statistical model, ML, and DL Multiple crops Multiple traits Li et al. (2024b)
CropGS‐hub Traditional statistical models and ML Multiple crops Multiple traits Chen et al. (2023)
BreedingAIDB ML Multiple crops Multiple traits Shen et al. (2024)
Crop‐GPA BERT, GCN Multiple crops Multiple traits Gao et al. (2024)
TrG2P Transfer learning Multiple crops Multiple traits Li et al. (2024c)
DeepCCR CNNs and bidirectional long short‐term memory Rice Multiple traits Ma et al. (2024)

BERT, bidirectional encoder representation from transformers; CNN, convolutional neural networks; DBN, deep belief network; DL, deep learning; GBM, genomic breeding machine; GCN, graph convolutional network; GS, genomic selection; ML, machine learning; MLP, multilayer perceptron; MTDL, multi‐trait DL; RF, random forest; RKHS, reproducing kernel Hilbert space; SVM, support vector machine.

Crop breeding entails the concurrent enhancement of various traits, some of which may exhibit correlations or share the same genetic basis. Multi‐trait ML and DL models outperformed their uni‐trait counterparts in predicting grain yield and protein content for wheat, with the RF algorithm and MLP neural networks emerging as the top‐performing models for these predictions (Sandhu et al., 2021a). A multi‐trait ML method named target‐oriented prioritization (TOP) was used to integrate omics traits to improve the prediction of multiple traits, and accurately identify improved candidate varieties that were closest to the ideotype (Yang et al., 2022). Furthermore, TrG2P obtained pre‐trained models by using the CNN algorithm to train models using genotypic non‐yield‐trait phenotypic data, and the convolutional layer parameters are transferred to the yield prediction task to obtain fine‐tuned models for enhanced accuracy (Li et al., 2024c). Collectively, AI models that integrate multi‐trait relationships showed enhanced prediction power compared with conventional models.

A key hurdle for using AI models in crop breeding is the lack of labeled genomic data paired with matching phenotypic information that is provided in a well organized format. Several smart breeding platforms integrated with AI models and accessible G2P paired data have recently been developed to accelerate genome‐designed breeding (Chen et al., 2023; Li et al., 2024b; Shen et al., 2024). For example, CropGS‐Hub provides six different models (BayesCpi, BayesL, BayesR, GBLUP, rrBLUP, and LightGBM) and encompasses a comprehensive collection of over 224 billion genotypic data and 434,000 phenotypic data generated from >30,000 individuals in 14 representative populations belonging to seven major crop species (Chen et al., 2023). BreedingAIDB integrates the LightGBM model with three different boosting types, and contains three databases, including the 143,477 rice G2P paired data for 41 distinct traits, the 284,395 soybean G2P paired data for 114 different traits, and the 12,654 maize G2P paired data for 18 different traits (Shen et al., 2024). The Smart Breeding Platform was proposed for the management and analysis of large‐scale genetic, genomic, and phenotypic data, allowing the user to conduct GWAS and GS using conventional, ML, and DL models (Li et al., 2024b) (Table 1). In addition, RiceNavi (including both web‐based version and package) was designed as a genome navigation system; it used rice quantitative trait nucleotides (QTNs) data to perform in silico breeding simulation, enabling QTN pyramiding and breeding route optimization (Wei et al., 2021).

Multi‐modal AI models

Phenotypic variation is associated with the complex interplay of genomic sequence, gene expressions, and environmental parameters. Therefore, predicting agronomic traits from multi‐modal data—mainly encompassing environmental (E), genotype (G), transcriptome, and phenome (P) data—has become increasingly popular and important (Table 2). For example, DeepG2P utilizes a one‐dimensional convolutional neural network (1D‐CNN) to integrate the SNPs, soil and weather data, improving the prediction performance of crop yields (Sharma et al., 2022). Phenotype–genotype multiple instance learning (PheGeMIL) uses an MLP to capture features from genotypes and employs a small residual network architecture to extract features from multispectral images and thermal images, reaching a 34.8% improvement over a genotype‐only linear baseline approach (Togninalli et al., 2023). Similarly, M2F‐Net used an MLP to capture the agrometeorological data and a pre‐trained DenseNet‐121 model to process red–green–blue (RGB) images of plants (Dhakshayani and Surendiran, 2023). Similarly, MMF‐MTL utilized an MLP to extract vegetation indices features and applied a ResNet50 model to process multispectral images (Cheng et al., 2024). These models typically employ an MLP or 1D‐CNN to capture genetic features, an MLP to handle numerical environmental data, and a CNN to assess image‐calibrated phenotypes. The extracted features are then integrated through various fusion techniques followed by an MLP to forecast crop characteristics. In all these studies, integrating multiple data sources leads to more accurate predictions than relying on a single data modality.

Table 2.

Multi‐modal AI models

Method, software Model for genetic data Model for environmental data Model for phenotypic data Crops Traits References
DeepG2P 1 D‐CNN

1 D‐CNN

& MLP

Maize Yield Sharma et al. (2022)
OpenDroneMap Alignment & Quickshift Maize Height Li et al. (2023)
QTL‐UAV QTL analysis CNN Rice Disease resistance Bai et al. (2023)
PheGeMIL MLP ResNet‐18 Wheat Yield Togninalli et al. (2023)
M2F‐Net MLP DenseNet‐121 Amaranth Over fertilization Dhakshayani and Surendiran (2023)
S‐DNet SPEI DenseNet‐121 Wheat Drought Yao et al. (2024)
MMDL MLP MLP Wheat Yield, weight Montesinos‐López et al. (2023)
MMF‐MTL MLP ResNet‐50 Wheat Growth Cheng et al. (2024)
MulDFNet Transformer Maize Growth Lou et al. (2024)
DEM Transformer Multiple crops Multiple traits Ren et al. (2024)
CropGPT Transformer Transformer Transformer Multiple crops Multiple traits Zhu et al. (2024)

1 D‐CNN, 1D convolutional neural networks; DEM, dual‐extraction modeling; G2P, genotype to phenotype; GPT, generative pre‐training transformer; M2F‐Net, multi‐modal fusion network; MLP, multilayer perceptron; MMDL, multi‐modal DL; MMF‐MTL, multi‐modal fusion and multi‐task DL; PheGeMIL, phenotype–genotype multiple instance learning; QTL, quantitative trait locus; SPEI, standardized precipitation evapotranspiration index; ResNet‐18, ResNet‐50, and DenseNet‐121 are deep convolutional neural networks; UAV, unmanned aerial vehicle.

The accumulation of BBD and the development of algorithms and technologies drive the continuous optimization and iteration of AI models. The transformer model, a DL model based on attention mechanism, can parallelly process diverse data and conveniently adapted to different tasks, showing easy scalability and powerful presentation capabilities. For instance, MulDFNet used a transformer model to process and integrate the hyperspectral imaging, vegetation indices and canopy height (Lou et al., 2024). Dual‐extraction modeling (DEM), a multi‐modal DL architecture incorporating a multi‐head self‐attention mechanism and fully connecting feedforward neural network (FFN), were designed to extract representative features from multi‐omics data, enabling accurate phenotypic prediction and efficient gene mining (Ren et al., 2024). In a parallel development, the CropGPT project put forth the idea of using the transformer architecture for processing and integrating genetic determinants, environmental parameters, and genotypic configurations (Zhu et al., 2024). In summary, multi‐modal AI models leverage multiple BBD types for precision prediction (Yu et al., 2023). However, transformer‐based multi‐modal models require more computational resources to train large amounts of data (Kaddour et al., 2023). In theory, the intricacy of a model's architecture necessitates a greater number of parameters and relies heavily on a substantial volume of high‐quality data for training (Kaplan et al., 2020). These latest model architectures will be capable of processing vast datasets in academic research, refining experimental designs, enhancing breeding efficiency, and promoting precision breeding.

Comparative analysis of AI methods

DL is a type of ML approach that belongs to a subfield of AI. Compared with conventional statistical methods, ML methods (such as SVM, RF, and LightGBM) do not require prior assumptions about the data and can effectively extract nonlinear relationships and capture non‐additive effects (Montesinos‐López et al., 2021). But they usually require complex data preprocessing before modeling, and a large number of noisy features can significantly reduce their predictive performance. In contrast, DL methods can automatically learn feature representations from data and have demonstrated outstanding predictive performance across various tasks. However, DL methods require large amounts of high‐quality data for training and substantial computational resources. As in the GS tasks, no significant differences in predictive performance were detected between conventional statistical methods and ML/DL models (Montesinos‐López et al., 2021). This can be attributed to several factors: (1) not all genomic data exhibit nonlinear characteristics; (2) dataset size and quality may not be sufficient for ML and DL to learn effectively; (3) the fixed DL architecture may not be optimal for the current dataset. Consequently, there is no universal model that consistently outperforms all others across a spectrum of tasks. The selection of the optimal model is contingent upon the specific characteristics of the data.

ARE WE SEEING THE DAWN OF INTELLIGENT PRECISION DESIGN BREEDING FOR CROP BREEDING?

AI model‐based GS, a classical large‐scale breeding model, showed good performance for phenotypic predictions. However, it relies on a large training population with rich phenotypic and genotypic data, which consumes immense computing power. It is necessary to propose an IPDB method suitable for developing countries or small‐sized and medium‐sized enterprises to make efficient use of resources in crop breeding. In contrast to using random markers spread across the entire genome for predictions, IPDB will be designed to use gene‐based functional information for precision prediction (Anilkumar et al., 2023). Instead of relying on conventional models to predict the effects of markers, IPDB should enable the precise estimation of gene effects by integrating gene‐based genotypic data and various BBD types and capturing the multi‐level interactions among genes and environmental factors. This approach will intelligently aggregate favorable alleles for all functional genes and generate an optimal digital plant through model implementation. IPDB will precisely predict breeding values based on gene‐based markers and AI models, and intelligently propose the best breeding materials that contain the desired genetic components.

A roadmap for IPDB

IPDB integrates biological techniques (BT), bioinformatics (IT), and breeding art (AT). It mainly consists of the following steps:

(1) Scientists combine AI and BBD to batch clone functional genes that are associated with agronomic traits, decipher the regulatory pathways of these functional genes, generate a functional gene dictionary, and design liquid DNA chips based on the sequences of these genes. Liquid chips are cheaper and more accurate for SNP detection than traditional solid microarrays.

(2) By genotyping breeding materials, breeders obtain all the informative genotypic data for their breeding stocks. After investigating the genetic relationships among their materials, breeders can create DH lines using their breeding materials or germplasm resources.

(3) Based on the genotypic data of all materials, IPDB predicts the performance of all possible hybrid combinations based on single‐modal (genotypic data) or multi‐modal methods (BBD). Guided by these predictions, breeders execute controlled crosses to produce the proposed hybrids and assess their performance across various environments.

(4) Breeders from different institutions can upload the genotypic data of their materials into the IPDB database to find other breeding lines that have combining ability with the materials from other breeders. In this manner, IPDB establishes connections between different breeders.

(5) Germplasm resources may contain key functional genes that were not detected in the first step or did not show functional variations in breeding materials. AI and BBD will help clone these genes and include them in updated liquid DNA chips.

(6) AI models will predict the best allele combinations for functional genes based on BBD, and scientists can then use gene‐editing technologies to edit the genome based on the results of AI predictions. Thus, new materials can be created specifically according to predefined breeding goals. The best breeding materials selected in the above steps will enter the next breeding cycle (Figure 3).

Figure 3.

Figure 3

Strategies designed for intelligence precision design breeding (IPDB)

AT, breeding art; BT, biological techniques; DH, doubled haploid; IT, bioinformatics.

CropGPT, a case for IPDB

CropGPT is a collaborative IPDB strategy. It relies on interdisciplinary team collaboration, involving the selection of high‐quality germplasm resources, batch cloning of functional genes with ML method, high‐throughput genotyping and phenotypic evaluation, and construction of an open platform providing IPDB service. In particular, it will develop independent pre‐trained encoders to process multi‐modal data, and the big data model will be continuously optimized through iteration to provide optimized breeding suggestions. CropGPT integrates AI and multiple genetic methods to handle various BBDs (Zhu et al., 2024). AI can facilitate the batch cloning of functional genes or QTGs by combining gene networks and ML method (Figure 4A Ⅱ) (Li et al., 2022b; Han et al., 2023), assisting in the predictions of global functional elements or domains by learning features from existing reference data (Figure 4A Ⅲ). The effects caused by the introduction of targeted, non‐native, sequence variants in these elements and domains can be predicted, which can be followed by the deployment of gene‐editing techniques such as knockouts, knock‐ins, and single‐base editing to generate the predicted sequence and new alleles (Figure 4A Ⅳ), possibly generating new phenotypes that are absent in natural germplasm. Artificial intelligence models predict the regulatory relationships among important functional genes based on constructed networks and give the best recommendations for gene combinations (Figure 4A V). At last, the best materials that combine the favorable alleles in important genes will be predicted, which will guide breeders in developing the suggested lines (Figure 4A VI). These materials/varieties can also be added into the existing germplasm resources and be used as starting materials for the next breeding cycle, with all of their associated information (genotype, phenotype, environmental data, gene networks) integrated into CropGPT, thus enhancing its performance (Figure 4A I).

Figure 4.

Figure 4

Combining BBD and AI for crop breeding

(A) CropGPT integrates AI and multiple BBD to support IPDB. (B) CropGPT also is a platform for communication and problem‐solving.

CropGPT is an open, shared, cooperative platform. It can collect the research results of global genetics and breeding literature about crops, integrate basic theoretical knowledge and phenotypic data from different crops, and generate a knowledge map or database to support crop breeding efforts across breeding units (Zhu et al., 2024). Breeders, breeding companies, and farmers can register different types of accounts and log into the platform, and they can seek help or answers from the platform according to their needs or problems, thus realizing real‐time communication among agricultural participants in different roles (Figure 4B). Farmers identify crop‐related problems in the field and provide their questions to the platform through the network. The platform evaluates the problems and provides a possible solution according to its stored knowledge bank. However, when the problem cannot be solved automatically, the platform will contact the breeders (or experts) and establish direct communication between the breeders and the farmers, thus providing alternative ways to solve the problems in crop production. Breeders can compare their breeding experience with the platform database and rapidly answer the farmers' questions; these data will further enrich the breeding database and expand the knowledge of breeders. Using feedback from farmers, breeding companies can optimize effective breeding methods and technologies based on the data in the database and the knowledge of breeders, thus assisting in the targeted improvement of current varieties (Figure 4B).

Comparing different breeding technologies and IPDB

The developments of genetic and biological technologies accelerate the development of breeding technologies, thus increasing breeding efficiency and scale while decreasing cost (Table 3). By contrast, conventional breeding is random, unpredictable, and has low‐throughput, and requires long breeding cycles. Although MAS speeds up breeding, it still takes at least five generations to complete one breeding cycle. Gene‐based breeding methods show obvious pertinence, controllability, and predictability. However, they are mainly used for specific genes, achieving only low throughput at high costs. Taking maize as an example, the cost of single‐gene transformation is approximately US$1,376.04 (http://www.wimibio.com/). Conventional model‐based breeding methods can predict phenotypes based on genotypic data, yet they are characterized by high costs and limited precision. The cost of genotyping one sample with a 20K DNA chip is approximately US$21 (https://www.molbreeding.com). Combining AI and existing BBD, IPDB uses target sequencing and liquid chips to achieve precision phenotypic predictions. It will show high throughput and accuracy, short breeding cycles, and lower costs (Table 3).

Table 3.

Comparison of different breeding technologies and the advantages of IPDB

Breeding approaches Supporting theories or technologies Characteristics Breeding efficiency Marker distribution Cost
Breeding based on phenotypic selection Breeding experience Random, unpredictable

Low‐throughput,

slow

NA NA
Breeding based on general genetic theories

Heterosis,

Mendelian genetics

Some predictability

Low‐throughput,

slow

NA NA
Breeding assisted by molecular marker

MAS, MARS,

MABC

Good controllability and predictability Low‐throughput One or several specific genes NA
Breeding by manipulating genes

Overexpression, RNAi,

gene editing

Good controllability and predictability

Low‐throughput,

slow

One or several specific genes Approximately US$1,376.04 per gene (maize)
Breeding using models and algorithms GS and upgraded GS with ML or DL Some predictability

High‐throughput,

fast

Whole genome Approximately US$8,289 or US$21,012 (3K or 20K chip; 1,000 individuals) (maize)
IPDB AI model and biological big data Good predictability

High‐throughput,

fast

Specific functional gene set (about 5,000 genes) Approximately US$6,907.50 (1,000 individuals) (maize)

DL, deep learning; GS, genomic selection; MABC, marker‐assisted backcrossing; MARS, marker‐assisted recurrent selection; MAS, marker‐assisted selection; ML, machine learning; RNAi, RNA interference.

The challenges of IPDB

IPDB faces several challenges in crop breeding:

(1) It is necessary to deepen our understanding of functional genomes and clone as many functional genes underlying each target trait as possible. Enriching the sets of gene‐based markers will be helpful for improving the predictive accuracy of IPDB. However, it is difficult to clone all functional genes associated with any trait.

(2) Due to the tight linkage of multiple genes, linkage drag, and the pleiotropism effect of key genes, the coordinated regulation of multiple traits cannot be precisely controlled. Moreover, the interactions of functional genes may respond to environmental factors, complicating the estimation accuracies of the effects of functional genes.

(3) The availability of high‐quality environmental and biological data is crucial for accurate phenotypic predictions. Therefore, new methods or tools are required to properly define (optimized recognition algorithms), accurately measure (advanced sensors), comprehensively collect (micro combined with macro), and accurately discriminate different types and sources of data (AI algorithms).

(4) Biological data have high dimensionality, show heterogeneity, and can present multi‐level structures, which makes it difficult to develop models that can accurately predict crop phenotypes. It is necessary to develop optimized algorithms or models for data integration and dimensionality reduction and identify the most informative data for model construction.

(5) AI models in IPDB need to maintain their generative capabilities to support information exchange platforms in the future, which requires more computing resources and massive data accumulation.

The above challenges may hinder the large‐scale implementation of IPDB in crop breeding. To overcome these challenges, researchers and breeders should collaborate to share breeding resources, develop innovative methods for rapid and comprehensive gene function analysis, and explore the regulatory network underlying agronomic traits. Moreover, we may need to develop new methods or tools for low‐cost genotyping and phenotyping, continuously optimize model performances through enhanced data training and loop iteration, or develop smaller and more efficient models (Shah et al., 2024). Importantly, IPDB requires policy and financial support to integrate resources and improve computing power, as well as publicity through multiple channels to promote the large‐scale application in crop breeding. In short, IPDB is very promising in meeting global food demands, but there is still a long way to go to before it fully realizes its potential.

CONFLICTS OF INTEREST

The authors declare no conflicts of interest.

AUTHOR CONTRIBUTIONS

W.Z., H.Z. and W.L. wrote the manuscript; L.L. supervised the content; and L.L. and H.Z. revised the paper. All authors read and approved of the contents of this paper.

ACKNOWLEDGEMENTS

The authors really appreciate the useful and inspiring discussion with Dr. Liang Li from the Institute of Crop Sciences, Chinese Academy of Agricultural Sciences (CAAS) and Dr. Xingming Fan from the Institute of Food Crops, Yunnan Academy of Agricultural Sciences. We also thank the high‐performance computing platform at the National Key Laboratory of Crop Genetic Improvement at Huazhong Agricultural University. The authors apologize to those whose work they were unable to cite owing to space constraints. This work was supported by the National Key Research and Development Program of China (2023YFF1000100) (to L.L. and W.L.) and the National Natural Science Foundation of China (32321005) (to L.L.).

Biographies

graphic file with name JIPB-67-722-g006.gif

graphic file with name JIPB-67-722-g003.gif

Zhu, W. , Li, W. , Zhang, H. , and Li, L. (2025). Big data and artificial intelligence‐aided crop breeding: Progress and prospects. J. Integr. Plant Biol. 67: 722–739.

Edited by: Zhizhong Gong, China Agricultural University, China

Contributor Information

Hongwei Zhang, Email: zhanghongwei@caas.cn.

Lin Li, Email: hzaulilin@mail.hzau.edu.cn.

REFERENCES

  1. Abnizova, I. , Stapel, C. , te Boekhorst, R. , Lee, J.T.H. , and Hemberg, M. (2024). Integrative analysis of transcriptomic and epigenomic data reveals distinct patterns for developmental and housekeeping gene regulation. BMC Biol. 22: 78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Abramson, J. , Adler, J. , Dunger, J. , Evans, R. , Green, T. , Pritzel, A. , Ronneberger, O. , Willmore, L. , Ballard, A.J. , Bambrick, J. , et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630: 493–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Alemu, A. , Åstrand, J. , Montesinos‐López, O.A. , Sánchez, J.I.Y. , Fernández‐González, J. , Tadesse, W. , Vetukuri, R.R. , Carlsson, A.S. , Ceplitis, A. , Crossa, J. , et al. (2024). Genomic selection in plant breeding: Key factors shaping two decades of progress. Mol. Plant 17: 552–578. [DOI] [PubMed] [Google Scholar]
  4. Alipanahi, B. , Delong, A. , Weirauch, M.T. , and Frey, B.J. (2015). Predicting the sequence specificities of DNA‐ and RNA‐binding proteins by deep learning. Nat. Biotechnol. 33: 831–838. [DOI] [PubMed] [Google Scholar]
  5. Anilkumar, C. , Azharudheen, T.P.M. , Sah, R.P. , Sunitha, N.C. , Devanna, B.N. , Marndi, B.C. , and Patra, B.C. (2023). Gene based markers improve precision of genome‐wide association studies and accuracy of genomic predictions in rice breeding. Heredity 130: 335–345. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Arrones, A. , Vilanova, S. , Plazas, M. , Mangino, G. , Pascual, L. , Díez, M.J. , Prohens, J. , and Gramazio, P. (2020). The dawn of the age of multi‐parent MAGIC populations in plant breeding: Novel powerful next‐generation resources for genetic analysis and selection of recombinant elite material. Biology (Basel) 9: 229. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Azodi, C.B. , Pardo, J. , VanBuren, R. , de los Campos, G. , and Shiu, S.H. (2020). Transcriptome‐based prediction of complex traits in maize. Plant Cell 32: 139–151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Bai, X.L. , Fang, H. , He, Y. , Zhang, J.N. , Tao, M.Z. , Wu, Q.G. , Yang, G.F. , Wei, Y.Z. , Tang, Y. , Tang, L. , et al. (2023). Dynamic UAV phenotyping for rice disease resistance analysis based on multisource data. Plant Phenomics 5: 0019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Battenfield, S.D. , Guzmán, C. , Gaynor, R.C. , Singh, R.P. , Peña, R.J. , Dreisigacker, S. , Fritz, A.K. , and Poland, J.A. (2016). Genomic selection for processing and end‐use quality traits in the CIMMYT spring bread wheat breeding program. Plant Genome 9: 2. [DOI] [PubMed] [Google Scholar]
  10. Ben Hassen, M. , Cao, T.V. , Bartholomé, J. , Orasen, G. , Colombi, C. , Rakotomalala, J. , Razafinimpiasa, L. , Bertone, C. , Biselli, C. , Volante, A. , et al. (2018). Rice diversity panel provides accurate genomic predictions for complex traits in the progenies of biparental crosses involving members of the panel. Theor. Appl. Genet. 131: 417–435. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Bi, Y. , Yassue, R.M. , Paul, P. , Dhatt, B.K. , Sandhu, J. , Do, P.T. , Walia, H. , Obata, T. , and Morota, G. (2023). Evaluating metabolic and genomic data for predicting grain traits under high night temperature stress in rice. G3‐Genes Genom. Genet. 13: jkad052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Biazzi, E. , Nazzicari, N. , Pecetti, L. , Brummer, E.C. , Palmonari, A. , Tava, A. , and Annicchiarico, P. (2017). Genome‐wide association mapping and genomic selection for Alfalfa (Medicago sativa) forage quality traits. PLoS One 12: e0169234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Chen, J.X. , Tan, C. , Zhu, M. , Zhang, C.Y. , Wang, Z.H. , Ni, X.M. , Liu, Y.L. , Wei, T. , Wei, X.F. , Fang, X.D. , et al. (2023). CropGS‐Hub: A comprehensive database of genotype and phenotype resources for genomic prediction in major crops. Nucleic Acids Res. 52: D1519–D1529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Chen, M.J. , Fan, W.J. , Ji, F.Y. , Hua, H. , Liu, J. , Yan, M.X. , Ma, Q.G. , Fan, J.J. , Wang, Q. , Zhang, S.F. , et al. (2021). Genome‐wide identification of agronomically important genes in outcrossing crops using OutcrossSeq. Mol. Plant 14: 556–570. [DOI] [PubMed] [Google Scholar]
  15. Cheng, Z.K. , Gu, X.B. , Du, Y.D. , Wei, C.Y. , Xu, Y. , Zhou, Z.H. , Li, W.L. , and Cai, W.J. (2024). Multi‐modal fusion and multi‐task deep learning for monitoring the growth of film‐mulched winter wheat. Precis. Agric. 2: 1–25. [Google Scholar]
  16. Clark, N.M. , Nolan, T.M. , Wang, P. , Song, G.Y. , Montes, C. , Valentine, C.T. , Guo, H.Q. , Sozzani, R. , Yin, Y.H. , and Walley, J.W. (2021). Integrated omics networks reveal the temporal signaling events of brassinosteroid response in. Nat. Commun. 12: 5858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Costello, Z. , and Garcia Martin, H. (2019). How to hallucinate functional proteins. arXiv: 1903.00458.
  18. Crow, J.F. (1998). 90 years ago: The beginning of hybrid maize. Genetics 148: 923–928. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Dan, Z. , Chen, Y. , Zhao, W. , Wang, Q. , and Huang, W. (2019). Metabolome‐based prediction of yield heterosis contributes to the breeding of elite rice. Life Sci. Alliance 3: e201900551. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Darrah, L.L. , and Zuber, M.S. (1986). 1985 United States farm maize germplasm base and commercial breeding strategies 1. Crop Sci. 26: 1109–1113. [Google Scholar]
  21. Deomano, E. , Jackson, P. , Wei, X.M. , Aitken, K. , Kota, R. , and Pérez‐Rodríguez, P. (2020). Genomic prediction of sugar content and cane yield in sugar cane clones in different stages of selection in a breeding program, with and without pedigree information. Mol. Breed. 40: 1–12. [Google Scholar]
  22. Dhakshayani, J. , and Surendiran, B. (2023). M2F‐Net: A deep learning‐based multimodal classification with high‐throughput phenotyping for identification of overabundance of fertilizers. Agriculture (Basel) 13: 1238. [Google Scholar]
  23. van Dijk, A.D.J. , Kootstra, G. , Kruijer, W. , and de Ridder, D. (2021). Machine learning in plant science and plant breeding. Iscience 24: 101890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Dong, H.R. , Huang, Y. , and Wang, K.J. (2021). The development of herbicide resistance crop plants using CRISPR/Cas9‐mediated gene editing. Genes (Basel) 12: 912. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Dwivedi, S.L. , Scheben, A. , Edwards, D. , Spillane, C. , and Ortiz, R. (2017). Assessing and exploiting functional diversity in germplasm pools to enhance abiotic stress adaptation and yield in cereals and food legumes. Front. Plant Sci. 8: 1461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Farooq, M.A. , Gao, S. , Hassan, M.A. , Huang, Z. , Rasheed, A. , Hearne, S. , Prasanna, B. , Li, X. , and Li, H. (2024). Artificial intelligence in plant breeding. Trends Genet. 40: 891–908. [DOI] [PubMed] [Google Scholar]
  27. Fox, R.J. , Davis, S.C. , Mundorff, E.C. , Newman, L.M. , Gavrilovic, V. , Ma, S.K. , Chung, L.M. , Ching, C. , Tam, S. , Muley, S. , et al. (2007). Improving catalytic function by ProSAR‐driven enzyme evolution. Nat. Biotechnol. 25: 338–344. [DOI] [PubMed] [Google Scholar]
  28. Fritsche‐Neto, R. , Galli, G. , Borges, K.L.R. , Costa‐Neto, G. , Alves, F.C. , Sabadin, F. , Lyra, D.H. , Morais, P.P.P. , de Andrade, L.R.B. , Granato, I. , et al. (2021). Optimizing genomic‐enabled prediction in small‐scale maize hybrid breeding programs: A roadmap review. Front. Plant Sci. 12: 658267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Fu, J.J. , Falke, K.C. , Thiemann, A. , Schrag, T.A. , Melchinger, A.E. , Scholten, S. , and Frisch, M. (2012). Partial least squares regression, support vector machine regression, and transcriptome‐based distances for prediction of maize hybrid performance with gene expression data. Theor. Appl. Genet. 124: 825–833. [DOI] [PubMed] [Google Scholar]
  30. Gao, C.X. (2021). Genome engineering for crop improvement and future agriculture. Cell 184: 1621–1635. [DOI] [PubMed] [Google Scholar]
  31. Gao, P.F. , Zhao, H.N. , Luo, Z. , Lin, Y.F. , Feng, W.J. , Li, Y.L. , Kong, F.J. , Li, X. , Fang, C. , and Wang, X.T. (2023). SoyDNGP: A web‐accessible deep learning framework for genomic prediction in soybean breeding. Brief. Bioinform. 24: bbad349. [DOI] [PubMed] [Google Scholar]
  32. Gao, Y.J. , Zhou, Q. , Luo, J.X. , Xia, C. , Zhang, Y.H. , and Yue, Z.Y. (2024). Crop‐GPA: An integrated platform of crop gene‐phenotype associations. NPJ Syst. Biol. Appl. 10: 15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Gemmer, M.R. , Richter, C. , Jiang, Y. , Schmutzer, T. , Raorane, M.L. , Junker, B. , Pillen, K. , and Maurer, A. (2020). Can metabolic prediction be an alternative to genomic prediction in barley? PLoS One 15: e0234052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Gesteiro, N. , Ordás, B. , Butrón, A. , de la Fuente, M. , Jiménez‐Galindo, J.C. , Samayoa, L.F. , Cao, A.A. , and Malvar, R.A. (2023). Genomic versus phenotypic selection to improve corn borer resistance and grain yield in maize. Front. Plant Sci. 14: 1162440. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Gezan, S.A. , Osorio, L.F. , Verma, S. , and Whitaker, V.M. (2017). An experimental validation of genomic selection in octoploid strawberry. Hortic. Res. 4: 16070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Gianola, D. , Okut, H. , Weigel, K.A. , and Rosa, G.J.M. (2011). Predicting complex quantitative traits with Bayesian neural networks: A case study with Jersey cows and wheat. BMC Genet. 12: 87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. González‐Camacho, J.M. , de los Campos, G. , Pérez, P. , Gianola, D. , Cairns, J.E. , Mahuku, G. , Babu, R. , and Crossa, J. (2012). Genome‐enabled prediction of genetic values using radial basis function neural networks. Theor. Appl. Genet. 125: 759–771. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Gorjanc, G. , Bijma, P. , and Hickey, J.M. (2015). Reliability of pedigree‐based and genomic evaluations in selected populations. Genet. Sel. Evol. 47: 65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Greenside, P. , Shimko, T. , Fordyce, P. , and Kundaje, A. (2018). Discovering epistatic feature interactions from neural network models of regulatory DNA sequences. Bioinformatics 34: 629–637. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Guo, Z.F. , Yang, Q. , Huang, F.F. , Zheng, H.J. , Sang, Z.Q. , Xu, Y.F. , Zhang, C. , Wu, K.S. , Tao, J.J. , Prasanna, B.M. , et al. (2021). Development of high‐resolution multiple‐SNP arrays for genetic analyses and molecular breeding through genotyping by target sequencing and liquid chip. Plant Commun. 2: 100230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Hadasch, S. , Simko, I. , Hayes, R.J. , Ogutu, J.O. , and Piepho, H.P. (2016). Comparing the predictive abilities of phenotypic and marker‐assisted selection methods in a biparental lettuce population. Plant Genome 9: 1. [DOI] [PubMed] [Google Scholar]
  42. Han, L.Q. , Zhong, W.S. , Qian, J. , Jin, M.L. , Tian, P. , Zhu, W.C. , Zhang, H.W. , Sun, Y.H. , Feng, J.W. , Liu, X.G. , et al. (2023). A multi‐omics integrative network map of maize. Nat. Genet. 55: 144–153. [DOI] [PubMed] [Google Scholar]
  43. Harfouche, A.L. , Jacobson, D.A. , Kainer, D. , Romero, J.C. , Harfouche, A.H. , Mugnozza, G.S. , Moshelion, M. , Tuskan, G.A. , Keurentjes, J.J.B. , and Altman, A. (2019). Accelerating climate resilient plant breeding by applying next‐generation artificial intelligence. Trends Biotechnol. 37: 1217–1235. [DOI] [PubMed] [Google Scholar]
  44. Hashemifar, S. , Neyshabur, B. , Khan, A.A. , and Xu, J.B. (2018). Predicting protein‐protein interactions through sequence‐based deep learning. Bioinformatics 34: 802–810. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Hawkins‐Hooker, A. , Depardieu, F. , Baur, S. , Couairon, G. , Chen, A. , and Bikard, D. (2021). Generating functional protein variants with variational autoencoders. PLoS Comput. Biol. 17: e1008736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. He, T.H. , and Li, C.D. (2020). Harness the power of genomic selection and the potential of germplasm in crop breeding for global food security in the era with rapid climate change. Crop J. 8: 688–700. [Google Scholar]
  47. Hufford, M.B. , Seetharam, A.S. , Woodhouse, M.R. , Chougule, K.M. , Ou, S.J. , Liu, J.N. , Ricci, W.A. , Guo, T.T. , Olson, A. , Qiu, Y.J. , et al. (2021). De novo assembly, annotation, and comparative analysis of 26 diverse maize genomes. Science 373: 655–662. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Huynh‐Thu, V.A. , Irrthum, A. , Wehenkel, L. , and Geurts, P. (2010). Inferring regulatory networks from expression data using tree‐based methods. PLoS One 5: e12776. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Jannink, J.L. , Lorenz, A.J. , and Iwata, H. (2010). Genomic selection in plant breeding: From theory to practice. Brief Funct. Genomics 9: 166–177. [DOI] [PubMed] [Google Scholar]
  50. Jansen, R.C. , and Nap, J.P. (2001). Genetical genomics: The added value from segregation. Trends Genet. 17: 388–391. [DOI] [PubMed] [Google Scholar]
  51. Jiao, C.Z. , Hao, C.Y. , Li, T. , Bohra, A. , Wang, L.F. , Hou, J. , Liu, H.X. , Liu, H. , Zhao, J. , Wang, Y.M. , et al. (2023). Fast integration and accumulation of beneficial breeding alleles through an AB‐NAMIC strategy in wheat. Plant Commun. 4: 100549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Jumper, J. , Evans, R. , Pritzel, A. , Green, T. , Figurnov, M. , Ronneberger, O. , Tunyasuvunakool, K. , Bates, R. , Zídek, A. , Potapenko, A. , et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596: 583–589. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Kaddour, J. , Harris, J. , Mozes, M. , Bradley, H. , Raileanu, R. , and McHardy, R. (2023). Challenges and applications of large language models. arXiv: 2307.10169.
  54. Kaplan, J. , McCandlish, S. , Henighan, T. , Brown, T.B. , Chess, B. , Child, R. , Gray, S. , Radford, A. , Wu, J. , and Amodei, D. (2020). Scaling laws for neural language models. arXiv: 2001.08361.
  55. Kasela, S. , Aguet, F. , Kim‐Hellmuth, S. , Brown, B.C. , Nachun, D.C. , Tracy, R.P. , Durda, P. , Liu, Y.M. , Taylor, K.D. , Johnson, W.C. , et al. (2024). Interaction molecular QTL mapping discovers cellular and environmental modifiers of genetic regulatory effects. Am. J. Hum. Genet. 111: 133–149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Khaki, S. , and Wang, L.Z. (2019). Crop yield prediction using deep neural networks. Front. Plant Sci. 10: 621. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Kidane, Y.G. , Gesesse, C.A. , Hailemariam, B.N. , Desta, E.A. , Mengistu, D.K. , Fadda, C. , Pè, M.E. , and Dell'Acqua, M. (2019). A large nested association mapping population for breeding and quantitative trait locus mapping in Ethiopian durum wheat. Plant Biotechnol. J. 17: 1380–1393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Kolosov, N. , Daly, M.J. , and Artomov, M. (2021). Prioritization of disease genes from GWAS using ensemble‐based positive‐unlabeled learning. Eur. J. Hum. Genet. 29: 1527–1535. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Lanubile, A. , Maschietto, V. , Borrelli, V.M. , Stagnati, L. , Logrieco, A.F. , and Marocco, A. (2017). Molecular basis of resistance to fusarium ear rot in maize. Front. Plant Sci. 8: 1774. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Lenaerts, B. , Collard, B.C.Y. , and Demont, M. (2019). Review: Improving global food security through accelerated plant breeding. Plant Sci. 287: 110207. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Li, B.S. , Sun, C. , Li, J.Y. , and Gao, C.X. (2024a). Targeted genome‐modification tools and their advanced applications in crop breeding. Nat. Rev. Genet. 24: 1–20. [DOI] [PubMed] [Google Scholar]
  62. Li, C.S. , Xiang, X.L. , Huang, Y.C. , Zhou, Y. , An, D. , Dong, J.Q. , Zhao, C.X. , Liu, H.J. , Li, Y.B. , Wang, Q. , et al. (2020a). Long‐read sequencing reveals genomic structural variations that underlie creation of quality protein maize. Nat. Commun. 11: 17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Li, H. , Li, X. , Zhang, P. , Feng, Y. , Mi, J. , Gao, S. , Sheng, L. , Ali, M. , Yang, Z. , Li, L. , et al. (2024b). Smart Breeding Platform: A web‐based tool for high‐throughput population genetics, phenomics, and genomic selection. Mol. Plant 17: 677–681. [DOI] [PubMed] [Google Scholar]
  64. Li, J. , Zhang, D. , Yang, F. , Zhang, Q. , Pan, S. , Zhao, X. , Zhang, Q. , Han, Y. , Yang, J. , and Wang, K. (2024c). TrG2P: A transfer learning‐based tool integrating multi‐trait data for accurate prediction of crop yield. Plant Commun. 15: 100975. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Li, X.R. , Guo, T.T. , Bai, G.H. , Zhang, Z.W. , See, D. , Marshall, J. , Garland‐Campbell, K.A. , and Yu, J.M. (2022a). Genetics‐inspired data‐driven approaches explain and predict crop performance fluctuations attributed to changing climatic conditions. Mol. Plant 15: 203–206. [DOI] [PubMed] [Google Scholar]
  66. Li, Y. , Wang, S. , Umarov, R. , Xie, B.Q. , Fan, M. , Li, L.H. , and Gao, X. (2018). DEEPre: Sequence‐based enzyme EC number prediction by deep learning. Bioinformatics 34: 760–769. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Li, Z. , Chen, X.X. , Shi, S.Q. , Zhang, H.W. , Wang, X. , Chen, H. , Li, W.F. , and Li, L. (2022b). DeepBSA: A deep‐learning algorithm improves bulked segregant analysis for dissecting complex traits. Mol. Plant 15: 1418–1427. [DOI] [PubMed] [Google Scholar]
  68. Li, Z.H. , Wang, P.C. , You, C.Y. , Yu, J.W. , Zhang, X.N. , Yan, F.L. , Ye, Z.X. , Shen, C. , Li, B.Q. , Guo, K. , et al. (2020b). Combined GWAS and eQTL analysis uncovers a genetic regulatory network orchestrating the initiation of secondary cell wall development in cotton. New Phytol. 226: 1738–1752. [DOI] [PubMed] [Google Scholar]
  69. Lima, F.D.E. , Willmitzer, L. , and Nikoloski, Z. (2018). Classification‐driven framework to predict maize hybrid field performance from metabolic profiles of young parental roots. PLoS One 13: e0196038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Lin, F. , Fan, J. , and Rhee, S.Y. (2019). QTG‐Finder: A machine‐learning based algorithm to prioritize causal genes of quantitative trait loci in arabidopsis and rice. G3‐Genes Genom. Genet. 9: 3129–3138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Liu, H.J. , Wang, X.Q. , Xiao, Y.J. , Luo, J.Y. , Qiao, F. , Yang, W.Y. , Zhang, R.Y. , Meng, Y.J. , Sun, J.M. , Yan, S.J. , et al. (2020). CUBIC: An atlas of genetic architecture promises directed maize improvement. Genome Biol. 21: 20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Liu, H.J. , and Yan, J.B. (2019). Crop genome‐wide association study: A harvest of biological relevance. Plant J. 97: 8–18. [DOI] [PubMed] [Google Scholar]
  73. Liu, S.Y. , Xu, F. , Xu, Y.T. , Wang, Q. , Yan, J. , Wang, J.Y. , Wang, X.B. , and Wang, X.F. (2022). MODAS: Exploring maize germplasm with multi‐omics data association studies. Sci. Bull. 67: 903–906. [DOI] [PubMed] [Google Scholar]
  74. Liu, Y. , Wang, D.L. , He, F. , Wang, J.X. , Joshi, T. , and Xu, D. (2019). Phenotype prediction and genome‐wide association study using deep convolutional neural network of soybean. Front. Genet. 10: 1091. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Lou, Z.X. , Quan, L.Z. , Sun, D. , Xia, F.L. , Li, H.L. , and Guo, Z.M. (2024). Multimodal deep fusion model based on Transformer and multi‐layer residuals for assessing the competitiveness of weeds in farmland ecosystems. Int. J. Appl. Earth Obs. Geoinf. 127: 103681. [Google Scholar]
  76. Luo, J. (2015). Metabolite‐based genome‐wide association studies in plants. Curr. Opin. Plant Biol. 24: 31–38. [DOI] [PubMed] [Google Scholar]
  77. Lyra, D.H. , Mendonça, L.D. , Galli, G. , Alves, F.C. , Granato, I.S.C. , and Fritsche‐Neto, R. (2017). Multi‐trait genomic prediction for nitrogen response indices in tropical maize hybrids. Mol. Breed. 37: 80. [Google Scholar]
  78. Ma, C. , Xin, M.M. , Feldmann, K.A. , and Wang, X.F. (2014). Machine learning‐based differential network analysis: A study of stress‐responsive transcriptomes in Arabidopsis. Plant Cell 26: 520–537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Ma, W.L. , Qiu, Z.X. , Song, J. , Li, J.J. , Cheng, Q. , Zhai, J.J. , and Ma, C. (2018). A deep convolutional neural network approach for predicting phenotypes from genotypes. Planta 248: 1307–1318. [DOI] [PubMed] [Google Scholar]
  80. Ma, X.D. , Wang, H. , Wu, S.Y. , Han, B. , Cui, D. , Liu, J. , Zhang, Q. , Xia, X.Z. , Song, P. , Tang, C.F. , et al. (2024). DeepCCR: Large‐scale genomics‐based deep learning method for improving rice breeding. Plant Biotechnol. J. 22: 2691–2693. [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Maccaferri, M. , Sanguineti, M.C. , Donini, P. , and Tuberosa, R. (2003). Microsatellite analysis reveals a progressive widening of the genetic basis in the elite durum wheat germplasm. Theor. Appl. Genet. 107: 783–797. [DOI] [PubMed] [Google Scholar]
  82. Maqbool, M.A. , Beshir, A. , and Khokhar, E.S. (2020). Doubled haploids in maize: Development, deployment, and challenges. Crop Sci. 60: 2815–2840. [Google Scholar]
  83. Massman, J.M. , Jung, H.J.G. , and Bernardo, R. (2013). Genomewide selection versus marker‐assisted recurrent selection to improve grain yield and stover‐quality traits for cellulosic ethanol in maize. Crop Sci. 53: 58–66. [Google Scholar]
  84. Mebratu, A. , Wegary, D. , Teklewold, A. , and Tarekegne, A. (2024). Testcross performance and combining ability of early‐medium maturing quality protein maize inbred lines in Eastern and Southern Africa. Sci. Rep. 14: 9151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Mejía‐Guerra, M.K. , and Buckler, E.S. (2019). A k‐mer grammar analysis to uncover maize regulatory architecture. BMC Plant Biol. 19: 13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. Meuwissen, T.H.E. , Hayes, B.J. , and Goddard, M.E. (2001). Prediction of total genetic value using genome‐wide dense marker maps. Genetics 157: 1819–1829. [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. Mirabello, C. , and Wallner, B. (2018). rawMSA: End‐to‐end deep learning makes protein sequence profiles and feature extraction obsolete. biorxiv: 394437.
  88. Moerman, T. , Santos, S.A. , González‐Blas, C.B. , Simm, J. , Moreau, Y. , Aerts, J. , and Aerts, S. (2019). GRNBoost2 and Arboreto: Efficient and scalable inference of gene regulatory networks. Bioinformatics 35: 2159–2161. [DOI] [PubMed] [Google Scholar]
  89. Montesinos‐López, A. , Montesinos‐López, O.A. , Gianola, D. , Crossa, J. , and Hernández‐Suárez, C.M. (2018a). Multi‐environment genomic prediction of plant traits using deep learners with dense architecture. G3‐Genes Genom. Genet. 8: 3813–3828. [DOI] [PMC free article] [PubMed] [Google Scholar]
  90. Montesinos‐López, A. , Rivera, C. , Pinto, F. , Piñera, F. , Gonzalez, D. , Reynolds, M. , Pérez‐Rodríguez, P. , Li, H. , Montesinos‐López, O.A. , and Crossa, J. (2023). Multimodal deep learning methods enhance genomic prediction of wheat breeding. G3‐Genes Genom. Genet. 13: jkad045. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. Montesinos‐López, O.A. , Martín‐Vallejo, J. , Crossa, J. , Gianola, D. , Hernández‐Suárez, C.M. , Montesinos‐López, A. , Juliana, P. , and Singh, R. (2019a). A Benchmarking between deep learning, support vector machine and Bayesian threshold best linear unbiased prediction for predicting ordinal traits in plant breeding. G3‐Genes Genom. Genet. 9: 601–618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  92. Montesinos‐López, O.A. , Montesinos‐López, A. , Crossa, J. , Gianola, D. , Hernández‐Suárez, C.M. , and Martín‐Vallejo, J. (2018b). Multi‐trait, multi‐environment deep learning modeling for genomic‐enabled prediction of plant traits. G3‐Genes Genom. Genet. 8: 3829–3840. [DOI] [PMC free article] [PubMed] [Google Scholar]
  93. Montesinos‐López, O.A. , Montesinos‐López, A. , Pérez‐Rodríguez, P. , Barrón‐López, J.A. , Martini, J.W.R. , Fajardo‐Flores, S.B. , Gaytan‐Lugo, L.S. , Santana‐Mancilla, P.C. , and Crossa, J. (2021). A review of deep learning applications for genomic selection. BMC Genomics 22: 19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  94. Montesinos‐López, O.A. , Montesinos‐López, A. , Tuberosa, R. , Maccaferri, M. , Sciara, G. , Ammar, K. , and Crossa, J. (2019b). Multi‐trait, multi‐environment genomic prediction of durum wheat with genomic best linear unbiased predictor and deep learning methods. Front. Plant Sci. 10: 1311. [DOI] [PMC free article] [PubMed] [Google Scholar]
  95. Ni, Y. , Aghamirzaie, D. , Elmarakeby, H. , Collakova, E. , Li, S. , Grene, R. , and Heath, L.S. (2016). A machine learning approach to predict gene regulatory networks in seed development in Arabidopsis. Front. Plant Sci. 7: 1936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  96. Nyine, M. , Uwimana, B. , Blavet, N. , Hribová, E. , Vanrespaille, H. , Batte, M. , Akech, V. , Brown, A. , Lorenzen, J. , Swennen, R. , et al. (2018). Genomic prediction in a multiploid crop: Genotype by environment interaction and allele dosage effects on predictive ability in banana. Plant Genome 11: 2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  97. Onogi, A. , Ideta, O. , Inoshita, Y. , Ebana, K. , Yoshioka, T. , Yamasaki, M. , and Iwata, H. (2015). Exploring the areas of applicability of whole‐genome prediction methods for Asian rice (Oryza sativa L.). Theor. Appl. Genet. 128: 41–53. [DOI] [PubMed] [Google Scholar]
  98. Paran, I. , and Michelmore, R.W. (1993). Development of reliable PCR‐based markers linked to downy mildew resistance genes in lettuce. Theor. Appl. Genet. 85: 985–993. [DOI] [PubMed] [Google Scholar]
  99. Park, S. , Kim, J.M. , Shin, W. , Han, S.W. , Jeon, M. , Jang, H.J. , Jang, I.S. , and Kang, J. (2018). BTNET: Boosted tree based gene regulatory network inference algorithm using time‐course measurement data. BMC Syst. Biol. 12: 20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  100. Pérez‐Rodríguez, P. , Flores‐Galarza, S. , Vaquera‐Huerta, H. , del Valle‐Paniagua, D.H. , Montesinos‐López, O.A. , and Crossa, J. (2020). Genome‐based prediction of Bayesian linear and non‐linear regression models for ordinal data. Plant Genome 13: e20021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  101. Pérez‐Rodríguez, P. , Gianola, D. , González‐Camacho, J.M. , Crossa, J. , Manes, Y. , and Dreisigacker, S. (2012). Comparison between linear and non‐parametric regression models for genome‐enabled prediction in wheat. G3‐Genes Genom. Genet. 2: 1595–1605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  102. Qin, Q. , and Feng, J.X. (2017). Imputation for transcription factor binding predictions based on deep learning. PLoS Comput. Biol. 13: e1005403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  103. Rachmatia, H. , Kusuma, W.A. , Hasibuan, L.S. (2017). Prediction of maize phenotype based on whole-genome single nucleotide polymorphisms using deep belief networks. Journal of Physics: Conference Series. 835: 012003. [Google Scholar]
  104. Ray, D.K. , Mueller, N.D. , West, P.C. , and Foley, J.A. (2013). Yield trends are insufficient to double global crop production by 2050. PLoS One 8: e66428. [DOI] [PMC free article] [PubMed] [Google Scholar]
  105. Ray, D.K. , Ramankutty, N. , Mueller, N.D. , West, P.C. , and Foley, J.A. (2012). Recent patterns of crop yield growth and stagnation. Nat. Commun. 3: 1293. [DOI] [PubMed] [Google Scholar]
  106. Ren, Y. , Wu, C. , Zhou, H. , Hu, X. , and Miao, Z. (2024). Dual‐extraction modeling: A multi‐modal deep‐learning architecture for phenotypic prediction and functional gene mining of complex traits. Plant Commun. 13: 101002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  107. Repecka, D. , Jauniskis, V. , Karpus, L. , Rembeza, E. , Rokaitis, I. , Zrimec, J. , Poviloniene, S. , Laurynenas, A. , Viknander, S. , Abuajwa, W. , et al. (2021). Expanding functional protein sequence spaces using generative adversarial networks. Nat. Mach. Intell. 3: 324–333. [Google Scholar]
  108. Sahu, P.K. , Sao, R. , Khute, I.K. , Baghel, S. , Patel, R.R.S. , Thada, A. , Parte, D. , Devi, Y.L. , Nair, S. , Kumar, V. (2023). Plant genetic resources: Conservation, evaluation and utilization in plant breeding. In: Raina, A. , Wani, M.R. , Laskar, R. A. , Tomlekova, N. , Khan, S. , eds. In Advanced Crop Improvement, Volume 2: Case Studies of Economically Important Crops (Cham: Springer; ), pp. 1–45. [Google Scholar]
  109. van de Sande, B. , Flerin, C. , Davie, K. , De Waegeneer, M. , Hulselmans, G. , Aibar, S. , Seurinck, R. , Saelens, W. , Cannoodt, R. , Rouchon, Q. , et al. (2020). A scalable SCENIC workflow for single‐cell gene regulatory network analysis. Nat. Protoc. 15: 2247–2276. [DOI] [PubMed] [Google Scholar]
  110. Sandhu, K. , Patil, S.S. , Pumphrey, M. , and Carter, A. (2021a). Multitrait machine‐ and deep‐learning models for genomic selection using spectral information in a wheat breeding program. Plant Genome 14: e20119. [DOI] [PubMed] [Google Scholar]
  111. Sandhu, K.S. , Mihalyov, P.D. , Lewien, M.J. , Pumphrey, M.O. , and Carter, A.H. (2021b). Combining genomic and phenomic information for predicting grain protein content and grain yield in spring wheat. Front. Plant Sci. 12: 613300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  112. Sasaki, A. , Ashikari, M. , Ueguchi‐Tanaka, M. , Itoh, H. , Nishimura, A. , Swapan, D. , Ishiyama, K. , Saito, T. , Kobayashi, M. , Khush, G.S. , et al. (2002). Green revolution: A mutant gibberellin‐synthesis gene in rice—New insight into the rice variant that helped to avert famine over thirty years ago. Nature 416: 701–702. [DOI] [PubMed] [Google Scholar]
  113. Scott, M.F. , Ladejobi, O. , Amer, S. , Bentley, A.R. , Biernaskie, J. , Boden, S.A. , Clark, M. , Dell'Acqua, M. , Dixon, L.E. , Filippi, C.V. , et al. (2020). Multi‐parent populations in crops: A toolbox integrating genomics and genetic /mapping with breeding. Heredity 125: 396–416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  114. Senapati, N. , Semenov, M.A. , Halford, N.G. , Hawkesford, M.J. , Asseng, S. , Cooper, M. , Ewert, F. , van Ittersum, M.K. , Martre, P. , Olesen, J.E. , et al. (2022). Global wheat production could benefit from closing the genetic yield gap. Nat. Food 3: 532–541. [DOI] [PubMed] [Google Scholar]
  115. Shah, J. , Bikshandi, G. , Zhang, Y. , Thakkar, V. , Ramani, P. , and Dao, T. (2024). Flashattention‐3: Fast and accurate attention with asynchrony and low‐precision. arXiv: 2407.08608.
  116. Shang, L. , Li, X. , He, H. , Yuan, Q. , Song, Y. , Wei, Z. , Lin, H. , Hu, M. , Zhao, F. , and Zhang, C. (2022). A super pan‐genomic landscape of rice. Cell Res. 32: 878–896. [DOI] [PMC free article] [PubMed] [Google Scholar]
  117. Sharma, S. , Partap, A. , Balaguer, M.A.d.L. , Malvar, S. , and Chandra, R. (2022). Deepg2p: Fusing multi‐modal data to improve crop production. arXiv:2211.05986.
  118. Shen, Z. , Shen, E. , Yang, K. , Fan, Z. , Zhu, Q.‐H. , Fan, L. , and Ye, C.‐Y. (2024). BreedingAIDB: A database integrating crop genome‐to‐phenotype paired data with machine learning tools applicable to breeding. Plant Commun. 3: 100894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  119. Spindel, J. , Begum, H. , Akdemir, D. , Virk, P. , Collard, B. , Redoña, E. , Atlin, G. , Jannink, J.L. , and McCouch, S.R. (2015). Genomic selection and association mapping in rice (Oryza sativa): Effect of trait genetic architecture, training population composition, marker number and statistical model on accuracy of rice genomic selection in elite, tropical rice breeding lines. PLoS Genet. 11: e1004982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  120. Tabashnik, B.E. , and Carrière, Y. (2017). Surge in insect resistance to transgenic crops and prospects for sustainability. Nat. Biotechnol. 35: 926–935. [DOI] [PubMed] [Google Scholar]
  121. Togninalli, M. , Wang, X. , Kucera, T. , Shrestha, S. , Juliana, P. , Mondal, S. , Pinto, F. , Govindan, V. , Crespo‐Herrera, L. , Huerta‐Espino, J. , et al. (2023). Multi‐modal deep learning improves grain yield prediction in wheat breeding by fusing genomics and phenomics. Bioinformatics 39: btad336. [DOI] [PMC free article] [PubMed] [Google Scholar]
  122. Tong, H. , and Nikoloski, Z. (2021). Machine learning approaches for crop improvement: Leveraging phenotypic and genotypic big data. J. Plant Physiol. 257: 153354. [DOI] [PubMed] [Google Scholar]
  123. Tran, N.H. , Zhang, X.L.L. , Xin, L. , Shan, B.Z. , and Li, M. (2017). De novo peptide sequencing by deep learning. Proc. Natl. Acad. Sci. U. S. A. 114: 8247–8252. [DOI] [PMC free article] [PubMed] [Google Scholar]
  124. Turner‐Hissong, S.D. , Mabry, M.E. , Beissinger, T.M. , Ross‐Ibarra, J. , and Pires, J.C. (2020). Evolutionary insights into plant breeding. Curr. Opin. Plant Biol. 54: 93–100. [DOI] [PubMed] [Google Scholar]
  125. Wallace, J.G. , Rodgers‐Melnick, E. , and Buckler, E.S. (2018). On the road to breeding 4.0: Unraveling the good, the bad, and the boring of crop quantitative genomics. Annu. Rev. Genet. 52: 421–444. [DOI] [PubMed] [Google Scholar]
  126. Wang, B.B. , Hou, M. , Shi, J.P. , Ku, L.X. , Song, W. , Li, C.H. , Ning, Q. , Li, X. , Li, C.Y. , Zhao, B.B. , et al. (2023a). De novo genome assembly and analyses of 12 founder inbred lines provide insights into maize heterosis. Nat. Genet. 55: 355. [DOI] [PubMed] [Google Scholar]
  127. Wang, C.S. , Tang, S.C. , Zhan, Q.L. , Hou, Q.Q. , Zhao, Y. , Zhao, Q. , Feng, Q. , Zhou, C.C. , Lyu, D.F. , Cui, L.L. , et al. (2019). Dissecting a heterotic gene through GradedPool‐Seq mapping informs a rice‐improvement strategy. Nat. Commun. 10: 2982. [DOI] [PMC free article] [PubMed] [Google Scholar]
  128. Wang, H. , Cimen, E. , Singh, N. , and Buckler, E. (2020). Deep learning for plant genomics and crop improvement. Curr. Opin. Plant Biol. 54: 34–41. [DOI] [PubMed] [Google Scholar]
  129. Wang, K.L. , Abid, M.A. , Rasheed, A. , Crossa, J. , Hearne, S. , and Li, H.H. (2023b). DNNGP, a deep neural network‐based method for genomic prediction using multi‐omics data in plants. Mol. Plant 16: 279–293. [DOI] [PubMed] [Google Scholar]
  130. Wang, M. , Tai, C. , Weinan, E. , and Wei, L.P. (2018a). DeFine: Deep convolutional neural networks accurately quantify intensities of transcription factor‐DNA binding and facilitate evaluation of functional non‐coding variants. Nucleic Acids Res. 46: e69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  131. Wang, X. , Han, L.Q. , Li, J. , Shang, X.Y. , Liu, Q. , Li, L. , and Zhang, H.W. (2023c). Next‐generation bulked segregant analysis for Breeding 4.0. Cell Rep. 42: 113039. [DOI] [PubMed] [Google Scholar]
  132. Wang, X. , Li, J. , Han, L.Q. , Liang, C.Y. , Li, J.X. , Shang, X.Y. , Miao, X.X. , Luo, Z. , Zhu, W.C. , Li, Z. , et al. (2023d). QTG‐Miner aids rapid dissection of the genetic base of tassel branch number in maize. Nat. Commun. 14: 5232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  133. Wang, X. , Xu, Y. , Hu, Z.L. , and Xu, C.W. (2018b). Genomic selection methods for crop improvement: Current status and prospects. Crop J. 6: 330–340. [Google Scholar]
  134. Washburn, J.D. , Mejia‐Guerra, M.K. , Ramstein, G. , Kremling, K.A. , Valluru, R. , Buckler, E.S. , and Wang, H. (2019). Evolutionarily informed deep learning methods for predicting relative transcript abundance from DNA sequence. Proc. Natl. Acad. Sci. U. S. A. 116: 5542–5549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  135. Wegenast, T. , Longin, C.F.H. , Utz, H.F. , Melchinger, A.E. , Maurer, H.P. , and Reif, J.C. (2008). Hybrid maize breeding with doubled haploids. IV. Number versus size of crosses and importance of parental selection in two‐stage selection for testcross performance. Theor. Appl. Genet. 117: 251–260. [DOI] [PubMed] [Google Scholar]
  136. Wei, X. , Qiu, J. , Yong, K.C. , Fan, J.J. , Zhang, Q. , Hua, H. , Liu, J. , Wang, Q. , Olsen, K.M. , Han, B. , et al. (2021). A quantitative genomics map of rice provides genetic insights and guides breeding. Nat. Genet. 53: 243–253. [DOI] [PubMed] [Google Scholar]
  137. Wolfe, M.D. , Del Carpio, D.P. , Alabi, O. , Ezenwaka, L.C. , Ikeogu, U.N. , Kayondo, I.S. , Lozano, R. , Okeke, U.G. , Ozimati, A.A. , Williams, E. , et al. (2017). Prospects for genomic selection in Cassava breeding. Plant Genome 10: 3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  138. Wu, C.L. , Zhang, Y.Y. , Ying, Z.W. , Li, L. , Wang, J. , Yu, H. , Zhang, M.C. , Feng, X.Z. , Wei, X.H. , and Xu, X.G. (2024a). A transformer‐based genomic prediction method fused with knowledge‐guided module. Brief. Bioinform. 25: bbad438. [DOI] [PMC free article] [PubMed] [Google Scholar]
  139. Wu, C.X. , Luo, J.Y. , and Xiao, Y.J. (2024b). Multi‐omics assists genomic prediction of maize yield with machine learning approaches. Mol. Breed. 44: 14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  140. Wu, L.M. , Han, L.Q. , Li, Q. , Wang, G.Y. , Zhang, H.W. , and Li, L. (2021). Using interactome big data to crack genetic mysteries and enhance future crop breeding. Mol. Plant 14: 77–94. [DOI] [PubMed] [Google Scholar]
  141. Xiao, Y.J. , Jiang, S.Q. , Cheng, Q. , Wang, X.Q. , Yan, J. , Zhang, R.Y. , Qiao, F. , Ma, C. , Luo, J.Y. , Li, W.Q. , et al. (2021). The genetic mechanism of heterosis utilization in maize improvement. Genome Biol. 22: 148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  142. Xu, S.Z. , Xu, Y. , Gong, L. , and Zhang, Q.F. (2016). Metabolomic prediction of yield in hybrid rice. Plant J. 88: 219–227. [DOI] [PubMed] [Google Scholar]
  143. Xu, Y. , Ma, K.X. , Zhao, Y. , Wang, X. , Zhou, K. , Yu, G.N. , Li, C. , Li, P.C. , Yang, Z.F. , Xu, C.W. , et al. (2021). Genomic selection: A breakthrough technology in rice breeding. Crop J. 9: 669–677. [Google Scholar]
  144. Xu, Y. , Wang, X. , Ding, X.W. , Zheng, X.F. , Yang, Z.F. , Xu, C.W. , and Hu, Z.L. (2018). Genomic selection of agronomic traits in hybrid rice using an NCII population. Rice 11: 32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  145. Yan, J. , and Wang, X.F. (2023). Machine learning bridges omics sciences and plant breeding. Trends Plant Sci. 28: 199–210. [DOI] [PubMed] [Google Scholar]
  146. Yan, J. , Xu, Y.T. , Cheng, Q. , Jiang, S.Q. , Wang, Q. , Xiao, Y.J. , Ma, C. , Yan, J.B. , and Wang, X.F. (2021). LightGBM: Accelerated genomically designed crop breeding through ensemble learning. Genome Biol. 22: 271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  147. Yang, W.Y. , Guo, T.T. , Luo, J.Y. , Zhang, R.Y. , Zhao, J.R. , Warburton, M.L. , Xiao, Y.J. , and Yan, J.B. (2022). Target‐oriented prioritization: targeted selection strategy by integrating organismal and molecular traits through predictive analytics in breeding. Genome Biol. 23: 80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  148. Yao, J. , Wu, Y. , Liu, J. , and Wang, H. (2024). Multimodal deep learning‐based drought monitoring research for winter wheat during critical growth stages. PLoS One 19: e0300746. [DOI] [PMC free article] [PubMed] [Google Scholar]
  149. Yu, Y. , Cheng, Q. , Wang, F. , Zhu, Y.L. , Shang, X.G. , Jones, A. , He, H.H. , and Song, Y.H. (2023). Crop/plant modeling supports plant breeding: I. Optimization of environmental factors in accelerating crop growth and development for speed breeding. Plant Phenomics 5: 0099. [DOI] [PMC free article] [PubMed] [Google Scholar]
  150. Zander, M. , Lewsey, M.G. , Clark, N.M. , Yin, L.L. , Bartlett, A. , Guzmán, J.P.S. , Hann, E. , Langford, A.E. , Jow, B. , Wise, A. , et al. (2020). Integrated multi‐omics framework of the plant response to jasmonic acid. Nat. Plants 6: 290–302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  151. Zebosi, B. , Ssengo, J. , Geadelmann, L.F. , Unger‐Wallace, E. , and Vollbrecht, E. (2024). An effective and safe maize seed chipping protocol using clipping pliers with applications in small‐scale genotyping and marker‐assisted breeding. bioRxiv: 2024.2004. 2001.587552. [DOI] [PMC free article] [PubMed]
  152. Zeng, D.L. , Tian, Z.X. , Rao, Y.C. , Dong, G.J. , Yang, Y.L. , Huang, L.C. , Leng, Y.J. , Xu, J. , Sun, C. , Zhang, G.H. , et al. (2017). Rational design of high‐yield and superior‐quality rice. Nat. Plants 3: 17031. [DOI] [PubMed] [Google Scholar]
  153. Zhang, L.K. , Duan, Y.F. , Zhang, Z.W. , Zhang, L. , Chen, S.M. , Cai, C.C. , Duan, S.G. , Zhang, K. , Li, G.C. , and Cheng, F. (2024). OcBSA: An NGS‐based bulk segregant analysis tool for outcross populations. Mol. Plant 17: 648–657. [DOI] [PubMed] [Google Scholar]
  154. Zhang, X.F. , Sallam, A. , Gao, L.L. , Kantarski, T. , Poland, J. , DeHaan, L.R. , Wyse, D.L. , and Anderson, J.A. (2016). Establishment and optimization of genomic selection to accelerate the domestication and improvement of intermediate wheatgrass. Plant Genome 9: 1. [DOI] [PubMed] [Google Scholar]
  155. Zhao, T. , Wu, H.Y. , Wang, X.T. , Zhao, Y.Y. , Wang, L.Y. , Pan, J.Y. , Mei, H. , Han, J. , Wang, S.Y. , Lu, K.N. , et al. (2023). Integration of eQTL and machine learning to dissect causal genes with pleiotropic effects in genetic regulation networks of seed cotton yield. Cell Rep. 42: 113111. [DOI] [PubMed] [Google Scholar]
  156. Zhou, J. , Theesfeld, C.L. , Yao, K. , Chen, K.M. , Wong, A.K. , and Troyanskaya, O.G. (2018). Deep learning sequence‐based ab initio prediction of variant effects on expression and disease risk. Nat. Genet. 50: 1171–1179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  157. Zhou, J. , and Troyanskaya, O.G. (2015). Predicting effects of noncoding variants with deep learning‐based sequence model. Nat. Methods 12: 931–934. [DOI] [PMC free article] [PubMed] [Google Scholar]
  158. Zhou, P. , Li, Z. , Magnusson, E. , Cano, F.G. , Crisp, P.A. , Noshay, J.M. , Grotewold, E. , Hirsch, C.N. , Briggs, S.P. , and Springer, N.M. (2020). Meta gene regulatory networks in maize highlight functionally relevant regulatory interactions. Plant Cell 32: 1377–1396. [DOI] [PMC free article] [PubMed] [Google Scholar]
  159. Zhou, P.H. , Tan, Y.F. , He, Y.Q. , Xu, C.G. , and Zhang, Q. (2003). Simultaneous improvement for four quality traits of Zhenshan 97, an elite parent of hybrid rice, by molecular marker‐assisted selection. Theor. Appl. Genet. 106: 326–331. [DOI] [PubMed] [Google Scholar]
  160. Zhu, W.C. , Han, R. , Shang, X.Y. , Zhou, T. , Liang, C.Y. , Qin, X.M. , Chen, H. , Feng, Z.W. , Zhang, H.W. , Fan, X.M. , et al. (2024). The CropGPT project: Call for a global, coordinated effort in precision design breeding driven by AI using biological big data. Mol. Plant 17: 215–218. [DOI] [PubMed] [Google Scholar]
  161. Zhu, W.C. , Miao, X.X. , Qian, J. , Chen, S.J. , Jin, Q.X. , Li, M.Z. , Han, L.Q. , Zhong, W.S. , Xie, D. , Shang, X.Y. , et al. (2023). A translatome‐transcriptome multi‐omics gene regulatory network reveals the complicated functional landscape of maize. Genome Biol. 24: 60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  162. Zhu, W.C. , Xu, J. , Chen, S.J. , Chen, J. , Liang, Y. , Zhang, C.J. , Li, Q. , Lai, J.S. , and Li, L. (2021). Large‐scale translatome profiling annotates the functional genome and reveals the key role of genic 3′ untranslated regions in translatomic variation in plants. Plant Commun. 2: 100181. [DOI] [PMC free article] [PubMed] [Google Scholar]
  163. Zhu, X.T. , Maurer, H.P. , Jenz, M. , Hahn, V. , Ruckelshausen, A. , Leiser, W.L. , and Würschum, T. (2022). The performance of phenomic selection depends on the genetic architecture of the target trait. Theor. Appl. Genet. 135: 653–665. [DOI] [PMC free article] [PubMed] [Google Scholar]
  164. Zingaretti, L.M. , Gezan, S.A. , Ferrao, L.F. , Osorio, L.F. , Monfort, A. , Muñoz, P.R. , Whitaker, V.M. , and Pérez‐Enciso, M. (2020). Exploring deep learning for complex trait genomic prediction in polyploid outcrossing species. Front. Plant Sci. 11: 25. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Journal of Integrative Plant Biology are provided here courtesy of Wiley

RESOURCES