Skip to main content
Nucleic Acids Research logoLink to Nucleic Acids Research
. 2025 Dec 3;54(D1):D1720–D1732. doi: 10.1093/nar/gkaf1260

Gramene 2025: expanded comparative genomics and pathway resources, integrated search, and pan-genome portals for crop research

Andrew Olson 1, Sunita Kumari 2, Xuehong Wei 3, Kapeel Chougule 4, Zhenyuan Lu 5, Marcela Karey Tello-Ruiz 6, Vivek Kumar 7, Peter Van Buren 8, Audra Olson 9, Catherine Kim 10, Janeen Braynen 11, Lifang Zhang 12, Sarah Dyer 13, Jorge Alvarez-Jarreta 14, Shradha Saraf 15, Bruno Contreras-Moreira 16, Guy Naamati 17, Christina Ernst 18, Irene Papatheodorou 19, Nancy George 20, Pankaj Jaiswal 21, Sushma Naithani 22, Parul Gupta 23, Justin Elser 24, Peter D’Eustachio 25, Sarah M Assmann 26, Ángel Ferrero-Serrano 27, Asher Pasha 28, Nicholas Provart 29, Nicholas Gladman 30,31, Doreen Ware 32,33,
PMCID: PMC12807754  PMID: 41335101

Abstract

Gramene (gramene.org) is a comprehensive reference database for comparative plant genomics and pathway analysis, integrating functional annotations, evidence-based curated pathways and their projections, and multi-omics datasets. Since our last report, Gramene has added crop-specific pan-genome portals for maize, sorghum, rice, and grapevine. These pan-genome portals host population-scale datasets and multiple assembled genomes per species, all anchored by shared reference genomes. Importantly, these portals now adopt standardized rsIDs for genetic variants, advancing FAIR data principles and enabling cross-database interoperability. The main site is now Gramene Plants, emphasizing its broad genome coverage. Release 69 features 233 reference genomes, curated pathways for 139 species, expression data from 1026 studies across 27 species, and genetic variation data mapped to 27 genomes from 19 species. Key updates to the integrated search functionality include embedded expression viewers from the Bio-Analytic Resource for Plant Biology and EMBL-EBI Expression Atlas, a literature-curated catalog of gene functions, and a new Germplasm tab linking accessions with loss-of-function alleles to seed repositories. These advances reinforce Gramene as a comprehensive platform for exploring plant genomic diversity, gene function, and evolutionary conservation across the Green Tree of Life and within key agricultural species.

Graphical Abstract

Graphical Abstract.

Graphical Abstract

Introduction

Plants possess remarkably complex genomes shaped by gene duplication, polyploidy, and transposable elements, features that underlie both adaptive evolution and targeted crop improvement. The widespread adoption of long-read sequencing and improved annotation workflows has led to a dramatic increase in high-quality reference genomes across the plant kingdom, with over 500 plant genomes published in 2024 alone, including 370 sequenced for the first time [1, 2]. Concurrently, the field has rapidly expanded sequencing-first assays that quantify gene expression, chromatin state and accessibility, DNA methylation, and three-dimensional chromatin organization, including single-cell and spatial assays [36]. These advances provide blueprints for understanding functional diversity within and across species. Yet, the scale and heterogeneity of these various omics data require robust platforms that integrate sequence, function, and evolutionary context into accessible and interoperable resources. Since 2001, the Gramene database (gramene.org) has served as a central knowledgebase for comparative plant genomics, supporting a broad community of researchers in genetics, breeding, systems biology, and evolutionary studies [7]. Over the site’s lifetime, Gramene publications and resources have been cited over 1500 times and accessed by users from over 100 countries, reflecting its global reach. Gramene integrates diverse genome-scale data and analysis tools, including the Ensembl Genome Browser for genome annotations and variation data, CLIMtools resources for querying environmental-variant associations, the Plant Reactome for pathways and regulatory interactions, and the EMBL-EBI Expression Atlas and Bio-Analytic Resource for Plant Biology (BAR) for gene expression across a range of tissues, growth and developmental phases, treatments, and cell types. [814].

A major focus of this update is the expansion of pan-genome support, including genome assembly and annotation pipelines. Pan-genome analysis provides a more complete view of intraspecific variation by comparing multiple whole-genome assemblies within a species, capturing structural variants and gene presence/absence variation that single reference genomes cannot fully represent. To support this, Gramene has launched dedicated pan-genome portals for maize with 56 maize genomes, including 25 NAM founder genomes [15]; rice with 28 rice genomes, including 16 Magic lines and numerous wild ancestors [1618]; grapevine with 29 grape reference genomes [19]; and sorghum with 85 sorghum cultivars [20], thus enabling exploration of structural variation, gene presence/absence, and gene family evolution. These pan-genome portals now incorporate standardized reference cluster SNP identifiers “rsIDs” as global identifiers for citing genetic variations to enhance reproducibility, cross-species interoperability, and FAIR (Findable, Accessible, Interoperable, Reusable) data stewardship [21].

In parallel, the main Gramene site has been rebranded as Gramene Plants, reflecting its expanded role as a comprehensive multi-species portal. The resource continues to add new and updated reference genome assemblies, genetic variation datasets, functional annotations, expression datasets, and reference pathways with their gene orthology-based pathway projections for additional species.

Beyond data integrations, the Gramene team actively contributes to the development and adoption of community standards through leadership roles in the AgBioData Consortium (www.agbiodata.org) [2226]. This consortium of agricultural biological databases works to consolidate standards and best practices for acquiring, displaying, and reusing genomic, genetic, and breeding data, including establishing working groups focused on pan-genomes, gene nomenclature, genetic variations, database interoperability, and metadata frameworks for emerging data types such as single-cell gene expression [2325]. Gramene invests in community capacity building to ensure that these standards and resources are broadly adopted through workshops, webinars, and student/postdoc training programs. Educational materials focus on FAIR data principles and provide training on biocuration, data sharing, and management best practices.

Together, these efforts position Gramene Plants and the crop pan-genome portals as both a comprehensive data portal and a community-driven standards hub, advancing scalability, translation, innovation, accessibility, and interoperability of plant genomic resources across the agricultural research ecosystem.

Materials and methods

This section describes the architecture, data integration workflows, and user interfaces that support Gramene Plants and the crop pan-genome portals. The resources combine multiple data types, including genome assemblies, annotations, genetic variation, gene expression, and pathways, within a unified technical framework. We outline the database infrastructure, ontology-based annotations, homolog inference, programmatic services, and search capabilities, as well as the specific approaches used for pan-genome integrations, rsID assignment, and pathway curation and projections.

Gramene database architecture

Gramene Plants and the crop pan-genome sites integrate genomes, genome annotations, variation data, pathways, and expression datasets through a unified, document-oriented framework (Fig. 1). This architecture follows a principle of polyglot persistence by employing multiple database technologies in parallel to store and index data in the formats best suited to their structure and use. Core resources include the Ensembl platform for genome assemblies, genomic variation, and comparative analyses; Plant Reactome for curated and projected pathways; and the EMBL-EBI Expression Atlas and BAR for gene expression. Data from these sources are imported into standardized MongoDB collections and Apache Solr search indexes, which in turn support programmatic access via APIs and interactive exploration through the Gramene Plants and crop-specific pan-genome portals.

Figure 1.

Figure 1.

Illustration of how input data and metadata are loaded into Gramene and crop-specific pan-genome portals and flow through web services to various user interfaces. First, biocurated data are loaded into specialized databases from Ensembl, Plant Reactome, EMBL-EBI Expression Atlas, and BAR. Each has its own dedicated web services and user interfaces. These primary data are integrated with various forms of metadata into MongoDB collections, indexed in Solr, and made accessible via web services to support the Gramene search interface and a range of integrated visualizations for our users. For Gramene, we use Drupal for the news feed and GitHub for release notes and user guides. In the case of SorghumBase, we have additional community information maintained in WordPress.

Ontology-based hierarchical annotations

Gramene and crop pan-genome portals incorporate structured annotations from controlled vocabularies, including Gene Ontology (GO) [27], Plant Ontology (PO), Trait Ontology (TO) [28], InterPro domains [29], Plant Reactome event hierarchies, and the NCBI taxonomy tree [30]. For each indexed gene, both its primary annotation terms and their ancestor terms (via is_a and part_of relationships) are included in the search index. This allows broad queries (e.g. “transcription factor activity”) to return genes annotated to more specific descendant terms.

To improve discoverability, a parallel Solr core of “search suggestions” is built from these vocabularies, as well as gene names, identifiers, and cross-references. Suggestions are grouped by category and include counts of matching genes per genome, enabling faceted and data-driven refinement of search queries.

Representative homologs

Since many gene models remain poorly annotated, functional inference often depends on homology to well-characterized genomes. Gramene integrates gene trees from Ensembl Compara [31] and applies a recursive scoring algorithm to assign representative homologs at internal nodes for search indexing. The scoring system prioritizes genes with curated names, functional descriptions, and experimental annotations supported by literature citations. At each binary node, the higher-scoring candidate is propagated upward, with scores adjusted by evolutionary distance to balance annotation quality against phylogenetic proximity. Representative homologs above the scoring threshold appear in search results with protein identity values, supporting functional inference from well-annotated homologs. These homology-based annotations and resources are accessible through both the web interfaces and programmatic APIs, facilitating integration into external computational workflows. The Plant Reactome, described later, uses this gene homology and additional in-house analyzed Inparanoid data for projecting reference pathway annotations.

Programmatic access

Programmatic access to Gramene Plants and crop pan-genome portals is provided through a set of RESTful APIs hosted at data.gramene.org. The search API is implemented as a Node.js service with Swagger-based documentation, providing unified access to MongoDB collections and Solr indexes. Available endpoints expose curated gene, pathway, expression, and variation datasets, as well as ontology metadata and genomic maps. Solr-based endpoints support indexed searching and auto-completion. Local services mimic EMBL-EBI APIs to ensure compatibility with the Ensembl browser and BLAST [32] interfaces (Supplementary Table S1). Building on these APIs, the user-facing search interface provides flexible tools for the discovery and refinement of results across genomes and datasets.

Search interface

The Gramene Plants and crop pan-genome portals are single-page React applications with Redux-based state management. The type-ahead search bar queries the/suggest endpoint, returning categorized filters that trigger/search requests. The interface displays the current set of filters, an interactive distribution of search results across indexed genomes, and a list of matching genes, with expandable details tabs. Embeddable visualization components from Plant Reactome, Expression Atlas, and BAR allow direct display of pathways and expression profiles. Users can iteratively refine search results by applying additional filters. When the default method for combining filters with “AND” does not suit a given query, they can be dynamically combined with Boolean logic to form more complex searches. For example, identifying C3H transcription factors requires selecting genes with the Znf_CCCH domain while excluding those containing Helicase or an RRM domain [33]. This query can be constructed through the interface by combining filters with NOT/OR/AND operators. The system dynamically constructs the corresponding Solr query, applying biologically precise searches beyond simple keywords.

In addition to the main multispecies portal, Gramene extends this infrastructure to crop pan-genome sites, enabling comparative exploration across populations.

Crop pan-genome infrastructure

Gramene’s crop-specific pan-genome portals leverage the same infrastructure as Gramene Plants. A shared set of anchor reference genomes is maintained across all sites. These anchor genomes host genetic variants, gene expression datasets, and curated functional annotations, and also serve as outgroups for comparative analyses, providing linkages among the different portals. Not all genes are indexed on a single Gramene portal, so it might be necessary to search Gramene Plants or a crop-specific pan-genome portal. To improve findability, we have included links below the matching search terms in the type-ahead search bar to allow users to easily perform the same query on one of our sister sites. Figure 2 shows the distribution of genes containing an NB-ARC domain on the maize pan-genome portal and four of the anchor genomes.

Figure 2.

Figure 2.

Taxonomic distribution of genes containing the NB-ARC domain on the maize pan-genome portal. This visualization shows the genomic distribution of these genes across their respective genomes. Each assembled chromosome is assigned a color from a palette of 10 hues, while unassembled regions are gray. The color saturation indicates the density of the search results.

Genomic variation and rsID assignment

A central feature of the crop pan-genome portals is their support for genomic variation data. Genomic variation data are stored in either Ensembl variation databases or as indexed VCF files. For the crop pan-genome portals, population study metadata are integrated into the database, while sample metadata (including standard germplasm identifiers and links to seed repositories) are stored in a dedicated MongoDB collection. In collaboration with Gramene, the European Variation Archive (EVA) has assigned all hosted genetic variants a permanent identifier, the rsID. EVA assigns rsIDs to variation data by matching an identical allele based on the conservation of its flanking sequence, ensuring continuity across reference genome builds. Stable identifiers improve database interoperability, reproducibility, and robust association of functional annotations and phenotypes with genetic variants.

Plant Reactome pathway curation and projection

Gramene curates biological pathways in Plant Reactome, linking genes and reactions to higher-level processes, and projects them across diverse plant species. Gene and pathway biocuration by biocurators involves critical review of the literature, re-analysis of omics data sets, harmonizing geneIDs, entity disambiguation, knowledge synthesis, ontology mapping, and illustration of reactions and pathways. When applicable, pathways are connected to other existing pathways. These manually curated reference pathways are projected to additional plant species based on gene-orthology relationships. Compara orthologs (from Ensembl Plants and Gramene Plants), along with a small set of orthologs not represented in Compara (from the in-house version of Inparanoid clustering of selected high-quality transcriptomes and genomes), are used to produce a series of “orthopair” files [34]. During the projection process, if at least one ortholog of the reference rice reaction is present in a species, the reaction and its parent pathway are included in the projection. Finally, APIs and Python scripts are used for quality checks, as previously described [9, 34, 35].

Results and discussion

Curation standards and FAIR access

The technology behind the Gramene infrastructure is only as valuable as the data it serves. Our biocurators ensure that genome assemblies, gene models, pathways, functional annotations, gene expression, and genetic variation datasets are accurate, consistent, and useful to the plant genomics community. Gramene adheres to FAIR principles [21] by adopting community standards, providing programmatic access to public datasets, and enabling bulk downloads of curated datasets across portals. Team members actively contribute to AgBioData leadership and working groups to shape best practices and promote standard adoption across plant biology and agriculture.

Outreach and community standards

To accelerate community uptake of standards for data and metadata, members of the Gramene team participate in and often chair AgBioData working groups. These efforts have resulted in recommendations that Gramene has implemented, including guidance on genotype/phenotype data standardization [23, 36] and updated perspectives on pan-genome tools and resources [37]. The team also meets regularly with its Scientific Advisory Board, whose feedback informs project priorities and ensures alignment with community needs. The team is also engaged with the NASA Plant Analysis Working Group, contributing to a meta-analysis of plant biology data in microgravity and space environments [38], and a decadal survey (2023–2032, Thriving in Space). We have also contributed to GO discussions [27].

Since the last report, the team has delivered numerous presentations and workshops at conferences, including the Annual International Biocuration Conference, Plant Biology conferences organized by the American Society of Plant Biologists, the annual International Plant and Animal Genome Conference, Advances in Genome Biology and Technology Agriculture, AgBioData consortium’s monthly webinar series, and NASA GeneLab’s Analysis Working Group Symposium. We have trained over 100 graduate students and postdoctorates in gene curation and annotations. We have trained 17 undergraduate students in the biocuration of genes and pathways. In the process, they learned to critically review gene structure, scientific literature, and developed knowledge synthesis skills. New training videos are available on the Gramene Plants site and specific pan-genome sites (https://gramene.org/videotutorials, https://sorghumbase.org/videos); we encourage user feedback and can arrange one-on-one sessions on request via our feedback form (https://www.gramene.org/feedback).

Genomes and analysis pipelines

Across the six releases since our last update, the number of genomes hosted in Gramene Plants has grown from 93 genomes (release 63) to 233 genomes (release 69), harmonized with select Ensembl Plants versions. In parallel, we have invested heavily in crop pan-genome resources for grapevine, maize, rice, and sorghum, for which the Gramene team performs end-to-end data management and analysis. Table 1 summarizes the content for Gramene Plants and the crop pan-genome portals. The set of standardized analysis pipelines starts with loading genome assemblies and annotations into Ensembl core MySQL databases. For each genome, InterProScan identifies protein domains and associated GO terms [39]. When gene sets lack a canonical transcript designation, we select a representative transcript per locus using TRaCE [40], balancing transcript/peptide length with domain evidence and, when available, expression support. Canonical protein-coding transcripts are then passed on to the Ensembl Compara pipeline to cluster gene families and reconstruct phylogenetic gene trees [31]. Whole-genome alignments and synteny are also generated for a subset of the genomes. Transposable elements are identified and annotated using ethylenediaminetetraacetic acid (v2.1.0) [41].

Table 1.

Gramene Plants and crop pan-genome site content overview

Site (URL) Latest release (date) Genomes Population data Expression studies
Gramene plants (gramene.org/) v69 (September 2025) 233 19 species 27 species
        6 species*
Grapevine (vitis.gramene.org) v4 (February 2024) 29 300K variants 17 lines 10 baseline
        23 differential
        1 study*
Maize (maize-pangenome.gramene.org) v5 (March 2025) 56 275M variants 1628 lines [4244] 33 baseline
        58 differential
        12 studies*
Rice (oryza.gramene.org) v8 (August 2024) 28 38M variants 3332 lines [45, 46] 15 baseline
        98 differential
        11 studies*
Sorghum (sorghumbase.org) v9 (January 2025) 88 12.8M variants 1628 mutant lines [4749] 11 baseline
        5 differential
        10 studies*
      101M variants 4147 lines [5054]  

*Expression studies available from the BAR Electronic Fluorescent Pictograph (eFP) Browser.

Curated gene function

Automatic GO annotations via InterProScan are complemented by manual curation of gene functions from long-running efforts at TAIR, RAP-DB, UniProt, and NCBI’s GeneRIF database [44, 5557]. These sources may also include TO or PO terms. At each release, we incorporate these annotations into Gramene and highlight them as representative homologs in the search interface. A link on the Gramene homepage triggers a search for these genes. In the search results, the Papers tab lists publications describing a gene’s function and provides a submission link for additional references to be submitted by the community. In release 69, 16 054 genes with curated references appear in 12 514 gene trees containing 2 096 093 of the 4 374 239 genes assembled into gene trees (48%). The species distribution in Supplementary Table S2 reflects curation priorities at NCBI [56] and might not capture the work of species-specific database curators. We intend to advocate for data exchange standards for literature and functional annotation among the AgBioData community and will work closely with collaborators at MaizeGDB [58] and other model organism databases to ensure we have updated functional annotations for all of the “anchor genomes.”

Plant Reactome biocurators extend gene functional annotation by linking evidence to reactions, pathways, molecular mechanisms and interactions, agronomic traits, mutant phenotypes, and allelic variation [34, 35]. Given extensive gene duplication and polyploidy in plant genomes, gene families can share the same GO annotations. This redundancy complicates genomic and transcriptomic analyses, limiting the accuracy of genotype–phenotype associations. We have developed methods to re-analyze publicly available gene expression and other high-throughput -omics data for improving gene functional annotations of large gene families [59, 60] and for synthesizing plant pathway knowledge [61, 62].

Genetic variation

Gramene Plants and Ensembl Plants host genetic variation data mapped by EVA to reference genomes registered with International Nucleotide Sequence Database Collection [63], currently totaling more than 355 million variants across 19 species. While published studies often target a common reference genome, interpretation can benefit from mapping variants to alternative genomes; however, cross-reference tracking across multiple reference genomes is challenging. To address this, the EVA variant remapping pipeline (https://github.com/EBIvariation/variant-remapping/blob/master/README.md) maps variants between genome assembly versions, and it has been applied to other closely related genomes using flanking sequences. In Gramene’s crop pan-genome portals, variants called on a species’ reference genome with rsIDs are projected to other related genomes within the crop site. Use of global stable identifiers enables the consistent cross-referencing of variants to gene, regulatory regions, and phenotypes, supporting provenance tracking, trait discovery, and interoperability within and across sites, including those from external resources.

We also import community-relevant population data sets into Gramene’s pan-genome variation databases, including sample metadata linked to germplasm repositories (Supplementary Table S3). All variants are annotated with Variant Effect Predictor (VEP) [64] to identify deleterious alleles. Accessions harboring predicted loss-of-function (LOF) alleles are indexed into the gene search, enabling (i) rapid identification of germplasm impacting genes of interest and (ii) queries for all genes with a LOF allele within a given accession or mutant line (Fig. 3).

Figure 3.

Figure 3.

The Germplasm tab on the Gramene Oryza pan-genome site (https://oryza.gramene.org?idList=Os08g0508800) displaying germplasm accessions harboring predicted LOF alleles within a lipoxygenase gene (Lox2, Os08g0508800). Seeds for some germplasm accessions are available at stock centers such as IRRI (https://www.irri.org) and GRIN (https://www.ars-grin.gov). Clicking the Search button will trigger a search for all genes with LOF alleles in the selected accession.

Gramene and the crop pan-genome sites host results from select GWAS [65] and QTL studies focused on crop phenotypes in collaboration with Sorghum OZ QTL Atlas [66]. Gramene also hosts a CLIMtools portal where users can explore environmental variables and their associations with genetic variation in geolocated Arabidopsis accessions and rice landraces [1214]. CLIMtools is a suite of R-shiny apps showing (i) the geographic distribution of sampled accessions and local climate descriptors, (ii) tables and visualizations of variants associated with a selected environmental variable, and (iii) a table of environmental variables associated with genetic variation within a given gene [12].

Gene expression data integration and visualization

Through partnerships with EMBL-EBI Expression Atlas (www.ebi.ac.uk/gxa) [11, 67, 68] and the BAR eFP Browsers (https://www.bar.utoronto.ca/) [69], Gramene curates, integrates, and provides visualizations of bulk expression data. Currently, 1026 studies spanning 26 plant species are sourced from the EBI Expression Atlas (125 baseline, 901 differential), with crop-specific subsets including maize (33/58), rice (15/98), grapevine (10/23), and sorghum (11/5) baseline/differential studies (Table 1). Expression Atlas employs stringent quality control measures (≥3 biological replicates per condition, contamination checks, metadata standards) for inclusion in the integrated RNA-seq Analysis Pipeline. These datasets are periodically reanalyzed when a new version of the genome annotation is available in Ensembl Plants. The BAR provides complementary functionality by accommodating expression studies that may not meet these stringent standards while maintaining scientific rigor, ensuring broader data accessibility and visualization capabilities through the eFP Browser interface.

Gramene’s embedded eFP Browsers cover six plant genomes: Arabidopsis (26 studies), maize (12), rice (11), sorghum (10), soybean (4), and grapevine (1). These enable intuitive exploration across anatomical tissues (roots, stems, leaves, panicles, seeds, etc.), developmental stages (seedlings to reproductive organs), and diverse abiotic and biotic stress environments (pathogen, chemical, temperature, drought, nutrient availability etc.).

Expression data are fully indexed for search and API access. In the user interface, the Expression tab offers three views: Paralogs, All Studies, and eFP Browser. Paralogs view compares target gene expression with predicted paralogs, identifying tissue-specific expression patterns and evolutionary conservation in specific baseline or differential expression studies (Fig. 4A). The All Studies view shows gene expression across tissues and developmental stages for all the baseline studies. The eFP Browser view renders organ-level pictographs with a linear color scale from zero (yellow) to the maximum expression (red) based on Transcripts Per Million (TPM) for rapid visual comparison across developmental stages and stress conditions (Fig. 4B).

Figure 4.

Figure 4.

Gene expression views in the Gramene Maize pan-genome site (https://maize-pangenome.gramene.org/?idList=Zm00001eb005920). (A) Heatmap of normalized gene expression across different tissues for the target gene (Lox9, Zm00001eb005920) (synonyms: GRMZM2G017616, Zm00001d027893, lipoxygenase9) and its 11 paralogous genes using the baseline study “RNA-seq of coding RNA from six different maize tissues (ear, embryo, endosperm, pollen, root, and tassel) from B73 strain” [70]. The interactive heatmap highlights the gene and tissue of interest when the cursor is hovered on the anatomogram or on the heatmap. For example, clicking on the heatmap of “Zm00001eb05450” under tassel tissue generates a popover with additional expression information (gene name, expression part, expression value, and number of biological replicates) and highlights the tissue in the anatomogram. (B) The eFP Browser tab renders eFP images to highlight gene expression across different plant structures that are derived from other studies. Here, the “Maize Kernel” study is selected to visualize the gene expression of the target gene (Lox9, Zm00001d027893) across different cellular compartments of the maize kernel [71].

The expression functionality facilitates the identification of tissue-specific expression patterns and gene redundancy, informing molecular characterization strategies.

Plant Reactome pathways

Plant Reactome, an open-source pathway knowledgebase, encompasses a wide range of biological processes, including plant metabolism, development, formation of plant structure, response to biotic and abiotic stimuli, hormone signaling, transport, and transcriptional networks, all within the context of the plant cell’s subcellular architecture [34]. It hosts 351 evidence-based, expert-curated, reference pathways from rice, and their gene-homology-based projections to 138 species (totaling >38 000 pathways, >1 118 000 reactions, and >263 000 proteins). Since the last NAR report [7] (Gramene Release 63), we added 45 new reference pathways and pathway projections for 41 additional species. Since our last Plant Reactome publication [62], we curated 10 new and 5 revised pathways, enhanced functional annotations for over 500 genes (from 114 publications), and added pathway projections for 10 new species. The pathway revision efforts focused on connecting metabolic pathways with developmental pathways, phenotypes, and agronomic traits. For example, the recently revised starch biosynthesis pathway illustrates the location of specific reactions (in amyloplasts, statoliths, and the cytosol) and their relationship with diverse functions of starch (e.g. gravity sensing, anther and seed development, and carbohydrate storage) in different plant parts [7275]. Recent biocuration of pathways associated with (i) abiotic stress response (e.g. gravity, drought, salinity, heat, submergence/flooding, photoperiod changes, nutrient deficiency), and (ii) plant growth and development (e.g. organ formation, flowering) are focused on linking genotype to phenotype to environment, and their association with traits and QTLs of agronomic importance. Figure 5 shows photoperiod-dependent regulation of florigen (Hd3a, RFT1) under short day and long day [7678] and illustrates the relationships of various genes associated with flowering-time QTLs.

Figure 5.

Figure 5.

Rice inflorescence development depends on the expression of the florigens Hd3a and RFT1, which is controlled by Hd1 (conserved across monocots and dicots) and Ehd1 (monocot-specific) in response to photoperiod. The expression of Hd1 and Ehd1 is controlled by many other transcription regulators (reactions highlighted in blue). (A) Under short-day conditions, Hd1 and Ehd1 positively regulate the expression of Hd3a and RFT1. (https://plantreactome.gramene.org/PathwayBrowser/#/R-OSA-8934108&PATH=R-OSA-9030769,R-OSA-9031669,R-OSA-8934088) (B) Under long-day conditions, RFT1 expression increases in the leaf phloem, and subsequently, the RFT1 protein moves from the leaves to the shoot apex, leading to the induction of flowering. Expression of Hd3a is suppressed under long days. Hd1, Ehd1, Hd3a, and RFT1 are encircled in red. (https://plantreactome.gramene.org/PathwayBrowser/#/R-OSA-8934036&PATH=R-OSA-9030769,R-OSA-9031669,R-OSA-8934088) [7678].

Researchers are using the Plant Reactome for pathway enrichment analysis to identify candidate genes associated with plant responses to pathogens, salinity, flooding, and low nitrogen [7982], and metabolome enrichment analysis, for example, to discover potential stress response biomarkers for early detection of stress signals [83].

Future directions

Gramene’s next phase addresses an aging infrastructure, rapidly increasing data within species, and the need to scale analyses across species. As Ensembl transitions to a new data model and tech stack (currently in beta) [8], we will update schemas, ETL pipelines, and APIs to deliver the new browser across Gramene Plants and the crop pan-genome portals in the next two years. Near-term priorities include enhancing existing crop pan-genome sites and expanding to additional crop communities as data availability and resources permit. In parallel, we will integrate CLIMtools more tightly with gene and variant pages, APIs, and downloads, and extend CLIMtools to one or two additional species [12, 13, 16]. Throughout, we will emphasize robust standards (including rsIDs), deeper incorporation of population and regulatory data, and streamlined access for analysis.

To improve functional and pathway annotations at scale, we will introduce curator-supervised, semi-automated literature screening and extraction for triage, identifier normalization, and relationship capture. Retrieval-augmented generation methods will be employed to minimize spurious inferences, with curators reviewing outputs and authors verifying extracted statements, when feasible (as demonstrated in WormBase [84]). Known challenges, including inconsistencies in gene nomenclature, ontology misalignment, and limited domain collections, will be addressed through the use of labeled training examples, refined guidelines, and harmonized terminology in collaboration with community partners. These methodological advances and best practices will be shared through AgBioData consortium working groups to benefit the broader agricultural database community and inform standardized approaches to automated curation across genomic resources.

We will publish analysis-ready data products, including versioned expression matrices, rsID-indexed variant tables, pathway membership and network exports, and batch download endpoints with clear provenance. Interoperability will be strengthened by extending rsID adoption across portals and tightening federation with external archives and germplasm repositories, enabling aggregation of functional evidence (e.g. VEP effects, predicted LOF variants, eQTLs) and enriched sample metadata to streamline cross-database queries and trait discovery. Finally, we will expand practical education materials (videos, tutorials, example notebooks) and collaborate with the agricultural community to adopt shared standards for pan-genomes, genetic variation, metadata frameworks, and emerging data types such as single-cell expression. Engagement will occur through AgBioData working groups, targeted workshops, and hackathons. We will continue to collaborate closely with peer resources (e.g. Ensembl [8], EVA [85], UniProt [57], TAIR [55], MaizeGDB [58], RAP-DB [44], Grapedia (COST innovators grant IG17111), grapegenomics.com, Expression Atlas [11], BAR [10], GRIN, Reactome [9]) and individual researchers to ensure timely data exchange, practical interoperability, and tools that meet the everyday needs of plant and agricultural scientists.

Supplementary Material

gkaf1260_Supplemental_File

Acknowledgements

We thank our user community and all data providers for making their datasets available for reuse within Gramene Plants and the crop pan-genome portals. We thank the European Variation Archive (EVA) for support with rsID assignment and variant remapping, and the USDA National Plant Germplasm System (GRIN https://www.ars-grin.gov), the International Rice Research Institute (IRRI https://www.irri.org), and community germplasm repositories for sample metadata and linkages. We acknowledge the AgBioData Consortium and the NASA GeneLab Plant Analysis Working Group for leadership on standards and data sharing. We thank the IT and Scientific Computing teams at CSHL, Oregon State University, and USDA-ARS SCINet for infrastructure and operational support, and colleagues at peer resources (e.g. UniProt [57], TAIR [86], RAP-DB [44], MaizeGDB [58], Grapedia, grapegenomics) and the many individual researchers who contributed annotations, feedback, and training materials. P.J. acknowledges the support while working at the National Science Foundation, USA.

Author contributions: Andrew Olson (Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources,Software, Supervision, Validation, Visualization, Writing—original draft, Writing—review & editing)(Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Resources [lead], Software [lead], Supervision [lead], Validation [lead], Visualization [lead], Writing—original draft [lead], Writing—review & editing [lead]), Xuehong Wei (Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Resources [lead], Software [lead], Validation [lead], Writing—original draft [equal], Writing—review & editing [equal]), Sunita Kumari (Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Resources [lead], Software [lead], Validation [lead], Writing—original draft [lead], Writing—review & editing [lead]), Kapeel Chougule (Conceptualization [lead], Data curation [lead], Investigation [lead], Methodology [lead], Resources [lead], Software [lead], Validation [lead], Writing—original draft [equal], Writing—review & editing [equal]), Zhenyuan Lu (Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Resources [lead], Software [lead], Validation [lead], Writing—review & editing [equal]), Marcela K. Tello-Ruiz (Data curation [equal], Formal analysis [equal], Investigation [equal], Methodology [equal], Resources [equal], Validation [equal], Writing—review & editing [equal]), Vivek Kumar (Data curation [equal], Formal analysis [equal], Software [equal], Validation [equal], Writing—review & editing [equal]), Peter Van Buren (Software [equal], Validation [equal], Data curation [supporting], Supervision, Writing—review & editing [equal]), Audra Olson (Data curation [equal], Investigation [equal], Resources [equal], Validation [equal], Writing—review & editing [equal]), Catherine Kim (Data curation [equal], Formal analysis [equal], Writing—review & editing [equal]), Janeen Braynen (Validation [equal], Writing—review & editing [equal]), Lifang Zhang (Validation [equal], Writing—review & editing [equal]), Sarah Dyer (Data curation [supporting], Formal analysis [supporting], Funding acquisition [supporting], Investigation [supporting], Methodology [supporting], Project administration [supporting], Software [supporting], Supervision [supporting], Writing—review & editing [supporting]), Jorge Alvarez-Jarreta (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Software [supporting], Writing—review & editing [supporting]), Shradha Saraf (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Software [supporting], Writing—review & editing [supporting]), Bruno Contreras-Moreira (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Software [supporting], Writing—review & editing [supporting]), Guy Naamati (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Software [supporting], Writing—review & editing [supporting]), Christina Ernst (Data curation [supporting], Formal analysis [supporting], Funding acquisition [supporting], Project administration [supporting], Supervision [supporting], Writing—review & editing [supporting]), Irene Papatheodorou (Data curation [supporting], Formal analysis [supporting], Funding acquisition [supporting], Supervision [supporting]), Nancy George (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Writing—review & editing [supporting]), Pankaj Jaiswal (Conceptualization [supporting], Data curation [supporting], Funding acquisition [supporting], Investigation [supporting], Methodology [supporting], Project administration [supporting], Resources [supporting], Supervision [supporting], Writing—review & editing [supporting]), Sushma Naithani (Data curation [supporting], Investigation [supporting], Methodology [supporting], Project administration [supporting], Supervision [supporting], Validation [supporting], Writing—original draft [supporting], Writing—review & editing [supporting]), Parul Gupta (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Resources [supporting], Validation [supporting], Writing—original draft [supporting], Writing—review & editing [supporting]), Justin Elser (Methodology [supporting], Resources [supporting], Software [supporting], Visualization [supporting]), Peter D’Eustachio (Data curation [supporting], Supervision [supporting], Writing—review & editing [supporting]), Sarah M. Assmann (Data curation [supporting], Formal analysis [supporting], Funding acquisition [supporting], Investigation [supporting], Methodology [supporting], Writing—review & editing [supporting]), Angel Ferrero-Serrano (Data curation [supporting], Formal analysis [supporting], Investigation [supporting], Methodology [supporting], Writing—review & editing [supporting]), Asher Pasha (Data curation [supporting], Software [supporting], Visualization [supporting]), Nicholas Provart (Supervision [supporting], Visualization [supporting], Writing—review & editing [supporting]), Nicholas Gladman (Funding acquisition [equal], Validation [equal], Writing—review & editing [lead]), Doreen Ware (Conceptualization [equal], Data curation [equal], Funding acquisition [equal], Investigation [equal], Methodology [equal], Project administration [equal], Resources [equal], Supervision [equal], Writing—original draft [equal], Writing—review & editing [lead]).

Contributor Information

Andrew Olson, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Sunita Kumari, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Xuehong Wei, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Kapeel Chougule, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Zhenyuan Lu, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Marcela Karey Tello-Ruiz, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Vivek Kumar, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Peter Van Buren, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Audra Olson, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Catherine Kim, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Janeen Braynen, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Lifang Zhang, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States.

Sarah Dyer, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.

Jorge Alvarez-Jarreta, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.

Shradha Saraf, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.

Bruno Contreras-Moreira, Department of Genetics and Plant Breeding, Estación Experimental de Aula Dei–Consejo Superior de Investigaciones Científicas, Zaragoza 50059, Spain.

Guy Naamati, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.

Christina Ernst, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom.

Irene Papatheodorou, Earlham Institute, Norwich Research Park, Norwich NR4 7UZ, United Kingdom.

Nancy George, Syngenta Crop Protection, Jealott’s Hill, Warfield, Bracknell, RG42 6EY, United Kingdom.

Pankaj Jaiswal, Department of Botany and Plant Pathology, Oregon State University, Corvallis, OR 97331, United States.

Sushma Naithani, Department of Botany and Plant Pathology, Oregon State University, Corvallis, OR 97331, United States.

Parul Gupta, Department of Botany and Plant Pathology, Oregon State University, Corvallis, OR 97331, United States.

Justin Elser, Department of Botany and Plant Pathology, Oregon State University, Corvallis, OR 97331, United States.

Peter D’Eustachio, Department of Biochemistry and Molecular Pharmacology, NYU Grossman School of Medicine, New York, NY 10016, United States.

Sarah M Assmann, Pennsylvania State University, 208 Mueller Laboratory, University Park, State College, PA 16802, United States.

Ángel Ferrero-Serrano, Pennsylvania State University, 208 Mueller Laboratory, University Park, State College, PA 16802, United States.

Asher Pasha, Department of Cell and Systems Biology/Centre for the Analysis of Genome Evolution and Function, University of Toronto, Toronto, ON M5S 3B2, Canada.

Nicholas Provart, Department of Cell and Systems Biology/Centre for the Analysis of Genome Evolution and Function, University of Toronto, Toronto, ON M5S 3B2, Canada.

Nicholas Gladman, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States; USDA-ARS Robert Holley Center, 538 Tower Road, Ithaca, NY 14853, United States.

Doreen Ware, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY 11724, United States; USDA-ARS Robert Holley Center, 538 Tower Road, Ithaca, NY 14853, United States.

Supplementary data

Supplementary data is available at NAR online.

Conflict of interest

None declared.

Funding

This work was supported by the United States Department of Agriculture grant USDA-ARS [8062-21000-051-000D to D.W., N.G.; 0201-88888-003-000D [SciNet]; 0201-88888-002-000D (SciNet)]; National Science Foundation, USA [2122358 to D.W.; 2122357 to S.A., A.F.-S.; 2029854 to P.J.]; National Aeronautics and Space Administration (NASA), USA [80NSSC22K0891 to S.N.; 80NSSC22K0855 to P.J.]; Defense Advanced Research Projects Agency (DARPA) [HR0011-23-9-0054 to P.J., S.N.]; US National Institutes of Health [S10OD028632-01 (CSHL Cluster); NHRGI U24 HG012198]; Natural Sciences and Engineering Research Council of Canada (NSERC) to N.J.P.; Wellcome Trust [grant numbers WT222155/Z/20/Z, 221401/Z/20/Z to C.E.]; European Molecular Biology Laboratory to C.E.; and Oregon State University. Funding to pay the Open Access publication charges for this article was provided by Cold Spring Harbor laboratory will pay for funding using funds from the USDA ARS 8062-21000-051-000D.

Data availability

All data described in this study are freely available. Interactive access is provided through the Gramene web interfaces, with programmatic access supported via public APIs at https://data.gramene.org. Bulk downloads of genomic and functional datasets are available from the Gramene Plants and crop-specific pan-genome FTP sites. All source code supporting these resources is openly maintained at https://github.com/warelab [87]. Plant Reactome data are available at https://plantreactome.gramene.org/download/currrent. We support the integration of the Plant Reactome API into third-party websites and provide programmatic access via a RESTful API to JSON data (https://plantreactome.gramene.org/ContentService). The opening pathway diagram can be set to a specific pathway or superpathway (e.g. the Gravitropism superpathway for the NASA Gene Lab).

References

  • 1. Pucker B, Irisarri I, de Vries Jet al. Plant genome sequence assembly in the era of long reads: progress, challenges and future directions. Quant Plant Bio. 2022;3:e5. 10.1017/qpb.2021.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Schwacke R, Bolger ME. Usadel B. PubPlant—a continuously updated online resource for sequenced and published plant genomes. Front Plant Sci. 2025;16:1603547. 10.3389/fpls.2025.1603547. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Yan F, Powell DR, Curtis DJet al. From reads to insight: a hitchhiker’s guide to ATAC-seq data analysis. Genome Biol. 2020;21:22. 10.1186/s13059-020-1929-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. ENCODE Project Consortium, Moore JE, Purcaro MJet al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature. 2020;583:699–710. 10.1038/s41586-020-2493-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. O’Malley RC, Huang S-SC, Song Let al. Cistrome and epicistrome features shape the regulatory DNA landscape. Cell. 2016;165:1280–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Pardo-Palacios FJ, Wang D, Reese Fet al. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification. Nat Methods. 2024;21:1349–63. 10.1038/s41592-024-02298-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Tello-Ruiz MK, Naithani S, Gupta Pet al. Gramene 2021: harnessing the power of comparative genomics and pathways for plant research. Nucleic Acids Res. 2021;49:D1452–63. 10.1093/nar/gkaa979. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Dyer SC, Austine-Orimoloye O, Azov AGet al. Ensembl 2025. Nucleic Acids Res. 2025;53:D948–57. 10.1093/nar/gkae1071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Gupta P, Naithani S, Preece Jet al. Plant Reactome and PubChem: the plant pathway and (Bio)chemical entity knowledgebases. In: Edwards D (ed.), Plant Bioinformatics: Methods and Protocols. New York, NY: Springer US; 2022, 511–25. [DOI] [PubMed] [Google Scholar]
  • 10. Sullivan A, Lombardo MN, Pasha Aet al. 20 years of the Bio-Analytic Resource for Plant Biology. Nucleic Acids Res. 2025;53:D1576–86. 10.1093/nar/gkae920. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Moreno P, Fexova S, George Net al. Expression Atlas update: gene and protein expression in multiple species. Nucleic Acids Res. 2022;50:D129–40. 10.1093/nar/gkab1030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Ferrero-Serrano Á, Sylvia MM, Forstmeier PCet al. Experimental demonstration and pan-structurome prediction of climate-associated riboSNitches in Arabidopsis. Genome Biol. 2022;23:101. 10.1186/s13059-022-02656-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Ferrero-Serrano Á, Chakravorty D, Kirven KJet al. Oryza CLIMtools: a genome-environment association resource reveals adaptive roles for heterotrimeric G proteins in the regulation of rice agronomic traits. Plant Commun. 2024;5:100813. 10.1016/j.xplc.2024.100813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Ferrero-Serrano Á, Assmann SM.. Phenotypic and genome-wide association with the local environment of Arabidopsis. Nat Ecol Evol. 2019;3:274–85. 10.1038/s41559-018-0754-5. [DOI] [PubMed] [Google Scholar]
  • 15. Hufford MB, Seetharam AS, Woodhouse MRet al. De novo assembly, annotation, and comparative analysis of 26 diverse maize genomes. Science. 2021;373:655–62. 10.1126/science.abg5289. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Wei S, Chougule K, Olson Aet al. GrameneOryza: a comprehensive resource for Oryza genomes, genetic variation, and functional data. Database (Oxford). 2025;2025:baaf021. 10.1093/database/baaf021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Zhou Y, Chebotarov D, Kudrna Det al. A platinum standard pan-genome resource that represents the population structure of Asian rice. Sci Data. 2020;7:113. 10.1038/s41597-020-0438-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Zhang J, Chen L-L, Xing Fet al. Extensive sequence divergence between the reference genomes of two elite indica rice varieties Zhenshan 97 and Minghui 63. Proc Natl Acad Sci USA. 2016;113:E5163–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Chougule K, Tello-Ruiz MK, Wei Set al. Pan genome resources for grapevine. Acta Hortic. 2024;1390:257–66. 10.17660/ActaHortic.2024.1390.31. [DOI] [Google Scholar]
  • 20. Gladman N, Olson A, Wei Set al. SorghumBase: a web-based portal for sorghum genetic information and community advancement. Planta. 2022;255:35. 10.1007/s00425-022-03821-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Wilkinson MD, Dumontier M, Aalbersberg IJJet al. The FAIR guiding principles for scientific data management and stewardship. Sci Data. 2016;3:160018. 10.1038/sdata.2016.18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. bay088Harper L, Campbell J, Cannon EKSet al. AgBioData consortium recommendations for sustainable genomics and genetics databases for agriculture. Database. 2018;2018:bay088. 10.1093/database/bay088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Deng CH, Naithani S, Kumari Set al. Genotype and phenotype data standardization, utilization and integration in the big data era for agricultural sciences. Database (Oxford). 2023;2023:baad088. 10.1093/database/baad088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Clarke JL, Cooper LD, Poelchau MFet al. Data sharing and ontology use among agricultural genetics, genomics, and breeding databases and resources of the AgBioData consortium. Database (Oxford). 2023;2023:1–32. 10.1093/database/baad076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Cannon EKS, Molik DC, Wright AJet al. Guidelines for gene and genome assembly nomenclature. Genetics. 2025;229:iyaf006. 10.1093/genetics/iyaf006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Marrano A, Cabugos L, Hafner Aet al. A teaching and training framework to promote findable, accessible, interoperable, and reusable data generation in agriculture. Database (Oxford). 2025;2025:baaf034. 10.1093/database/baaf034. [DOI] [Google Scholar]
  • 27. Gene Ontology Consortium, Aleksander SA, Balhoff Jet al. The Gene Ontology knowledgebase in 2023. Genetics. 2023;224:iyad031. 10.1093/genetics/iyad031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Cooper L, Elser J, Laporte M-Aet al. Planteome 2024 update: reference ontologies and knowledgebase for plant biology. Nucleic Acids Res. 2024;52:D1548–55. 10.1093/nar/gkad1028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Blum M, Andreeva A, Florentino LCet al. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res. 2025;53:D444–56. 10.1093/nar/gkae1082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Schoch CL, Ciufo S, Domrachev Met al. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database (Oxford). 2020;2020:baaa062. 10.1093/database/baaa062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Vilella AJ, Severin J, Ureta-Vidal Aet al. EnsemblCompara GeneTrees: complete, duplication-aware phylogenetic trees in vertebrates. Genome Res. 2009;19:327–35. 10.1101/gr.073585.107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Camacho C, Coulouris G, Avagyan Vet al. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10:421. 10.1186/1471-2105-10-421. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Tian F, Yang D-C, Meng Y-Qet al. PlantRegMap: charting functional regulatory maps in plants. Nucleic Acids Res. 2020;48:D1104–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Naithani S, Gupta P, Preece Jet al. Plant Reactome: a knowledgebase and resource for comparative pathway analysis. Nucleic Acids Res. 2020;48:D1093–103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Naithani S, Preece J, D’Eustachio Pet al. Plant Reactome: a resource for plant pathways and comparative analysis. Nucleic Acids Res. 2017;45:D1029–39. 10.1093/nar/gkw932. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Callwood J, Celebioglu B, Gladman Net al. The need for robust, FAIR phenomic databases supporting agricultural efficiency and resiliency. Sci Public Policy. 2025;scaf039. 10.1093/scipol/scaf039. [DOI] [Google Scholar]
  • 37. Naithani S, Deng CH, Sahu SKet al. Exploring pan-genomes: an overview of resources and tools for unraveling structure, function, and evolution of crop genes and genomes. Biomolecules. 2023;13:1403. 10.3390/biom13091403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Barker R, Kruse CPS, Johnson Cet al. Meta-analysis of the space flight and microgravity response of the Arabidopsis plant transcriptome. NPJ Microgravity. 2023;9:21. 10.1038/s41526-023-00247-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Jones P, Binns D, Chang H-Yet al. InterProScan 5: genome-scale protein function classification. Bioinformatics. 2014;30:1236–40. 10.1093/bioinformatics/btu031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Olson AJ, Ware D. Ranked choice voting for representative transcripts with TRaCE. Bioinformatics. 2021;38:261–4. 10.1093/bioinformatics/btab542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Ou S, Su W, Liao Yet al. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol. 2019;20:275. 10.1186/s13059-019-1905-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. 3,000 rice genomes project . The 3,000 rice genomes project. Gigascience. 2014;3:7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. GSOR Collection: USDA Mini-Core Collection : USDA ARS. https://www.ars.usda.gov/southeast-area/stuttgart-ar/dale-bumpers-national-rice-research-center/docs/mini-core-collection(26 September 2025, date last accessed). [Google Scholar]
  • 44. Sakai H, Lee SS, Tanaka Tet al. Rice annotation project database (RAP-DB): an integrative and interactive database for rice genomics. Plant Cell Physiol. 2013;54:e6. 10.1093/pcp/pcs183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Andorf CM, Ross-Ibarra J, Seetharam ASet al. A unified VCF dataset from nearly 1,500 diverse maize accessions and resources to explore the genomic landscape of maize. G3. 2025;15:jkae281. 10.1093/g3journal/jkae281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Chia J-M, Song C, Bradbury PJet al. Maize HapMap2 identifies extant variation from a genome in flux. Nat Genet. 2012;44:803–7. 10.1038/ng.2313. [DOI] [PubMed] [Google Scholar]
  • 47. Addo-Quaye C, Tuinstra M, Carraro Net al. Whole-genome sequence accuracy is improved by replication in a population of mutagenized sorghum. G3. 2018;8:1079–94. 10.1534/g3.117.300301. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Jiao Y, Burke J, Chopra Ret al. A sorghum mutant resource as an efficient platform for gene discovery in grasses. Plant Cell. 2016;28:1551–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Jiao Y, Nigam D, Barry Ket al. A large sequenced mutant library - valuable reverse genetic resource that covers 98% of sorghum genes. Plant J. 2024;117:1543–57. 10.1111/tpj.16582. [DOI] [PubMed] [Google Scholar]
  • 50. Lozano R, Gazave E, Dos Santos JPRet al. Comparative evolutionary genetics of deleterious load in sorghum and maize. Nat Plants. 2021;7:17–24. 10.1038/s41477-020-00834-5. [DOI] [PubMed] [Google Scholar]
  • 51. Boatwright JL, Sapkota S, Jin Het al. Sorghum association panel whole-genome sequencing establishes cornerstone resource for dissecting genomic diversity. Plant J. 2022;111:888–904. 10.1111/tpj.15853. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Kumar N, Boatwright JL, Boyles REet al. Identification of pleiotropic loci mediating structural and non-structural carbohydrate accumulation within the sorghum bioenergy association panel using high-throughput markers. Front Plant Sci. 2024;15:1356619. 10.3389/fpls.2024.1356619. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Lasky JR, Upadhyaya HD, Ramu Pet al. Genome-environment associations in sorghum landraces predict adaptive traits. Sci Adv. 2015;1:e1400218. 10.1126/sciadv.1400218. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. SGT . In: Mysite. https://www.globalsorghuminitiative.org(29 September 2025, date last accessed). [Google Scholar]
  • 55. Reiser L, Bakker E, Subramaniam Set al. The Arabidopsis information resource in 2024. Genetics. 2024;227:iyae027. 10.1093/genetics/iyae027. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Jimeno-Yepes AJ, Sticco JC, Mork JGet al. GeneRIF indexing: sentence selection based on machine learning. BMC Bioinformatics. 2013;14:171. 10.1186/1471-2105-14-171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. UniProt Consortium . UniProt: the universal protein knowledgebase in 2025. Nucleic Acids Res. 2025;53:D609–17. 10.1093/nar/gkae1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Woodhouse MR, Cannon EK, Portwood JLet al. A pan-genomic approach to genome databases using maize as a model system. BMC Plant Biol. 2021;21:385. 10.1186/s12870-021-03173-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Naithani S, Komath SS, Nonomura Aet al. Plant lectins and their many roles: carbohydrate-binding and beyond. J Plant Physiol. 2021;266:153531. 10.1016/j.jplph.2021.153531. [DOI] [PubMed] [Google Scholar]
  • 60. Naithani S, Dikeman D, Garg Pet al. Beyond Gene Ontology (GO): using biocuration approach to improve the gene nomenclature and functional annotation of rice S-domain kinase subfamily. PeerJ. 2021;9:e11052. 10.7717/peerj.11052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Naithani S, Mohanty B, Elser Jet al. Biocuration of a transcription factors network involved in submergence tolerance during seed germination and coleoptile elongation in rice (Oryza sativa). Plants. 2023;12:2146. 10.3390/plants12112146. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Gupta P, Elser J, Hooks Eet al. Plant Reactome knowledgebase: empowering plant pathway exploration and OMICS data analysis. Nucleic Acids Res. 2024;52:D1538–47. 10.1093/nar/gkad1052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Karsch-Mizrachi I, Arita M, Burdett Tet al. The international nucleotide sequence database collaboration (INSDC): enhancing global participation. Nucleic Acids Res. 2025;53:D62–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64. McLaren W, Gil L, Hunt SEet al. The ensembl Variant Effect Predictor. Genome Biol. 2016;17:122. 10.1186/s13059-016-0974-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Mural RV, Grzybowski M, Miao Cet al. Meta-analysis identifies pleiotropic loci controlling phenotypic trade-offs in sorghum. Genetics. 2021;218:iyab087. 10.1093/genetics/iyab087. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Mace E, Innes D, Hunt Cet al. The Sorghum QTL Atlas: a powerful tool for trait dissection, comparative genomics and crop improvement. Theor Appl Genet. 2019;132:751–66. 10.1007/s00122-018-3212-5. [DOI] [PubMed] [Google Scholar]
  • 67. George N, Fexova S, Fuentes AMet al. Expression Atlas update: insights from sequencing data at both bulk and single cell level. Nucleic Acids Res. 2024;52:D107–14. 10.1093/nar/gkad1021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68. Papatheodorou I, Moreno P, Manning Jet al. Expression Atlas update: from tissues to single cells. Nucleic Acids Res. 2020;48:D77–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Winter D, Vinegar B, Nahal Het al. “Electronic Fluorescent Pictograph” browser for exploring and analyzing large-scale biological data sets. PLoS One. 2007;2:e718. 10.1371/journal.pone.0000718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Wang B, Tseng E, Regulski Met al. Unveiling the complexity of the maize transcriptome by single-molecule long-read sequencing. Nat Commun. 2016;7:11708. 10.1038/ncomms11708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71. Doll NM, Just J, Brunaud Vet al. Transcriptomics at maize embryo/endosperm interfaces identifies a transcriptionally distinct endosperm subdomain adjacent to the embryo scutellum. Plant Cell. 2020;32:833–52. 10.1105/tpc.19.00756. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Sack FD. Plant gravity sensing. Int Rev Cytol. 1991;127:193–252. [DOI] [PubMed] [Google Scholar]
  • 73. Blancaflor EB, Fasano JM, Gilroy S.. Mapping the functional roles of cap cells in the response of Arabidopsis primary roots to gravity. Plant Physiol. 1998;116:213–22. 10.1104/pp.116.1.213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74. Hayashi M, Crofts N, Oitome NFet al. Analyses of starch biosynthetic protein complexes and starch properties from developing mutant rice seeds with minimal starch synthase activities. BMC Plant Biol. 2018;18:59. 10.1186/s12870-018-1270-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Qu A, Xu Y, Yu Xet al. Sporophytic control of anther development and male fertility by glucose-6-phosphate/phosphate translocator 1 (OsGPT1) in rice. Journal of Genetics and Genomics. 2021;48:695–705. 10.1016/j.jgg.2021.04.013. [DOI] [PubMed] [Google Scholar]
  • 76. Yano M, Katayose Y, Ashikari Met al. Hd1, a major photoperiod sensitivity quantitative trait locus in rice, is closely related to the Arabidopsis flowering time gene CONSTANS. Plant Cell. 2000;12:2473–83. 10.1105/tpc.12.12.2473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77. Doi K, Izawa T, Fuse Tet al. Ehd1, a B-type response regulator in rice, confers short-day promotion of flowering and controls FT-like gene expression independently of Hd1. Genes Dev. 2004;18:926–36. 10.1101/gad.1189604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78. Komiya R, Ikegami A, Tamaki Set al. Hd3a and RFT1 are essential for flowering in rice. Development. 2008;135:767–74. 10.1242/dev.008631. [DOI] [PubMed] [Google Scholar]
  • 79. Sun T, Ma N, Jiao Yet al. TaCAMTA4 negatively regulates H2O2-dependent wheat leaf rust resistance by activating catalase 1 expression. Plant Physiol. 2024;196:2078–88. 10.1093/plphys/kiae443. [DOI] [PubMed] [Google Scholar]
  • 80. Kruasuwan W, Lohmaneeratana K, Munnoch JTet al. Transcriptome landscapes of salt-susceptible rice cultivar IR29 associated with a plant growth promoting endophytic Streptomyces. Rice. 2023;16:6. 10.1186/s12284-023-00622-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81. Subudhi PK, Garcia RS, Coronejo Set al. Comparative transcriptomics of rice genotypes with contrasting responses to nitrogen stress reveals genes influencing nitrogen uptake through the regulation of root architecture. Int J Mol Sci. 2020;21:5759. 10.3390/ijms21165759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82. Murray SC, Verhoef A, Adak Aet al. Detecting novel plant pathogen threats to food system security by integrating the Plant Reactome and remote sensing. Curr Opin Plant Biol. 2025;83:102684. 10.1016/j.pbi.2024.102684. [DOI] [PubMed] [Google Scholar]
  • 83. Dussarrat T, Prigent S, Latorre Cet al. Predictive metabolomics of multiple Atacama plant species unveils a core set of generic metabolites for extreme climate resilience. New Phytol. 2022;234:1614–28. 10.1111/nph.18095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84. Arnaboldi V, Raciti D, Van Auken Ket al. Text mining meets community curation: a newly designed curation platform to improve author experience and participation at WormBase. Database (Oxford). 2020;2020:baaa006. 10.1093/database/baaa006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85. Cezard T, Cunningham F, Hunt SEet al. The European Variation Archive: a FAIR resource of genomic variation for all species. Nucleic Acids Res. 2022;50:D1216–20. 10.1093/nar/gkab960. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Berardini TZ, Reiser L, Li Det al. The Arabidopsis information resource: making and mining the “gold standard” annotated reference plant genome. Genesis. 2015;53:474–85. 10.1002/dvg.22877. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87. Olson A. warelab/gramene-meta: Gramene Zenodo. Zenodo; 2025. 10.5281/ZENODO.17362645. [DOI]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Citations

  1. Olson A. warelab/gramene-meta: Gramene Zenodo. Zenodo; 2025. 10.5281/ZENODO.17362645. [DOI]

Supplementary Materials

gkaf1260_Supplemental_File

Data Availability Statement

All data described in this study are freely available. Interactive access is provided through the Gramene web interfaces, with programmatic access supported via public APIs at https://data.gramene.org. Bulk downloads of genomic and functional datasets are available from the Gramene Plants and crop-specific pan-genome FTP sites. All source code supporting these resources is openly maintained at https://github.com/warelab [87]. Plant Reactome data are available at https://plantreactome.gramene.org/download/currrent. We support the integration of the Plant Reactome API into third-party websites and provide programmatic access via a RESTful API to JSON data (https://plantreactome.gramene.org/ContentService). The opening pathway diagram can be set to a specific pathway or superpathway (e.g. the Gravitropism superpathway for the NASA Gene Lab).


Articles from Nucleic Acids Research are provided here courtesy of Oxford University Press

RESOURCES