Abstract
RNA editing, a post-transcriptional modification in plant mitochondria and plastids, is essential for environmental adaptation and diverse physiological processes. Despite extensive identification of RNA editing factors, primarily including pentatricopeptide repeat (PPR), multiple organelle RNA editing factor (MORF), organelle RNA recognition motif-containing (ORRM), and organelle zinc finger (OZ) proteins, their evolutionary history remains poorly understood. Here, we perform kingdom-wide evolutionary analyses across 364 high-quality Archaeplastida genomes and find massive PPR gene expansions in early-diverging land plants, predominantly driven by dispersed duplication associated with retroposition. Furthermore, integrative analyses imply that DYW subgroup PPR genes have been horizontally transferred from plants to bdelloid rotifers. MORF proteins, accessory partners of PPRs, possess MORF hallmark domains structurally similar to protein-folding peptidase S8 propeptide/proteinase inhibitor I9 domains, suggesting a role in protein folding during RNA editing. Considering diverse domain compositions, we reclassify MORF, ORRM, and OZ proteins and uncover prevalent hallmark domain fusions. Together, these findings illuminate the kingdom-wide evolution of plant RNA editing machinery.
Keywords: RNA editing, PPR, MORF, ORRM, OZ, evolution
Introduction
RNA editing, a post-transcriptional modification primarily involving cytidine-to-uridine (C-to-U) and uridine-to-cytidine (U-to-C) alterations in mitochondria and plastids, serves as a pivotal layer of genetic regulation in plants (Castandet and Araya 2011; Hao et al. 2021). This process is essential for adaptation to diverse environments (Fujii and Small 2011; Small et al. 2020; Hu et al. 2023) and various physiological processes (Wang et al. 2019b), including development (Sun et al. 2018), growth (Li et al. 2025b), flowering (Shi et al. 2016b), fruit ripening (Yang et al. 2017b), and resistance to stresses (Yuan and Liu 2012; Luo et al. 2022). Editing factors, a group of nuclear-encoded proteins, are assembled into an editing complex to mediate RNA editing (Small et al. 2020). Given the remarkable diversity of editing factors in plants, elucidating their evolution is essential for understanding how the RNA editing machinery emerged and diversified (Takenaka et al. 2013; Sun et al. 2016; Knoop 2023; Wang and Takenaka 2025).
Decades of work have been made to identify RNA editing factors, primarily including pentatricopeptide repeat (PPR), multiple organelle RNA editing factor (MORF; also known as RIP or DAG-like protein), organelle RNA recognition motif-containing (ORRM), and organelle zinc (OZ) finger proteins (Sun et al. 2016; Luo et al. 2017), each characterized by hallmark domains or motifs that contribute distinctly to regulating RNA editing (Yan et al. 2018). Among them, PPR proteins constitute the largest family of RNA editing factors (Li et al. 2019), specifically recognizing RNA sequences through P/PLS motif tracts and catalyzing C-to-U and U-to-C editing via carboxyl-terminal (C-terminal) DYW and DW:KP domains, respectively (Ichinose and Sugita 2016; Wang and Tan 2025). Furthermore, PPR proteins are classified into two subfamilies, P and PLS (Yan et al. 2018). Specifically, those in the P subfamily play multiple roles in diverse molecular processes, including RNA editing (Leu et al. 2016; Zhao et al. 2025), as well as transcription (Barkan and Small 2014), RNA stabilization (Pfalz et al. 2009), RNA cleavage (Hattori et al. 2007), RNA splicing (Yang et al. 2023a), and translation (Prikryl et al. 2011), while those in the PLS subfamily are predominantly dedicated to RNA editing (Small et al. 2020). Moreover, the PLS subfamily can be further divided into PLS, E (E1 and E2), E+, and DYW subgroups (Wang and Tan 2025). The four subgroups contain P, L, and S motifs in common, and the latter three contain specific motifs, namely, E, E+, and DYW motifs, respectively (Wang and Tan 2025). Notably, PPR proteins in the DYW subgroup, sequentially comprising P, L, S, E, and DYW motifs, are thought to be canonical (Cheng et al. 2016). In contrast, PPR proteins in the PLS, E, and E+ subgroups are regarded as C-terminally truncated forms (Lurin et al. 2004; Cheng et al. 2016). PPRs' accessory proteins, including MORF, ORRM, and OZ, which are named after their MORF, RRM, Ran Binding Protein 2 (RanBP2) type Zinc finger (Znf) hallmark domains, respectively, coordinate RNA editing through interactions both with PPR proteins and among themselves (Sun et al. 2013, 2015; Shi et al. 2016a, 2016b; Yan et al. 2017; Small et al. 2020; Gipson et al. 2022; Yang et al. 2022; Wang et al. 2023, 2025c). Beyond their hallmark domains, these proteins also incorporate diverse additional domains, such as Glycine-Rich (GR) domain in MORF proteins (Takenaka et al. 2012; Haag et al. 2017; Li et al. 2025a) and PPR/MORF domains in ORRM proteins (Sugita et al. 2016).
Building on these advances, studies have increasingly focused on the evolutionary histories of PPR, MORF, ORRM, and OZ genes (Shikanai 2006; Sun et al. 2016; Small et al. 2020; Baudry et al. 2024; Wang and Takenaka 2025). The PPR gene family is one of the largest gene families in plants (Wang and Tan 2025), ranging from tens of members in early-diverging chlorophytes to hundreds or thousands in land plants (Fujii and Small 2011; Small et al. 2020) and thereby reflecting massive expansions (Fujii and Small 2011; Gutmann et al. 2020). PPR genes also occur beyond plants, with P subfamily members widespread across eukaryotes and occasionally detected in prokaryotes (Manna and Barth 2013; Manna 2015). Notably, DYW subgroup PPR genes are observed in multiple nonplant eukaryote lineages (Knoop and Rüdinger 2010; Manna et al. 2013; Schallenberg-Rüdinger et al. 2013; Schaap et al. 2015), including Amoebozoa, Discoba, Fungi, Metazoa, and Malawimonadidae, presumably resulting from horizontal gene transfer (HGT) (Fu et al. 2014). In contrast, MORF genes are present exclusively in angiosperms and gymnosperms (Gutmann et al. 2020); ORRM genes are broadly distributed across Viridiplantae lineages including chlorophytes, mosses, liverworts, and angiosperms (Sugita et al. 2016); and OZ genes are sporadically distributed in gymnosperms, mosses, and lycophytes, but are absent from chlorophytes (Sun et al. 2015). Despite these insights, a comprehensive evolutionary characterization of these RNA editing factors across Archaeplastida remains unexplored.
Toward this end, we collect a comprehensive list of genomes across Archaeplastida, systematically identify PPR, MORF, ORRM, and OZ genes and reconstruct an evolutionary landscape of plant RNA editing factors. Building on the landscape, we perform phylogenomic analyses to elucidate the mechanisms underlying PPR expansions and investigate the extent and evolutionary significance of HGT in DYW subgroup PPR genes. Through sequence and structural alignments, we explore the conservation and diversity of hallmark domains of MORF proteins. Finally, we reclassify MORF, ORRM, and OZ proteins based on their domain composition.
Results
Evolutionary landscape of PPR, MORF, ORRM, and OZ genes across Archaeplastida
To systematically investigate the evolution of PPR, MORF, ORRM, and OZ genes, we obtained a total of 364 high-quality genomes across Archaeplastida that span early-diverging aquatic algae to highly derived land plants (for details, see the Materials and methods section), including glaucophytes (n = 1), rhodophytes (n = 7), prasinodermophytes (n = 1), chlorophytes (n = 33), charophytes (n = 7), hornworts (n = 22), mosses (n = 13), liverworts (n = 4), lycophytes (n = 6), ferns (n = 5), gymnosperms (n = 13), and angiosperms (n = 262) (Table S1). Based on these genomes, we first identify 138 single-copy orthologous gene families (Table S2) and then build a plant phylogenomic tree (Fig. 1a and Fig. S1) that is topologically congruent with previous studies (Puttick et al. 2018; Wu et al. 2024). Next, we systematically identify 212,779 PPR, 2,759 MORF, 7,760 ORRM, and 829 OZ genes (Tables S3 and S4) and evolutionarily characterize these genes in the phylogenomic tree as built above. Noticeably, we find that PPR genes are scarcely present in glaucophytes, rhodophytes, prasinodermophytes, chlorophytes, and charophytes, ranging only from 8 to 182 (Fig. 1b and Table S4). In contrast, hundreds to thousands of PPR genes are detected in land plants, particularly certain hornworts, mosses, and lycophytes, followed by gymnosperms, ferns, and angiosperms. Remarkably, two lycophytes, namely, Selaginella moellendorffii and Isoetes sinensis, have the highest number of PPR genes (n = 3,861 and 3,256, respectively) among the 364 Archaeplastida species in our study (Fig. 1b and Table S4). On the contrary, liverworts are short of PPR genes, less than 100 in count (Table S4).
Figure 1.
The evolutionary landscape of PPR, MORF, ORRM, and OZ genes across Archaeplastida. a) The maximum-likelihood phylogenomic tree of 364 Archaeplastida species, including glaucophytes (n = 1), rhodophytes (n = 7), prasinodermophytes (n = 1), chlorophytes (n = 33), charophytes (n = 7), hornworts (n = 22), mosses (n = 13), liverworts (n = 4), lycophytes (n = 6), ferns (n = 5), gymnosperms (n = 13), and angiosperms (n = 262). The best-fit model of each single-copy orthologous gene tree is automatically selected according to Bayesian Information Criterion in IQ-TREE. All gene trees (n = 138) are merged and the final species tree is inferred by ASTRAL. Bootstraps are from 1,000 replicates. The distribution of b) PPR, c) MORF, d) ORRM, and e) OZ gene across Archaeplastida lineages. Two lycophytes, S. moellendorffii and I. sinensis, have the highest number of PPR genes (n = 3,861 and 3,256, respectively). Orange and purple bars indicate gene expansions occurring simultaneously across all four gene families, and within individual families, respectively. Filled and open black circles denote species that have undergone genome diploidization/polyploidization (D/P), and those with unknown genome expansion, respectively. D/P: Arachis hypogaea, Atropa belladonna, B. napus, C. rotundifolia, C. sativus, Helianthus tuberosus, Papaver somniferum, P. cablin, Pontederia crassipes, T. aestivum, Zanthoxylum bungeanum, and Zingiber officinale. Unknown D/P: C. pinnatifida, Iris pallida, and M. fusca.
Compared to PPR genes, the other three types of genes present distinct patterns. Specifically, MORF genes are found only in seed plants, including angiosperms and gymnosperms (Fig. 1c), in agreement with a previous study (Gutmann et al. 2020). In addition, seed plants generally harbor about ten MORF genes, whereas some species, such as Brassica napus, Pogostemon cablin, and Cissus rotundifolia, have more MORF genes (n = 40, 39, and 37, respectively; Table S4). Conversely, ORRM genes are ubiquitous across Archaeplastida and exhibit a gradual increase from glaucophytes to angiosperms, indicating their ancient origin (Fig. 1d). Albeit much fewer in number than the rest three types of genes, OZ genes are enriched in angiosperms, with a minor presence in chlorophytes, hornworts, mosses, liverworts, lycophytes, ferns, and gymnosperms (Fig. 1e). Besides, many angiosperms that have undergone diploidization or polyploidization, such as P. cablin (Shen et al. 2022), Crocus sativus (Xu et al. 2024), and Triticum aestivum (Wang et al. 2025b), exhibit concurrent expansions of PPR, MORF, ORRM, and OZ genes (Fig. 1 and Table S4). In contrast, among these four types of genes, MORF genes are the only ones markedly expanded in drought-resistant C. rotundifolia relative to other angiosperms, whereas OZ genes are the only ones significantly expanded in Crataegus pinnatifida and Malus fusca. Overall, PPR genes exhibit dramatic expansions in land plants, in contrast to MORF, ORRM, and OZ genes.
PPR gene expansions in early-diverging land plants
Gene expansion can be achieved by gene duplication with four different modes, viz., whole genome duplication (WGD), tandem duplication (TD), proximal duplication (PD), and dispersed duplication (DD), which are classified according to genomic loci of duplicated gene pairs (Freeling 2009; Panchy et al. 2016; Wang et al. 2022). To explore the underlying mechanisms of PPR gene expansions in early-diverging land plants, we select high-quality chromosome-level genomes, including four hornworts (Anthoceros agrestis, 2,308 PPRs; Notothylas orbicularis, 1,772 PPRs; Paraphymatoceros pearsonii, 1,583 PPRs; and Phymatoceros phymatodes, 1,627 PPRs), one moss (Takakia lepidozioides, 2,016 PPRs), and two lycophytes (I. sinensis, 3,256 PPRs; and Huperzia asiatica, 1,583 PPRs) (Table S4), with the charophyte Zygnema circumcarinatum (67 PPRs) serving as the phylogenetic outgroup. Furthermore, we select well-annotated PPR genes, accounting for 93.40% to 100% of our identified PPRs and detect that 90.04% to 95.89% of them form duplicated gene pairs based on sequence similarity analysis (Table S5). Meanwhile, we infer WGDs in these selected plants through comprehensive analyses of gene synteny and synonymous substitution rates (Ks) of homologous gene pairs (for details, see the Materials and methods section). Accordingly, we find that Z. circumcarinatum and four hornworts lack evidence of WGD, whereas T. lepidozioides and I. sinensis each experienced one WGD, and H. asiatica underwent two WGDs (Figs. S2 to S9). These results are consistent with previous studies (Li et al. 2020, 2024; Zhang et al. 2020; Hu et al. 2023; Feng et al. 2024; Schafran et al. 2025), except that I. sinensis has been reported to have experienced two WGDs based solely on Ks distributions (Cui et al. 2023). Consequently, the integration of inferred WGD events with comparison of genomic loci of duplicated PPR gene pairs (for details see the Materials and methods section) indicates that DD is primarily responsible for PPR gene expansions in the selected species, followed by WGD in I. sinensis and H. asiatica (Fig. 2a and Table S5).
Figure 2.
Mechanisms of PPR gene expansions. a) Numbers of duplicated PPR genes (original and copied) derived from WGD, TD, PD, and DD in seven representative species (A. agrestis, N. orbicularis, P. pearsonii, P. phymatodes, T. lepidozioides, I. sinensis, and H. asiatica) and the outgroup (Z.circumcarinatum). A red five-pointed star denotes a WGD event. WGD, whole genome duplication; TD, tandem duplication; PD, proximal duplication; DD, dispersed duplication. b) Percentage of intronless DD PPR gene pairs. Numbers to the left and right of the colon indicate intron counts in each DD PPR gene pair. Categories “0:0”, “0:n”, “1:n”, and “n:n (difference > 10)” denote intron loss in DD PPR gene pairs. c) E2 and d) DYW subgroup PPR gene pairs (DD) flanked by similar LTR/Copia and LTR/Gypsy, respectively. e) Percentages of intronless DD PPR gene pairs with polyA tails and flanked by LTRs.
Furthermore, it has been reported that DD is related to retroposition (Panchy et al. 2016; Wang et al. 2022), a process that the complementary DNA transcript of a gene is inserted back into the genome, thus producing a retrogene (Zhang et al. 2014). In general, the retrogene of a duplicated pair features intron loss and a polyadenylation (polyA) tail, and the pair is flanked by short direct repeats (e.g. transposable element, TE) (Zhang et al. 2014; Schrader and Schmitz 2019). Thus, we examined these features of DD PPR gene pairs as indicators of retroposition. Specifically, we compare intron counts of PPR genes in each dispersed duplicate pair and find that 74.78% to 93.21% of DD-derived PPR genes in hornworts and lycophytes have undergone intron loss, followed by the moss T. lepidozioides with 51.98% (Table S6 and Fig. 2b; for details, see the Materials and methods section). Furthermore, we investigate the presence of polyA tails downstream of the 3′ untranslated regions (3′UTRs) of intronless DD-derived PPR genes and find that 97.99% to 99.26% possess polyA tails in lycophytes, 70.37% to 75.61% in hornworts, and 81.55% in the moss T. lepidozioides (Fig. S10). Meanwhile, we identify TEs within 20 kb upstream and downstream of intronless DD PPR gene pairs, showing that 70.52% to 99.93% of them are flanked by TEs in lycophytes, 26.56% to 68.31% in hornworts, and 80.11% in T. lepidozioides (Fig. S11). Notably, most of these TEs are long terminal repeats (LTRs), including LTR/Copia, LTR/Gypsy, and LTR/unknown (Table S7). In addition, we present two DD PPR gene pairs with intron loss that are flanked by the same type of LTR with high sequence similarity, viz., E2 subgroup PPR genes, HaA01G1373 and HaA43G0090 in H. asiatica flanked by LTR/Copia (96.43% upstream and 100% downstream) (Fig. 2c), and DYW subgroup PPR genes, Tle3c03864.1 and Tle3c03878.1 in T. lepidozioides by LTR/Gypsy (99.86% upstream and 100% downstream) (Fig. 2d). Furthermore, co-occurrence of these three retroposition features is detected in 69.93% to 97.92% of DD PPR gene pairs in lycophytes, 17.75% to 51.86% in hornworts, and 64.97% in T. lepidozioides (Fig. 2e), suggesting lineage-specific degeneration of retroposition signals. Collectively, these results indicate that PPR gene expansions in early-diverging land plants are primarily driven by DD associated with retroposition.
Horizontal transfer of DYW subgroup PPR genes from plants to bdelloid rotifers
As PPR genes consist of the P and PLS subfamilies, we explore their phylogenetic patterns by reconstructing an evolutionary landscape across Archaeplastida (Fig. S12). We find that the P subfamily PPR genes are ubiquitously distributed in Archaeplastida, with their numbers progressively increasing from early-diverging glaucophytes to highly derived angiosperms. Conversely, the PLS subfamily, including PLS, E, E+, DYW, and DYW:KP subgroups, is confined to land plants, with rare occurrence in liverworts and some mosses. Among C-terminally truncated PPR proteins, the PLS and E subgroups are numerous in lycophytes, while the E+ subgroup is exclusively enriched in hornworts (Fig. S12). The DYW subgroup is broadly distributed in land plants, with a maximum of 750 members observed in the moss T. lepidozioides. In contrast, its variant, the DYW:KP subgroup, is restricted to hornworts, lycophytes, and ferns (Table S3 and Fig. S12), conforming with a previous study (Gutmann et al. 2020).
Inspired by the previous findings that DYW subgroup PPR genes have also been detected in nonplants (Schallenberg-Rüdinger et al. 2013; Fu et al. 2014), we next investigate the evolutionary relationships of the DYW and DYW:KP subgroups between plants and nonplants. Given that repetitive and variable PLS motif triplets may hinder reliable phylogenetic inference and DYW domain catalyzes C-to-U editing via its PG box (covering the first two β-strands), the zinc ion binding signature HxE(x)nCxxC, and C-terminal DYW tripeptide (Hayes et al. 2013; Boussardon et al. 2014; Takenaka et al. 2021; Yang et al. 2023b), we extract 6,646 DYW and 3,761 DYW:KP domains of 68 land plants as representatives of their respective subgroup PPR genes (Table S8). Based on these plant domains, we further survey nonplant DYW and DYW:KP domains across 40 lineages spanning eukaryotes, prokaryotes, and viruses (Table S9; for details, see the Materials and methods section). In total, we detect 1,666 DYW and 6 DYW:KP domains in Metazoa, 31 DYW in Rhizaria, 22 DYW and 1 DYW:KP in Discoba, 15 DYW in Fungi, 9 DYW in Alveolata, 3 DYW in Haptista, and 1 DYW in Amoebozoa (Table S10), exhibiting 21.21% to 80% sequence similarity to their plant counterparts (Table S11). To reduce domain redundancy, we further select 1,151 DYW and 356 DYW:KP domains from six representative species spanning major land plant lineages, including P. pearsonii (hornwort), T. lepidozioides (moss), I. sinensis (lycophyte), Alsophila spinulosa (fern), Gnetum montanum (gymnosperm), and Amborella trichopoda (angiosperm), together with all nonplant DYW and DYW:KP domains for phylogenetic reconstruction (Table S10). Notably, three major clusters (I to III) of nonplant DYW domains are embedded within plant lineages, suggesting multiple waves of HGT from plants to nonplants (Fig. 3).
Figure 3.
The maximum-likelihood phylogenetic tree of DYW and DYW:KP domains in plants and nonplants. Three predominant clusters (I, II, and III) of nonplant DYW domains are embedded within plants. Among nonplant lineages, Metazoa contain the largest number of DYW domains. Importantly, cluster IV of Metazoa contains numerous DYW domains, which are all from bdelloid rotifers, including A. vaga, A. ricciae, A. steineri, D. carnosus, P. roseola, R. magnacalcarata, R. socialis, R. sordida, R. sp. “Silwood1”, and R. sp. “Silwood2”. Best-fit model (IQ-TREE): JTT + R10. Bootstraps are from 1,000 replicates.
To explore the potential role of DYW subgroup PPR genes in nonplants, we select a representative DYW domain from each lineage and compare its sequence and structure with those of plant donors. On the one hand, most of these nonplant DYW domains contain the conserved HxE(x)nCxxC motif, except for the HxE(x)nGxxN in Fungi, whereas the PG box (approximately 15 amino acids [aa] beginning with proline and glycine [Hayes et al. 2013]) and DYW tripeptide are variable (Fig. 4a and b; Fig. S13). On the other hand, these nonplant DYW domains exhibit significant structural similarity to those of their plant donors (root mean square deviation, RMSD ≤ 2.5 Å [Bolhuis 2006]), such as the DYW domain in the bdelloid rotifer Adineta vaga (Metazoa) and its counterpart in the moss T. lepidozioides (Fig. 4b and Fig. S13). Surprisingly, Metazoa (cluster IV) possess the largest number of DYW domains (Fig. 3), all of which occur in bdelloid rotifers that are resistant to ultraviolet (UV) radiation and inhabit diverse environments, such as mosses, lichens, soils, and deserts (Hespeels et al. 2023; Wang et al. 2025a). Notably, these bdelloid rotifers comprise A. vaga, Adineta ricciae, Adineta steineri, Didymodactylos carnosus, Philodina roseola, Rotaria magnacalcarata, Rotaria socialis, Rotaria sordida, Rotaria sp. “Silwood1”, and Rotaria sp. “Silwood2”. Among them, we find that A. vaga harbors lots of PPR genes (n = 191), nearly one-third of which (n = 60) are members of the DYW subgroup (Table S12). The majority of these DYW subgroup PPR genes (58/60) are expressed (Fig. 4c), suggesting their potential roles in biological processes. Intriguingly, the moss T. lepidozioides, a potential donor of DYW subgroup PPR genes to A. vaga, is also resistant to UV radiation (Hu et al. 2023). Consistent with this resistance, its DYW genes are upregulated during recovery from UV irradiation (Fig. 4d and Table S13).
Figure 4.
HGT of DYW domains from plants to bdelloid rotifers. a) Sequence alignments between plant and nonplant DYW domains. Amino acids in rectangles denote PG box, HxE(x)nCxxC, and DYW tripeptide. b) Protein structure alignment between DYW domain (protein ID: Tle1c03524.1) in the moss T. lepidozioides and its counterpart (protein ID: UJR34389.1) in the bdelloid rotifer A. vaga. c) Expression of PPR genes (DYW, E+, PLS, and P subgroups) in whole tissues of A. vaga. TPM, transcripts per million. d) Expression of DYW subgroup PPR genes in T. lepidozioides under control condition and after 2-h UV irradiation, followed by 0, 6, 12, and 24-h recovery.
Evolutionary conservation and diversity of MORF domains
MORF proteins, exclusive to seed plants (Fig. 1c) and comprising nine members (MORF1 to MORF9) in A. thaliana, are characterized by their MORF hallmark domains and are essential for PPR-mediated RNA editing (Takenaka et al. 2012; Haag et al. 2017; Li et al. 2025a). Moreover, MORF proteins function as essential components of RNA editing complexes, engaging in complex physical interactions with themselves and various editing factors, such as PPR, OZ1, ORRM1, and ISE2 proteins (Sandoval et al. 2019). To explore the evolutionary conservation of MORF domains, we extract 334 instances from 33 representative seed plants (Table S14) and identify their homologs across 45 lineages spanning eukaryotes, prokaryotes, and viruses (Table S9). We find that MORF domains are homologous to peptidase S8 propeptide/proteinase inhibitor I9 domains (also known as PA domain [Haag et al. 2017]) in metazoa, fungi, bryophytes, lycophytes, and ferns (Fig. 5a and Table S15), extending the previous identification of PA domains in mosses and ferns (Luo et al. 2017). Furthermore, we compare the structures of the MORF domain and its homologous PA domain with the highest sequence similarity in each of these five lineages. We find that although they have relatively low sequence similarity (33.33% to 41.94%), their structures are significantly similar (RMSD ≤ 2.5 Å), with RMSD values varying from 0.705 to 1.039 Å (Fig. 5b and c; Fig. S14). Notably, MORF domains from MORF3 and MORF8 proteins share the same α-helix and β-strand architecture as PA domains (Fig. 5c and Fig. S14). Given the importance of PA domain in assisting zymogens and mature peptidase folding (Li et al. 1995; Jain et al. 1998), these results imply that MORF domains from MORF3 and MORF8 proteins may modulate protein folding during the regulation of RNA editing.
Figure 5.
Conservation and diversity of MORF domains. a) The maximum-likelihood phylogenetic tree of MORF and PA domains. The tree scale bar represents the number of substitutions per site. Model (FastTree): Jones–Taylor–Thornton, CAT approximation with 20 rate categories. b) Sequence and c) structure alignments between MORF3 domain (protein ID: Pd03G21970A) from Pinus densiflora and its homologous PA domain (protein ID: XP_018897422.1) from Bemisia tabaci. d) The maximum-likelihood phylogenetic tree of MORF1 to MORF9 domains from angiosperms and gymnosperms. Five nonseed plant PA domains with highest sequence similarity to MORF domains are used as the outgroup (protein IDs: Dicom.22G070000.1.p, HaB36G042, Phcar.S3G189600.t1, Phsp.C2G192800.t1, and XP_024533311.1). Best-fit model (IQ-TREE): JTT + R8. Bootstraps are from 1,000 replicates.
To investigate the evolutionary diversity of MORF domains, we classify them into MORF1 to MORF9 as identified in A. thaliana and survey PA domains across our collected 364 high-quality Archaeplastida genomes. Accordingly, PA domains are ubiquitous across major lineages, including glaucophytes, charophytes, mosses, liverworts, hornworts, lycophytes, ferns, gymnosperms, and angiosperms (Fig. S15). Furthermore, we reconstruct a phylogenetic tree of the 334 MORF domains, using five nonseed plant PA domains with the highest sequence similarity as the outgroup. Consequently, MORF2, MORF5, and MORF6 domains cluster into early-diverging clades in the phylogenetic tree, indicating their earlier origins relative to other MORF domains in seed plants (Fig. 5d). Moreover, most MORF domains (MORF3 to 6, 8, and 9) are broadly distributed across angiosperms and gymnosperms, suggesting that MORF diversification predates the major divergences among seed plant lineages. This observation is consistent with previous phylogenetic and comparative analyses on MORF genes across different species (Luo et al. 2017). At the gene level, we analyze the expression of MORF1 to MORF9 genes, each harboring a single MORF hallmark domain, across multiple tissues of the gymnosperm Torreya grandis and the angiosperm A. thaliana. Notably, among these MORF genes, MORF8 and MORF9 display higher expression across tissues in both species, consistent with observations in Apium graveolens and Populus trichocarpa (Wang et al. 2019a; Liu et al. 2024), indicating their conservation in expression (Fig. S16).
Hallmark domain fusions among MORF, ORRM, OZ, and PPR proteins
While MORF proteins contain the MORF hallmark domain, the presence of additional domains, such as the GR domain in MORF1, MORF4, and MORF8 proteins (Li et al. 2025a), suggests the novel domain composition and underscores the need for an updated classification of MORF proteins. Thus, in terms of domain composition, we classify our identified 2,759 MORF proteins into four groups: (1) single-pure MORF (SP-MORF, n = 1,916), containing a single MORF domain; (2) single-mixed MORF (SM-MORF, n = 544), containing a single MORF domain plus one or more non-MORF domains; (3) multiple-pure MORF (MP-MORF, n = 42), containing multiple MORF domains; and (4) multiple-mixed MORF (MM-MORF, n = 257), containing multiple MORF domains plus one or more non-MORF domains (Fig. 6a and Table S16). Furthermore, we reconstruct the evolutionary landscape of these reclassified MORF proteins across 262 angiosperms and 13 gymnosperms (Fig. 6b). Specifically, SP-MORF proteins are widespread in gymnosperms and angiosperms with larger numbers, whereas SM-MORF proteins are abundant in angiosperms but rare in gymnosperms. Conversely, MP-MORF proteins occur sparsely in both lineages and MM-MORF proteins are only observed in angiosperms. Importantly, these reclassified MORF proteins exhibit a wide range of domain compositions. Among them, 492 SM-MORF and 9 MM-MORF proteins harbor a GR domain alongside MORF1 to 6, 8, and 9 domains in angiosperms, whereas in the gymnosperm T. grandis, a single SM-MORF protein contains a GR domain co-occurring with MORF9 domain, extending the previously observed MORF-GR domain architectures in MORF1, MORF4, and MORF8 proteins (Li et al. 2025a). Surprisingly, we find that MORF domains and P motifs are fused in two SM-MORF proteins, viz., AY_017782_RA in the angiosperm Aquilaria yunnanensis and KAH9320946.1 in the gymnosperm Taxus chinensis (Table S16). Moreover, MORF domains and the RRM hallmark domains of ORRM are fused in SM- and MM-MORF proteins in angiosperms. Specifically, these domain fusions fall into four types: 1-MORF & 1-RRM (n = 23), 1-MORF & 3-RRM (n = 1), 2-MORF & 1-RRM (n = 243), and 3-MORF & 1-RRM (n = 3) (Fig. S17 and Table S16), where the leading number indicates the copy number of each corresponding domain. Notably, the most abundant fused type, 2-MORF & 1-RRM, is represented by the RNA editing factor ORRM1 (Searing et al. 2020).
Figure 6.
Reclassification and diversity of PPR accessory proteins. a) Reclassification of MORF proteins based on domain composition. Blue and green bars represent MORF1 to MORF9 and other domains, respectively. SP-MORF, single-pure MORF; SM-MORF, single-mixed MORF; MP-MORF, multiple-pure MORF; MM-MORF, multiple-mixed MORF. b) The phylogenetic tree of reclassified MORF proteins from angiosperms and gymnosperms. c) Fusions among PPR, MORF, RRM, and RanBP2 Znf domains. Numbers in parentheses indicate domain counts. Proteins marked with solid and dashed arrows are localized to mitochondria/plastids and other subcellular compartments, respectively.
ORRM proteins, specifically, ORRM1 to ORRM6, have been identified in several angiosperms as mediators of RNA editing (Shi et al. 2015, 2016a). To refine the classification of the ORRM family within the broader RRM domain-containing proteins and identify novel members, we define ORRM proteins by considering both RRM domain architecture and organellar localization (for details, see the Materials and methods section). Consistent with the MORF reclassification framework, we recategorize our identified 7,760 ORRM proteins into SP-ORRM (n = 2,689), SM-ORRM (n = 1,751), MP-ORRM (n = 2,948), and MM-ORRM (n = 372) (Table S17). Furthermore, we reconstruct their evolutionary landscape across Archaeplastida and find that SP-, SM-, and MP-ORRM proteins exhibit a gradual increase from glaucophytes to angiosperms with larger number, while MM-ORRM proteins show a sparse distribution (Fig. S18). In addition to previously identified ORRM1 to ORRM6 with a single RRM domain, MP-ORRM proteins, such as CP29A, CP31A, and CP31B (which also have been characterized as members of the chloroplast ribonucleoprotein family), each containing two RRM domains, also affect RNA editing (Shi et al. 2016a). Intriguingly, numerous RRM and PPR domains are fused in SM- and MM-ORRM proteins across Archaeplastida (Fig. S19). Specifically, SM-ORRM proteins contain the highest number of RRM-P fusions in angiosperms (n = 171), followed by hornworts (n = 6), mosses (n = 4), liverworts (n = 4), charophytes (n = 3), gymnosperms (n = 3), lycophytes (n = 1), and ferns (n = 1). Moreover, SM-ORRM proteins exhibit one RRM-DYW fusion in hornworts and angiosperms and three RRM-E fusions in angiosperms (Table S17). Conversely, MM-ORRM proteins have only five RRM-P and one RRM-E fusions in angiosperms. In addition, fusions of RRM and the RanBP2 Znf hallmark domain of OZ protein are detected in proteins that are not targeted to mitochondria or plastids (Table S18). Similarly, we reclassify OZ proteins and explore their evolution using the same approach applied to MORF and ORRM proteins. Consequently, we identify 761 MP-OZ, 61 MM-OZ, 6 SP-OZ, and 1 SM-OZ proteins (Table S19). Based on their evolutionary landscape across Archaeplastida, we find that MP-OZ proteins are enriched in angiosperms (n = 725), followed by mosses (n = 14), lycophytes (n = 6), ferns (n = 6), gymnosperms (n = 4), chlorophytes (n = 3), and liverworts (n = 3) (Fig. S20). In addition, 60 MM-OZ, 6 SP-OZ, and 1 SM-OZ proteins are found in angiosperms, while a single MM-OZ occurs in hornworts. Notably, we detect that RanBP2 Znf and PPR domains (P and E) are fused in four angiosperm proteins, despite their absence of targeting to mitochondria or plastids (Table S20). Collectively, our results indicate that hallmark domains of PPR, MORF, ORRM, and OZ proteins can be fused with more diverse patterns than previously thought (Fig. 6c).
Discussion
Our kingdom-wide characterization of RNA editing factors across Archaeplastida reveals a complex evolutionary trajectory defined by massive gene expansions, HGT, and the diversity and fusion of hallmark domains. PPR genes, the primary mediators of plant RNA editing, exhibit dramatic expansions in hornworts, mosses, and lycophytes. We demonstrate that these expansions are predominantly driven by DD associated with retroposition, evidenced by signatures such as intron loss, polyA tails, and flanking TEs. In contrast, WGD plays a subsidiary role in PPR expansion in lycophytes and contributes little to that in the moss T. lepidozioides (Fig. 2a). Notably, T. lepidozioides exhibits predominant DD but limited retroposition features, suggesting large-scale gene fractionation after an ancient WGD event (Liang and Schnable 2018). These results imply that retroposition, as well as WGD, provided the genomic plasticity necessary for early land plants to diversify their RNA editing machinery during the colonization of harsh terrestrial environments (Fujii and Small 2011).
The evolutionary reach of plant RNA editing extends far beyond Archaeplastida. The discovery of DYW subgroup PPR genes in diverse nonplant lineages—including Alveolata, Amoebozoa, Discoba, Fungi, Haptista, Metazoa, and Rhizaria—suggests that DYW-mediated C-to-U editing is a broadly distributed eukaryotic mechanism, which is more common than previously recognized (Rüdinger et al. 2011; Fu et al. 2014; Yang et al. 2017a). Noticeably, high expression of these PPR genes in both recipient bdelloid rotifers (Metazoa) and their putative plant donors (e.g. T. lepidozioides) might be associated with adaptation to UV radiation, as highly expressed DYW subgroup PPR proteins may mediate the editing of increased RNAs during recovery from UV irradiation. However, the driving force of HGT of DYW subgroup PPR genes from plants to bdelloid rotifers remains elusive. On the one hand, inspired by the origin of RNA editing in plants, constructive neutral evolution may partially account for the origin of the HGT (Fig. S21), as DYW-mediated RNA editing can alleviate mRNA-level defects—potentially induced by UV irradiation—thereby enabling the fixation of DNA mutations (Covello and Gray 1993; Gray 2012). On the other hand, positive selection may partially explain the occurrence of HGT of DYW subgroup PPR genes, since RNA editing mediated by these genes can promote translation by creating start codons or removing stop codons (Fauskee et al. 2025; Kwok van der Giezen et al. 2025). This process potentially generates new functional proteins that might confer an adaptive advantage in extreme environments (Takenaka et al. 2013), such as UV radiation.
MORF proteins, restricted to seed plants, display significant structural similarity between their hallmark MORF domains and PA domains of other proteins. Specifically, MORF3 and MORF8 domains more closely resemble PA domains associated with protein folding, whereas MORF1 domain has been reported to be more structurally similar to an N-terminal ferredoxin-like domain that confers RNA substrate positioning in bacterial 4-thio-uracil tRNA synthetases (Haag et al. 2017). Our results suggest that the structural diversity of MORF domains underpins their functional versatility in RNA editing regulation.
Prompted by the structural diversity observed among MORF domains, we extended our analysis to MORF, ORRM, and OZ proteins, systematically reclassifying them based on domain composition. This reclassification revealed not only a broader repertoire of potential RNA editing factors but also extensive hallmark domain fusions among PPR, MORF, ORRM, and OZ proteins (except between MORF and OZ), mirroring their reported interactions (Sun et al. 2013; Yan et al. 2018; Gipson et al. 2022; Li et al. 2025a; Wang et al. 2025c) and exemplifying the general link between domain fusion and protein–protein networking (Enright et al. 1999; Marcotte et al. 1999a, 1999b). Based on these findings, we propose an evolutionary model for land plant RNA editosomes, characterized by a diversification in composition from hornworts to angiosperms that potentially underpins the maintenance of efficient RNA editing (Fig. 7). Specifically, editosomes in hornworts and mosses primarily comprise PPR and ORRM proteins. Although we identified a single MM-OZ (containing 1 WSS1-like metalloprotease and 2 RanBP2 Znf domains) protein in hornworts and 14 MP-OZ (containing 3 RanBP2 Znf domains) proteins in mosses (Table S19), both architectures are distinct from the editing factor OZ1 with two RanBP2 Znf domains (Gipson et al. 2022), leaving their roles in RNA editing regulation elusive. In contrast, editosomes in liverworts, lycophytes, and ferns exhibit increased complexity, predominantly incorporating PPR, ORRM, and OZ1-like proteins. Moreover, hornworts, lycophytes, and ferns harbor distinct PPR proteins specialized for mediating C-to-U and U-to-C RNA editing (Knoop 2023). In addition to PPR, ORRM, and OZ proteins, the recruitment of MORF proteins as novel components further expands the combinatorial complexity of editosomes in gymnosperms and angiosperms. Importantly, extensive studies have demonstrated that angiosperm editosomes are characterized by highly heterogeneous combinatorial patterns of PPR, MORF, ORRM, and OZ proteins, as exemplified by the previously reported editosomes: (i) E+ subgroup PPR protein, ZmNUWA (P subfamily PPR protein), ZmDYW2A/2B (DYW subgroup PPR protein), ZmMORF, ZmORRM1, and ZmOZ1 (Wang et al. 2024); (ii) E subgroup PPR protein, GRP23 (P subfamily PPR protein), MEF8/MEF8S (DYW subgroup PPR protein), and MORF proteins (Yang et al. 2022); (iii) E+ subgroup PPR protein, NUWA (P subfamily PPR protein), DYW2 (DYW subgroup PPR protein), and MORF proteins (Yang et al. 2022); (iv) DEK53 (E subgroup PPR protein), MORF1, ORRM2, ORRM3, and ORRM4 (Dai et al. 2020); and (v) AtECB2 (DYW subgroup PPR protein), HEMC (porphobilinogen deaminase), MORF2, and MORF8 (Huang et al. 2017). Beyond the well-characterized PPR, MORF, ORRM, and OZ families, editosomes across various land plant lineages may potentially involve other types of proteins. Taken together, these computationally derived insights provide a refined understanding of the genomic innovations that have shaped the plant RNA machinery.
Figure 7.
A proposed model for the evolutionary diversification of plant RNA editing factors. Editing factor composition varies from hornworts to angiosperms, potentially maintaining efficient RNA editing. The question mark indicates that the function of OZ protein remains elusive in hornworts and mosses. Deep-rose PPR, DYW or DYW:KP subgroup PPR protein; red-orange PPR, DYW subgroup PPR protein; C => U, C-to-U RNA editing; C <=> U, C-to-U or U-to-C RNA editing.
Materials and methods
Collection and quality control of Archaeplastida genomes
To estimate the numbers of PPR, MORF, ORRM, and OZ genes and reconstruct their evolutionary landscapes, 364 Archaeplastida high-quality genomes were retrieved from public databases (Table S1). To ensure accurate gene counts, only the longest protein isoform of each gene was retained. The completeness of annotated gene sets was estimated by Benchmarking Universal Single-Copy Orthologue (BUSCO) (version 5.8.2) with three datasets chlorophyta_odb10, viridiplantae_odb10, and eukaryota_odb10 (Manni et al. 2021). As these datasets contain abundant sequences of chlorophytes and angiosperms, genomes in these two lineages with BUSCO scores (“complete”) greater than 90% were retained. For other lineages, chromosome-level genomes with BUSCO scores (“complete”) greater than 85%; otherwise, the top 20% genomes with the highest BUSCO scores were retained within each lineage.
Phylogenomic tree reconstruction of Archaeplastida
To reconstruct the phylogenomic tree of 364 Archaeplastida species, single-copy orthologous gene families were inferred by OrthoFinder (version 3.0.1b1) (Emms and Kelly 2019), and 138 single-copy orthologous gene families covering at least 75% species were retained. These single-copy orthologous gene families were aligned by MAFFT (version 7.453) (Katoh and Standley 2013) with the parameter “–auto,” and poorly aligned regions were identified and trimmed by BMGE (version 1.12) with BLOSUM30 (Criscuolo and Gribaldo 2010). Gene trees were constructed by IQ-TREE (version 2.1.2) (Minh et al. 2020), with the best-fit substitution model for each gene automatically selected according to Bayesian Information Criterion, and using parameters “-B 1000 -bnni -m MFP -T AUTO -ntmax 20.” The resulting trees were subsequently merged, and the species tree was inferred by ASTRAL (version 5.7.8) (Mirarab et al. 2014) using Cyanophora paradoxa as the outgroup.
Identification of PPR, MORF, ORRM, and OZ proteins
The hallmark domains and motifs of PPR (P, P1, P2, S1, S2, SS, L1, L2, E1, E2, E+, DYW, and DYW:KP), MORF (MORF), ORRM (RRM), and OZ (RanBP2 Znf) proteins were identified using hmmsearch in HMMER (version 3.2.1) (Eddy 2011) with corresponding domain and motif profiles. PPR hallmark domain and motif profiles were obtained from a previous study (Gutmann et al. 2020) and searched using hmmsearch with parameters “–domtblout -noali -E 0.1.” All hmmsearch scores for P, P1, P2, S1, S2, L1, L2, E1, E2, and E+ domains or motifs were greater than 0, with SS motif greater than 10, and DYW/DYW:KP domain greater than 30. As most PPR hallmark domain-containing proteins localize to mitochondria and/or plastids (Meng et al. 2024; Wang and Tan 2025), they are defined as PPR proteins. In addition, PPR proteins were required to contain at least three of P, P1, P2, L1, L2, or SS motifs, unless they contained a DYW or DYW:KP domain.
The MORF domain profile, constructed using hmmbuild in HMMER (version 3.2.1) with nine MORF domains from A. thaliana, was searched using hmmsearch with parameters “–domtblout -noali -E 1e-10.” As MORF domain was homologous to PA domain, only hits with hmmsearch score greater than 36 were kept. The RRM domain profile was obtained from Pfam (PF00076.26) (Mistry et al. 2021) and searched using hmmsearch with parameters “–domtblout -noali -E 1e-10.” The hmmsearch score was greater than 0. GR domains, present in some MORF and ORRM proteins, were defined as regions containing at least eight glycines within a 20-amino-acid sliding window. The RanBP2 Znf profile was obtained from Pfam (PF00641.22) and searched using hmmsearch with the same parameters as for the RRM domain. As MORF, ORRM, and OZ proteins are localized to mitochondrion and/or plastids, the subcellular localization of the identified hallmark domain-containing proteins was predicted using DeepLoc (version 2.1) (Ødum et al. 2024) with “Accurate” mode, TargetP (version 2) (Armenteros et al. 2019), Plant-mSub (Sahu et al. 2020) (https://bioinfo.usu.edu/Plant-mSubP), and LOCALIZER (Sperschneider et al. 2017) (https://localizer.csiro.au). Only localization results supported by at least two tools were retained. The domain composition was identified using InterPro (https://www.ebi.ac.uk/interpro) (Blum et al. 2025).
WGD inference
To infer WGD, intra-genome alignments were conducted using BLASTP in the Basic Local Alignment Search Tool (BLAST, version 2.12.0+) (Camacho et al. 2009) with an E-value cutoff of 1e-5. Based on the alignments, homologous dotplots were generated using WGDI (Sun et al. 2022), and synteny blocks were inferred using MCScanX (Wang et al. 2012b). Ks values of homologs in synteny blocks were calculated using KaKs_Calulator (version 3.0) (Zhang 2022) and ParaAT (version 2.0) (Zhang et al. 2012) with parameters “-p proc -m mafft -f axt -g –k.” The median Ks value for each synteny block was calculated. Collectively, Ks peak counts and synteny blocks in the homologous dotplot were used to infer the number of WGD events.
Identification and classification of duplicated gene pairs
To identify duplicated gene pairs, all protein sequences within a genome were aligned using BLASTP (BLAST, version 2.12.0+) with an E-value cutoff of 1e-5. The highest-scoring alignment for each gene was considered a duplicated pair. Alignments covering more than 60% of the longer gene (query or subject) were required to have a sequence similarity exceeding 20%. Alternatively, for alignments covering less than 60% of the longer gene, the alignment was required to span at least 105 amino acids with a sequence similarity exceeded 40%. Otherwise, the query and subject genes were considered singletons. Based on genomic loci, duplicated gene pairs were classified into four categories: tandem pairs (adjacent genes), proximal pairs (within ten loci on the same chromosome), dispersed pairs (more than ten loci apart on the same chromosome or on different chromosomes), and WGD pairs (located in synteny blocks).
TE annotation
TEs were identified by a combination of ab initio and homology-based methods. For ab initio prediction, a consensus sequence library was built using RepeatModeler (version 2.0.6) (Flynn et al. 2020) with the parameter “-LTRStruct.” An LTR library was then constructed using LTR_Finder (version 1.07) (Xu and Wang 2007)/LTR_FINDER_parallel (version 1.1) (Ou and Jiang 2019), LTRharvest (version 1.6.6) (Ellinghaus et al. 2008) with parameters “-minlenltr 100 -maxlenltr 7000”, and LTR_retriever (version 2.9.0) (Ou and Jiang 2018). Both libraries were used to annotate TEs using RepeatMasker (version 4.1.8) (Chen 2004). For homology-based detection, TEs were identified using RepeatMasker (version 4.1.8) with Dfam (version 3.9) (Storer et al. 2021) database. TE annotations obtained from both methods were then combined.
Retroposition inference
Retroposition, underlying DD gene pairs, was characterized by intron loss, polyA enrichment downstream of the 3′UTR, and flanking TEs. First, intron loss in one DD gene pair was required to meet one of following criteria: (i) both genes had no intron; (ii) a gene had no intron, while the other had at least one intron; (iii) considering intron retention, a gene contained one intron, while the other had at least one intron; (iv) both genes had more than one intron, differing by more than ten introns. Second, polyA enrichment downstream of the 3′UTR of the intronless gene (or polyT enrichment upstream of the 5′UTR) was defined as at least eight adenines/thymines within a 10-base window. Genomic annotations were used to detect polyA/T enrichment in regions extending 1 kb downstream of the 3′UTR or upstream of the 5′UTR. If the 3′UTR or 5′UTR of a gene was not annotated, the shortest corresponding UTR in the genome was used. If no 3′UTR or 5′UTR was annotated in the genome, they were defined as the 100-bp regions downstream of the stop codon and upstream of the start codon, respectively. Third, DD gene pairs flanked by TEs within 20 kb upstream and downstream were defined as TE-flanked.
Identification of DYW and DYW:KP domain homologs and HGT inference
To explore HGT of plant DYW and DYW:KP domains, 10,407 protein sequences (hmmsearch score ≥ 30, length ≥ 130 aa) of 68 representative species were collected (Table S8). DYW and DYW:KP domain homologs were searched in the NCBI Non-redundant Protein Sequence Database (NR) (acquisition date: 2025.01.19) using BLASTP (BLAST, version 2.12.0+) with an E-value cutoff of 1e-5. Protein sequences retrieved from NR originated from 40 lineages (Table S9). Homologs exhibiting more than 20% sequence similarity to plant DYW and DYW:KP domains were further confirmed via hmmsearch in HMMER (version 3.2.1) using a score cutoff of 30. To avoid contamination, nonplant DYW and DYW:KP domains were retained only if present in at least two species within the corresponding lineage. Otherwise, nonplant DYW and DYW:KP domains detected in a single species were retained only if present in at least one additional genome assembly.
To reconstruct the phylogenetic tree of DYW and DYW:KP domains from plants and nonplants, domains from six representative plants (P. pearsonii, T. lepidozioides, I. sinensis, A. spinulosa, G. montanum, and A. trichopoda) and all nonplant counterparts were aligned using MAFFT (version 7.525), with poorly aligned regions trimmed by trimAl (version 1.5.rev0) (Capella-Gutiérrez et al. 2009) using parameters “-gt 0.7 -cons 60 -st 0.1”. The phylogenetic tree was inferred using IQ-TREE (version 2.1.2) (Best-fit model: JTT + R10) with parameters “-B 1000 -bnni -m MFP -T AUTO –safe” and visualized with iTOL (version 7) (Letunic and Bork 2024). Bootstraps were from 1,000 replicates. Multiple sequence alignment of DYW domains from plants and nonplants was performed using MEGA (version 11) (Tamura et al. 2021) and visualized with Jalview (version 2.11.4.0) (Procter et al. 2021).
Identification of MORF domain homologs
To identify homologs of plant MORF domains, 334 protein sequences (hmmsearch score ≥ 0, length ≥ 86 aa) from 33 representative species were collected (Table S14). These MORF domain homologs searched in the NCBI NR using BLASTP (BLAST, version 2.12.0+) with an E-value cutoff of 1e-5. The protein sequences retrieved from NR originated from 45 lineages (Table S9). Homologs with more than 20% similarity to plant MORF domains were retained. The MORF domains and their homologs were aligned using MAFFT (version 7.525), and the phylogenetic tree was constructed using FastTree (Price et al. 2010) (Model: Jones–Taylor–Thornton, CAT approximation with 20 rate categories) with parameters “-pseudo -gamma” and visualized by iTOL (version 7).
Phylogenetic analysis of MORF domains
To reconstruct the phylogenetic tree of MORF domains, all MORF domains from gymnosperms and angiosperms were selected. These domains were aligned to MORF1∼MORF9 domains from A. thaliana to assign their types. Five PA domains from nonseed plants (Dicom.22G070000.1.p, HaB36G042, Phcar.S3G189600.t1, Phsp.C2G192800.t1, XP_024533311.1), exhibiting the highest similarity to MORF domains, were used as the outgroup. All MORF domains, along with the outgroup sequences, were aligned using MAFFT (version 7.525). The phylogenetic tree was inferred using IQ-TREE (version 2.1.2) (Best-fit model: JTT + R8) with parameters “-B 1000 -bnni -m MFP -T AUTO” and visualized by iTOL (version 7) (Letunic and Bork 2024). Bootstraps were from 1,000 replicates.
Analysis of PPR and MORF gene expression
To calculate gene expression, reference genome indexes were built using hisat2-build in HISAT2 (version 2.0.5) (Kim et al. 2019) and rsem-prepare-reference in RSEM (version 1.3.1) (Li and Dewey 2011). Raw RNA sequencing data were quality-filtered using fastp (version 0.20.1) (Chen et al. 2018) with parameters “-q 20 -u 40 -l 15 -g -x -r -W 4 -M 20 -w 12”. Strand specificity was inferred using infer_experiment.py in RSeQC (version 2.6.4) (Wang et al. 2012a). Reads mapping was conducted using STAR (version 2.7.1a) (Dobin et al. 2013). Gene expression matrix was generated using rsem-calculate-expression in RSEM (version 1.3.1) with parameters “–keep-intermediate-files –time –star –append-names –output-genome-bam –sort-bam-by-coordinate”.
Protein structure prediction and alignment
Protein structures were predicted using ColabFold (https://colab.research.google.com/github/sokrypton/ColabFold) (Mirdita et al. 2022), and different protein structures were aligned using “Matchmaker” in ChimeraX (version 1.9) (Meng et al. 2023).
Supplementary Material
Acknowledgments
We thank Shuhui Song, Lina Ma, Chao Zhang, Tianyi Xu, Tongtong Zhu, Yetong Chang, Shuangyang Wu, Lin Liu, Zhao Li, Yang Zhang, Xing Zheng, Yue Qi, Xinyu Zhou, Wenzhuo Cheng, Yuxin Qin, Miaomiao Wang, Shiting Wang, Lingjie Wang, Zihan Wang, Kehua Ma, Pan Li, Yiran Zhan, Zheng Luo, and Dechang Yang for their valuable discussions and advice.
Contributor Information
Ming Chen, National Genomics Data Center, China National Center for Bioinformation, Beijing, China; Beijing Key Laboratory of Intelligent Governance and Application of Biological Big Data, China National Center for Bioinformation, Beijing, China; Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China; School of Future Technology and College of Life Sciences, University of Chinese Academy of Sciences, Beijing, China.
Qingming Qu, State Key Laboratory of Cellular Stress Biology, School of Life Sciences, Xiamen University, Xiamen, China.
Zhang Zhang, National Genomics Data Center, China National Center for Bioinformation, Beijing, China; Beijing Key Laboratory of Intelligent Governance and Application of Biological Big Data, China National Center for Bioinformation, Beijing, China; Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China; School of Future Technology and College of Life Sciences, University of Chinese Academy of Sciences, Beijing, China.
Author contributions
M.C. collected and analyzed the data, and wrote the manuscript. Z.Z. conceived, initiated and supervised this study. Q.Q. and Z.Z. revised the manuscript.
Supplementary material
Supplementary material is available at Molecular Biology and Evolution online.
Funding
This work was supported by National Natural Science Foundation of China (T2425005).
Data availability
All data used and generated in this study are available at the Open Archive for Miscellaneous Data (https://ngdc.cncb.ac.cn/omix, OMIX: OMIX012585) (Zhang et al. 2025) in the National Genomics Data Center (NGDC), China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences (CNCB-NGDC Members and Partners 2025). The computational codes used in this study have been deposited in GitHub (https://github.com/Bioinfo-Ming/PlantEditosomeEvolution).
References
- Armenteros JJA et al. Detecting sequence signals in targeting peptides using deep learning. Life Sci Alliance. 2019:2:e201900429. 10.26508/lsa.201900429. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barkan A, Small I. Pentatricopeptide repeat proteins in plants. Annu Rev Plant Biol. 2014:65:415–442. 10.1146/annurev-arplant-050213-040159. [DOI] [PubMed] [Google Scholar]
- Baudry K, Monachello D, Castandet B, Majeran W, Lurin C. Dissecting the molecular puzzle of the editosome core in Arabidopsis organelles. Plant Sci. 2024:344:112101. 10.1016/j.plantsci.2024.112101. [DOI] [PubMed] [Google Scholar]
- Blum M et al. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res. 2025:53:D444–D456. 10.1093/nar/gkae1082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bolhuis PG. Sampling kinetic protein folding pathways using all-atom models. Lecture Notes Phys. 2006:703:393–433. 10.1007/3-540-35273-2_11. [DOI] [Google Scholar]
- Boussardon C et al. The cytidine deaminase signature HxE(x)nCxxC of DYW1 binds zinc and is necessary for RNA editing of ndhD-1. New Phytol. 2014:203:1090–1095. 10.1111/nph.12928. [DOI] [PubMed] [Google Scholar]
- Camacho C et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009:10:421. 10.1186/1471-2105-10-421. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009:25:1972–1973. 10.1093/bioinformatics/btp348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Castandet B, Araya A. RNA editing in plant organelles. Why make it easy? Biochemistry (Mosc). 2011:76:924–931. 10.1134/S0006297911080086. [DOI] [PubMed] [Google Scholar]
- Chen N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics. 2004:Chapter 4:4.10.11–14.10.14. 10.1002/0471250953.bi0410s05. [DOI] [PubMed] [Google Scholar]
- Chen S, Zhou Y, Chen Y, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018:34:i884–i890. 10.1093/bioinformatics/bty560. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cheng S et al. Redefining the structural motifs that determine RNA binding and RNA editing by pentatricopeptide repeat proteins in land plants. Plant J. 2016:85:532–547. 10.1111/tpj.13121. [DOI] [PubMed] [Google Scholar]
- CNCB-NGDC Members and Partners . Database resources of the national genomics data center, China National Center for Bioinformation in 2025. Nucleic Acids Res. 2025:53:D30–D44. 10.1093/nar/gkae978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Covello P, Gray M. On the evolution of RNA editing. Trends Genet. 1993:9:265–268. 10.1016/0168-9525(93)90011-6. [DOI] [PubMed] [Google Scholar]
- Criscuolo A, Gribaldo S. BMGE (Block mapping and gathering with entropy): a new software for selection of phylogenetic informative regions from multiple sequence alignments. BMC Evol Biol. 2010:10:210. 10.1186/1471-2148-10-210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cui J et al. Chromosome-level reference genome of tetraploid Isoetes sinensis provides insights into evolution and adaption of lycophytes. Gigascience. 2023:12:giad079. 10.1093/gigascience/giad079. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dai D et al. Maize pentatricopeptide repeat protein DEK53 is required for mitochondrial RNA editing at multiple sites and seed development. J Exp Bot. 2020:71:6246–6261. 10.1093/jxb/eraa348. [DOI] [PubMed] [Google Scholar]
- Dobin A et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013:29:15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Eddy SR. Accelerated profile HMM searches. PLoS Comput Biol. 2011:7:e1002195. 10.1371/journal.pcbi.1002195. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ellinghaus D, Kurtz S, Willhoeft U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics. 2008:9:18. 10.1186/1471-2105-9-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Emms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019:20:238. 10.1186/s13059-019-1832-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Enright AJ, Iliopoulos I, Kyrpides NC, Ouzounis CA. Protein interaction maps for complete genomes based on gene fusion events. Nature. 1999:402:86–90. 10.1038/47056. [DOI] [PubMed] [Google Scholar]
- Fauskee BD, Kuo LY, Heath TA, Xie PJ, Pryer KM. Comparative phylogenetic analyses of RNA editing in fern plastomes suggest possible adaptive innovations. New Phytol. 2025:247:2945–2963. 10.1111/nph.70244. [DOI] [PubMed] [Google Scholar]
- Feng X et al. Genomes of multicellular algal sisters to land plants illuminate signaling network evolution. Nat Genet. 2024:56:1018–1031. 10.1038/s41588-024-01737-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Flynn JM et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci U S A. 2020:117:9451–9457. 10.1073/pnas.1921046117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Freeling M. Bias in plant gene content following different sorts of duplication: tandem, whole-genome, segmental, or by transposition. Annu Rev Plant Biol. 2009:60:433–453. 10.1146/annurev.arplant.043008.092122. [DOI] [PubMed] [Google Scholar]
- Fu C-J, Sheikh S, Miao W, Andersson SG, Baldauf SL. Missing genes, multiple ORFs, and C-to-U type RNA editing in Acrasis kona (Heterolobosea, Excavata) mitochondrial DNA. Genome Biol Evol. 2014:6:2240–2257. 10.1093/gbe/evu180. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fujii S, Small I. The evolution of RNA editing and pentatricopeptide repeat genes. New Phytol. 2011:191:37–47. 10.1111/j.1469-8137.2011.03746.x. [DOI] [PubMed] [Google Scholar]
- Gipson AB, Hanson MR, Bentolila S. The RanBP2 zinc finger domains of chloroplast RNA editing factor OZ1 are required for protein–protein interactions and conversion of C to U. Plant J. 2022:109:215–226. 10.1111/tpj.15569. [DOI] [PubMed] [Google Scholar]
- Gray MW. Evolutionary origin of RNA editing. Biochemistry. 2012:51:5235–5242. 10.1021/bi300419r. [DOI] [PubMed] [Google Scholar]
- Gutmann B et al. The expansion and diversification of pentatricopeptide repeat RNA-editing factors in plants. Mol Plant. 2020:13:215–230. 10.1016/j.molp.2019.11.002. [DOI] [PubMed] [Google Scholar]
- Haag S et al. Crystal structures of the Arabidopsis thaliana organellar RNA editing factors MORF1 and MORF9. Nucleic Acids Res. 2017:45:4915–4928. 10.1093/nar/gkx099. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hao W et al. RNA editing and its roles in plant organelles. Front Genet. 2021:12:757109. 10.3389/fgene.2021.757109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hattori M, Miyake H, Sugita M. A pentatricopeptide repeat protein is required for RNA processing of clpP pre-mRNA in moss chloroplasts. J Biol Chem. 2007:282:10773–10782. 10.1074/jbc.M608034200. [DOI] [PubMed] [Google Scholar]
- Hayes ML, Giang K, Berhane B, Mulligan RM. Identification of two pentatricopeptide repeat genes required for RNA editing and zinc binding by C-terminal cytidine deaminase-like domains. J Biol Chem. 2013:288:36519–36529. 10.1074/jbc.M113.485755. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hespeels B et al. Back to the roots, desiccation and radiation resistances are ancestral characters in bdelloid rotifers. BMC Biol. 2023:21:72. 10.1186/s12915-023-01554-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hu R et al. Adaptive evolution of the enigmatic Takakia now facing climate change in Tibet. Cell. 2023:186:3558–3576.e17. 10.1016/j.cell.2023.07.003. [DOI] [PubMed] [Google Scholar]
- Huang C et al. Porphobilinogen deaminase HEMC interacts with the PPR-protein AtECB2 for chloroplast RNA editing. Plant J. 2017:92:546–556. 10.1111/tpj.13672. [DOI] [PubMed] [Google Scholar]
- Ichinose M, Sugita M. RNA editing and its molecular mechanism in plant organelles. Genes (Basel). 2016:8:5. 10.3390/genes8010005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jain SC, Shinde U, Li Y, Inouye M, Berman HM. The crystal structure of an autoprocessed Ser221Cys-subtilisin E-propeptide complex at 2.0 Å resolution. J Mol Biol. 1998:284:137–144. 10.1006/jmbi.1998.2161. [DOI] [PubMed] [Google Scholar]
- Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013:30:772–780. 10.1093/molbev/mst010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019:37:907–915. 10.1038/s41587-019-0201-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Knoop V. C-to-U and U-to-C: RNA editing in plant organelles and beyond. J Exp Bot. 2023:74:2273–2294. 10.1093/jxb/erac488. [DOI] [PubMed] [Google Scholar]
- Knoop V, Rüdinger M. DYW-type PPR proteins in a heterolobosean protist: plant RNA editing factors involved in an ancient horizontal gene transfer? FEBS Lett. 2010:584:4287–4291. 10.1016/j.febslet.2010.09.041. [DOI] [PubMed] [Google Scholar]
- Kwok van der Giezen FM et al. High conservation of translation-enabling RNA editing sites in hyper-editing ferns implies they are not selectively neutral. Mol Biol Evol. 2025:42:msaf241. 10.1093/molbev/msaf241. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Letunic I, Bork P. Interactive tree of life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool. Nucleic Acids Res. 2024:52:W78–W82. 10.1093/nar/gkae268. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Leu K-C, Hsieh M-H, Wang H-J, Hsieh H-L, Jauh G-Y. Distinct role of Arabidopsis mitochondrial P-type pentatricopeptide repeat protein-modulating editing protein, PPME, in nad1 RNA editing. RNA Biol. 2016:13:593–604. 10.1080/15476286.2016.1184384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li B, Dewey CN. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011:12:323. 10.1186/1471-2105-12-323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li C et al. Extraordinary preservation of gene collinearity over three hundred million years revealed in homosporous lycophytes. Proc Natl Acad Sci U S A. 2024:121:e2312607121. 10.1073/pnas.2312607121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li F-W et al. Anthoceros genomes illuminate the origin of land plants and the unique biology of hornworts. Nat Plants. 2020:6:259–272. 10.1038/s41477-020-0618-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li J, Yuan J, Jing Y, Lin R. MORF proteins: a small family regulating organellar RNA editing and beyond. J Integr Plant Biol. 2025a:67:2532–2544. 10.1111/jipb.13967. [DOI] [PubMed] [Google Scholar]
- Li J et al. The pentatricopeptide repeat protein OsPPR674 regulates rice growth and drought sensitivity by modulating RNA editing of the mitochondrial transcript ccmC. Int J Mol Sci. 2025b:26:2646. 10.3390/ijms26062646. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li M et al. Plant editosome database: a curated database of RNA editosome in plants. Nucleic Acids Res. 2019:47:D170–D174. 10.1093/nar/gky1026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li Y, Hu Z, Jordan F, Inouye M. Functional analysis of the propeptide of subtilisin E as an intramolecular chaperone for protein folding: refolding and inhibitory abilities of propeptide mutants. J Biol Chem. 1995:270:25127–25132. 10.1074/jbc.270.42.25127. [DOI] [PubMed] [Google Scholar]
- Liang Z, Schnable JC. Functional divergence between subgenomes and gene pairs after whole genome duplications. Mol Plant. 2018:11:388–397. 10.1016/j.molp.2017.12.010. [DOI] [PubMed] [Google Scholar]
- Liu P-Z et al. Genome-wide identification and expression analysis of the MORF gene family in celery reveals their potential role in chloroplast development. J Genet Eng Biotechnol. 2024:22:100443. 10.1016/j.jgeb.2024.100443. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Luo M et al. Functional divergence and origin of the DAG-like gene family in plants. Sci Rep. 2017:7:5688. 10.1038/s41598-017-05961-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Luo Z et al. Pentatricopeptide repeat gene-mediated mitochondrial RNA editing impacts on rice drought tolerance. Front Plant Sci. 2022:13:926285. 10.3389/fpls.2022.926285. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lurin C et al. Genome-wide analysis of Arabidopsis pentatricopeptide repeat proteins reveals their essential role in organelle biogenesis. Plant Cell. 2004:16:2089–2103. 10.1105/tpc.104.022236. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Manna S. An overview of pentatricopeptide repeat proteins and their applications. Biochimie. 2015:113:93–99. 10.1016/j.biochi.2015.04.004. [DOI] [PubMed] [Google Scholar]
- Manna S, Barth C. Identification of a novel pentatricopeptide repeat subfamily with a C-terminal domain of bacterial origin acquired via ancient horizontal gene transfer. BMC Res Notes. 2013:6:525. 10.1186/1756-0500-6-525. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Manna S, Brewster J, Barth C. Identification of pentatricopeptide repeat proteins in the model organism Dictyostelium discoideum. Int J Genomics. 2013:2013:586498. 10.1155/2013/586498. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Manni M, Berkeley MR, Seppey M, Simão FA, Zdobnov EM. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021:38:4647–4654. 10.1093/molbev/msab199. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Marcotte EM, Pellegrini M, Thompson MJ, Yeates TO, Eisenberg D. A combined algorithm for genome-wide prediction of protein function. Nature. 1999a:402:83–86. 10.1038/47048. [DOI] [PubMed] [Google Scholar]
- Marcotte EM et al. Detecting protein function and protein-protein interactions from genome sequences. Science. 1999b:285:751–753. 10.1126/science.285.5428.751. [DOI] [PubMed] [Google Scholar]
- Meng EC et al. UCSF ChimeraX: tools for structure building and analysis. Protein Sci. 2023:32:e4792. 10.1002/pro.4792. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meng L et al. PPR proteins in plants: roles, mechanisms, and prospects for rice research. Front Plant Sci. 2024:15:1416742. 10.3389/fpls.2024.1416742. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Minh BQ et al. IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol. 2020:37:1530–1534. 10.1093/molbev/msaa015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mirarab S et al. ASTRAL: genome-scale coalescent-based species tree estimation. Bioinformatics. 2014:30:i541–i548. 10.1093/bioinformatics/btu462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mirdita M et al. ColabFold: making protein folding accessible to all. Nat Methods. 2022:19:679–682. 10.1038/s41592-022-01488-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mistry J et al. Pfam: the protein families database in 2021. Nucleic Acids Res. 2021:49:D412–D419. 10.1093/nar/gkaa913. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ødum MT et al. DeepLoc 2.1: multi-label membrane protein type prediction using protein language models. Nucleic Acids Res. 2024:52:W215–W220. 10.1093/nar/gkae237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ou S, Jiang N. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 2018:176:1410–1422. 10.1104/pp.17.01310. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ou S, Jiang N. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mob DNA. 2019:10:48. 10.1186/s13100-019-0193-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Panchy N, Lehti-Shiu M, Shiu S-H. Evolution of gene duplication in plants. Plant Physiol. 2016:171:2294–2316. 10.1104/pp.16.00523. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pfalz J, Bayraktar OA, Prikryl J, Barkan A. Site-specific binding of a PPR protein defines and stabilizes 5′ and 3′ mRNA termini in chloroplasts. EMBO J. 2009:28:2042–2052. 10.1038/emboj.2009.121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Price MN, Dehal PS, Arkin AP. FastTree 2-approximately maximum-likelihood trees for large alignments. PLoS One. 2010:5:e9490. 10.1371/journal.pone.0009490. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Prikryl J, Rojas M, Schuster G, Barkan A. Mechanism of RNA stabilization and translational activation by a pentatricopeptide repeat protein. Proc Natl Acad Sci U S A. 2011:108:415–420. 10.1073/pnas.1012076108. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Procter JB et al. Alignment of biological sequences with Jalview. Springer US; 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Puttick MN et al. The interrelationships of land plants and the nature of the ancestral embryophyte. Curr Biol. 2018:28:733–745.e2. 10.1016/j.cub.2018.01.063. [DOI] [PubMed] [Google Scholar]
- Rüdinger M, Fritz-Laylin L, Polsakiewicz M, Knoop V. Plant-type mitochondrial RNA editing in the protist Naegleria gruberi. RNA. 2011:17:2058–2062. 10.1261/rna.02962911. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sahu SS, Loaiza CD, Kaundal R. Plant-mSubP: a computational framework for the prediction of single-and multi-target protein subcellular localization using integrated machine-learning approaches. AoB Plants. 2020:12:plz068. 10.1093/aobpla/plz068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sandoval R et al. Stable native RIP9 complexes associate with C-to-U RNA editing activity, PPRs, RIPs, OZ1, ORRM1 and ISE2. Plant J. 2019:99:1116–1126. 10.1111/tpj.14384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schaap P et al. The Physarum polycephalum genome reveals extensive use of prokaryotic two-component and metazoan-type tyrosine kinase signaling. Genome Biol Evol. 2015:8:109–125. 10.1093/gbe/evv237. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schafran P et al. Pan-phylum genomes of hornworts reveal conserved autosomes but dynamic accessory and sex chromosomes. Nat Plants. 2025:11:49–62. 10.1038/s41477-024-01883-w. [DOI] [PubMed] [Google Scholar]
- Schallenberg-Rüdinger M, Lenz H, Polsakiewicz M, Gott JM, Knoop V. A survey of PPR proteins identifies DYW domains like those of land plant RNA editing factors in diverse eukaryotes. RNA Biol. 2013:10:1549–1556. 10.4161/rna.25755. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schrader L, Schmitz J. The impact of transposable elements in adaptive evolution. Mol Ecol. 2019:28:1537–1549. 10.1111/mec.14794. [DOI] [PubMed] [Google Scholar]
- Searing AM, Satyanarayan MB, O’ Donnell JP, Lu Y. Two organelle RNA recognition motif proteins affect distinct sets of RNA editing sites in the Arabidopsis thaliana plastid. Plant Direct. 2020:4:e00213. 10.1002/pld3.213. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shen Y et al. Chromosome-level and haplotype-resolved genome provides insight into the tetraploid hybrid origin of patchouli. Nat Commun. 2022:13:3511. 10.1038/s41467-022-31121-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shi X, Bentolila S, Hanson MR. Organelle RNA recognition motif-containing (ORRM) proteins are plastid and mitochondrial editing factors in Arabidopsis. Plant Signal Behav. 2016a:11:e1167299. 10.1080/15592324.2016.1167299. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shi X, Germain A, Hanson MR, Bentolila S. RNA recognition motif-containing protein ORRM4 broadly affects mitochondrial RNA editing and impacts plant development and flowering. Plant Physiol. 2016b:170:294–309. 10.1104/pp.15.01280. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shi X, Hanson MR, Bentolila S. Two RNA recognition motif-containing proteins are plant mitochondrial editing factors. Nucleic Acids Res. 2015:43:3814–3825. 10.1093/nar/gkv245. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shikanai T. RNA editing in plant organelles: machinery, physiological function and evolution. Cell Mol Life Sci. 2006:63:698–708. 10.1007/s00018-005-5449-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Small ID, Schallenberg-Rüdinger M, Takenaka M, Mireau H, Ostersetzer-Biran O. Plant organellar RNA editing: what 30 years of research has revealed. Plant J. 2020:101:1040–1056. 10.1111/tpj.14578. [DOI] [PubMed] [Google Scholar]
- Sperschneider J et al. LOCALIZER: subcellular localization prediction of both plant and effector proteins in the plant cell. Sci Rep. 2017:7:44598. 10.1038/srep44598. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Storer J, Hubley R, Rosen J, Wheeler TJ, Smit AF. The Dfam community resource of transposable element families, sequence models, and genome annotations. Mob DNA. 2021:12:2. 10.1186/s13100-020-00230-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sugita M, Uchiyama H, Ichinose M. RRM and PPR proteins network as a global regulator for plastid RNA metabolism. Endocytobiosis Cell Res. 2016:27:41–44. [Google Scholar]
- Sun P et al. WGDI: a user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol Plant. 2022:15:1841–1851. 10.1016/j.molp.2022.10.018. [DOI] [PubMed] [Google Scholar]
- Sun T, Bentolila S, Hanson MR. The unexpected diversity of plant organelle RNA editosomes. Trends Plant Sci. 2016:21:962–973. 10.1016/j.tplants.2016.07.005. [DOI] [PubMed] [Google Scholar]
- Sun T et al. An RNA recognition motif-containing protein is required for plastid RNA editing in Arabidopsis and maize. Proc Natl Acad Sci U S A. 2013:110:E1169–E1178. 10.1073/pnas.1220162110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sun T et al. A zinc finger motif-containing protein is essential for chloroplast RNA editing. PLoS Genet. 2015:11:e1005028. 10.1371/journal.pgen.1005028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sun YK, Gutmann B, Yap A, Kindgren P, Small I. Editing of chloroplast rps14 by PPR editing factor EMB2261 is essential for Arabidopsis development. Front Plant Sci. 2018:9:841. 10.3389/fpls.2018.00841. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Takenaka M, Zehrmann A, Verbitskiy D, Haertel B, Brennicke A. RNA editing in plants and its evolution. Annu Rev Genet. 2013:47:335–352. 10.1146/annurev-genet-111212-133519. [DOI] [PubMed] [Google Scholar]
- Takenaka M et al. Multiple organellar RNA editing factor (MORF) family proteins are required for RNA editing in mitochondria and plastids of plants. Proc Natl Acad Sci U S A. 2012:109:5104–5109. 10.1073/pnas.1202452109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Takenaka M et al. DYW domain structures imply an unusual regulation principle in plant organellar RNA editing catalysis. Nat Catal. 2021:4:510–522. 10.1038/s41929-021-00633-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tamura K, Stecher G, Kumar S. MEGA11: molecular evolutionary genetics analysis version 11. Mol Biol Evol. 2021:38:3022–3027. 10.1093/molbev/msab120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang D et al. Genome-wide analysis of multiple organellar RNA editing factor family in poplar reveals evolution and roles in drought stress. Int J Mol Sci. 2019a:20:1425. 10.3390/ijms20061425. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang L, Wang S, Li W. RSeQC: quality control of RNA-seq experiments. Bioinformatics. 2012a:28:2184–2185. 10.1093/bioinformatics/bts356. [DOI] [PubMed] [Google Scholar]
- Wang T, Takenaka M. The molecular basis and evolution of the organellar RNA editosome by complementary DYW deaminases in seed plants. Plant Physiol. 2025:197:kiaf142. 10.1093/plphys/kiaf142. [DOI] [PubMed] [Google Scholar]
- Wang W, Yang Y, Örstan A, He Z, Wang Q. Species diversity of bdelloid rotifer (Rotifera, Bdelloidea) in different areas in China, with a description of two new species. ZooKeys. 2025a:1229:1–23. 10.3897/zookeys.1229.113385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang X et al. A recent burst of gene duplications in Triticeae. Plant Commun. 2022:3:100268. 10.1016/j.xplc.2021.100268. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Y, Tan B-C. Pentatricopeptide repeat proteins in plants: cellular functions, action mechanisms, and potential applications. Plant Commun. 2025:6:101203. 10.1016/j.xplc.2024.101203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Y et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 2012b:40:e49. 10.1093/nar/gkr1293. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Y et al. Maize PPR-E proteins mediate RNA C-to-U editing in mitochondria by recruiting the trans deaminase PCW1. Plant Cell. 2023:35:529–551. 10.1093/plcell/koac298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Y et al. Multiple factors interact in editing of PPR-E+-targeted sites in maize mitochondria and plastids. Plant Commun. 2024:5:100836. 10.1016/j.xplc.2024.100836. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Y et al. Empty Pericarp21 encodes a novel PPR-DYW protein that is required for mitochondrial RNA editing at multiple sites, complexes I and V biogenesis, and seed development in maize. PLoS Genet. 2019b:15:e1008305. 10.1371/journal.pgen.1008305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang Z et al. Near-complete assembly and comprehensive annotation of the wheat Chinese spring genome. Mol Plant. 2025b:18:892–907. 10.1016/j.molp.2025.02.002. [DOI] [PubMed] [Google Scholar]
- Wang Z et al. OsOZ1, a zinc finger motif-containing protein, is essential for chloroplast development through RNA editing and intron splicing in rice. Plant Growth Regul. 2025c:105:449–461. 10.1007/s10725-025-01283-w. [DOI] [Google Scholar]
- Wu W et al. Genomic convergence in terrestrial root plants through tandem duplication in response to soil microbial pressures. Cell Rep. 2024:43:114786. 10.1016/j.celrep.2024.114786. [DOI] [PubMed] [Google Scholar]
- Xu Z, Wang H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007:35:W265–W268. 10.1093/nar/gkm286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Xu Z et al. Crocus genome reveals the evolutionary origin of crocin biosynthesis. Acta Pharm Sin B. 2024:14:1878–1891. 10.1016/j.apsb.2023.12.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yan J, Zhang Q, Yin P. RNA editing machinery in plant organelles. Sci China Life Sci. 2018:61:162–169. 10.1007/s11427-017-9170-3. [DOI] [PubMed] [Google Scholar]
- Yan J et al. MORF9 increases the RNA-binding activity of PLS-type pentatricopeptide repeat protein in plastid RNA editing. Nat Plants. 2017:3:17037. 10.1038/nplants.2017.37. [DOI] [PubMed] [Google Scholar]
- Yang H et al. Rice FLOURY ENDOSPERM22, encoding a pentatricopeptide repeat protein, is involved in both mitochondrial RNA splicing and editing and is crucial for endosperm development. J Integr Plant Biol. 2023a:65:755–771. 10.1111/jipb.13402. [DOI] [PubMed] [Google Scholar]
- Yang J, Harding T, Kamikawa R, Simpson AG, Roger AJ. Mitochondrial genome evolution and a novel RNA editing system in deep-branching heteroloboseids. Genome Biol Evol. 2017a:9:1161–1174. 10.1093/gbe/evx086. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang Y et al. DYW cytidine deaminase domains have a long-range impact on RNA recognition by the PPR array of chimeric plant C-to-U RNA editing factors and strongly affect target selection. Plant J. 2023b:116:840–854. 10.1111/tpj.16412. [DOI] [PubMed] [Google Scholar]
- Yang Y et al. The RNA editing factor SlORRM4 is required for normal fruit ripening in tomato. Plant Physiol. 2017b:175:1690–1702. 10.1104/pp.17.01265. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang Y-Z et al. GRP23 plays a core role in E-type editosomes via interacting with MORFs and atypical PPR-DYWs in Arabidopsis mitochondria. Proc Natl Acad Sci U S A. 2022:119:e2210978119. 10.1073/pnas.2210978119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yuan H, Liu D. Functional disruption of the pentatricopeptide protein SLG1 affects mitochondrial RNA editing, plant development, and responses to abiotic stresses in Arabidopsis. Plant J. 2012:70:432–444. 10.1111/j.1365-313X.2011.04883.x. [DOI] [PubMed] [Google Scholar]
- Zhang C, Gschwend AR, Ouyang Y, Long M. Evolution of gene structural complexity: an alternative-splicing-based model accounts for intron-containing retrogenes. Plant Physiol. 2014:165:412–423. 10.1104/pp.113.231696. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang J et al. The hornwort genome and early land plant evolution. Nat Plants. 2020:6:107–118. 10.1038/s41477-019-0588-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang S et al. The GSA family in 2025: a broadened sharing platform for multi-omics and multimodal data. Genomics Proteomics Bioinformatics. 2025:23:qzaf072. 10.1093/gpbjnl/qzaf072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z. Kaks_calculator 3.0: calculating selective pressure on coding and non-coding sequences. Genomics Proteomics Bioinformatics. 2022:20:536–540. 10.1016/j.gpb.2021.12.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Z et al. ParaAT: a parallel tool for constructing multiple protein-coding DNA alignments. Biochem Biophys Res Commun. 2012:419:779–781. 10.1016/j.bbrc.2012.02.101. [DOI] [PubMed] [Google Scholar]
- Zhao B et al. A six-repeat PPR protein WPR directly binds target RNAs and coordinates chloroplast RNA processing via dual recruitment of MORF1, MORF8b, and CAF2 proteins in rice. Plant J. 2025:124:e70526. 10.1111/tpj.70526. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data used and generated in this study are available at the Open Archive for Miscellaneous Data (https://ngdc.cncb.ac.cn/omix, OMIX: OMIX012585) (Zhang et al. 2025) in the National Genomics Data Center (NGDC), China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences (CNCB-NGDC Members and Partners 2025). The computational codes used in this study have been deposited in GitHub (https://github.com/Bioinfo-Ming/PlantEditosomeEvolution).







