SUMMARY
Agrobacterium transfers DNA into plant cells, leading to tumors, hairy roots (HR), and natural genetically modified organisms (nGMOs). Transferred DNAs (T‐DNAs) from agrobacteria and T‐DNA‐derived cellular T‐DNAs (cT‐DNAs) from nGMOs vary considerably and may carry up to 15 different genes. Among these, opine synthase (ops) genes encode the synthesis of opines used as nutrients by the agrobacteria. Earlier studies predicted large numbers of naturally transformed plant species, but only few have been identified and studied so far. We therefore developed a general method to detect cT‐DNAs in all publicly available whole genome sequences (WGS) and Sequence Read Archive (SRA) data from land plants. To avoid false positives, we only retained DNA sequences coding for T‐DNA proteins. A total of 2614 nGMO species were identified, most are eudicots. However, cT‐DNAs were also found in 82 mosses and 75 ferns, showing that Agrobacterium can also generate natural transformants among the early land plants. Analysis of 149 cT‐DNA maps revealed different types of T‐DNAs. Most notably, these included small T‐DNAs (mini T‐DNAs) with a single opine synthase gene. Mini T‐DNAs are not expected to induce tumors or HRs. The predominance of mini cT‐DNAs in mosses and ferns, and the presence of more complex cT‐DNAs in spermatophytes, indicate that mini T‐DNAs represent the earliest types of T‐DNA. Our study also detected unusual T‐DNA integration patterns, with multiple copies spread out over several hundreds of kilobases.
Keywords: horizontal gene transfer, nGMOs, Agrobacterium , mini T‐DNAs
Significance Statement
Horizontal gene transfer from Agrobacterium has contributed to the genomes of numerous plant species, but its frequency and evolutionary history across land plants have remained poorly understood. By systematically mining whole‐genome and Sequence Read Archive data across embryophytes, we identified 2614 natural genetically modified plant species, greatly expanding the known diversity of natural Agrobacterium‐mediated transformation. Importantly, we detected cellular T‐DNAs in mosses, liverworts, hornworts, and ferns, indicating that Agrobacterium‐mediated transformation is not restricted to flowering plants and may date to early land plant evolution. Comparative analysis further revealed a widespread class of mini T‐DNAs, often carrying a single opine synthase gene, suggesting that these simple T‐DNAs may represent an ancient form of T‐DNA and providing new insight into the evolution of the Agrobacterium–plant interaction. This work establishes a broad framework for investigating the evolutionary, ecological, and functional significance of natural T‐DNA integration in plants.
INTRODUCTION
Agrobacterium‐mediated horizontal gene transfer (HGT) involves a unique mechanism of DNA transfer, leading to stable insertion in the nuclear genome. The transferred DNA (T‐DNA) typically results in hairy roots (HR) or crown gall tumors (Gelvin, 2017; Hooykaas, 2023; Nester, 2014; Zhu et al., 2000). T‐DNAs can carry from one up to 15 genes (Otten, 2021; Weisberg et al., 2020), which encode growth induction and modification, and opine synthesis (opine synthase or ops genes). Opines are small molecules used by agrobacteria as nutrients. It has been proposed that the induction of opine synthesis is the driving force behind the evolution of the T‐DNA transfer mechanism (Petit & Tempé, 1985). A number of plant species contain T‐DNAs (cellular T‐DNAs or cT‐DNAs) and are called natural genetically modified organisms (nGMOs) or natural transformants. Among these, several Nicotiana species (Furner et al., 1986; White et al., 1983), common toadflax (Linaria vulgaris; Matveeva et al., 2012), and Ipomoea (Kyndt et al., 2015) have been reported. A first systematic study using whole genome sequences (WGS) detected 23 nGMO species (Matveeva & Otten, 2019), and by extrapolation predicted about 10 000 naturally transformed plant species. Subsequent studies investigated four genera in detail: Camellia (Chen et al., 2022, 2023), Vaccinium (Zhidkin et al., 2023), Arachis (Bogomaz et al., 2024), and Diospyros (Otten et al., 2025). These studies, contrary to the 2019 study, included SRA and WGS data (SRA+WGS search). Although SRA data are less complete than WGS data, they can reveal significantly more nGMOs than WGS data alone. We therefore adopted the SRA+WGS approach for a search targeting all of the embryophytes (14 000 instead of four genera). Comparative analysis from the vast amount of new data allowed us to gain more insight into the history of Agrobacterium‐mediated HGT events, most notably through the discovery of small T‐DNAs with a single T‐DNA gene (mini T‐DNAs) and their distribution across vascular and nonvascular plants.
RESULTS
Searching cT‐DNAs in WGS and SRA data from land plants
The first systematic search for nGMOs (Matveeva & Otten, 2019) in angiosperms used WGS data, but no SRA data. Presently there are 3930 Terabytes of SRA data for land plants. We therefore devised a method to extract putative land plant cT‐DNA sequences from both WGS and SRA data (‘Materials and Methods’ section; Figure 1), and retained those corresponding to protein‐coding regions, using a set of T‐DNA/cT‐DNA proteins (Table S1) as query. Particular cT‐DNA inserts will be named by using the first two letters of the genus in which they were first discovered, followed by T (for cT‐DNA) and a letter for the insert, starting with A. Thus, CaTA is the first cT‐DNA insert (A) discovered in Camellia, DiTD the fourth one (D) from Diospyros. cT‐DNA alleles are indicated by a number (like DiTD‐2). Older published names, like IbT‐DNA1 (for Ipomoea batatas T‐DNA 1) are kept to avoid confusion. WGS contig numbers are from the National Center for Biotechnology Information (NCBI), but with simplified numbering. For example, JBNGMI010000139.1 becomes JBNGMI‐139.
Figure 1.

Overview of the process of identifying new natural genetically modified organisms (nGMOs).
All available land plant whole genome sequences (WGS) and Sequence Read Archive (SRA) datasets were searched using T‐DNA nucleotide (NTfuse6) and protein (Table S1) queries. Positive reads were assembled into contigs and subjected to BLASTX analysis to identify cT‐DNA genes. Potential nGMOs were further filtered to exclude false positives (e.g., hairy root lines, distant bacterial and plant homologs). In addition, to exclude potential Agrobacterium contamination, all positive WGS and SRA accessions were further screened using queries based on several Agrobacterium virulence (vir) genes, following the approach described in Matveeva and Otten (2019). Verified cases were compared with published nGMOs to distinguish newly identified nGMOs from old ones. cT‐DNA maps from old nGMOs and T‐DNA maps from Rhizobium/Agrobacterium were redrawn at the same format (Figure S2A). Newly identified nGMOs were mapped from WGS data (Figure S2B). SRA data were summarized according to their T‐DNA genes (Table S2). Because of the very large number of publicly available SRA data for monocots, Fabaceae, Solanaceae, and Brassicaceae, a maximum of 20 SRR accessions per species were selected for analysis, the Dioscoreaceae family was analyzed using all available SRR accessions.
Nonvascular embryophytes
Nonvascular embryophytes consist of Bryophyta (mosses, 12 700 species), Anthocerotophyta (hornworts, 225), and Marchantiophyta (liverworts, 9000). Species numbers in this paper are from Christenhusz & Byng, (2016). Analysis of the corresponding WGS and SRA data from the three groups yielded 77, one and four nGMOs, respectively. These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table 1. All, except four, contain a cucumopine synthase (cus) gene. WGS sequences from Ptychostomum pallens, Polytrichum commune, and Aerobryopsis subdivergens show direct cT‐DNA repeats, separated by 12, 12, and 32 kb, respectively. Hypopterygium flavolimbatum cT‐DNA HyTA (JBNGMI‐139:498 534–505 091, 6558 nt) could be defined by comparison with JBNGMJ from the cT‐DNA‐less Hypopterygium elatum. JBNGMI‐139:494 797–498 533 aligns with JBNGMJ‐4450:251 913–255 682, and JBNGMI‐139:505 092–505 245 with JBNGMJ‐4450:257 557–257 708. Thus, HyTA caused a 1875 nt deletion upon insertion.
Table 1.
nGMO species from nonvascular embryophytes (A: Anthocerotophyta, B: Bryophyta, M: Marchantiophyta), and ferns (L: Lycopodiophyta and P: Pteridophyta). Data are from WGS and SRA data
| Order | Family | Genus and species | Genes | WGS/SRR | |||
|---|---|---|---|---|---|---|---|
| 1 | Ferns | L | Isoetales | Isoetacea | Isoetes valida | ags | SRR25755400, SRR25755356 |
| 2 | Ferns | L | Lycopodiales | Lycopodiaceae | Huperzia selago | a | SRR28958735, SRR28958738 |
| 3 | Ferns | L | Lycopodiales | Lycopodiaceae | Huperzia serrata | b | SRR28958736, SRR18149617 |
| 4 | Ferns | L | Selaginellales | Selaginellaceae | Selaginella erythropus | mas2′, mas1′, ags | SRR15348965, SRR15348964, SRR15348963 |
| 5 | Ferns | L | Selaginellales | Selaginellaceae | Selaginella martensii | vis | SRR31771695, SRR31771696, SRR31771697 |
| 6 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium alternifolium | susL | ERR14012285 |
| 7 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium komarovii | mas1′, ags | SRR14381415, SRR8772343 |
| 8 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium trichomanes subsp. trichomanes | susL | ERR14009220, ERR5554902, ERR5232099 |
| 9 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium trichomanes subsp. inexpectans | susL | ERR5529585 |
| 10 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium trichomanes subsp. quadrivalens | susL | ERR14012664 |
| 11 | Ferns | P | Polypodiales | Aspleniaceae | Asplenium × adulterinum | susL | ERR14047057, ERR5529300 |
| 12 | Ferns | P | Polypodiales | Aspleniaceae | Hymenasplenium sp. | mas1′ | SRR6920712 |
| 13 | Ferns | P | Polypodiales | Athyriaceae | Deparia giraldii | susL | SRR31694523 |
| 14 | Ferns | P | Polypodiales | Athyriaceae | Deparia lancea | ags | SRR6899396 |
| 15 | Ferns | P | Polypodiales | Athyriaceae | Diplazium unilobum | ags | SRR16974227 |
| 16 | Ferns | P | Polypodiales | Blechnaceae | Blechnopsis orientalis | cus | SRR28681840, SRR27869072 |
| 17 | Ferns | P | Polypodiales | Blechnaceae | Stenochlaena palustris | vis | ERR12670167, ERR12670177 |
| 18 | Ferns | P | Cyatheales | Cyatheaceae | Alsophila costularia | ags | SRR27937298 |
| 19 | Ferns | P | Cyatheales | Cyatheaceae | Alsophila latebrosa | vis, ags | SRR19887278 |
| 20 | Ferns | P | Cyatheales | Cyatheaceae | Alsophila metteniana | ags | SRR19887263 |
| 21 | Ferns | P | Cyatheales | Cyatheaceae | Alsophila spinulosa | c, ags, orf2, rolA | SRR28161632, SRR28161616, SRR28161618 |
| 22 | Ferns | P | Cyatheales | Cyatheaceae | Cyathea glabra | ags | SRR19887387 |
| 23 | Ferns | P | Cyatheales | Cyatheaceae | Gymnosphaera | acs, ags | SRR11076226, SRR11075945, SRR19887515 |
| 24 | Ferns | P | Cyatheales | Cyatheaceae | Gymnosphaera andersonii | acs | SRR19887516 |
| 25 | Ferns | P | Polypodiales | Davalliaceae | Davallia denticulata | cus | ERR12670369 |
| 26 | Ferns | P | Polypodiales | Dennstaedtiaceae | Dennstaedtia scandens | acs | SRR22250942 |
| 27 | Ferns | P | Cyatheales | Dicksoniaceae | Dicksonia lanata | ocs | SRR8580726 |
| 28 | Ferns | P | Polypodiales | Dryopteridaceae | Arachniodes davalliiformis | cus | DRR591047 |
| 29 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris affinis | e | ERR14042545 |
| 30 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris campyloptera | vis | SRR9050854 |
| 31 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris celsa | cus | SRR14320985 |
| 32 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris crassirhizoma | acs | SRR13447702 |
| 33 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris filix‐mas | e | SRR12518786 |
| 34 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris intermedia | ocs | SRR9050845 |
| 35 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris keysseriana | acs | SRR18497099 |
| 36 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris marginalis | mis,e | SRR11229720 |
| 37 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris pseudocaenopteris | mas2′ | SRR2103701 |
| 38 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris remota | vis | ERR14047111 |
| 39 | Ferns | P | Polypodiales | Dryopteridaceae | Dryopteris tyrrhena | vis | ERR14041301 |
| 40 | Ferns | P | Polypodiales | Dryopteridaceae | Lomagramma matthewii | susL | SRR2103704 |
| 41 | Ferns | P | Polypodiales | Dryopteridaceae | Polystichum acrostichoides | vis | SRR18053988, JAOYMV01191549 |
| 42 | Ferns | P | Polypodiales | Dryopteridaceae | Polystichum aculeatum | mas1′ | ERR14030239 |
| 43 | Ferns | P | Polypodiales | Dryopteridaceae | Rumohra adiantiformis | mas1′, ags | SRR30618658 |
| 44 | Ferns | P | Hymenophyllales | Hymenophyllaceae | Abrodictyum obscurum | orf14, orf511, susL | ERR12670379, ERR12670380, SRR32258258 |
| 45 | Ferns | P | Schizaeales | Lygodiaceae | Lygodium japonicum | cus | SRR29127758 |
| 46 | Ferns | P | Schizaeales | Lygodiaceae | Lygodium palmatum | orf14 | SRR26863426 |
| 47 | Ferns | P | Schizaeales | Lygodiaceae | Lygodium salicifolium | acs | SRR7121778 |
| 48 | Ferns | P | Maratttiales | Marattiaceae | Danaea nodosa | mas1′, mas2′ | ERR2041196 |
| 49 | Ferns | P | Polypodiales | Onocleaceae | Onoclea sensibilis var. interrupta | acs | SRR31693313 |
| 50 | Ferns | P | Polypodiales | Polypodiaceae | Goniophlebium amoenum | cus | SRR8185331 |
| 51 | Ferns | P | Polypodiales | Polypodiaceae | Lemmaphyllum drymoglossoides | acs | SRR23238769 |
| 52 | Ferns | P | Polypodiales | Polypodiaceae | Loxogramme chinensis | mas2′ | SRR2103729 |
| 53 | Ferns | P | Polypodiales | Polypodiaceae | Platycerium wallichii | susL | SRR14802567 |
| 54 | Ferns | P | Polypodiales | Polypodiaceae | Pyrrosia piloselloides | ocs | SRR26086252 |
| 55 | Ferns | P | Polypodiales | Pteridaceae | Adiantum nelumboides | see map 147 to 149 | JAKNSL020005917.1 |
| 56 | Ferns | P | Polypodiales | Pteridaceae | Adiantum reniforme | mas1′, mas2′ | SRR32258808 |
| 57 | Ferns | P | Polypodiales | Pteridaceae | Antrophyum callifolium | susL | SRR2103739 |
| 58 | Ferns | P | Polypodiales | Pteridaceae | Antrophyum formosanum | susL | SRR32258181, SRR32212872 |
| 59 | Ferns | P | Polypodiales | Pteridaceae | Ceratopteris pteridoides | cus | SRR23096549 |
| 60 | Ferns | P | Polypodiales | Pteridaceae | Ceratopteris richardii | vis, cus | JAIKUY010002021, 2022, 938 |
| 61 | Ferns | P | Polypodiales | Pteridaceae | Ceratopteris thalictroides | vis | SRR16588741 |
| 62 | Ferns | P | Polypodiales | Pteridaceae | Cyrtomium falcatum | vis | SRR6727965 |
| 63 | Ferns | P | Polypodiales | Pteridaceae | Haplopteris elongata | cus | SRR14048925 |
| 64 | Ferns | P | Polypodiales | Pteridaceae | Haplopteris ensiformis | cus | SRR20678395 |
| 65 | Ferns | P | Polypodiales | Pteridaceae | Haplopteris heterophylla | cus | SRR6920718 |
| 66 | Ferns | P | Polypodiales | Pteridaceae | Hemionitis arifolia | mas1′ | SRR2103738 |
| 67 | Ferns | P | Polypodiales | Pteridaceae | Parahemionitis cordata | mas1′ | ERR2040933 |
| 68 | Ferns | P | Polypodiales | Pteridaceae | Pteris biaurita | mas1′, mas2′ | SRR7121768 |
| 69 | Ferns | P | Salviniales | Salviniaceae | Salvinia molesta | mas1′, ags | SRR23404241 |
| 70 | Ferns | P | Schizaeales | Schizaeaceae | Actinostachys digitata | cus, susL, orf14, orf511 | ERR12670432, ERR12670423, ERR12670425 |
| 71 | Ferns | P | Polypodiales | Tectariaceae | Tectaria polymorpha | vis | SRR2103745 |
| 72 | Ferns | P | Polypodiales | Thelypteridaceae | Abacopteris gymnopteridifrons | vis | SRR22806495 |
| 73 | Ferns | P | Polypodiales | Thelypteridaceae | Amblovenatum opulentum | vis | SRR26157139, SRR26157140 |
| 74 | Ferns | P | Polypodiales | Thelypteridaceae | Grypothrix megacuspis | acs | SRR18496617 |
| 75 | Ferns | P | Polypodiales | Thelypteridaceae | Gymnocarpium robertianum | orf8 | ERR14029864 |
| 76 | Ferns | P | Polypodiales | Thelypteridaceae | Thelypteris parasitica | vis | SRR7121612 |
| 77 | Ferns | P | Polypodiales | Woodsiaceae | Matteuccia struthiopteris | vis | SRR31694452 |
| 78 | Mosses | A | Anthocerotales | Anthocerotaceae | Folioceros fuciformis | mas1 | SRR29281613 |
| 79 | Mosses | B | Hypnales | Amblystegiaceae | Drepanocladus polygamus | cus | SRR26398964 |
| 80 | Mosses | B | Hypnales | Amblystegiaceae | Hygrohypnum luridum | cus | SRR26398954 |
| 81 | Mosses | B | Hypnales | Amblystegiaceae | Sciaromiopsis sinensis | cus | SRR15179250 |
| 82 | Mosses | B | Hypnales | Anomodontaceae | Pseudanomodon attenuatus | cus | SRR33208610, SRR33208612, SRR33208608 |
| 83 | Mosses | B | Bryales | Bartramiaceae | Bartramia ithyphylla | cus | SRR33208738 |
| 84 | Mosses | B | Hypnales | Brachytheciaceae | Brachythecium laetum | cus | SRR33208359, SRR33208361, SRR33208632 |
| 85 | Mosses | B | Hypnales | Brachytheciaceae | Bryhnia novae‐angliae | cus | SRR33208521, SRR33208529, SRR33208528 |
| 86 | Mosses | B | Hypnales | Brachytheciaceae | Kindbergia praelonga | cus | SRR33208255, SRR33208261, SRR33208260 |
| 87 | Mosses | B | Hypnales | Brachytheciaceae | Myuroclada maximowiczii | cus | SRR13605954 |
| 88 | Mosses | B | Bryales | Bryaceae | Anomobryum julaceum | cus | SRR24580469 |
| 89 | Mosses | B | Bryales | Bryaceae | Bryum yuennanense | cus | SRR31709169 |
| 90 | Mosses | B | Bryales | Bryaceae | Ptychostomum cyclophyllum | cus | ERR15381807 |
| 91 | Mosses | B | Bryales | Bryaceae | Ptychostomum pallens | cus | ERR15378477, ERR15381808, OZ377804.1 |
| 92 | Mosses | B | Hypnales | Calliergonaceae | Sarmentypnum sarmentosum | cus | ERR13725965 |
| 93 | Mosses | B | Hypnales | Catagoniaceae | Catagonium nitens | cus | SRR33208480, SRR33208469, SRR33208435 |
| 94 | Mosses | B | Hypnales | Climaciaceae | Climacium americanum | cus | JBNGMP |
| 95 | Mosses | B | Hypnales | Climaciaceae | Climacium dendroides | cus | ERR11242531, ERR10934078, ERR10934077 |
| 96 | Mosses | B | Dicranales | Dicranaceae | Campylopus introflexus | cus | ERR15378467, ERR15381794 |
| 97 | Mosses | B | Dicranales | Ditrichaceae | Pleuridium rhynchostegium | cus | DRR379964 |
| 98 | Mosses | B | Encalyptales | Encalyptaceae | Encalypta ciliata | cus | ERR15659579, SRR33208205 |
| 99 | Mosses | B | Hypnales | Entodontaceae | Entodon concinnus | cus | JBNGNB |
| 100 | Mosses | B | Dicranales | Fissidentaceae | Fissidens adianthoides | cus | ERR15381760 |
| 101 | Mosses | B | Funariales | Funariaceae | Physcomitrellopsis africana | cus | SRR26586950, SRR26596311, SRR26596310 |
| 102 | Mosses | B | Hedwigiales | Hedwigiaceae | Hedwigia ciliata | cus | ERR5232302 |
| 103 | Mosses | B | Hypnales | Hylocomiaceae | Hylocomiadelphus triquetrus | cus | ERR9793174, ERR9793175, ERR9793176 |
| 104 | Mosses | B | Hypnales | Hylocomiaceae | Hylocomium splendens | cus | SRR2518082, SRR25553778, SRR25553779 |
| 105 | Mosses | B | Hypnales | Hylocomiaceae | Loeskeobryum brevirostre | cus | ERR15551479, ERR13731937, ERR13725996 |
| 106 | Mosses | B | Hypnales | Hylocomiaceae | Pleurozium schreberi | cus | SRR2513357 |
| 107 | Mosses | B | Hypnales | Hylocomiaceae | Rhytidiadelphus loreus | cus | ERR6895899, ERR6895898, ERR6895897 |
| 108 | Mosses | B | Hypnales | Hylocomiaceae | Rhytidiadelphus subpinnatus | cus | JBNGMQ |
| 109 | Mosses | B | Hypnales | Hylocomiaceae | Rhytidiopsis robusta | cus | SRR2518096 |
| 110 | Mosses | B | Hypnales | Hypnaceae | Hyocomium armoricum | cus | ERR13389723, ERR13382531 |
| 111 | Mosses | B | Hypopterygiales | Hypopterygiaceae | Hypopterygium fauriei | cus | SRR31755810 |
| 112 | Mosses | B | Hypopterygiales | Hypopterygiaceae | Hypopterygium flavolimbatum | cus | JBNGMI |
| 113 | Mosses | B | Hypnales | Leucodontaceae | Antitrichia curtipendula | cus | SRR2518092 |
| 114 | Mosses | B | Hypnales | Meteoriaceae | Aerobryopsis subdivergens | cus | JBNGMS |
| 115 | Mosses | B | Leucodontales | Meteoriaceae | Barbella flagellifera | cus | SRR23095850 |
| 116 | Mosses | B | Bryales | Mniaceae | Plagiomnium ciliare | cus | SRR33208530 |
| 117 | Mosses | B | Leucodontales | Neckaraceae | Homaliodendron scalpellifolium | cus | JBNGMX |
| 118 | Mosses | B | Orthotrichales | Orthotrichaceae | Orthotrichum anomalum | cus | SRR33208788 |
| 119 | Mosses | B | Orthotrichales | Orthotrichaceae | Orthotrichum diaphanum | cus | ERR15745620 |
| 120 | Mosses | B | Orthotrichales | Orthotrichaceae | Plenogemma phyllantha | cus | ERR15405273 |
| 121 | Mosses | B | Polytrichales | Polytrichaceae | Oligotrichum hercynicum | cus | ERR15610518 |
| 122 | Mosses | B | Polytrichales | Polytrichaceae | Pogonatum subfuscatum | cus | SRR33208732 |
| 123 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichastrum ohioense | cus | SRR33208602, JBNGKP |
| 124 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichum commune | cus | JBNGKQ |
| 125 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichum formosum | cus | ERR12721078, ERR12721079, SRR8707304 |
| 126 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichum juniperinum | cus | ERR15606502, ERR15610523 |
| 127 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichum piliferum | cus | ERR15378475, ERR15381805, ERR15381806 |
| 128 | Mosses | B | Polytrichales | Polytrichaceae | Polytrichum strictum | cus | JBNGKO |
| 129 | Mosses | B | Pottiales | Pottiaceae | Anoectangium aestivum | cus | ERR15656951, ERR15659573 |
| 130 | Mosses | B | Pottiales | Pottiaceae | Bryoerythrophyllum caledonicum | cus | ERR12370311, ERR12356316 |
| 131 | Mosses | B | Pottiales | Pottiaceae | Bryoerythrophyllum recurvirostrum | cus | SRR24580705 |
| 132 | Mosses | B | Pottiales | Pottiaceae | Hymenostylium aurantiacum | cus | JBNGLN |
| 133 | Mosses | B | Pottiales | Pottiaceae | Hyophila propagulifera | cus | DRR428591 |
| 134 | Mosses | B | Pottiales | Pottiaceae | Syntrichia ruralis | cus | SRR27955587 |
| 135 | Mosses | B | Hypnales | Pseudoleskeaceae | Lescuraea plicata | cus | SRR24580403 |
| 136 | Mosses | B | Hypnales | Regmadontaceae | Regmatodon declinatus | cus | SRR33208207, SRR33208208, SRR33208209 |
| 137 | Mosses | B | Rhabdoweisiales | Rhabdoweisiaceae | Dicranoweisia cirrata | cus | ERR15381799 |
| 138 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum affine | cus | SRR18184152 |
| 139 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum angustifolium | cus | SRR6966991, SRR6968021, SRR6966992 |
| 140 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum capillifolium | cus | SRR6965943, SRR6965944, SRR6965384 |
| 141 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum compactum | cus | SRR6973028, ERR4781406, ERR4778815 |
| 142 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum contortum | cus | ERR12259833 |
| 143 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum divinum | cus | SRR18184163, SRR18184156 |
| 144 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum fallax | cus | SRR18184003, SRR18183987 |
| 145 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum fuscum | cus | SRR6966970, SRR6967721, SRR6966971 |
| 146 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum lindberghii | cus | ERR15126493 |
| 147 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum magellanicum | cus | SRR22501966, SRR22501971, SRR22501943 |
| 148 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum palustre | cus | SRR6964393, ERR13650026 |
| 149 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum papillosum | cus | SRR6966183, SRR6966184 |
| 150 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum portoricense | cus | SRR18427469 |
| 151 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum rubellum | cus | SRR6966977, SRR6965089, SRR6965088 |
| 152 | Mosses | B | Sphagnales | Sphagnaceae | Sphagnum squarrosum | cus | SRR6973636, SRR6973635, ERR4780966 |
| 153 | Mosses | B | Splachnales | Splachnaceae | Tayloria subglabra | cus | SRR33208468 |
| 154 | Mosses | B | Hypnales | Thuidiaceae | Thuidium delicatulum | cus | ERR12374262 |
| 155 | Mosses | B | Hypnales | Thuidiaceae | Thuidium tamariscinum | cus | ERR6895937, ERR6895936, ERR6895935 |
| 156 | Mosses | M | Lophoziales | Anastrophyllaceae | Neoorthocaulis attenuatus | plast | ERR14995244, ERR15165952 |
| 157 | Mosses | M | Porellales | Porellaceae | Porella caespitans | cus | JBNGJO |
| 158 | Mosses | M | Jungermanniales | Solenostomataceae | Solenostoma hyalinum | acs | SRR25319279 |
| 159 | Mosses | M | Jungermanniales | Solenostomataceae | Solenostoma erectum | ags | SRR33208228 |
acs‐orf2‐orf3n‐orf8‐rolB‐orf14‐mas2′‐mas1′‐ags‐vis.
acs‐orf2‐orf3n‐orf8‐rolA‐rolB‐1215like‐mas1′‐ags‐cus.
Earlier we have shown for Camellia (Chen et al., 2023) and Diospyros (Otten et al., 2025) that cT‐DNA insertions can precede speciation. The Bryophyta provide another example of this. The P. commune assembly JBNGKQ‐5857 shows a small cT‐DNA with a cus gene. Sequences from the NCBI core_nt collection from Atrichum undulatum (no cT‐DNA), and numerous moss sequences from the WGS database show that the insert (called PoTA) is limited by coordinates JBNGKQ‐5857:1 251 718–1 252 882. The same insert is also found in Polytrichastrum ohioense (JBNGKP‐18738) and Polytrichastrum strictum (JBNGKO‐15772). DNA sequence homology between the three Polytrichum species extends well beyond PoTA, showing that all three have the same PoTA insert. Thus, PoTA was inserted before speciation. Conversely, Entodon concinnus JBNGNB‐6 shows 85% identity with PoTA JBNGKQ‐5857:1 251 715–1 252 823, but diverges on both sides. Therefore, its insert (EnTA) differs from PoTA. Similarly, Homaliodendron scalpellifolium JBNGMX‐38774 has a cT‐DNA similar to those of PoTA and EnTA, but diverges beyond, and is therefore still another insert (HoTA). The present data represent many new opportunities for pre‐speciation insertion analysis, which remain to be explored. A phylogenetic tree of the predicted Cus proteins shows three distinct groups, only found in mosses and ferns (Figure 2).
Figure 2.

Phylogenetic tree for Cus proteins from ferns (green), mosses (red), eudicots (yellow), Agrobacterium/Rhizobium (magenta), other bacteria (purple), and fungi (blue).
I, II, III: Cus groups specific for mosses and ferns. Numbers on branches: bootstrap values. Reconstr: reconstructed from fragments, part: partial sequence.
A search of the clusteredNR protein collection on the NCBI BLASTP site, using all available Cus sequences as queries, also detected Cus sequences in seven fungal species. As in the case of mosses and ferns, this small number of species suggests that these cus genes did not originate from fungi but were acquired from Agrobacterium. This will require further investigation.
The cT‐DNA sequences in Bryophyta, Anthocerotophyta, and Marchantiophyta show that Agrobacterium can transform nonvascular plants, and that these plant species can regenerate into nGMOs under natural conditions. We next investigated the Lycopodiophyta and Polypodiophyta.
Lycopodiophyta and Polypodiophyta
Vascular embryophytes comprise the Lycopodiophyta (1290), Polypodiophyta (10 560), and Spermatophyta (296 463). We found five nGMOs in Lycopodiophyta and 70 in Polypodiophyta (Table 1). These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table 1. WGS data were used to construct cT‐DNA maps. T‐DNA/cT‐DNA maps in this paper are numbered in bold from 1 to 166; genes oriented from right to left T‐DNA border are indicated in bold. Fern cT‐DNA maps are shown in 145–149 (Figure 3a). Remarkably, contigs JAIKUY‐2021 (145) and JAIKUY‐2022 (146) carry 8 and 17 vitopine synthase (vis) genes, spread out over 245 and 423 kb, respectively. 2021‐vis2, 2021‐vis5, and 2022‐vis1 ORFs are intact. The 2022 vis copies are part of larger repeats (Figure 3b), suggesting amplification after insertion. Ceratopteris richardii JAIKUY‐938 (not shown) contains a single cus gene. Adiantum nelumboides (Figure 3a) contains AdTA‐1 (147), AdTA‐2 (148), and AdTB (149). The AdTA‐1 cT‐DNA sequences are interspersed with 144 kb of plant DNA (containing a 16 kb repeat), which is highly unusual for cT‐DNAs. AdTA‐2 is a short version of AdTA‐1. AdTB‐ipt is distantly related to AdTA‐ipt (64% identity) and probably derived from another T‐DNA. AdTA and AdTB genes ipt, 6b, vis, and d are unusual for Rhizobium rhizogenes, but common in Agrobacterium tumefaciens. The only intact AdTA gene is susL (L,L‐succinamopine synthase). These data show that ferns can be transformed by Agrobacterium and generate nGMOs in nature. We next investigated the spermatophytes.
Figure 3.

Fern cT‐DNA maps.
(a) Ceratopteris richardii (JAIKUY‐2021, 145), Ceratopteris richardii (JAIKUY‐2022, 146), Adiantum nelumboides, AdTA‐1 (JAKNSL‐5917, 147), Adiantum nelumboides, AdTA‐2 (JAKNSL‐0548, 148), and Adiantum nelumboides, AdTB (JAKNSL‐7169, 149). Repeats in 147 are indicated by red arrows. *: intact ORF.
(b) Repeat analysis for 17 vis copies from JAIKUY‐2022:1–450 000. Matrix analysis shows the extent of the repeats. Top: Query: 2022:1–12 000 (containing vis1). Long repeats occur around vis2, vis8, vis10, smaller ones around vis3, vis5, vis7, vis9, vis11, and vis12, and very small ones around vis4, vis6, and v13 to vis17. Bottom: Query: 2022:415 000–440 000 (containing an incomplete vis gene, vis17): small repeats for vis1 to vis3, vis5, vis7 to vis12, larger repeats for vis13 to vis17. Thus, the inserts can be divided into two groups: vis1 to vis12, and vis13 to vis17. vis coding sequences in red.
nGMOs in spermatophytes
Spermatophytes consist of gymnosperms (1079) and angiosperms (295 403). Their nGMO data are shown in Table S2. In gymnosperms, we found five nGMOs, in angiosperms 2452. The Angiosperm Phylogeny Group (APG; Group et al., 2016) recognizes 64 orders and 416 families. They comprise basal angiosperms: Amborellales, Nymphaeales, and Austrobaileyales (183 species) and core angiosperms (or Mesangiospermae): magnoliids, Chloranthales, monocots, Ceratophyllales, and eudicots (295 220 species).
Two nGMOs were detected in basal angiosperms, 68 in monocots (74 300 species), 2338 in eudicots (210 000 species), and 44 in the remaining groups (10 920 species). These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table S2. 2270/2452 angiosperm nGMOs were exclusively found in SRA data, 13 times more than from WGS searches. Only 219/2452 angiosperm nGMOs were reported previously (Table S2, marked in bold). Thus, the present results represent an 11‐fold increase. A distribution of the nGMOs over the spermatophyte families and their genera is shown in Figure S1.
Estimation of nGMO species among land plants
The large increase in nGMO numbers compared to the 2019 study (Matveeva & Otten, 2019), allows a better estimate of the total number of nGMO species among land plants. For this, we mainly used WGS data, as most SRA sequences are incomplete. The present NCBI data contain 2435 WGS sequences for land plants (after correction for multiple WGS data from the same species), 195 of which (8%) are nGMOs. However, 105 additional species without cT‐DNAs in their WGS sequences do show cT‐DNA sequences in non‐WGS data from other accessions of the same species, indicating their cT‐DNA sequences are not fixed. This leads to a new estimation of 300/2435 species (12.3%), which extrapolates to 40 700 nGMOs among the 330 200 land plant species.
The new nGMO data were used to investigate five areas of interest: extension of previous nGMO studies, tracing the origin of the cT‐DNA in Vaccinium, identification of nGMOs used by humans, presence of nGMOs in eudicot sister groups, and comparative structural analysis of fully mapped cT‐DNAs.
Data presented in the next three subchapters are shown in Figure S2 and Table S2.
Extension of previous nGMO studies
The present data extend the results of previous detailed studies on the genera Nicotiana, Linaria, Vaccinium, Ipomoea, Camellia, Diospyros, and Arachis.
In Nicotiana, we found 13 additional nGMO species. A new cT‐DNA (TF, 23) was found in N. otophora (SRR25726134, SRR35059995). In Linaria, three additional nGMOs were found: L. angustissima, L. japonica, and L. repens. Linaria vulgaris cT‐DNA maps were constructed for EU735069.2 (28) and WGS CAOJCA (29, 30). Another well‐studied genus is Vaccinium. Vaccinium macrocarpon (cranberry) contains a single plast‐like gene (1215‐like or rolB/rolC). rolB/rolC was reported in 26 Vaccinium species and Agapetes serpens (Zhidkin et al., 2023). Our present study revealed 1215‐like sequences in as many as 28/46 genera from the Vaccinioideae subfamily (see also next chapter).
Two cT‐DNAs (IbT‐DNA1 and IbT‐DNA2) were reported in five Ipomoea species (Kyndt et al., 2015; Quispe‐Huamanquispe et al., 2019; Yan et al., 2024). Here we report 53 additional Ipomoea nGMO species, and a new cT‐DNA (IbT‐DNA3) in I. batatas and I. trifida. Ipomoea cT‐DNAs are shown in 31–46.
Camellia has been reported to contain 71 nGMO species (containing different combinations of cT‐DNAs CaTA to CaTN; Chen et al., 2023). The cT‐DNA ancestors, except those of CaTJ and CaTK (both incomplete), are shown in 47–58. Sixty‐eight new Camellia nGMO species were detected in the present study.
Diospyros (persimmon) has 39 nGMOs (Otten et al., 2025). This study adds D. apiculata, D. armata, D. hoyleana, D. major, and D. sonorae.
The tetraploid species Arachis hypogaea (peanut, 2n = 4x = 40) carries ancestor genomes A and B, with cus and mas2′ sequences (Matveeva & Otten, 2019). Twenty‐three Arachis nGMO species were identified (Bogomaz et al., 2024). The authors noted an intact cus gene, cus gene remnants and mas2′ sequences (A. hypogaea), ags (A. macedoi, A. pusilla), and mas1′ (A. appressipila, A. rigonii). We detected five additional species, used WGS SDMP (A. hypogaea) to recover cT‐DNAs ArTA, ArTB, ArTC, and ArTD (163–166), and determined the distribution of these four cT‐DNAs among various Arachis species (Table S3).
Tracing the origin of the Vaccinium VaTA insert
The unexpected presence of 1215‐like sequences in a large number of Ericaceae species belonging to different genera raised the question whether all were derived from the same transformation event or not. In order to answer this question, two kinds of information are needed: defining the limits of the insert, and the nature of the flanking plant sequences. We first determined the extent of the cT‐DNA insert in Vaccinium ovatum (assembly JBVKUD‐2), by identifying the surrounding plant sequences, easily recognizable because of their repeated nature. This showed that the insert (subsequently called VaTA) was located at JBVKUD‐2:25 046 610–25 047 986. VaTA was then used as a query to all WGS and SRA sequences from Ericaceae. In case of sufficient coverage, the hits extend into the flanking plant sequences. If these align with those flanking VaTA, the accession carries the same VaTA insert. Using this approach, we retrieved VaTA inserts and their flanking sequences from Vaccinium and the other genera. Ninety‐six Vaccinium species contain the 1215‐like gene, 51 showed enough coverage for VaTA border analysis, all contained VaTA. Outside the genus Vaccinium, VaTA and its flanking sequences were found in Gaultheria (G. prostrata, shallon, baccata, reticulata, cinerea, and angustifolia), Agapetes (A. malipoensis, forrestii, serpens, and polifolia), Polyclita (P. turbinata), Notopora (N. schomburgkii), and Psammisia (P. fissilis). Thus, the VaTA cT‐DNA was inserted in the common ancestor of these six genera. The root of the Vaccinieae tribe (with Vaccinium, Agapetes, Polyclita, Notopora, and Psammisia) has been estimated at 30 million years ago (Becker et al., 2024). As VaTA is also present in tribe Gaultherieae (Gaultheria) which split off earlier, the VaTA insertion must have occurred before. The conservation of the 1215‐like open reading frame noted before (Zhidkin et al., 2023) is all the more remarkable in view of the long history of the VaTA insert. In the next part, we provide data on new nGMOs used by humans.
New nGMOs among plants used by humans
Among nGMOs used for food and drinks, pomelo (Citrus maxima) carries ags and mas2′ (Matveeva & Otten, 2019). Here we assembled CiTA (124) from JBJGYR‐8 (C. maxima cv. ZPY). Interestingly, C. maxima assemblies JAQPSH, JAUJEG, JBKACN, JBKJAM, JBKJAN, JBLANY, and MKYQ lack CiTA, showing that CiTA is non‐fixed. Similarly, three C. sinensis cultivars (SRR11681090, SRR4089854, and SRR12300375) are nGMOs, but 22 others are not. Eighteen additional Citrus nGMO species and hybrids were found. Among these, citron (C. medica), mandarin (C. reticulata), kumquat (C. hindsii, C. japonica), bitter orange (C. aurantium), satsuma mandarin (C. unshiu), and grapefruit (Citrus × paradisi).
We also found cT‐DNA sequences in coffee (Coffea arabica, C. canephora, and C. liberica), avocado (Persea americana), cinnamon (Cinnamomum aromaticum), strawberry (Fragaria viridis and others), sea buckthorn (Hippophae rhamnoides), icecream bean (Inga edulis), cassava/manioc (Manihot esculenta), black pepper (Piper nigrum), apricot (Prunus armeniaca), raspberry (Rubus ideaeus), dewberry (Rubus caesius), and sorrel (Rumex acetosa).
cT‐DNAs in eudicot sister groups
Aside from eudicots, core angiosperms contain magnoliids, Chloranthales, Ceratophyllales (Figure 4, ‘others’, 10 290 species) and monocots. In the ‘others’ group, we found 41 (magnoliids), 3 (Chloranthales), and 0 (Ceratophyllales) nGMOs. Monocots contain 11 orders, 77 families, 2700 genera, and 74 273 species. Matveeva and Otten (2019) reported cT‐DNAs in Musa acuminata (banana) and greater yam (Dioscorea alata, WGS CZHE). We found no further Musa cT‐DNA sequences, but the Dioscorea results were confirmed and considerably extended. Whereas the CZHE contigs are small, JAFBII (D. alata, cv. TDa95/00328) has fully assembled sequences, attributed to each of its 20 chromosomes (Bredeson et al., 2022). DcTA (6315 nt, 141, located on chromosome 8) carries a single cus gene, interrupted by a 4920 nt plant sequence, PL1. Long reads from SRR13615784 showed that TDa95/00328 carries two different DcTA alleles, one with an intact cus gene (DcTA‐1, 1331 nt) and one with PL1 (DcTA‐2). Bioprojects PRJNA918625 and PRJNA666450 contain sequences from 134 D. alata cultivars. All carry DcTA and show five structural variants (140–144) with different PL insertions (Table S4). DcTA‐2 occurs in 6/134 accessions and is therefore not representative for D. alata. Only seven accessions lack intact cus genes.
Figure 4.

Overview of plant groups with natural genetically modified organisms (nGMOs). Estimated species numbers for each group (Christenhusz & Byng, 2016) in black, nGMO numbers based on Sequence Read Archive (SRA) and whole genome sequences (WGS) data in red.
Group marked by ‘others’: magnoliids, Chloranthales, and Ceratophyllales. Photographs show representative nGMOs: Polytrichum commune (Bryophyta), Folioceros fuciformis (Anthocerotophyta), Solenostoma hyalinum (Marchantiophyta), Dryopteris affinis (Pteridophyta), Huperzia selago (Lycopodiophyta), Juniperus formosana (gymnosperms), Euryale ferox (basal angiosperms), Dioscorea alata (monocots), and Dianthus caryophyllus (eudicots).
cus sequences were also found in three additional Dioscorea species. The related species Trichopus zeylanicus has ags sequences. Finally, SRA data from 67 additional monocots (9 orders, 20 families, 52 genera) also contain cT‐DNA sequences (Table S2).
Diversity of cT‐DNA structures
The large number of cT‐DNAs recovered here allowed us to estimate the frequency of occurrence of the different cT‐DNA genes (Table S2). This showed that some (like susD, 6b, p5, ipt, c', d, and e) are rare, whereas others (like orf13, orf14, susL, and cus) are very frequent. The unusual genes are more typical for A. tumefaciens.
In order to study the different cT‐DNA structures in more detail, the WGS data were used to construct cT‐DNA maps. Using a standard format, we redrew relevant T‐DNA maps (1–17), earlier published cT‐DNA maps (18–73), and new cT‐DNA maps (74–166). Four groups (A–D, Table S5) can be distinguished.
Group A cT‐DNAs (like DcTA, 140) contain a single gene (‘mini cT‐DNAs’). Only one mini T‐DNA is known from agrobacteria (Otten, 2021): Agrobacterium 1641 has 1641‐T2 (vis), as well as 1641‐T1 (acs, susD, and 6b). 87/183 from the investigated WGS show mini cT‐DNAs, 75 of which carry ops genes. Thus, this type of T‐DNA is very frequent, contrary to what could be expected on the basis of Agrobacterium T‐DNA studies. Mini cT‐DNAs can be present in multiple copies, as in C. richardii (vis, 146), Urtica dioica (vis, 155, 156), Silene latifolia (cus, 157), and Paulownia fortunei (vis, 91). Twenty‐one WGS sequences from Caryophyllaceae show a mini cT‐DNA with a single cus gene. Vaccinium and related genera have a mini cT‐DNA with a single 1215‐like gene (see above).
Group B cT‐DNAs resemble the TL‐DNA from R. rhizogenes LMG152 (3), with acs‐orf2‐orf3n‐orf8‐rolA‐rolB‐rolC‐orf13‐orf13a‐orf14‐susL1‐susL2 . T‐DNA variant (5) has gene c between orf3n and orf8, and gene plast at the place of rolC. Others have susL1‐susL2 replaced by mas2′‐mas1′ (11), susL (4, 6, 7, 12), cus (1), or mis (2, 5). Some cT‐DNAs from Group B have nos at the right border (58, 59, 65). rolC can be replaced by plast (56, 86, 97, 113, 116, 139), or accompanied by plast or d (77, 100, 104). rolB can be replaced by rolB TR (48, 53, 58, 92, 99, 121, 125). In T‐DNAs, acs can be replaced by orf358 (12), most likely an ops gene (Chen et al., 2023). In cT‐DNAs, acs (most cases) or orf358 (54, 74) can be replaced by vis (81, 82, 138), ocs (22, 63, 114), or orf358 ‐ ocs (62). Thus, a more general TL structure is ( ops )‐ ops‐orf2‐orf3n‐( c )‐orf8‐rolA‐ rolB‐(or rolB TR )‐rolC‐(d or plast)‐orf13‐orf13a‐orf14‐ops ‐( ops ). Gene c from DiTE ( ocs‐c‐orf8, 63) is inverted with respect to its orientation in Group B ( c‐orf8). DiTD (62) is a combination of two Group B T‐DNAs (T1 and T2).
Group C cT‐DNAs are similar to LMG152‐TR (8) with iaaH ‐iaaM‐ rolB TR ‐ mas2′‐mas1′‐ ags . T‐DNA variants lack ags (9) or have iaaH ‐iaaM‐ vis (10). cT‐DNAs TE1 (24) and DiTN (72) have iaaH ‐iaaM‐ 6b ‐ vis . Group C sequences can be found linked to Group B sequences (B+C, Table S5).
Group D contains cT‐DNAs with orf511. orf511 has not yet been found in agrobacteria, but a non‐characterized bacterium from the Paracoccaceae potentially encodes an Orf511 homolog (hypothetical protein MCU0909928.1), its function is unknown. The orf511 gene has been identified on TD from N. tomentosiformis (20), CaTD (50), CaTF (52), CaTM (57), and CaTN (58) from Camellia, and DiTG (65) and DiTH (66) from Diospyros. It also occurs in cT‐DNAs 80, 86, 97, and 117, and is generally accompanied by orf14 and one or two susL genes: orf14‐orf511‐ susL ‐( susL ). This DNA may be related to (3) by insertion of orf511 in the orf14‐ susL ‐ susL fragment. orf511 is a common T‐DNA gene, as it appears in 188 species from 59 genera (Table 1; Table S2). Two fern species (Actinostachys digitata, Abrodictyum obscurum) contain orf511, orf14, and susL; these remain to be mapped. A few cT‐DNAs could not be classified (Table S5, not defined).
cT‐DNA repeats
Most cT‐DNA inserts are relatively simple inverted repeats. Direct repeats are rare; examples are JBJYGR‐8 (124), EU735069.2 (28), and JBHWAD‐60 (159). More complex repeats occur in BSXM‐40 (31, six repeats) or CANTUR‐331 (160) with four direct repeats (331‐A to 331‐D) in a total region of 600 kb. 331‐A to 331‐D appear to result from amplification of cT‐DNA inserts and surrounding plant sequences. CANTUR‐344 (161) and CANTUR‐1107 (162) each have one copy. The original cT‐DNA insert contains four fragments from the original T‐DNA. Altogether, CANTUR (Mallotus nudiflorus) contains 356 kb of cT‐DNA. CBDBZS (Camellia meiocarpa) also contains complex cT‐DNA repeats (150–153). Silene latifolia JBLOPE‐1 has 11 cus copies spread out over 146 kb (157). These are not part of longer repeats but seem to be independent insertions. The origin of such multiple insertions clustered within a small chromosomal region is unknown. Urtica dioica CAXLOR‐222 (155) and CAXLOR‐56 (156) have multiple cT‐DNAs with a single vis gene within 250 and 225 kb, respectively; these are part of longer repeats. Overall, the cT‐DNA repeat structures appear highly variable.
DISCUSSION
A first systematic cT‐DNA search in public databases from eudicots and monocots (an estimated 284 300 species) found 23 nGMO species in 275 WGS sequences (Matveeva & Otten, 2019). By searching more recent WGS data, adding SRA data, and extending the search to all land plants, we obtained a list of 2614 nGMO species. Most of their cT‐DNAs remain to be completed. However, we believe that the present list can serve as a reliable starting point for further investigations. For example, in the genus Solanum, we detected 62 nGMO species. Solanum has about 1500–2000 species, most of which remain unsequenced. Phylogenetic studies in this genus (see, e.g., Messeder et al., 2024) may benefit from detailed analysis of the cT‐DNA distribution patterns, as each insertion marks the start of a new lineage, and thus provides an independent way to confirm phylogenetic patterns. The overall phylogenetic distribution of the presently known nGMOs is shown in Figure 4. It should be noted that many plant species remain unsequenced. Our estimation of 40 700 nGMO species in land plants is probably underestimated, as many cT‐DNAs are non‐fixed, and noncoding sequences were excluded in our study. The presence of cT‐DNAs in ferns and mosses shows that these spore‐bearing land plants can be transformed by Agrobacterium, and regenerate into nGMOs, suggesting that Agrobacterium transformation could date back to early stages in land plant evolution. This is also indicated by the large sequence divergence of the plast‐like T‐DNA genes, which suggests an early origin of natural plant transformation (Otten, 2018).
Although many sequences recovered here are incomplete, these could already be used to extend earlier data on nGMO genera like Vaccinium, Ipomoea, Arachis, Dioscorea, and Citrus. The presence of cT‐DNA in various plants used for popular food and drinks, like cassava, apricots, avocado, strawberry, and coffee (to name a few) could lead to further studies on the possible contribution of cT‐DNA genes to the properties of these economically important plants. WGS data were used to construct a large collection of cT‐DNA maps, which defined four structural groups. Group A contains mini cT‐DNAs, mostly with a single ops gene. They are found in mosses, ferns, Dioscorea, and eudicots. Most likely, Agrobacterium strains with mini T‐DNAs are unable to induce tumors or HRs. Thus, they would most likely escape the attention of plant pathologists, which could explain their absence in bacterial collections and the lack of their sequences in DNA databases.
In view of the broad occurrence of cus genes in land plants, it seems possible, at first sight, that they are of plant origin, and were at some stage transferred to Agrobacterium. However, we believe this is unlikely. Recently, 138 whole genomes (accession codes: JBNGxx) were obtained from mosses and liverworts (Dong et al., 2024). Only 10 (Figure 2), belonging to different families, contain the cus gene. Also, the Cus protein from Porella (Marchantiophyta) clusters with Cus from Polytrichum (Bryophyta), in spite of the fact that these two species are quite distant (Figure 2). This patchy pattern is not what one would expect if the cus gene would be of plant origin, but is more likely to result from transformation of a limited number of species by related Agrobacterium strains. Most importantly, an ‘empty insertion site’ (a site without cT‐DNA insert) in H. elatum aligns with the sequences surrounding the HyTA insert in H. flavolimbatum, making it likely that this empty site represents the original, non‐modified site, and that H. flavolimbatum acquired HyTA by Agrobacterium‐mediated transformation.
A search of the NCBI clusteredNR protein collection, using various Cus sequences as queries, detected seven Cus sequences in fungi. Such a low number makes it unlikely that these are of fungal origin, but (as in the case of plants) argues in favor of Agrobacterium‐mediated transformation.
The mini cT‐DNAs from Group A may be explained in three ways. They could result from infection by strains with mini T‐DNAs, by partial transfer of longer T‐DNAs, or from deletions of longer cT‐DNAs after integration. Possibility two and three are expected to leave some traces of neighboring genes, as the deletion process is expected to be random. Because such traces have not been observed so far, we favor the first hypothesis. Of course, final proof will require the discovery of Agrobacterium strains with mini‐T‐DNAs.
T‐DNAs from groups B and C could have been derived from mini T‐DNAs by incorporating HR‐inducing genes. Group D T‐DNAs (carrying orf511, orf14, susL) are unknown in agrobacteria but common in nGMOs. orf14 plays a minor role in HR initiation and growth (Aoki & Syōno, 1999; Otten, 2018) but can induce significant dwarfing (Favero et al., 2022). Like for orf511, its mode of action remains unknown, but the large number of orf14 and orf511 genes already identified in nGMOs might be exploited for further investigations. Possibly, strains with group D T‐DNAs are also non‐symptomatic.
The distribution of Group A–D structures among land plants suggests their possible evolutionary origin. Mosses carry mini cT‐DNAs with a single cus gene; these could be the earliest cT‐DNAs. Ferns show mini T‐DNAs with cus and other ops genes, indicating these appeared later. About half of the eudicots have complex cT‐DNAs; the others carry mini cT‐DNAs. Examples of eudicots with cus mini cT‐DNAs are found in genera Cerastium (19 accessions), Dianthus (17), Gypsophila (9), Moehringia (8), Schiedea (17), Silene (85), Stellaria (19), and Stylosanthes (10). Genera with other mini cT‐DNAs are: Cuscuta (mis, 14), Juglans‐Platycarya‐Pterocarya‐Cyclocarya (susL, 17), and Urtica (vis, 8). We predict that species from these 14 genera are able to regenerate nGMOs from single transformed cells, without a need for an intermediate HR stage. Vaccinioideae with their 1215‐like genes may also be included in this group, but its effect on transformed cells is unknown. The widespread occurrence of mini cT‐DNAs with ops genes strongly supports the notion that opines are the ‘raison d'être’ of the Agrobacterium–plant interaction (Petit & Tempé, 1985). However, it remains to be demonstrated whether opine synthase genes from nGMOs still play a role in Agrobacterium biology, or have been recruited for other functions instead.
In general, the conservation of cT‐DNA open reading frames in present‐day nGMO species suggests that they play some role in nGMOs. The VaTA 1215‐like gene from Vaccinium and other, related genera is a particularly striking case, as most accessions have an intact open reading frame (Zhidkin et al., 2023), in spite of the fact that the insert is over 30 Mio years old. Overexpression studies of T‐DNA and cT‐DNA genes in heterologous systems have shown various phenotypic effects, suggesting similar effects in the respective nGMOs (Chen & Otten, 2017). However, defining a biological role will require comparative studies of wild‐type nGMOs and isogenic mutants. A recent paper describes the knock‐out of two IbT‐DNA2 genes in sweet potato: rolB/C (a plast gene) and rolD (an opine synthase gene). The authors showed a reduction of chlorogenic acid and polyphenol content in the rolB/C mutant and a reduction in biomass and downregulation of cell cycle‐related genes in the rolD mutant (Shkryl et al., 2025). Further studies will be required to determine the molecular mode of action of these genes. Among the cT‐DNA genes, the opine synthase genes will be the easiest ones to study, as their general function is well known. Deoxyfructosylglutamine (santhopine) has been found in N. tabacum (Chen et al., 2016) and mikimopine in Cuscuta (Zhang et al., 2020), but no mutants have been studied.
cT‐DNAs mostly consist of simple, partial inverted repeats. However, direct repeats and complex repeats also occur. Complex repeats can result from ligation of T‐DNA fragments before insertion, multiple insertions during transformation, or insertion followed by amplification. T‐DNA inserts from laboratory experiments show similar variation (De Buck et al., 2009; Gelvin, 2003; Jupe et al., 2019; Kleinboelting et al., 2015; Kralemann et al., 2022; Pucker et al., 2021). A special feature of nGMOs is the presence of multiple inserts in a large chromosomal region, as observed for the mini cT‐DNAs from C. richardii, U. dioica, and S. latifolia. Possibly, strains with mini T‐DNA have different T‐DNA transfer properties, leading to such unusual insertion patterns.
This study, and other investigations, show that cT‐DNAs and their corresponding T‐DNAs could become rich sources of genes with economically useful effects on plant growth and metabolism. Wild‐type R. rhizogenes strains have been used to improve ornamental plants (Favero et al., 2022). This approach may be extended to vegetable crops. Indeed, nGMOs show that both Agrobacterium‐mediated T‐DNA transfer and the subsequent generation of transformants are natural phenomena. These processes have occurred for millions of years in all land plants and involve large numbers of cultivated species.
MATERIALS AND METHODS
Whole genome sequencing data analysis
In order to identify and extract cT‐DNA sequences from WGS data, we first collected taxonomic IDs (taxid) from the NCBI Taxonomy Browser for Spermatophyta (txid 58024), Acrogymnospermae (txid 1437180), Polypodiophyta (or Polypodiopsida, txid 241806), Lycopodiophyta (or Lycopodiopsida, txid 1521260), Bryophyta (txid 3208), Marchantiophyta (txid 3195), and Anthocerotophyta (txid 13809) (Schoch et al., 2020).
Using taxid2wgs.pl (https://ftp.ncbi.nlm.nih.gov/blast/WGS_TOOLS/taxid2wgs.pl), we generated alias files for each group and downloaded the WGS datasets. These were searched with blastn_vdb (Altschul et al., 1990) to find homologous sequences. For this, we used the Ntfuse6 DNA query, a concatenated sequence composed of selected T‐DNA and cT‐DNA sequences (Otten et al., 2025). In order to avoid false positives, the DNA sequences were checked by BLASTX analysis for bona fide T‐DNA proteins using a set of 380 T‐DNA and cT‐DNA proteins (Table S1). Long WGS numbers were abbreviated. For example, JAIKUY010002022.1 was converted to JAIKUY‐2022.
Retrieval and analysis of SRA datasets
To retrieve the full set of SRA data from all land plant groups (except monocots and three large eudicot families, see below), we used the Entrez Direct tool to obtain the accession numbers of the target group, followed by the SRA‐toolkit to download the raw sequencing data from the SRA database at the NCBI (https://www.ncbi.nlm.nih.gov/sra/). As the data generated from different experiments and projects may vary significantly, and the default maximum download file size for the ‘prefetch’ command is 20 GB, the ‘‐‐max‐size’ parameter was set to its maximum value. We downloaded all available SRA data for land plant groups, with the exception of monocots, Fabaceae, Brassicaceae, and Solanaceae, where the number of SRR accessions was extremely large. For these, we used a simplified approach by randomly selecting up to 20 SRR accessions per species. The SRA data of the Dioscoreae family were searched completely.
For searching T‐DNA gene homologs in all sequencing data of land plant groups, we used the ‘blastn_vdb’ command from BLAST+ 2.15.0, with the Ntfuse6 sequence as the query and the following parameters: ‘‐gapopen 5 ‐gapextend 2 ‐word_size 11 ‐evalue 0.05 ‐penalty ‐3 ‐reward 2’.
Recovery of cT‐DNA sequences and assembly
The data archived in the Sequence Read Archive (SRA) can be either paired‐end sequences or single‐end sequences, depending on the experimental design and requirements. Paired‐end reads can enhance the credibility and accuracy of the sequences. However, in some cases, the original ‘blastn_vdb’ output may only contain single‐end reads. To address the issue of such reads, caution is needed in extracting sequences. Specifically, the blast output distinguishes between paired‐end sequences with the suffixes ‘.1’ or ‘.2’. After removing the sequence name suffixes and performing the ‘uniq’ operation, sequences are then extracted separately from the paired‐end data. After recovering cT‐DNA homologous sequences from SRA data, CodonCode Aligner (CCA, version 11.0.2, CodonCode Corporation) was used for sequence assembly. Variations in experimental design and sequencing protocols can significantly affect data quality, resulting in differences in the recovery of homologous sequences. Because CCA assembly uses the De Bruijn graph‐based method, datasets with low read counts often yield poor‐quality assemblies. Therefore, we opted not to assemble SRA data containing fewer than 100 reads; instead, these reads were analyzed directly with BLASTX. For datasets with more than 100 reads, we first generated contigs by CCA assembly and subsequently subjected these to BLASTX analysis. Unassembled reads were analyzed in the same way. Contigs obtained from CodonCode Aligner were utilized as queries in BLASTX to identify the closest protein homologs. For contigs less than 1 kb, the protein with the lowest E‐value was selected; for longer contigs, the Graphic Summary view in online BLASTX was used to examine the closest proteins. It is important to note that in many cases, no full cT‐DNAs could be assembled. The combined WGS and SRA data approach (Figure 1) allowed us to identify and catalog the presence of Agrobacterium/Rhizobium‐derived T‐DNA sequences in various plant genomes.
Exclusion of false positives
In the SRA, some datasets are from species transformed by R. rhizogenes, such HR lines should be excluded. So far, only four A. rhizogenes strains have been used, LMG152 (15 834, A4 or HR‐I), LMG63 (8196), 1724, and NCPPB2659 (K599) (Ying et al., 2023). To detect HR lines, all recovered sequences were compared to the T‐DNAs of these four A. rhizogenes strains by BLASTN, and excluded if they showed more than 99% identity. The second group of contaminants was due to weakly homologous sequences from plant genomes. Indeed, T‐DNA proteins Orf8, SusL, Mas1′, Ags, Ipt, IaaH, and IaaM (Table S1) may show weak similarity to embryophyte proteins (less than 40% identity). Only a 45 amino acid IaaM fragment (MCX‐IaaM coordinates 252–297) may have 65% identity. As nucleotide sequences are less conserved than protein sequences, the chance of finding DNA sequences from these plant homologs in our initial search with DNA query Ntfuse6 is low. Nevertheless, to avoid false positives, we checked all proteins with identity levels below 50% by BLASTP against the NCBI ClusteredNR database. In case the best hits were from plants without cT‐DNAs, the accession was discarded. Bacteria other than agrobacteria may also yield false positives with weak DNA and protein homology. For these SRRs, we compared the recovered DNA sequences to the NCBI core nucleotide database with BLASTN, and removed accessions whose top hits matched sequences from bacteria outside the agrobacteria group. Finally, to exclude Agrobacterium contamination of sequenced plant material, we tested the positive SRR and WGS accessions with a DNA query based on Agrobacterium virulence genes, as indicated in Matveeva and Otten (2019).
Phylogenetic tree construction
Cus protein sequences were aligned using MUSCLE (Edgar, 2004). Maximum likelihood phylogenetic trees were generated using IQ‐TREE2 with 1000 bootstrap replicates (Minh et al., 2020). Trees were visualized and annotated in iTOL (Letunic & Bork, 2021).
Matrix analysis
Complex repeats were analyzed by matrix analysis on the National Library of Medicine (NLM) blast site, using BLASTN with the ‘highly similar sequences’ setting; the matrix figure was recovered and analyzed for repeats.
AUTHOR CONTRIBUTIONS
KC and LO conceived and designed research. HL, YH, SC, ZH, YL, JH, and YY provided bioinformatics analysis. The first draft of the manuscript was written by LO and HL. The manuscript was revised by LO, KC, WW, and JL. All authors read and approved the final document.
CONFLICT OF INTEREST
None of the authors have a conflict of interest to disclose.
Supporting information
Figure S1. nGMO distribution on the evolutionary tree of the 425 spermatophyte families.
Figure S2. Maps of T‐DNAs in Agrobacterium/Rhizobium and of cT‐DNAs in nGMOs.
Table S1. Protein query sequences for searching cT‐DNA‐located coding sequences by the BLASTX approach.
Table S2. cT‐DNA genes in spermatophytes (gymnosperms and angiosperms), based on WGS and SRA data.
Table S3. List of cT‐DNAs ArTA to ArTD in WGS sequences of Arachis.
Table S4. DcTA types from 134 different Dioscorea alata accessions.
Table S5. Groups of cT‐DNA structures based on WGS data.
ACKNOWLEDGMENTS
This work was supported by the National Natural Science Foundation of China (32370382 to KC). It was also supported by a special fund for scientific research of Shanghai landscaping and city appearance administrative bureau (G262408 and G242406 to KC) and a special fund for scientific research of national botanical gardens to benefit sustainable development (2026 to HL). The authors have no competing interests to declare that are relevant to the content of this article. The authors declare that no human and/or animal material, data, or cell lines were used in this study. We thank Todd Blevins for discussion and comments.
Contributor Information
Léon Otten, Email: otten@unistra.fr.
Ke Chen, Email: kchen@cemps.ac.cn.
DATA AVAILABILITY STATEMENT
The datasets generated and/or analyzed during the current study are accessible via the following GitHub repository: https://github.com/HLiuprojects/nGMO‐cT‐DNA‐sequences.
REFERENCES
- Altschul, S.F. , Gish, W. , Miller, W. , Myers, E.W. & Lipman, D.J. (1990) Basic local alignment search tool. Journal of Molecular Biology, 215, 403–410. [DOI] [PubMed] [Google Scholar]
- Aoki, S. & Syōno, K. (1999) Function of Ngrol genes in the evolution of Nicotiana glauca: conservation of the function of NgORF13 and NgORF14 after ancient infection by an Agrobacterium rhizogenes‐like ancestor. Plant and Cell Physiology, 40, 222–230. [Google Scholar]
- Becker, A.L. , Crowl, A.A. , Luteyn, J.L. , Chanderbali, A.S. , Judd, W.S. , Manos, P.S. et al. (2024) A global blueberry phylogeny: evolution, diversification, and biogeography of Vaccinieae (Ericaceae). Molecular Phylogenetics and Evolution, 201, 108202. [DOI] [PubMed] [Google Scholar]
- Bogomaz, O.D. , Bemova, V.D. , Mirgorodskii, N.A. & Matveeva, T.V. (2024) Evolutionary fate of the opine synthesis genes in the Arachis L. genomes. Biology‐Basel, 13, 601. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bredeson, J.V. , Lyons, J.B. , Oniyinde, I.O. , Okereke, N.R. , Kolade, O. , Nnabue, I. et al. (2022) Chromosome evolution and the genetic basis of agronomically important traits in greater yam. Nature Communications, 13, 2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen, K. & Otten, L. (2017) Natural Agrobacterium transformants: recent results and some theoretical considerations. Frontiers in Plant Science, 8, e1600. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen, K. , Dorlhac de Borne, F. , Julio, E. , Obszynski, J. , Pale, P. & Otten, L. (2016) Root‐specific expression of opine genes and opine accumulation in some cultivars of the naturally occurring GMO Nicotiana tabacum . Plant Journal, 87, 258–269. [DOI] [PubMed] [Google Scholar]
- Chen, K. , Liu, H. , Blevins, T. , Hao, J. & Otten, L. (2023) Extensive natural Agrobacterium‐induced transformation in the genus Camellia . Planta, 258, 81. [DOI] [PubMed] [Google Scholar]
- Chen, K. , Zhurbenko, P. , Danilov, L. , Matveeva, T. & Otten, L. (2022) Conservation of an Agrobacterium cT‐DNA insert in Camellia section Thea reveals the ancient origin of tea plants from a genetically modified ancestor. Frontiers in Plant Science, 13, 997762. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Christenhusz, M.J.M. & Byng, J.W. (2016) The number of known plants species in the world and its annual increase. Phytotaxa, 261, 201–217. [Google Scholar]
- De Buck, S. , Podevin, N. , Nolf, J. , Jacobs, A. & Depicker, A. (2009) The T‐DNA integration pattern in Arabidopsis transformants is highly determined by the transformed target cell. Plant Journal, 60, 134–145. [DOI] [PubMed] [Google Scholar]
- Dong, S. , Wang, S. , Li, L. , Yu, J. , Zhang, Y. , Xue, J. et al. (2024) Bryophytes hold a larger gene family space than vascular plants. Nature Genetics, 57, 2562–2569. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Edgar, R.C. (2004) MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32, 1792–1797. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Favero, B.T. , Tan, Y. , Chen, X. , Muller, R. & Lutken, H. (2022) Kalanchoe blossfeldiana naturally transformed with Rhizobium rhizogenes exhibits superior root phenotype. Plant Science, 321, 111323. [DOI] [PubMed] [Google Scholar]
- Furner, I.J. , Huffman, G.A. , Amasino, R.M. , Garfinkel, D.J. , Gordon, M.P. & Nester, E.W. (1986) An Agrobacterium transformation in the evolution of the genus Nicotiana . Nature, 319, 422–427. [Google Scholar]
- Gelvin, S.B. (2003) Agrobacterium‐mediated plant transformation: the biology behind the “gene‐jockeying” tool. Microbiology and Molecular Biology Reviews, 67, 16–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gelvin, S.B. (2017) Integration of Agrobacterium T‐DNA into the plant genome. Annual Review of Genetics, 51, 195–217. [DOI] [PubMed] [Google Scholar]
- Group, T.A.P. , Chase, M.W. , Christenhusz, M.J.M. , Fay, M.F. , Byng, J.W. , Judd, W.S. et al. (2016) An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG IV. Botanical Journal of the Linnean Society, 181, 1–20. [Google Scholar]
- Hooykaas, P.J.J. (2023) The Ti plasmid, driver of Agrobacterium pathogenesis. Phytopathology, 113, 594–604. [DOI] [PubMed] [Google Scholar]
- Jupe, F. , Rivkin, A.C. , Michael, T.P. , Zander, M. , Motley, S.T. , Sandoval, J.P. et al. (2019) The complex architecture and epigenomic impact of plant T‐DNA insertions. PLoS Genetics, 15, e1007819. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kleinboelting, N. , Huep, G. , Appelhagen, I. , Viehoever, P. , Li, Y. & Weisshaar, B. (2015) The structural features of thousands of T‐DNA insertion sites are consistent with a double‐strand break repair‐based insertion mechanism. Molecular Plant, 8, 1651–1664. [DOI] [PubMed] [Google Scholar]
- Kralemann, L.E.M. , de Pater, S. , Shen, H. , Kloet, S.L. , van Schendel, R. , Hooykaas, P.J.J. et al. (2022) Distinct mechanisms for genomic attachment of the 5′ and 3′ ends of Agrobacterium T‐DNA in plants. Nature Plants, 8, 526–534. [DOI] [PubMed] [Google Scholar]
- Kyndt, T. , Quispe, D. , Zhai, H. , Jarret, R. , Ghislain, M. , Liu, Q. et al. (2015) The genome of cultivated sweet potato contains Agrobacterium T‐DNAs with expressed genes: an example of a naturally transgenic food crop. Proceedings of the National Academy of Sciences of the United States of America, 112, 5844–5849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Letunic, I. & Bork, P. (2021) Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Research, 49, W293–W296. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Matveeva, T.V. & Otten, L. (2019) Widespread occurrence of natural genetic transformation of plants by Agrobacterium . Plant Molecular Biology, 101, 415–437. [DOI] [PubMed] [Google Scholar]
- Matveeva, T.V. , Bogomaz, D.I. , Pavlova, O.A. , Nester, E.W. & Lutova, L.A. (2012) Horizontal gene transfer from genus Agrobacterium to the plant Linaria in nature. Molecular Plant‐Microbe Interactions, 25, 1542–1551. [DOI] [PubMed] [Google Scholar]
- Messeder, J.V.S. , Carlo, T.A. , Zhang, G. , Tovar, J.D. , Arana, C. , Huang, J. et al. (2024) A highly resolved nuclear phylogeny uncovers strong phylogenetic conservation and correlated evolution of fruit color and size in Solanum L. New Phytologist, 243, 765–780. [DOI] [PubMed] [Google Scholar]
- Minh, B.Q. , Schmidt, H.A. , Chernomor, O. , Schrempf, D. , Woodhams, M.D. , von Haeseler, A. et al. (2020) IQ‐TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution, 37, 1530–1534. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nester, E.W. (2014) Agrobacterium: nature's genetic engineer. Frontiers of Plant Science, 5, 730. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Otten, L. (2018) The Agrobacterium phenotypic plasticity (Plast) genes. Current Topics in Microbiology and Immunology, 418, 375–419. [DOI] [PubMed] [Google Scholar]
- Otten, L. (2021) T‐DNA regions from 350 Agrobacterium genomes: maps and phylogeny. Plant Molecular Biology, 106, 239–258. [DOI] [PubMed] [Google Scholar]
- Otten, L. , Liu, H. , Meeprom, N. , Linan, A. , Puglisi, C. & Chen, K. (2025) Accumulation of numerous cellular T‐DNA sequences in the genus Diospyros by multiple rounds of natural transformation. Plant Journal, 122, e70202. [DOI] [PubMed] [Google Scholar]
- Petit, A. & Tempé, J. (1985) The function of T‐DNA in nature. In: Molecular form and function of the plant genome. New York: Plenum Press, pp. 625–636. [Google Scholar]
- Pucker, B. , Kleinbolting, N. & Weisshaar, B. (2021) Large‐scale genomic rearrangements in selected Arabidopsis thaliana T‐DNA lines are caused by T‐DNA insertion mutagenesis. BMC Genomics, 22, 599. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Quispe‐Huamanquispe, D.G. , Gheysen, G. , Yang, J. , Jarret, R. , Rossel, G. & Kreuze, J.F. (2019) The horizontal gene transfer of Agrobacterium T‐DNAs into the series Batatas (genus Ipomoea) genome is not confined to hexaploid sweetpotato. Science Reports, 9, 12584. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schoch, C.L. , Ciufo, S. , Domrachev, M. , Hotton, C.L. , Kannan, S. , Khovanskaya, R. et al. (2020) NCBI taxonomy: a comprehensive update on curation, resources and tools. Database: The Journal of Biological Databases and Curation, 2020, baaa062. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shkryl, Y. , Yaroshenko, Y. , Grigorchuk, V. , Bulgakov, V. & Yugay, Y. (2025) Functional analysis of naturally integrated rol genes in sweet potato via CRISPR/Cas9 genome editing. Plants, 14, 3708. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weisberg, A.J. , Davis, E.W. , Tabima, J. , Belcher, M.S. , Miller, M. , Kuo, C.H. et al. (2020) Unexpected conservation and global transmission of agrobacterial virulence plasmids. Science, 368, eaba5256. [DOI] [PubMed] [Google Scholar]
- White, F.F. , Garfinkel, D.J. , Huffman, G.A. , Gordon, M.P. & Nester, E.W. (1983) Sequences homologous to Agrobacterium rhizogenes T‐DNA in the genomes of uninfected plants. Nature, 301, 348–350. [Google Scholar]
- Yan, M. , Li, M. , Wang, Y. , Wang, X. , Moeinzadeh, M.H. , Quispe‐Huamanquispe, D.G. et al. (2024) Haplotype‐based phylogenetic analysis and population genomics uncover the origin and domestication of sweetpotato. Molecular Plant, 17, 277–296. [DOI] [PubMed] [Google Scholar]
- Ying, W. , Wen, G. , Xu, W. , Liu, H. , Ding, W. , Zheng, L. et al. (2023) Agrobacterium rhizogenes: paving the road to research and breeding for woody plants. Frontiers in Plant Science, 14, 1196561. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang, Y. , Wang, D. , Wang, Y. , Dong, H. , Yuan, Y. , Yang, W. et al. (2020) Parasitic plant dodder (Cuscuta spp.): a new natural Agrobacterium‐to‐plant horizontal gene transfer species. Science China. Life Sciences, 63, 312–316. [DOI] [PubMed] [Google Scholar]
- Zhidkin, R. , Zhurbenko, P. , Bogomaz, O. , Gorodilova, E. , Katsapov, I. , Antropov, D. et al. (2023) Biodiversity of rolB/C‐like natural transgene in the genus Vaccinium L. and its application for phylogenetic studies. International Journal of Molecular Sciences, 24, 6923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhu, J. , Oger, P.M. , Schrammeijer, B. , Hooykaas, P.J. , Farrand, S.K. & Winans, S.C. (2000) The bases of crown gall tumorigenesis. Journal of Bacteriology, 182, 3885–3895. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figure S1. nGMO distribution on the evolutionary tree of the 425 spermatophyte families.
Figure S2. Maps of T‐DNAs in Agrobacterium/Rhizobium and of cT‐DNAs in nGMOs.
Table S1. Protein query sequences for searching cT‐DNA‐located coding sequences by the BLASTX approach.
Table S2. cT‐DNA genes in spermatophytes (gymnosperms and angiosperms), based on WGS and SRA data.
Table S3. List of cT‐DNAs ArTA to ArTD in WGS sequences of Arachis.
Table S4. DcTA types from 134 different Dioscorea alata accessions.
Table S5. Groups of cT‐DNA structures based on WGS data.
Data Availability Statement
The datasets generated and/or analyzed during the current study are accessible via the following GitHub repository: https://github.com/HLiuprojects/nGMO‐cT‐DNA‐sequences.
