Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 17;127(4):e71087. doi: 10.1111/tpj.71087

Embryophyte‐wide detection of natural Agrobacterium‐mediated horizontal gene transfer reveals an ancient role for mini T‐DNAs

Hai Liu 1,2, Weiqing Wang 1,2, Jiani Li 1,2, Yaxi He 3, Shengyuan Cao 3, Zenghui Hu 4, Yuanyuan Liu 2, Jie Hao 2,3, Yuehong Yan 1,3, Léon Otten 5,, Ke Chen 1,2,3,
PMCID: PMC13480146  PMID: 42606503

SUMMARY

Agrobacterium transfers DNA into plant cells, leading to tumors, hairy roots (HR), and natural genetically modified organisms (nGMOs). Transferred DNAs (T‐DNAs) from agrobacteria and T‐DNA‐derived cellular T‐DNAs (cT‐DNAs) from nGMOs vary considerably and may carry up to 15 different genes. Among these, opine synthase (ops) genes encode the synthesis of opines used as nutrients by the agrobacteria. Earlier studies predicted large numbers of naturally transformed plant species, but only few have been identified and studied so far. We therefore developed a general method to detect cT‐DNAs in all publicly available whole genome sequences (WGS) and Sequence Read Archive (SRA) data from land plants. To avoid false positives, we only retained DNA sequences coding for T‐DNA proteins. A total of 2614 nGMO species were identified, most are eudicots. However, cT‐DNAs were also found in 82 mosses and 75 ferns, showing that Agrobacterium can also generate natural transformants among the early land plants. Analysis of 149 cT‐DNA maps revealed different types of T‐DNAs. Most notably, these included small T‐DNAs (mini T‐DNAs) with a single opine synthase gene. Mini T‐DNAs are not expected to induce tumors or HRs. The predominance of mini cT‐DNAs in mosses and ferns, and the presence of more complex cT‐DNAs in spermatophytes, indicate that mini T‐DNAs represent the earliest types of T‐DNA. Our study also detected unusual T‐DNA integration patterns, with multiple copies spread out over several hundreds of kilobases.

Keywords: horizontal gene transfer, nGMOs, Agrobacterium , mini T‐DNAs

Significance Statement

Horizontal gene transfer from Agrobacterium has contributed to the genomes of numerous plant species, but its frequency and evolutionary history across land plants have remained poorly understood. By systematically mining whole‐genome and Sequence Read Archive data across embryophytes, we identified 2614 natural genetically modified plant species, greatly expanding the known diversity of natural Agrobacterium‐mediated transformation. Importantly, we detected cellular T‐DNAs in mosses, liverworts, hornworts, and ferns, indicating that Agrobacterium‐mediated transformation is not restricted to flowering plants and may date to early land plant evolution. Comparative analysis further revealed a widespread class of mini T‐DNAs, often carrying a single opine synthase gene, suggesting that these simple T‐DNAs may represent an ancient form of T‐DNA and providing new insight into the evolution of the Agrobacterium–plant interaction. This work establishes a broad framework for investigating the evolutionary, ecological, and functional significance of natural T‐DNA integration in plants.

INTRODUCTION

Agrobacterium‐mediated horizontal gene transfer (HGT) involves a unique mechanism of DNA transfer, leading to stable insertion in the nuclear genome. The transferred DNA (T‐DNA) typically results in hairy roots (HR) or crown gall tumors (Gelvin, 2017; Hooykaas, 2023; Nester, 2014; Zhu et al., 2000). T‐DNAs can carry from one up to 15 genes (Otten, 2021; Weisberg et al., 2020), which encode growth induction and modification, and opine synthesis (opine synthase or ops genes). Opines are small molecules used by agrobacteria as nutrients. It has been proposed that the induction of opine synthesis is the driving force behind the evolution of the T‐DNA transfer mechanism (Petit & Tempé, 1985). A number of plant species contain T‐DNAs (cellular T‐DNAs or cT‐DNAs) and are called natural genetically modified organisms (nGMOs) or natural transformants. Among these, several Nicotiana species (Furner et al., 1986; White et al., 1983), common toadflax (Linaria vulgaris; Matveeva et al., 2012), and Ipomoea (Kyndt et al., 2015) have been reported. A first systematic study using whole genome sequences (WGS) detected 23 nGMO species (Matveeva & Otten, 2019), and by extrapolation predicted about 10 000 naturally transformed plant species. Subsequent studies investigated four genera in detail: Camellia (Chen et al., 2022, 2023), Vaccinium (Zhidkin et al., 2023), Arachis (Bogomaz et al., 2024), and Diospyros (Otten et al., 2025). These studies, contrary to the 2019 study, included SRA and WGS data (SRA+WGS search). Although SRA data are less complete than WGS data, they can reveal significantly more nGMOs than WGS data alone. We therefore adopted the SRA+WGS approach for a search targeting all of the embryophytes (14 000 instead of four genera). Comparative analysis from the vast amount of new data allowed us to gain more insight into the history of Agrobacterium‐mediated HGT events, most notably through the discovery of small T‐DNAs with a single T‐DNA gene (mini T‐DNAs) and their distribution across vascular and nonvascular plants.

RESULTS

Searching cT‐DNAs in WGS and SRA data from land plants

The first systematic search for nGMOs (Matveeva & Otten, 2019) in angiosperms used WGS data, but no SRA data. Presently there are 3930 Terabytes of SRA data for land plants. We therefore devised a method to extract putative land plant cT‐DNA sequences from both WGS and SRA data (‘Materials and Methods’ section; Figure 1), and retained those corresponding to protein‐coding regions, using a set of T‐DNA/cT‐DNA proteins (Table S1) as query. Particular cT‐DNA inserts will be named by using the first two letters of the genus in which they were first discovered, followed by T (for cT‐DNA) and a letter for the insert, starting with A. Thus, CaTA is the first cT‐DNA insert (A) discovered in Camellia, DiTD the fourth one (D) from Diospyros. cT‐DNA alleles are indicated by a number (like DiTD‐2). Older published names, like IbT‐DNA1 (for Ipomoea batatas T‐DNA 1) are kept to avoid confusion. WGS contig numbers are from the National Center for Biotechnology Information (NCBI), but with simplified numbering. For example, JBNGMI010000139.1 becomes JBNGMI‐139.

Figure 1.

Figure 1

Overview of the process of identifying new natural genetically modified organisms (nGMOs).

All available land plant whole genome sequences (WGS) and Sequence Read Archive (SRA) datasets were searched using T‐DNA nucleotide (NTfuse6) and protein (Table S1) queries. Positive reads were assembled into contigs and subjected to BLASTX analysis to identify cT‐DNA genes. Potential nGMOs were further filtered to exclude false positives (e.g., hairy root lines, distant bacterial and plant homologs). In addition, to exclude potential Agrobacterium contamination, all positive WGS and SRA accessions were further screened using queries based on several Agrobacterium virulence (vir) genes, following the approach described in Matveeva and Otten (2019). Verified cases were compared with published nGMOs to distinguish newly identified nGMOs from old ones. cT‐DNA maps from old nGMOs and T‐DNA maps from Rhizobium/Agrobacterium were redrawn at the same format (Figure S2A). Newly identified nGMOs were mapped from WGS data (Figure S2B). SRA data were summarized according to their T‐DNA genes (Table S2). Because of the very large number of publicly available SRA data for monocots, Fabaceae, Solanaceae, and Brassicaceae, a maximum of 20 SRR accessions per species were selected for analysis, the Dioscoreaceae family was analyzed using all available SRR accessions.

Nonvascular embryophytes

Nonvascular embryophytes consist of Bryophyta (mosses, 12 700 species), Anthocerotophyta (hornworts, 225), and Marchantiophyta (liverworts, 9000). Species numbers in this paper are from Christenhusz & Byng, (2016). Analysis of the corresponding WGS and SRA data from the three groups yielded 77, one and four nGMOs, respectively. These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table 1. All, except four, contain a cucumopine synthase (cus) gene. WGS sequences from Ptychostomum pallens, Polytrichum commune, and Aerobryopsis subdivergens show direct cT‐DNA repeats, separated by 12, 12, and 32 kb, respectively. Hypopterygium flavolimbatum cT‐DNA HyTA (JBNGMI‐139:498 534–505 091, 6558 nt) could be defined by comparison with JBNGMJ from the cT‐DNA‐less Hypopterygium elatum. JBNGMI‐139:494 797–498 533 aligns with JBNGMJ‐4450:251 913–255 682, and JBNGMI‐139:505 092–505 245 with JBNGMJ‐4450:257 557–257 708. Thus, HyTA caused a 1875 nt deletion upon insertion.

Table 1.

nGMO species from nonvascular embryophytes (A: Anthocerotophyta, B: Bryophyta, M: Marchantiophyta), and ferns (L: Lycopodiophyta and P: Pteridophyta). Data are from WGS and SRA data

Order Family Genus and species Genes WGS/SRR
1 Ferns L Isoetales Isoetacea Isoetes valida ags SRR25755400, SRR25755356
2 Ferns L Lycopodiales Lycopodiaceae Huperzia selago a SRR28958735, SRR28958738
3 Ferns L Lycopodiales Lycopodiaceae Huperzia serrata b SRR28958736, SRR18149617
4 Ferns L Selaginellales Selaginellaceae Selaginella erythropus mas2′, mas1′, ags SRR15348965, SRR15348964, SRR15348963
5 Ferns L Selaginellales Selaginellaceae Selaginella martensii vis SRR31771695, SRR31771696, SRR31771697
6 Ferns P Polypodiales Aspleniaceae Asplenium alternifolium susL ERR14012285
7 Ferns P Polypodiales Aspleniaceae Asplenium komarovii mas1′, ags SRR14381415, SRR8772343
8 Ferns P Polypodiales Aspleniaceae Asplenium trichomanes subsp. trichomanes susL ERR14009220, ERR5554902, ERR5232099
9 Ferns P Polypodiales Aspleniaceae Asplenium trichomanes subsp. inexpectans susL ERR5529585
10 Ferns P Polypodiales Aspleniaceae Asplenium trichomanes subsp. quadrivalens susL ERR14012664
11 Ferns P Polypodiales Aspleniaceae Asplenium × adulterinum susL ERR14047057, ERR5529300
12 Ferns P Polypodiales Aspleniaceae Hymenasplenium sp. mas1 SRR6920712
13 Ferns P Polypodiales Athyriaceae Deparia giraldii susL SRR31694523
14 Ferns P Polypodiales Athyriaceae Deparia lancea ags SRR6899396
15 Ferns P Polypodiales Athyriaceae Diplazium unilobum ags SRR16974227
16 Ferns P Polypodiales Blechnaceae Blechnopsis orientalis cus SRR28681840, SRR27869072
17 Ferns P Polypodiales Blechnaceae Stenochlaena palustris vis ERR12670167, ERR12670177
18 Ferns P Cyatheales Cyatheaceae Alsophila costularia ags SRR27937298
19 Ferns P Cyatheales Cyatheaceae Alsophila latebrosa vis, ags SRR19887278
20 Ferns P Cyatheales Cyatheaceae Alsophila metteniana ags SRR19887263
21 Ferns P Cyatheales Cyatheaceae Alsophila spinulosa c, ags, orf2, rolA SRR28161632, SRR28161616, SRR28161618
22 Ferns P Cyatheales Cyatheaceae Cyathea glabra ags SRR19887387
23 Ferns P Cyatheales Cyatheaceae Gymnosphaera acs, ags SRR11076226, SRR11075945, SRR19887515
24 Ferns P Cyatheales Cyatheaceae Gymnosphaera andersonii acs SRR19887516
25 Ferns P Polypodiales Davalliaceae Davallia denticulata cus ERR12670369
26 Ferns P Polypodiales Dennstaedtiaceae Dennstaedtia scandens acs SRR22250942
27 Ferns P Cyatheales Dicksoniaceae Dicksonia lanata ocs SRR8580726
28 Ferns P Polypodiales Dryopteridaceae Arachniodes davalliiformis cus DRR591047
29 Ferns P Polypodiales Dryopteridaceae Dryopteris affinis e ERR14042545
30 Ferns P Polypodiales Dryopteridaceae Dryopteris campyloptera vis SRR9050854
31 Ferns P Polypodiales Dryopteridaceae Dryopteris celsa cus SRR14320985
32 Ferns P Polypodiales Dryopteridaceae Dryopteris crassirhizoma acs SRR13447702
33 Ferns P Polypodiales Dryopteridaceae Dryopteris filix‐mas e SRR12518786
34 Ferns P Polypodiales Dryopteridaceae Dryopteris intermedia ocs SRR9050845
35 Ferns P Polypodiales Dryopteridaceae Dryopteris keysseriana acs SRR18497099
36 Ferns P Polypodiales Dryopteridaceae Dryopteris marginalis mis,e SRR11229720
37 Ferns P Polypodiales Dryopteridaceae Dryopteris pseudocaenopteris mas2 SRR2103701
38 Ferns P Polypodiales Dryopteridaceae Dryopteris remota vis ERR14047111
39 Ferns P Polypodiales Dryopteridaceae Dryopteris tyrrhena vis ERR14041301
40 Ferns P Polypodiales Dryopteridaceae Lomagramma matthewii susL SRR2103704
41 Ferns P Polypodiales Dryopteridaceae Polystichum acrostichoides vis SRR18053988, JAOYMV01191549
42 Ferns P Polypodiales Dryopteridaceae Polystichum aculeatum mas1 ERR14030239
43 Ferns P Polypodiales Dryopteridaceae Rumohra adiantiformis mas1′, ags SRR30618658
44 Ferns P Hymenophyllales Hymenophyllaceae Abrodictyum obscurum orf14, orf511, susL ERR12670379, ERR12670380, SRR32258258
45 Ferns P Schizaeales Lygodiaceae Lygodium japonicum cus SRR29127758
46 Ferns P Schizaeales Lygodiaceae Lygodium palmatum orf14 SRR26863426
47 Ferns P Schizaeales Lygodiaceae Lygodium salicifolium acs SRR7121778
48 Ferns P Maratttiales Marattiaceae Danaea nodosa mas1′, mas2 ERR2041196
49 Ferns P Polypodiales Onocleaceae Onoclea sensibilis var. interrupta acs SRR31693313
50 Ferns P Polypodiales Polypodiaceae Goniophlebium amoenum cus SRR8185331
51 Ferns P Polypodiales Polypodiaceae Lemmaphyllum drymoglossoides acs SRR23238769
52 Ferns P Polypodiales Polypodiaceae Loxogramme chinensis mas2 SRR2103729
53 Ferns P Polypodiales Polypodiaceae Platycerium wallichii susL SRR14802567
54 Ferns P Polypodiales Polypodiaceae Pyrrosia piloselloides ocs SRR26086252
55 Ferns P Polypodiales Pteridaceae Adiantum nelumboides see map 147 to 149 JAKNSL020005917.1
56 Ferns P Polypodiales Pteridaceae Adiantum reniforme mas1′, mas2 SRR32258808
57 Ferns P Polypodiales Pteridaceae Antrophyum callifolium susL SRR2103739
58 Ferns P Polypodiales Pteridaceae Antrophyum formosanum susL SRR32258181, SRR32212872
59 Ferns P Polypodiales Pteridaceae Ceratopteris pteridoides cus SRR23096549
60 Ferns P Polypodiales Pteridaceae Ceratopteris richardii vis, cus JAIKUY010002021, 2022, 938
61 Ferns P Polypodiales Pteridaceae Ceratopteris thalictroides vis SRR16588741
62 Ferns P Polypodiales Pteridaceae Cyrtomium falcatum vis SRR6727965
63 Ferns P Polypodiales Pteridaceae Haplopteris elongata cus SRR14048925
64 Ferns P Polypodiales Pteridaceae Haplopteris ensiformis cus SRR20678395
65 Ferns P Polypodiales Pteridaceae Haplopteris heterophylla cus SRR6920718
66 Ferns P Polypodiales Pteridaceae Hemionitis arifolia mas1 SRR2103738
67 Ferns P Polypodiales Pteridaceae Parahemionitis cordata mas1 ERR2040933
68 Ferns P Polypodiales Pteridaceae Pteris biaurita mas1′, mas2 SRR7121768
69 Ferns P Salviniales Salviniaceae Salvinia molesta mas1′, ags SRR23404241
70 Ferns P Schizaeales Schizaeaceae Actinostachys digitata cus, susL, orf14, orf511 ERR12670432, ERR12670423, ERR12670425
71 Ferns P Polypodiales Tectariaceae Tectaria polymorpha vis SRR2103745
72 Ferns P Polypodiales Thelypteridaceae Abacopteris gymnopteridifrons vis SRR22806495
73 Ferns P Polypodiales Thelypteridaceae Amblovenatum opulentum vis SRR26157139, SRR26157140
74 Ferns P Polypodiales Thelypteridaceae Grypothrix megacuspis acs SRR18496617
75 Ferns P Polypodiales Thelypteridaceae Gymnocarpium robertianum orf8 ERR14029864
76 Ferns P Polypodiales Thelypteridaceae Thelypteris parasitica vis SRR7121612
77 Ferns P Polypodiales Woodsiaceae Matteuccia struthiopteris vis SRR31694452
78 Mosses A Anthocerotales Anthocerotaceae Folioceros fuciformis mas1 SRR29281613
79 Mosses B Hypnales Amblystegiaceae Drepanocladus polygamus cus SRR26398964
80 Mosses B Hypnales Amblystegiaceae Hygrohypnum luridum cus SRR26398954
81 Mosses B Hypnales Amblystegiaceae Sciaromiopsis sinensis cus SRR15179250
82 Mosses B Hypnales Anomodontaceae Pseudanomodon attenuatus cus SRR33208610, SRR33208612, SRR33208608
83 Mosses B Bryales Bartramiaceae Bartramia ithyphylla cus SRR33208738
84 Mosses B Hypnales Brachytheciaceae Brachythecium laetum cus SRR33208359, SRR33208361, SRR33208632
85 Mosses B Hypnales Brachytheciaceae Bryhnia novae‐angliae cus SRR33208521, SRR33208529, SRR33208528
86 Mosses B Hypnales Brachytheciaceae Kindbergia praelonga cus SRR33208255, SRR33208261, SRR33208260
87 Mosses B Hypnales Brachytheciaceae Myuroclada maximowiczii cus SRR13605954
88 Mosses B Bryales Bryaceae Anomobryum julaceum cus SRR24580469
89 Mosses B Bryales Bryaceae Bryum yuennanense cus SRR31709169
90 Mosses B Bryales Bryaceae Ptychostomum cyclophyllum cus ERR15381807
91 Mosses B Bryales Bryaceae Ptychostomum pallens cus ERR15378477, ERR15381808, OZ377804.1
92 Mosses B Hypnales Calliergonaceae Sarmentypnum sarmentosum cus ERR13725965
93 Mosses B Hypnales Catagoniaceae Catagonium nitens cus SRR33208480, SRR33208469, SRR33208435
94 Mosses B Hypnales Climaciaceae Climacium americanum cus JBNGMP
95 Mosses B Hypnales Climaciaceae Climacium dendroides cus ERR11242531, ERR10934078, ERR10934077
96 Mosses B Dicranales Dicranaceae Campylopus introflexus cus ERR15378467, ERR15381794
97 Mosses B Dicranales Ditrichaceae Pleuridium rhynchostegium cus DRR379964
98 Mosses B Encalyptales Encalyptaceae Encalypta ciliata cus ERR15659579, SRR33208205
99 Mosses B Hypnales Entodontaceae Entodon concinnus cus JBNGNB
100 Mosses B Dicranales Fissidentaceae Fissidens adianthoides cus ERR15381760
101 Mosses B Funariales Funariaceae Physcomitrellopsis africana cus SRR26586950, SRR26596311, SRR26596310
102 Mosses B Hedwigiales Hedwigiaceae Hedwigia ciliata cus ERR5232302
103 Mosses B Hypnales Hylocomiaceae Hylocomiadelphus triquetrus cus ERR9793174, ERR9793175, ERR9793176
104 Mosses B Hypnales Hylocomiaceae Hylocomium splendens cus SRR2518082, SRR25553778, SRR25553779
105 Mosses B Hypnales Hylocomiaceae Loeskeobryum brevirostre cus ERR15551479, ERR13731937, ERR13725996
106 Mosses B Hypnales Hylocomiaceae Pleurozium schreberi cus SRR2513357
107 Mosses B Hypnales Hylocomiaceae Rhytidiadelphus loreus cus ERR6895899, ERR6895898, ERR6895897
108 Mosses B Hypnales Hylocomiaceae Rhytidiadelphus subpinnatus cus JBNGMQ
109 Mosses B Hypnales Hylocomiaceae Rhytidiopsis robusta cus SRR2518096
110 Mosses B Hypnales Hypnaceae Hyocomium armoricum cus ERR13389723, ERR13382531
111 Mosses B Hypopterygiales Hypopterygiaceae Hypopterygium fauriei cus SRR31755810
112 Mosses B Hypopterygiales Hypopterygiaceae Hypopterygium flavolimbatum cus JBNGMI
113 Mosses B Hypnales Leucodontaceae Antitrichia curtipendula cus SRR2518092
114 Mosses B Hypnales Meteoriaceae Aerobryopsis subdivergens cus JBNGMS
115 Mosses B Leucodontales Meteoriaceae Barbella flagellifera cus SRR23095850
116 Mosses B Bryales Mniaceae Plagiomnium ciliare cus SRR33208530
117 Mosses B Leucodontales Neckaraceae Homaliodendron scalpellifolium cus JBNGMX
118 Mosses B Orthotrichales Orthotrichaceae Orthotrichum anomalum cus SRR33208788
119 Mosses B Orthotrichales Orthotrichaceae Orthotrichum diaphanum cus ERR15745620
120 Mosses B Orthotrichales Orthotrichaceae Plenogemma phyllantha cus ERR15405273
121 Mosses B Polytrichales Polytrichaceae Oligotrichum hercynicum cus ERR15610518
122 Mosses B Polytrichales Polytrichaceae Pogonatum subfuscatum cus SRR33208732
123 Mosses B Polytrichales Polytrichaceae Polytrichastrum ohioense cus SRR33208602, JBNGKP
124 Mosses B Polytrichales Polytrichaceae Polytrichum commune cus JBNGKQ
125 Mosses B Polytrichales Polytrichaceae Polytrichum formosum cus ERR12721078, ERR12721079, SRR8707304
126 Mosses B Polytrichales Polytrichaceae Polytrichum juniperinum cus ERR15606502, ERR15610523
127 Mosses B Polytrichales Polytrichaceae Polytrichum piliferum cus ERR15378475, ERR15381805, ERR15381806
128 Mosses B Polytrichales Polytrichaceae Polytrichum strictum cus JBNGKO
129 Mosses B Pottiales Pottiaceae Anoectangium aestivum cus ERR15656951, ERR15659573
130 Mosses B Pottiales Pottiaceae Bryoerythrophyllum caledonicum cus ERR12370311, ERR12356316
131 Mosses B Pottiales Pottiaceae Bryoerythrophyllum recurvirostrum cus SRR24580705
132 Mosses B Pottiales Pottiaceae Hymenostylium aurantiacum cus JBNGLN
133 Mosses B Pottiales Pottiaceae Hyophila propagulifera cus DRR428591
134 Mosses B Pottiales Pottiaceae Syntrichia ruralis cus SRR27955587
135 Mosses B Hypnales Pseudoleskeaceae Lescuraea plicata cus SRR24580403
136 Mosses B Hypnales Regmadontaceae Regmatodon declinatus cus SRR33208207, SRR33208208, SRR33208209
137 Mosses B Rhabdoweisiales Rhabdoweisiaceae Dicranoweisia cirrata cus ERR15381799
138 Mosses B Sphagnales Sphagnaceae Sphagnum affine cus SRR18184152
139 Mosses B Sphagnales Sphagnaceae Sphagnum angustifolium cus SRR6966991, SRR6968021, SRR6966992
140 Mosses B Sphagnales Sphagnaceae Sphagnum capillifolium cus SRR6965943, SRR6965944, SRR6965384
141 Mosses B Sphagnales Sphagnaceae Sphagnum compactum cus SRR6973028, ERR4781406, ERR4778815
142 Mosses B Sphagnales Sphagnaceae Sphagnum contortum cus ERR12259833
143 Mosses B Sphagnales Sphagnaceae Sphagnum divinum cus SRR18184163, SRR18184156
144 Mosses B Sphagnales Sphagnaceae Sphagnum fallax cus SRR18184003, SRR18183987
145 Mosses B Sphagnales Sphagnaceae Sphagnum fuscum cus SRR6966970, SRR6967721, SRR6966971
146 Mosses B Sphagnales Sphagnaceae Sphagnum lindberghii cus ERR15126493
147 Mosses B Sphagnales Sphagnaceae Sphagnum magellanicum cus SRR22501966, SRR22501971, SRR22501943
148 Mosses B Sphagnales Sphagnaceae Sphagnum palustre cus SRR6964393, ERR13650026
149 Mosses B Sphagnales Sphagnaceae Sphagnum papillosum cus SRR6966183, SRR6966184
150 Mosses B Sphagnales Sphagnaceae Sphagnum portoricense cus SRR18427469
151 Mosses B Sphagnales Sphagnaceae Sphagnum rubellum cus SRR6966977, SRR6965089, SRR6965088
152 Mosses B Sphagnales Sphagnaceae Sphagnum squarrosum cus SRR6973636, SRR6973635, ERR4780966
153 Mosses B Splachnales Splachnaceae Tayloria subglabra cus SRR33208468
154 Mosses B Hypnales Thuidiaceae Thuidium delicatulum cus ERR12374262
155 Mosses B Hypnales Thuidiaceae Thuidium tamariscinum cus ERR6895937, ERR6895936, ERR6895935
156 Mosses M Lophoziales Anastrophyllaceae Neoorthocaulis attenuatus plast ERR14995244, ERR15165952
157 Mosses M Porellales Porellaceae Porella caespitans cus JBNGJO
158 Mosses M Jungermanniales Solenostomataceae Solenostoma hyalinum acs SRR25319279
159 Mosses M Jungermanniales Solenostomataceae Solenostoma erectum ags SRR33208228
a

acs‐orf2‐orf3n‐orf8‐rolB‐orf14‐mas2‐mas1‐ags‐vis.

b

acs‐orf2‐orf3n‐orf8‐rolA‐rolB‐1215like‐mas1‐ags‐cus.

Earlier we have shown for Camellia (Chen et al., 2023) and Diospyros (Otten et al., 2025) that cT‐DNA insertions can precede speciation. The Bryophyta provide another example of this. The P. commune assembly JBNGKQ‐5857 shows a small cT‐DNA with a cus gene. Sequences from the NCBI core_nt collection from Atrichum undulatum (no cT‐DNA), and numerous moss sequences from the WGS database show that the insert (called PoTA) is limited by coordinates JBNGKQ‐5857:1 251 718–1 252 882. The same insert is also found in Polytrichastrum ohioense (JBNGKP‐18738) and Polytrichastrum strictum (JBNGKO‐15772). DNA sequence homology between the three Polytrichum species extends well beyond PoTA, showing that all three have the same PoTA insert. Thus, PoTA was inserted before speciation. Conversely, Entodon concinnus JBNGNB‐6 shows 85% identity with PoTA JBNGKQ‐5857:1 251 715–1 252 823, but diverges on both sides. Therefore, its insert (EnTA) differs from PoTA. Similarly, Homaliodendron scalpellifolium JBNGMX‐38774 has a cT‐DNA similar to those of PoTA and EnTA, but diverges beyond, and is therefore still another insert (HoTA). The present data represent many new opportunities for pre‐speciation insertion analysis, which remain to be explored. A phylogenetic tree of the predicted Cus proteins shows three distinct groups, only found in mosses and ferns (Figure 2).

Figure 2.

Figure 2

Phylogenetic tree for Cus proteins from ferns (green), mosses (red), eudicots (yellow), Agrobacterium/Rhizobium (magenta), other bacteria (purple), and fungi (blue).

I, II, III: Cus groups specific for mosses and ferns. Numbers on branches: bootstrap values. Reconstr: reconstructed from fragments, part: partial sequence.

A search of the clusteredNR protein collection on the NCBI BLASTP site, using all available Cus sequences as queries, also detected Cus sequences in seven fungal species. As in the case of mosses and ferns, this small number of species suggests that these cus genes did not originate from fungi but were acquired from Agrobacterium. This will require further investigation.

The cT‐DNA sequences in Bryophyta, Anthocerotophyta, and Marchantiophyta show that Agrobacterium can transform nonvascular plants, and that these plant species can regenerate into nGMOs under natural conditions. We next investigated the Lycopodiophyta and Polypodiophyta.

Lycopodiophyta and Polypodiophyta

Vascular embryophytes comprise the Lycopodiophyta (1290), Polypodiophyta (10 560), and Spermatophyta (296 463). We found five nGMOs in Lycopodiophyta and 70 in Polypodiophyta (Table 1). These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table 1. WGS data were used to construct cT‐DNA maps. T‐DNA/cT‐DNA maps in this paper are numbered in bold from 1 to 166; genes oriented from right to left T‐DNA border are indicated in bold. Fern cT‐DNA maps are shown in 145149 (Figure 3a). Remarkably, contigs JAIKUY‐2021 (145) and JAIKUY‐2022 (146) carry 8 and 17 vitopine synthase (vis) genes, spread out over 245 and 423 kb, respectively. 2021‐vis2, 2021‐vis5, and 2022‐vis1 ORFs are intact. The 2022 vis copies are part of larger repeats (Figure 3b), suggesting amplification after insertion. Ceratopteris richardii JAIKUY‐938 (not shown) contains a single cus gene. Adiantum nelumboides (Figure 3a) contains AdTA‐1 (147), AdTA‐2 (148), and AdTB (149). The AdTA‐1 cT‐DNA sequences are interspersed with 144 kb of plant DNA (containing a 16 kb repeat), which is highly unusual for cT‐DNAs. AdTA‐2 is a short version of AdTA‐1. AdTB‐ipt is distantly related to AdTA‐ipt (64% identity) and probably derived from another T‐DNA. AdTA and AdTB genes ipt, 6b, vis, and d are unusual for Rhizobium rhizogenes, but common in Agrobacterium tumefaciens. The only intact AdTA gene is susL (L,L‐succinamopine synthase). These data show that ferns can be transformed by Agrobacterium and generate nGMOs in nature. We next investigated the spermatophytes.

Figure 3.

Figure 3

Fern cT‐DNA maps.

(a) Ceratopteris richardii (JAIKUY‐2021, 145), Ceratopteris richardii (JAIKUY‐2022, 146), Adiantum nelumboides, AdTA‐1 (JAKNSL‐5917, 147), Adiantum nelumboides, AdTA‐2 (JAKNSL‐0548, 148), and Adiantum nelumboides, AdTB (JAKNSL‐7169, 149). Repeats in 147 are indicated by red arrows. *: intact ORF.

(b) Repeat analysis for 17 vis copies from JAIKUY‐2022:1–450 000. Matrix analysis shows the extent of the repeats. Top: Query: 2022:1–12 000 (containing vis1). Long repeats occur around vis2, vis8, vis10, smaller ones around vis3, vis5, vis7, vis9, vis11, and vis12, and very small ones around vis4, vis6, and v13 to vis17. Bottom: Query: 2022:415 000–440 000 (containing an incomplete vis gene, vis17): small repeats for vis1 to vis3, vis5, vis7 to vis12, larger repeats for vis13 to vis17. Thus, the inserts can be divided into two groups: vis1 to vis12, and vis13 to vis17. vis coding sequences in red.

nGMOs in spermatophytes

Spermatophytes consist of gymnosperms (1079) and angiosperms (295 403). Their nGMO data are shown in Table S2. In gymnosperms, we found five nGMOs, in angiosperms 2452. The Angiosperm Phylogeny Group (APG; Group et al., 2016) recognizes 64 orders and 416 families. They comprise basal angiosperms: Amborellales, Nymphaeales, and Austrobaileyales (183 species) and core angiosperms (or Mesangiospermae): magnoliids, Chloranthales, monocots, Ceratophyllales, and eudicots (295 220 species).

Two nGMOs were detected in basal angiosperms, 68 in monocots (74 300 species), 2338 in eudicots (210 000 species), and 44 in the remaining groups (10 920 species). These species, their taxonomic affiliations, and the type of cT‐DNA genes they contain are shown in Table S2. 2270/2452 angiosperm nGMOs were exclusively found in SRA data, 13 times more than from WGS searches. Only 219/2452 angiosperm nGMOs were reported previously (Table S2, marked in bold). Thus, the present results represent an 11‐fold increase. A distribution of the nGMOs over the spermatophyte families and their genera is shown in Figure S1.

Estimation of nGMO species among land plants

The large increase in nGMO numbers compared to the 2019 study (Matveeva & Otten, 2019), allows a better estimate of the total number of nGMO species among land plants. For this, we mainly used WGS data, as most SRA sequences are incomplete. The present NCBI data contain 2435 WGS sequences for land plants (after correction for multiple WGS data from the same species), 195 of which (8%) are nGMOs. However, 105 additional species without cT‐DNAs in their WGS sequences do show cT‐DNA sequences in non‐WGS data from other accessions of the same species, indicating their cT‐DNA sequences are not fixed. This leads to a new estimation of 300/2435 species (12.3%), which extrapolates to 40 700 nGMOs among the 330 200 land plant species.

The new nGMO data were used to investigate five areas of interest: extension of previous nGMO studies, tracing the origin of the cT‐DNA in Vaccinium, identification of nGMOs used by humans, presence of nGMOs in eudicot sister groups, and comparative structural analysis of fully mapped cT‐DNAs.

Data presented in the next three subchapters are shown in Figure S2 and Table S2.

Extension of previous nGMO studies

The present data extend the results of previous detailed studies on the genera Nicotiana, Linaria, Vaccinium, Ipomoea, Camellia, Diospyros, and Arachis.

In Nicotiana, we found 13 additional nGMO species. A new cT‐DNA (TF, 23) was found in N. otophora (SRR25726134, SRR35059995). In Linaria, three additional nGMOs were found: L. angustissima, L. japonica, and L. repens. Linaria vulgaris cT‐DNA maps were constructed for EU735069.2 (28) and WGS CAOJCA (29, 30). Another well‐studied genus is Vaccinium. Vaccinium macrocarpon (cranberry) contains a single plast‐like gene (1215‐like or rolB/rolC). rolB/rolC was reported in 26 Vaccinium species and Agapetes serpens (Zhidkin et al., 2023). Our present study revealed 1215‐like sequences in as many as 28/46 genera from the Vaccinioideae subfamily (see also next chapter).

Two cT‐DNAs (IbT‐DNA1 and IbT‐DNA2) were reported in five Ipomoea species (Kyndt et al., 2015; Quispe‐Huamanquispe et al., 2019; Yan et al., 2024). Here we report 53 additional Ipomoea nGMO species, and a new cT‐DNA (IbT‐DNA3) in I. batatas and I. trifida. Ipomoea cT‐DNAs are shown in 3146.

Camellia has been reported to contain 71 nGMO species (containing different combinations of cT‐DNAs CaTA to CaTN; Chen et al., 2023). The cT‐DNA ancestors, except those of CaTJ and CaTK (both incomplete), are shown in 4758. Sixty‐eight new Camellia nGMO species were detected in the present study.

Diospyros (persimmon) has 39 nGMOs (Otten et al., 2025). This study adds D. apiculata, D. armata, D. hoyleana, D. major, and D. sonorae.

The tetraploid species Arachis hypogaea (peanut, 2n = 4x = 40) carries ancestor genomes A and B, with cus and mas2′ sequences (Matveeva & Otten, 2019). Twenty‐three Arachis nGMO species were identified (Bogomaz et al., 2024). The authors noted an intact cus gene, cus gene remnants and mas2′ sequences (A. hypogaea), ags (A. macedoi, A. pusilla), and mas1′ (A. appressipila, A. rigonii). We detected five additional species, used WGS SDMP (A. hypogaea) to recover cT‐DNAs ArTA, ArTB, ArTC, and ArTD (163166), and determined the distribution of these four cT‐DNAs among various Arachis species (Table S3).

Tracing the origin of the Vaccinium VaTA insert

The unexpected presence of 1215‐like sequences in a large number of Ericaceae species belonging to different genera raised the question whether all were derived from the same transformation event or not. In order to answer this question, two kinds of information are needed: defining the limits of the insert, and the nature of the flanking plant sequences. We first determined the extent of the cT‐DNA insert in Vaccinium ovatum (assembly JBVKUD‐2), by identifying the surrounding plant sequences, easily recognizable because of their repeated nature. This showed that the insert (subsequently called VaTA) was located at JBVKUD‐2:25 046 610–25 047 986. VaTA was then used as a query to all WGS and SRA sequences from Ericaceae. In case of sufficient coverage, the hits extend into the flanking plant sequences. If these align with those flanking VaTA, the accession carries the same VaTA insert. Using this approach, we retrieved VaTA inserts and their flanking sequences from Vaccinium and the other genera. Ninety‐six Vaccinium species contain the 1215‐like gene, 51 showed enough coverage for VaTA border analysis, all contained VaTA. Outside the genus Vaccinium, VaTA and its flanking sequences were found in Gaultheria (G. prostrata, shallon, baccata, reticulata, cinerea, and angustifolia), Agapetes (A. malipoensis, forrestii, serpens, and polifolia), Polyclita (P. turbinata), Notopora (N. schomburgkii), and Psammisia (P. fissilis). Thus, the VaTA cT‐DNA was inserted in the common ancestor of these six genera. The root of the Vaccinieae tribe (with Vaccinium, Agapetes, Polyclita, Notopora, and Psammisia) has been estimated at 30 million years ago (Becker et al., 2024). As VaTA is also present in tribe Gaultherieae (Gaultheria) which split off earlier, the VaTA insertion must have occurred before. The conservation of the 1215‐like open reading frame noted before (Zhidkin et al., 2023) is all the more remarkable in view of the long history of the VaTA insert. In the next part, we provide data on new nGMOs used by humans.

New nGMOs among plants used by humans

Among nGMOs used for food and drinks, pomelo (Citrus maxima) carries ags and mas2′ (Matveeva & Otten, 2019). Here we assembled CiTA (124) from JBJGYR‐8 (C. maxima cv. ZPY). Interestingly, C. maxima assemblies JAQPSH, JAUJEG, JBKACN, JBKJAM, JBKJAN, JBLANY, and MKYQ lack CiTA, showing that CiTA is non‐fixed. Similarly, three C. sinensis cultivars (SRR11681090, SRR4089854, and SRR12300375) are nGMOs, but 22 others are not. Eighteen additional Citrus nGMO species and hybrids were found. Among these, citron (C. medica), mandarin (C. reticulata), kumquat (C. hindsii, C. japonica), bitter orange (C. aurantium), satsuma mandarin (C. unshiu), and grapefruit (Citrus × paradisi).

We also found cT‐DNA sequences in coffee (Coffea arabica, C. canephora, and C. liberica), avocado (Persea americana), cinnamon (Cinnamomum aromaticum), strawberry (Fragaria viridis and others), sea buckthorn (Hippophae rhamnoides), icecream bean (Inga edulis), cassava/manioc (Manihot esculenta), black pepper (Piper nigrum), apricot (Prunus armeniaca), raspberry (Rubus ideaeus), dewberry (Rubus caesius), and sorrel (Rumex acetosa).

cT‐DNAs in eudicot sister groups

Aside from eudicots, core angiosperms contain magnoliids, Chloranthales, Ceratophyllales (Figure 4, ‘others’, 10 290 species) and monocots. In the ‘others’ group, we found 41 (magnoliids), 3 (Chloranthales), and 0 (Ceratophyllales) nGMOs. Monocots contain 11 orders, 77 families, 2700 genera, and 74 273 species. Matveeva and Otten (2019) reported cT‐DNAs in Musa acuminata (banana) and greater yam (Dioscorea alata, WGS CZHE). We found no further Musa cT‐DNA sequences, but the Dioscorea results were confirmed and considerably extended. Whereas the CZHE contigs are small, JAFBII (D. alata, cv. TDa95/00328) has fully assembled sequences, attributed to each of its 20 chromosomes (Bredeson et al., 2022). DcTA (6315 nt, 141, located on chromosome 8) carries a single cus gene, interrupted by a 4920 nt plant sequence, PL1. Long reads from SRR13615784 showed that TDa95/00328 carries two different DcTA alleles, one with an intact cus gene (DcTA‐1, 1331 nt) and one with PL1 (DcTA‐2). Bioprojects PRJNA918625 and PRJNA666450 contain sequences from 134 D. alata cultivars. All carry DcTA and show five structural variants (140144) with different PL insertions (Table S4). DcTA‐2 occurs in 6/134 accessions and is therefore not representative for D. alata. Only seven accessions lack intact cus genes.

Figure 4.

Figure 4

Overview of plant groups with natural genetically modified organisms (nGMOs). Estimated species numbers for each group (Christenhusz & Byng, 2016) in black, nGMO numbers based on Sequence Read Archive (SRA) and whole genome sequences (WGS) data in red.

Group marked by ‘others’: magnoliids, Chloranthales, and Ceratophyllales. Photographs show representative nGMOs: Polytrichum commune (Bryophyta), Folioceros fuciformis (Anthocerotophyta), Solenostoma hyalinum (Marchantiophyta), Dryopteris affinis (Pteridophyta), Huperzia selago (Lycopodiophyta), Juniperus formosana (gymnosperms), Euryale ferox (basal angiosperms), Dioscorea alata (monocots), and Dianthus caryophyllus (eudicots).

cus sequences were also found in three additional Dioscorea species. The related species Trichopus zeylanicus has ags sequences. Finally, SRA data from 67 additional monocots (9 orders, 20 families, 52 genera) also contain cT‐DNA sequences (Table S2).

Diversity of cT‐DNA structures

The large number of cT‐DNAs recovered here allowed us to estimate the frequency of occurrence of the different cT‐DNA genes (Table S2). This showed that some (like susD, 6b, p5, ipt, c', d, and e) are rare, whereas others (like orf13, orf14, susL, and cus) are very frequent. The unusual genes are more typical for A. tumefaciens.

In order to study the different cT‐DNA structures in more detail, the WGS data were used to construct cT‐DNA maps. Using a standard format, we redrew relevant T‐DNA maps (117), earlier published cT‐DNA maps (1873), and new cT‐DNA maps (74166). Four groups (A–D, Table S5) can be distinguished.

Group A cT‐DNAs (like DcTA, 140) contain a single gene (‘mini cT‐DNAs’). Only one mini T‐DNA is known from agrobacteria (Otten, 2021): Agrobacterium 1641 has 1641‐T2 (vis), as well as 1641‐T1 (acs, susD, and 6b). 87/183 from the investigated WGS show mini cT‐DNAs, 75 of which carry ops genes. Thus, this type of T‐DNA is very frequent, contrary to what could be expected on the basis of Agrobacterium T‐DNA studies. Mini cT‐DNAs can be present in multiple copies, as in C. richardii (vis, 146), Urtica dioica (vis, 155, 156), Silene latifolia (cus, 157), and Paulownia fortunei (vis, 91). Twenty‐one WGS sequences from Caryophyllaceae show a mini cT‐DNA with a single cus gene. Vaccinium and related genera have a mini cT‐DNA with a single 1215‐like gene (see above).

Group B cT‐DNAs resemble the TL‐DNA from R. rhizogenes LMG152 (3), with acsorf2‐orf3n‐orf8‐rolA‐rolB‐rolC‐orf13‐orf13a‐orf14‐susL1susL2 . T‐DNA variant (5) has gene c between orf3n and orf8, and gene plast at the place of rolC. Others have susL1susL2 replaced by mas2mas1′ (11), susL (4, 6, 7, 12), cus (1), or mis (2, 5). Some cT‐DNAs from Group B have nos at the right border (58, 59, 65). rolC can be replaced by plast (56, 86, 97, 113, 116, 139), or accompanied by plast or d (77, 100, 104). rolB can be replaced by rolB TR (48, 53, 58, 92, 99, 121, 125). In T‐DNAs, acs can be replaced by orf358 (12), most likely an ops gene (Chen et al., 2023). In cT‐DNAs, acs (most cases) or orf358 (54, 74) can be replaced by vis (81, 82, 138), ocs (22, 63, 114), or orf358 ocs (62). Thus, a more general TL structure is ( ops )‐ opsorf2‐orf3n‐( c )‐orf8‐rolA‐ rolB‐(or rolB TR )‐rolC‐(d or plast)‐orf13‐orf13a‐orf14‐ops ‐( ops ). Gene c from DiTE ( ocs‐c‐orf8, 63) is inverted with respect to its orientation in Group B ( c‐orf8). DiTD (62) is a combination of two Group B T‐DNAs (T1 and T2).

Group C cT‐DNAs are similar to LMG152‐TR (8) with iaaH iaaM rolB TR mas2mas1′‐ ags . T‐DNA variants lack ags (9) or have iaaH iaaM vis (10). cT‐DNAs TE1 (24) and DiTN (72) have iaaH iaaM 6b vis . Group C sequences can be found linked to Group B sequences (B+C, Table S5).

Group D contains cT‐DNAs with orf511. orf511 has not yet been found in agrobacteria, but a non‐characterized bacterium from the Paracoccaceae potentially encodes an Orf511 homolog (hypothetical protein MCU0909928.1), its function is unknown. The orf511 gene has been identified on TD from N. tomentosiformis (20), CaTD (50), CaTF (52), CaTM (57), and CaTN (58) from Camellia, and DiTG (65) and DiTH (66) from Diospyros. It also occurs in cT‐DNAs 80, 86, 97, and 117, and is generally accompanied by orf14 and one or two susL genes: orf14orf511 susL ‐( susL ). This DNA may be related to (3) by insertion of orf511 in the orf14 susL susL fragment. orf511 is a common T‐DNA gene, as it appears in 188 species from 59 genera (Table 1; Table S2). Two fern species (Actinostachys digitata, Abrodictyum obscurum) contain orf511, orf14, and susL; these remain to be mapped. A few cT‐DNAs could not be classified (Table S5, not defined).

cT‐DNA repeats

Most cT‐DNA inserts are relatively simple inverted repeats. Direct repeats are rare; examples are JBJYGR‐8 (124), EU735069.2 (28), and JBHWAD‐60 (159). More complex repeats occur in BSXM‐40 (31, six repeats) or CANTUR‐331 (160) with four direct repeats (331‐A to 331‐D) in a total region of 600 kb. 331‐A to 331‐D appear to result from amplification of cT‐DNA inserts and surrounding plant sequences. CANTUR‐344 (161) and CANTUR‐1107 (162) each have one copy. The original cT‐DNA insert contains four fragments from the original T‐DNA. Altogether, CANTUR (Mallotus nudiflorus) contains 356 kb of cT‐DNA. CBDBZS (Camellia meiocarpa) also contains complex cT‐DNA repeats (150153). Silene latifolia JBLOPE‐1 has 11 cus copies spread out over 146 kb (157). These are not part of longer repeats but seem to be independent insertions. The origin of such multiple insertions clustered within a small chromosomal region is unknown. Urtica dioica CAXLOR‐222 (155) and CAXLOR‐56 (156) have multiple cT‐DNAs with a single vis gene within 250 and 225 kb, respectively; these are part of longer repeats. Overall, the cT‐DNA repeat structures appear highly variable.

DISCUSSION

A first systematic cT‐DNA search in public databases from eudicots and monocots (an estimated 284 300 species) found 23 nGMO species in 275 WGS sequences (Matveeva & Otten, 2019). By searching more recent WGS data, adding SRA data, and extending the search to all land plants, we obtained a list of 2614 nGMO species. Most of their cT‐DNAs remain to be completed. However, we believe that the present list can serve as a reliable starting point for further investigations. For example, in the genus Solanum, we detected 62 nGMO species. Solanum has about 1500–2000 species, most of which remain unsequenced. Phylogenetic studies in this genus (see, e.g., Messeder et al., 2024) may benefit from detailed analysis of the cT‐DNA distribution patterns, as each insertion marks the start of a new lineage, and thus provides an independent way to confirm phylogenetic patterns. The overall phylogenetic distribution of the presently known nGMOs is shown in Figure 4. It should be noted that many plant species remain unsequenced. Our estimation of 40 700 nGMO species in land plants is probably underestimated, as many cT‐DNAs are non‐fixed, and noncoding sequences were excluded in our study. The presence of cT‐DNAs in ferns and mosses shows that these spore‐bearing land plants can be transformed by Agrobacterium, and regenerate into nGMOs, suggesting that Agrobacterium transformation could date back to early stages in land plant evolution. This is also indicated by the large sequence divergence of the plast‐like T‐DNA genes, which suggests an early origin of natural plant transformation (Otten, 2018).

Although many sequences recovered here are incomplete, these could already be used to extend earlier data on nGMO genera like Vaccinium, Ipomoea, Arachis, Dioscorea, and Citrus. The presence of cT‐DNA in various plants used for popular food and drinks, like cassava, apricots, avocado, strawberry, and coffee (to name a few) could lead to further studies on the possible contribution of cT‐DNA genes to the properties of these economically important plants. WGS data were used to construct a large collection of cT‐DNA maps, which defined four structural groups. Group A contains mini cT‐DNAs, mostly with a single ops gene. They are found in mosses, ferns, Dioscorea, and eudicots. Most likely, Agrobacterium strains with mini T‐DNAs are unable to induce tumors or HRs. Thus, they would most likely escape the attention of plant pathologists, which could explain their absence in bacterial collections and the lack of their sequences in DNA databases.

In view of the broad occurrence of cus genes in land plants, it seems possible, at first sight, that they are of plant origin, and were at some stage transferred to Agrobacterium. However, we believe this is unlikely. Recently, 138 whole genomes (accession codes: JBNGxx) were obtained from mosses and liverworts (Dong et al., 2024). Only 10 (Figure 2), belonging to different families, contain the cus gene. Also, the Cus protein from Porella (Marchantiophyta) clusters with Cus from Polytrichum (Bryophyta), in spite of the fact that these two species are quite distant (Figure 2). This patchy pattern is not what one would expect if the cus gene would be of plant origin, but is more likely to result from transformation of a limited number of species by related Agrobacterium strains. Most importantly, an ‘empty insertion site’ (a site without cT‐DNA insert) in H. elatum aligns with the sequences surrounding the HyTA insert in H. flavolimbatum, making it likely that this empty site represents the original, non‐modified site, and that H. flavolimbatum acquired HyTA by Agrobacterium‐mediated transformation.

A search of the NCBI clusteredNR protein collection, using various Cus sequences as queries, detected seven Cus sequences in fungi. Such a low number makes it unlikely that these are of fungal origin, but (as in the case of plants) argues in favor of Agrobacterium‐mediated transformation.

The mini cT‐DNAs from Group A may be explained in three ways. They could result from infection by strains with mini T‐DNAs, by partial transfer of longer T‐DNAs, or from deletions of longer cT‐DNAs after integration. Possibility two and three are expected to leave some traces of neighboring genes, as the deletion process is expected to be random. Because such traces have not been observed so far, we favor the first hypothesis. Of course, final proof will require the discovery of Agrobacterium strains with mini‐T‐DNAs.

T‐DNAs from groups B and C could have been derived from mini T‐DNAs by incorporating HR‐inducing genes. Group D T‐DNAs (carrying orf511, orf14, susL) are unknown in agrobacteria but common in nGMOs. orf14 plays a minor role in HR initiation and growth (Aoki & Syōno, 1999; Otten, 2018) but can induce significant dwarfing (Favero et al., 2022). Like for orf511, its mode of action remains unknown, but the large number of orf14 and orf511 genes already identified in nGMOs might be exploited for further investigations. Possibly, strains with group D T‐DNAs are also non‐symptomatic.

The distribution of Group A–D structures among land plants suggests their possible evolutionary origin. Mosses carry mini cT‐DNAs with a single cus gene; these could be the earliest cT‐DNAs. Ferns show mini T‐DNAs with cus and other ops genes, indicating these appeared later. About half of the eudicots have complex cT‐DNAs; the others carry mini cT‐DNAs. Examples of eudicots with cus mini cT‐DNAs are found in genera Cerastium (19 accessions), Dianthus (17), Gypsophila (9), Moehringia (8), Schiedea (17), Silene (85), Stellaria (19), and Stylosanthes (10). Genera with other mini cT‐DNAs are: Cuscuta (mis, 14), JuglansPlatycaryaPterocarya‐Cyclocarya (susL, 17), and Urtica (vis, 8). We predict that species from these 14 genera are able to regenerate nGMOs from single transformed cells, without a need for an intermediate HR stage. Vaccinioideae with their 1215‐like genes may also be included in this group, but its effect on transformed cells is unknown. The widespread occurrence of mini cT‐DNAs with ops genes strongly supports the notion that opines are the ‘raison d'être’ of the Agrobacterium–plant interaction (Petit & Tempé, 1985). However, it remains to be demonstrated whether opine synthase genes from nGMOs still play a role in Agrobacterium biology, or have been recruited for other functions instead.

In general, the conservation of cT‐DNA open reading frames in present‐day nGMO species suggests that they play some role in nGMOs. The VaTA 1215‐like gene from Vaccinium and other, related genera is a particularly striking case, as most accessions have an intact open reading frame (Zhidkin et al., 2023), in spite of the fact that the insert is over 30 Mio years old. Overexpression studies of T‐DNA and cT‐DNA genes in heterologous systems have shown various phenotypic effects, suggesting similar effects in the respective nGMOs (Chen & Otten, 2017). However, defining a biological role will require comparative studies of wild‐type nGMOs and isogenic mutants. A recent paper describes the knock‐out of two IbT‐DNA2 genes in sweet potato: rolB/C (a plast gene) and rolD (an opine synthase gene). The authors showed a reduction of chlorogenic acid and polyphenol content in the rolB/C mutant and a reduction in biomass and downregulation of cell cycle‐related genes in the rolD mutant (Shkryl et al., 2025). Further studies will be required to determine the molecular mode of action of these genes. Among the cT‐DNA genes, the opine synthase genes will be the easiest ones to study, as their general function is well known. Deoxyfructosylglutamine (santhopine) has been found in N. tabacum (Chen et al., 2016) and mikimopine in Cuscuta (Zhang et al., 2020), but no mutants have been studied.

cT‐DNAs mostly consist of simple, partial inverted repeats. However, direct repeats and complex repeats also occur. Complex repeats can result from ligation of T‐DNA fragments before insertion, multiple insertions during transformation, or insertion followed by amplification. T‐DNA inserts from laboratory experiments show similar variation (De Buck et al., 2009; Gelvin, 2003; Jupe et al., 2019; Kleinboelting et al., 2015; Kralemann et al., 2022; Pucker et al., 2021). A special feature of nGMOs is the presence of multiple inserts in a large chromosomal region, as observed for the mini cT‐DNAs from C. richardii, U. dioica, and S. latifolia. Possibly, strains with mini T‐DNA have different T‐DNA transfer properties, leading to such unusual insertion patterns.

This study, and other investigations, show that cT‐DNAs and their corresponding T‐DNAs could become rich sources of genes with economically useful effects on plant growth and metabolism. Wild‐type R. rhizogenes strains have been used to improve ornamental plants (Favero et al., 2022). This approach may be extended to vegetable crops. Indeed, nGMOs show that both Agrobacterium‐mediated T‐DNA transfer and the subsequent generation of transformants are natural phenomena. These processes have occurred for millions of years in all land plants and involve large numbers of cultivated species.

MATERIALS AND METHODS

Whole genome sequencing data analysis

In order to identify and extract cT‐DNA sequences from WGS data, we first collected taxonomic IDs (taxid) from the NCBI Taxonomy Browser for Spermatophyta (txid 58024), Acrogymnospermae (txid 1437180), Polypodiophyta (or Polypodiopsida, txid 241806), Lycopodiophyta (or Lycopodiopsida, txid 1521260), Bryophyta (txid 3208), Marchantiophyta (txid 3195), and Anthocerotophyta (txid 13809) (Schoch et al., 2020).

Using taxid2wgs.pl (https://ftp.ncbi.nlm.nih.gov/blast/WGS_TOOLS/taxid2wgs.pl), we generated alias files for each group and downloaded the WGS datasets. These were searched with blastn_vdb (Altschul et al., 1990) to find homologous sequences. For this, we used the Ntfuse6 DNA query, a concatenated sequence composed of selected T‐DNA and cT‐DNA sequences (Otten et al., 2025). In order to avoid false positives, the DNA sequences were checked by BLASTX analysis for bona fide T‐DNA proteins using a set of 380 T‐DNA and cT‐DNA proteins (Table S1). Long WGS numbers were abbreviated. For example, JAIKUY010002022.1 was converted to JAIKUY‐2022.

Retrieval and analysis of SRA datasets

To retrieve the full set of SRA data from all land plant groups (except monocots and three large eudicot families, see below), we used the Entrez Direct tool to obtain the accession numbers of the target group, followed by the SRA‐toolkit to download the raw sequencing data from the SRA database at the NCBI (https://www.ncbi.nlm.nih.gov/sra/). As the data generated from different experiments and projects may vary significantly, and the default maximum download file size for the ‘prefetch’ command is 20 GB, the ‘‐‐max‐size’ parameter was set to its maximum value. We downloaded all available SRA data for land plant groups, with the exception of monocots, Fabaceae, Brassicaceae, and Solanaceae, where the number of SRR accessions was extremely large. For these, we used a simplified approach by randomly selecting up to 20 SRR accessions per species. The SRA data of the Dioscoreae family were searched completely.

For searching T‐DNA gene homologs in all sequencing data of land plant groups, we used the ‘blastn_vdb’ command from BLAST+ 2.15.0, with the Ntfuse6 sequence as the query and the following parameters: ‘‐gapopen 5 ‐gapextend 2 ‐word_size 11 ‐evalue 0.05 ‐penalty ‐3 ‐reward 2’.

Recovery of cT‐DNA sequences and assembly

The data archived in the Sequence Read Archive (SRA) can be either paired‐end sequences or single‐end sequences, depending on the experimental design and requirements. Paired‐end reads can enhance the credibility and accuracy of the sequences. However, in some cases, the original ‘blastn_vdb’ output may only contain single‐end reads. To address the issue of such reads, caution is needed in extracting sequences. Specifically, the blast output distinguishes between paired‐end sequences with the suffixes ‘.1’ or ‘.2’. After removing the sequence name suffixes and performing the ‘uniq’ operation, sequences are then extracted separately from the paired‐end data. After recovering cT‐DNA homologous sequences from SRA data, CodonCode Aligner (CCA, version 11.0.2, CodonCode Corporation) was used for sequence assembly. Variations in experimental design and sequencing protocols can significantly affect data quality, resulting in differences in the recovery of homologous sequences. Because CCA assembly uses the De Bruijn graph‐based method, datasets with low read counts often yield poor‐quality assemblies. Therefore, we opted not to assemble SRA data containing fewer than 100 reads; instead, these reads were analyzed directly with BLASTX. For datasets with more than 100 reads, we first generated contigs by CCA assembly and subsequently subjected these to BLASTX analysis. Unassembled reads were analyzed in the same way. Contigs obtained from CodonCode Aligner were utilized as queries in BLASTX to identify the closest protein homologs. For contigs less than 1 kb, the protein with the lowest E‐value was selected; for longer contigs, the Graphic Summary view in online BLASTX was used to examine the closest proteins. It is important to note that in many cases, no full cT‐DNAs could be assembled. The combined WGS and SRA data approach (Figure 1) allowed us to identify and catalog the presence of Agrobacterium/Rhizobium‐derived T‐DNA sequences in various plant genomes.

Exclusion of false positives

In the SRA, some datasets are from species transformed by R. rhizogenes, such HR lines should be excluded. So far, only four A. rhizogenes strains have been used, LMG152 (15 834, A4 or HR‐I), LMG63 (8196), 1724, and NCPPB2659 (K599) (Ying et al., 2023). To detect HR lines, all recovered sequences were compared to the T‐DNAs of these four A. rhizogenes strains by BLASTN, and excluded if they showed more than 99% identity. The second group of contaminants was due to weakly homologous sequences from plant genomes. Indeed, T‐DNA proteins Orf8, SusL, Mas1′, Ags, Ipt, IaaH, and IaaM (Table S1) may show weak similarity to embryophyte proteins (less than 40% identity). Only a 45 amino acid IaaM fragment (MCX‐IaaM coordinates 252–297) may have 65% identity. As nucleotide sequences are less conserved than protein sequences, the chance of finding DNA sequences from these plant homologs in our initial search with DNA query Ntfuse6 is low. Nevertheless, to avoid false positives, we checked all proteins with identity levels below 50% by BLASTP against the NCBI ClusteredNR database. In case the best hits were from plants without cT‐DNAs, the accession was discarded. Bacteria other than agrobacteria may also yield false positives with weak DNA and protein homology. For these SRRs, we compared the recovered DNA sequences to the NCBI core nucleotide database with BLASTN, and removed accessions whose top hits matched sequences from bacteria outside the agrobacteria group. Finally, to exclude Agrobacterium contamination of sequenced plant material, we tested the positive SRR and WGS accessions with a DNA query based on Agrobacterium virulence genes, as indicated in Matveeva and Otten (2019).

Phylogenetic tree construction

Cus protein sequences were aligned using MUSCLE (Edgar, 2004). Maximum likelihood phylogenetic trees were generated using IQ‐TREE2 with 1000 bootstrap replicates (Minh et al., 2020). Trees were visualized and annotated in iTOL (Letunic & Bork, 2021).

Matrix analysis

Complex repeats were analyzed by matrix analysis on the National Library of Medicine (NLM) blast site, using BLASTN with the ‘highly similar sequences’ setting; the matrix figure was recovered and analyzed for repeats.

AUTHOR CONTRIBUTIONS

KC and LO conceived and designed research. HL, YH, SC, ZH, YL, JH, and YY provided bioinformatics analysis. The first draft of the manuscript was written by LO and HL. The manuscript was revised by LO, KC, WW, and JL. All authors read and approved the final document.

CONFLICT OF INTEREST

None of the authors have a conflict of interest to disclose.

Supporting information

Figure S1. nGMO distribution on the evolutionary tree of the 425 spermatophyte families.

Figure S2. Maps of T‐DNAs in Agrobacterium/Rhizobium and of cT‐DNAs in nGMOs.

TPJ-127-0-s001.rar (2.6MB, rar)

Table S1. Protein query sequences for searching cT‐DNA‐located coding sequences by the BLASTX approach.

Table S2. cT‐DNA genes in spermatophytes (gymnosperms and angiosperms), based on WGS and SRA data.

Table S3. List of cT‐DNAs ArTA to ArTD in WGS sequences of Arachis.

Table S4. DcTA types from 134 different Dioscorea alata accessions.

Table S5. Groups of cT‐DNA structures based on WGS data.

TPJ-127-0-s002.rar (507.2KB, rar)

ACKNOWLEDGMENTS

This work was supported by the National Natural Science Foundation of China (32370382 to KC). It was also supported by a special fund for scientific research of Shanghai landscaping and city appearance administrative bureau (G262408 and G242406 to KC) and a special fund for scientific research of national botanical gardens to benefit sustainable development (2026 to HL). The authors have no competing interests to declare that are relevant to the content of this article. The authors declare that no human and/or animal material, data, or cell lines were used in this study. We thank Todd Blevins for discussion and comments.

Contributor Information

Léon Otten, Email: otten@unistra.fr.

Ke Chen, Email: kchen@cemps.ac.cn.

DATA AVAILABILITY STATEMENT

The datasets generated and/or analyzed during the current study are accessible via the following GitHub repository: https://github.com/HLiuprojects/nGMO‐cT‐DNA‐sequences.

REFERENCES

  1. Altschul, S.F. , Gish, W. , Miller, W. , Myers, E.W. & Lipman, D.J. (1990) Basic local alignment search tool. Journal of Molecular Biology, 215, 403–410. [DOI] [PubMed] [Google Scholar]
  2. Aoki, S. & Syōno, K. (1999) Function of Ngrol genes in the evolution of Nicotiana glauca: conservation of the function of NgORF13 and NgORF14 after ancient infection by an Agrobacterium rhizogenes‐like ancestor. Plant and Cell Physiology, 40, 222–230. [Google Scholar]
  3. Becker, A.L. , Crowl, A.A. , Luteyn, J.L. , Chanderbali, A.S. , Judd, W.S. , Manos, P.S. et al. (2024) A global blueberry phylogeny: evolution, diversification, and biogeography of Vaccinieae (Ericaceae). Molecular Phylogenetics and Evolution, 201, 108202. [DOI] [PubMed] [Google Scholar]
  4. Bogomaz, O.D. , Bemova, V.D. , Mirgorodskii, N.A. & Matveeva, T.V. (2024) Evolutionary fate of the opine synthesis genes in the Arachis L. genomes. Biology‐Basel, 13, 601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Bredeson, J.V. , Lyons, J.B. , Oniyinde, I.O. , Okereke, N.R. , Kolade, O. , Nnabue, I. et al. (2022) Chromosome evolution and the genetic basis of agronomically important traits in greater yam. Nature Communications, 13, 2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Chen, K. & Otten, L. (2017) Natural Agrobacterium transformants: recent results and some theoretical considerations. Frontiers in Plant Science, 8, e1600. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Chen, K. , Dorlhac de Borne, F. , Julio, E. , Obszynski, J. , Pale, P. & Otten, L. (2016) Root‐specific expression of opine genes and opine accumulation in some cultivars of the naturally occurring GMO Nicotiana tabacum . Plant Journal, 87, 258–269. [DOI] [PubMed] [Google Scholar]
  8. Chen, K. , Liu, H. , Blevins, T. , Hao, J. & Otten, L. (2023) Extensive natural Agrobacterium‐induced transformation in the genus Camellia . Planta, 258, 81. [DOI] [PubMed] [Google Scholar]
  9. Chen, K. , Zhurbenko, P. , Danilov, L. , Matveeva, T. & Otten, L. (2022) Conservation of an Agrobacterium cT‐DNA insert in Camellia section Thea reveals the ancient origin of tea plants from a genetically modified ancestor. Frontiers in Plant Science, 13, 997762. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Christenhusz, M.J.M. & Byng, J.W. (2016) The number of known plants species in the world and its annual increase. Phytotaxa, 261, 201–217. [Google Scholar]
  11. De Buck, S. , Podevin, N. , Nolf, J. , Jacobs, A. & Depicker, A. (2009) The T‐DNA integration pattern in Arabidopsis transformants is highly determined by the transformed target cell. Plant Journal, 60, 134–145. [DOI] [PubMed] [Google Scholar]
  12. Dong, S. , Wang, S. , Li, L. , Yu, J. , Zhang, Y. , Xue, J. et al. (2024) Bryophytes hold a larger gene family space than vascular plants. Nature Genetics, 57, 2562–2569. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Edgar, R.C. (2004) MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32, 1792–1797. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Favero, B.T. , Tan, Y. , Chen, X. , Muller, R. & Lutken, H. (2022) Kalanchoe blossfeldiana naturally transformed with Rhizobium rhizogenes exhibits superior root phenotype. Plant Science, 321, 111323. [DOI] [PubMed] [Google Scholar]
  15. Furner, I.J. , Huffman, G.A. , Amasino, R.M. , Garfinkel, D.J. , Gordon, M.P. & Nester, E.W. (1986) An Agrobacterium transformation in the evolution of the genus Nicotiana . Nature, 319, 422–427. [Google Scholar]
  16. Gelvin, S.B. (2003) Agrobacterium‐mediated plant transformation: the biology behind the “gene‐jockeying” tool. Microbiology and Molecular Biology Reviews, 67, 16–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Gelvin, S.B. (2017) Integration of Agrobacterium T‐DNA into the plant genome. Annual Review of Genetics, 51, 195–217. [DOI] [PubMed] [Google Scholar]
  18. Group, T.A.P. , Chase, M.W. , Christenhusz, M.J.M. , Fay, M.F. , Byng, J.W. , Judd, W.S. et al. (2016) An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG IV. Botanical Journal of the Linnean Society, 181, 1–20. [Google Scholar]
  19. Hooykaas, P.J.J. (2023) The Ti plasmid, driver of Agrobacterium pathogenesis. Phytopathology, 113, 594–604. [DOI] [PubMed] [Google Scholar]
  20. Jupe, F. , Rivkin, A.C. , Michael, T.P. , Zander, M. , Motley, S.T. , Sandoval, J.P. et al. (2019) The complex architecture and epigenomic impact of plant T‐DNA insertions. PLoS Genetics, 15, e1007819. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Kleinboelting, N. , Huep, G. , Appelhagen, I. , Viehoever, P. , Li, Y. & Weisshaar, B. (2015) The structural features of thousands of T‐DNA insertion sites are consistent with a double‐strand break repair‐based insertion mechanism. Molecular Plant, 8, 1651–1664. [DOI] [PubMed] [Google Scholar]
  22. Kralemann, L.E.M. , de Pater, S. , Shen, H. , Kloet, S.L. , van Schendel, R. , Hooykaas, P.J.J. et al. (2022) Distinct mechanisms for genomic attachment of the 5′ and 3′ ends of Agrobacterium T‐DNA in plants. Nature Plants, 8, 526–534. [DOI] [PubMed] [Google Scholar]
  23. Kyndt, T. , Quispe, D. , Zhai, H. , Jarret, R. , Ghislain, M. , Liu, Q. et al. (2015) The genome of cultivated sweet potato contains Agrobacterium T‐DNAs with expressed genes: an example of a naturally transgenic food crop. Proceedings of the National Academy of Sciences of the United States of America, 112, 5844–5849. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Letunic, I. & Bork, P. (2021) Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Research, 49, W293–W296. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Matveeva, T.V. & Otten, L. (2019) Widespread occurrence of natural genetic transformation of plants by Agrobacterium . Plant Molecular Biology, 101, 415–437. [DOI] [PubMed] [Google Scholar]
  26. Matveeva, T.V. , Bogomaz, D.I. , Pavlova, O.A. , Nester, E.W. & Lutova, L.A. (2012) Horizontal gene transfer from genus Agrobacterium to the plant Linaria in nature. Molecular Plant‐Microbe Interactions, 25, 1542–1551. [DOI] [PubMed] [Google Scholar]
  27. Messeder, J.V.S. , Carlo, T.A. , Zhang, G. , Tovar, J.D. , Arana, C. , Huang, J. et al. (2024) A highly resolved nuclear phylogeny uncovers strong phylogenetic conservation and correlated evolution of fruit color and size in Solanum L. New Phytologist, 243, 765–780. [DOI] [PubMed] [Google Scholar]
  28. Minh, B.Q. , Schmidt, H.A. , Chernomor, O. , Schrempf, D. , Woodhams, M.D. , von Haeseler, A. et al. (2020) IQ‐TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution, 37, 1530–1534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Nester, E.W. (2014) Agrobacterium: nature's genetic engineer. Frontiers of Plant Science, 5, 730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Otten, L. (2018) The Agrobacterium phenotypic plasticity (Plast) genes. Current Topics in Microbiology and Immunology, 418, 375–419. [DOI] [PubMed] [Google Scholar]
  31. Otten, L. (2021) T‐DNA regions from 350 Agrobacterium genomes: maps and phylogeny. Plant Molecular Biology, 106, 239–258. [DOI] [PubMed] [Google Scholar]
  32. Otten, L. , Liu, H. , Meeprom, N. , Linan, A. , Puglisi, C. & Chen, K. (2025) Accumulation of numerous cellular T‐DNA sequences in the genus Diospyros by multiple rounds of natural transformation. Plant Journal, 122, e70202. [DOI] [PubMed] [Google Scholar]
  33. Petit, A. & Tempé, J. (1985) The function of T‐DNA in nature. In: Molecular form and function of the plant genome. New York: Plenum Press, pp. 625–636. [Google Scholar]
  34. Pucker, B. , Kleinbolting, N. & Weisshaar, B. (2021) Large‐scale genomic rearrangements in selected Arabidopsis thaliana T‐DNA lines are caused by T‐DNA insertion mutagenesis. BMC Genomics, 22, 599. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Quispe‐Huamanquispe, D.G. , Gheysen, G. , Yang, J. , Jarret, R. , Rossel, G. & Kreuze, J.F. (2019) The horizontal gene transfer of Agrobacterium T‐DNAs into the series Batatas (genus Ipomoea) genome is not confined to hexaploid sweetpotato. Science Reports, 9, 12584. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Schoch, C.L. , Ciufo, S. , Domrachev, M. , Hotton, C.L. , Kannan, S. , Khovanskaya, R. et al. (2020) NCBI taxonomy: a comprehensive update on curation, resources and tools. Database: The Journal of Biological Databases and Curation, 2020, baaa062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Shkryl, Y. , Yaroshenko, Y. , Grigorchuk, V. , Bulgakov, V. & Yugay, Y. (2025) Functional analysis of naturally integrated rol genes in sweet potato via CRISPR/Cas9 genome editing. Plants, 14, 3708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Weisberg, A.J. , Davis, E.W. , Tabima, J. , Belcher, M.S. , Miller, M. , Kuo, C.H. et al. (2020) Unexpected conservation and global transmission of agrobacterial virulence plasmids. Science, 368, eaba5256. [DOI] [PubMed] [Google Scholar]
  39. White, F.F. , Garfinkel, D.J. , Huffman, G.A. , Gordon, M.P. & Nester, E.W. (1983) Sequences homologous to Agrobacterium rhizogenes T‐DNA in the genomes of uninfected plants. Nature, 301, 348–350. [Google Scholar]
  40. Yan, M. , Li, M. , Wang, Y. , Wang, X. , Moeinzadeh, M.H. , Quispe‐Huamanquispe, D.G. et al. (2024) Haplotype‐based phylogenetic analysis and population genomics uncover the origin and domestication of sweetpotato. Molecular Plant, 17, 277–296. [DOI] [PubMed] [Google Scholar]
  41. Ying, W. , Wen, G. , Xu, W. , Liu, H. , Ding, W. , Zheng, L. et al. (2023) Agrobacterium rhizogenes: paving the road to research and breeding for woody plants. Frontiers in Plant Science, 14, 1196561. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Zhang, Y. , Wang, D. , Wang, Y. , Dong, H. , Yuan, Y. , Yang, W. et al. (2020) Parasitic plant dodder (Cuscuta spp.): a new natural Agrobacterium‐to‐plant horizontal gene transfer species. Science China. Life Sciences, 63, 312–316. [DOI] [PubMed] [Google Scholar]
  43. Zhidkin, R. , Zhurbenko, P. , Bogomaz, O. , Gorodilova, E. , Katsapov, I. , Antropov, D. et al. (2023) Biodiversity of rolB/C‐like natural transgene in the genus Vaccinium L. and its application for phylogenetic studies. International Journal of Molecular Sciences, 24, 6923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Zhu, J. , Oger, P.M. , Schrammeijer, B. , Hooykaas, P.J. , Farrand, S.K. & Winans, S.C. (2000) The bases of crown gall tumorigenesis. Journal of Bacteriology, 182, 3885–3895. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1. nGMO distribution on the evolutionary tree of the 425 spermatophyte families.

Figure S2. Maps of T‐DNAs in Agrobacterium/Rhizobium and of cT‐DNAs in nGMOs.

TPJ-127-0-s001.rar (2.6MB, rar)

Table S1. Protein query sequences for searching cT‐DNA‐located coding sequences by the BLASTX approach.

Table S2. cT‐DNA genes in spermatophytes (gymnosperms and angiosperms), based on WGS and SRA data.

Table S3. List of cT‐DNAs ArTA to ArTD in WGS sequences of Arachis.

Table S4. DcTA types from 134 different Dioscorea alata accessions.

Table S5. Groups of cT‐DNA structures based on WGS data.

TPJ-127-0-s002.rar (507.2KB, rar)

Data Availability Statement

The datasets generated and/or analyzed during the current study are accessible via the following GitHub repository: https://github.com/HLiuprojects/nGMO‐cT‐DNA‐sequences.


Articles from The Plant Journal are provided here courtesy of Wiley

RESOURCES