Skip to main content
Physiology and Molecular Biology of Plants logoLink to Physiology and Molecular Biology of Plants
. 2020 Jan 1;26(3):409–418. doi: 10.1007/s12298-019-00736-7

Chloroplast genome of an extremely endangered conifer Thuja sutchuenensis Franch.: gene organization, comparative and phylogenetic analysis

Tao Yu 1, Bing-Hong Huang 2, Yuyang Zhang 1, Pei-Chun Liao 2,, Jun-Qing Li 1,
PMCID: PMC7078402  PMID: 32205919

Abstract

Thuja sutchuenensis is a critically endangered tertiary relict species of Cupressaceae from southwestern China. We sequenced the complete chloroplast (cp) genome of T. sutchuenensis, showing the genome content of 129,776 bp, 118 unique genes including 82 unique protein-coding genes, 32 tRNA genes, and 4 rRNA genes. The genome structures, gene order, and GC content are similar to other typical gymnosperm cp genomes. Thirty-eight simple sequence repeats were identified in the T. sutchuenensis cp genome. We also found an apparent inversion between trnT and psbK between genera Thuja and Thujopsis. In addition, positive selection signals were detected in seven genes with high Ka/Ks ratios. The reconstructed phylogeny based on locally collinear blocks of cp genomes among 21 gymnosperms species is similar to previous inferences. We also inferred a Late-Miocene divergence between T. sutchuenensis and T. standishii, according to the dating of ~ 11.05 Mya by cp genomes. These results will be helpful for future studies of Cupressaceae phylogeny as well as studies in population genetics, systematics, and cp genetic engineering.

Electronic supplementary material

The online version of this article (10.1007/s12298-019-00736-7) contains supplementary material, which is available to authorized users.

Keywords: Thuja sutchuenensis, Chloroplast genome, Sequence divergence, Non-synonymous substitution, Phylogenomics

Introduction

The chloroplast (cp) is a semiautonomous organelle in plants, encoding several key proteins involved in photosynthesis and interactions between plants and the surrounding environment (Saski et al. 2007; Daniell et al. 2016). With the rapid development of high-throughput sequencing technologies, an increasing number of sequenced cp genomes can easily be acquired from online databases such as the National Center for Biotechnology Information (NCBI, http://www.ncbi.nlm.nih.gov/genomes/) (Daniell et al. 2016). The size of the cp genome in almost all vascular plant species ranges between 120 and 160 kb in length, and contains about 130 genes (Chumley et al. 2006). In most gymnosperm plants, the cp genome is inherited from the patrilineal lineage, and exhibits very little or no recombination (Yi et al. 2013). Due to the relatively simple features of the cp genome, the cp sequences are commonly used as DNA barcodes for genetic identification, for systematic plant studies, and studies of plant biodiversity, biogeography, adaptation, etc. (Wambugu et al. 2015; Brozynska et al. 2016).

Thuja sutchuenensis is a tertiary relict species of Cupressaceae (Tang et al. 2015). It is found in Sichuan and Chongqing Provinces in China, with an altitudinal range between 800 and 2100 m (mostly between 1000 and 1500 m) (Qiaoping et al. 2015). This species was firstly discovered by a French botanist, Paul Guillaume Farges, who collected specimens from 1892 to 1900 and described it in 1899. It was not seen again for almost 100 years. T. sutchuenensis was then classified as extinct in the wild (EW) until it was rediscovered in northeastern Chongqing in 1999 (Qiaoping et al. 2015). It was reassessed as critically endangered (CR) in 2003 (Qiaoping et al. 2015; Gao et al. 2018). Southwest China is rich in vegetation, but after decades of excessive deforestation and habitat destruction, the original forest of T. sutchuenensis was fragmented, and only a small area remains. Protection of this critically endangered plant requires adequate basic research.

The study of conservation genetics and the reconstruction of evolutionary history depend on suitable molecular markers. However, molecular research on T. sutchuenensis is scarce. Previous population genetics studies based on inter simple sequence repeats (ISSR) can roughly describe genetic diversity but cannot comprehensively illustrate evolutionary history due to the limitation of dominant markers and small numbers of loci used (Liu et al. 2013), particularly considering the large size of the T. sutchuenensis genome [1C = 12.10 pg (Zonneveld 2012)]. The large genome size makes the development of ample nuclear markers difficult. Instead, it is easier to obtain homologous genes from small-sized cp genomes, and they are easier to sequence and provide comparable genetic information. In this study, we used an Illumina Miseq Platform to assemble the cp genomes of T. sutchuenensis in order to (1) deepen the understanding of the cp genome structure of T. sutchuenensis and (2) to develop effective molecular markers for T. sutchuenensis that can be applied to conservation genetics and evolutionary inference. We also reconstructed a phylogenomic tree with other published cp genomes of Cupressaceae species in order to acquire more robust phylogenetic inference in the subfamily Cupressideae.

Materials and methods

Plant materials and DNA sequencing

Young leaves from T. sutchuenensis trees in the Chinese Academy of Forestry were collected and dried immediately with silica gel for preservation. Whole genomic DNA was extracted using the plant genome DNA extraction kit (TIANGEN, Beijing, China) based on the manufacturer’s protocol. The average insert sizes were approximately 350 bp. Paired-end Library preparations were constructed using the Illumina Miseq platform according to the Illumina standard method at Megagenomics Company (Beijing, China). First, DNA sequence fragmentation by ultrasound, then terminal repair phosphorylation: DNA polymerase is used to repair fragmented DNA into dsDNA at the flat terminal, and T4 poly nucleotide kinase phosphorylates the 5 ‘terminal, and add A prominent A tail to the 3’ end of dsDNA. Connect sequencing joints, and appropriate fragments with relatively concentrated sizes were screened according to the requirements of subsequent analysis. Finally, enrich the target fragment and obtain the library for sequencing (Meyer and Kircher 2010). We generated 9.5 Gb of total data with a 150-bp average read length.

Genome assembly, annotation

High-quality data were filtered from raw sequence data using the NGSQC (Dai et al. 2010) with default settings. Clean reads were assembled using MITObim v1.8 (Hahn et al. 2013) and NOVOplasty (Dierckxsens et al. 2016) and using the cp genome of T. standishii as the reference (Qu et al. 2017). Samtools (Li et al. 2009) was used to recheck the sequences based on the degree of coverage. Gaps between the plastomic contigs were filled up using Sanger sequencing. The specific primers used for PCR are shown in Supplementary Table S1. We annotated protein-coding genes using Cpgavas (Yong and Zheng 2012) and checked gene boundaries by comparing T. standishii cp genomes using the BLASTN software in NCBI (http://www.ncbi.nlm.nih.gov). The error-corrected SQN-file of T. sutchuenensis cp genome sequence was submitted to GenBank (accession number: MH784400). The circular gene map of the cp genome (Fig. 1) was drawn by the Organella Genome DRAW software (OGDRAW v1.2.) (Lohse et al. 2007).

Fig. 1.

Fig. 1

Map of the chloroplast genome of Thuja sutchuenensis. Genes inside and outside of the circle are transcribed in the clockwise and counterclockwise directions, respectively. Genes belonging to different functional groups are color-coded. The dashed area in the inner circle indicates the GC content of the chloroplast genomes (colour figure online)

Identifying cp SSRs

Simple sequence repeats (SSRs) in T. sutchuenensis and T. standishii cp genomes were detected by MISA Perl script (http://pgrc.ipk-gatersleben.de/misa/). The parameter of minimum repeat units was set at 10 for mono-, 6 for di-, and 5 for tri- to hexanucleotides.

Genome-wide homologous comparison and divergence of coding gene sequences

Whole cp genomes of T. sutchuenensis, T. standishii, and Thujopsis dolabrata were aligned in MAUVE (Darling et al. 2004) under default settings to test rearrangement events. Coding genes in each cp genome were determined and aligned by MAFFT (Katoh et al. 2005). To identify positive selection of T. sutchuenensis and T. standishii cp genome, synonymous (Ks) and non-synonymous substitution rate (Ka) of each coding gene was calculated in DnaSP 5.0 software (Librado and Rozas 2009).

Phylogenetic analysis and divergence time estimation

Twenty-one cp genomes of Cupressideae were used to reconstruct the phylogenomic tree. Species of subfamilies Taiwanioideae (Taiwania crypyomerioides and T. flousiana), Cunninghamhioideae (Cunninghamia lanceolata), Sequoioideae (Metasequoia glyptostroboides), and Taxodioideae (Cryptomeria japonica) were selected as outgroups. Accession numbers of all used cp genome sequences are listed in Table S2. We used HomBlocks (Bi et al. 2017) to determine locally collinear blocks (LCBs) among cp genomes and used Circoletto (http://tools.bat.infspire.org/circoletto/) (Darzentas 2010) to correspond analysis between locally collinear block sequences and T. sutchuenensis genes sequence. PhyML (Guindon et al. 2009) and MrBayes (Huelsenbeck and Ronquist 2001) were used to reconstruct the phylogenetic tree based on ML and BI, respectively. The best substitution model was determined according to the Akaike information criterion (AIC), as suggested by Modeltest in MEGA7 (Kumar et al. 2016). The approximate likelihood-ratio test (aLRT) was used to evaluate the supporting values of branches in the ML tree (Anisimova and Gascuel 2006). For the BI tree, two parallel Markov chain Monte Carlo (MCMC) simulations of 10 million generations that sampled every 1000 generations, with 25% burn-in was used for obtaining the consensus tree. Divergence time was estimated using the BEAST program under a relaxed clock model (Drummond and Rambaut 2007). Three internal fossil data crowns of Cupressoideae 157.2 (Mya), ThujaThujopsis clade 58.5 (Mya), and Juniperus 33.8 (Mya) were used as calibration points (Kangshan et al. 2012). MCMC procedures had a burn-in of 100 million iterations. The default settings were adopted for other parameters when performing BEAST analysis (Drummond and Rambaut 2007). Chronogram was drawn using FigTree v1.4.0 (http://tree.bio.ed.ac.uk/).

Results and discussion

The overall features of cp DNA of T. sutchuenensis

The gene map for the T. sutchuenensis cp genome is shown in Fig. 1. The complete cp genome of T. sutchuenensis is 129,776 bp, which is 729 bp shorter than that of T. standishii. The cp genome contained 118 unique genes, including 82 unique protein-coding genes, 32 tRNA genes, and 4 rRNA genes (Tables 1 and 2). Among these genes, 14 genes harbored a single intron (trnA-UGC, trnG-UCC, trnI-GAU, trnK-UUU, trnL-UAA, trnV-UAC, atpF, ndhA, ndhB, petB, petD, rpoC1, rpl2, rpl16) and two genes (ycf3 and rps12) harbored two introns (Table 3). The overall GC content of the T. sutchuenensis cp genome was 34.2%, which is similar to that of T. standishii. So did other genomic features, including the gene content, gene order, introns, and intergenic spacers.

Table 1.

Chloroplast genome characteristics of T. sutchuenensis and T. standishii

Thuja sutchuenensis Thuja standishii
Genome size (bp) 129,776 130,505
GC contents 34.4% 34.2%
Number of genes 82 82
Number of tRNA 32 32
Number of rRNA 4 4

Table 2.

Genes present in the T. sutchuenensis chloroplast genome

Group of gene Genes name
Photostsyem I psaA, psaB, psaC, psaI, psaJ, psaM,ycf3**, ycf4
Photostsyem II psbA, psbB, psbC, psbD, psbE, psbF, psbH, psbI, psbJ, psbK, psbL, psbM, psbN, psbT, psbZ
Cytochrome b/f complex petA, petB*, petD*, petG, petL, petN
ATP synthase atpA, atpB, atpE, atpF*, atpH, atpI
Chlorophyll biosynthesis chlB, chlL, chlN
NADH dehydrogenase ndhA*, ndhB*, ndhC, ndhD, ndhE, ndhF, ndhG, ndhH, ndhI, ndhJ, ndhK
RubisCO large subunit rbcL
RNA polymerase rpoA, rpoB, rpoC1*, rpoC2
Ribosomal proteins (SSU) rps2, rps3, rps4, rps7, rps8, rps11, rps12**,T, rps14, rps15, rps18, rps19
Ribosomal proteins (LSU) rpl2*, rpl14, rpl16*, rpl20, rpl22, rpl23, rpl32, rpl33, rpl36
Other gene clpP, matK, accD, ccsA, infA, cemA, ycf1, ycf2
Transfer RNAs 32 tRNAs (six contain a single intron)
Ribosomal RNAs rrn4.5, rrn5, rrn16, rrn23

A single asterisk (*) preceding gene names indicate intron-containing genes, and double asterisks (**) preceding gene names indicate two introns in the gene; T trans-splicing of the related gene

Table 3.

The genes with introns in the T. sutchuenensis chloroplast genome and the length of the exons and introns

Gene Exon I (bp) Intron I (bp) Exon II (bp) Intron II (bp) Exon III (bp)
atpF 145 672 230
ndhA 558 759 549
ndhB 723 699 756
petB 6 788 642
petD 8 677 493
rpl16 9 860 411
rpl2 400 643 431
rpoC1 471 748 1674
rps12 114 232 526 26
ycf3 126 765 228 705 156
trnA-UGC 38 776 35
trnG-UCC 24 749 48
trnI-GAU 37 904 35
trnK-UUU 37 2447 35
trnL-UAA 35 494 50
trnV-UAC 35 521 37

The cp genome of each Thuja species has 38 SSRs, comprised of motifs of mono- to penta-nucleotides. Most of the mononucleotide- and dinucleotide-motifs are comprised of A/T repeats, which are consistent with the view that cpSSRs are attributed to AT richness (Liu et al. 2018). There are 24 SSRs shared between cp genomes of T. sutchuenensis and T. standishii, which comprise 14 mononucleotide-, one dinucleotide-, two trinucleotide-, six tetranucleotide-, and one pentanucleotide-motif SSRs (Table 4; Fig. 2). Among these common SSRs, eight are exons of coding genes, five in introns, and the remaining nine are in non-coding regions. These SSRs can function as underlying biomarkers for studies of genetic diversity.

Table 4.

List of simple sequence repeats (SSRs) shared in T. sutchuenensis and T. standishii chloroplast genome

SSR type Base Location
p1 A/T rps4
p1 A/T psbMpetN
p1 A/T petNtrnC
p1 A/T rpoC1 intron
p1 A/T rps2atpI
p1 A/T atpIatpH
p1 A/T atpF intron
p1 A/T atpF intron
p1 A/T rpl32ndhF
p1 A/T rps19
p1 A/T rps8
p1 A/T rpl36rps11
p1 A/T petB intron
p1 A/T ndhCtrnV
P2 TA/AT rrn16trnV
P3 TTC/GAA matKchlB
P3 GAA/TTC rpoB
P4 TCCA/TGGA trnK intron
P4 TACT/AGTA rpoB
P4 TTTC/GAAA ndhF
P4 CTAC/GTAG rrn23
P4 CTTG/CAAG trnIrrn16
P4 TTAA/AATT trnIycf2
P5 ATATT/AATAT chlL

Fig. 2.

Fig. 2

Distribution of SSRs present in T. sutchuenensis and T. standishii chloroplast genomes

Rearrangements of cp genomes in Thuja

The extent of structural rearrangements between cp genomes of genera Thuja and Thujopsis can be visualized in the Mauve genome alignment. One large inversion was found between genes trnQ-UUQ and trnQ-UUG (Fig. 3). The inversion changes the orders of trnT, rps4, trnS, ycf3, psaA, psaB, rps14, trnfM, trnG, psbZ, trnS, psbC, psbD, trnT, trnE, trnY, trnD, psbM, petN, trnC, rpoB, rpoC1, rpoC2, rps2, atpI, atpH, atpF, atpA, trnR, trnG, psaM, trnS, psbI and psbK in T. sutchuenensis and T. standishii from the sister genus Thujopsis. Inversion is an important mechanism in plant evolution that forms and maintains interspecific differentiation. The standing inversion may predominate over new mutations in maintaining local adaptation, although not absolutely (He and Knowles 2017). The inversion of the plastid genome is associated with the origin of plant groups, for example, two cp inversion events of the family Asteraceae likely occurred around the same time during the origin of Asteraceae (Kim et al. 2005). This long-fragment inversion characterized the cp genome of Thuja, that can be used as a very reliable phylogenetic marker (Jansen and Palmer 1987; Doyle et al. 1992, 1996; Kim et al. 2005). Since inverted repeats (IRs) play critical roles in stabilizing cp genomes against major structural variations, and the loss of IRs could result in shorter intergenic spacers (Wu and Chaw 2014), more gene loss, and structural rearrangements (Zheng et al. 2016). One example of this is Pinaceae evolution, with short repeats that can increase the diversity of cpDNA variety and complement the reduced IRs (Wu et al. 2011). Thus, loss of the large IRs in Thuja and Thujopsis may be the main cause of rearrangements in gene block order.

Fig. 3.

Fig. 3

Synteny and rearrangements detected in Thuja and Thujopsis chloroplast genomes using the Mauve multiple-genome alignment. Colored outlined blocks surround regions of the genome sequence that align with part of another genome. The colored bars inside the blocks are related to the level of sequence similarities. Lines link blocks with homology between two genomes. The red box appears as an inverted region (colour figure online)

Evolution of T. sutchuenensis and T. standishii

To pinpoint whether any genes of the cp genome underwent adaptive evolution in T. sutchuenensis, we compared the cp genome between T. sutchuenensis and the closely related species T. standishii. Our analysis showed that the average Ka/Ks ratio of 82 protein genes was 0.256 between two cp genomes, which are similar to previous studies of Haberlea rhodopensis cp genome gene evolution (Ivanova et al. 2017). Among these 82 genes, there are 72 genes reveal a low synonymous substitution rate (Ks < 0.01 between two species), suggesting recent species divergence or genetic constraint between T. sutchuenensis and T. standishii. Since Ka/Ks are primarily gene-specific, we calculated the Ka/Ks of each gene one by one. Seven genes, accD, ycf2, ycf1, rpl22, rps18, rpoC1, and rps15, have an estimated Ka/Ks > 1. We further drew a plot of Ka/Ks against Ks to assess the timing of the adaptive divergence of these genes. Overall, the distribution of these seven genes is not left-skewed in Ka/Ksversus Ks plot (Fig. 4), excluding the recent adaptive divergence of the cp genome between T. sutchuenensis and T. standishii.

Fig. 4.

Fig. 4

Gene-specific Ks and Ka/Ks values between the chloroplast genomes of Thuja sutchuenensis and T. standishii

The accD gene, which has a basic metabolic function in producing acetyl-CoA carboxylase, has the highest Ka/Ks ratio in the cp genome. It plays multiple ecological functions of the tolerance and resistance of stresses, insect predation, and pathogens (Sablok et al. 2016), explaining the signature of positive selection it presents. The genes of ycf1 and ycf2 showed the second and third highest Ka/Ks. Although their functions are unclear, the positive-selection signal was also reported in buckwheat (Cho et al. 2015), sesame (Zhang et al. 2013), Pinus (Parks et al. 2009), and orchids (Neubig et al. 2009). The ycf1 and ycf2, with their high evolutionary rate, are further suggested as molecular markers for phylogenetic reconstruction (Neubig et al. 2009; Parks et al. 2009). Three ribosomal protein genes rpl22, rps18, and rps15 also reveal a high Ka/Ks ratio, and have been demonstrated to play versatile roles in plant growth and development (Fleischmann et al. 2011). Thuja sutchuenensis and T. standishii are distributed in southwestern China and southern Japan, respectively. Highly differential environmental and growth conditions between these two species (Kusumi et al. 2006; Tang et al. 2015) may facilitate the adaptive divergence of these cp genes. The mean annual temperature of the major habitat areas of T. sutchuenensis in the Dabashan Nature Reserve of Chengkou County is 14 °C, with 2.7 °C in the coldest month (January) (Tang et al. 2015). In contrast, T. standishii grows primarily in regions where the climate is relatively cool and humid, and the winter is snowy and much colder (Kusumi et al. 2006). Such climatic differences may lead to the differential efficiency of physiological and pathological processes of the cp genome, particularly the ribosomal protein genes (Xiang et al. 2015), between these two Thuja species. However, we cannot preclude the possibility of false positives since these genes do not deviate too far from neutrality (Ka/Ks = 1).

Phylogenetic analysis and estimation of divergence time

The cp genome sequence is useful for systematics. Specific relationships of the Cupressoideae remain obscure due to complicated evolutionary history. Numerous studies have tried to resolve intergeneric phylogenetic relationships (Mao et al. 2012; Yang et al. 2012; Qu et al. 2017). To acquire a reasonable phylogeny and molecular dating of Cupressoideae, we reconstruct a phylogenetic tree based on cp genomes. A total of 95,128 bp length LCB sequences were recognized by HomBlocks. The LCB sequence contains all of the coding genes, rRNAs genes, and a part of tRNAs. The tRNAs of trnL-CAA, trnI-CAU, trnQ-UUG, trnN-GUU, trnT-GGU, trnQ-UUG in T. sutchuenensis didn’t include in LCB sequence (Fig. S1 in the Supplementary materials). The reconstructed phylogenetic topology of the ML and BI analysis had high bootstrap supports that can provide a proper evolutionary placement for T. sutchuenensis.

The highly supportive clustering of T. sutchuenensis and T. standishii indicates efficient discrimination of the cp genome between Thuja and other genera. In our reconstructed phylogeny, the TujaTujopsis cluster diverged earlier from other genera in Cupressoideae clade, following in sequence Chamaecyparis, clade CalocedrusPlatycladus, and clade Juniperus (J)–Cupressus (C)–HesperocyparisCallitropsis (HCX, here the X, Xanthocyparis, is not included) (Fig. 5), which is similar to Qu et al.’s (2017) inference. Our inference, however, partially disagrees with the inference of “(J, (C, (HCX)))” by Qu et al. (2017), but we agree with their hypothesis that the long-branch attraction may blur the phylogenetic relationships among these genera, particularly for the relatively low supporting values in the J–C cluster (Fig. 5).

Fig. 5.

Fig. 5

Phylogenetic tree constructed by maximum likelihood (ML) and Bayesian inference (BI) methods based on chloroplast genome locally collinear blocks sequence of 23 Cupressoideae. The phylogenetic tree was drawn using Cunninghamia lanceolata, Taiwania crypyomerioides, and T. flousiana as outgroup. Nodes with 100/100 support values not showed. ML support values were in the front

Previous studies of T. sutchuenensis were based on several DNA segments (Jian-Hua and Xiang 2005; Yang et al. 2012), which may be heavily biased in estimating the true branch length. To compensate for this deficiency, we applied the LCBs of complete cp genome sequences to estimate the splitting times between T. sutchuenensis and relatives. Divergence time within the Cupressaceae between subclades J and C-HCX was estimated at about 49.72 million years ago (Mya), and the split between C and HCX about 46.29 Mya (Fig. 6), which is similar to the estimation by Qu et al. (2017). The splitting time between genera Thuja and Thujopsis is 58.46 Mya (Fig. 6), which is also consistent with Qu et al.’s (2017) estimation at 58.5 Mya. Our study further estimated the divergence time of T. sutchuenensis and T. standishii, which is roughly at the base of the Late Miocene (11.05 Mya, Fig. 6). Our dating is almost consistent, although slightly earlier, with the previous estimation based on nuclear ribosomal DNA internal transcribed spacer (9.53 ± 2.30 Mya) (Jian-Hua and Xiang 2005). This stage was roughly after the Maximum Transgression in the Tortonian Period (Late Miocene), when the climate was warmer than it is now (Ogg et al. 2016). Warmer climates may result in population decline for these cooler adapted species, accelerating the differentiation of populations and/or species. Although this inference still needs further verification, the estimates given are reasonable.

Fig. 6.

Fig. 6

Bayesian chronogram for Cupressoideae and related subfamilies. Estimates of divergence times for major clades are listed, and blue bars show the 95% highest posterior density (HPD) of relevant nodes (colour figure online)

Conclusions

In this study, we present the whole cp genome of endangered species T. sutchuenensis. The genome size and genomic contents are similar to other Cupressaceae species. We also found an inversion between trnT and psbK between Thuja and Thujopsis. In addition, seven genes underlying positive selection were inferred by high Ka/Ks ratio. We also found 38 interspecific cpSSRs between T. sutchuenensis and T. standishii and estimated a splitting time at approximately 11.05 Mya according to the cp genome sequences. In addition to validating previous research results with cp genome, this study provides additional information that can facilitate fundamental research in systematics, population genetics, and applications for the DNA barcoding and genetic engineering of this species.

Electronic supplementary material

Below is the link to the electronic supplementary material.

12298_2019_736_MOESM1_ESM.png (714.6KB, png)

Fig. S1 Corresponding analysis between locally collinear blocks sequence and T. sutchuenensis genes sequences. (PNG 714 kb)

Acknowledgements

This research was financially supported by the program ‘‘Reintroduction Technologies and Demonstration of Extremely Rare Wild Plant Population’’ of National Key Research and Development Program (2016YFC0503106) and subsidized by the Ministry of Science and Technology of Taiwan (Grant Nos.: MOST 105–2628-B-003–001-MY3 and MOST 105–2628-B-003–002-MY3) and National Taiwan Normal University (NTNU).

Data archiving

The complete chloroplast genome sequence data of Thuja sutchuenensis has been submitted to GenBank of NCBI with accession number MH784400. Raw data used for the assembly of the chloroplast genome was submitted to NCBI, the SRA accession: PRJNA578239. The data will be available publically after the acceptance of the manuscript.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Pei-Chun Liao, Email: pcliao@ntnu.edu.tw.

Jun-Qing Li, Email: lijq@bjfu.edu.cn.

References

  1. Anisimova M, Gascuel O. Approximate likelihood-ratio test for branches: a fast, accurate, and powerful alternative. Syst Biol. 2006;55:539–552. doi: 10.1080/10635150600755453. [DOI] [PubMed] [Google Scholar]
  2. Bi G, Mao Y, Xing Q, Cao M. HomBlocks: a multiple-alignment construction pipeline for organelle phylogenomics based on locally collinear block searching. Genomics. 2017;110:18–22. doi: 10.1016/j.ygeno.2017.08.001. [DOI] [PubMed] [Google Scholar]
  3. Brozynska M, Furtado A, Henry RJ. Genomics of crop wild relatives: expanding the gene pool for crop improvement. Plant Biotechnol J. 2016;14:1070–1085. doi: 10.1111/pbi.12454. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Cho KS, Yun BK, Yoon YH, et al. Complete chloroplast genome sequence of tartary buckwheat (Fagopyrum tataricum) and comparative analysis with common buckwheat (F. esculentum) PLoS ONE. 2015;10:e0125332. doi: 10.1371/journal.pone.0125332. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Chumley TW, Palmer JD, Mower JP, et al. The complete chloroplast genome sequence of Pelargonium × hortorum: organization and evolution of the largest and most highly rearranged chloroplast genome of land plants. Mol Biol Evol. 2006;23:2175–2190. doi: 10.1093/molbev/msl089. [DOI] [PubMed] [Google Scholar]
  6. Dai M, Thompson RC, Maher C, et al. NGSQC: cross-platform quality analysis pipeline for deep sequencing data. BMC Genom. 2010;11(Suppl 4):S7. doi: 10.1186/1471-2164-11-S4-S7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Daniell H, Lin CS, Ming Y, Chang WJ. Chloroplast genomes: diversity, evolution, and applications in genetic engineering. Genome Biol. 2016;17:134. doi: 10.1186/s13059-016-1004-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Darling ACE, Mau B, Blattner FR, Perna NT. Mauve: multiple alignment of conserved genomic sequence with rearrangements. Genome Res. 2004;14:1394–1403. doi: 10.1101/gr.2289704. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Darzentas N. Circoletto: visualizing sequence similarity with Circos. Bioinformatics. 2010;26:2620–2621. doi: 10.1093/bioinformatics/btq484. [DOI] [PubMed] [Google Scholar]
  10. Dierckxsens N, Mardulyn P, Smits G. NOVOPlasty: de novo assembly of organelle genomes from whole genome data. Nucl Acids Res. 2016;45:e18. doi: 10.1093/nar/gkw955. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Doyle JJ, Davis JI, Soreng RJ, et al. Chloroplast DNA inversions and the origin of the grass family (Poaceae) Proc Natl Acad Sci USA. 1992;89:7722–7726. doi: 10.1073/pnas.89.16.7722. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Doyle JJ, Doyle JL, Ballenger JA, Palmer JD. The distribution and phylogenetic significance of a 50-kb chloroplast DNA inversion in the flowering plant family leguminosae. Mol Phylogenetics Evol. 1996;5:429–438. doi: 10.1006/mpev.1996.0038. [DOI] [PubMed] [Google Scholar]
  13. Drummond AJ, Rambaut A. BEAST: Bayesian evolutionary analysis by sampling trees. BMC Evol Biol. 2007;7:214. doi: 10.1186/1471-2148-7-214. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Fleischmann TT, Scharff LB, Alkatib S, et al. Nonessential plastid-encoded ribosomal proteins in tobacco: a developmental role for plastid translation and implications for reductive genome evolution. Plant Cell. 2011;23:3137–3155. doi: 10.1105/tpc.111.088906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Gao J, Liu Y, Bogonovich M. Habitat is more important than climate and animal richness at shaping latitudinal variation in plant diversity in China. Biodivers Conserv. 2018 doi: 10.1007/s10531-018-1620-0. [DOI] [Google Scholar]
  16. Guindon S, Dufayard JF, Hordijk W, et al. PhyML: fast and accurate phylogeny reconstruction by maximum likelihood. Infect Genet Evol. 2009;9:384–385. [Google Scholar]
  17. Hahn C, Bachmann L, Chevreux B. Reconstructing mitochondrial genomes directly from genomic next-generation sequencing reads—a baiting and iterative mapping approach. Nucl Acids Res. 2013;41:e129. doi: 10.1093/nar/gkt371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. He Q, Knowles LL (2017) Rapid adaptation with gene flow via a reservoir of chromosomal inversion variation? bioRxiv: 150771. 10.1101/150771
  19. Huelsenbeck JP, Ronquist F. MRBAYES: Bayesian inference of phylogenetic trees. Bioinformatics. 2001;17:754–755. doi: 10.1093/bioinformatics/17.8.754. [DOI] [PubMed] [Google Scholar]
  20. Ivanova Z, Sablok G, Daskalova E, et al. Chloroplast genome analysis of resurrection tertiary relict Haberlea rhodopensis highlights genes important for desiccation stress response. Front Plant Sci. 2017;8:204. doi: 10.3389/fpls.2017.00204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Jansen RK, Palmer JD. A chloroplast DNA inversion marks an ancient evolutionary split in the sunflower family (Asteraceae) Proc Natl Acad Sci USA. 1987;84:5818–5822. doi: 10.1073/pnas.84.16.5818. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Jian-Hua LI, Xiang QP. Phylogeny and biogeography of Thuja L. (Cupressaceae), an Eastern Asian and North American Disjunct Genus. J Integr Plant Biol. 2005;47:651–659. doi: 10.1111/j.1744-7909.2005.00087.x. [DOI] [Google Scholar]
  23. Katoh K, Kuma K, Toh H, Miyata T. MAFFT version 5: improvement in accuracy of multiple sequence alignment. Nucl Acids Res. 2005;33:511–518. doi: 10.1093/nar/gki198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Kim KJ, Choi KS, Jansen RK. Two chloroplast DNA inversions originated simultaneously during the early evolution of the sunflower family (Asteraceae) Mol Biol Evol. 2005;22:1783–1792. doi: 10.1093/molbev/msi174. [DOI] [PubMed] [Google Scholar]
  25. Kumar S, Stecher G, Tamura K. MEGA7: molecular evolutionary genetics analysis version 7.0 for bigger datasets. Mol Biol Evol. 2016;33:1870. doi: 10.1093/molbev/msw054. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Kusumi J, Sato A, Tachida H. Relaxation of functional constraint on light-independent protochlorophyllide oxidoreductase in Thuja. Mol Biol Evol. 2006;23:941–948. doi: 10.1093/molbev/msj097. [DOI] [PubMed] [Google Scholar]
  27. Li H, Handsaker B, Wysoker A, et al. The sequence alignment/map (SAM) format and SAMtools. Bioinformatics. 2009;25:1653–1654. doi: 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Librado P, Rozas J. DnaSP v5: a software for comprehensive analysis of DNA polymorphism data. Bioinformatics. 2009;25:1451–1452. doi: 10.1093/bioinformatics/btp187. [DOI] [PubMed] [Google Scholar]
  29. Liu J, Shi S, et al. Genetic diversity of the critically endangered Thuja sutchuenensis revealed by ISSR markers and the implications for conservation. Int J Mol Sci. 2013;14:14860–14871. doi: 10.3390/ijms140714860. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Liu X, Li Y, Yang H, Zhou B. Chloroplast genome of the folk medicine and vegetable plant Talinum paniculatum (Jacq.) Gaertn.: gene organization, comparative and phylogenetic analysis. Molecules. 2018;23:857. doi: 10.3390/molecules23040857. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Lohse M, Drechsel O, Bock R. OrganellarGenomeDRAW (OGDRAW): a tool for the easy generation of high-quality custom graphical maps of plastid and mitochondrial genomes. Curr Genet. 2007;52:267–274. doi: 10.1007/s00294-007-0161-y. [DOI] [PubMed] [Google Scholar]
  32. Mao K, Milne RI, Zhang L, et al. Distribution of living Cupressaceae reflects the breakup of Pangea. Proc Natl Acad Sci USA. 2012;109:7793–7798. doi: 10.1073/pnas.1114319109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Meyer M, Kircher M. Illumina sequencing library preparation for highly multiplexed target capture and sequencing. Cold Spring Harb Protoc. 2010;2010:pdb.prot5448. doi: 10.1101/pdb.prot5448. [DOI] [PubMed] [Google Scholar]
  34. Neubig KM, Whitten WM, Carlsward BS, et al. Phylogenetic utility of ycf 1 in orchids: a plastid gene more variable than matK. Plant Syst Evol. 2009;277:75–84. doi: 10.1007/s00606-008-0105-0. [DOI] [Google Scholar]
  35. Ogg JG, Ogg G, Gradstein FM. A concise geologic time scale: 2016. Amsterdam: Elsevier; 2016. [Google Scholar]
  36. Parks M, Cronn R, Liston A. Increasing phylogenetic resolution at low taxonomic levels using massively parallel sequencing of chloroplast genomes. BMC Biol. 2009;7:84. doi: 10.1186/1741-7007-7-84. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Qiaoping X, Fajon A, Zhenyu L, et al. Thuja sutchuenensis: a rediscovered species of the Cupressaceae. Bot J Linn Soc. 2015;139:305–310. doi: 10.1046/j.1095-8339.2002.00055.x. [DOI] [Google Scholar]
  38. Qu XJ, Jin JJ, Chaw SM, et al. Multiple measures could alleviate long-branch attraction in phylogenomic reconstruction of Cupressoideae (Cupressaceae) Sci Rep. 2017;7:41005. doi: 10.1038/srep41005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Sablok G, Mudunuri SB, Edwards D, Ralph PJ. Chloroplast genomics: expanding resources for an evolutionary conserved miniature molecule with enigmatic applications. Curr Plant Biol. 2016;7:34–38. doi: 10.1016/j.cpb.2016.12.004. [DOI] [Google Scholar]
  40. Saski C, Lee SB, Fjellheim S, et al. Complete chloroplast genome sequences of Hordeum vulgare, Sorghum bicolor and Agrostis stolonifera, and comparative analyses with other grass genomes. Theor Appl Genet. 2007;115:591. doi: 10.1007/s00122-007-0595-0. [DOI] [PubMed] [Google Scholar]
  41. Tang CQ, Yang Y, Ohsawa M, et al. Community structure and survival of tertiary relict Thuja sutchuenensis (Cupressaceae) in the subtropical daba mountains, Southwestern China. PLoS ONE. 2015;10:e0125307. doi: 10.1371/journal.pone.0125307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Wambugu PW, Brozynska M, Furtado A, et al. Relationships of wild and domesticated rices (Oryza AA genome species) based upon whole chloroplast genome sequences. Sci Rep. 2015;5:13957. doi: 10.1038/srep13957. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Wu CS, Chaw SM. Highly rearranged and size-variable chloroplast genomes in conifers II clade (cupressophytes): evolution towards shorter intergenic spacers. Plant Biotechnol J. 2014;12:344–353. doi: 10.1111/pbi.12141. [DOI] [PubMed] [Google Scholar]
  44. Wu CS, Lin CP, Hsu CY, et al. Comparative chloroplast genomes of Pinaceae: insights into the mechanism of diversified genomic organizations. Genome Biol Evol. 2011;3:309–319. doi: 10.1093/gbe/evr026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Xiang Z, Wen-Juan L, Jun-Ming L, et al. Ribosomal proteins: functions beyond the ribosome. J Mol Cell Biol. 2015;7:92–104. doi: 10.1093/jmcb/mjv014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Yang ZY, Ran JH, Wang XQ. Three genome-based phylogeny of Cupressaceae s.l.: further evidence for the evolution of gymnosperms and Southern Hemisphere biogeography. Mol Phylogenetics Evol. 2012;64:452–470. doi: 10.1016/j.ympev.2012.05.004. [DOI] [PubMed] [Google Scholar]
  47. Yi X, Gao L, Wang B, et al. The complete chloroplast genome sequence of Cephalotaxus oliveri (Cephalotaxaceae): evolutionary comparison of Cephalotaxus chloroplast DNAs and insights into the loss of inverted repeat copies in gymnosperms. Genome Biol Evol. 2013;5:688–698. doi: 10.1093/gbe/evt042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Yong Y, Zheng X. CpGAVAS, an integrated web server for the annotation, visualization, analysis, and GenBank submission of completely sequenced chloroplast genome sequences. BMC Genom. 2012;13:715. doi: 10.1186/1471-2164-13-715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Zhang H, Li C, Miao H, Xiong S. Insights from the complete chloroplast genome into the evolution of Sesamum indicum L. PLoS ONE. 2013;8:e80508. doi: 10.1371/journal.pone.0080508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Zheng W, Chen J, Hao Z, Shi J. Comparative analysis of the chloroplast genomic information of Cunninghamia lanceolata (Lamb.) Hook with sibling species from the Genera Cryptomeria D. Don, Taiwania Hayata, and Calocedrus Kurz. Int J Mol Sci. 2016;17:1084. doi: 10.3390/ijms17071084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Zonneveld BJM. Conifer genome sizes of 172 species, covering 64 of 67 genera, range from 8 to 72 picogram. Nord J Bot. 2012;30:490–502. doi: 10.1111/j.1756-1051.2012.01516.x. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12298_2019_736_MOESM1_ESM.png (714.6KB, png)

Fig. S1 Corresponding analysis between locally collinear blocks sequence and T. sutchuenensis genes sequences. (PNG 714 kb)


Articles from Physiology and Molecular Biology of Plants are provided here courtesy of Springer

RESOURCES