Abstract
Background
Rubus belong to the Rosaceae family, and they are commonly used as nutritious foods and medicinal plants in China. The fruits of Rubus plants are rich in nutrients and are renowned for their excellent antioxidant properties. Many Rubus plants can be processed into Chinese medicinal materials, with the fruits of Rubus chingii being a typical example. However, due to the similar appearance of Rubus fruits, they cannot be well distinguished after processing. Therefore, this study seeks to characterize the chloroplast genomes of 28 Rubus species and examine variations in their sequences, so as to provide valid data support for the identification of closely related species and the analysis of their genetic evolutionary relationships.
Results
Our research found that the chloroplast genomes of the 28 Rubus species possess a canonical quadripartite organization, comprising one Large Single Copy (LSC) region, a pair of inverted repeat (IR) regions(IRa and IRb), and one Small Single Copy (SSC) region. Their overall size varied between 155,470 and 156,429 bp, showing only slight fluctuations, and the GC contents ranged from 36.91% to 37.28%. The RSCU analysis revealed that UUA was the most frequently used synonymous codon, whereas its synonymous counterparts CUC and CUG exhibited the lowest RSCU values. Shrinkage and expansion phenomena were primarily observed at the junctions between the SSC region and IR region in the chloroplast genomes. Repeat sequences demonstrated that A/T base pairs were most prevalent in simple sequence repeats (SSRs). Nucleotide polymorphism analysis indicated that there were multiple regions with high PI values in the chloroplast genes of the 28 Rubus species, namely trnH-psbA, rps16-trnQ 1, rps16-trnQ 2, trnS-trnG, petN-psbM, petA-psbJ, psbE-petL and rpl32-trnL. Codon usage analysis showed that most genes had Ka/Ks values less than one. Phylogenetic analysis showed that these 28 genomes were mainly divided into two major clades: Rubus chingii was placed in Clade B and was most closely related to Rubus corchorifolius. We further validated the molecular markers rpl32-trnL and rps16-trnQ2, which showed efficiency in species identification.
Conclusion
These findings enrich our knowledge of Rubus chloroplast genomes, and offer valuable molecular data for its taxonomic classification, evolutionary exploration and species discrimination.
Supplementary Information
The online version contains supplementary material available at https://doi.org/10.1186/s12864-026-13120-z.
Keywords: Rubus, Chloroplast genome, Genomic structure, Phylogenetic analysis
Introduction
The Rubus L. is a nutritive plant commonly used as a functional food and medicine in China, belonging to the Rosaceae family, widely recognized for its nutritional and medicinal value in China. This genus comprises numerous species, including Rubus chingii, Rubus lambertianus, Rubus idaeus, Rubus corchorifolius, Rubus rosifolius, and Rubus parvifolius among others [1]. The fruits of Rubus species are highly nutritious and are renowned for their exceptional antioxidant properties, which surpass those of most other fruits [2]. Due to their health benefits, Rubus fruits are often referred to as “fruits of life.” Among these species, R. chingii holds particular significance as a traditional Chinese medicinal herb, primarily cultivated in provinces such as Zhejiang, Jiangsu, Jiangxi, Anhui, Guangxi, and Fujian [3]. The fruit of R. chingii is rich in bioactive compounds, including ellagic acid, flavonoids, and tannins [4]. Recent studies have highlighted its diverse pharmacological properties, such as antioxidant, neuroprotective, hepatoprotective, and anticancer activities [5–7].
Chloroplasts are critical organelles in plant cells, possessing their own independent genome. Research has consistently demonstrated that the chloroplast genome structure is highly conserved in most flowering plants, typically exhibiting a quadripartite organization consisting of a large single-copy (LSC) region, a small single-copy (SSC) region, and two inverted repeat (IR) regions. The chloroplast genome also shows significant rate heterogeneity between coding and non-coding sequences. Due to its low sequencing cost, minimal assembly errors, and high conservation, the chloroplast genome has become a valuable tool for studying plant phylogeny, taxonomy, and species identification. Over the past three decades, comparative analyses of chloroplast genomes have been increasingly utilized to resolve phylogenetic relationships, clarify taxonomic classifications, and delineate species boundaries in plants.
The genus Rubus presents a significant phylogenetic challenge within flowering plants, yet it also serves as an excellent system for investigating plant reproductive evolution. Species boundaries within Rubus are often indistinct, characterized by extensive morphological variation and convergent evolution. These complexities are further compounded by widespread polyploidy, hybridization, and apomixis [8]. In this study, we selected 28 Rubus species for chloroplast genome comparative analysis, including one newly sequenced Rubus chingii obtained in our experiment, and the remaining 27 chloroplast genome sequences downloaded from the NCBI GenBank database.
Materials and methods
DNA sequencing and data collection
The plant material of Rubus chingii used in this research were collected from cultivated fields in Hangzhou County (29.97°N, 119.53°E), Zhejiang Province, China. Healthy young leaves of Rubus chingii were collected and immediately flash-frozen in liquid nitrogen, followed by storage at ultra-low temperature freezer (-80℃). Genomic DNA was extracted using a kit-based method (FastPure Plant DNA Isolation Mini Kit-BOX2, Vazyme, Cat. DC104-01). The isolated DNA was randomly fragmented into smaller segments with an Ultrasound Covaris instrument, generating a series of short DNA fragments. The fragmented DNA was subsequently purified, end-repaired, and subjected to 3’ end A-tailing. Isolated DNA’s quality and concentration were verified using agarose gel electrophoresis and ultra micro nucleic acid analyzer (K5500plus), with subsequent PCR amplification to construct a sequencing library. The library was initially quality-checked and, upon qualification, sequenced using the Illumina HiSeq platform in a paired-end (PE150) format. All experimental procedures followed the standard protocols provided by Sangon Biotech (Shanghai, China), Including sample quality assessment, sequencing library construction, library quality control, and library-based sequencing.
Other 27 chloroplast genome sequences were published Rubus species were downloaded from NCBI, including Rubus coreanus、Rubus pinfaensis、Rubus takesimensis、Rubus crataegifolius、Rubus boninensis、Rubus trifidus、Rubus niveus、Rubus lambertianus、Rubus columellaris、Rubus peltatus、Rubus crassifolius、Rubus setchuenensis、Rubus sachalinensiss、Rubus glandulosopunctatus、Rubus rubroangustifolius、Rubus taiwanicola、Rubus parvifolius、Rubus occidentalis、Rubus conduplicatus、Rubus hirsutus、Rubus innominatus、Rubus corchorifolius、Rubus idaeus、Rubus calycacanthus、Rubus fockeanus、Rubus irenaeus、Rubus rosifolius. The MAFFT v7.0 software [9] was employed to align the sequences of 28 complete chloroplast genomes. Subsequently, these aligned sequences were manually inspected and adjusted for the following analysis.
Chloroplast assembly and annotation
The assembly of the chloroplast genome of Rubus chingii was carried out using clean data with GetOrganille v1.7.2 [10]. Assembled sequences were analyzed via BLAST on NCBI (https://blast.ncbi.nlm.nih.gov/Blast.cgi), with CPGAVAS2’s default parameters (http://47.96.249.172: 16019/analyzer/annotate) [11]. The annotation results were further manually curated and polished using Apollo v1.11.8 [12] to obtain the final annotated plastome. To evaluate assembly quality, clean reads were mapped back to the assembled chloroplast genome. The average sequencing depth was 3838.82×, and the genome-wide coverage reached 100%, with no gaps or ambiguous bases. These metrics confirmed the high reliability and completeness of the assembly. The final chloroplast genome map was visualized using the online tool Chloroplot (https://irscope.shinyapps.io/Chloroplot/).
RSCU analysis
Protein-coding sequences (CDS) were extracted from the chloroplast genomes of the 28 Rubus accessions included in this study. To minimize analytical bias and ensure the robustness of codon usage inference, redundant sequences were first removed. The remaining CDS were then filtered to retain only those with a total length divisible by three, consistent with the requirement of intact open reading frames. Only sequences longer than 300 bp were retained to avoid stochastic variation from short fragments, and all sequences were further verified to contain canonical start and stop codons. The relative synonymous codon usage (RSCU) values were subsequently calculated using MEGA 11 software. Than we visualized the relevant results using the RSCU-Plot Shiny app (https://pcg-lab.shinyapps.io/RSCU-Plot/) [13–16].
Long repeats and SSRs identification
Repetitive sequences in the chloroplast genomes from 28 Rubus species were investigated to identify simple sequence repeats (SSRs) and long repeat sequences. For the identification of dispersed repeats, encompassing tandem, forward, reverse, palindromic, and complement types, the online software REPuter (https://bibiserv.cebitec.uni-bielefeld.de/reputer/) [17] was utilized, with a minimum repeat length of 30 bp and a Hamming distance of 3 specified as parameters. SSR detection was performed via the Python3-based CPStools program [18]. The screening criteria were set as mononucleotide repeats ≥ 8, dinucleotide repeats ≥ 5, trinucleotide repeats ≥ 4, and tetranucleotide, pentanucleotide, and hexanucleotide repeats each ≥ 3. All identified SSR loci were further divided into intergenic spacer (IGS), coding gene, and intronic regions based on their genomic locations. Tandem repeats were characterized with Tandem Repeat Finder (TRF; https://tandem.bu.edu/trf/trf.html) under default parameter settings.
Comparative genomic analysis
The Rubus comparative chloroplast genome analysis included the examination of Inverted Repeat (IR) boundary variations. The shrinkage and expansion phenomena of IR regions relative to the LSC and SSC regions were visualized via the online program IRScope (https://irscope.shinyapps.io/irapp/) [19], with a comparative schematic diagram generated to illustrate structural divergence. For sequence variation analysis, GBK-format chloroplast genome files of the 28 Rubus species were programmatically converted to mVISTA-compatible formats using Python 3.10.1 [20], followed by uploading the annotated sequences to the mVISTA platform (http://genome.lbl.gov/vista/mvista/submit.shtml) [21]. Comparative characterization of complete chloroplast genomes among 28 Rubus species was conducted via the Shuffle-LAGAN algorithm to delineate genomic variations.
Selective pressure analysis
To evaluate the selective pressure acting on chloroplast protein-coding genes among 28 Rubus species, all valid coding sequences (CDS) were firstly extracted from the assembled chloroplast genomes using PhyloSuite software. The retrieved CDS sequences were subsequently aligned via the MAFFT and ensure homologous sequence consistency. Alignments were converted to Axt format for Ka/Ks calculation using the Ka/Ks calculation module of Python3-based CPStools [18]. Finally, the nonsynonymous substitution rate (Ka), synonymous substitution rate (Ks), and their ratio (Ka/Ks) were calculated for each homologous gene pair to quantify the evolutionary selection patterns of chloroplast genes across Rubus species.
Phylogenetic analysis
To analyze the phylogenetic placement of the 28 Rubus species, we performed phylogenetic analysis based on entire chloroplast genome sequences, with Fragaria pentaphylla (NC_034347) and Fragaria nilgerrensis (MK560340) serving as the outgroup. The complete chloroplast genome of Rubus chingii was obtained through sequencing, while the remaining sequences were obtained through NCBI. Multiple sequence alignment was conducted using MAFFT v7.0, and maximum likelihood (ML) analysis was employed to reconstruct the phylogeny. The best-fit substitution model for phylogenetic inference, selected via IQTREE v1.6.8 based on the Bayesian Information Criterion (BIC), was TVM + F+I+G4 [22]. Branch support was assessed with 1000 bootstrap replicates.
Validation of hypervariable region IGS markers
Four original plant samples belonging to Rubus chingii, Rubus corchorifolius, Rubus lambertianus and Rubus hirsutus were collected from Hangzhou and Ningbo. Total genomic DNA was isolated from all individual specimens using FastPure Plant DNA Isolation Mini Kit-BOX2 (Vazyme, Cat. DC104-01).Target primers located in conserved flanking regions were designed based on the screened hypervariable IGS fragments. The Primer Design module embedded in Geneious software was applied to screen primers with proper fragment length, and all synthesized primers were obtained from Sangon Biotech (Shanghai) Co., Ltd.Subsequent PCR amplification was carried out under the following thermal cycling conditions: initial denaturation at 95 °C for 5 min; 95 °C for 30 s, 55 °C for 35 s and 72 °C for 30 s, 35 cycles; final extension at 72 °C for 10 min, and 4 °C hold.
Raw sequencing peak graphs and nucleotide sequences were assembled into consensus sequences in Geneious software via de novo assembly strategy, meanwhile low-quality reads and redundant primer sequences were trimmed off. Multiple sequence alignment and manual sequence correction were conducted using MAFFT program. Combining the newly acquired hypervariable IGS sequences of chloroplast genome in this study with previously published homologous sequences, phylogenetic trees were constructed to clarify the phylogenetic distribution characteristics among different Rubus species.
Results
Genome characteristics
Through high-throughput sequencing, our study successfully obtained the complete chloroplast genome of Rubus chingii (GenBank: PQ 510071), which exhibits a typical quadripartite structure common to angiosperms. The circular genome spans 155,555 bp in length (Fig. 1), comprising a LSC region of 85,331 bp and a SSC region of 18,704 bp, separated by a pair of IR regions (IRa and IRb) each measuring 25,760 bp. Plastome annotation identified 130 functional genes in this chloroplast genome: 85 are protein-coding (CDS), 37 are tRNA genes, and 8 are rRNA genes, a gene complement consistent with most angiosperm plastids (Table 1). Comparative analysis was conducted with 27 additional Rubus chloroplast genomes retrieved from NCBI. Across the 27 examined Rubus species, the overall size of chloroplast genomes from 155,470 to 156,429 bp, reflecting relative structural conservation within the genus. Structural parameters showed limited variation: LSC regions (85,029–86,130 bp), SSC regions (18,697 − 18,876 bp), and IR regions (25,737 − 26,016 bp for both IRa and IRb) demonstrated high sequence conservation (Table 2). Notably, the R. chingii chloroplast genome represents the second shortest among the newly sequenced specimens. This minimal size variation underscores the remarkable structural conservation of chloroplast genomes within Rubus species, consistent with the slow evolutionary rate characteristic of plant organellar genomes.
Fig. 1.

Chloroplast genome map of Rubus chingii
Table 1.
List of genes annotated in the chloroplast genome of Rubus chingii
| Gene function | Group of genes | Gene name |
|---|---|---|
| Photosynthesis related | Subunit of ATP synthase | atpA, atpF, atpH, atpI, atpE, atpB |
| Subunit of NADH-plastoquinone oxidoreductase | ndhJ, ndhK, ndhC, ndhB ndhF, ndhD, ndhE, ndhI, ndhA, ndhH, ndhB | |
| Photosystem I | psaB, psaA, psaI, psaJ, psaC | |
| Photosystem II | psbA, psbK, psbI, psbM, psbD, psbC, psbZ, psbJ, psbL, psbF, psbE, psbB, psbT, psbN, psbH | |
| Subunit of ribulose 1,5-bisphosphate carboxylase/oxygenase | rbcL | |
| Cytochrome b6/f complex | petN, petL, petG, petD | |
| Cytochrome f | petA | |
| Cytochrome b6 | petB | |
| Self-replication related | Subunit of RNA polymerase | rpoC2, rpoC1, rpoB, rpoA |
| Ribosomal RNA genes | rrn16, rrn23, rrn4.5, rrn5 rrn16, rrn23, rrn4.5, rrn5 | |
| Transfer RNA genes | trnH-GUG, trnK-UUU, trnQ-UUG, trnS-GCU, trnG-UCC, trnR-UCU, trnC-GCA, trnD-GUC, trnY-GUA, trnE-UUC, trnT-GGU, trnS-UGA, trnG-GCC, trnfM-CAU, trnS-GGA, trnT-UGU, trnL-UAA, trnF-GAA, trnV-UAC, trnM-CAU, trnW-CCA, trnP-UGG, trnI-CAU, trnL-CAA, trnV-GAC, trnI-GAU, trnA-UGC, trnR-ACG, trnN-GUU trnL-UAG, trnI-CAU, trnL-CAA, trnV-GAC, trnI-GAU, trnA-UGC, trnR-ACG, trnN-GUU | |
| Ribosomal protein | rps16, rps2, rps14, rps4, rpl33, rps18, rpl20, rps12, rps11, rpl36, rps8, rpl14, rpl16, rps3, rpl22, rps19, rpl2, rpl32, rps7 rps15, rps12, rpl2, rpl32, rps7 | |
| Other genes | Maturase gene | matK |
| Subunit of acetyl-CoA carboxylase | accD | |
| Envelope membrane protein | cemA | |
| cytochrome c heme attachment protein | ccsA | |
| ATP-dependent Clp protease proteolytic subunit | clpP | |
| Unkown function | Translation initiation factor | infA |
| Conserved open reading frame | ycf3, ycf4, ycf2, ycf15 ycf1, ycf2, ycf15 |
Table 2.
Summary of Rubus chloroplast genome features
| Species | Accession number | Genome size(bp) | AT% | GC% | ||||
|---|---|---|---|---|---|---|---|---|
| LSC | IRb | SSC | IRa | Total | ||||
| Rubus coreanus | MT478114.1 | 85,029 | 25,993 | 18,770 | 25,993 | 155,785 | 62.73 | 37.27 |
| Rubus pinfaensis | MZ352081.1 | 85,211 | 25,797 | 18,718 | 25,797 | 155,523 | 62.87 | 37.13 |
| Rubus takesimensis | NC 037991.1 | 85,430 | 25,781 | 18,768 | 25,781 | 155,760 | 62.90 | 37.10 |
| Rubus crataegifolius | NC 039704.1 | 85,402 | 25,781 | 18,750 | 25,781 | 155,714 | 62.89 | 37.11 |
| Rubus boninensis | NC 046015.1 | 85,438 | 25,793 | 18,783 | 25,793 | 155,807 | 62.93 | 37.07 |
| Rubus trifidus | NC 046585.1 | 85,466 | 25,799 | 18,759 | 25,799 | 155,823 | 62.91 | 37.09 |
| Rubus niveus | NC 056930.1 | 85,175 | 25,996 | 18,779 | 25,996 | 155,946 | 62.74 | 37.26 |
| Rubus lambertianus | NC 056931.1 | 85,879 | 25,782 | 18,873 | 25,782 | 156,316 | 62.84 | 37.16 |
| Rubus columellaris | NC 056932.1 | 85,160 | 25,796 | 18,718 | 25,796 | 155,470 | 62.86 | 37.14 |
| Rubus peltatus | NC 056937.1 | 85,329 | 25,737 | 18,779 | 25,737 | 155,582 | 63.09 | 36.91 |
| Rubus crassifolius | NC 056941.1 | 85,834 | 25,777 | 18,853 | 25,777 | 156,241 | 62.83 | 37.17 |
| Rubus setchuenensis | NC 056946.1 | 85,879 | 25,771 | 18,874 | 25,771 | 156,295 | 62.83 | 37.17 |
| Rubus sachalinensis | NC 056965.1 | 85,077 | 25,996 | 18,718 | 25,996 | 155,787 | 62.76 | 37.24 |
| Rubus glandulosopunctatus | NC 057624.1 | 85,526 | 25,749 | 18,718 | 25,749 | 155,742 | 63.02 | 36.98 |
| Rubus rubroangustifolius | NC 057629.1 | 85,371 | 25,749 | 18,697 | 25,749 | 155,566 | 62.99 | 37.01 |
| Rubus taiwanicola | NC 057631.1 | 85,408 | 25,742 | 18,724 | 25,742 | 155,616 | 63.02 | 36.98 |
| Rubus parvifolius | NC 060617.1 | 85,125 | 26,016 | 18,749 | 26,016 | 155,906 | 62.72 | 37.28 |
| Rubus occidentalis | NC 060646.1 | 85,889 | 25,751 | 18,862 | 25,751 | 156,253 | 62.82 | 37.18 |
| Rubus conduplicatus | NC 064143.1 | 85,432 | 25,788 | 18,723 | 25,788 | 155,731 | 62.91 | 37.09 |
| Rubus chingii | PQ 510,071 | 85,331 | 25,760 | 18,704 | 25,760 | 155,555 | 62.96 | 37.04 |
| Rubus rosifolius | PP566893.1 | 85,758 | 25,763 | 18,722 | 25,763 | 156,006 | 62.98 | 37.02 |
| Rubus irenaeus | PP566888.1 | 85,814 | 25,771 | 18,876 | 25,771 | 156,232 | 62.82 | 37.18 |
| Rubus fockeanus | PP566886.1 | 86,130 | 25,742 | 18,815 | 25,742 | 156,429 | 62.88 | 37.12 |
| Rubus calycacanthus | PP566882.1 | 85,883 | 25,782 | 18,874 | 25,782 | 156,321 | 62.84 | 37.16 |
| Rubus idaeus | OR698909.1 | 85,036 | 25,967 | 18,717 | 25,967 | 155,687 | 62.75 | 37.25 |
| Rubus corchorifolius | OP581010.1 | 85,298 | 25,760 | 18,700 | 25,760 | 155,518 | 62.96 | 37.04 |
| Rubus innominatus | OK127883.1 | 85,095 | 25,993 | 18,793 | 25,993 | 155,874 | 62.76 | 37.24 |
| Rubus hirsutuss | OK127882.1 | 85,784 | 25,763 | 18,710 | 25,763 | 156,020 | 62.99 | 37.01 |
Comparative analysis of 27 Rubus species revealed interspecific variations in chloroplast genome sizes (155,470 − 156,429 bp), potentially attribuTab. to evolutionary dynamics in inverted repeat (IR) region expansion/contraction (Table 2). The total GC content exhibited narrow variation (36.91–42.75%) with no statistically significant differences among species. This conserved GC profile, particularly in coding regions where sequence composition typically reflects phylogenetic relationships, further corroborates the high genomic synteny observed across the genus.
Analysis of RSCU
We characterized the codon usage patterns of protein-coding genes across chloroplast genomes from 28 Rubus species (Fig. 2). A total of 64 codons corresponding to 20 amino acids and three termination codons were identified in all sequenced genomes. UUA exhibited the highest usage frequency, with its relative synonymous codon usage (RSCU) values ranging from 1.95 to 2.05. Conversely, among the six synonymous codons for leucine, CUC and CUG showed the lowest usage frequencies, with RSCU values ranging from 0.34 to 0.41 and 0.35 to 0.39, respectively. Among three stop codons, UAA showed predominant usage, with RSCU values fluctuating between 1.29 and 1.80. Notably, the RSCU value of tryptophan-specific codon UGG was consistently maintained at 1.0 across all accessions. This phenomenon is attributed to its monodegenerate property without synonymous codon alternatives, which is consistent with previous findings reported in other plant chloroplast genomic studies.
Fig. 2.

Comparison of Relative Synonymous Codon Usage (RSCU) patterns across the 28 Rubus chloroplast genomes
Analysis of IR contraction
When we characterized key structural features of Rubus chloroplast genomes, we identified an approximate 959 bp sequence gap across the chloroplast genomes. Meanwhile, analysis the boundary maps of single copy (SC) and inverted repeat (IR) region across the 28 studied species confirmed that their cp. genomes conform to the four part structure common to most plants ( LSC, SSC, and two IR regions that separate them), where differences in genome size are primarily attributable to expansion and contraction dynamics of the IR and SSC regions [23]. In our paper, we examined chloroplast genomes of 28 Rubus species to assess IR boundary expansion and contraction (Fig. 3). The results showed that, at the IRb-LSC boundary, the rps19 gene of all 28 species was located in the LSC region with a 14–25 bp extension to IRb, and rpl2 gene was present at IRb region. At the IRa-SSC boundary, the ycf1 gene was present in all 28 Rubus species, with substantial variation in its position relative to the boundary. In 27 species, the ycf1 gene spaned the IRa-SSC boundary, extending into the IRa region. Notably, in Rubus coreanus, the ycf1 gene was entirely located within the SSC region, with a 193 bp distance from the SSC-IRa boundary.The IRa-LSC boundary closely resembled the LSC-IRb boundary in structure, with the trnH gene nearest the junction (0–2 bp from the boundary).
Fig. 3.

Comparison of the LSC, SSC, and IR boundary regions across the 28 Rubus chloroplast genomes
Repeat sequence analysis
A total of 1483 tandem repeats and 616 dispersed repeats were identified across the chloroplast genomes of the 28 Rubus species (Fig. 4). The identified dispersed repeats consisted of 616 inverted repeats accounting for 41.54%, 448 palindromic repeats for 30.21%, 224 forward repeats for 15.10%, and 112 complement repeats for 7.55% of the total dispersed repeats. Furthermore, the genomic distribution patterns of these repeat sequences exhibited limited interspecific variations in copy number and structural diversity.
Fig. 4.

Analyses of repeat sequences in 28 Rubus chloroplast genomes
Simple sequence repeats (SSRs), widely utilized as molecular markers, consist of tandemly repeated 1–6 nucleotide motifs [24, 25], and their evolutionary significance in population genetics and phylogenetic studies has been well established [26]. A total of 42,385 SSRs were identified from the chloroplast genomes of 28 Rubus species (Fig. 5B; Table S2). Among these loci, mononucleotide repeats were the most abundant, accounting for 94.00% (39,840) of all detected SSRs, with A/T repeats far outnumbering G/C repeats. Dinucleotide repeats ranked second, comprising 3.54% (1,500) of the total SSRs, followed by tetranucleotide repeats (1.53%, 648 loci) and trinucleotide repeats (0.77%, 325 loci). By contrast, pentanucleotide and hexanucleotide repeats were rare, with only 54 and 18 loci detected, respectively. Consistent with the universal characteristics of angiosperm chloroplast genomes [27, 28], the SSRs identified in Rubus chloroplast genomes exhibited obvious base preference and were overwhelmingly dominated by polyA and polyT motifs, which is a typical feature widely observed in numerous plant species [29, 30].
Fig. 5.

Analysis of simple sequence repeats (SSRs) in the Rubus chloroplast genomes. Note: A Genomic regional distribution of identified SSRs. The number and proportion of SSR loci distributed in intergenic spacer (IGS), coding gene, and intronic regions were statistically summarized. B Composition and classification of SSR motifs across Rubus species, showing the abundance variation of different repeat types
Genomic regional distribution analysis revealed distinct positional enrichment of these SSRs (Fig. 5A; Table S2). In detail, 25,226 SSRs (59.52%) were located in intergenic spacer (IGS) regions, 8,120 SSRs (19.16%) were distributed in intronic regions, and 9,039 SSRs (21.33%) were annotated in coding gene regions. The predominant enrichment of SSRs in IGS regions represents a conserved evolutionary pattern in Rosaceae species. Owing to weaker selective constraints in non-coding IGS regions, these genomic regions accumulate more sequence variations and polymorphic SSR loci compared with highly conserved coding and intronic regions.
Identification of highly variable regions
A comprehensive analysis of nucleotide polymorphisms across 28 chloroplast genomes was performed via DnaSP software, with simultaneous calculation of nucleotide diversity (Pi). This work uncovered 24,705 polymorphic sites, distributed as 19,271 in the LSC, 3127 in the IR, and 2307 in the SSC region. The SSC region displayed the greatest average Pi (0.997), whereas the LSC regions showed the lowest mean diversity (0.1; Tab. S3). This study sorted the pi values and screened highly variable sites, which are as follows: seven in the LSC (trnH-psbA, rps16-trnQ 1, rps16-trnQ 2, trnS-trnG, petN-psbM, petA-psbJ, psbE-petL ), and one in the SSC (rpl32-trnL), loci that hold promise as genus specific molecular markers (Fig. 6). Interesting, the LSC regions demonstrated the most pronounced nucleotide polymorphism among all genomic regions, whereas the IR regions displayed the lowest polymorphism, reflecting its evolutionary conservation in IR regions. The eight highly variable loci identified have the potential to be used as molecular markers for species identification in Rubus.
Fig. 6.

Sliding window analysis of the whole cp. genomes of 28 Rubus plants. Note: Window length: 100 bp, step size: 25 bp. X-axis, the position of the midpoint of a window; Y-axis, nucleotide diversity of each window; Eight highly variable loci are labeled, including seven in the LSC region (① trnH-psbA, ② rps16-trnQ 1, ③ rps16-trnQ 2, ④ trnS-trnG, ⑤ petN-psbM, ⑥ petA-psbJ, ⑦ psbE-petL) and one in the SSC region (⑧ rpl32-trnL)
Comparative architectural profiling of chloroplast genomes in Rubus
Using mVISTA, the chloroplast genomes of 28 Rubus species were compared by our study (Fig. 7), with R. chingii as the reference, the rusults revealing clear sequence variations. Overall, the genomes exhibited high conservation and the most coding genes showing minimal divergence. Analyses of the 28 Rubus chloroplast genomes indicated that major genomic variation was concentrated in the LSC and SSC regions, with noticeable sequence gaps occurring primarily in non-coding regions rather than protein-coding or rRNA loci. This discovery is consistent with other plant species.
Fig. 7.

Visualization of genome alignment of the complete chloroplast genome of 28 Rubus. Note: The cp. genome of R. chingii was used as reference. X-axis indicates the position of the sequence in the whole cp. genome. Y-axis represents the similarity of the aligned region (shown percent identity: 50–100%)
Comparative selective pressure profiling of chloroplast genomes in Rubus
Nucleotide substitution patterns of synonymous and nonsynonymous sites are crucial for uncovering the dynamic evolutionary characteristics of functional genes. Here, the Ka and Ks substitution rates of 60 shared protein-coding genes retrieved from the chloroplast genomes of 28 Rubus species were calculated to evaluate their selection regimes (Fig. 8). Our quantitative assessment showed that the majority of examined genes exhibited Ka/Ks ratios lower than 1, indicative of predominant purifying selection acting on these chloroplast coding sequences. Such strong selective constraint contributes to the functional conservation of core chloroplast genes across Rubus species. Additionally, variable Ka/Ks values were detected among different gene orthologs, implying distinct evolutionary rates and heterogeneous selection pressures acting on individual chloroplast genes.
Fig. 8.

Comparison of Ka/Ks ratios of protein-coding genes across 28 Rubus species
Phylogenetic analysis
To distinguish the phylogenetic position of R. chingii, a reference was made to the chloroplast sequences of 28 other Rubus plants already published by NCBI, and the whole genome sequences were selected for phylogenetic analysis. The results were as follows: the support rate for clustering was high, with a test score of 100% for 25 nodes, indicating a high reliability of the clustering results. As shown in Fig. 9, the 28 plants can be divided into two main categories. The first category includes 13 plant species, containing Rubus idaeus、Rubus sachalinensiss、Rubus niveus、Rubus innominatus、Rubus parvifolius、Rubus coreanus、Rubus occidentalis、Rubus fockeanus、Rubus irenaeus、Rubus setchuenensis、Rubus crassifolius、Rubus calycacanthus and Rubus lambertianus. The second clade encompasses 15 species: Rubus columellaris、Rubus pinfaensis、Rubus peltatus、Rubus trifidus、Rubus boninensis、Rubus conduplicatus、Rubus crataegifolius、Rubus takesimensis、Rubus chingii、Rubus corchorifolius、Rubus hirsutus、Rubus rosifolius、Rubus taiwanicola、Rubus rubroangustifolius and Rubus glandulosopunctatus. The clustering diagram demonstrated that Rubus chingii and Rubus corchorifolius possessed the nearest phylogenetic relationship within the genus.
Fig. 9.

Phylogenetic trees based on the whole-genome sequences of chloroplast genomes of 28 Rubus species
Screening and identification analysis of hypervariable IGS
Based on the sequence divergence analysis of hypervariable segments, three loci with the highest genetic variation were screened out as candidate verification regions, namely rpl32-trnL, rps16-trnQ 2 and trnS-trnG. Specific primers targeting these three fragments were subsequently designed. All candidate primer pairs were further evaluated and screened using Primer Premier 5 software, and three qualified primer combinations were finally obtained (Table 3).
Table 3.
Information of hypervariable IGS primers
| Primer | Primer sequence(5’to 3’) | Annealing temperature(℃) |
|---|---|---|
| rpl32-trnL F | TTTCTACCGGTAGCTCAAAAAG | 54 |
| rpl32-trnL R | ATTATCGAGTCAAGTCGATGGG | |
| rps16-trnQ 2 F | TTTTTTATCATCAGTAAGAGAGAGA | 50 |
| rps16-trnQ 2 R | GTCCCGCTATTCGGAGG | |
| trnS-trnG F | CCACTCAGCCATCTCTCCTC | 56 |
| trnS-trnG R | AGGTTTTTGACCTATACGTCTCG |
Phylogenetic analyses were performed by combining the newly acquired amplicon sequences with previously released homologous sequences, and the corresponding phylogenetic trees were established (Fig. 10). The topological results revealed that both rpl32-trnL and rps16-trnQ 2 fragments possessed sufficient discriminatory power to clearly distinguish four studied Rubus species, including Rubus chingii, Rubus corchorifolius, Rubus lambertianus and Rubus hirsutus. In contrast, the trnS-trnG region could only successfully separate Rubus lambertianus from the other congeners.
Fig. 10.

Phylogenetic trees reconstructed from sequences amplified by different IGS hypervariable region primers and published reference sequences. Note: A Tree constructed with rpl32-trnL amplified fragments; (B) Tree based on rps16-trnQ 2 amplified sequences; (C) Tree derived from trnS-trnG amplified products
Discuss
Genomic architecture and comparative analysis
The complete chloroplast genome of Rubus chingii was sequenced, assembled, and annotated in our paper. Subsequently, we integrated our experimental chloroplast genome datasets with those of 27 additional species downloaded from NCBI, laying the foundation for comparative genomic analysis. By comparing sequences, the results demonstrated that 28 chloroplast genomes exhibited a conserved quadripartite structure and demonstrated no significant variations in genomic size. GC content ranged here from 36.91% to 42.75%, marginally surpassing the 36.3% average in terrestrial plant chloroplast genomes [31]. Elevated GC levels are likely to strengthen chloroplast genome stability, as GC pairs (sustained by three hydrogen bonds, in contrast to two for AT pairs) confer enhanced robustness against thermal denaturation and ecological stress challenges [32].
After calculating the relative synonymous codon usage (RSCU) values of protein-coding genes from chloroplast genomes across 28 Rubus species, distinct codon usage bias was observed among these species. A total of 31 codons presented RSCU values higher than 1.0, among which thirty codons terminated with A or U, while only UUG ended with G. This distribution feature was consistent with the previous finding of high AT content in these plastomes.In addition, another 31 codons had RSCU values lower than 1.0. Within these low-frequency codons, merely UGA, CUA and AUA possessed A/U terminal bases, and the remaining twenty-eight codons ended with G/C bases. These results revealed that protein-coding genes in Rubus chloroplast genomes tended to prefer codons ending with A/U, which was in line with the conclusions reported in Epimedium species [33]. Such preference for AT-ending codons, especially in highly expressed genes, may be associated with environmental adaptation, and contribute to the improvement of translational efficiency and mRNA structural stability [34–36].
Reduction of IR regions caused variable cpDNA sizes in Rubus
Extensive prior work has corroborated that IR region expansion and contraction constitute key determinants of chloroplast genome size [37, 38]. For the 28 Rubus chloroplast genomes analyzed in this study, we detected evidence of contraction and expansion at the IR boundaries, where the most notable variations took place in the IR and Small Single Copy (SSC) regions. Within these two regions, multiple genes alternated at the boundaries, the biggest change among them is the ycf1 gene. Such boundary shifts are accompanied by diverse modifications affecting the length, position, and pseudogenization of the ycf1 gene [39]. Owing to its high sequence lability, ycf1 is frequently regarded as a promising molecular marker [40]. Although the ycf1 gene exists in both IR and LSC regions in 28 Rubus, genes located at their boundary junctions showed distinct degrees of positional deviation. This lays the foundation for subsequent screening of candidate genes in Rubus for molecular identification.
Importance of chloroplast molecular markers in population genetics
We analyzed three types of repeats, including long repeats, tandem repeats, and simple repeats. Twenty-eight Rubus samples contained all three types: simple repeats were the most abundant among them, followed by tandem repeats, while long repeats exhibited almost no diversity. Among simple repeats, (A/T) sequences were dominant. This pronounced A/T bias likely reflects the replication mechanism and structural stability of chloroplast genomes [41]. The prevalence of AT-enriched SSRs reflects the intrinsic compositional bias of chloroplast genomes, and indicates that slipped-strand mispairing tends to occur more frequently in AT-rich genomic regions. Furthermore, this base preference is also shaped by natural mutation bias and long-term selective pressure during evolutionary adaptation [42]. These variable SSR loci can serve as valuable molecular markers for subsequent species identification and population genetic analysis. Additionally, such preference for A/T-rich mononucleotide repeats may also be associated with the evolution of chloroplast genomes and their optimization for environmental adaptation [43]. In our analysis results, we observed that the abundance of repeat sequences decreased as their length increased. Studies have indicated that in regions with high A + T content, the pressure on GC content grows as A + T content rises, this leads to a predisposition for GC base pair loss within these genomic [44].
In addition to abundant mononucleotide repeats, multiple types of polynucleotide SSR motifs were also identified across the chloroplast genomes of Rubus species. Several highly polymorphic SSR loci with promising discriminatory potential were screened in this study. Specifically, variable intergenic SSR motifs were detected in IGS regions, including the TCCTAA repeat located in the trnS-rps4 spacer of Rubus trifidus and the GTATAA motif within the trnT-trnL intergenic sequence of Rubus peltatus. Furthermore, unique genic SSRs were also discovered in coding regions, such as the ATTCCC motif embedded in the hypervariable ycf1 gene of Rubus innominatus. These polymorphic chloroplast SSR loci represent robust molecular candidates for DNA barcoding and fine-scale phylogeographic investigations, laying a solid theoretical foundation for species delimitation and evolutionary analysis of Rubus species.
Based on the analysis of nucleotide diversity, eight genomic segments with relatively high Pi values were detected among the 28 Rubus accessions, corresponding to the intergenic regions trnH-psbA, rps16-trnQ 1, rps16-trnQ 2, trnS-trnG, petN-psbM, petA-psbJ, psbE-petL, and rpl32-trnL. These regions, which show high levels of sequence variation, represent potential candidate molecular markers that may be applicable for future population genetic studies in Rubus. Furthermore, these variable regions could provide preliminary insights for the development of Rubus specific DNA barcodes, and may have potential utility in the authentication of medicinal Rubus chingii.
Our evolutionary selection pressure analysis revealed that the majority of chloroplast protein-coding genes across the examined Rubus species exhibited Ka/Ks ratios < 1, indicative of pervasive purifying selection. Nevertheless, a subset of functional genes were identified with Ka/Ks ratios ≥ 1, suggesting the occurrence of positive selection. These adaptive genes primarily included ccsA, petB, rps4, matK, ycf1, and ycf2. The differential elevation in Ka/Ks ratios implied heterogeneous selective constraints acting on distinct chloroplast genes, which were presumably linked to their functional roles in facilitating species-specific adaptive differentiation within Rubus. Functional annotation demonstrated that ccsA modulates the maturation and synthesis of chloroplast cytochrome c proteins, whereas petB encodes the b6 subunit of the Cytb6f complex; both genes are integral components of the photosynthetic electron transport system and are indispensable for maintaining stable photosynthetic functions [45]. The rps4 gene contributes to the assembly of the chloroplast small ribosomal subunit and modulates the protein translation machinery in chloroplasts. The matK gene encodes maturase K, a core regulator responsible for intron splicing and post-transcriptional modification of chloroplast genes.Furthermore, ycf1 and ycf2 were conserved hypothetical chloroplast genes characterized by rapid evolutionary rates and abundant sequence variation, thus, they have the potential to serve as candidate molecular markers, which had been widely confirmed in previous botanical studies [46, 47].
Phylogenetic relationship analysis from complete chloroplast genome data
The chloroplast genome possesses the characteristics of abundant gene content, conserved structural organization, reduced evolutionary rates, and elevated copy counts. It has long been the primary focus of research in phylogeny and molecular evolutionary research [48]. Early studies on chloroplast genome phylogeny were depended on single-gene fragments [49]; however, single-gene fragments carry insufficient informative sites, leading to low bootstrap support for many branches [50]. With the progressive expansion of datasets, phylogenetic resolution and bootstrap support of phylogenetic analyses based on multi-gene concatenated sequences have been significantly enhanced [51], and this approach has been widely adopted [52]. Nevertheless, with the growing availability of complete chloroplast genome data for Rubus species, we consequently conducted NJ (Neighbor-Joining) tree analysis on the chloroplast genomes of 28 Rubus species. These 28 genomes were mainly clustered into two major clades: Rubus chingii was positioned in Clade B and exhibited the closest phylogenetic relationship with Rubus corchorifolius. This study’s outcomes establish a foundational basis for follow up studies on further classification, germplasm identification, and genetic relationship assessment of Rubus species. Furthermore, phylogenetic incongruence between plastid and nuclear gene sequences is a common phenomenon, because the topological relationships inferred from chloroplast data are not fully consistent with those derived from nuclear genes. Such discordance is mainly attributed to incomplete lineage sorting and frequent interspecific hybridization events. Given the maternal inheritance pattern of chloroplast genomes, chloroplast-based phylogeny can effectively avoid genotyping confusion caused by interspecific hybridization in nuclear gene analyses [53]. This study was outcomes establish a foundational basis for follow up studies on further classification, germplasm identification, and genetic relationship assessment of Rubus species.
Applicability of screened hypervariable IGS markers in species identification
Based on the sequence variation analysis, abundant highly divergent intergenic regions were detected across the 28 Rubus species examined. Conserved flanking sequences of these variable loci were retrieved, and three pairs of specific primers targeting rpl32-trnL, rps16-trnQ2 and trnS-trnG were successfully developed.Experimental amplification results revealed that the rpl32-trnL and rps16-trnQ2 fragments possessed reliable taxonomic resolution to distinguish the four investigated Rubus species, demonstrating great potential as effective molecular diagnostic markers for this genus. Although the trnS-trnG fragment failed to fully separate all tested species, it could distinctly identify Rubus lambertianus.Previous investigations on other plant taxa have demonstrated that chloroplast intergenic spacer regions generally harbor rich sequence variations, including rpl32-trnL, rps16-trnQ, petN-psbM and trnE-rpoB [54]. This can be explained by the high sequence variation of IGS regions, their rapid evolutionary rate and weak functional constraints contribute to their good performance in species identification [55].In summary, the rpl32-trnL and rps16-trnQ2 regions can clearly differentiate Rubus chingii, Rubus corchorifolius, Rubus lambertianus and Rubus hirsutus. Both molecular fragments are promising candidate barcodes for species identification and phylogenetic research of Rubus plants.
Conclusions
Alongside the rapid progression of high-throughput sequencing platforms in recent decades, an increasing number of researchers have turned their attention to chloroplast genome analysis. Herein, we performed comprehensive comparative chloroplast genomic profiling of 28 Rubus species, encompassing the newly generated sequencing data from the current work. Results from our analyses established that the chloroplast genomes of 28 Rubus species showed high grade stability in genomic length, gene content, gene structure, and gene arrangement. Angiosperms typically exhibit a dominant maternal plastid inheritance mode, and this reproductive trait is what maintains the conserved structural organization of their chloroplast genomes. Our genome wide comparative surveys demonstrated that plastid genome differentiation is concentrated predominantly in the LSC and SSC regions, a trend that aligns with prior insights from angiosperm chloroplast genomic studies. We additionally conducted targeted SSR analysis, and identified unique gene SSRs within coding regions, the ATTCCC motif located in the hypervariable ycf1 gene of Rubus innominatus. Building on this, We found that regions including trnH-psbA, rps16-trnQ 1, rps16-trnQ 2, trnS-trnG, petN-psbM, petA-psbJ, psbE-petL and rpl32-trnL exhibited relatively high sequence divergence. Further verification indicated that rpl32-trnL and rps16-trnQ2 possess moderate discriminatory capacity, and can serve as candidate molecular markers for Rubus species.
Supplementary Information
Acknowledgements
Not applicable.
Authors' contributions
Lixia Yuan and Furong Wang designed the study and revised the manuscript. Juan Wang, Tianling Lou, Hao Wu and Lishuang Wu performed the experiments. Yangjian Chen, Jingjing Li, , Yiyuan Luo and Bin Chen collected data and samples in the field.
Funding
Zhejiang Province Public Welfare Program (ZCLTGN24C0201); Natural Science Foundation of Ningbo (2024J181, 2024J251); School-enterprise Cooperation Project from Domestic Visiting Engineers of Colleges and Universities (FG20240052024); Ningbo youth science and technology innovation leading talent project (2024QL064); Zhejiang Provincial Public Welfare Program (LGN22H280009, LTGN23H280003).
Data availability
The complete chloroplast genome of *Rubus chingii* have been deposited in the NCBI repository, https://www.ncbi.nlm.nih.gov/nuccore/PQ510071.1/ . The remaining 27 chloroplast genome sequences used for comparative analysis are available in NCBI, with their corresponding accession numbers provided in Table 2 of this article.
Declarations
Ethics approval and consent to participate
Not applicable. No specific permits were necessary for the collection of specimens of this study.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Li JJ, Du LF, He Y, et al. Chemical constituents and biological activities of plants from the genus Rubus. Chem Biodivers. 2015;12:1809–47. 10.1002/cbdv.201400307. [DOI] [PubMed] [Google Scholar]
- 2.Wang J, Zhang X, Yu J, et al. Constituents of the fruits of Rubus chingii Hu and their neuroprotective effects on human neuroblastoma SH-SY5Y cells. Food Res Int. 2023;173(Pt 1):113255. 10.1016/j.foodres.2023.113255. [DOI] [PubMed] [Google Scholar]
- 3.Yu GH, Luo ZQ, Wang WB, et al. Rubus chingii Hu: A Review of the Phytochemistry and Pharmacology. Front Pharmacol. 2019;10:799. 10.3389/fphar.2019.00799. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Hua YJ, Xie F, Mao KJ, et al. Insights into the metabolite profiles of Rubus chingii Hu at different developmental stages of fruit. J Sep Sci. 2023;46(16):e2300264. 10.1002/jssc.202300264. [DOI] [PubMed] [Google Scholar]
- 5.Ke H, Bao T, Chen W. New function of polysaccharide from Rubus chingii Hu: protective effect against ethyl carbamate induced cytotoxicity. J Sci Food Agric. 2021;101(8):3156–64. 10.1002/jsfa.10944. [DOI] [PubMed] [Google Scholar]
- 6.Ke H, Bao T, Chen W. Polysaccharide from Rubus chingii Hu afford protection against palmitic acid-induced lipotoxicity in human hepatocytes. Int J Biol Macromol. 2019;133:1063–71. [DOI] [PubMed] [Google Scholar]
- 7.Zhang TT, Lu CL, Jiang JG et al. Bioactivities and extraction optimization of crude polysaccharides from the fruits and leaves of Rubus chingii Hu. Carbohydr Polym. 2015, 130:307 – 15. 10.1016/j.carbpol.2015.05.012. Epub 2015 May 18. PMID: 26076631. [DOI] [PubMed]
- 8.Sochor M, VaŠUt RJ, Sharbel TF, et al. How just a few makes a lot: Speciation via reticulation and apomixis on example of European brambles (Rubus subgen. Rubus, Rosaceae). Mol Phylogenetics Evol. 2015;89:13–27. 10.1016/j.ympev. [DOI] [PubMed] [Google Scholar]
- 9.Katoh K, Standley DM. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 2013;30(4):772–80. 10.1093/molbev/mst010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Jin JJ, Yu WB, Yang JB et al. GetOrganelle: a fast and versatile toolkit for accurate de novo assembly of organelle genomes. Genome Biol. 2020; 21(1): 241. dio: 10.1186/s13059-020-02154-5. [DOI] [PMC free article] [PubMed]
- 11.Shi L, Chen H, Jiang M, et al. CPGAVAS2, an integrated plastome sequence annotator and analyzer. Nucleic Acids Res. 2019;47. 10.1093/nar/gkz345. [DOI] [PMC free article] [PubMed]
- 12.Lewis SE, Searle SM, Harris N, et al. Apollo: a sequence annotation editor. Genome Biol. 2002;3(12). 10.1186/gb-2002-3-12-research0082. [DOI] [PMC free article] [PubMed]
- 13.Hejazi FA, Mohammadi P, Soorni A. Comparative chloroplast genomics of Teucrium species reveals genome evolution, phylogenetic relationships, and candidate molecular markers. Sci Rep. 2025;15:44318–31. 10.1038/s41598-025-29339-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Diani GS, Soorni A. Plastome diversity and phylogenomic analysis of Satureja (Lamiaceae): uncovering evolutionary patterns and diagnostic markers. Planta. 2025;262(6):149. 10.1007/s00425-025-04875-y. [DOI] [PubMed] [Google Scholar]
- 15.Akrami AM, Meratian ES, Soorni A. Decoding the chloroplast genomes of five Iranian Salvia species: insights into genomic structure, phylogenetic relationships, and molecular marker development. BMC Genomics. 2025;26(1):545–531. 10.1186/s12864-025-11729-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Soorni A, Golchini MM. Complete Chloroplast genome of Mentha aquatica reveals hypervariable regions and resolves phylogenetic position within the genus Mentha. Mol Biol Rep. 2025;52(1):677–92. 10.1007/s11033-025-10789-5. [DOI] [PubMed] [Google Scholar]
- 17.Kurtz S, Choudhuri JV, Ohlebusch E, et al. REPuter: the manifold applications of repeat analysis on a genomic scale. Nucleic Acids Res. 2001;29(22):4633–42. 10.1093/nar/29.22.4633. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Huang L, Yu H, Wang Z, et al. CPStools: A package for analyzing chloroplast genome sequences. IMetaOmics. 2024;1(2):e25. 10.1002/imo2.25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Amiryousefi A, Hyvönen J, Poczai P. IRscope: an online program to visualize the junction sites of chloroplast genomes. Bioinformatics. 2018;34(17):3030–1. 10.1093/bioinformatics/bty220. [DOI] [PubMed] [Google Scholar]
- 20.Frazer KA, Pachter L, Poliakov A, et al. VISTA: computational tools for comparative genomics. Nucleic Acids Res. 2004;1(7). 10.1093/nar/gkh458. [DOI] [PMC free article] [PubMed]
- 21.Mayor C, Brudno M, Schwartz JR, et al. VISTA: visualizing global DNA sequence alignments of arbitrary length. Bioinformatics. 2000;16(11):1046–7. 10.1093/bioinformatics/16.11.1046. [DOI] [PubMed] [Google Scholar]
- 22.Minh BQ, Schmidt HA, Chernomor O, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol. 2020;37(8):1530. 10.1093/molbev/msaa015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Danchun Z, Jiajun T, Xiaoxia D, et al. Analysis of the chloroplast genome and phylogenetic evolution of Bidens pilosa. BMC Genomics. 2023;24(1):113. 10.1186/s12864-023-09195-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Xiaxia L, Lijun Q, Birong C, et al. SSR markers development and their application in genetic diversity evaluation of garlic (Allium sativum) germplasm. Plant Divers. 2022;44(5):481–91. 10.1016/j.pld.2021.08.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Mengying Y, Yanan Y, Yaju G, et al. The complete chloroplast genome of Solanum sisymbriifolium (Solanaceae), the wild eggplant. Mitochondrial DNA Part B. 2022;7(5):886–8. 10.1080/23802359.2022.2077667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Nie X, Lv S, Zhang Y, et al. Complete chloroplast genome sequence of a major invasive species, crofton weed (Ageratina adenophora). PLoS ONE. 2012;7(5):e36869. 10.1371/journal.pone.0036869. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Yang J, Hu G, Hu G. Comparative genomics and phylogenetic relationships of two endemic and endangered species (Handeliodendron bodinieri and Eurycorymbus cavaleriei) of two monotypic genera within Sapindales. BMC Genomics. 2022;23(1):27. 10.1186/s12864-021-08259-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Xu K, Lin C, Lee SY, et al. Comparative analysis of complete Ilex (Aquifoliaceae) chloroplast genomes: insights into evolutionary dynamics and phylogenetic relationships. BMC Genomics. 2022;23(1):203. 10.1186/s12864-022-08397-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Ren T, Xun L, Jia Y, et al. Insights into the plastome evolution and phylogenetic relationships of Lepidium (Brassicaceae). BMC Plant Biol. 2025;25(1):1489. 10.1186/s12870-025-07524-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Du Y, Luo Y, Wang Y, et al. Comparative Analysis of the Complete Chloroplast Genomes of Eight Salvia Medicinal Species: Insights into the Deep Phylogeny of Salvia in East Asia. Curr Issues Mol Biol. 2025;47(7):493. 10.3390/cimb47070493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Guo YY, Yang JX, Li HK, et al. Chloroplast genomes of two species of Cypripedium: Expanded genome size and proliferation of AT-biased repeat sequences. Front Plant Sci. 2021;12:609729. 10.3389/fpls.2021.609729. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Gao R, Wang W, Huang Q, et al. Complete chloroplast genome sequence of Dryopteris fragrans (L.) Schott and the repeat structures against the thermal environment. Sci Rep. 2018;8:166351–62. 10.1038/s41598-018-35007-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Wang Y, Jiang D, Guo K, et al. Comparative analysis of codon usage patterns in chloroplast genomes of ten Epimedium species. BMC Genom Data. 2023;24(1):3. 10.1186/s12863-023-01104-x . PMID: 36624369; PMCID: PMC9830715. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Zhang Y, Ma Y, Gao J, et al. Comparative analysis of codon usage patterns in Chloroplast genomes of Maple (Genus Acer). Biochem genet. 2025 10.1007/s10528-025-11292-z . [DOI] [PubMed] [Google Scholar]
- 35.Xiao M, Hu X, Li Y, et al. Comparative analysis of codon usage patterns in the chloroplast genomes of nine forage legumes. Physiol Mol Biol Plants. 2024;30(2):153–66. 10.1007/s12298-024-01421-0 . Epub 2024 Mar 9. PMID: 38623162; PMCID: PMC11016040. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Fang J, Hu Y, Hu Z. Comparative analysis of codon usage patterns in 16 chloroplast genomes of suborder Halimedineae. BMC Genomics. 2024;25(1):945. 10.1186/s12864-024-10825-x . PMID: 39379800; PMCID: PMC11459826. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Gong L, Ding X, Guan W, et al. Comparative chloroplast genome analyses of Amomum: insights into evolutionary history and species identification. BMC Plant Biol. 2022;22(1):520. 10.1186/s12870-022-03898-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Song Y, Zhao W, Xu J, et al. Chloroplast Genome Evolution and species Identification of Styrax (Styracaceae). Biomed Res Int. 2022;5364094. 10.1155/2022/5364094. [DOI] [PMC free article] [PubMed]
- 39.Li L, Yang W, Liu Q, et al. Comparative analysis of 18 chloroplast genomes reveals genomic diversity and evolutionary dynamics in subtribe Malaxidinae (Orchidaceae). BMC Plant Biol. 2025;25(1):1013. 10.1186/s12870-025-06772-8 . PMID: 40753185; PMCID: PMC12317619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Yang M, Wang L, Xin G. Anatomy and the chloroplast genome of Bischofia plants reveal important phylogenetic relationship and the genetic diversity in phyllanthaceae. Sci Rep. 2026. 10.1038/s41598-026-52125-2. Epub ahead of print. PMID: 42128939. [DOI] [PMC free article] [PubMed]
- 41.Zhang D, Tu J, Ding X, et al. Analysis of the chloroplast genome and phylogenetic evolution of Bidens pilosa. BMC Genomics. 2023;24(1):113. 10.1186/s12864-023-09195-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Yan R, Abdullah, Jia J, et al. Characterization of Verbesina encelioides (Asteroideae, Asteraceae) Chloroplast Genome and Phylogenetic Insights. Ecol Evol. 2026;16(1):e72897. 10.1002/ece3.72897 . PMID: 41488798; PMCID: PMC12758953. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Jiang D, Cai X, Gong M, et al. Complete chloroplast genomes provide insights into evolution and phylogeny of Zingiber (Zingiberaceae). BMC Genomics. 2023;24(1):397. 10.1186/s12864-023-09393-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Morton BR, Bi IV, McMullen MD, et al. Variation in mutation dynamics across the maize genome as a function of regional and flanking base composition. Genetics. 2006;172(1):569–77. 10.1534/genetics.105.049916. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Jadhao KR, Dey P. Phylogenetic Analysis of strawberry, Fragaria x ananassa (Rosaceae: Fragaria) using cytochrome C heme attachment protein (ccsA) gene[J].Research Biotica, 2021. 10.54083/resbio/3.3.2021.145-153. [DOI]
- 46.Dong W, Xu C, Li C, et al. ycf1, the most promising plastid DNA barcode of land plants. Sci Rep. 2015;5:8348. 10.1038/srep08348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Jin LH, Gui LS, Da MZ. Molecular evolution and phylogeny of the angiosperm ycf2 gene[J]. J Syst Evol. 2010;48(4):240–8. 10.1111/j.1759-6831.2010. 00080.x. [DOI] [Google Scholar]
- 48.Lv Z, Yang Y, Hou H, et al. Genetic diversity analysis of proso millet (Panicum miliaceum L.) germplasm resources based on phenotypic traits and SSR markers. Front Plant Sci. 2025;16:1649200. 10.3389/fpls.2025.1649200. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Song YQ, Kang KL, Chen J, et al. Molecular Structure, Comparative Analysis, and Phylogenetic Insights into the Complete Chloroplast Genomes ofFissidens crispulus. Genes (Basel). 2025;16(9):1103. 10.3390/genes16091103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Shi S, Zhang Z, Lin X, et al. Comparative and Phylogenetic Analysis of the Complete Chloroplast Genomes of Lithocarpus Species (Fagaceae) in South China. Genes (Basel). 2025;16(6):616. 10.3390/genes16060616. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Rokas A, Williams BL, King N, et al. Genome-scale approaches to resolving incongruence in molecular phylogenies. Nature. 2003;425(6960):798–804. 10.1038/nature02053. [DOI] [PubMed] [Google Scholar]
- 52.Qiu YL, Lee J, Bernasconi-Quadroni F, et al. The earliest angiosperms: evidence from mitochondrial, plastid and nuclear genomes. Nature. 1999;402(6760):404–7. 10.1038/46536. [DOI] [PubMed] [Google Scholar]
- 53.Gao XF, Xiong XH, Boufford DE et al. Phylogeny of the Diploid Species of Rubus (Rosaceae). Genes (Basel). 2023, 14(6):1152. 10.3390/genes14061152. PMID: 37372332; PMCID: PMC10298701. [DOI] [PMC free article] [PubMed]
- 54.Chen C, Luo D, Wang Z, et al. Complete chloroplast genomes of eight Artemisia species: Comparative analysis, molecular identification, and phylogenetic analysis. Plant Biol J. 2024;26:257–69. 10.1111/plb.13608. [DOI] [PubMed] [Google Scholar]
- 55.Ghotbi FS, Soorni A. Chloroplast phylogenomics and barcode discovery in medicinal Stachys species reveal evolutionary relationships and adaptive signatures. Ecol Evol. 2026;16:e73618. 10.1002/ece3.73618 . PMID: 42083652; PMCID: PMC1313585 9. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The complete chloroplast genome of *Rubus chingii* have been deposited in the NCBI repository, https://www.ncbi.nlm.nih.gov/nuccore/PQ510071.1/ . The remaining 27 chloroplast genome sequences used for comparative analysis are available in NCBI, with their corresponding accession numbers provided in Table 2 of this article.
