Skip to main content
Genome Biology logoLink to Genome Biology
. 2026 Jul 26;27:237. doi: 10.1186/s13059-026-04179-8

The hidden sources of spurious fusion transcripts in plants

Xi-Tong Zhu 1,#, Xinyan Lu 2,#, Zengxin Zhang 2,#, Qian Tang 1,#, Mengting Liu 2, Fan Xia 2, Xiaoyu Zhang 2, Sanz-Jimenez Pablo 2, Run Zhou 2, Huan Li 2,3, Yidan Ouyang 2,✉, Ling-Ling Chen 1,4,✉
PMCID: PMC13401787  PMID: 42503510

Abstract

Background

Fusion transcripts, first characterized in cancer, have been increasingly reported in plants with the expansion of next-generation sequencing. However, their prevalence and biological relevance remain highly debated, particularly given the technical challenges associated with their detection.

Results

Here, by integrating multiple high-quality long-read RNA sequencing datasets from rice, we present a systematic assessment of fusion transcript detection in plants and demonstrate that almost all detected fusion transcripts arise from technical and analytical artifacts rather than genuine biological events. Mechanistically, we identify short homologous sequence mediated template switching during reverse transcription as the predominant source of spurious fusions, especially in PCR-based workflows. Additional contributors include misalignment, reference genome bias, and gene misannotation. We further uncover recurrent artifact hotspots that explain the non-random distribution of fusion signals. Through redesigned in vitro and in vivo validation experiments, we demonstrate that commonly detected fusion signals lack reproducibility and do not reflect true transcriptomic events. Importantly, we establish a gold-standard validation pipeline prioritizing long-read direct RNA data, reference-aware mapping, and rigorous experimental validation to establish new reproducibility criteria for identifying authentic fusion transcripts.

Conclusions

Our study provides a comprehensive, plant-focused experimental dissection of fusion transcript artifacts across sequencing platforms. These findings challenge prevailing assumptions about the abundance of fusion transcripts in plants and establish a robust framework for their reliable identification, with broad implications for transcriptomics studies in complex genomes.

Graphical Abstract

graphic file with name 13059_2026_4179_Figa_HTML.jpg

Supplementary Information

The online version contains supplementary material available at 10.1186/s13059-026-04179-8.

Keywords: Fusion transcripts, False-positive fusions, Short homologous sequence, Long-read RNA sequencing, Direct RNA sequencing, Technical artifacts, Plant transcriptomics

Background

Fusion transcripts, also referred to as chimeric RNAs, contain sequences derived from two distinct genes. They can arise through two conceptually different routes: DNA-level fusions, in which separate genes become physically joined in the genome through structural rearrangements such as inversion or translocation, and RNA-level fusions, in which genes remain separate in the genome but their transcripts are joined through processes such as trans-splicing or readthrough transcription (Fig. 1a) [1, 2], Initially discovered in cancer and often associated with tumorigenesis [3], fusion transcripts have also been reported in plants largely through short-read RNA sequencing data (RNAseq) [4–7]. However, we observe several puzzling trends that are inconsistent or difficult to reconcile. First, the reported number of fusion transcripts varies dramatically across species, from 82,969 in A. thaliana [5] to fewer than 1,500 in Zea mays [7], and even more evident for the same species (e.g., 201 to 82,969 in A. thaliana) [4, 6]. This enthusiasm led to the deposition of nearly 300,000 putative fusion transcripts in databases like AtFusionDB [5] and PFusionDB [8]. Second, many fusions involve neighboring genes, which are more plausibly attributed to gene model mis-annotation [9–11]. Third, validation strategies have often relied on PCR-based techniques without critical controls such as equimolar cDNA mixtures of fusion partners, obscuring the experimental artifact [12]. Finally, DNA- and RNA-level fusions are seldom distinguished in most existing studies, contributing to persistent conceptual confusion in the field.

Fig. 1.

Fig. 1

Origins and characteristics of potential fusion transcripts in rice using three long-read RNA sequencing protocols. a Two major categories of fusion transcripts based on their origin. DNA-level fusions arise when two separate genes become physically linked in the genome through structural rearrangements such as deletion, translocation, inversion, or duplication, typically producing fusion genes. RNA-level fusions occur without underlying genomic rearrangements; instead, genes remain separate in the genome and their transcripts are joined during or after transcription through mechanisms such as trans-splicing between distinct genes, and can be further illustrated according to the relationship between partner genes, including sense/antisense fusions, allelic fusions, and non-allelic fusions (intra- or inter-chromosomal). b Classification of fusion transcripts based on supported reads and breakpoint positions. Fusion transcripts were classified into four major categories: High-Confidence (HC), Low-Confidence (LC) and adjacent fusion transcripts (Potential readthrough, PRT). Within the HC group, further differentiation is made based on whether the fusion breakpoint lies at an exon boundary (HC-a) or not (HC-b). c Number of different types of fusion transcripts identified using three long-read protocols, PacBio IsoSeq, Nanopore direct RNA sequencing (dRNAseq), and Nanopore PCR-amplified cDNA sequencing (cDNAseq). d Number and proportion of inter- and intra-chromosomal fusions, based on the chromosomal locations of their partner genes. e Density plot comparing expression levels of partner genes involved in fusion transcripts versus randomly selected gene pairs. Dotted lines indicate the mean expression values of fusion-associated and control gene pairs. P-value was calculated using a permutation test (N = 1,000). f Overlap between identified fusion partners and interaction datasets. Five types of chromatin or transcriptomic interaction datasets from MH63 were used, including DNA-DNA interactions marked by H3K4me3 and RNAPII, DNA-RNA interactions marked by H3K4me3 and H3K9ac, and RNA-RNA interactions

Plant genomes exhibit remarkable plasticity compared to animals, with high levels of intraspecific variation driven by frequent polyploidization, transposable element activity, and structural diversity [13, 14]. The advent of third-generation long-read RNA sequencing technologies provides opportunities to assess fusion transcript authenticity and complexity in plants [15, 16], yet systematic evaluation remains lacking. Here, we comprehensively catalogue, quantify, and mechanistically dissect fusions across multiple sequencing platforms including PacBio IsoSeq (IsoSeq), Nanopore direct RNA sequencing (dRNAseq), Nanopore PCR-amplified cDNA sequencing (cDNAseq), and short-read RNAseq, multiple species including rice and A. thaliana, and multiple rice genotypes. Strikingly, fusion candidates showed minimal overlap across platforms and replicates, with recurrent hotspot loci driven by short homologous sequence (SHS) or misalignments, while annotation- and genome-aware comparisons revealed that many fusions arise from annotation errors or reference biases. By revealing the hidden complexity of spurious fusions, we provide a framework to reassess their biological relevance and establishes benchmarks for rigorous validation of authentic plant fusion transcripts.

Results

The prevalence of spurious fusion transcripts across sequencing platforms

To rigorously investigate fusion transcripts and their origins, we generated high-quality long-read sequencing datasets from rice cultivar Minghui 63 (MH63, Oryza sativa ssp. xian/indica) using three complementary platforms: IsoSeq, dRNAseq, and cDNAseq. These datasets showed high reproducibility, with an average inter-sample correlation of 0.86, and 77.49% of genes were consistently detected across all three platforms, establishing a robust foundation for reliable fusion transcripts analysis (Additional file 1: Fig. S1; Additional file 2: Table S1).

We identified 725, 146 and 4,164 fusion transcripts in IsoSeq, dRNAseq, and cDNAseq, respectively (Additional file 2: Table S2), with partner genes (genes involved in fusion transcripts formation) exhibiting no chromosomal positional bias (R = 0.98, compared with overall genomic gene distribution, P = 7.06e-09) (Additional file 1: Fig. S2). These fusions were categorized into High-Confidence (HC, supported reads > 1), Low-Confidence (LC, supported reads = 1), and Potential readthrough (PRT, adjacent gene fusions), with HC further subdivided into HC-a (exonic boundaries breakpoints) and HC-b (non-exonic boundary breakpoints) (Fig. 1b, c). Surprisingly, only two HC fusions were consistently detected between dRNAseq and cDNAseq datasets. Reproducibility analysis showed limited overlap in dRNAseq (1/146) and cDNAseq (19/4164) biological replicates, irrespective of the tools used (JAFFAL [17] or LongGF [18]), suggesting that the observed challenges in the reproducibility are not primarily caused by algorithmic differences, but instead intrinsic to the data itself (Additional file 1: Fig. S3).

Physical proximity in nuclear 3D space has been hypothesized as a prerequisite for fusion transcript formation, particularly those generated from trans-splicing [2, 19]. Despite a high proportion of identified fusion transcripts were derived from inter-chromosome genes (71.92%−91.45%; Fig. 1d) and exhibited elevated expression (Fig. 1e), only 102 fusions overlapped with millions of DNA–DNA (H3K4me3, RNAPII) [20], DNA–RNA (H3K4me3, H3K9ac) [21, 22], and RNA–RNA [22] genomic or transcriptomic interaction datasets (48 from IsoSeq, 11 from dRNAseq, and 43 from cDNAseq) (Fig. 1f; Additional file 2: Table S3). This minimal overlap suggests that nuclear 3D spatial proximity is not a dominant mechanism for fusion formation.

To assess the generalizability of our findings beyond MH63, we extended our analysis to publicly available dRNAseq datasets from six tissues of Oryza sativa ssp. geng/japonica cv. Nipponbare (NIP) [23] and two tissues of A. thaliana [11]. Although the number of fusion transcripts varied across tissues and species (9 to 631), only a small fraction (0% to 2.42%) of these fusions was reproducible across biological replicates (Additional file 1: Fig. S4), reinforcing our earlier observations in MH63 that most fusions lack consistent support and may represent technical noise.

Short homologous sequence-mediated template switching as a major source of false-positive fusion transcripts

Fusion breakpoints were categorized into four types: SHS (breakpoints covered by short homologous sequence), Joined (directly linked), Unknown (linked by non-partner sequences), and Misaligned (failure to realign sequences to fusion reads using alternative aligners) (Fig. 2a; Methods). Analysis revealed that SHS-mediated events accounted for 75.43% of all fusion transcripts (Fig. 2b), with particularly high prevalence in PCR-based methods (IsoSeq 75.45%; cDNAseq 77.59%). These SHS sequences displayed a strong length bias (91.65% ≥ 4bp) but similar structural motifs to random sequences (Divergence: 0.023, Pmin = 0.25). Notably, SHS-mediated fusion partners exhibited significant tissue-specific expression (P < 0.001), suggesting that their detection was also tissue-specific (Additional file 1: Fig. S5; Additional file 2: Table S4).

Fig. 2.

Fig. 2

Experimental validation of short homologous sequence (SHS)-mediated fusion artifacts in rice. a Classification of fusion transcript junctions based on reciprocal alignment analysis. Four types of junctions were identified, including SHS, Joined (directly linked), Unknown (concatenated with bases not presenting in both partner genes), and Misaligned (either partner locus failed to align to the breakpoints on fusion reads). See Methods for details. b Distribution of fusion breakpoints across the four categories (SHS, Unknown, Misaligned, and Joined). c Example of a false-positive SHS-mediated fusion (MH63-SHS1). Tracks illustrate the partner genes, fusion reads, breakpoints, and the 12bp SHS. d Schematic of the in vitro validation strategy for MH63-SHS1. Primer sets were designed to amplify partner genes with or without the 12bp SHS. Amplified products were mixed to generate SHS-included and SHS-excluded groups, alongside the MH63 cDNA control. Fusion-specific primers were used for final PCR. e Gel electrophoresis of MH63-SHS1 validation. Expected PCR products (labeled above lanes) confirm SHS-dependency of MH63-SHS1. f Validation of MH63-SHS1 using different reverse transcriptases (AMV and M-MLV). DNA marker (left) indicates fragment sizes. g Fluorescent reporter constructs for MH63-SHS1 partner genes. Gene A (OsMH_01G0028500) was fused to N-terminal GFP, and Gene B (OsMH_02G0018200) to C-terminal mCherry. These constructs were then transformed into rice protoplasts. h Flow cytometry analysis and qRT-PCR of transformed protoplasts. Four cell populations were sorted: untransformed (control), GFP-only, mCherry-only, and co-transformed (GFP + mCherry). Left: Flow cytometry gating; right: quantification of fluorescence markers. Scale bar: 10 μm. i In vivo protoplast assay for MH63-SHS1 detection. Top: specific primers targeting the fluorescent reporter sequence to detect transfected fusions. Bottom: PCR amplification of MH63-SHS1 in different cells using AMV or M-MLV RT. Marker: Trans2K Plus DNA Marker (TransGen Biotech, BM111). Error bars: ± S.E.M

To directly test whether SHS was required for the generation of fusion signals and whether these signals arose during reverse transcription, we selected a fusion named MH63-SHS1 (OsMH_01G0028500#OsMH_02G0018200) with 12bp SHS (Fig. 2c) and constructed two types of artificial cDNA mixtures alongside the MH63 cDNA native control: SHS-included group in which PCR-amplified cDNA of the two fusion partners retained the SHS and SHS-removed group in which the SHS was excluded, both mixed at equal molar ratios (Fig. 2d). Fusion products were amplified only in the SHS-included group and the native MH63 cDNA, but not in the SHS-removed group, demonstrating that SHS was necessary for generating MH63-SHS1 in vitro (Fig. 2e). Furthermore, using different reverse transcriptase showed that fusion signals were present with avian myeloblastosis virus (AMV) but abolished with high-purity moloney murine leukemia virus (M-MLV) reverse transcriptase, indicating that SHS-dependent fusion arose from reverse-transcription–mediated template switching rather than representing genuine biological fusion events (Fig. 2f). The reduced RNase H activity of M-MLV may minimize RNA template degradation, thereby decreasing the occurrence of template switching events caused by fragmented templates.

We next established an in vivo rice protoplast system to test whether SHS-mediated fusion signals can be observed in a physiological cellular context in which genuine fusions were not expected. Two partners of MH63-SHS1 were independently introduced into protoplasts via single or co-transformation, with one partner fused to an N-terminal GFP and another to a C-terminal mCherry tag (Fig. 2g). Following fluorescence-activated cell sorting, untransformed, single-transformed (expressing GFP or mCherry), and co-transformed cells were collected (Fig. 2h). Using specific primers targeting the fluorescent reporter sequence, we explicitly traced fusion signals from exogenous partners and found that fusion-specific PCR products were detected in both co-transformed cells and single-transformed mixtures when AMV reverse transcriptase was used (Fig. 2i). Given that fusion detection was restricted to transfected transcripts and no genuine fusion events should occur in the single-transformed mixtures, our findings indicate that identifying a fusion transcript using PCR-based methods does not necessarily confirm its formation. Collectively, by decoupling fusion detection from fusion existence, this complementary in vivo validation demonstrates that SHS-mediated fusion signals can emerge even without true fusions, indicating they are technical artifacts inherent to the detection process.

Complex origin of false-positive non-SHS-mediated fusion transcripts

To further explore mechanisms behind non-SHS-mediated fusions, we systematically analyzed the characteristics of three remaining breakpoint types: Unknown, Joined, and Misaligned. Unknown-type fusions showed the highest expression of partner genes (Median: 6.12, P < 0.0001), and featured an average 26bp junction sequence at breakpoints, with GC content consistent to random genomic regions (Fig. 3a-c). Notably, junction sequences with length > 8bp showed non-specific mapping to their partners and nearly all of them (338 out of 339) could be mapped to the nuclear or organellar genomes, or the transcriptome, suggesting that those fusions likely arise from the ligation of nuclease-resistant genomic or transcriptomic fragments during RNA sample preparation [24] (Fig. 3d).

Fig. 3.

Fig. 3

In-depth analysis of fusion transcripts with Unknown, Joined and Misaligned breakpoint types. a Comparison of partner gene expression levels (Transcripts Per Million, TPM) across four breakpoint types and randomly paired expressed genes. P-values were calculated using the Wilcoxon test. b Length distribution of sequence flanking Unknown-type breakpoints. Dashed line marks the mean length (26bp). c Comparison of GC content between Unknown-type sequences and randomly permuted genomic sequences. P-value was calculated using the Wilcoxon test. d Number of Unknown-type sequences (> 8bp) with perfect matches to MH63 nuclear genome, transcriptome, or organellar genome of MH63. The right bar indicates the number of junctions that could be aligned to one of their partner genes. e Classification of Joined-type fusions by breakpoints position (exonic/exon-boundary) and mechanistic subtypes (eight categories; see Additional file 1: Method S8). f Representative examples of Joined-type fusion subtypes, including alignment artifacts (AlignmentError), potential structural variants (Potential SV), transposon-associated fusions (Helitron), and cases with truncated alignments (ShortAlignment; red dashed box highlights minimal 5' alignment). g Comparison of alignment coverage and aligned length for the minimum-end of fusions between “ShortAlignment” and other Joined-type fusions. Wilcoxon test P-value was labeled. h Enrichment of Helitron transposons near breakpoints of “within exon” versus “exon boundary” fusions, relative to random controls (permutation test P-value, N = 1,000). i Breakpoint classification consistency between raw and error-corrected dRNAseq reads (Wilcoxon test P-values). Venn diagram shows shared vs. unique Misaligned fusions. j Case study of a Misaligned fusion where the 5' partner gene fails to properly align to the fusion junction (BLAST E-value < 10e-5). k Network analysis of the Misaligned 5' partner gene from (j), demonstrating its false-positive participation in multiple fusion events. Node colors reflect expression levels (TPM) of associated genes

Manual inspection of Joined-type fusions revealed that 55% harbored previously undetected SHS upon reanalysis with alternative alignment tools such as Mummer [25] or BLAT [26] (Additional file 2: Table S5), highlighting software-dependent limitations in SHS detection sensitivity. The remaining fusions were further classified into four distinct groups by genomic features, including 17 alignment errors, 33 potential structural variants (SVs), 92 short alignment artifacts with poor alignment quality and limited coverage at one breakpoint, and 14 Helitron-associated fusions (Fig. 3e-g). Notably, fusion transcripts with breakpoints at canonical splice sites showed significant enrichment (P < 0.001) of Helitron insertions around their junctions compared to those within exons (Fig. 3h). This suggests that putative trans-splicing fusions, typically characterized by breakpoints at exon boundaries, are intimately linked to TEs activity, indicating their origin as DNA-level variations.

Systematic analysis of three reproducible HC-a fusions (Joined type) revealed their false-positive identities and tight association with DNA-level variants. OsMH_05G0422200#OsMH_09G0065600 (MH63-HC1) were initially classified as HC-a event. However, all supporting reads could be mapped contiguously to a single locus (contig ptg000012l) in the draft MH63 assembly, rather than two separate loci (Additional file 1: Fig. S6). This single locus was not fully assembled in the final MH63 reference genome, leading to the erroneous fusion call. Further comparative analysis using Zhenshan 97 (ZS97, Oryza sativa ssp. xian/indica) genomes revealed that this locus was flanked by an intact Helitron (Additional file 1: Fig. S7; Additional file 2: Table S6), suggesting that the fragmented alignments to chromosomes 5 and 9 likely resulted from stepwise Helitron-mediated gene duplications [27] (Additional file 1: Fig. S8). Similarly, the two HC-a fusions from NIP dRNAseq dataset were also identified as artifacts: one caused by a Helitron-driven structural rearrangement and another by a potential genomic inversion (Additional file 1: Fig. S9).

PCR-free dRNAseq showed higher proportion of Misaligned breakpoints (52.70%), and correction of basecalling errors drastically reduced false-positive calls (Fig. 3i). Comparison of normalized alignment scores, minimal alignment length, coverage of partner genes, and sequence similarity between two partners (Additional file 1: Fig. S10) consistently demonstrated that Misaligned-type fusions showed significantly lower mapping quality (67.7 vs 129 compared with Unknown-type) but contained more similar genes. A case showed erroneous alignment of the 5' partner (OsMH_11G0283700) to fusion reads. This gene recurrently participated in numerous Misaligned-derived fusion events (Fig. 3j, k). These results highlight the importance of breakpoint-based classification for accurate interpretation of fusion transcript.

To further assess genes that are prone to generating artifact fusions, we constructed genome-wide fusion networks incorporating all fusion partners (Additional file 2: Table S7), where node degree represented the number of unique fusion events per gene. Despite their artificial origin, these networks displayed highly non-random topologies (Fig. 4a). While most partners (82.07%) were involved in single fusion, a small fraction (1.11%) showed exceptionally high connectivity (degree ≥ 10) (Fig. 4b).

Fig. 4.

Fig. 4

Network analysis of false-positive fusion transcript partner genes reveals key topological and expression features. a Degree distribution of partner genes in the fusion network, where degree represents the number of unique fusion events per gene. b Ranked distribution of gene degrees (top 100 genes shown) across sequencing protocols, highlighting highly connected partners. c Cross-platform comparison of the top 10 high-degree genes, showing their fusion degree, expression levels (TPM), and fusion type composition. Those genes involved in multiple fusion events spanning different breakpoint types, predominantly SHS in IsoSeq and cDNAseq, and Misaligned in dRNAseq. d Correlation between gene degree and expression level, with linear regression (lm method) demonstrating the relationship between transcript abundance and fusion frequency. e Top 10 gene modules (by size) identified across datasets, with module rankings (left) and cumulative gene frequency distributions (right). Modular analysis identified densely interconnected gene clusters consisting of multiple partner genes, representing recurrent artifact-prone gene pairs. f The types and corresponding counts of fusion transcripts involving OsMH_10G0000800, the highest-degree gene identified in the IsoSeq dataset. g Schematic representation of the different types of fusion transcripts formed by gene OsMH_10G0000800, including SHS, Joined, Unknown, and Misaligned fusions

Analysis of the top 10 high-degree genes across datasets showed that, while these genes exhibited moderate to high expression levels (Fig. 4c), their expression levels correlated only moderately with network degree (R = 0.43; Fig. 4d), suggesting that transcriptional activity alone cannot explain their frequent misidentification. Modular analysis also identified densely interconnected gene clusters representing recurrent artifact-prone gene pairs (Fig. 4e). A representative case, OsMH_10G0000800, participated in 266 fusion events across all breakpoint types, with SHS being predominant (75.56%) (Fig. 4f, g). Its apparent fusion promiscuity may result from both elevated expression level (TPM: 1701.4) and sequence features such as higher GC content (56.15% vs 43.64% compared to genomic average). These results suggest that those recurrent “hub” genes in fusion network can serve as reliable indicators to flag technical artifacts in plant transcriptome.

Fusion detection is sensitive to annotation quality and reference genome selection

To determine whether misannotated adjacent genes mimic fusion transcripts, we developed an Arriba [28] algorithm-based pipeline to systematically identify fusion transcripts from four short-read RNAseq datasets (flag leaf, leaf, root, panicle) [29, 30] (Additional file 1: Fig. S11; Additional file 2: Table S8, S9). We found 62 neighbor-gene fusions, 74% of which were actually single transcription units that had been incorrectly split by annotation errors (Fig. 5a). Using a targeted PCR assay with five primer pairs that designed to distinguish between true fusion transcripts and misannotated single genes (Fig. 5b), we confirmed that cases like OsMH_01G0159500#OsMH_01G0159600 (Fig. 5c) and two other randomly selected fusions were annotation artifacts, as partner genes actually belong to a single locus (Additional file 1: Fig. S12).

Fig. 5.

Fig. 5

Identification and validation of annotation-derived fusion artifacts from short-read RNAseq data. a Summary of adjacent, inverted, and translocated fusion transcripts identified from short-read RNAseq data across four MH63 tissues. The left pie chart indicates the proportion of adjacent gene fusions supported by single long read. The right pie chart shows the proportion of inversion/translocation fusions with breakpoints overlapping canonical splice sites. b PCR validation strategy for an adjacent fusion. Top: Gene models of partners (OsMH_01G0159500 and OsMH_01G0159600) with breakpoints and transposable elements (grey). Bottom: Transcriptome evidence (dRNAseq/IsoSeq/RNAseq) and validation primers (F1/R1 [red], F2/R2 [purple], F3/R3 [blue]). c Gel electrophoresis confirming the adjacent fusion in (B) as a single misannotated gene. Expected bands for individual primers (F1R1, F2R2) and fusion product (F1R2) are shown. Marker: Trans2K Plus DNA Marker (TransGen Biotech, BM111)

We also identified 19 fusion transcripts with translocation- or inversion-like patterns from short reads, six of which contained breakpoints at splice sites (Additional file 1: Fig. S13a). Comparison with long-read data revealed critical discrepancies in these fusions. In one representative case (OsMH_01G0555700#OsMH_10G0374200), a 12bp SHS detected in long-read sequencing was missed in short-read data due to misannotated exon–intron boundary (Additional file 1: Fig. S13b), highlighting the importance of long-read sequencing in resolving annotation-derived artifacts.

Comparative analysis of four annotation sets, original MH63 reference, manually curated MH63 reference (MH63_man), MSU7.0, and IRGSP1.0, revealed substantial variation in gene model annotations, particularly in split (single gene annotated as multiple loci) and fused (multiple genes annotated as one locus) genes (Additional file 1: Fig. S13c-d; Additional file 2: Table S10, S11). Among the fused genes differing between MH63 and MH63_man, 1,317 out of 1,628 (80.90%) were supported by single long read, suggesting that these fusions are also likely annotation-derived artifacts (Additional file 1: Fig. S13e).

We also observed that fusion detection is highly dependent on selected reference genome. Alignment of fusion-supporting reads to divergent genomes often produced false-positive calls due to split or discordant mapping, whereas realignment to the correct genome yielded contiguous reads, indicating these signals may reflect genomic divergence rather than true fusions (Additional file 1: Fig. S14a). To quantify this phenomenon, we realigned fusions from ZS97 dRNAseq data, which was originally mapped to MH63 reference genome, back to the ZS97 reference. This revealed that 82.47% (160/194) of previously identified fusion reads were false positives arising from inter-genotype structural variations (Additional file 1: Fig. S14b; Additional file 2: Table S12). All four HC-a type fusions were completely resolved as artifacts upon ZS97 realignment, with each mapping to a single genomic locus (Additional file 1: Fig. S15, S16).

Absence of reliable allelic and non-allelic fusions in hybrids

Hybrid systems, which inherently combine RNAs from genetically distinct parents, enable simultaneous examination of biological fusion events and technical artifacts arising from RNA intermixing. Using a well-characterized hybrid system involving parental lines MH63 and ZS97 and their hybrid Shanyou 63 (SY63), we systematically identified allelic and non-allelic fusions using trans-mate pairs analysis, which are defined as reads with two mates (or two segments) uniquely and perfectly aligned to distinct parental genomes (Additional file 1: Fig. S17, Method S13; Additional file 2: Table S13, S14). While thousands of allelic and non-allelic fusions were detected, minimal overlap was observed between leaf and root tissues (52.83% for allelic, 0.57% for non-allelic fusions) and between RNAseq and IsoSeq data (34.65% for allelic, 0.76% for non-allelic fusions), indicating that most candidates are non-reproducible false positives (Additional file 1: Fig. S18).

A strong correlation was observed between trans-mate pairs from parental RNA mixture and SY63 (R = 0.76, P < 2.2e-16). Fusion events detected in both the parental mixture control and hybrid sample are definitively technical artifacts, primarily caused by template switching during reverse transcription or cDNA amplification [31]. Based on detectability across samples and replicates, we classified hybrid fusions into four categories: Template Switching (co-detected in control and hybrid); Hybrid-singleRepeat (detected only in one hybrid replicate); Control-only (exclusive to the control), and Manual-review (hybrid-specific, requiring further inspection) (Additional file 1: Fig. S19a). Among non-allelic fusions, 22.46% were Template Switching artifacts (a case illustrated in Additional file 1: Fig. S19b), 26.26% were Hybrid-singleRepeat, and 49.76% were Control-only (Additional file 2: Table S15-S17). Manual curation of Manual-review (1.51%) identified sequencing errors (e.g., discordant SNP signatures between paired reads or reference genome) and alignment artifacts (e.g., multi-mapping in repetitive regions) (Additional file 1: Fig. S19c), but ultimately failed to validate any reliable fusions.

Allelic fusions also exhibited characteristics indicative of their false-positive nature. Those fusions displayed no genomic position bias. Their partner genes showed significantly higher expression levels (TPM 34.62 vs 0.28), gene length (2,033bp vs 1,081bp), and exon number (6 vs 3) compared to random genes. And no statistically significant or biologically interpretable tissue-specific GO terms were identified (Additional file 1: Fig. S20). Similar to non-allelic fusions, the majority of allelic fusions consisted of Template Switching (36.59%), Hybrid-singleRepeat (29.64%) and Control-only (32.34%) artifacts. Among Manual-review fusions, 42.5% were reclassified as Template Switching artifacts in different tissues; for example, one was Manual-review in leaf but Template Switching in root, showing poor cross-tissue consistency (Additional file 1: Fig. S21; Additional file 2: Table S18-S20). Experimental validation using SNP-specific primers on two allelic fusion candidates (one Template Switching and one Manual-review/others) in parents, hybrid and mixed parental cDNA controls failed to detect authentic fusion products (Additional file 1: Fig. S22), confirming their origin as technical artifacts.

Discussion

By integrating multiple long-read datasets, we demonstrated that truly reliable plant fusion transcripts are exceptionally rare, especially those generated from trans-splicing. Most observed cases were SHS-mediated fusion artifacts, particularly in PCR-based methods [24] (Additional file 1: Fig. S23). Other hidden sources, such as misalignment, annotation quality, and reference choice also contributed largely to false positives. From these observations, we proposed the following key considerations for authentic fusion identification in plants under normal or stressed conditions (Fig. 6):

  1. Prioritize long-read sequencing, particularly dRNAseq, offers direct sequencing of native RNA molecules without PCR amplification [32–34], greatly reducing SHS-mediated fusion artifacts.

  2. Use a genotype-matched reference genome and high-quality annotations to minimize mapping errors and distinguish RNA fusions from DNA-level genomic rearrangements [35].

  3. Incorporate rigorous validation controls in well-designed PCR assays to improve accessibility, cost-effectiveness, and interpretability.

Fig. 6.

Fig. 6

Systematic validation framework for distinguishing potential fusion transcripts in plants. This framework includes: sequencing platform selection, accurate breakpoint classification, genome-aware alignment, annotation-aware interpretation, and control-based experimental validation. The recommended criterions and benefits of each corresponding measure are outlined in the accompanying panel. This systematic framework aims to precisely identify and categorize spurious fusion transcripts in plants, which is essential for accurately elucidating the functions of genuine fusion events

Major technological breakthrough such as advances in sequencing often solves existing questions but simultaneously revealing new challenges. The explosion of reported fusion transcript following the rise of short-read sequencing exemplifies this dilemma, as PCR-based methods introduced unexpected technical artifacts that complicate identification of fusions. Emerging long-read technologies, capable of sequencing full-length native RNA molecules, now offer a solution by dramatically improving breakpoint accuracy and revealing previously misclassified fusions. This evolution reflects more than simple linear progress; it represents a recursive cycle in which each generation of tools rectifies and refines the limitations of the last.

Conclusions

In summary, while plant fusion transcripts hold considerable potential for revealing novel regulatory mechanisms or evolutionary events [36, 37], their biological significance remains uncertain without stringent quality control and validation. Our study offers both a cautionary perspective and a solution-oriented path forward, thereby establishing a foundation for more rigorous and reliable exploration of fusion events in plant transcriptomes.

Methods

Plant materials and RNA extraction

All fusion transcripts were analyzed using long-read datasets generated from two Oryza sativa ssp. xian/indica varieties, Minghui 63 (MH63) and Zhenshan 97 (ZS97). Plants were cultivated under standard field conditions at the experimental farm of Huazhong Agricultural University in Wuhan, China. For each variety, the uppermost fully expanded flag leaves on the main culm were sampled 7 days after heading. Samples were immediately flash-frozen in liquid nitrogen upon collection and stored at − 80 °C until RNA extraction. Each sample was collected in two biological repeats except IsoSeq. Total RNA was extracted using Plant RNA Reagent (Thermo Fisher Scientific) and genomic DNA were removed using DNase I (M0303L, NEB). Three distinct long-read RNA sequencing platforms were employed: PacBio IsoSeq, Nanopore direct RNA sequencing (dRNAseq), and Nanopore PCR-amplified cDNA sequencing (cDNAseq). The long-read library preparation and sequencing, data processing can be found in Supplementary Methods.

Identification and classification of putative fusion transcripts from long-read datasets

Putative fusion transcripts were identified independently in each long-read dataset using a modified version of the JAFFAL (v2.3) pipeline [17]. A set of custom scripts was developed to generate the required JAFFAL database files including transcript sequences and metadata based on the MH63 or ZS97 reference genome assemblies and their respective gene annotations (http://rice.hzau.edu.cn/rice_rs3/). We further modified the JAFFAL pipeline to ensure compatibility with the MH63 annotation format and only reads with total alignment coverage greater than 90% were retained. Fusion transcripts in long-read datasets were then identified using the customized JAFFAL with parameters -p genome = MH63 -p annotation = RS3. A fusion transcript was defined based on its partner genes and the coordinates of the genomic breakpoints on both partners. This standardized definition allowed for consistent merging of fusion calls across different datasets and the identification of shared events. We developed a confidence scoring system to categorize fusion transcripts based on supporting read count and breakpoint position. Specifically, fusion transcripts supported by ≥ 2 long reads were classified as High-Confidence (HC), whereas those supported by a single read were designated as Low-Confidence (LC). Among HC fusions, those with breakpoints exactly at annotated exon boundaries were termed HC-a, while those whose breakpoints fell within exons were labeled HC-b. Additionally, fusions occurring between adjacent annotated genes were classified as Potential readthrough (PRT). All fusion transcripts were also classified based on the chromosomal origin of their partner genes: inter-chromosomal fusions involve genes on different chromosomes, while intra-chromosomal fusions occur between genes on the same chromosome. In addition to JAFFAL, we also used LongGF [18] with default parameters to independently identify fusion transcripts from the dRNAseq data.

Breakpoint classification of fusion transcripts

To investigate the potential origins of false-positive fusion transcripts, we performed a detailed classification of fusion breakpoints based on identified fusion transcript sequences. For each fusion event, we first extracted the genomic sequences corresponding to both partner genes, including full gene loci. These partner sequences were then aligned to the corresponding fusion transcript using Blastn “short” mode (-task blastn-short -perc_identity 90 -outfmt 6 -evalue 1e-5) [38] to capture precise breakpoint characteristics. Based on the alignment patterns at and around the breakpoint, we classified each fusion transcript into one of four categories: (i) SHS (short homologous sequence), where both partner gene sequences aligned to the fusion transcript with overlapping regions at the breakpoint, indicative of short sequence homology that may mediate template switching; (ii) Joined, where the two partner gene sequences were directly adjacent at the breakpoint with no overlapping bases; (iii) Unknown, where both partner gene sequences could be mapped to the fusion transcript, but no partner sequence covered the region surrounding the breakpoint; and (iv) Misaligned, where one or both partner gene sequences failed to properly align to the breakpoint region, suggesting spurious mapping or read mis-assignment. For transcripts classified as SHS or Unknown, the overlapping or unmapped sequences around the breakpoints were extracted for downstream characteristic analysis.

Construction and analysis of networks from false-positive fusion transcripts

To explore the topological properties of partner genes involved in false-positive fusion transcripts, we constructed networks using partners of each fusion transcript. In each dataset (IsoSeq, dRNAseq, cDNAseq), nodes represent genes, and undirected edges were drawn between gene partners that formed a fusion transcript. Networks were built using the igraph R package (graph_from_data_frame, directed = F) [39]. For each network, we calculated key topological metrics, including degree (number of fusions each gene participated in) and module composition. Top ten genes with largest degree for each dataset were defined as high-degree genes and extracted to analyze their expression and breakpoint types. Fusion gene modules were identified using the Louvain community detection algorithm implemented in igraph (cluster_louvain). For each module, we recorded the number of unique genes and cumulative module frequency. Modules involving specific genes were visualized using Gephi [40]. To investigate the relationship between gene connectivity and expression, we correlated each gene’s degree with its TPM expression using Pearson correlation and linear regression (lm) models. For representative high-degree genes, we annotated all fusion events they involved in by breakpoint types (SHS, Joined, Unknown and Misaligned) and further visualized using IGV [41] and JBrowse2 [42].

Identification of fusion transcripts from short-read RNAseq data

We implemented a customized filtering pipeline based on the Arriba algorithm (v2.3.0) [28] to identify reliable fusion transcripts from short-read RNAseq data (root, young leaf, panicle and flag leaf of MH63). Given that short-read fusion detection is known to produce a high number of false-positive fusion reads [28], stringent filtering was essential to reduce noise and retain high confidence events. We applied multiple criteria to filter fusion reads using Arriba with default parameters, including removal of redundant reads, retention of fragments with insert sizes > 10 kb, mismatched alignment P-value < 0.01, unique mapping of both fusion partners, and more than two supporting split reads. Subsequently, additional filters were applied based on breakpoint properties and characteristics of partner genes. Fusions with breakpoints located in tandemly repeated gene clusters or within a single gene model were excluded. Partner genes with expression levels in the top 1% or sequence identity exceeding 30% were discarded to avoid false positives caused by template-switching and ambiguous mapping. Adjacent breakpoints were cluster to avoid redundant calls, and only fusion events reproducibly detected in both biological replicates were retained for downstream analyze. Each retained fusion event was further annotated according to Arriba’s classification scheme, which includes fusion type (e.g., adjacent gene fusion, translocation, inversion), confidence level, and whether the breakpoint falls within canonical splice sites.

In vitro PCR validation of fusion transcripts

Total RNA was extracted from MH63 rice tissues using TRIzol™ Reagent (Invitrogen) according to the manufacturer’s protocol. For first-strand cDNA synthesis, three reverse transcriptase (RT) systems with distinct RNase H activities were employed to assess their potential influence on artificial fusion transcript formation: (i) Engineered Moloney Murine Leukemia Virus Reverse Transcriptase (EM-MLV, modified to enhance cDNA synthesis efficiency), (ii) Natural Avian Myeloblastosis Virus Reverse Transcriptase (AMV, NEB), and (iii) Natural Moloney Murine Leukemia Virus Reverse Transcriptase (M-MLV, Invitrogen). Reverse transcription was performed using 1 μg of total RNA and oligo(dT)₁₅ primers at 42 °C for AMV or 37 °C for M-MLV/EM-MLV, following manufacturer’s recommended protocol.

To investigate whether SHS promoted false-positive fusion transcript formation, two primer sets were used: one containing SHS segments and one lacking SHS segments near the junction. The first-round PCR was performed using primer combinations MH63-SHS1-F1/R2 and MH63-SHS1-F2/R1 or MH63-SHS1-QF/QR with MH63 cDNA (synthesized by EM-MLV, AMV or M-MLV) as the template. PCR products were gel-purified using the TIANGEN DNA Extraction Kit and verified by Sanger sequencing following standard procedures. Equimolar purified gene A (OsMH_01G0028500) and gene B (OsMH_02G0018200) PCR products were then mixed and diluted 1:100 as templates in the second-round PCR. The second-round PCR employed the MH63-SHS1-F1/R2 primer pair (identical to those used in the first round) to amplify potential fusion transcripts using different equimolar templates: (i) SHS-included templates, (ii) SHS-removed templates and (iii) MH63 cDNA templates. Standard and touchdown PCR were carried out using KOD One™ PCR Master Mix (Sigma-Aldrich). Touchdown cycling parameters included an initial denaturation at 95 °C for 2 min, followed by 10 cycles of annealing from 68 °C to 58 °C (–1°C per cycle), and 25 additional cycles at 58 °C to enhance amplification specificity. PCR products were analyzed by agarose gel electrophoresis and sequenced to confirm fusion junctions. The Trans2K Plus DNA Marker (TransGen Biotech, catalog no. BM111) were used the molecular size standard. All experiments were performed in three independent replicates. Primer sequences for all targets are provided in Additional file 2: Table S21.

In vivo FACS (fluorescence-activated cell sorting)-sorted dual-fluorescence PCR

Full-length coding sequences of OsMH_01G0028500 (designated MH63-SHS1-A) and OsMH_02G0018200 (MH63-SHS1-B) were amplified from MH63 cDNA using primers containing 20-bp homologous recombination arms corresponding to the KpnI and BamHI restriction sites of pCAMBIA1300-Ubiquitin vectors. The coding fragments were fused to their respective fluorescent reporters (eGFP for MH63-SHS1-A and mCherry for MH63-SHS1-B) via nested PCR, and cloned into the over-expression vector through homologous recombination. All constructs were sequence-verified and purified using the QIAGEN Plasmid Midi Kit.

Rice (MH63) protoplasts were isolated from 14-day-old etiolated seedlings following. Four transfection groups were prepared: (i) eGFP-A only, (ii) B-mCherry only, (iii) co-transfection with equimolar amounts of eGFP-A and B-mCherry plasmids, and (iv) control (equal volume of sterile water). Protoplast transfection was performed via polyethylene glycol (PEG)-mediated transformation using a solution containing 10% (w/v) PEG and 100mM CaCl2. All transformations were conducted in triplicate, and protoplasts were incubated for 16 h at 28 °C in complete darkness to allow transient expression. After incubation, protoplasts were sorted by fluorescence-activated cell sorting (FACS) using a BD FACS Aria II SORP flow cytometer equipped with FITC and mCherry channels and a 100 μm nozzle. Four distinct populations were collected: eGFP⁻/mCherry⁻ (control), eGFP⁺/mCherry⁻ (A-only), eGFP⁻/mCherry⁺ (B-only), and eGFP⁺/mCherry⁺ (co-expression). Approximately 1 × 10⁶ cells were collected per population. Fluorescence expression and subcellular localization were confirmed by confocal microscopy (Nikon, AX R NSPARC, Japan).

Total RNA was the extracted from equivalent cell numbers of each sorted population using TRIzol™ Reagent (Invitrogen), followed by reverse transcription using either AMV or M-MLV reverse transcriptase, according to manufacturer’s instructions. qRT-PCR was performed using SYBR Green ROX dye with gene-specific primers for MH63-SHS1-A, MH63-SHS1-B, eGFP, and mCherry to verify expression of the introduced constructs. The Ubiquitin gene was used as an internal control for normalization. To evaluate whether SHS induced false-positive fusion transcript formation in vivo, cDNA templates from (i) control, (ii) equimolar mixed samples (A-only and B-only cDNAs mixed post-transcription), and (iii) co-transfected eGFP-A + B-mCherry cells were normalized to equal concentration. These cDNAs were subjected to PCR and qRT-PCR amplification using primer sets targeting gene A, gene B, and the predicted fusion transcript. PCR products were analyzed by agarose gel electrophoresis and verified by Sanger sequencing. qRT-PCR was further performed to quantify fusion transcript abundance across samples derived from different reverse transcriptases and transfection conditions. Relative expression levels were calculated using the 2^–ΔΔCt method. Statistical significance was evaluated using Student’s t-test. All experiments were performed in three independent biological replicates. Primers and plasmid construction details are listed in Additional file 2: Table S21.

Software and computational resources

Additional computational tools were used for sequence processing, genome and transcriptome analysis, statistical testing, and data visualization. Raw short- and long-read sequencing data were processed and assessed using Guppy (https://nanoporetech.com/software/other/guppy), fastp [43], NanoPack [44], FMLRC2 [45], and IsoSeq3 (https://github.com/ylipacbio/IsoSeq3). Reads were aligned to the corresponding reference genomes or transcriptomes using BWA [46], HISAT2 [47], STAR [48], or minimap2 [49], whereas AnchorWave [50] was used for gene-anchored whole-genome alignment and comparative genomic analysis. SAMtools [51], BEDTools [52], SeqKit [53], and deepTools [54] were used for alignment manipulation, genomic interval operations, sequence processing, and the generation of sequencing-coverage tracks, respectively.

Transcript abundance was quantified using featureCounts [55], RSEM [56], or Salmon [57]. Trinity [58] was used for the construction of expression matrix. Long-read transcript models were identified, quantified, and quality controlled using TALON [59] and SQANTI3 [60]. Draft genome assemblies were generated or examined using hifiasm [61], whereas genome annotation files were standardized, manipulated, and compared using AGAT [62]. Repetitive sequences and Helitron-like transposable elements were annotated using RepeatMasker [63] and HELIANO [64], respectively.

Downstream data processing and statistical analyses were conducted in R using tidyverse [65], GenomicFeatures [66], and regioneR [67]. Figures were generated using ggplot2 [68], ggpubr [69], circlize [70], ComplexHeatmap [71], and DiffLogo [72]. Detailed software versions, analysis-specific applications, parameters, and resource links are provided in Additional file 1: Method S1-S16 and Additional file 2: Table S22.

Supplementary Information

13059_2026_4179_MOESM1_ESM.pdf (8.6MB, pdf)

Additional file 1: Supplementary Figures and Supplementary Methods. Contains Fig. S1-S23, and Method S1-S16.

13059_2026_4179_MOESM2_ESM.xlsx (2MB, xlsx)

Additional file 2: Supplementary Tables. Contains Table S1-S22 in a multi-tab Excel workbook.

Acknowledgements

The authors would like to thank the National Key Laboratory of Crop Genetic Improvement, Huazhong Agricultural University for providing the bioinformatics computing resources. The authors would also like to thank Engineer Jianing Zhang and the National Key Laboratory of Agricultural Microbiology Core Facility, Huazhong Agricultural University for their assistance and technical support with confocal microscopy.

Peer review information

Wenjing She was the primary editor of this article and managed its editorial process and peer review in collaboration with the rest of the editorial team. This manuscript was first reviewed at another journal. Only the peer-review history at Genome Biology is available in the online version of this article.

Authors' contributions

L.-L.C. and Y.O. conceived and designed the experiments. X.-T.Z., X.L., Z. Z., Q.T., R.Z., H.L., and S.J. P. analyzed the sequencing data. X. L., Z. Z., M.L., F.X., and X.Z performed the experimental validation. X.-T.Z. and Q.T. prepared the figures. X.-T.Z., X. L., Y.O., and L.-L.C. wrote the manuscript. All authors read and approved the final manuscript.

Funding

This work was supported by the National Natural Science Foundation of China (U24A20369, 32341031), Guangxi Natural Science Foundation (2024GXNSFGA010003), Guangxi Science and Technology Major Program (guikeAA23062085), the Fundamental Research Program of Hubei Province (2024AFE001), the Fundamental Research Funds for the Central Universities (2662025SKPY007), and the Earmarked Fund for China Agriculture Research System (CARS-01).

Data availability

The Fastq files of long- and short-read RNA sequencing data for each sample have been deposited in the Sequence Read Archive at the National Center for Biotechnology Information (NCBI) under the accession number PRJNA1291274 [73]. Custom scripts are available on GitHub (https://github.com/Zhuxitong/Spurious-Fusion-Transcripts-in-Plants) [74] and Zenodo (10.5281/zenodo.19881401) [75]. The source code is released under the OSI-approved MIT License.

Several publicly available datasets were also used in this study. Short-read RNAseq data from the flag leaf of MH63 were obtained from the NCBI under BioProject PRJNA550148 [76], whereas RNAseq data from the panicle, leaf, and root of MH63 were obtained from the National Genomics Data Center (NGDC) under accession CRA002886 [77]. Direct RNA sequencing data from six tissues of the rice cultivar Nipponbare were obtained from NCBI under BioProject PRJNA752930 [78], and direct RNA sequencing data from two tissues of A. thaliana were obtained from the Gene Expression Omnibus (GEO) under accession GSE144828 [79]. Chromatin and nucleic acid interaction datasets from MH63 were obtained from the GEO database. These included H3K4me3- and RNA polymerase II-associated DNA–DNA interaction data under accession GSE131202 [80], H3K4me3-associated ChRD–PET DNA–RNA interaction data under accession GSE163119 [81], and H3K9ac-associated TaDRIM-seq DNA–RNA and RNA–RNA interaction data under accession GSE234228 [82]. PacBio HiFi genomic reads for MH63 and ZS97 were obtained from the NGDC under accession PRJCA005549 [83, 84], and HiFi genomic reads for Nipponbare were obtained under accession PRJCA018610 [85, 86].

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Xi-Tong Zhu, Xinyan Lu, Zengxin Zhang and Qian Tang contributed equally to this work.

Contributor Information

Yidan Ouyang, Email: diana1983941@mail.hzau.edu.cn.

Ling-Ling Chen, Email: llchen@gxu.edu.cn.

References

  • 1.Dorney R, Dhungel BP, Rasko JEJ, Hebbard L, Schmitz U. Recent advances in cancer fusion transcript detection. Brief Bioinform. 2023;24:bbac519. 10.1093/bib/bbac519. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Lei Q, Li C, Zuo Z, Huang C, Cheng H, Zhou R. Evolutionary insights into RNA trans-splicing in vertebrates. Genome Biol Evol. 2016;8:562–77. 10.1093/gbe/evw025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Li Z, Qin F, Li H. Chimeric RNAs and their implications in cancer. Curr Opin Genet Dev. 2018;48:36–43. 10.1016/j.gde.2017.10.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Cong J, Zhang S, Zhang Q, Yu X, Huang J, Wei X, et al. Conserved features and diversity attributes of chimeric RNAs across accessions in four plants. Plant Biotechnol J. 2024;22:3151–63. 10.1111/pbi.14437. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Singh A, Zahra S, Das D, Kumar S. AtFusionDB: a database of fusion transcripts in Arabidopsis thaliana. Database (Oxford). 2019;2019:bay135. 10.1093/database/bay135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chitkara P, Singh A, Gangwar R, Bhardwaj R, Zahra S, Arora S, et al. The landscape of fusion transcripts in plants: a new insight into genome complexity. BMC Plant Biol. 2024;24:1–18. 10.1186/s12870-024-05900-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Wang B, Tseng E, Regulski M, Clark TA, Hon T, Jiao Y, et al. Unveiling the complexity of the maize transcriptome by single-molecule long-read sequencing. Nat Commun. 2016;7:11708. 10.1038/ncomms11708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Arya A, Arora S, Hamid F, Kumar S. PFusionDB: a comprehensive database of plant-specific fusion transcripts. 3 Biotech. 2024;14:282. 10.1007/s13205-024-04132-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zhang Y, Li H, Shen Y, Wang S, Tian L, Yin H, et al. Readthrough events in plants reveal plasticity of stop codons. Cell Rep. 2024;43:113723. 10.1016/j.celrep.2024.113723. [DOI] [PubMed] [Google Scholar]
  • 10.Chen M-X, Zhu F-Y, Gao B, Ma K-L, Zhang Y, Fernie AR, et al. Full-length transcript-based proteogenomics of rice improves its genome and proteome annotation. Plant Physiol. 2020;182:1510–26. 10.1104/pp.19.00430. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Zhang S, Li R, Zhang L, Chen S, Xie M, Yang L, et al. New insights into Arabidopsis transcriptome complexity revealed by direct sequencing of native RNAs. Nucleic Acids Res. 2020;48:7700–11. 10.1093/nar/gkaa588. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Houseley J, Tollervey D. Apparent non-canonical trans-splicing is generated by reverse iranscriptase in vitro. PLoS ONE. 2010;5:e12271. 10.1371/journal.pone.0012271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Song B, Buckler ES, Stitzer MC. New whole-genome alignment tools are needed for tapping into plant diversity. Trends Plant Sci. 2024;29:355–69. 10.1016/j.tplants.2023.08.013. [DOI] [PubMed] [Google Scholar]
  • 14.Wang S, Mank JE, Ortiz-Barrientos D, Rieseberg LH. Genome architecture and speciation in plants and animals. Mol Ecol. 2025;34:e70004. 10.1111/mec.70004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Garalde DR, Snell EA, Jachimowicz D, Sipos B, Lloyd JH, Bruce M, et al. Highly parallel direct RNA sequencing on an array of nanopores. Nat Methods. 2018;15:201–6. 10.1038/nmeth.4577. [DOI] [PubMed] [Google Scholar]
  • 16.Zhang T, Li J, Tang C, Wu Y, Wu H, Zhu XT, et al. Nanopore direct RNA sequencing and the epitranscriptome: advances in mapping native RNA landscapes. iMeta. 2026;5:e70136. 10.1002/imt2.70136.
  • 17.Davidson NM, Chen Y, Sadras T, Ryland GL, Blombery P, Ekert PG, et al. JAFFAL: detecting fusion genes with long-read transcriptome sequencing. Genome Biol. 2022;23:10. 10.1186/s13059-021-02588-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Liu Q, Hu Y, Stucky A, Fang L, Zhong JF, Wang K. LongGF: computational algorithm and software tool for fast and accurate detection of gene fusions by long-read transcriptome sequencing. BMC Genomics. 2020;21:793. 10.1186/s12864-020-07207-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Liu C, Zhang Y, Li X, Jia Y, Li F, Li J, et al. Evidence of constraint in the 3D genome for trans-splicing in human cells. Sci China Life Sci. 2020;63:1380–93. 10.1007/s11427-019-1609-6. [DOI] [PubMed] [Google Scholar]
  • 20.Zhao L, Wang S, Cao Z, Ouyang W, Zhang Q, Xie L, et al. Chromatin loops associated with active genes and heterochromatin shape rice genome architecture for transcriptional regulation. Nat Commun. 2019;10:3640. 10.1038/s41467-019-11535-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Xiao Q, Huang X, Zhang Y, Xu W, Yang Y, Zhang Q, et al. The landscape of promoter-centred RNA–DNA interactions in rice. Nat Plants. 2022;8:157–70. 10.1038/s41477-021-01089-4. [DOI] [PubMed] [Google Scholar]
  • 22.Ding C, Chen G, Luan S, Gao R, Fan Y, Zhang Y, et al. Simultaneous profiling of chromatin-associated RNA at targeted DNA loci and RNA-RNA interactions through TaDRIM-seq. Nat Commun. 2025;16:1500. 10.1038/s41467-024-53534-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Yu F, Qi H, Gao L, Luo S, Njeri Damaris R, Ke Y, et al. Identifying RNA modifications by direct RNA sequencing reveals complexity of epitranscriptomic dynamics in rice. Genomics Proteomics Bioinformatics. 2023;21:788–804. 10.1016/j.gpb.2023.02.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Peng Z, Yuan C, Zellmer L, Liu S, Xu N, Liao DJ. Hypothesis: artifacts, including spurious chimeric RNAs with a short homologous sequence, caused by consecutive reverse transcriptions and endogenous random primers. J Cancer. 2015;6:555. 10.7150/jca.11997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Marçais G, Delcher AL, Phillippy AM, Coston R, Salzberg SL, Zimin A. MUMmer4: a fast and versatile genome alignment system. PLoS Comput Biol. 2018;14:e1005944. 10.1371/journal.pcbi.1005944. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kent WJ. BLAT–the BLAST-like alignment tool. Genome Res. 2002;12:656–64. 10.1101/gr.229202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ramakrishnan M, Satish L, Sharma A, Kurungara Vinod K, Emamverdian A, Zhou M, et al. Transposable elements in plants: recent advancements, tools and prospects. Plant Mol Biol Rep. 2022;40:628–45. 10.1007/s11105-022-01342-w. [Google Scholar]
  • 28.Uhrig S, Ellermann J, Walther T, Burkhardt P, Fröhlich M, Hutter B, et al. Accurate and efficient detection of gene fusions from RNA sequencing data. Genome Res. 2021;31:448–60. 10.1101/gr.257246.119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Shao L, Xing F, Xu CH, Zhang QH, Che J, Wang XM, et al. Patterns of genome-wide allele-specific expression in hybrid rice and the implications on the genetic basis of heterosis. Proc Natl Acad Sci U S A. 2019;116:5653–8. 10.1073/pnas.1820513116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhu XT, Zhou R, Che J, Zheng YY, Tahir ul Qamar M, Feng JW, et al. Ribosome profiling reveals the translational landscape and allele-specific translational efficiency in rice. Plant Commun. 2023;4:100457. 10.1016/j.xplc.2022.100457. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.McManus CJ, Duff MO, Eipper-Mains J, Graveley BR. Global analysis of trans-splicing in Drosophila. Proc Natl Acad Sci U S A. 2010;107:12975–9. 10.1073/pnas.1007586107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Jain M, Abu-Shumays R, Olsen HE, Akeson M. Advances in nanopore direct RNA sequencing. Nat Methods. 2022;19:1160–4. 10.1038/s41592-022-01633-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Zhu XT, Sanz-Jimenez P, Ning XT, Tahir ul Qamar M, Chen LL. Direct RNA sequencing in plants: practical applications and future perspectives. Plant Commun. 2024;5:101064. 10.1016/j.xplc.2024.101064. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang T, Li H, Jiang M, Hou H, Gao Y, Li Y, et al. Nanopore sequencing: flourishing in its teenage years. J Genet Genom. 2024;51:1361–74. 10.1016/j.jgg.2024.09.007. [DOI] [PubMed] [Google Scholar]
  • 35.Zhou Y, Zhang C, Zhang L, Ye Q, Liu N, Wang M, et al. Gene fusion as an important mechanism to generate new genes in the genus Oryza. Genome Biol. 2022;23:130. 10.1186/s13059-022-02696-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Arora S, Hamid F, Kumar S. Fusion transcripts in plants: hidden layer of transcriptome complexity. Trends Plant Sci. 2025;30:229–31. 10.1016/j.tplants.2024.12.004. [DOI] [PubMed] [Google Scholar]
  • 37.Hamid F, Arora S, Kumar S. Breaking and making genes: the genesis of novel traits in plants. New Phytol. 2026;249:2746–59. 10.1111/nph.70844. [DOI] [PubMed] [Google Scholar]
  • 38.Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10:421. 10.1186/1471-2105-10-421. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Csardi G, Nepusz T. The igraph software package for complex network research. Complex Systems. 2006;1695:1–9. https://igraph.org.
  • 40.Bastian M, Heymann S, Jacomy M. Gephi: an open source software for exploring and manipulating networks. Proc Int AAAI Conf. 2009;3:361–2. 10.1609/icwsm.v3i1.13937.
  • 41.Robinson JT, Thorvaldsdóttir H, Winckler W, Guttman M, Lander ES, Getz G, et al. Integrative genomics viewer. Nat Biotechnol. 2011;29:24–6. 10.1038/nbt.1754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Diesh C, Stevens GJ, Xie P, De Jesus MT, Hershberg EA, Leung A, et al. JBrowse 2: a modular genome browser with views of synteny and structural variation. Genome Biol. 2023;24:74. 10.1186/s13059-023-02914-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Chen SF, Zhou YQ, Chen YR, Gu J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018;34:i884–90. 10.1093/bioinformatics/bty560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.De Coster W, Rademakers R. Nanopack2: population-scale evaluation of long-read sequencing data. Bioinformatics. 2023;39:btad311. 10.1093/bioinformatics/btad311. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Mak QXC, Wick RR, Holt JM, Wang JR. Polishing de novo Nanopore assemblies of bacteria and eukaryotes with FMLRC2. Mol Biol Evol. 2023;40:msad048. 10.1093/molbev/msad048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Li H, Durbin R. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics. 2009;25:1754–60. 10.1093/bioinformatics/btp324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Kim D, Langmead B, Salzberg SL. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015;12:357–60. 10.1038/nmeth.3317. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018;34:3094–100. 10.1093/bioinformatics/bty191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Song BX, Marco-Sola S, Moreto M, Johnson L, Buckler ES, Stitzer MC. AnchorWave: sensitive alignment of genomes with high sequence diversity, extensive structural polymorphism, and whole-genome duplication. Proc Natl Acad Sci U S A. 2022;119:e2113075119. 10.1073/pnas.2113075119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The sequence alignment/map format and SAMtools. Bioinformatics. 2009;25:2078–9. 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–2. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Shen W, Sipos B, Zhao L. SeqKit2: a Swiss army knife for sequence and alignment processing. iMeta. 2024;3:e191. 10.1002/imt2.191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Ramírez F, Ryan DP, Grüning B, Bhardwaj V, Kilpert F, Richter AS, et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Res. 2016;44:W160–5. 10.1093/nar/gkw257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Liao Y, Smyth GK, Shi W. FeatureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30:923–30. 10.1093/bioinformatics/btt656. [DOI] [PubMed] [Google Scholar]
  • 56.Li B, Dewey CN. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformics. 2011;12:323. 10.1186/1471-2105-12-323. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14:417–9. 10.1038/nmeth.4197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Haas BJ, Papanicolaou A, Yassour M, Grabherr M, Blood PD, Bowden J, et al. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc. 2013;8:1494–512. 10.1038/nprot.2013.084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wyman, D, Balderrama-Gutierrez G, Reese F, Jiang S, Rahmanian S, Forner S, et al. A technology-agnostic long-read analysis pipeline for transcriptome discovery and quantification. bioRxiv. 2020;672931. 10.1101/672931.
  • 60.Pardo-Palacios FJ, Arzalluz-Luque A, Kondratova L, Salguero P, Mestre-Tomás J, Amorín R, et al. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms. Nat Methods. 2024;21:793–7. 10.1038/s41592-024-02229-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Cheng H, Concepcion GT, Feng X, Zhang H, Li H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18:170–5. 10.1038/s41592-020-01056-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Dainat J. Another Gtf/Gff Analysis Toolkit (AGAT): resolve interoperability issues and accomplish more with your annotations. Plant and animal genome XXIX conference. 2022. https://zenodo.org/records/13799920.
  • 63.Tarailo-Graovac M, Chen N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinform. 2009;25:4.10.1-4.10.14. 10.1002/0471250953.bi0410s25. [DOI] [PubMed] [Google Scholar]
  • 64.Li Z, Gilbert C, Peng H, Pollet N. Discovery of numerous novel Helitron-like elements in eukaryote genomes using HELIANO. Nucleic Acids Res. 2024;52:e79. 10.1093/nar/gkae679. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Wickham H, Averick M, Bryan J, Chang W, McGowan LDA, François R, et al. Welcome to the Tidyverse. J Open Source Softw. 2019;4:1686. 10.21105/joss.01686. [Google Scholar]
  • 66.Lawrence M, Huber W, Pagès H, Aboyoun P, Carlson M, Gentleman R, et al. Software for computing and annotating genomic ranges. PLoS Comput Biol. 2013;9:e1003118. 10.1371/journal.pcbi.1003118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Gel B, Díez-Villanueva A, Serra E, Buschbeck M, Peinado MA, Malinverni R. regioneR: an R/Bioconductor package for the association analysis of genomic regions based on permutation tests. Bioinformatics. 2016;32:289–91. 10.1093/bioinformatics/btv562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Wickham H. ggplot2: elegant graphics for data analysis. Springer-Verlag New York. 2016;ISBN 978–3–319–24277–4. https://ggplot2.tidyverse.org/.
  • 69.Kassambara A. ggpubr: ggplot2 based publication ready plots. R package. 2020. https://rpkgs.datanovia.com/ggpubr/.
  • 70.Gu Z, Gu L, Eils R, Schlesner M, Brors B. circlize implements and enhances circular visualization in R. Bioinformatics. 2014;30:2811–2. 10.1093/bioinformatics/btu393. [DOI] [PubMed] [Google Scholar]
  • 71.Gu ZG, Eils R, Schlesner M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics. 2016;32:2847–9. 10.1093/bioinformatics/btw313. [DOI] [PubMed] [Google Scholar]
  • 72.Nettling M, Treutler H, Grau J, Keilwagen J, Posch S, Grosse I. DiffLogo: a comparative visualization of sequence motifs. BMC Bioinformics. 2015;16:387. 10.1186/s12859-015-0767-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Zhu XT, Lu X, Zhang Z, Tang Q, Liu M, Xia F, et al. The hidden sources of spurious fusion transcripts in plants. Datasets. NCBI BioProject. 2026. https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1291274/. [DOI] [PMC free article] [PubMed]
  • 74.Zhu XT, Lu X, Zhang Z, Tang Q, Liu M, Xia F, et al. The hidden sources of spurious fusion transcripts in plants. GitHub. 2026. https://github.com/Zhuxitong/Spurious-Fusion-Transcripts-in-Plants. [DOI] [PMC free article] [PubMed]
  • 75.Zhu XT, Lu X, Zhang Z, Tang Q, Liu M, Xia F, et al. The hidden sources of spurious fusion transcripts in plants. Zenodo. 2026. 10.5281/zenodo.19881401. [DOI] [PMC free article] [PubMed]
  • 76.Shao L, Xing F, Xu CH, Zhang QH, Che J, Wang XM, et al. Patterns of genome-wide allele-specific expression in hybrid rice and the implications on the genetic basis of heterosis. Datasets. NCBI BioProject. 2019. https://www.ncbi.nlm.nih.gov/bioproject/PRJNA550148/. [DOI] [PMC free article] [PubMed]
  • 77.Zhu XT, Zhou R, Che J, Zheng YY, Tahir ul Qamar M, Feng JW, et al. Ribosome profiling reveals the translational landscape and allele-specific translational efficiency in rice. Datasets. NGDC Genome Sequence Archive. 2023. https://ngdc.cncb.ac.cn/gsa/browse/CRA002886. [DOI] [PMC free article] [PubMed]
  • 78.Yu F, Qi H, Gao L, Luo S, Njeri Damaris R, Ke Y, et al. Identifying RNA modifications by direct RNA sequencing reveals complexity of epitranscriptomic dynamics in rice. Datasets. NCBI BioProject. 2022. https://www.ncbi.nlm.nih.gov/bioproject/PRJNA752930/. [DOI] [PMC free article] [PubMed]
  • 79.Zhang S, Li R, Zhang L, Chen S, Xie M, Yang L, et al. New insights into Arabidopsis transcriptome complexity revealed by direct sequencing of native RNAs. Datasets. Gene Expression Omnibus. 2020. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE144828. [DOI] [PMC free article] [PubMed]
  • 80.Zhao L, Wang S, Cao Z, Ouyang W, Zhang Q, Xie L, et al. Chromatin loops associated with active genes and heterochromatin shape rice genome architecture for transcriptional regulation. Datasets. Gene Expression Omnibus. 2019. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE131202. [DOI] [PMC free article] [PubMed]
  • 81.Xiao Q, Huang X, Zhang Y, Xu W, Yang Y, Zhang Q, et al. The landscape of promoter-centred RNA–DNA interactions in rice. Datasets. Gene Expression Omnibus. 2021. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE163119. [DOI] [PubMed]
  • 82.Ding C, Chen G, Luan S, Gao R, Fan Y, Zhang Y, et al. Simultaneous profiling of chromatin-associated RNA at targeted DNA loci and RNA-RNA Interactions through TaDRIM-seq. Datasets. Gene Expression Omnibus. 2023. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE234228. [DOI] [PMC free article] [PubMed]
  • 83.Song JM, Xie WZ, Wang S, Guo YX, Koo DH, Kudrna D, et al. Two gap-free reference genomes and a global view of the centromere architecture in rice. Mol Plant. 2021;14:1757–67. 10.1016/j.molp.2021.06.018. [DOI] [PubMed] [Google Scholar]
  • 84.Song JM, Xie WZ, Wang S, Guo YX, Koo DH, Kudrna D, et al. Two gap-free reference genomes and a global view of the centromere architecture in rice. Datasets. NGDC BioProject. 2021. https://ngdc.cncb.ac.cn/bioproject/browse/PRJCA005549. [DOI] [PubMed]
  • 85.Shang L, He W, Wang T, Yang Y, Xu Q, Zhao X, et al. A complete assembly of the rice Nipponbare reference genome. Mol Plant. 2023;16:1232–6. 10.1016/j.molp.2023.08.003. [DOI] [PubMed] [Google Scholar]
  • 86.Shang L, He W, Wang T, Yang Y, Xu Q, Zhao X, et al. A complete assembly of the rice Nipponbare reference genome. Datasets. NGDC BioProject. 2023. https://ngdc.cncb.ac.cn/bioproject/browse/PRJCA018610. [DOI] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

13059_2026_4179_MOESM1_ESM.pdf (8.6MB, pdf)

Additional file 1: Supplementary Figures and Supplementary Methods. Contains Fig. S1-S23, and Method S1-S16.

13059_2026_4179_MOESM2_ESM.xlsx (2MB, xlsx)

Additional file 2: Supplementary Tables. Contains Table S1-S22 in a multi-tab Excel workbook.

Data Availability Statement

The Fastq files of long- and short-read RNA sequencing data for each sample have been deposited in the Sequence Read Archive at the National Center for Biotechnology Information (NCBI) under the accession number PRJNA1291274 [73]. Custom scripts are available on GitHub (https://github.com/Zhuxitong/Spurious-Fusion-Transcripts-in-Plants) [74] and Zenodo (10.5281/zenodo.19881401) [75]. The source code is released under the OSI-approved MIT License.

Several publicly available datasets were also used in this study. Short-read RNAseq data from the flag leaf of MH63 were obtained from the NCBI under BioProject PRJNA550148 [76], whereas RNAseq data from the panicle, leaf, and root of MH63 were obtained from the National Genomics Data Center (NGDC) under accession CRA002886 [77]. Direct RNA sequencing data from six tissues of the rice cultivar Nipponbare were obtained from NCBI under BioProject PRJNA752930 [78], and direct RNA sequencing data from two tissues of A. thaliana were obtained from the Gene Expression Omnibus (GEO) under accession GSE144828 [79]. Chromatin and nucleic acid interaction datasets from MH63 were obtained from the GEO database. These included H3K4me3- and RNA polymerase II-associated DNA–DNA interaction data under accession GSE131202 [80], H3K4me3-associated ChRD–PET DNA–RNA interaction data under accession GSE163119 [81], and H3K9ac-associated TaDRIM-seq DNA–RNA and RNA–RNA interaction data under accession GSE234228 [82]. PacBio HiFi genomic reads for MH63 and ZS97 were obtained from the NGDC under accession PRJCA005549 [83, 84], and HiFi genomic reads for Nipponbare were obtained under accession PRJCA018610 [85, 86].


Articles from Genome Biology are provided here courtesy of BMC

RESOURCES