Abstract
One of the main challenges of whole transcriptome sequencing is the difficulty in detecting and quantifying low-to-moderate abundance transcripts. Methods that address this are either complicated to scale or customize; long-range PCR is problematic to scale, and probe hybridization panels are expensive to customize. In this study, we developed an RNA-guided CRISPR–Cas9 nuclease-based enrichment strategy combined with long-read sequencing, which achieved up to 60-fold enrichment of the target. Our findings demonstrate that the CRISPR–Cas system is a highly effective method for customizable long-read sequencing of target transcripts, which preserves estimation of relative abundance.
Graphical Abstract
Graphical Abstract.

Introduction
In eukaryotic cells, a single gene can produce multiple messenger RNA (mRNA) products through cell type-specific processes, including alternative transcription initiation, splicing, intron retention, and polyadenylation [1–3]. In humans, ~95% of all protein-coding genes undergo alternative splicing to varying degrees [4]. Each isoform can contain a unique set of features that influence protein production differently [5], thereby diversifying the transcriptome and proteome. Switching between transcript isoforms has been detected in several diseases and infections [6–8], including infections with Severe Acute Respiratory Coronavirus 2 (SARS-CoV-2) [9]. Thus, examining transcriptome diversity not only at the gene level but also at the isoform level is necessary for fully understanding disease pathogenesis.
Long-read RNA sequencing has been broadly used to understand gene expression in different cell types and to resolve full-length transcript isoforms as well as other complex RNA processing events [10–13]. Pacific Biosciences (PacBio) [14] and Oxford Nanopore Technologies (ONT) [15] are two dominant platforms that have overcome major limitations caused by short-read sequencing by generating reads longer than 10 kb and directly sequencing full-length transcript molecules end-to-end instead of short reads at 150–500 bp [16, 17]. ONT sequencing platform provides two PCR-free protocols to generate long-read RNA sequencing, direct cDNA sequencing, and direct RNA sequencing (dRNA-seq). While direct cDNA sequencing requires reverse transcription, dRNA-seq allows native RNA molecules to be translocated through nanopores, avoiding both reverse transcription and PCR amplification [18]. However, the process often involves the sequencing of the whole transcriptome, resulting in sequencing reads being disproportionately allocated to highly expressed transcripts. Consequently, low-abundance isoforms receive limited coverage, reducing confidence in their characterization and quantification. Thus, the lowly expressed and rare isoforms are frequently overlooked from bulk sequencing due to insufficient per-transcript coverage. Moreover, many diagnostic assays only require information on expression of a subset of transcripts. Several targeted sequencing approaches have been developed to allow deeper sequencing of specific transcripts within complex transcriptomes. Pre-amplifying the target transcripts by using long-range RT-PCR followed by long-read sequencing can generate full-length transcripts. Yet, primer cross-reactivity and amplification bias lead to the difficulty of scaling up to large gene panels [19]. Hybridization capture-based enrichment using biotinylated [20, 21] and synthesizing capture oligos [22] have been recently introduced. These methods rely on the generation of oligos with high tiling density across target transcripts, which is either expensive when ordered from commercial suppliers or requires complex experimental protocols to create. Moreover, adding or subtracting transcripts from an existing panel is not straightforward. Finally, the hybridization workflows can be time-consuming.
Clustered regularly interspaced short palindromic repeats (CRISPR)/CRISPR-associated protein 9 (Cas9), CRISPR–Cas9 is a powerful genome editing and cloning tool [23, 24]. With high specificity, it is rapidly deployed to enrich specific genomic fragments prior sequencing for both research and clinical applications [25]. Examples include CATE-seq, CISMR [26], CRISPR-Cap [27], CRISPR-DS [28], Context-Seq [29], DASH [30], FLASH [31], and in vitro enChIP [32]. ONT long-read sequencing platform first adapted CRISPR–Cas9 to enrich chromosomal DNA and introduced CATCH (Cas9-assisted targeting of chromosome segments) sequencing [33]. Cas9 nuclease cleaves the regions of interest which are then separated by size selection using pulsed field gel electrophoresis (PFGE). However, this approach generates biases due to the follow-up PCR amplification of the target to compensate for the low recovery. Later, nanopore Cas9-targeted sequencing (nCAT) directly ligates adapters to Cas9-cleaved DNA products [34]. This method has successfully enriched several DNA targets. However, the potential of applying this method to enrich RNA transcripts has not yet been explored.
Here, we employed nanopore Cas9-guided targeted sequencing (nCAT) to enrich specific RNA transcripts within complex transcriptomes of different human cell lines with and without SARS-CoV-2-infections. We first validated the method on five human genes with a wide range of expression levels. We then extended the enrichment strategy to SARS-CoV-2 infected cells to characterize viral transcripts and identify the expression transcript isoforms of human cells with the infection.
Materials and methods
Cell culture and viral infection
For this study, all cell lines were maintained at 37°C with 5% CO2. The culture of A549 (human lung carcinoma epithelial) was performed in F12 media (Gibco) supplemented with 10% FCS, GlutaMAX, and sodium pyruvate. Cell culture of Caco-2 (human intestinal epithelial cells, ATCC Cat#HTB-37; RRID: CVCL_0025) has been described in previous studies [9, 35]. Briefly, Caco-2 cells were cultured in Dulbecco’s modified Eagle medium (DMEM) (Media Preparation Unit, Peter Doherty Institute) supplemented with 1× non-essential amino acids (Sigma–Aldrich), 20 mM HEPES, 2 mM L-glutamine, 1× GlutaMAX 2 μg/ml Fungizone solution, 26.6 μg/ml gentamicin, 100 IU/ml penicillin, 100 μg/ml streptomycin, and 20% FBS. For viral infection, the cell line was seeded in 4× six-well tissue culture plates and infected with SARS-CoV-2 (Australia/VIC01/2020) virus at a multiplicity of infection (MOI) of 0.1. The infected cells were harvested at 2-, 24-, and 48-h post-infection (hpi).
RNA extraction, cDNA synthesis, and purification
For this study, cultured cells were harvested by centrifugation. The cell pellets were then placed on dry ice to snap-freeze. These frozen cells were stored at −80°C before RNA extraction. RNA was extracted following manufacturer’s instructions for the RNeasy Mini Kit (QIAGEN). High integrity of extracted RNA (RIN score ≥ 9.0) was then treated with DNase from the Turbo DNA-free kit (Invitrogen) to reduce DNA contamination and purified by RNAclean XP magnetic beads (Beckman Coulter). Maxima H minus Double-Stranded cDNA synthesis (Thermoscientific – K2561) was used to synthesize cDNA for Cas9-targeted sequencing. The concentration and fragmentation of extracted RNA and synthesized cDNA were measured by Nanodrop 2000C (Thermo Fisher Scientific), 4200 TapeStation System (Agilent Technology), and Qubit 4 Flourometer (Invitrogen). RNA samples of control and infected cells harvested at 2, 24 and 48 hpi from Caco-2 were obtained from our previous study with similar sample preparation protocol [9].
Cas9-targeted sequencing
Design of gRNAs
All crRNAs were designed using CHOPCHOP (version 3) web tool to select target sites specific for CRISPR/Cas9 nanopore enrichment (https://chopchop.cbu.uib.no/about). Human genome (Homo sapiens – hg38/GRCh38) and SARS-CoV-2 genome (Australia/VIC01/2020) served as the references. The sequences of crRNAs are provided in Supplementary Table S2.
CHOPCHOP incorporates several methods to evaluate and calculate the cleavage efficiency by Cas9 nuclease. The default setting is “G20,” which prioritizes a guanine at position 20. CHOPCHOP ranks gRNAs according to: (i) efficiency score, (ii) number of off-targets and whether they have mismatches, (iii) existence of self-complementarity regions longer than 3 nt, (iv) GC-content: sgRNAs are most effective with a GC-content between 40% and 70%, and (v) location of sgRNA within a gene. To preserve nearly full-length transcripts, we prioritized the location of gRNAs over the cutting efficiency.
Cas9 cleaving and library preparation
The preparation for Cas9 nuclease cutting and library preparation for sequencing followed the recommended protocol from Oxford Nanopore Technologies with some modifications. In brief, all crRNAs were pooled and combined with tracrRNA (trans-activating crRNA) in Duplex Buffer to generate gRNAs. HiFi Cas9 Nuclease V3 (IDT) was mixed with gRNAs in 1× CutSmart buffer (NEB) to generate Cas9 ribonucleoprotein complex (RNP). Approximately 1 μg of pooled synthesized cDNA was dephosphorylated by Quick calf intestinal alkaline phosphatase (CIP). Dephosphorylated cDNA was cut by RNP and dA-tailing by NEB Tag polymerase, followed by library preparation. Ligation sequencing Kits, including SQK-LSK109 and SQK-LSK114, were utilized to prepare sequencing libraries for R9.4.1 and R10.4.1 flow cells, respectively. Samples were then run on a MinION or GridION sequencer. A detailed description of runs (flow cells, gRNA, sequencer) is provided in Supplementary Table S4.
Sequencing analysis
Basecalling and read alignment of nanopore sequencing data
For both R9.4.1 and R10.4.1 chemistry samples, Dorado (v0.7.0) was used to perform the base-calling. Model dna_r9.4.1_e8_sup@v3.6 was applied to R9.4.1 samples and model dna_r10.4.1_e8.2_400bps_sup@v5.0.0 was applied to R10.4.1 samples. A parameter “--min-qscore 7” was set to filter out low-quality reads. Samtools (v1.16.1) was used to convert the output ubam to fastq. Reads were mapped to human reference genome (Homo_sapiens GRCh38 release 100) and SARS-CoV-2 genome (SARS-CoV-2/human/AUS/VIC01/2020) using Minimap2 (v2.24) with parameter “-y --eqx -uf -ax splice -k 12” for human and “-aL -x map-ont” for SARS-CoV-2. On-target reads were defined as those that aligned within the start position and end position of a gene. Average coverage and read length analyses were obtained by Samtools.
For whole transcriptome direct cDNA and dRNA-seq, the data for A549, Caco-2 mock control, and Caco-2 infected with SARS-CoV-2 harvested after 2, 24, and 48 hpi was obtained from previous study [9]. Guppy (v3.5.2) was used to perform the basecalling with configuration file dna_r9.4.1_450bps_hac.cfg. Similarly, reads alignment, coverage, and read length analysis were performed using the same methods as previous analysis. Noting that 2 hpi is often incubation time for SARS-CoV-2 [36–38], there was no SARS-CoV-2 detected after 2 hpi in direct cDNA, dRNA-seq, and Cas9 enrichment (<0.00005%, Supplementary Table S4). Thus, we excluded 2 hpi from the main data analysis.
Transcript analysis
For human, gene-level read counts were generated by FeatureCounts [39] (v2.0.6). For SARS-CoV-2, our in-house tool, npTranscript (v1.0; https://github.com/lachlancoin/npTranscript), was utilized to assign reads to genome and sub-genome transcripts. Salmon [40] (v1.10.1) was used for isoform quantification with parameter “--noErrorModel -l U” with transcriptome reference (GENCODE release 35). For each sample, we calculated an on-target rate by dividing the number of reads mapped to target by the total number of reads mapped to all features of the reference genome. Fold enrichment on a given sample was calculated by dividing the on-target rate for Cas9-targeted sequencing by the on-target rate for the non-enrichment sample.
Modeling the impact of Cas9 cutting efficiency on enrichment
We calculated the correlation between CHOP CHOP estimated Cas9 cutting efficiency and observed log-fold change in expression estimated by FeatureCounts. Given the strong positive correlation, we then investigated whether we could improve our estimates of the unenriched CPM from enriched CPM based on Cas9 cutting efficiency. We fitted a linear model for log2-fold change based on CHOP CHOP predicted cutting efficiency, using leave-one-out cross validation (Supplementary Table S7). We utillized this model to predict fold change for the held-out sample, using the following formula: predicted fold Change = 2^(intercept + slope *efficiency), and then predicted unenriched_cpm = enriched_cpm/predicted fold change.
Statistics and illustration
All statistics and plots were performed by using GraphPad Prism (v10.4.1) for MacOS and RStudio (v2024.04.2).
Results
High enrichment of target transcripts by CRISPR–Cas9 targeted sequencing
To investigate the enrichment efficiency, we first selected five human genes encoded for FTL (ferritin light chain – Chromosome 19), TMSB10 (thymosin beta 10 – Chromosome 2), TMSB4X (thymosin beta 4 X-linked – Chromosome X), GAPDH (glyceraldehyde-3-phosphate dehydrogenase – Chromosome 12), and ALDH1A1 (aldehyde dehydrogenase – Chromosome 9), respectively. We classified these genes into low, medium, and high expression levels based on the published results from whole transcriptome direct cDNA sequencing in a well-characterized A549 cell line [9] (Fig. 1A and Supplementary Table S1). We designed all CRISPR RNAs (crRNAs) to bind to the first exon at 5′ ends or the last exon at 3′ ends and generate a single-end cut for ligating sequencing adapters (Fig. 1B, “Materials and methods”, Supplementary Table S2, Supplementary Fig. S1). crRNAs were pooled before Cas9 nuclease cutting and adaptor ligation for sequencing.
Figure 1.

Enrichment of human transcripts in A549 cell line. (A) Five target transcripts with a wide range of expression. (B) Cas9-targeted sequencing: cDNA is synthesized by using a poly(T) primer. gRNA binds to the target site allowing Cas9 nuclease cutting. Adapters are ligated to phosphate groups on both strands of the cDNA. Created in BioRender. Nguyen, A. (2026) https://BioRender.com/cyh6hg6. (C) The proportion of reads in each transcript in enriched and non-enriched cDNA sequencing experiments. An inset shows the total target proportion. Cas9-targeted sequencing significantly increases the proportion of target transcripts among the transcriptome (two-tailed t-test, P = .0002) (Supplementary Table S3). (D) Alignment coverage of each transcript. All transcripts have even coverage in all known exons. Dashed red lines indicate gRNA binding sites. (E) Read counts at the gene level of all mapped reads in enrichment and non-enrichment sequencing. Each dot represents the read count of specific transcript. Columns represent human chromosomes. Cas9-targeted sequencing reduces sequencing noise compared with whole transcriptome sequencing and promotes the identification of target transcripts. (F) Highly correlated expression between enrichment and non-enrichment sequencing at gene level (Pearson’s correlation). (G) 30 highly abundant transcripts in enriched and unenriched samples. Red color indicates off-targets, while blue color represents on-target transcripts.
We compared our enrichment data with the published direct cDNA sequencing of the same cell lines [9]. Our approach achieved significantly higher coverage across all on-target transcripts by using ~1 μg of poly(T)28 synthesized cDNA samples and running on MinION flow cells (both R9.4.1 and R10.4.1 versions) (two-tailed t-test, P = .0002) (Fig. 1C, “Materials and Methods”, Supplementary Table S3). On average, target mapping coverage ranged from 300× to 2200× (Supplementary Table S4). We then assigned all mapped reads to genomic features and evaluated the expression of each target at the gene level. In total, >50% of feature-assigned reads were on-target in the target enriched sequencing, compared to ~2% in unenriched sequencing (Supplementary Table S3). The gene exhibiting the highest fold change of coverage was GAPDH (62-fold change), while TMSB10 showed the lowest enrichment with a 23-fold change (Supplementary Table S3). Reference lengths of the targets range from 461 bp to 2096 bp, which is slightly longer than the majority read length generated by our method. Despite this, our method generated even depth for all exons of the targets (Fig. 1D and Supplementary Fig. S1). In each target chromosome, Cas9-guided enrichment amplified the targets while reducing noise caused by other high-expression genes, which were abundantly found in unenriched sequencing (Fig. 1E). The expression of the targets was highly correlated with the estimated abundance by whole transcriptome cDNA sequencing (Pearson’s correlation of 0.93, Fig. 1F).
Next, we investigated off-target reads in our experiment by extracting the 30 most abundant transcripts. We considered two possibilities: if most of the off-targets were found in both enrichment and non-enrichment methods and expressed at high levels, this would indicate the random ligation of the adapters; on the other hand, we would find the enrichment of the off-targets, which are only available in Cas9-enrichment data. We found that 13 out of 25 (52%) of the off-targets were also present among the most abundant transcripts in non-enrichment sequencing data (Fig. 1G and Supplementary Table S5). This suggests that off-targets are primarily due to random ligation, which is similar to that observed in Cas9-targeted sequencing to enrich chromosomal DNA [34].
Cas9-targeted sequencing amplifies target signals for isoform identification
Since we obtained substantial high depth for target transcripts in the previous experiment, here, we applied the Cas9-targeted approach to focus on mRNA with well-known isoforms. KPNA2 (karyopherin subunit alpha 2) is a member of the nuclear transporter family and is involved in the delivery of numerous cargo proteins from the cytoplasm to the nucleus, which has been shown to contribute to cancer development [41, 42]. Several predicted isoforms have been reported in which ENST00000300459 (KPNA2-201, 1977 nucleotides, GC = 44.76%) and ENST00000537025 (KPNA2-202, 2693 nucleotides, GC = 47.95%) were highly expressed and showed differential transcript usage during viral infection in specific cell types [9]. The two variants are distinct in their lengths, particularly in the first exon (Fig. 2A). Thus, we designed a gRNA binding to the second exon, which generates two fragments, a long fragment including a portion of exon 2 and exon 3–11; and a short fragment including exon 1 and a part of exon 2 (Fig. 2A).
Figure 2.

Identification of KPNA2 isoforms using targeted sequencing. (A) Two highly expressed transcripts in human cell lines. (B) Gene level counts for KPNA2 in enrichment and non-enrichment samples. We obtained higher counts for enrichment samples compared to whole transcriptome cDNA and dRNA samples. (C) Isoform expression in infected and non-infected cells. There were no differences in isoform abundance between infected and non-infected cells. (D) IGV (Integrative genomics viewer) screenshot displaying mapping coverage of KPNA2. Reads were colored by read strand, blue for negative strand, pink for positive strand. Red arrow indicates the Cas9 nuclease cutting site.
We first re-analyzed whole transcriptome data generated by both direct cDNA sequencing and dRNA-seq from previous study, which utilized the same A549 and Caco-2 cell lines [9, 35] (Materials and methods). In these cell lines, the KPNA2 gene is expressed at comparably low levels. Particularly, at gene level, only 0.09% of reads mapped to this transcript in the A549 sample. Similarly, Caco-2 control and Caco-2 infected with SARS-CoV-2 samples expressed the KPNA2 gene at 0.03%. By using the Cas9-targeted system, we obtained a significant increase in KPNA2 coverage (two-tailed t-test, P = .0008) (Fig. 2B)—an average of 62-fold higher coverage of KPNA2 transcript compared to unenriched samples.
To identify the proportion of the isoforms in these samples, we quantified the isoform expression using Salmon (Materials and methods). In alignment with previous study [9], we found that KPNA2-201 is dominantly detected in both Caco-2 and A549 cell lines, and the relative proportions of these isoforms were preserved across enrichment and non-enrichment methods (Fig. 2C and Supplementary Table S6). Interestingly, we were able to detect the reads that covered exon 1 of KPNA2-202, which is rarely seen in cDNA sequencing and full-length dRNA-seq (Fig. 2D and Supplementary Fig. S2).
Enrichment of SARS-CoV-2 RNA improves the identification of diverse viral sub-genomes in infected cells
Viral infections, such as those with SARS-CoV-2, cause a broad spectrum of disease states with widely variable severity and symptoms. Deep genomic sequencing is critical for early treatment and understanding the viral distribution and circulation in vivo. However, it is difficult to obtain high-depth sequencing viral due to the high proportion of host RNA and the potential for RNA degradation. Thus, we applied Cas9 enrichment to target both human and viral transcripts. The SARS-CoV-2 genome is a singular, non-segmented large positive sense-stranded RNA with a length of ~30 kb. From the genome, a set of plus-strand sub-genomic RNAs (sgRNAs) are generated with a common leader sequence of 70 nucleotides located at 5′-end and a section at the 3′ end of the genome [43]. These sgRNAs encode the following viral proteins: structural proteins—spike (S), envelope (E), membrane (M) and nucleocapsid protein (N)—and several accessory proteins—3a, 6, 7a, 7b, 8, and 10. To determine the efficiency of Cas9-targeted sequencing enriching SARS-CoV-2 transcriptome in infected samples, we designed a crRNA that targets the 5′-end of the leader sequence to capture all possible viral structural transcripts and accessory transcripts (Fig. 3A, Materials and methods). We then compared our enrichment data with data generated from whole transcriptome dRNA-seq and cDNA sequencing of the same infected samples.
Figure 3.

Enrichment of SARS-CoV-2 transcripts in human-infected cells. (A) SARS-CoV-2 genome alignment. Mapped reads covered all ORFs with substantial reads mapping at 5′ leader sequence and all sub-genome ORFs: enriched (blue), unenriched: cDNA (orange) and dRNA-seq (red). (B) Proportion of human reads and viral reads. While unenriched samples maintained high human proportions, the enrichment method increased the proportion of viral reads. (C) The abundance of different classes of sgRNA in enriched and unenriched samples of infected cells at different time points. Cas9 enrichment method recovered all transcripts with higher coverage compared to whole transcriptome cDNA and dRNA sequencing. (D) Highly correlated abundance of sub-genome transcripts between Cas9 enrichment and non-enrichment (Pearson’s correlation).
We first mapped all enriched reads to the SARS-CoV-2 and human genomes to identify the fraction of viral reads in the samples (Materials and methods). We were able to cover the whole genome of SARS-CoV-2 (Fig. 3A). Particularly, in Caco-2 infected cells harvested at 48 hpi, the median read length was 4972 nt, and 63 reads were mapped to the genome with lengths longer than 25 kb. We observed high coverage for the region of the leader sequence and at the 3′ end of the viral genome (Fig. 3A). Moreover, we noted a difference in the length of the leader sequence between enriched and unenriched samples due to the cutting of Cas9 nuclease. There was uniform coverage across the body of each annotated open reading frame (ORF), with a sharp drop at the end of the ORF. This observation is consistent with the previous studies and supports the accepted model of SARS-CoV-2 discontinuous transcription, where most of the viral sgRNAs are the result of the fusion of an ORF with the 5′ leader sequence [43].
We then identified the proportion of viral reads compared to human reads. We found a significant increase in viral reads compared to non-enrichment methods (two-tailed t-test, P = .0156). Particularly at 24 hpi, we observed >40% of reads mapping to SARS-CoV-2 after enrichment, compared to only 2% without enrichment (20-fold increase) (Fig. 3B). We did not detect any viral reads in the control sample using Cas9 enrichment (Supplementary Table S4).
Next, we assigned all mapped reads to specific viral sgRNAs to examine the transcript expression (Materials and methods). We obtained significantly higher read counts for each sgRNA compared to non-enrichment samples and covered all sgRNAs that were found in both direct cDNA sequencing and dRNA-seq (Fig. 3C). Our enrichment data was highly correlated with whole-transcriptomic sequencing data (Fig. 3D) and exhibited a similar expression pattern of structural and accessory proteins with nucleocapsid N as the most highly expressed (Fig. 3C).
Discussion
Full-length RNA sequencing facilitates the identification of transcript isoforms. However, the limitation is that low-abundance transcripts are often missed due to the lack of depth. Targeted transcript long-read sequencing leverages the ability to sequence transcript molecules end-to-end, while circumventing the weakness of limited sequencing yield and low transcript coverage. Although there are existing solutions for targeted long-read RNA sequencing, these are either associated with high costs; or are complicated to set up and implement. Here, we employed targeted CRISPR–Cas9 (nCAT) sequencing strategy to enrich the preselected transcripts within transcriptomes. We tested the method in several cell lines with a wide range of transcript abundance. Our comprehensive benchmark analysis indicated consistently high on-target rate and fold enrichment across all samples. We showed that Cas9-targeted sequencing substantially improved the sensitivity of detecting target transcripts as well as the identification of mRNA isoforms. At the same time, the estimated abundances of target transcripts correlated highly with the whole transcriptome cDNA sequencing and dRNA-seq. We also showed that the Cas9-targeted sequencing was able to enrich transcripts across species by testing the method on SARS-CoV-2-infected cell lines at different time points. This targeted sequencing strategy provided clear evidence of the expression of viral accessory transcripts, which were missed in whole transcriptome cDNA sequencing. Overall, these results indicate that Cas9-targeted sequencing provides a robust tool for the quantification of specific transcripts as well as isoform confirmation and possibly new isoform discovery.
Several approaches have successfully adapted CRISPR–Cas9 for enriching genomic sequences, following up with short-read and/or long-read sequencing [27–32]. Advantages and disadvantages of these methods have been reviewed elsewhere [25]. At the transcript level, however, the use of CRISPR–Cas9 for targeted sequencing that provides full-length reads with significant depth remains scarce. An alternative method such as TEQUILA-seq [21], which uses customized biotinylated capture oligos, has solved the problems of primer cross-binding as well as reduced the amplification bias associated with long-range PCR-based methods. Moreover, this strategy significantly reduces the cost and improves efficiency compared to earlier methods in the same category (e.g. CLS [22], ORF Capture-Seq [20]). TEQUILA-seq achieved on-target rate of >80% with ~280-fold enrichment, which is higher than our targeted Cas9 sequencing (∼60% and the highest fold change of 62). However, it is noted that TEQUILA-seq still involves two PCR steps before and after capture. In addition, a significant number of long probes (120–150 nt) are required to be able to cover all annotated exons of a gene. In contrast, our CRISPR–Cas9 strategy is PCR-free and requires a minimal number of gRNAs; in this study, only one gRNA per transcript was used.
CRISPR–Cas systems have been extensively utilized, developed, and commercialized to detect human pathogenic DNA and RNA viruses, including SARS-CoV-2 [44–46], influenza [47], HPV [48], and Zika virus [49]. These approaches offer flexibility, sensitivity and applicable to clinical samples. However, their broader application for investigating viral genomic and sub-genomic structures, as well as viral expression profiles during pandemics or population-scale surveillance, remains limited by the need for amplification steps and the inability to integrate next-generation sequencing technologies [50, 51]. Here, we demonstrate that nCAT efficiently enriches SARS-CoV-2 from samples containing abundance of host materials, reduces sequencing noise, and enables analyses to characterize viral genomic architecture and expression patterns (Fig. 3). Furthermore, the ease of customizing and multiplexing facilitates not only viral detection but also the discovery of virus-host interactions.
Our work has several limitations that warrant further studies. First, the major challenge of Nanopore long-read sequencing platform is the requirement of significant amounts of purified RNA/DNA input. While this issue is often manageable for culturable samples, it limits ONT’s application on rare and difficult-to-obtain specimens, such as clinical specimens or environmental samples, particularly when PCR amplification is avoided. In this study, we utilized a minimum recommended concentration of ~1 μg of cDNA to guarantee the sequencing depth and coverage for each target transcript. However, depending on number of targets and their abundance, additional optimization studies using lower input are required. Second, although we validated the efficiency of Cas9-targeting sequencing across several transcripts with varying expression levels, we did not investigate rare transcripts with TPM < 10. In this study, KPNA2, which has been reported in several malignancies, including cervical cancer [52], liver cancer [53], ovarian cancer [54], and prostate cancer [55], is expressed as two isoforms, KPNA2-201 and KPNA2-202. The expression of KPNA2 gene is > 10 TPM; however, the expression level of each isoform is considerably low [~10 TPM in direct cDNA sequencing (Supplementary Table S6)]. Using our approach, we achieved the fold-change of >60 times for each isoform compared to cDNA sequencing, highlighting the potential of this method for detection and validation of low- and rare-abundance transcripts. Third, the average read length generated by nCAT is shorter than that obtained from dRNA or cDNA sequencing (Supplementary Fig. S3). This limitation arises from the CRISPR–Cas mechanism, which relies on the binding of gRNA and Cas9-mediated cleavage of the target sequence to generate a free 5′ phosphate group for adapter ligation. To preserve nearly full-length transcripts, gRNA’s binding sites should be ideally designed within the first exon. However, this strategy may prevent the enrichment if the alternative RNA processing generated isoforms lacking the first exon, as observed for KPNA2. In such cases, gRNA binding sites located in the middle exons are preferable. When transcript structures are unknown, multiple binding sites should be examined. However, combining multiple gRNAs for one target in a single reaction could cause fragmentation, thereby complicating downstream analysis. Lastly, while transcript characteristics such as GC content, transcript length, and abundance are not the main factors causing fold-change variation, we found a positive correlation between fold-change and CHOPCHOP’s efficiency score (Supplementary Table S7 and Supplementary Fig. S4). Since fold-change and efficiency score are highly correlated, a desirable feature for an enrichment strategy is the ability to predict the abundance before enrichment (i.e. in the primary sample) from the Cas9-enriched CPM, or vice versa. We found that by modelling the effect of predicted cutting efficiency on log fold change, we could predict the original abundance from Cas9-enriched abundance with a squared correlation coefficient of 0.98 (Supplementary Fig. S5, using leave-one-out cross validation to avoid overfitting). Future studies with more targets will be able to refine the relationship between Cas9 cutting efficiency and abundance fold change.
In summary, Cas9-targeted sequencing enables the enrichment of specific transcripts within complex transcriptomes. Here, we illustrated the application of the Cas9 method to enrich a small scale of gene panels and only focused on human and viral transcripts. However, Cas-9 targeted sequencing is customizable to be applied to any gene or set of genes of interest for focused discovery and quantification of transcript isoforms. We anticipate that the strategy for applying Cas9-targeted sequencing will have a broad impact, facilitate the significant cost reduction as well as straightforward implementation, and therefore will pave the way for wide adaptation in diverse biomedical and clinical settings.
Supplementary Material
Acknowledgements
The graphical abstract was created in BioRender.com. Nguyen, A. (2026) https://BioRender.com/sj281mr
Author contributions: An N.T. Nguyen (Formal analysis [equal], Investigation [equal], Methodology [equal], Resources [equal], Validation [equal], Writing – original draft [lead]), Jianshu Zhang (Data curation [equal], Formal analysis [equal], Resources [equal], Software [equal], Writing – review & editing [equal]), Shuxin Zhang (Formal analysis [equal], Methodology [equal], Writing – review & editing [equal]), Miranda E. Pitt (Formal analysis [equal], Investigation [equal], Methodology [equal], Writing – review & editing [equal]), Devika Ganesamoorthy (Conceptualization [equal], Methodology [equal], Writing – review & editing [equal]), Svenja Fritzlar (Formal analysis [equal], Writing – review & editing [equal]), Jessie J.-Y. Chang (Formal analysis [equal], Investigation [equal], Methodology [equal], Writing – review & editing [equal]), Sarah L. Londrigan (Supervision [equal], Writing – review & editing [equal]) and Lachlan J.M. Coin (Supervision [lead], Writing – review & editing [equal], Resources [equal], Conceptualization [equal], Methodology [equal], Software [equal]).
Contributor Information
An N T Nguyen, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Jianshu Zhang, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Shuxin Zhang, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Miranda E Pitt, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia; Australian Institute for Microbiology and Infection, University of Technology Sydney, Ultimo, New South Wales 2007, Australia.
Devika Ganesamoorthy, Children’s Intensive Care Research Program, Child Health Research Centre, The University of Queensland, Brisbane 4101, Australia.
Svenja Fritzlar, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Jessie J -Y Chang, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Sarah L Londrigan, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Lachlan J M Coin, Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, Victoria 3000, Australia.
Supplementary data
Supplementary data is available at NAR Genomics & Bioinformatics online.
Conflict of interest
None declared.
Funding
This project was funded by GNT1195743 from National Health and Medical Research Council (LJMC) and DP260104501 from Australian Research Council.
Data availability
Sequencing data for this study can be found in NCBI with accession number PRJNA1244050. Data for direct cDNA sequencing and dRNA-seq obtained from previous studies can be found in online repositories: 1. https://www.ncbi.nlm.nih.gov/, PRJNA675370. 2. https://figshare.com/10.6084/m9.figshare.17139995. 3. https://figshare.com/10.6084/m9.figshare.16841794. 4. https://figshare.com/10.6084/m9.figshare.17140007. Our in-house tool, npTranscript (v1.0), is available on GitHub (https://github.com/lachlancoin/npTranscript) and Figshare (https://doi.org/10.26188/33256032).
References
- 1. Barbosa-Morais NL, Irimia M, Pan Q et al. The evolutionary landscape of alternative splicing in vertebrate species. Science. 2012;338:1587–93. 10.1126/science.1230612 [DOI] [PubMed] [Google Scholar]
- 2. Melé M, Ferreira PG, Reverter F et al. The human transcriptome across tissues and individuals. Science. 2015;348:660–5. 10.1126/science.aaa0355 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Zheng S, Black DL. Alternative pre-mRNA splicing in neurons: growing up and extending its reach. Trends Genet. 2013;29:442–8. 10.1016/j.tig.2013.04.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Pan Q, Shai O, Lee LJ et al. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing. Nat Genet. 2008;40:1413–5. 10.1038/ng.259 [DOI] [PubMed] [Google Scholar]
- 5. Floor SN, Doudna JA. Tunable protein synthesis by transcript isoforms in human cells. eLife. 2016;5:e10921. 10.7554/eLife.10921 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Scotti MM, Swanson MS. RNA mis-splicing in disease. Nat Rev Genet. 2016;17:19–32. 10.1038/nrg.2015.3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. van den Hoogenhof MM, Pinto YM, Creemers EE. RNA splicing: regulation and dysregulation in the heart. Circ Res. 2016;118:454–68. 10.1161/circresaha.115.307872 [DOI] [PubMed] [Google Scholar]
- 8. Lara-Pezzi E, Gómez-Salinero J, Gatto A et al. The alternative heart: impact of alternative splicing in heart disease. J of Cardiovasc Trans Res. 2013;6:945–55. 10.1007/s12265-013-9482-z [DOI] [PubMed] [Google Scholar]
- 9. Chang JJ-Y, Gleeson J, Rawlinson D et al. Long-read RNA sequencing identifies polyadenylation elongation and differential transcript usage of host transcripts during SARS-CoV-2 in vitro infection. Front Immunol. 2022;13:832223. 10.3389/fimmu.2022.832223 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Maher CA, Kumar-Sinha C, Cao X et al. Transcriptome sequencing to detect gene fusions in cancer. Nature. 2009;458:97–101. 10.1038/nature07638 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Tang AD, Soulette CM, van Baren MJ et al. Full-length transcript characterization of SF3B1 mutation in chronic lymphocytic leukemia reveals downregulation of retained introns. Nat Commun. 2020;11:1438. 10.1038/s41467-020-15171-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Wenger AM, Peluso P, Rowell WJ et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol. 2019;37:1155–62. 10.1038/s41587-019-0217-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Stark R, Grzelak M, Hadfield J. RNA sequencing: the teenage years. Nat Rev Genet. 2019;20:631–56. 10.1038/s41576-019-0150-2 [DOI] [PubMed] [Google Scholar]
- 14. Eid J, Fehr A, Gray J et al. Real-time DNA sequencing from single polymerase molecules. Science. 2009;323:133–8. 10.1126/science.1162986 [DOI] [PubMed] [Google Scholar]
- 15. Mikheyev AS, Tin MMY. A first look at the Oxford Nanopore MinION sequencer. Mol Ecol Resour. 2014;14:1097–102. 10.1111/1755-0998.12324 [DOI] [PubMed] [Google Scholar]
- 16. Pollard MO, Gurdasani D, Mentzer AJ et al. Long reads: their purpose and place. Hum Mol Genet. 2018;27:R234–41. 10.1093/hmg/ddy177 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Amarasinghe SL, Su S, Dong X et al. Opportunities and challenges in long-read sequencing data analysis. Genome Biol. 2020;21:30. 10.1186/s13059-020-1935-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Jain M, Abu-Shumays R, Olsen HE et al. Advances in nanopore direct RNA sequencing. Nat Methods. 2022;19:1160–4. 10.1038/s41592-022-01633-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Clark MB, Wrzesinski T, Garcia AB et al. Long-read sequencing reveals the complex splicing profile of the psychiatric risk gene CACNA1C in human brain. Mol Psychiatry. 2020;25:37–47. 10.1038/s41380-019-0583-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Sheynkman GM, Tuttle KS, Laval F et al. ORF Capture-Seq as a versatile method for targeted identification of full-length isoforms. Nat Commun. 2020;11:2326. 10.1038/s41467-020-16174-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Wang F, Xu Y, Wang R et al. TEQUILA-seq: a versatile and low-cost method for targeted long-read RNA sequencing. Nat Commun. 2023;14:4760. 10.1038/s41467-023-40083-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Lagarde J, Uszczynska-Ratajczak B, Carbonell S et al. High-throughput annotation of full-length long noncoding RNAs with capture long-read sequencing. Nat Genet. 2017;49:1731–40. 10.1038/ng.3988 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Cox DBT, Platt RJ, Zhang F. Therapeutic genome editing: prospects and challenges. Nat Med. 2015;21:121–31. 10.1038/nm.3793 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Li T, Yang Y, Qi H et al. CRISPR/Cas9 therapeutics: progress and prospects. Sig Transduct Target Ther. 2023;8:36. 10.1038/s41392-023-01309-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Malekshoar M, Azimi SA, Kaki A et al. CRISPR–Cas9 targeted enrichment and next-generation sequencing for mutation detection. J Mol Diagn. 2023;25:249–62. 10.1016/j.jmoldx.2023.01.010 [DOI] [PubMed] [Google Scholar]
- 26. PE B-B, Mueller JL. CRISPR-mediated isolation of specific megabase segments of genomic DNA. Nucleic Acids Res. 2017;45:e165. 10.1093/nar/gkx749 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Lee J, Lim H, Jang H et al. CRISPR-Cap: multiplexed double-stranded DNA enrichment based on the CRISPR system. Nucleic Acids Res. 2019;47:e1. 10.1093/nar/gky820 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Nachmanson D, Lian S, Schmidt EK et al. Targeted genome fragmentation with CRISPR/Cas9 enables fast and efficient enrichment of small genomic regions and ultra-accurate sequencing with low DNA input (CRISPR-DS). Genome Res. 2018;28:1589–99. 10.1101/gr.235291.118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Fuhrmeister ER, Kim S, Mairal SA et al. Context-Seq: CRISPR–Cas9 targeted nanopore sequencing for transmission dynamics of antimicrobial resistance. Nat Commun. 2025;16:5898. 10.1038/s41467-025-60491-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Gu W, Crawford ED, O’Donovan BD et al. Depletion of abundant sequences by hybridization (DASH): using Cas9 to remove unwanted high-abundance species in sequencing libraries and molecular counting applications. Genome Biol. 2016;17:41. 10.1186/s13059-016-0904-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Quan J, Langelier C, Kuchta A et al. FLASH: a next-generation CRISPR diagnostic for multiplexed detection of antimicrobial resistance sequences. Nucleic Acids Res. 2019;47:e83. 10.1093/nar/gkz418 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Fujita T, Fujii H. Efficient isolation of specific genomic regions and identification of associated proteins by engineered DNA-binding molecule-mediated chromatin immunoprecipitation (enChIP) using CRISPR. Biochem Biophys Res Commun. 2013;439:132–6. 10.1016/j.bbrc.2013.08.013 [DOI] [PubMed] [Google Scholar]
- 33. Gabrieli T, Sharim H, Fridman D et al. Selective nanopore sequencing of human BRCA1 by Cas9-assisted targeting of chromosome segments (CATCH). Nucleic Acids Res. 2018;46:e87. 10.1093/nar/gky411 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Gilpatrick T, Lee I, Graham JE et al. Targeted nanopore sequencing with Cas9-guided adapter ligation. Nat Biotechnol. 2020;38:433–8. 10.1038/s41587-020-0407-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Chang JJY, Rawlinson D, Pitt ME et al. Transcriptional and epi-transcriptional dynamics of SARS-CoV-2 during cellular infection. Cell Rep. 2021;35:109108. 10.1016/j.celrep.2021.109108 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Mautner L, Hoyos M, Dangel A et al. Replication kinetics and infectivity of SARS-CoV-2 variants of concern in common cell culture models. Virol J. 2022;19:76. 10.1186/s12985-022-01802-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Bojkova D, Klann K, Koch B et al. Proteomics of SARS-CoV-2-infected host cells reveals therapy targets. Nature. 2020;583:469–72. 10.1038/s41586-020-2332-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Chu H, Chan JF, Yuen TT et al. Comparative tropism, replication kinetics, and cell damage profiling of SARS-CoV-2 and SARS-CoV with implications for clinical manifestations, transmissibility, and laboratory studies of COVID-19: an observational study. Lancet Microbe. 2020;1:e14–23. 10.1016/S2666-5247(20)30004-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Liao Y, Smyth GK, Shi W featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30:923–30. 10.1093/bioinformatics/btt656 [DOI] [PubMed] [Google Scholar]
- 40. Patro R, Duggal G, Love MI et al. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14:417–9. 10.1038/nmeth.4197 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Liao WC, Lin TJ, Liu YC et al. Nuclear accumulation of KPNA2 impacts radioresistance through positive regulation of the PLSCR1-STAT1 loop in lung adenocarcinoma. Cancer Sci. 2022;113:205–20. 10.1111/cas.15197 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Pasha T, Zatorska A, Sharipov D et al. Karyopherin abnormalities in neurodegenerative proteinopathies. Brain. 2021;144:2915–32. 10.1093/brain/awab201 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Sola I, Almazán F, Zúñiga S et al. Continuous and discontinuous RNA synthesis in coronaviruses. Annu Rev Virol. 2015;2:265–88. 10.1146/annurev-virology-100114-055218 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Patchsung M, Jantarug K, Pattama A et al. Clinical validation of a Cas13-based assay for the detection of SARS-CoV-2 RNA. Nat Biomed Eng. 2020;4:1140–9. 10.1038/s41551-020-00603-x [DOI] [PubMed] [Google Scholar]
- 45. Kumar M, Gulati S, Ansari AH et al. FnCas9-based CRISPR diagnostic for rapid and accurate detection of major SARS-CoV-2 variants on a paper strip. eLife. 2021;10:e67130. 10.7554/eLife.67130 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Ali Z, Sanchez E, Tehseen M et al. Bio-SCAN: a CRISPR/dCas9-based lateral flow assay for rapid, specific, and sensitive detection of SARS-CoV-2. ACS Synth Biol. 2022;11:406–19. 10.1021/acssynbio.1c00499 [DOI] [PubMed] [Google Scholar]
- 47. Freije CA, Myhrvold C, Boehm CK et al. Programmable inhibition and detection of RNA viruses using Cas13. Mol Cell. 2019;76:826–37. 10.1016/j.molcel.2019.09.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Zhang B, Wang Q, Xu X et al. Detection of target DNA with a novel Cas9/sgRNAs-associated reverse PCR (CARP) technique. Anal Bioanal Chem. 2018;410:2889–900. 10.1007/s00216-018-0873-5 [DOI] [PubMed] [Google Scholar]
- 49. Myhrvold C, Freije CA, Gootenberg JS et al. Field-deployable viral diagnostics using CRISPR–Cas13. Science. 2018;360:444–8. 10.1126/science.aas8836 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Das PK, Aich M, Adu P et al. CRISPR-cas diagnostics (CRISPR-Dx) of viral pathogens in low and limited resource areas: current state and future directions. TrAC, Trends Anal Chem. 2026;195:118590. 10.1016/j.trac.2025.118590 [DOI] [Google Scholar]
- 51. Lamb CH, Kang B, Myhrvold C. Multiplexed CRISPR-based methods for pathogen nucleic acid detection. Curr Opin Biomed Eng. 2023;27:100471. 10.1016/j.cobme.2023.100471 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. van der Watt PJ, Maske CP, Hendricks DT et al. The karyopherin proteins, Crm1 and karyopherin beta1, are overexpressed in cervical cancer and are critical for cancer cell survival and proliferation. Int J Cancer. 2009;124:1829–40. 10.1002/ijc.24146 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Yoshitake K, Tanaka S, Mogushi K et al. Importin-alpha1 as a novel prognostic target for hepatocellular carcinoma. Ann Surg Oncol. 2011;18:2093–103. 10.1245/s10434-011-1569-7 [DOI] [PubMed] [Google Scholar]
- 54. Zheng M, Tang L, Huang L et al. Overexpression of karyopherin-2 in epithelial ovarian cancer and correlation with poor prognosis. Obstet Gynecol. 2010;116:884–91. 10.1097/AOG.0b013e3181f104ce [DOI] [PubMed] [Google Scholar]
- 55. Mortezavi A, Hermanns T, Seifert HH et al. KPNA2 expression is an independent adverse predictor of biochemical recurrence after radical prostatectomy. Clin Cancer Res. 2011;17:1111–21. 10.1158/1078-0432.CCR-10-0081 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Sequencing data for this study can be found in NCBI with accession number PRJNA1244050. Data for direct cDNA sequencing and dRNA-seq obtained from previous studies can be found in online repositories: 1. https://www.ncbi.nlm.nih.gov/, PRJNA675370. 2. https://figshare.com/10.6084/m9.figshare.17139995. 3. https://figshare.com/10.6084/m9.figshare.16841794. 4. https://figshare.com/10.6084/m9.figshare.17140007. Our in-house tool, npTranscript (v1.0), is available on GitHub (https://github.com/lachlancoin/npTranscript) and Figshare (https://doi.org/10.26188/33256032).
