Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Sep 16.
Published in final edited form as: Cell Rep. 2025 Jul 25;44(8):116075. doi: 10.1016/j.celrep.2025.116075

SUGP1 loss drives SF3B1 hotspot mutant missplicing in cancer

Peiqi Xing 1,2,5, Pedro Bak-Gordon 3,5, Jindou Xie 1,2,4,5, Jian Zhang 3,5, Zhaoqi Liu 1,2,4,*, James L Manley 3,6,*
PMCID: PMC12434867  NIHMSID: NIHMS2107270  PMID: 40714635

SUMMARY

SF3B1 is the most frequently mutated splicing factor in cancer. Such mutations cause missplicing by promoting aberrant 3′ splice site usage; however, how this occurs mechanistically remains controversial. To address this issue, we employed a computational screen of 600 splicing-related proteins to identify those whose reduced expression recapitulates mutant SF3B1-induced splicing dysregulation. Strikingly, our analysis reveals only two proteins whose knockdown or knockout reproduces this effect. Extending our previous findings, loss of the G-patch protein SUGP1 recapitulates almost all splicing defects induced by SF3B1 hotspot mutations. Unexpectedly, loss of the RNA helicase Aquarius (AQR) reproduces ∼40% of these defects. However, we find that AQR knockdown causes significant SUGP1 missplicing and reduced SUGP1 levels, suggesting that AQR loss reproduces mutant SF3B1 splicing defects only indirectly. This study advances our understanding of missplicing caused by oncogenic SF3B1 mutations and highlights the fundamental role of SUGP1 in this process.

Graphical Abstract

graphic file with name nihms-2107270-f0001.jpg

In brief

Xing et al. report the first high-confidence list of cryptic 3′ss used by SF3B1-mutant cells. To address the underlying mechanism, they analyzed whether knockout/knockdown of any of 600 RBPs recapitulated these events. Strikingly, only one, SUGP1, did so, implicating SUGP1 loss from mutant spliceosomes as the sole driver of missplicing.

INTRODUCTION

RNA splicing entails a two-step reaction by which non-coding introns are excised from precursor messenger RNAs (pre-mRNAs) and the remaining exons are ligated to generate mature mRNA transcripts.1 More than 95% of human transcripts are subject to the process of alternative splicing, whereby pre-mRNAs are differentially spliced to generate multiple proteins with functional diversity.2,3 Splicing occurs in a large and dynamic structure called the spliceosome, consisting of over 100 proteins and five small nuclear RNAs (snRNAs) assembled into small nuclear ribonucleoproteins (snRNPs).4 As a fundamental cellular process, dysregulation of mRNA splicing has been implicated in multiple diseases, including cancer, in which it can contribute to tumor initiation, progression, and treatment resistance.5

Multiple genes encoding splicing factors are recurrently mutated in cancer, highlighting splicing dysregulation as a genetic driver of oncogenesis.6 Among these, mutations in the gene encoding the core splicing factor SF3B1 are the most common, frequently occurring as “hotspot” heterozygous point mutations affecting specific amino acid residues concentrated in the HEAT repeat domain (SF3B1HEAT), specifically repeats H4–H7.79 SF3B1 hotspot mutations promote the recognition of aberrant branchpoints, resulting in increased usage of cryptic 3′ splice sites (3′ss) typically located 10–30 nucleotides (nt) upstream from canonical 3′ss.1012 While other types of missplicing have been reported,13 these constitute a minority, and cryptic 3′ss are the most prevalent. This change in splicing is consistent with the role of SF3B1 as a core component of the U2 snRNP complex, which is responsible for branch site (BS) recognition and spliceosome assembly during the early stages of splicing.10 Functional studies of the downstream splicing defects arising in SF3B1-mutant cancers, including myelodysplastic syndromes (MDS),1416 chronic lymphocytic leukemia (CLL),17 uveal melanoma (UVM),13 and breast18 and pancreatic cancer,19 have elucidated specific aberrant splicing events critical for the establishment and maintenance of these cancers. However, the precise mechanism(s) by which SF3B1 mutations lead to cryptic alternative 3′ss selection and the key interacting partner proteins involved in this process remain controversial.

Previous systematic analyses of RNA splicing profiles across The Cancer Genome Atlas (TCGA) database revealed that somatic mutations in SUGP1 (SURP and G-patch domain containing 1) mimic the mutant SF3B1 splicing dysregulation.20,21 Importantly, these findings are consistent with earlier studies showing that SF3B1 hotspot mutations disrupt an interaction between SF3B1 and SUGP1 that is necessary for correct BS recognition during splicing and that SUGP1 knockdown can recapitulate splicing defects observed in SF3B1 mutant cells.22 Subsequently, the regions of SUGP1 encompassing the previously identified cancer mutations, which flank its G-patch domain, were found to interact directly with SF3B1 HEAT repeats harboring cancer-associated hotspot mutations, resulting in a conformational change of SUGP1 that exposes its G-patch domain to interact with and thereby activate the DEAH box helicase DHX15.23,24 Thus, cancer-associated hotspot mutations in SF3B1 or SUGP1 prevent activation of DHX15, which in turn leads to aberrant use of upstream branch points and cryptic 3′ss characteristic of tumors with mutant SF3B1. However, whether this mechanism entirely explains mutant SF3B1 missplicing remains unclear.

Recent studies have suggested that a number of splicing factors lacking known mutations in cancer are potentially linked to mutant SF3B1 splicing dysregulation. For example, Benbarche et al. found that the G-patch domain-containing protein GPATCH8 can exert an antagonistic effect on SUGP1 by competing with SUGP1 for the same binding region of DHX15. Silencing GPATCH8 enhanced the DHX15/SUGP1 interaction and corrected one-third of mutant SF3B1 missplicing defects.25 In addition, two members of the DEAD-box ATPase family, DDX42 and DDX46 (also known as PRP5), have been reported to compete for binding to SF3B1 at the hotspot region,26,27 and SF3B1 cancer mutations were found to weaken these interactions.28 However, knockdown (KD) of DDX46 did not induce cryptic 3′ss usage in any top targets of mutant SF3B1.24 Notwithstanding, it remains to be determined if loss of DDX42/46, or in fact DHX15, can also recapitulate the missplicing induced by SF3B1 mutations, as SUGP1 loss or mutation does.

As indicated above, computational screening of splicing changes has contributed to a better understanding of mutant SF3B1 splicing dysregulation. Even though TCGA offers rich genomic information on numerous tumor types, instances of mutations that can mimic mutant SF3B1 splicing defects are rare.21 Beyond mutations, altered expression levels of RNA-binding proteins (RBPs) can also contribute to cancer-associated splicing abnormalities.2931 Thus, further computational screening may reveal additional RBPs/splicing factors whose expression changes reproduce the aberrant splicing patterns induced by SF3B1 mutations, thereby providing new clues for fully elucidating the mechanism underlying mutant SF3B1 missplicing.

In this study, we first employed a computational approach to define a high-confidence list of cryptic 3′ss events characteristic of SF3B1 mutations across different cancer types and cell lines. Based on these events, we conducted a comprehensive screen to evaluate whether RBP/splicing factor loss reproduced the splicing defects induced by SF3B1 mutations. Our findings identify loss of SUGP1 as the underlying driver of mutant SF3B1 missplicing, underscoring the indispensable role of SUGP1 in SF3B1 function during BS recognition and 3′ss selection that is perturbed by SF3B1 hotspot mutations in cancer.

RESULTS

Identification of a comprehensive high-confidence list of cryptic 3′ss used in SF3B1-mutant cells

As discussed above, multiple studies have shown that SF3B1 hotspot mutations typically result in missplicing arising from the use of cryptic 3′ss in affected transcripts. These analyses used RNA from a variety of sources, ranging from patient samples to CRISPR-generated cell lines, and identified hundreds to thousands of transcripts, depending on technical parameters.5 Although several common targets have been identified and validated, the differences in biological samples and computational analyses have prevented the establishment of a comprehensive list of high-confidence mutant SF3B1 targets. Such a list would be of value for many studies, including our efforts here to identify splicing factors whose function might be relevant to mutant SF3B1-induced missplicing.

To achieve the above goals, we designed a computational framework to identify and quantify cryptic 3′ss utilization using 62 samples containing SF3B1 mutations (SF3B1MUT) and 56 SF3B1 wild-type controls (SF3B1WT) from six different datasets, including, for example, MDS, CLL, and UVM (Figure 1A and Table S1). We first identified a total pool of 22,158 cryptic 3′ss candidates with at least 200 cryptic read counts across all SF3B1MUT samples. Next, to capture authentic 3′ss changes resulting from SF3B1 mutations, we compared the percent-spliced-in (PSI) values of each event between SF3B1MUT and SF3B1WT. Cryptic events associated with a p value < 0.05 (t test) and a ΔPSI > 0.1 in three or more of the six cohorts and with a distance of 3–100 nt between cryptic and canonical 3′ss were selected (n = 316). Last, we manually examined every event using Integrative Genomics Viewer (IGV). This analysis identified 295 cryptic 3′ss events as high-confidence recurrent splicing changes arising from SF3B1 mutations (Table S2). We also validated the specificity of the 295 cryptic 3′ss events by unsupervised clustering and principal component analysis (PCA), both of which clearly distinguished SF3B1MUT from SF3B1WT (Figures S1A and S1B).

Figure 1. Identification of SF/RBPs whose expression loss recapitulates cryptic 3′ss usage by SF3B1 hotspot mutations.

Figure 1.

(A) Schema of the computational pipeline used to identify cryptic 3′ss events induced by SF3B1 mutations (left) and to evaluate effects of SF/RBP loss on these cryptic 3′ss events (right).

(B and C) Scatterplot representations of SF/RBPs whose loss of expression positively correlated with cryptic 3′ss usage. The horizontal axis shows the averaged absolute PSI change values between SF/RBP KD/KO samples and controls, and the vertical axis shows the percentage of recapitulated events. Analysis was performed on the 123 core (B) and 295 SF3B1MUT-specific (C) events.

(D) Principal component analysis (PCA) projection of all RNA-seq samples using the PSI values of 295 cryptic 3′ss events.

(E) Hierarchical clustering using Euclidean distance and heatmap analysis of the usage of 295 SF3B1MUT-specific events in all SF/RBP KD/KO, SF3B1MUT, and control samples. Rows represent the 295 SF3B1MUT-specific events, while columns represent the samples. Matrix values correspond to raw PSI values. Each vertical bar represents the mean of the PSI values of the 295 events in each of the SF/RBP loss and control samples. The row annotation bar plot on the right of the heatmap indicates whether the event is in the core 123 event list.

The above analysis utilized samples with different SF3B1 mutations and cancer types, and an interesting question is whether there are mutation/cancer-type-specific differences in splicing patterns. As shown in Figure S1A, we indeed observed such differences within our defined 295-event list. To address systematically whether splicing profiles differ by mutation type or cancer context, we generated eight distinct lists of SF3B1-induced cryptic 3′ss events (defined as p < 0.05 and |ΔPSI| > 0.1), stratified by either mutation type, including the two major hotspots, K700E (n = 31) and R625C/H (n = 18), or cancer cohort, specifically MDS (n = 14), CLL (n = 13), UVM (n = 13), and SKCM (n = 16), as well as the K562/Nalm-6 cell lines. For each list, we assessed the proportion of events overlapping with our 295-event list (Figure S1C). K700E mutation exhibited a higher overlap than R625C/H, and SKCM showed greater overlap compared to other cancer types, likely due to larger sample sizes that enhance statistical power. Notably, the K562/Nalm-6 cell lines exhibited the highest recapitulation rates overall, likely due to reduced transcriptional heterogeneity across replicates.

To minimize potential confounding effects of sample size and cancer/cell line context on SF3B1 mutation-specific splicing profiles, we additionally defined a “core” set of 123 cryptic 3′ss events that met the significance and effect size thresholds (p < 0.05 and |ΔPSI| > 0.1) across all six, as opposed to only three, cohorts (Table S3). This high-confidence list filters out context-specific noise and prioritizes splicing events consistently associated with SF3B1 mutations, thereby enhancing the robustness of our downstream analyses.

As expected and consistent with previous studies,1012 a large majority of the 295 identified cryptic 3′ss were located 10 to 30 nt upstream of the respective canonical 3′ss (Figure S1D), although a small fraction were located further upstream and, notably, 21% were located downstream. To characterize the cryptic 3′ss further, we performed sequence motif analysis on the nucleotide sequences ±50 bp from both cryptic and associated canonical 3′ss of the 295 events, using 500 randomly selected canonical 3′ss without cryptic events as controls (Figure S1E). Consistent with previous observations,11,14,17 a relatively short and weak polypyrimidine tract interrupted with adenosines was found upstream of cryptic 3′ss compared to canonical and control sites (Figure S1E). All together, we have compiled the first robust lists of cryptic 3′ss that are highly specific to and characteristic of SF3B1MUT cells.

Depletion of only two splicing-related proteins out of 600 recapitulates mutant SF3B1-induced missplicing

We next used the above lists of cryptic 3′ss induced by mutant SF3B1 as benchmarks to assess the ability of splicing factor (SF)/RBP loss to reproduce the splicing dysregulation brought about by SF3B1 mutations. To this end, we collected RNA-sequencing (RNA-seq) data on 600 SF/RBP knockdown (KD)/knockout (KO) samples from ENCODE and GEO databases (Table S4). We then calculated the percentage of cryptic 3′ss that could be recapitulated by SF/RBP loss (Figure 1A and Table S5) and ranked all SF/RBPs based on the percentage of recapitulated events (Table S6). As a positive control, we also performed this analysis on each of the SF3B1MUT samples used to generate the event list. We found that all of these SF3B1MUT samples consistently captured most cryptic 3′ss events, with percentages varying from 40% to 98% for the core list and 34%–85% for the 295-event list (Figures 1B and 1C). The variation could be due to multiple factors, including the different SF3B1-mutated residues, which are known to produce distinctive splicing patterns,32 the variant allele frequency, the affected cancer types as well as variations between individual samples. This observation further supports our strategy of utilizing as many SF3B1MUT samples as possible for generation of the lists of cryptic 3′ss events that occur across different cohorts. Notably, SF3B1 KO/KD did not significantly recapitulate these splicing defects (Table S6), consistent with the fact that SF3B1 mutations are neomorphic change-of-function mutations rather than loss-of-function.32

Among the SF/RBP samples, exceeding 2,000 in total, KD/KO of only two spliceosomal proteins recapitulated the SF3B1MUT events to a significant extent. Strikingly, both SUGP1 KO samples reproduced a high percentage of the splicing defects, ∼95% of the 123 core events and ∼80% of the 295 events. Indeed, these values were very similar to the top-ranked SF3B1MUT samples themselves (Figures 1B and 1C), indicating that SUGP1 loss underlies nearly all SF3B1MUT missplicing.

While our missplicing benchmark focuses on cryptic 3′ss usage, given such events represents the dominant splicing abnormality caused by SF3B1MUT, we also explored the contribution of SUGP1 loss to less frequent splicing alterations, such as skipped exon (SE) and intron retention (IR) events, which have also been validated as SF3B1MUT specific in prior studies.13 To this end, we conducted rMATS analysis across six independent SF3B1MUT cohorts. Using the same filtering criteria as for our 295-event list, we identified 43 SE and 24 IR events associated with SF3B1MUT (Figure 2A and Table S13). Assessing their recapitulation in SUGP1 KO K562 cells revealed an approximately 60% overlap (Figure 2B), including key examples such as BRD9 poison exon inclusion and MTERFD3 intron removal13 (Figure 2C). Several explanations are plausible for the subset of SE and IR events not reproduced upon SUGP1 loss. For example, these events may be highly cell type specific and thus not manifest in K562 cells, the model used for SUGP1 knockout. Additionally, SE and IR events may arise from SF3B1 mutations as indirect consequences of broader transcriptomic or regulatory perturbations, rather than the disruption of branchpoint recognition that occurs upon SUGP1 loss.

Figure 2. SUGP1 KO recapitulates ∼60% of SF3B1MUT-specific SE and IR events.

Figure 2.

(A) Schema of the computational pipeline used to identify SE and IR events induced by SF3B1 mutations.

(B) Scatterplot depicts ΔPSI values of SE and IR events in SF3B1MUT cells compared with SF3B1MUT and SUGP1 KO cells. The horizontal axis shows the ΔPSI value between SF3B1MUT and WT samples, while the vertical axis shows the ΔPSI value between SUGP1 KO and control samples. The pie plot (lower-right) indicates the proportion of events recapitulated (brown) or not recapitulated (blue) by SUGP1 KO.

(C) IGV plots of SE and IR events recapitulated by SUGP1 loss in samples with SF3B1 mutations, SUGP1 loss, or WT/control.

Unexpectedly, the second SF/RBP that recapitulated SF3B1MUT-induced missplicing when depleted was the RNA helicase Aquarius (AQR). AQR KD was found to reproduce 28%–42% for the 123-event list (Figure 1B), and 26%–34% for the 295-event list (Figure 1C and Table S6). We also applied unsupervised PCA to the 295-event list across all SF/RBP loss, SF3B1MUT, and SF3B1WT samples, which gave a similar conclusion (Figure 1D). In addition, to obtain a global view of missplicing of the 295 3′ss events, we performed unsupervised hierarchical clustering on all the samples (Figure 1E). Overall, SF3B1 mutations, SUGP1 KO, and AQR KD samples were clustered closer together and with significantly higher PSI than observed with all other SF/RBP KD/KO and WT samples (Figure 1E), indicative of the similarity between these samples in cryptic 3′ss usage.

AQR loss recapitulates SF3B1MUT-specific cryptic 3′ss events by downregulating SUGP1

As mentioned above, our previous studies found that cancer-associated mutations in SUGP1, as well as SUGP1 KD, mimic the missplicing caused by SF3B1 mutations.20,21 Although others have suggested alternative mechanisms, it was not entirely unexpected that our analysis revealed the striking overlap between SUGP1 loss and SF3B1 mutations in inducing cryptic 3′ss use. However, it was completely unanticipated that AQR, a protein with limited evidence of involvement in BS recognition,33 was the only other SF/RBP whose loss mimicked SF3B1MUT missplicing (Figures 1B1E). To validate this finding, we first used unsupervised hierarchical clustering and PCA with AQR KD samples from K562 and HEK293T cells, which revealed that AQR loss recapitulated 39% of SF3B1MUT-specific cryptic 3′ss in the 123-event core list and 32% in the 295-event list (Figures 3A, 3B, and S2AS2C). We next examined experimentally whether AQR KD induced SF3B1MUT cryptic 3′ss use, with a panel of top target transcripts identified above as examples (Figure 3C). To do this, we used two independent small interfering RNAs (siRNAs) to deplete AQR in K562 cells (KD efficiencies shown in Figure 3D). Reverse-transcription PCR (RT-PCR) showed that AQR KD indeed robustly induced use of cryptic 3′ss, recapitulating the splicing pattern we observed in SF3B1K700E K562 cells22 (Figure 3E).

Figure 3. AQR KD recapitulates ∼40% of SF3B1MUT-specific cryptic 3′ss events and induces SUGP1 mRNA missplicing.

Figure 3.

(A) Hierarchical clustering and heatmap analysis of the usage of 123 core events in AQR KD samples and corresponding control samples. Rows represent the 123 core events, while columns represent the samples.

(B) Scatterplot displaying the ΔPSI values for the 123 core events in AQR KD samples compared with SF3B1MUT samples. The horizontal axis shows the ΔPSI value between SF3B1MUT and control samples, while the vertical axis shows the ΔPSI value between the AQR KD sample and control sample. The pie plot (upper left) indicates the proportion of 3′ss events recapitulated (brown) or not recapitulated (blue) by AQR KD.

(C) IGV plots of cryptic 3′ss event recapitulated by AQR loss in samples with SF3B1 mutations, SUGP1 loss, AQR loss, or WT. PSI values are labeled for each track.

(D) Western blotting of protein extracts for SUGP1 from either CRISPR-engineered K562 cells with WT (W) or K700E (K) SF3B1 or parental K562 cells treated with a negative control siRNA (siC) or one of two independent siRNAs targeting AQR.

(E) RT-PCR of cryptic 3′ss (white arrow) and canonical 3′ss (black arrow) produced from splicing of the indicated transcripts.

(F) IGV plots of cryptic 3′ss usage in SUGP1 intron 6 in WT, SF3B1K700E, SUGP1 KO, and AQR KD samples. PSI values are shown for each track.

(G) Scatterplot representation of differentially spliced 3′ss changes on U2 complex genes between AQR KD and controls showing the magnitude (difference of PSI; x axis) and significance (−log10(q value); y axis).

(H) RT-PCR of the cryptic 3′ss and canonical 3′ss produced from splicing of SUGP1 intron 6. Black arrowheads indicate the canonical 3′ splice site (canonical 3′ss), and the white arrowheads indicate the cryptic 3′ splice site (crytpic 3′ss).

(I) Same RT-PCR as in (H) but with or without emetine treatment. Black arrowheads indicate the canonical 3′ splice site (canonical 3′ss), and the white arrowheads indicate the cryptic 3′ splice site (crytpic 3′ss).

(J) Scatterplot depicting the correlation between AQR KD and SUGP1 KO on the SF3B1MUT-specific events recapitulated by AQR KD. The horizontal axis shows the PSI changes between SUGP1 KO and control samples, while the vertical axis shows the PSI changes between AQR KD and control samples. The indicated R value and p value were determined by Pearson correlation.

Previous studies have provided evidence that AQR KD can lead to thousands of altered splicing events.34,35 Indeed, our analysis demonstrated that AQR depletion resulted in 2,407 cryptic 3′ss (Table S7; note that KD/KO of none of the other 600 proteins except SUGP1 resulted in more than 200 cryptic 3′ss), 68 cryptic 5′sss, 806 mutually exclusive exons (MXEs), 121 retained introns (RIs), and 3,919 skipped exons (SEs) (Figure S2D). This raised the possibility that the ∼40% of SF3B1MUT-missplicing events recapitulated by AQR KD arose only by coincidence, reflecting chance overlap. If this were the case, then a similar degree of overlap would be expected to occur with any other sample in which splicing has been disrupted. To test this, we utilized RNA-seq data from CRISPR-engineered K562 cells with hotspot mutations in U2AF1 (three cases of S34F vs. three WTs) and SRSF2 (four cases of P95H vs. four WTs), which are the two most commonly mutated SF genes in cancer besides SF3B136 (Table S8). We first identified the top 295 events that were most specific to the SRSF2 or U2AF1 mutant samples (Figures S2E and S2F) and then compared the correlation between the PSI changes of the 295 events from all three spliceosomal mutations and AQR KD. The results showed that the splicing changes of AQR KD significantly correlated only with the splicing defects detected in the SF3B1MUT samples (Figure S2G). These findings indicate that the overlap we detected between AQR KD and SF3B1MUT samples is specific.

Since AQR has not been shown to function directly with SF3B1 during BS recognition, we wondered whether AQR loss recapitulated SF3B1MUT splicing defects indirectly. An interesting possibility was that this might have resulted from downregulation of SUGP1. To test this hypothesis, we performed western blot (WB) analysis of SUGP1 in AQR KD cells and indeed found a significant decrease (∼50%–60%) of SUGP1 protein levels upon AQR KD (Figures 3D and S2H). One explanation for this, consistent with the massive missplicing induced by AQR KD, is that SUGP1 transcripts were misspliced in AQR KD cells, resulting in reduced levels of SUGP1 mRNA. Using the same computational pipeline for 3′ss event identification as above, we examined the splicing changes in transcripts encoding U2 complex-associated splicing factors annotated in the UniProt database.37 Strikingly, the top target we identified in AQR KD cells was a cryptic 3′ss in SUGP1 intron 6, usage of which led to insertion of a 58-nt intronic sequence in the mRNA that introduced a pre-mature termination codon (PTC) (Figures 3F3H). The misspliced SUGP1 mRNA was subject to nonsense-mediated decay (NMD), as the RT-PCR products encompassing the cryptic 3′ss were increased upon NMD inhibition by emetine37 (Figure 3I). Notably, AQR KD induced use of this novel cryptic 3′ss by a mechanism different from SF3B1 mutations, as SF3B1 mutations did not induce this event (Figure 3F). Finally, consistent with the above, we found a significant positive correlation of ΔPSI between SUGP1 loss and AQR loss on AQR KD-recapitulated SF3B1MUT 3′ss events, which accounted for 32.2% of the 295 events (Figure 3J).

Our results strongly suggest that AQR loss mimics SF3B1MUT splicing by reducing SUGP1 protein levels. To provide further evidence for this hypothesis, we next investigated whether exogenous expression of SUGP1 restores natural 3′ss usage following AQR KD. To this end, we electroporated K562 cells with SUGP1-expressing or backbone control plasmids in the presence of a control siRNA or either of the two siRNAs targeting AQR used above (Figure S3A). RNA was isolated, and splicing of three SF3B1/SUGP1 target transcripts (TOR1AIP2, MED6, and WASHC5) was analyzed by RT-PCR. Consistent with our expectations, exogenous expression of SUGP1 in AQR-KD cells significantly reduced usage of cryptic 3′ss in all three target transcripts (Figures S3B and S3C). These findings confirm that cryptic 3′ss usage of mutSF3B1 target transcripts induced by AQR loss indeed resulted from lower SUGP1 protein levels.

In addition to SUGP1, multiple transcripts misspliced following AQR KD are involved in RNA processing. One example is TFIP11 (Figure 4A), which encodes a G-patch protein that, like SUGP1, functions in splicing by activating DHX15.38 Interestingly, another is the transcript encoding Senataxin (SETX), an RNA: DNA helicase that functions to resolve cotranscriptional R loops.39 SETX was in fact one of the most misspliced transcripts following AQR KD, with > 40% of transcripts using a cryptic 3′ss in intron 23 (Figure 4B). This missplicing, which was not observed in SF3B1MUT or SUGP1 KD cells, generated a PTC and reduced protein levels by ∼80% (Figures 4C and 4D). Notably, it has been proposed that AQR is directly involved in resolving R loops, as the abundance of these structures is increased upon AQR KD.4043 However, AQR has only been shown to possess helicase activity capable of unwinding RNA:RNA hybrids33 and is thought to be localized exclusively in spliceosomal complexes.35 Our data thus suggest that AQR loss leads only indirectly to an increase in R loops, due to SETX downregulation.

Figure 4. AQR KD causes missplicing of SETX transcripts and reduced protein levels.

Figure 4.

(A) Scatterplot representation of differentially spliced 3′ss on all transcripts between AQR KD and controls showing the magnitude (difference of PSI; x axis) and significance (−log10(q value); y axis). Identities of select transcripts are indicated.

(B) IGV plots of the cryptic 3′ss event in SETX intron 23 in WT, SF3B1K700E, SUGP1 KO, and AQR KD samples. PSI values are shown for each track.

(C) Western blot analysis of SETX, AQR, and TUB in extracts from AQR KD K562 cells. SETX/TUB and AQR/TUB ratios from three independent experiments are highlighted.

(D) Quantification of the SETX/TUB ratio relative to siC from WB in (C). Bars represent the mean ± SEM (n = 3). ***p < 0.001.

SUGP1 is the only G-patch protein whose loss recapitulates splicing defects induced by SF3B1MUT

A recent report found that another G-patch-containing protein, GPATCH8, is required for a subset of the missplicing events induced by SF3B1 mutations. GPATCH8 was found to exert an antagonistic effect on SUGP1, and silencing of GPATCH8 in SF3B1-mutated cells corrected one-third of SF3B1MUT-dependent splicing defects.25 Thus, we next sought to compare the effect of SUGP1 or GPATCH8 loss on the cryptic 3′ss changes observed in SF3B1MUT cells. We compared the PSI distributions of these events along an axis of SF3B1WT, SF3B1MUT, and SUGP1 KO in SF3B1MUT cells. Interestingly, we observed a monotonic increase of the overall PSI values along this axis, in which SUGP1 loss in SF3B1MUT cells led to further increased 3′ss usage in about 85% of the 123 core events (Figures 5A and 5B) and 78% of the 295 events (Figures S4A and S4B) compared to SF3B1 mutation alone (Table S9). Consistent with results reported in the previous study,28 this phenomenon was opposite to that observed with GPATCH8 KO. Compared to SF3B1 mutation alone, GPATCH8 KO in SF3B1MUT cells led to decreased 3′ss usage in about 64% of the 123 core events (Figures 5C and 5D) and 66% of the 295 events (Figures S4C and S4D). This rescue proportion was higher than the ∼30% reported by Benbarche et al. upon GPATCH8 KO.25 A likely reason for this discrepancy is that we selected the most prominent splicing events from different SF3B1MUT datasets, including only cryptic 3′ss events, which are not only the most common type of splicing change in SF3B1MUT cells but also the most likely to reflect direct effects.

Figure 5. Distinct effects of G-patch protein loss on SF3B1MUT-specific cryptic 3′ss events.

Figure 5.

(A) Sankey diagram illustrating the changes in PSI values of 123 core events induced by SF3B1 mutation along an axis of SF3B1WT, SF3B1MUT, and SUGP1 KO in SF3B1MUT cells. Different colors represent the range of PSI values.

(B) Scatterplot depicts ΔPSI values of 123 core events in SF3B1MUT cells compared with SF3B1MUT and SUGP1 KO cells. The horizontal axis shows the ΔPSI value between SF3B1MUT and control samples, while the vertical axis shows the ΔPSI value between SF3B1MUT and SUGP1 KO and control samples. Pie plot (lower right) indicates proportions of increasingly (brown) and decreasingly (blue) used cryptic 3′ss in SF3B1MUT and SUGP1 KO samples compared to SF3B1MUT alone sample.

(C) Similar to (A) but for GPATCH8.

(D) Similar to (B) but for GPATCH8.

(E) Hierarchical clustering and heatmap analysis of the ΔPSI values of 123 core events in samples with KO/KD of G-patch protein-encoding genes. Rows represent the 123 core events, while columns represent the samples. The column annotation bar plot at the top of the heatmap indicates the averaged ΔPSI values across the 123 core events.

(F) Scatterplot representations of G patch-containing proteins (Gpatch gene) whose expression loss positively correlated with cryptic 3′ss usage. The horizontal axis shows the average ΔPSI values between G-patch protein KD/KO samples and controls, and the vertical axis shows the percentage of recapitulated events. Analysis was performed on the 123 core events.

Given that DHX helicases such as DHX15 can be activated by multiple G-patch-containing proteins,44 we wondered if loss of any G-PATCH protein other than SUGP1 might also mimic SF3B1MUT missplicing. To examine this possibility, we analyzed splicing patterns induced by loss of all G patch-containing RBPs for which appropriate RNA-seq data were available and found that SUGP1 was the only one whose loss recapitulated the splicing changes induced by mutant SF3B1MUT (Figures 5E and 5F for the 123 core events; Figures S4E and S4F for 295 events).

Loss of DHX15 but not DDX46 or DDX42 recapitulates SF3B1MUT missplicing

As mentioned above, two DDX-type RNA helicases, DDX46 and DDX42, are known to associate with human U2 snRNP and to interact with SF3B1 at the region where hotspot mutations occur.2628 Given that SF3B1 cancer mutations can disrupt the DDX46 interaction, it has been suggested that DDX46 loss from the spliceosome may underlie SF3B1MUT-induced missplicing.31 However, we previously demonstrated experimentally that DDX46 KD did not recapitulate missplicing of any of a panel of top SF3B1MUT targets.24 But since only a small number of transcripts were analyzed, we wished to examine a larger cohort of SF3B1MUT targets. For this, we utilized our core list of 123 transcripts and found that loss of DDX46 or DDX42 reproduced only 9.1% or 8.1% of SF3B1MUT-induced missplicing events, respectively (Figures 6A and 6B). Compared to AQR and SUGP1 loss, the effects of DDX46 and DDX42 loss on recapitulating SF3B1MUT-induced missplicing were minimal (Figure 6D), likely reflecting chance overlap and essentially identical to the overlap observed with all other SF/RBPs tested (Figure 6D). Additionally, we manually examined the raw RNA-seq data for several well-known SF3B1MUT splicing targets, and neither DDX46 nor DDX42 loss caused any splicing defects among these transcripts (Figure 6F). These findings, together with those described above, highlight the unique role of SUGP1 in SF3B1MUT-induced missplicing.

Figure 6. Loss of DHX15 but not DDX42/46 partially recapitulates SF3B1MUT-specific cryptic 3′ss usage.

Figure 6.

(A) Scatterplot representing the ΔPSI values for 123 core events in DDX46 KD samples compared with SF3B1MUT samples. The horizontal axis shows the ΔPSI value between DDX46 KD and control samples, while the vertical axis shows the ΔPSI value between SF3B1MUT sample and control sample. Pie plot (upper left) indicates the proportion of 3′ss events recapitulated (brown) and not recapitulated (blue) by DDX46.

(B) Similar to (A) but for DDX42.

(C) Similar to (A) but for DHX15.

(D) Each vertical bar represents the mean of the PSI values of the 123 core events in each of the SF/RBP loss and control samples.

(E) Scatterplot representations of SF/RBPs whose expression loss positively correlated with cryptic 3′ss usage. The horizontal axis shows the average absolute ΔPSI value between RBP KD/KO samples and controls, and the vertical axis shows the percentage of recapitulated events. Analysis was performed on the 123 core events.

(F) IGV plots of SF3B1MUT-specific cryptic 3′ss events for the indicated transcripts detected in WT, SF3BK700E, SUGP1 KO, DHX15 depletion, DDX46 KD, and DDX42 KO. PSI values are shown for each track. DHX15 depletion was induced using degron tag technology.24

As mentioned above, previous studies have indicated that the DHX helicase DHX15 functions in concert with SUGP1 to ensure accurate BS selection in SF3B1MUT-target introns. However, DHX15 was not among the 600 SF/RBP KD/KO samples in the ENCODE and GEO databases we used for our analyses. We therefore examined whether DHX15 loss recapitulated SF3B1MUT-induced missplicing using six DHX15-depleted samples and controls from the study of Feng et al.45 (Table S10). Notably, DHX15 was very efficiently depleted in these experiments using conditional degron tags, rather than KD/KO. Using the same computational parameters as above, we found that DHX15 depletion recapitulated 22%–27% of the missplicing events in our 123-event core list (Figure 6C and Tables S11 and S12). Albeit not as robust as observed with SUGP1 or AQR loss (Figures 6D6F), these findings support the role of DHX15 in SF3B1MUT-induced missplicing, as suggested previously.23,24

DISCUSSION

In this study, we generated lists of high-confidence cryptic 3′ss events induced by SF3B1 hotspot mutations. Whereas other studies have identified SF3B1MUT target transcripts,1017 our lists are unique in that they are comprehensive and rely on data from numerous studies involving a variety of cell and cancer types. These lists, which captured all previously studied cryptic 3′ss targeted by SF3B1MUT, including MAP3K7, TMEM14C, ABCB7, ORAI2, PPOX, and PPP2R5A, thus represent a valuable resource to drive the biological and functional interrogation of cancer-specific mRNA isoforms. We used this new benchmark to perform an unbiased computational screen of 600 SF/RBPs to identify any whose loss, by KD or KO, recapitulates mutant SF3B1 splicing dysregulation. The results revealed that only loss of SUGP1, and to a lesser extent AQR, likely functioning indirectly, replicated the cryptic 3′ss use associated with SF3B1 mutations, while a separate analysis identified DHX15, previously implicated in SUGP1 function. Below, we discuss these findings and their significance for both mechanisms necessary for correct BS recognition during splicing, as well as how cancer-associated mutations affect this process and cause missplicing.

We have provided here strong evidence that loss of SUGP1 from pre-spliceosomal complexes is the primary, if not exclusive, cause of SF3B1MUT-induced missplicing in cells harboring SF3B1 hotspot mutations. This conclusion is based on the nearly complete overlap between the cryptic 3′sss activated by SUGP1 KO and those used in SF3B1MUT cells. Several previous studies are consistent with this result. Our initial study, which provided the first evidence that SF3B1 cancer-associated hotspot mutations alter splicing by disrupting interaction with SUGP1, showed by RT-PCR that missplicing of a panel of SF3B1MUT target transcripts could be recapitulated by SUGP1 KD.22 Alsafadi et al., in their study identifying cancer-associated mutations in SUGP1, found by RNA-seq that SUGP1 KD recapitulated splicing defects brought about by overexpression of the SF3B1K700E in HEK293 cells.20 Similarly, Feng et al.,45 in a study providing evidence supporting a functional interaction between SUGP1 and DHX15, also observed a significant overlap in missplicing between SUGP1 KD cells and cells expressing SF3B1K700E. None of these studies, however, addressed the question of whether SUGP1 loss is exclusively, at least among the 600 RBPs we systematically evaluated, responsible for SF3B1MUT missplicing, as we have shown here. Moreover, our findings are consistent with our previous work showing that SUGP1 loss or somatic mutations in SUGP1 can mimic the splicing dysregulation observed in SF3B1 mutant tumors.2224 Notably, overexpression of SUGP1 partially rescues these splicing defects in SF3B1-K700E mutant cells.22 These results, supported by RT-PCR validation, provide functional evidence for the causal role of SUGP1 loss in mediating SF3B1 mutation-induced missplicing and highlight SUGP1 as the key effector of the aberrant splicing program associated with SF3B1 mutations.

As mentioned above, the RNA helicases DDX46 and DDX42 have also been suggested to function in SF3B1MUT-induced missplicing. The two helicases have been shown to compete for binding to the hotspot region of SF3B1 through their N-terminal domains during formation of the branch stem loop of U2snRNA, allowing formation of the U2-BS RNA duplex.2628 Mutations altering the hotspot region in SF3B1, such as K700E, disrupt the interaction with DDX46/DDX42.28 However, previous studies16,20,21 found no evidence that the genes encoding either of these two helicases harbor cancer-associated mutations that mimic the cryptic 3′ splicing caused by SF3B1 mutations. This contrasts with the identification of such mutations in SUGP1,20,21 the consequences of which are now understood mechanistically.24 Moreover, as shown here, neither DDX46 KD nor DDX42 KO recapitulated SF3B1 mutation-associated splicing dysregulation. Thus, not all proteins interacting with the SF3B1 hotspot region contribute to the aberrant BS selection resulting from SF3B1 mutations. One explanation for this is that DDX46/DDX42 are recruited to interact with SF3B1 at a later stage in splicing, after discharge of SUGP1 from the spliceosome.24 Once BS identification is complete, facilitated by SUGP1/DHX15, changes in DDX46/DDX42 levels will no longer impact BS recognition.46

Previous studies provided conflicting results concerning the role of DHX15 in mediating the effects of mutant SF3B1. Our study that established the interaction between SUGP1 and DHX15 revealed an ∼13% overlap between SF3B1K700E and DHX15 KD samples,23 while Feng et al. reported no significant overlap between their DHX15 degron-depleted samples and another unspecified SF3B1K700E sample.45 Why these studies gave disparate results and differed from those reported here is unclear but likely reflects the nature and size of the datasets analyzed. In any case, while the SF3B1MUT-like missplicing we detected here in the DHX15-depleted samples was greater than observed with any other SF/RBP except SUGP1 or AQR, an interesting question is why DHX15 loss did not result in the extensive missplicing caused by SUGP1 depletion. While additional studies will be required to address this, it could reflect the fact that DHX15 is involved in several cellular processes44 and can interact with multiple G-patch proteins,47 possibly suppressing certain SUGP1-dependent missplicing events, as we discussed previously.23 Another possibility is that a different DHX helicase might partially compensate for DHX15 loss.

It was initially surprising that AQR was the only protein other than SUGP1, and to a lesser degree DHX15, whose loss significantly recapitulated SF3B1MUT-induced missplicing. AQR is incorporated into spliceosomes as part of a pentameric intron-binding complex that associates with U2snRNP and is essential for splicing in vitro.33 However, based on cryo-EM and biochemical data, AQR appears to function after BS recognition, facilitating transfer of the BS-U2snRNA hybrid to the catalytic center of the spliceosome and release of SF3A and –B complexes from the spliceosome.35 Thus, a direct role for AQR in activation of cryptic 3′ss in SF3B1MUT cells seemed unlikely, and our finding that SUGP1 pre-mRNA is among the many thousands of misspliced transcripts in AQR-depleted cells provides a parsimonious explanation for the SF3B1MUT-like missplicing found in these cells. Indeed, the 50%–60% reduction of SUGP1 protein upon AQR KD is consistent with the fact that only ∼40% of cryptic 3′ss events induced by SF3B1MUT were recapitulated by AQR KD. It is also notable that although AQR loss leads to SUGP1 downregulation by inducing the use of a cryptic 3′ss in SUGP1 pre-mRNA, this site is not used in SF3B1MUT cells. This finding indicates that AQR loss induces missplicing by a mechanism different from that of SF3B1 hotspot mutations, consistent with the structural studies noted above.48

In summary, our study has provided a comprehensive list of core splicing events that are dysregulated by cancer-driving mutations in SF3B1 in multiple cancer types, which should be of considerable value in elucidating the phenotypic consequences of these mutations. Using this list, we found that among hundreds of splicing-related proteins, KD or KO of only two, SUGP1 and AQR, replicated the missplicing events induced by mutant SF3B1. However, because AQR functions indirectly, by reducing SUGP1 protein levels, our findings indicate that SUGP1 loss from mutant spliceosomes is the sole driver of cancer-associated splicing dysregulation induced by SF3B1 hotspot mutations.

Limitations of the study

While our findings strongly support a central role for SUGP1 loss in mediating the splicing defects induced by SF3B1 hotspot mutations, particularly the activation of cryptic 3′ss, it is difficult to rule out conclusively that there are exceptions. First, although our functional screen encompassed ∼600 RBPs from the ENCODE project and GEO datasets, these datasets are inherently constrained by variability in KO/KD efficiency, sequencing depth, and incomplete coverage of all splicing regulators. Second, while SUGP1 loss robustly recapitulates a large proportion of SF3B1MUT-associated cryptic 3′ss events, its ability to mimic other classes of SF3B1MUT-specific splicing alterations is less comprehensive. In our analysis, ∼40% of non-3′ss events, specifically, SE and IR events, were not reproduced in SUGP1 knockout K562 cells. This could suggest another mechanism than SUGP1 loss for these events. Alternatively, it is possible that some or all of them arise indirectly as secondary consequences of broader transcriptomic perturbations.

RESOURCE AVAILABILITY

Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, James L. Manley (jlm2@columbia.edu).

Materials availability

This study did not generate new, unique reagents.

Data and code availability

All data reported in this paper will be shared by the lead contact upon request.

This paper does not report original code. The new RNA-sequencing data regarding the effect of the U2AF1S34F mutation on splicing has been uploaded to the Gene Expression Omnibus (GEO)49 under accession number GEO: GSE285785.

STAR★METHODS

EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

Cell culture

CRISPR-engineered K562 cells with WT or K700E mutant SF3B1 (three clones each) were generated in our previous study.22 These CRISPR cells and parental K562 were cultured in Iscove’s Modified Dulbecco’s Medium (IMDM; Gibco, cat # 12440-053) supplemented with 10% fetal bovine serum (FBS; Seradigm, cat # 89510-186) in a 37°C, 5% CO2 incubator.

Pan-cancer SF3B1-mutated sample collection

We collected RNA sequencing data from both patient and cell line samples with and without SF3B1 hotspot mutations from the TCGA (https://portal.gdc.cancer.gov) and GEO databases (https://www.ncbi.nlm.nih.gov/geo). In total, we collected 63 SF3B1 mutant samples and 56 WT samples, encompassing four tumor types (MDS, CLL, UVM and SKCM) and two cell lines (K562 and Nalm6). Detailed sample information is provided in Table S1.

RBP KO/KD sample collection

All RNA-seq data for RBP KD and KO samples (including shRNA, siRNA, CRISPR and CRISPRi RNA-seq) were downloaded from the ENCODE project phase III database (https://www.encodeproject.org/publications/4928df3e-6995-4401-b643-84980bb94057/). Additionally, we collected RNA-seq data for 187 KD/KO samples targeting 35 RBPs from the GEO database. These RBPs were identified as physically interacting with SF3B1 using mass spectrometry in our previous work.22 In total, we compiled 2,151 RBP-related KD or KO samples and 298 matched control samples, covering 600 RBPs and 18 cell types. Detailed sample information is provided in Table S4.

METHOD DETAILS

Identification of cryptic 3′ss events

We designed a computational analysis pipeline to identify and quantify the usage of all annotated and novel cryptic splice junctions, as well as canonical 3′ss associated with cryptic 3′ ss. Briefly, FASTQ files of the RNA-seq data for SF3B1MUT and SF3B1WT sample were aligned to human genome (hg19) using STAR50 version 2.7.11a, with the reference genome annotation GTF file downloaded from GENCODE database in release 19 (https://www.gencodegenes.org/human/release_19.html). Then, junction reads counts from the STAR output file (SJ.out.tab) was merged into a matrix, with each row representing a splice junction and each column indicating a sample. Low-abundance junctions with fewer than 200 reads (summed up across all samples) were filtered out. For the remaining junctions, we compared each one to known splice junctions in the reference GTF file. A junction was defined as alternatively spliced if there is another junction shared the same annotated 3′ or 5′ end. The associated canonical junction was then identified if it shared the same annotated 3′ or 5′ end with the cryptic splice junction and had the maximum supporting reads in SF3B1WT samples. We then determined the relative position of each alternative 3′ or 5′ splice site to the associated canonical 3′ or 5′ splice site based on their locations on the same transcript strand. Cryptic splicing events where the distance was less than or equal to 3 base pairs (bp) or greater than 100 bp were filtered out. Next, we calculated the PSI values from the raw junction read counts. We performed t-tests to obtain p-values by comparing the PSI values between SF3B1MUT and SF3B1WT samples across six SF3B1MUT-cohorts. Additionally, we calculated ΔPSI, defined as the difference between the mean PSI of SF3B1MUT and SF3B1WT, across six cohorts. Finally, after manual verification by IGV, we retained only those cryptic 3′ss events with a p-value < 0.05 and an absolute ΔPSI >0.1 in more than 2 out of 6 cohorts, and defined this set as the SF3B1MUT-specific events (295 list). Additionally, cryptic 3′ss events with a p-value < 0.05 in more than 2 cohorts and an absolute ΔPSI >0.1 in all 6 cohorts were defined as the core list of events induced by SF3B1 mutant (123 list). For cryptic 3′ss events induced by SUGP1 loss, we applied the same workflow on SUGP1 KO RNA-seq data from GSE242092 and control RNA-seq data from GSE187356, with a total junction read count threshold of 10. After calculating ΔPSI and p-value between the SUGP1 KO and control samples, we sorted the cryptic 3′ss events by p-value and retained the top 295 cryptic 3′ss events specific to SUGP1 loss.

Evaluation of effects of RBP loss on recapitulating SF3B1MUT specific cryptic 3′ss events

RBP KD and KO samples were processed by aligning the FASTQ files to human genome (hg19) using STAR version 2.7.11a, employing the reference genome annotation GTF file downloaded from GENCODE (https://www.gencodegenes.org/human/release_19.html). Next, for each sample, we extracted alternative junction read counts and canonical junction read counts of 295 SF3B1MUT-specific cryptic 3′ss events from corresponding SJ.out.tab to calculate the PSI value (PSIRBP). For each of these 295 cryptic 3′ss events, we summed the alterative and canonical junction reads, respectively, across all control samples to establish the background PSI value (PSICon). Then, a cryptic 3′ss event was deemed to be recapitulated by RBP KD/KO if the absolute value of dPSI (PSIRBP- PSICon) ≥0.1 and the event exhibits the same splicing direction in RBP KD/KO and SF3B1MUT. Lastly, all RBP KD/KO samples were visualized based on the proportion of recapitulated events and mean absolute dPSI (|PSIRBP- PSICon|).

Identification the sequence characteristics of SF3B1MUT-specific cryptic 3′ss events

To identify the sequence characteristics of SF3B1MUT-specific cryptic 3′ss events, we used bedtools2.28 to obtain the nucleotide sequences 50bp upstream and downstream of both the cryptic and associated canonical 3′splice site from our list of 295 events. As a control, we selected 500 introns where no cryptic 3ss usage was detected (p-value = 1 between SF3B1 mutant and WT). For both the aberrant splicing sequences and the control sequences, we calculated the nucleotide composition of adenine (A), thymine (T), guanine (G), and cytosine (C) at each position. We then employed Fisher’s exact test to determine the significance of enrichment for each nucleotide in the aberrant group compared to the control group. Finally, we visualized the nucleotide percentages at each position using the R package ‘ggseqlogo’.

Heatmap and hierarchical clustering

To explore the potential similarity of SF3B1MUT-specific cryptic 3′ss events alteration between SF3B1 mutants and different RBP KD/KO, principal component analysis and unsupervised hierarchical clustering were performed using 295/123 alternative 3′ss event induced by SF3B1MUT. We utilized the ‘prcomp’ command from the R package ‘stats’ for principal component analysis. For clustering analysis, we employed the ‘Heatmap’ command from the R package ‘ComplexHeatmap’, using Euclidean distance as the metric.

Identification of alternative splicing events induced by AQR loss

To identify cryptic 3′ and 5′ splice sites (3′ss/5′ss), we applied the analytical pipeline described in the “Identification of cryptic 3′ss events” STAR Methods section to 14 AQR-knockdown (KD) and 14 matched control RNA-seq samples (sourced from ENCODE and GSE195668; Table S2). For other splicing event types—including skipped exons (SE), retained introns (RI), and mutually exclusive exons (MXE)—we performed rMATS52 analysis. All events were called using consistent statistical thresholds (p < 0.05 and |ΔPSI| > 0.1) as specified throughout the manuscript.

Identification of alternative splicing events induced by SRSF2MUT and U2AF1MUT

For SRSF250 FASTQ files were downloaded from the GEO database via access id GSE71299, and aligned to the human genome (hg19) using STAR version 2.7.11a, with reference genome annotation GTF file downloaded from GENCODE database in release 19 (https://www.gencodegenes.org/human/release_19.html). Then, rMATs (ref) (version 4.1.1) was used to identify the differential exon skipping (ES) events between SRSF2 mutant and WT (4 vs. 4). We extracted the inclusion level (IncLevel) values from the [Junction Counts and Exon Counts] (JCEC) file to create a matrix with IncLevel as the variable across samples. Subsequently, we recalculated the p-values using a t-test between SRSF2MUT and WT to determine the inclusion level difference. Events were filtered by a p-value of 0.05 and an inclusion level difference of 10%. Additionally, we collected SRSF2MUT-specific events that were experimentally validated from the ASCancer Atlas database (https://ngdc.cncb.ac.cn/ascancer/home). After manual confirmation using IGV, 42 events were retained. Finally, by sorting based on p-values and including the 42 experimentally validated events, a total of 295 SRSF2MUT-specific AS events were obtained (Figure S2D). For U2AF1, we applied the same workflow as above on three cases of U2AF1MUT and three cases of U2AF1WT RNA-seq data to obtain 295 U2AF1MUT specific AS events (Figure S2E).

Evaluation of effects of SUGP1 loss on recapitulating SF3B1MUT specific skipped exon (SE) and intron retention (IR) events

We first used the BAM files generated by STAR50 (version 2.7.11a) to run rMATS52 and identify SF3B1MUT-associated SE and IR events (Figure 2A). Specifically, based on the PSI values calculated by rMATS for all SF3B1 mutant and wild-type samples, we computed the p-value (t-test) and ΔPSI for each cohort, where ΔPSI is defined as the difference between the mean PSI of SF3B1MUT and SF3B1WT samples. We retained only those SE/IR events with a p-value <0.05 and an absolute ΔPSI >0.1 in more than 2 out of 6 cohorts, and defined this set as the original SF3B1MUT -specific events (195 SE events and 95 IR events). Because our evaluation of recapitulation by SUGP1 loss required PSI values based on precise event coordinates (rather than rMATS output), we re-calculated PSI using SJ and BAM files. For SE events, PSI was computed from the relevant junction read counts in the SJ.out. tab file. For IR events, PSI was calculated as: PSI = Intron-inclusion reads/(Intron-inclusion reads + Intron-exclusion reads). Intron-exclusion reads were extracted from the SJ file, while intron-inclusion reads were counted from the BAM files using SAMtools,53 based on read coverage across ±5 bp flanking the splice junctions (Figure 2A). Based on the re-quantified PSI values, we applied the same filtering criteria (p < 0.05 and |ΔPSI| > 0.1 in ≥3 cohorts), followed by manual validation in IGV, resulting in a high-confidence set of 43 SE and 24 IR events.

To assess SUGP1 loss-mediated recapitulation, we calculated PSI values in control and SUGP1KO samples using the same approach. An event was considered recapitulated if dPSI (PSISUGP1−KO-PSICon) ≥0.1 and the event exhibits the same splicing direction in SUGP1KO and SF3B1MUT.

Knockdown and rescue experiments

K562 cells were electroporated with two independent AQR siRNAs using the Neon Transfection System (Thermo Fisher Sci) with the Neon 100uL tip kit that includes the R buffer for electroporation (Thermo Fisher Sci). 1.5 million cells were resuspended in 100uL of R buffer, containing 1uL of 100uM siRNAs. Transfection was performed following the manufacturer’s protocol. Electroporated cells were immediately placed in 2 mL complete media for 24h for a final siRNA concentration of 50uM. After 24h, cells were washed with PBS and electroporated again with the same siRNAs (second round) and placed in 2 mL of complete media for another 24h at a 50uM concentration. Cells were harvested at 48h after the second round of electroporation for downstream assays. The sequences of the two siRNAs targeting AQR were: siAQR-1 sense strand (5′-GAGAUUGUCAAAUCAAGGUdTdT-3′), siAQR-1 antisense strand (5′-AC CUUGAUUUGACAAUCUCdTdT-3′), siAQR-2 sense strand (5′-GGCGCUGGUUUAAUACCAUdTdT-3′), and siAQR-2 antisense strand (5′-AUGGUAUUAAACCAGCGCCdTdT-3′). The sequences of the sense and antisense strands of the negative control siRNA (siC) were listed in our previous study.22 For emetine treatment, a final concentration of 100 μg/mL emetine dihydrochloride was added to the cells at 36h post-second-round siRNA electroporation. After 12h of incubation, the cells were harvested for RNA extraction. For the exogenous SUGP1 rescue experiments, 1.5 million K562 leukemia cells per condition were electroporated twice in consecutive days according to the same protocol mentioned above with 6ug of N-terminally HA-tagged SUGP1 plasmid22 or backbone control in the presence of either 50uM of control siRNA (siC) or two siRNAs targeting the AQR mRNA (siAQR-1 and siAQR-2). 48h after the second electroporation, cells were collected and total RNA was extracted using TRIzol purification.

Western blotting

Western blotting was performed as described22 with the following primary antibodies: anti-AQR (1:1,000, Bethyl Laboratories, A302-547A-T), anti-SUGP1 (1:1,000, Sigma-Aldrich, HPA004890), anti-ACTIN (1:2,000, Sigma-Aldrich, A2066), anti-SETX (1:1000, Bethyl Laboratories, A301-105A), anti-HA (1:1,000, ABM, G166). Secondary antibodies used were Goat anti-Rabbit IgG (H + L) Secondary Antibody, HRP (ThermoFisher, cat #31460) and Goat anti-Mouse IgM (Heavy chain) Secondary Antibody, HRP (ThermoFisher, cat# 62-6820). Membranes were detected using the Pierce ECL Western Blotting Substrate chemiluminescent reagent (ThermoFisher cat# 1859698). The fluorescence secondary antibody used was Donkey anti-Rabbit IgG (LI-COR, 926-68073, 1:5,000). Fluorescence and chemiluminescent signals were detected using the ChemiDoc Imaging System (Bio-Rad).

Reverse transcription-polymerase chain reaction

Total RNA was extracted with TRIzol (Thermo Fisher Scientific). In a 20-μL reaction, 2 μg total RNA was reverse-transcribed using 0.3 μL Maxima Reverse Transcriptase (Thermo Fisher Scientific) and 50 pmol oligo-dT primer. The cDNA was diluted (1:10), and 3 μL were used for PCR. PCR products were subjected to agarose gel electrophoresis with ethidium bromide staining, followed by gel imaging with the ChemiDoc Imaging System (Bio-Rad). Primers used were MED6 forward, 5′-CTGGAGTGATCTATCAGGCACC-3′; MED6 reverse, 5′-GCCACCAATACCCTTTGGAAGG-3′; TOR1AIP2 forward, 5′-GTTGGTCTCGAACTCCTGGCTT-3′; TOR1AIP2 reverse 5′-TGAAGAGTCTGGGAATGCGTCAC-3′; WASHC5 forward, 5′-GGGCAGATGCAGATTCTGAGAC-3′; WASHC5 reverse, 5′-GTGAAGGGTCCT GATAGTGGG-3′; ZNF410 forward, 5′-ACAGGAAGGTATCATTGGCTCTG-3′; ZNF410 reverse, 5′-TCTACAAGCCCCTGTCCCAATG-3′; SUGP1 forward, 5′-AAAGGAAGCACAGAAGTCGCAG-3′; and SUGP1 reverse, 5′-TTCTCACGGTTGTTCTGGAGGG-3′. Primers for GCC2 and MAP3K7 were listed in our previous study.22

QUANTIFICATION AND STATISTICAL ANALYSIS

Statistical analysis of western blots and RT-PCR assays

Chemiluminescent/fluorescence signals were analyzed using ImageJ and the band intensity of the target proteins (SUGP1, AQR and SETX) was divided by the band intensity of the housekeeping proteins used (ACTB and TUB). The ratios obtained from siAQR-1 and siAQR-2 were then normalized to the ratios of siControl samples. The normalized ratios from 3 independent experiments were then analyzed using an ordinary one-way ANOVA with Dunnet’s multiple comparison test wherein the mean of each column (siAQR-1 and siAQR-2) was compared to the mean of a control column (siC). For the SUGP1 rescue experiments, band intensity of each splicing event (lower band - normal splicing, top band – cryptic 3′ss splicing) was measured using ImageJ and divided by the PCR product size, since ethidium bromide binds DNA double-stranded PCR products in an equimolar ratio. The normalized band intensities were then used to calculate the percent spliced-in (PSI) ratio by dividing the top band intensity by the total intensity of both bands. These values were analyzed in GraphPad Prism in a two-way ANOVA with Tukey’s multiple comparison test in which siAQR-1-EV and siAQR-2-EV PSI were compared with siAQR-1-SUGP1 and siAQR-2-SUGP1, respectively. P-values for the multiple tests are reported as asterisks (*p ≤ 0.05, **p ≤ 0.01, ***p ≤ 0.001, ****p ≤ 0.0001) (Table S14, Figures S2H, S3C, and 4D).

Statistical analysis of alternative splicing events

Differential splicing analysis was performed using unpaired Student’s t-tests comparing: (i) splicing factor mutants vs. wild-type, and (ii) RBP-loss vs. control samples. Significant splicing events were defined as those with p < 0.05 and |ΔPSI|>0.1 (Figures 1A, 2A, S2E, and S2F).

Correlation analysis

All correlation analyses in this study were performed using Pearson correlation. Statistical significance was defined as p < 0.05 (Figures 3J and S2G).

Supplementary Material

1
2
3
4
5
6
7
9
10
11
12
13
14
15
16

Supplemental information can be found online at https://doi.org/10.1016/j.celrep.2025.116075.

KEY RESOURCES TABLE

REAGENT or RESOURCE SOURCE IDENTIFIER
Antibodies
anti-AQR Bethyl Laboratories A302-547A-T
anti-ACTIN Sigma A2066
anti-HA rabbit polyclonal Abm G166
anti-SUGP1 Sigma HPA004890
anti-SETX Bethyl Laboratories A301-105A
Goat anti-Rabbit IgG (H + L) ThermoFisher 31460
Goat anti-Mouse IgM (Heavy chain) ThermoFisher 62-6820
Donkey anti-Rabbit IgG LI-COR 926-68073

Chemicals, peptides, and recombinant proteins
Pierce ECL Western Blotting Substrate chemiluminescent reagent Thermo Fisher Scientific 1859698
M-PER Mammalian Protein Extraction Reagent Thermo Fisher Scientific 78505
TRIzol Thermo Fisher Scientific 15596018

Deposited data
RNA-sequencing data This paper GEO: GSE285785

Experimental models: Cell lines
HEK293T ATCC CRL-3216
K562 ATCC CCL-243
CRISPR-engineered K562 cells with WT or K700E mutant SF3B1 Lieu et al., 202215 N/A

Oligonucleotides
siC sense, UUCUCCGAACGUGUCACGUdTdT Zhang et al., 201922 N/A
siC antisense, ACGUGACACGUUCGGAGAAdTdT Zhang et al., 201922 N/A
siAQR-1 sense, GAGAUUGUCAAAUCAAGGUdTdT This paper N/A
siAQR-1 antisense, ACCUUGAUUUGACAAUCUCdTdT This paper N/A
siAQR-2 sense, GGCGCUGGUUUAAUACCAUdTdT This paper N/A
siAQR-2 antisense, AUGGUAUUAAACCAGCGCCdTdT This paper N/A

Recombinant DNA
p3xFLAG-CMV-14-HA-SUGP1 Zhang et al., 201922 N/A

Software and algorithms
STAR (2.7.11a) Dobin et al., 201350 https://github.com/alexdobin/STAR
ImageJ NIH https://imagej.nih.gov/ij/
SAMtools Li et al., 200951 https://samtools.sourceforge.net/
rMATS (4.1.1) Shen et al., 201452 N/A
Excel Microsoft N/A
R software R Core Team, 2013 https://www.R-project.org/
Prism GraphPad N/A
Illustrator Adobe N/A

Highlights.

  • High-confidence list of missplicing events induced by hotspot SF3B1 mutations

  • SUGP1 KO reproduces ∼95% of cryptic 3′ss induced by SF3B1 mutations

  • AQR KD recapitulates ∼40% of mutant SF3B1 missplicing by downregulating SUGP1

  • Loss of DHX15 but not DDX42/46 partially recapitulates mutant SF3B1 missplicing

ACKNOWLEDGMENTS

This work was supported by grants from the Beijing Natural Science Foundation (grant Z220012) to Z.L. and from the NIH (grant R35 GM118136) to J.L.M., as well as by grants to Z.L. from the National Natural Science Foundation of China (grant 32170565) and the Chinese Academy of Sciences Hundred Talents Program. We thank Dr. Yen Lieu for her help during the early stages of this work.

Footnotes

DECLARATION OF INTERESTS

The authors declare no conflicts of interest.

REFERENCES

  • 1.Wilkinson ME, Charenton C, and Nagai K (2020). RNA Splicing by the Spliceosome. Annu. Rev. Biochem 89, 359–388. [DOI] [PubMed] [Google Scholar]
  • 2.Pan Q, Shai O, Lee LJ, Frey BJ, and Blencowe BJ (2008). Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing. Nat. Genet 40, 1413–1415. [DOI] [PubMed] [Google Scholar]
  • 3.Wang ET, Sandberg R, Luo S, Khrebtukova I, Zhang L, Mayr C, Kingsmore SF, Schroth GP, and Burge CB (2008). Alternative isoform regulation in human tissue transcriptomes. Nature 456, 470–476. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wahl MC, Will CL, and Lührmann R (2009). The Spliceosome: Design Principles of a Dynamic RNP Machine. Cell 136, 701–718. [DOI] [PubMed] [Google Scholar]
  • 5.Bradley RK, and Anczuków O (2023). RNA splicing dysregulation and the hallmarks of cancer. Nat. Rev. Cancer 23, 135–155. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Seiler M, Peng S, Agrawal AA, Palacino J, Teng T, Zhu P, Smith PG, Cancer Genome Atlas Research Network; Buonamici S, and Yu L (2018). Somatic Mutational Landscape of Splicing Factor Genes and Their Functional Consequences across 33 Cancer Types. Cell Rep 23, 282–296.e4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Yoshida K, Sanada M, Shiraishi Y, Nowak D, Nagata Y, Yamamoto R, Sato Y, Sato-Otsubo A, Kon A, Nagasaki M, et al. (2011). Frequent pathway mutations of splicing machinery in myelodysplasia. Nature 478, 64–69. [DOI] [PubMed] [Google Scholar]
  • 8.Papaemmanuil E, Cazzola M, Boultwood J, Malcovati L, Vyas P, Bowen D, Pellagatti A, Wainscoat JS, Hellstrom-Lindberg E, Gamba-corti-Passerini C, et al. (2011). Somatic SF3B1 Mutation in Myelodysplasia with Ring Sideroblasts. N. Engl. J. Med 365, 1384–1395. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Cretu C, Schmitzová J, Ponce-Salvatierra A, Dybkov O, De Laurentiis EI, Sharma K, Will CL, Urlaub H, Lührmann R, and Pena V (2016). Molecular Architecture of SF3b and Structural Consequences of Its Cancer-Related Mutations. Mol. Cell 64, 307–319. [DOI] [PubMed] [Google Scholar]
  • 10.Alsafadi S, Houy A, Battistella A, Popova T, Wassef M, Henry E, Tirode F, Constantinou A, Piperno-Neumann S, Roman-Roman S, et al. (2016). Cancer-associated SF3B1 mutations affect alternative splicing by promoting alternative branchpoint usage. Nat. Commun 7, 10615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Darman RB, Seiler M, Agrawal AA, Lim KH, Peng S, Aird D, Bailey SL, Bhavsar EB, Chan B, Colla S, et al. (2015). Cancer-Associated SF3B1 Hotspot Mutations Induce Cryptic 3′ Splice Site Selection through Use of a Different Branch Point. Cell Rep 13, 1033–1045. [DOI] [PubMed] [Google Scholar]
  • 12.DeBoever C, Ghia EM, Shepard PJ, Rassenti L, Barrett CL, Jepsen K, Jamieson CHM, Carson D, Kipps TJ, and Frazer KA (2015). Transcriptome Sequencing Reveals Potential Mechanism of Cryptic 3′ Splice Site Selection in SF3B1-mutated Cancers. PLoS Comput. Biol 11, e1004105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Inoue D, Chew GL, Liu B, Michel BC, Pangallo J, D’Avino AR, Hitchman T, North K, Lee SCW, Bitner L, et al. (2019). Spliceosomal disruption of the non-canonical BAF complex in cancer. Nature 574, 432–436. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Obeng EA, Chappell RJ, Seiler M, Chen MC, Campagna DR, Schmidt PJ, Schneider RK, Lord AM, Wang L, Gambe RG, et al. (2016). Physiologic Expression of Sf3b1K700E Causes Impaired Erythropoiesis, Aberrant Splicing, and Sensitivity to Therapeutic Spliceosome Modulation. Cancer Cell 30, 404–417. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lieu YK, Liu Z, Ali AM, Wei X, Penson A, Zhang J, An X, Rabadan R, Raza A, Manley JL, and Mukherjee S (2022). SF3B1 mutant-induced missplicing of MAP3K7 causes anemia in myelodysplastic syndromes. Proc. Natl. Acad. Sci 119, e2111703119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Clough CA, Pangallo J, Sarchi M, Ilagan JO, North K, Bergantinos R, Stolla MC, Naru J, Nugent P, Kim E, et al. (2022). Coordinated missplicing of TMEM14C and ABCB7 causes ring sideroblast formation in SF3B1-mutant myelodysplastic syndrome. Blood 139, 2038–2049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Liu Z, Yoshimi A, Wang J, Cho H, Chun-Wei Lee S, Ki M, Bitner L, Chu T, Shah H, Liu B, et al. (2020). Mutations in the RNA Splicing Factor SF3B1 Promote Tumorigenesis through MYC Stabilization. Cancer Discov 10, 806–821. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Liu B, Liu Z, Chen S, Ki M, Erickson C, Reis-Filho JS, Durham BH, Chang Q, de Stanchina E, Sun Y, et al. (2021). Mutant SF3B1 promotes AKT- and NF-κB–driven mammary tumorigenesis. J. Clin. Investig 131, e138315. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Simmler P, Ioannidi EI, Mengis T, Marquart KF, Asawa S, Van-Lehmann K, Kahles A, Thomas T, Schwerdel C, Aceto N, et al. (2023). Mutant SF3B1 promotes malignancy in PDAC. eLife 12, e80683. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Alsafadi S, Dayot S, Tarin M, Houy A, Bellanger D, Cornella M, Wassef M, Waterfall JJ, Lehnert E, Roman-Roman S, et al. (2021). Genetic alterations of SUGP1 mimic mutant-SF3B1 splice pattern in lung adenocarcinoma and other cancers. Oncogene 40, 85–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Liu Z, Zhang J, Sun Y, Perea-Chamblee TE, Manley JL, and Rabadan R (2020). Pan-cancer analysis identifies mutations in SUGP1 that recapitulate mutant SF3B1 splicing dysregulation. Proc. Natl. Acad. Sci 117, 10305–10312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Zhang J, Ali AM, Lieu YK, Liu Z, Gao J, Rabadan R, Raza A, Mukherjee S, and Manley JL (2019). Disease-Causing Mutations in SF3B1 Alter Splicing by Disrupting Interaction with SUGP1. Mol. Cell 76, 82–95.e7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Zhang J, Huang J, Xu K, Xing P, Huang Y, Liu Z, Tong L, and Manley JL (2022). DHX15 is involved in SUGP1-mediated RNA missplicing by mutant SF3B1 in cancer. Proc. Natl. Acad. Sci 119, e2216712119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhang J, Xie J, Huang J, Liu X, Xu R, Tholen J, Galej WP, Tong L, Manley JL, and Liu Z (2023). Characterization of the SF3B1–SUGP1 interface reveals how numerous cancer mutations cause mRNA missplicing. Genes Dev 37, 968–983. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Benbarche S, Pineda JMB, Galvis LB, Biswas J, Liu B, Wang E, Zhang Q, Hogg SJ, Lyttle K, Dahi A, et al. (2024). GPATCH8 modulates mutant SF3B1 mis-splicing and pathogenicity in hematologic malignancies. Mol. Cell 84, 1886–1903.e10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Will CL, Urlaub H, Achsel T, Gentzel M, Wilm M, and Lührmann R (2002). Characterization of novel SF3b and 17S U2 snRNP proteins, including a human Prp5p homologue and an SF3b DEAD-box protein. EMBO J 21, 4978–4988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yang F, Bian T, Zhan X, Chen Z, Xing Z, Larsen NA, Zhang X, and Shi Y (2023). Mechanisms of the RNA helicases DDX42 and DDX46 in human U2 snRNP assembly. Nat. Commun 14, 897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Zhang X, Zhan X, Bian T, Yang F, Li P, Lu Y, Xing Z, Fan R, Zhang QC, and Shi Y (2024). Structural insights into branch site proof-reading by human spliceosome. Nat. Struct. Mol. Biol 31, 835–845. [DOI] [PubMed] [Google Scholar]
  • 29.David CJ, Chen M, Assanah M, Canoll P, and Manley JL (2010). HnRNP proteins controlled by c-Myc deregulate pyruvate kinase mRNA splicing in cancer. Nature 463, 364–368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Jbara A, Lin KT, Stossel C, Siegfried Z, Shqerat H, Amar-Schwartz A, Elyada E, Mogilevsky M, Raitses-Gurevich M, Johnson JL, et al. (2023). RBFOX2 modulates a metastatic signature of alternative splicing in pancreatic cancer. Nature 617, 147–153. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Zhou X, Wang R, Li X, Yu L, Hua D, Sun C, Shi C, Luo W, Rao C, Jiang Z, et al. (2019). Splicing factor SRSF1 promotes gliomagenesis via oncogenic splice-switching of MYO1B. J. Clin. Investig 129, 676–693. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Canbezdi C, Tarin M, Houy A, Bellanger D, Popova T, Stern MH, Roman-Roman S, and Alsafadi S (2021). Functional and conformational impact of cancer-associated SF3B1 mutations depends on the position and the charge of amino acid substitution. Comput. Struct. Biotechnol. J 19, 1361–1370. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.De I, Bessonov S, Hofele R, dos Santos K, Will CL, Urlaub H, Lührmann R, and Pena V (2015). The RNA helicase Aquarius exhibits structural adaptations mediating its recruitment to spliceosomes. Nat. Struct. Mol. Biol 22, 138–144. [DOI] [PubMed] [Google Scholar]
  • 34.Van Nostrand EL, Freese P, Pratt GA, Wang X, Wei X, Xiao R, Blue SM, Chen JY, Cody NAL, Dominguez D, et al. (2020). A large-scale binding and functional map of human RNA-binding proteins. Nature 583, 711–719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Van Nostrand EL, Pratt GA, Yee BA, Wheeler EC, Blue SM, Mueller J, Park SS, Garcia KE, Gelboin-Burkhart C, Nguyen TB, et al. (2020). Principles of RNA processing from analysis of enhanced CLIP maps for 150 RNA binding proteins. Genome Biol 21, 90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Zhang Q, Ai Y, and Abdel-Wahab O (2024). Molecular impact of mutations in RNA splicing factors in cancer. Mol. Cell 84, 3667–3680. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.UniProt Consortium (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res 51, D523–D531. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Yoshimoto R, Kataoka N, Okawa K, and Ohno M (2009). Isolation and characterization of post-splicing lariat-intron complexes. Nucleic Acids Res 37, 891–902. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hasanova Z, Klapstova V, Porrua O, Stefl R, and Sebesta M (2023). Human senataxin is a bona fide R-loop resolving enzyme and transcription termination factor. Nucleic Acids Res 51, 2818–2837. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sollier J, Stork CT, García-Rubio ML, Paulsen RD, Aguilera A, and Cimprich KA (2014). Transcription-coupled nucleotide excision repair factors promote R-loop-induced genome instability. Mol. Cell 56, 777–785. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Sakasai R, Isono M, Wakasugi M, Hashimoto M, Sunatani Y, Matsui T, Shibata A, Matsunaga T, and Iwabuchi K (2017). Aquarius is required for proper CtIP expression and homologous recombination repair. Sci. Rep 7, 13808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Cui H, Shi Q, Macarios CM, and Schimmel P (2024). Metabolic regulation of mRNA splicing. Trends Cell Biol 34, 756–770. [DOI] [PubMed] [Google Scholar]
  • 43.Saha S, Yang X, Huang SYN, Agama K, Baechler SA, Sun Y, Zhang H, Saha LK, Su S, Jenkins LM, et al. (2022). Resolution of R-loops by topoisomerase III-β (TOP3B) in coordination with the DEAD-box helicase DDX5. Cell Rep 40, 111067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Bohnsack KE, Kanwal N, and Bohnsack MT (2022). Prp43/DHX15 exemplify RNA helicase multifunctionality in the gene expression network. Nucleic Acids Res 50, 9012–9022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Feng Q, Krick K, Chu J, and Burge CB (2023). Splicing quality control mediated by DHX15 and its G-patch activator SUGP1. Cell Rep 42, 113223. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Bak-Gordon P, and Manley JL (2025). SF3B1: From core splicing factor to oncogenic driver. RNA In Press 31, 314–332. 10.1261/rna.080368.124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Bohnsack KE, Ficner R, Bohnsack MT, and Jonas S (2021). Regulation of DEAH-box RNA helicases by G-patch proteins. Biol. Chem 402, 561–579. [DOI] [PubMed] [Google Scholar]
  • 48.Schmitzová J, Cretu C, Dienemann C, Urlaub H, and Pena V (2023). Structural basis of catalytic activation in human splicing. Nature 617, 842–850. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Edgar R, Domrachev M, and Lash AE (2002). Gene Expression Omnibus: NCBI gene expression and hybridization array data repository. Nucleic Acids Res 30, 207–210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, and Gingeras TR (2013). STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, and Durbin R; 1000 Genome Project Data Processing Subgroup (2009). 1000 Genome Project Data Processing Subgroup. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Shen S, Park JW, Lu ZX, Lin L, Henry MD, Wu YN, Zhou Q, and Xing Y (2014). rMATS: robust and flexible detection of differential alternative splicing from replicate RNA-Seq data. Proc. Natl. Acad. Sci 111, E5593–E5601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Zhang J, Lieu YK, Ali AM, Penson A, Reggio KS, Rabadan R, Raza A, Mukherjee S, and Manley JL (2015). Disease-associated mutation in SRSF2 misregulates splicing by altering RNA-binding affinities. Proc. Natl. Acad. Sci. USA 112, E4726–E4734. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1
2
3
4
5
6
7
9
10
11
12
13
14
15
16

Data Availability Statement

All data reported in this paper will be shared by the lead contact upon request.

This paper does not report original code. The new RNA-sequencing data regarding the effect of the U2AF1S34F mutation on splicing has been uploaded to the Gene Expression Omnibus (GEO)49 under accession number GEO: GSE285785.

RESOURCES