Summary
Variants with intermediate functional effects—neither fully disruptive nor functionally neutral—represent an under-recognized source of genetic complexity and define a functional gray zone that complicates variant classification. Here, we address this issue using GT>GC (+2T>C) 5′ splice-site variants as a tractable model, as approximately 15%–18% of such substitutions retain variable levels of residual wild-type (WT) transcript. Using residual WT transcript as a quantitative functional readout, we first revisited disease-associated GT>GC variants previously shown to retain substantial WT transcript, including SPINK1 c.194+2T>C, HBB c.315+2T>C, and BRCA2 c.8331+2T>C, illustrating how intermediate splicing effects complicate clinical interpretation across distinct genes and disease contexts. We then performed a locus-wide assessment of all 26 theoretically possible GT>GC substitutions in CFTR, integrating SpliceAI delta donor-loss scores with classifications from expert-curated databases. Minigene splicing analyses of four selected CFTR variants, together with full-length and minigene analyses of a BAP1 GT>GC variant with conflicting clinical interpretations, revealed heterogeneous and context-dependent splicing outcomes, underscoring both inter-assay variability and the inherent limitations of commonly used splicing assay systems. Collectively, our findings indicate that GT>GC variants capable of generating appreciable residual WT transcript exemplify a broader class of intermediate-effect alleles that expose the limitations of both computational prediction and experimental assessment. These observations highlight the need for classification frameworks that incorporate quantitative functional data and better capture the continuum of variant effects.
Keywords: ACMG/AMP variant classification guidelines, canonical splice-site variants, ClinVar, GT>GC, +2T>C, intermediate functional effects, loss of function, LoF variants, residual wild-type transcript, SpliceAI, variant effect spectrum, variants of uncertain significance, VUS
Intermediate-effect variants blur the line between pathogenic and benign. Using GT>GC (+2T>C) 5′ splice-site variants as a model, this study shows how residual wild-type transcript generation complicates prediction and interpretation, underscoring the need for variant classification frameworks that capture quantitative effects along a continuum.
Introduction
Accurate classification of genetic variants remains a central challenge in human genetics.1,2 The increasing use of a genotype-first or genomic-first approach,3,4 enabled by exome and genome sequencing, has greatly magnified this challenge. In many cases, the mode of inheritance becomes apparent only after genetic analysis, and disorders presumed to follow a simple Mendelian pattern may instead have a complex, multifactorial, or oligogenic basis. This complexity underscores the limitations of the widely used American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) framework.5 Its dichotomization of variants as either “pathogenic” or “benign,” although useful for clinical consistency, cannot account for the full spectrum of variant effects that contribute to human genetic disease.3,6,7,8,9 It should be noted that the current ACMG/AMP category “variant of uncertain significance (VUS)”5 is not a biological entity but a temporary designation expected to resolve as either pathogenic or benign once sufficient evidence becomes available.
For clarity, we focus here on variants that act through a loss-of-function (LoF) mechanism. In contrast to variants at the two extremes—those causing complete or near-complete loss of function of the affected allele, typically responsible for classic Mendelian disease phenotypes, and those retaining complete or near-complete normal function—variants with intermediate effects are challenging to interpret because they lie outside the scope of binary pathogenicity models. This challenge is illustrated by the multiple partial LoF variants identified in SPINK1 (MIM: 167790).8 Despite their clear biological reality and clinical relevance, such variants remain underexplored and are not adequately addressed by any widely established classification frameworks, which offer limited guidance for interpreting quantitative or intermediate functional effects and thereby risk systematic over- or underclassification.
In this study, we explore GT>GC (+2T>C) 5′ splice-site variants as a tractable model to investigate how intermediate functional effects complicate classification. Although substitutions at canonical GT donor or AG acceptor sites typically abolish normal splicing,10,11 a subset of GT>GC substitutions defy this expectation. Specifically, a meta-analysis of 45 disease-associated variants, together with full-length gene-splicing assay (FLGSA) data from 103 nucleotide substitutions, showed that 15%–18% retain residual (alternatively referred to as “leaky”) wild-type (WT) transcript, in some cases at levels approaching 80% of normal.12 This variability can be clinically significant. For instance, in cystic fibrosis (MIM: 219700), as little as 5% normal cystic fibrosis transmembrane conductance regulator (CFTR) activity may be sufficient to prevent pulmonary disease,13,14 whereas in the hemophilias (MIM: 306700 and 306900), factor VIII or IX levels above 5% may lead only to mild bleeding diatheses.15,16
The ability of some GT>GC 5′ splice-site variants to generate residual WT transcript has a clear biological basis. Splicing at U2-type 5′ splice sites depends on base-pairing between a degenerate 9-bp motif and the 5′ end of U1 small nuclear RNA (snRNA), such that overall splice-site strength—not solely the invariant GT dinucleotide at positions +1_2—contributes to splicing efficiency.17 Consistent with this principle, noncanonical GC donor sites occur naturally as WT in approximately 1% of human introns.18,19,20 Systematic analyses of GT>GC variants retaining residual WT transcript, including both disease-associated and artificially engineered variants, have shown that, within the context of the degenerate 9-bp motif, the corresponding C alleles tend to exhibit enhanced complementarity at surrounding positions compared with the G alleles, likely compensating for reduced base-pairing at the +2 position.12 Nevertheless, splice-site sequence context represents only one determinant of residual splicing efficiency, and additional regulatory features are likely to contribute.17,21 In this regard, the substantial discrepancies observed in splicing outcomes of GT>GC variants between full-length and minigene splicing assays underscore the importance of broader genomic context in splicing regulation.22 However, a comprehensive mechanistic framework accounting for this variability remains incomplete and is beyond the scope of the present study.
Here, using residual WT transcript levels as a quantitative measure, this study illustrates the difficulties inherent in classifying GT>GC variants that generate substantial yet incomplete amounts of WT transcript. By integrating transcript-level evidence, locus-wide analyses, assay-context comparisons, and database classifications, we highlight the challenges posed by intermediate functional effects and reinforce the need for classification frameworks capable of capturing the full continuum of variant effects.6,8
Material and methods
Identification and selection of GT>GC variants
Our 2019 study12 served as the starting point. From the FLGSA-analyzed GT>GC variants reported therein, we first selected variants that both retained residual WT transcript and were registered in ClinVar23 (https://www.ncbi.nlm.nih.gov/clinvar/; last accessed January 30, 2026) for analysis.
To identify additional disease-associated GT>GC variants with evidence of residual WT transcript generation beyond those analyzed in our 2019 study,12 we conducted a targeted literature search using Google Scholar and PubMed with the keywords “+2T>C” or “GT>GC.” This search was complemented by cross-checking entries in the Professional Version of the Human Gene Mutation Database (HGMD; https://www.hgmd.cf.ac.uk/)24 and by cascade reference tracking and cross-referencing within the peer-reviewed literature. This literature search was last updated on January 30, 2026.
In contrast to our previous work,12 which focused on RNA analyses performed on material derived from affected homozygous, hemizygous, or compound heterozygous individuals, we additionally screened individuals heterozygous for the variant in whom residual WT transcript derived from the variant allele was unequivocally demonstrated.
At the single-gene level, all 26 theoretically possible GT>GC variants in CFTR (MIM: 602421) were subjected to initial evaluation.
Finally, a BAP1 (MIM: 603089) variant, c.783+2T>C (GenBank: NM_004656.4, NG_031859.1), identified during the literature search as having conflicting classifications,25,26 was included for full-length and minigene splicing analyses.
In silico splicing predictions
In silico predictions were obtained using SpliceAI27 via the SpliceAI Lookup tool (https://spliceailookup.broadinstitute.org/; last accessed January 30, 2026) with the following settings: genome build hg38, GENCODE comprehensive annotation, maximum distance of 10,000 bp, and display of both REF and ALT scores. Among the four delta scores generated—Δ acceptor loss (ΔAL), Δ donor loss (ΔDL), Δ acceptor gain (ΔAG), and Δ donor gain (ΔDG)—the ΔDL score was used as a heuristic prioritization metric for downstream analyses, given that (1) all variants examined involved substitutions at the +2 position of canonical donor splice sites and (2) the analyses focused on variants’ ability to generate residual WT transcript. The remaining delta scores were examined only in selected instances to assess concordance between SpliceAI predictions and observed splicing outcomes.
Variant classification or description in public databases
Selected variants were queried in ClinVar (https://www.ncbi.nlm.nih.gov/clinvar/)23 to determine their classification status. The Clinical and Functional Translation of CFTR database (CFTR2; http://www.cftr2.org/)28 and the CFTR-France database (https://cftr.chu-montpellier.fr/)29 were consulted to evaluate CFTR variant classifications. Variant information from the Cystic Fibrosis Mutation Database (CFTR1; http://www.genet.sickkids.on.ca/) was used to assist in the interpretation of two CFTR variants that were not registered in either CFTR2 or CFTR-France. Variant classification or description data from all databases were last accessed on January 30, 2026.
Minigene splicing analysis of selected CFTR GT>GC variants
Variant selection
Of the 26 possible CFTR GT>GC variants (GenBank: NM_000492.4, NG_016465.4), three had a SpliceAI ΔDL score of ≤0.56, whereas all remaining variants had ΔDL scores >0.80. These three variants—c.2490+2T>C, c.3717+2T>C, and c.4242+2T>C—together with one additional variant, c.3873+2T>C, which had a high SpliceAI ΔDL score (0.98) but had previously been reported to generate residual WT transcript,30 were selected for minigene splicing analysis.
Construction of pET01 and pSPL3 minigene expression vectors
For each selected +2T>C variant, the corresponding WT genomic sequences cloned into the pET01 and pSPL3 exon trapping vectors22 were always identical. Each CFTR WT insert comprised 213–217 bp from the 3′ end of intron N − 1, the entire exon N, and 213–216 bp from the 5′ end of intron N (N = variant-affected intron). Cloning of WT constructs and introduction of the selected +2T>C variants, followed by Sanger confirmation, were performed by GENEWIZ Biotech (Suzhou, China).
Cell culture, transfection, RT-PCR, and sequencing
Experiments were carried out essentially as previously described.22 Human embryonic kidney 293T (HEK293T) cells were maintained in DMEM (Gibco) supplemented with 10% fetal calf serum (Procell). Cells (3.5 × 105/well) were seeded in 6-well plates 24 h before transfection. Per well, 2.5 μg of WT or variant plasmid were transfected using HieffTrans Universal Transfection Reagent (Yeasen). Forty-eight hours later, total RNA was extracted with the FastPure Cell/Tissue Total RNA Isolation Kit V2 (+gDNA wiper) (Vazyme). Reverse transcription (RT) was carried out using the HiScript III 1st Strand cDNA Synthesis Kit (Vazyme), incorporating 2 μL of 5× gDNA Wiper Mix, 2 μL of 10× RT Mix, 2 μL of HiScript III Enzyme Mix, 1 μL of Oligo(dT)20-VN, and 1 μg of total RNA. RT-PCR was performed in a 25-μL reaction mixture containing 12.5 μL of 2× Taq Master Mix (Vazyme), 1 μL of cDNA, and 0.4 μM each primer. The primer pairs were as follows: pET01 vector—forward 5′-GAGGGATCCGCTTCCTGGCCC-3′, reverse 5′-CTCCCGGGCCACCTCCAGTGCC-3′; pSPL3 vector—forward 5′-TCTGAGTCACCTGGACAACC-3′, reverse 5′-ATCTCAGTGGTATTTGTGAGC-3′. The PCR program comprised an initial denaturation step at 94°C for 5 min, followed by 35 cycles of denaturation at 94°C for 30 s, annealing at 55°C for 30 s, and extension at 72°C for 1 min, with a final extension step at 72°C for 7 min. RT-PCR products were separated on agarose gels, excised, and purified (Omega Bio-Tek Gel Extraction Kit). Sequencing primers were the same as those used for RT-PCR, and reactions were performed with the BigDye Terminator v.3.1 Cycle Sequencing Kit (Applied Biosystems).
Full-length and minigene splicing analysis of BAP1 c.783+2T>C
The full-length genomic BAP1 WT insert, spanning from the translation initiation codon to the termination codon (GenBank: NM_004656.4, NG_031859.1), was amplified from genomic DNA of a normal control using long-range PCR. The resulting fragment was cloned into the pcDNA3.1/V5-His-TOPO vector. These procedures, as well as subsequent transfection and RT-PCR, were performed essentially as previously described.12
Minigene splicing analysis followed the same procedures as for CFTR variants, except that two WT constructs (and their corresponding variant constructs) were generated. The BAP1 exon 9 insert comprised 337 bp from intron 8, exon 9, and 302 bp from intron 9; whereas the BAP1 exon 8–10 insert comprised 197 bp from intron 7, exon 8, intron 8, exon 9, intron 9, exon 10, and 243 bp from intron 10.
Results
Analytical framework and study design
Building on our 2019 study, which showed that approximately 15%–18% of GT>GC variants retain some residual WT transcript,12 we analyzed three successive variant categories (Figure 1). First, we re-examined GT>GC variants previously described in the 2019 study, encompassing both disease-associated variants and those characterized using FLGSA. Second, we considered additional disease-associated GT>GC variants that generate residual WT transcript beyond those analyzed in the 2019 study. Across these two categories, we assessed residual WT transcript levels alongside genetic and clinical findings, SpliceAI ΔDL scores, and ClinVar classifications. Third, we evaluated all theoretically possible GT>GC variants in CFTR and selected four variants for minigene splicing analysis. In addition, we attempted full-length and minigene splicing analyses of the BAP1 c.783+2T>C variant, which has conflicting classifications.25,26 By integrating findings from these analyses, we delineated the classification challenges posed by variants with intermediate functional effects, hereafter referred to conceptually as gray-zone variants.
Figure 1.
Schematic overview of the analytical framework
GT>GC variants were analyzed in three successive categories, together with full-length and minigene analyses of a BAP1 variant with conflicting classifications. For the first two variant categories, variants were evaluated across multiple features (yellow shading). The red box highlights the central issue addressed in this study.
Comparing disease-associated and FLGSA-analyzed GT>GC variants
In our 2019 study,12 the proportion of GT>GC variants retaining residual WT transcript was similar between disease-associated and FLGSA-analyzed variants: 15.6% (7 of 45) versus 18.4% (19 of 103). However, the magnitude of residual WT transcript levels differed markedly. The seven disease-associated variants generating residual WT transcript yielded only low levels, ranging from trace amounts up to 15% of normal. By contrast, among the ten FLGSA-analyzed nucleotide substitutions that generated exclusively WT transcripts (the remaining nine also produced additional bands), eight (80%) showed residual WT transcript levels ranging from 13.7% to 84% of normal, with only two falling below 5%.
The consistently low residual WT transcript levels observed for the seven disease-associated variants most likely reflect clinical ascertainment bias, as noted previously.12 Nonetheless, among the four variants retaining 10%–15% residual WT transcript, three—homozygous CAV3 (MIM: 601253) c.114+2T>C (GenBank: NM_033337.3, NG_008797.2),31 hemizygous CD40LG (MIM: 300386) c.346+2T>C (GenBank: NM_000074.3, NG_007280.1),32 and hemizygous DMD (MIM: 300377) c.8027+2T>C (GenBank: NM_004006.3, NG_012232.1)33—were associated with milder clinical phenotypes; the remaining variant, homozygous SPINK1 c.194+2T>C (GenBank: NM_001379610.1, NG_008356.2), is discussed below.
The corresponding full-length WT constructs for the 103 FLGSA-analyzed substitutions were invariably preselected to yield a single or quasi-single RT-PCR band of the expected WT transcript size, consistent with their respective reference mRNA sequences.12 We also analyzed HBB (MIM: 141900) c.315+2T>C (GenBank: NM_000518.5, NG_059281.1)34 using FLGSA, despite its full-length WT construct producing two strong bands—one corresponding to the WT transcript and the other representing an aberrantly spliced transcript. This variant also generated residual WT transcript in FLGSA,12 bringing the total number of residual WT transcript-generating GT>GC variants identified by FLGSA to 20.
Detailed analysis of three FLGSA-analyzed GT>GC variants
Of the 20 GT>GC variants previously shown by FLGSA12 to generate residual WT transcript, only three were registered in ClinVar. These three variants were selected for detailed analysis.
SPINK1 c.194+2T>C
Among GT>GC variants, SPINK1 c.194+2T>C (GenBank: NM_001379610.1, NG_008356.2) is distinctive in that transcript expression has been assessed both in vivo (using RNA from an affected homozygote)35 and in vitro using FLGSA.12 Both approaches yielded concordant results, indicating approximately 10% residual WT transcript from the variant allele.36 Consistently, SpliceAI predicts a ΔDL score of 0.33 for SPINK1 c.194+2T>C.
Our in-depth analysis of SPINK1-related chronic pancreatitis8 indicated that this ∼10% residual WT transcript places c.194+2T>C in the gray zone, conferring a well-defined genetic effect that is too weak to meet criteria for classification as “pathogenic” under current ACMG/AMP guidelines5 but too strong to be accommodated within the proposed “risk allele” category.7 We therefore suggested classifying this variant as “predisposing,” reflecting its partial LoF nature and intermediate genetic impact.8 ClinVar currently lists this variant as “pathogenic/likely pathogenic” with a two-star review status (VCV000132142.62; based on 12 out of 20 submissions). However, convergent evidence from functional assays, in silico prediction, and clinical and genetic data supports its role as a proof of concept highlighting the need for expanded variant classification frameworks.8
HBB c.315+2T>C
This variant was identified during routine thalassemia screening in a 34-year-old individual who exhibited hypochromia and microcytosis consistent with a thalassemia trait rather than β-thalassemia major (MIM: 613985). The individual was compound heterozygous for HBB c.315+2T>C and HBB c.79G>A (GenBank: NM_000518.5, NG_059281.1) (p.Glu27Lys).34 The original authors reported that HBB c.315+2T>C had a limited effect on splicing efficiency, stating that it “results in expression that is approximately two-thirds of the normal allele.”34 However, this estimate was based on hemoglobin quantitation and was likely confounded by the trans-inherited HBB c.79G>A (p.Glu27Lys) variant. The latter likely retains substantial normal function, as most individuals homozygous for HBB c.79G>A (p.Glu27Lys) exhibit only mild hypochromic microcytic anemia.37
Our FLGSA analysis confirmed that HBB c.315+2T>C generates residual WT transcript.12 Using ImageJ (https://imagej.net/ij/), we compared the intensity of the WT transcript band derived from the variant allele with that derived from the WT allele in the original gel.12 The variant allele was estimated to retain approximately 30% residual WT transcript. Notably, the orthologous rabbit variant also generated residual WT transcript,38,39 supporting conservation of this partial splicing effect across species. The SpliceAI ΔDL score for HBB c.315+2T>C was 0.72.
HBB alleles associated with β-thalassemia have traditionally been classified as β0 (complete loss of β-globin synthesis), β+ (moderate to severe reduction), or β++ (“silent,” minimal reduction).40 HBB c.315+2T>C was originally described as a mild β+ allele,34 consistent with our FLGSA finding of residual WT transcript generation.12 It is currently listed in ClinVar as “pathogenic” (VCV000869355.2) based on a single submission (review status: zero stars), a designation that may warrant re-evaluation in light of its residual WT transcript generation and the mild hematological phenotype observed in the 34-year-old compound heterozygote.
MGP c.94+2T>C
MGP (MIM: 154870) c.94+2T>C (GenBank: NM_000900.5, NG_023331.2) was shown by FLGSA12 to retain approximately 80% residual WT transcript, whereas SpliceAI predicted a high ΔDL score of 0.97. ClinVar lists MGP c.94+2T>C as “likely pathogenic” (VCV001987659.2; review status: one star), based on a single clinical testing submission (SCV003031993.2) with unknown affected status. Together, these observations highlight a discrepancy between predicted splice disruption, experimental splicing outcome, and current database classification for MGP c.94+2T>C.
Together, these three variants, summarized in Table 1, exemplify how GT>GC variants retaining residual WT transcript can pose significant challenges for variant classification.
Table 1.
Examples of GT>GC variants generating substantial residual WT transcripts and their ClinVar classifications
| Varianta | Reference mRNA and genomic sequences, GenBank: | hg38 coordinate | Residual WT transcriptb | Assessment method | SpliceAI ΔDL score | Clinical/genetic findings | ClinVar classification (review status) |
|---|---|---|---|---|---|---|---|
| SPINK1 c.194+2T>C | NM_001379610.1, NG_008356.2 | chr5:147828020A>G | 10% | RT-PCR of material derived from an affected homozygote35; FLGSA12 | 0.33 | well-defined genetic effect that does not meet criteria for a pathogenic classification8 | pathogenic/likely pathogenic (2★) |
| HBB c.315+2T>C | NM_000518.5, NG_059281.1 | chr11:5226575A>G | 30% | FLGSA12 | 0.72 | originally described as a mild β+-thalassemia allele34 | pathogenic (0★) |
| MGP c.94+2T>C | NM_000900.5, NG_023331.2 | chr12:14884211A>G | 80% | FLGSA12 | 0.97 | no confirmed disease cases reported | likely pathogenic (1★) |
| BRCA2 c.8331+2T>C | NM_000059.4, NG_012772.3 | chr13:32363535T>C | 61% | allele-specific RT-PCR of material derived from an affected heterozygote41 | 0.93 | BRCA2 history-weighting algorithm42 indicated a benign profile41; BRCA2 exhibits tolerance to reduced transcript dosage43 | pathogenic/likely pathogenic (2★) |
| DOCK8 c.3234+2T>C | NM_203447.4, NG_017007.1 | chr9:399261T>C | predominant WT transcript | RT-PCR and protein analyses of T cells from an affected homozygote44 | 0.49 | disease explained by an unlinked homozygous IFNAR1 nonsense variant44 | conflicting classifications of pathogenicity (2★) |
| NF1 c.888+2T>C | NM_000267.4, NG_009018.1 | chr17:31182667T>C | only WT transcript detected | RNA sequencing of the heterozygous individual and appropriately matched controls in parallel45 | 0.18 | the heterozygous individual had no clinical features consistent with neurofibromatosis type 145 | conflicting classifications of pathogenicity (1★) |
ΔDL, delta donor loss; FLGSA, full-length gene-splicing assay; RT-PCR, reverse transcription-PCR; WT, wild-type.
Variants are arranged in order of mention in the main text.
Percentage of WT transcript generated from the variant allele relative to that from the WT allele.
Additional disease-associated GT>GC variants
We next screened disease-associated GT>GC variants beyond those analyzed in our 2019 study, focusing on those for which residual WT transcript generation had been documented in material derived from affected individuals. All candidate reports were carefully reviewed, and only those in which residual WT transcript generation had been unequivocally demonstrated were included, with particular attention being paid to individuals heterozygous for the variant.
For example, heterozygous RB1 (MIM: 614041) c.2489+2T>C (GenBank: NM_000321.3, NG_009009.1)46 was excluded from further analysis because residual WT transcript generation from the variant allele had not been unequivocally demonstrated. This variant was identified in an individual diagnosed with retinoblastoma at 9 months of age and was considered “benign” by the original authors based on the observation that RT-PCR analysis of a blood sample from the affected individual revealed only WT transcript.46 However, aberrantly spliced transcripts derived from the variant allele may not have been detected due to nonsense-mediated mRNA decay (NMD).47 Indeed, RB1 c.2489+2T>C was predicted by SpliceAI to have a ΔAL score of 0.24 and a ΔDL score of 0.96, suggesting exon 23 skipping. Skipping of the 164-bp exon 23 would introduce a premature stop codon and could therefore be expected to trigger NMD. Thus, the possibility that the observed WT transcript was derived exclusively from the WT allele could not be excluded.
Our literature search identified seven additional variants in which residual WT transcript generation was unequivocally demonstrated. While this search was designed to be comprehensive, it cannot be considered exhaustive. Among these seven variants, homozygous COL6A1 (MIM: 120220) c.227+2T>C (GenBank: NM_001848.3, NG_008674.1),48 homozygous SBDS (MIM: 607444) c.258+2T>C (GenBank: NM_016038.4, NG_007277.1),49 and compound heterozygous SRP68 (MIM: 604858) c.184+2T>C (GenBank: NM_014230.4, NC_000017.11)—the second allele being a large deletion removing exon 1 of SRP6850—produced only ∼2%–5% residual WT transcript, with SpliceAI ΔDL scores of 0.88, 0.93, and 0.37, respectively. Heterozygous TNXB (MIM: 600985) c.12469+2T>C (GenBank: NM_001365276.2, NG_008337.2) retained a moderate level of WT transcript by allele-specific RT-PCR, consistent with its very low ΔDL score of 0.01, and has been associated with moderate Ehlers-Danlos-like manifestations in congenital adrenal hyperplasia.51 The remaining three variants are discussed below.
BRCA2 c.8331+2T>C
BRCA2 (MIM: 600185) c.8331+2T>C (GenBank: NM_000059.4, NG_012772.3) is associated with a high SpliceAI ΔDL score of 0.93 and is listed in ClinVar as “pathogenic/likely pathogenic” (VCV000267692.71) with a two-star review status (based on 15 out of 17 submissions). However, employing a cis-inherited heterozygous variant (BRCA2 c.7242A>G) as an allelic marker, allele-specific RT-PCR performed in an individual heterozygous for BRCA2 c.8331+2T>C demonstrated that 38% of the WT transcript originated from the variant allele and 62% from the WT allele.41 This corresponds to the variant allele retaining the capacity to generate approximately 61% of normal WT transcript levels. A validated BRCA1/2 history-weighting algorithm, which compares observed cancer histories with expectations for pathogenic versus benign variants,42 yielded a benign call.41
In addition, BRCA2 appears relatively tolerant of reduced transcript dosage, and one study suggested that variants producing more than 5% residual WT transcript should, in the absence of other supporting evidence, be considered VUS.43 While the original report reclassified c.8331+2T>C from “pathogenic” to a “VUS,”41 the combined functional, genetic, and clinical evidence appears to favor a benign interpretation under current ACMG/AMP principles.5
DOCK8 c.3234+2T>C
Biallelic LoF variants in DOCK8 (MIM: 611432) cause autosomal-recessive hyper-IgE syndrome type 2 (MIM: 243700). DOCK8 c.3234+2T>C (GenBank: NM_203447.4, NG_017007.1) was reported in the homozygous state in an infant with severe viral infections.44 T cells from this individual showed a predominance of WT transcript, normal DOCK8 protein accumulation, intact cellular function, and no DOCK8 deficiency phenotype. The clinical presentation was instead explained by a homozygous IFNAR1 (MIM: 107450) nonsense variant, c.1156G>T (GenBank: NM_000629.3) (p.Glu386Ter). On this basis, the original authors concluded that DOCK8 c.3234+2T>C was likely neutral with respect to the disease phenotype in the studied individual.44 This variant was predicted by SpliceAI to have a modest ΔDL score of 0.49 and is currently listed in ClinVar as having “conflicting classifications of pathogenicity” with a two-star review status (VCV000573664.15).
NF1 c.888+2T>C
Heterozygous NF1 (MIM: 613113) c.888+2T>C (GenBank: NM_000267.4, NG_009018.1), located in intron 8, was identified in an individual with early-onset breast cancer but without reported clinical features consistent with neurofibromatosis type 1 (MIM: 162200).45 RNA sequencing performed in this individual, together with a normal control and an individual heterozygous for NF1 c.888+1G>A, indicated that NF1 c.888+2T>C does not measurably disrupt splicing.45
Sashimi plot analysis revealed no aberrant splicing events exceeding a PSI (percent splicing index) threshold of 2% in either the NF1 c.888+2T>C heterozygote or the control individual. By contrast, the individual heterozygous for NF1 c.888+1G>A analyzed in parallel exhibited two aberrantly spliced transcripts in addition to WT transcript, namely an in-frame transcript retaining the first 60 bp of intron 8 and an out-of-frame transcript resulting from exon 8 skipping. Both aberrant splice products were supported by substantial read counts and were consistent with the high SpliceAI-predicted ΔDG score of 0.93 and ΔDL score of 0.95 for NF1 c.888+1G>A (Table 2).
Table 2.
RNA-sequencing findings and SpliceAI Δ scores for two NF1 splice-site variants
| Samples analyzeda | Variant hg38 coordinate | RNA-sequencing resultsb |
ΔAL | ΔDLc | ΔAG | ΔDGc | |
|---|---|---|---|---|---|---|---|
| Detected transcript(s) | Read number | ||||||
| NF1 c.888+2T>C heterozygote | chr17:31182667T>C | WT transcript | 19,081 | 0.05 | 0.18 (−2 bp) | 0.00 | 0.75 (58 bp) |
| NF1 c.888+1G>A heterozygote | chr17:31182666G>A | WT transcript | 4,201 | 0.05 | 0.95 (−1 bp) | 0.00 | 0.93 (59 bp) |
| aberrant transcript retaining the first 60 bp of intron 8 (in-frame) | 2,440 | ||||||
| aberrant transcript with exon 8 skipped (out-of-frame) | 550 | ||||||
| Normal control | – | WT transcript | 11,323 | – | – | – | – |
ΔAL, delta acceptor loss; ΔDL, delta donor loss; ΔAG, delta acceptor gain; ΔDG, delta donor gain; WT, wild-type.
Original data from Zimmermann et al.45 Reference mRNA and genomic sequences for NF1 are GenBank: NM_000267.4 and NG_009018.1.
Splice events with PSI (percent splicing index) <2% were filtered out.
Values in parentheses indicate the distance (bp) from the variant position. Negative values, distance 5′ to the variant position; positive values, distance 3′ to the variant position.
By comparison, SpliceAI predicted markedly lower splice-disrupting scores for NF1 c.888+2T>C, with a ΔDG score of 0.75 and a ΔDL score of only 0.18 (Table 2). Given that both variants affect the same donor splice site of intron 8, any splice-altering effect of NF1 c.888+2T>C may be expected to generate one or both of the aberrant transcripts observed for NF1 c.888+1G>A. The absence of such aberrant transcripts in the NF1 c.888+2T>C heterozygote, in the context of robust sequencing depth across all three individuals (Table 2), further supports a lack of detectable splice disruption.
The original authors interpreted NF1 c.888+2T>C as a benign variant.45 ClinVar currently lists this variant as having “conflicting classifications of pathogenicity” (VCV001067639.6; review status: one star), with one submission classifying it as “pathogenic” and another as of “uncertain significance.”
Together, these variants further illustrate how GT>GC variants retaining residual WT transcript can complicate variant interpretation across diverse genes, inheritance contexts, and clinical phenotypes. These variants are also summarized in Table 1.
Locus-wide analysis of CFTR GT>GC variants
Having examined the classification challenges posed by GT>GC variants that retain WT transcript across multiple genes, we next explored how frequent and systematic this phenomenon might be within a single locus. We selected CFTR, in which biallelic LoF variants cause cystic fibrosis,52,53 as a model system. CFTR harbors 26 GT donor splice sites (Table S1), is among the most extensively studied loci in human genetics, and benefits from expert-curated databases, including CFTR228 and CFTR-France,29 that provide authoritative variant annotations.
Building upon our previous36 and current findings, we observed that although some GT>GC variants retaining substantial residual WT transcript are associated with high SpliceAI ΔDL scores, most exhibit low ΔDL scores. We therefore used SpliceAI ΔDL scores as an initial heuristic for this analysis, complemented by cross-examination of variant classifications in CFTR2 and CFTR-France (Figure 2A). Notably, all 16 CFTR GT>GC variants that we grouped as “pathogenic” had SpliceAI ΔDL scores greater than 0.80. These included 13 designated as “CF-causing” in CFTR2, one absent from CFTR2 but listed as “likely pathogenic” in CFTR-France (c.3468+2T>C), and two absent from both CFTR2 and CFTR-France but classified by us as “pathogenic” based on clinical and genetic data in CFTR1 (c.2908+2T>C and c.2988+2T>C).
Figure 2.
Analysis of GT>GC substitutions in CFTR
(A) Flowchart summarizing the 26 possible CFTR GT>GC substitutions categorized by SpliceAI ΔDL scores and database information (see Table S1 for details). Variants with low ΔDL scores (≤0.56) were selected for minigene splicing assays.
(B) Four CFTR GT>GC variants analyzed by minigene splicing assays. Under “Database classification and reporting,” note that, unlike CFTR2 and CFTR-France, CFTR1 does not provide formal variant classifications. Under “Residual WT transcript,” “yes” and “no” indicate whether the variant allele generated residual WT transcript, as assessed by a previously reported expression minigene (EMG) assay30 and by minigene splicing assays performed in the present study. Results from this study are shown separately for the two assay systems (pET01 and pSPL3; see Figure 3), with “No in both assays” indicating absence of detectable residual WT transcript from the variant allele in either system. CBAVD, congenital bilateral absence of the vas deferens; CF, cystic fibrosis; CCP, conflicting classifications of pathogenicity; EMG, expression minigene; ΔDL, delta donor loss; P, pathogenic; P/LP, pathogenic/likely pathogenic; VUS, variant of uncertain significance; WT, wild-type.
Eight CFTR GT>GC variants were not registered in any of these three databases. Of these, seven had SpliceAI ΔDL scores above 0.80, whereas one variant had a score of 0.56 (c.3717+2T>C). The remaining two CFTR GT>GC variants—one listed as a VUS in CFTR-France (c.2490+2T>C) and one reported in CFTR1 in two individuals with suspected cystic fibrosis (c.4242+2T>C)—had SpliceAI ΔDL scores below 0.56 (Figures 2A and 2B). Full details for all 26 CFTR GT>GC variants, including the rationale for our classifying two as “pathogenic,” are provided in Table S1.
Among the three CFTR variants with relatively low SpliceAI ΔDL scores, only c.4242+2T>C had previously undergone functional analysis. Using an expression minigene (EMG) system comprising the full CFTR coding sequence and intronic regions tailored to the variant under study, generation of residual WT transcript from the c.4242+2T>C allele was confirmed by Sanger sequencing, and mature CFTR protein was detected in transfected cells by immunoblotting.30 The only other CFTR GT>GC variant reported to yield WT transcript is c.3873+2T>C, although this was mentioned only briefly and without accompanying experimental data.30 Notably, c.3873+2T>C is classified as “CF-causing” in both CFTR2 and CFTR-France, consistent with its high SpliceAI ΔDL score (0.98) (Figure 2B).
Using SpliceAI ΔDL scores as a triage heuristic rather than a definitive predictor, we prioritized CFTR GT>GC variants with relatively low or intermediate scores for further functional evaluation. This strategy was designed to enrich for variants potentially capable of retaining residual WT transcript while recognizing that SpliceAI cannot reliably discriminate between complete and partial splice disruption, particularly for non-all-or-none effects. We therefore subjected the three variants with low SpliceAI ΔDL scores, together with c.3873+2T>C, to minigene assays using both pET01 and pSPL3 vectors (Figures 3 and S1). The results for each variant are detailed below.
Figure 3.
Minigene splicing analysis of four selected CFTR GT>GC substitutions using pET01 and pSPL3 vectors
Each variant (A–D) was analyzed with its preceding exon and short flanking intronic sequences cloned into the respective reporter vectors. Arrows indicate bands corresponding to the normally spliced transcript, in accordance with the CFTR mRNA reference sequence GenBank: NM_000492.4. Additional bands are annotated in Figure S1.
c.2490+2T>C (intron 14)
WT constructs in both minigene vectors generated multiple RT-PCR products, including the expected WT transcript. Notably, residual WT transcript derived from the variant allele was detected in the pET01 vector (Figure 3A), consistent with the modest SpliceAI ΔDL score of 0.35. Although c.2490+2T>C is not classified by CFTR2, it is listed in CFTR-France as a VUS, having been reported in two asymptomatic homozygotes. In CFTR1, the variant was described in the heterozygous state in the partner of an individual affected by congenital bilateral absence of the vas deferens (Figure 2B).
ClinVar lists c.2490+2T>C as having “conflicting classifications of pathogenicity” (VCV000370078.23; review status: one star), based on six clinical-testing submissions (three “pathogenic,” two “likely pathogenic,” and one of “uncertain significance”). Of these submissions, only one specified a known affected status, reporting homozygosity for c.2490+2T>C in an individual with fetal cystic hygroma (SCV001737053.1). By contrast, another submission classified the variant as “likely pathogenic” while explicitly noting the absence of reported occurrences in individuals affected with cystic fibrosis and the lack of experimental evidence demonstrating an impact on CFTR protein function (SCV001363801.5). Taken together, the available clinical, genetic, in silico, and experimental data argue against c.2490+2T>C representing a pathogenic CFTR variant in the context of classic cystic fibrosis.
Several methodological and interpretive considerations regarding the minigene splicing assays merit clarification. First, because both WT and variant constructs generated multiple RT-PCR products, we did not attempt to quantify the proportion of residual WT transcript produced from the variant allele in the pET01 context. Second, the discrepant outcomes observed between the pET01 and pSPL3 systems underscore the strong context dependence of minigene-based splicing assays. Differences in vector architecture, exon-intron configuration, and flanking sequence length can influence splice-site recognition and the relative use of competing splice sites. Notably, detection of residual WT transcript in the pET01 system is consistent with the relatively low SpliceAI ΔDL score (0.35), whereas the aberrant transcript observed in the pSPL3 system, which retains the first four bases of intron 14 (Figure S1), is concordant with the high SpliceAI ΔDG score (0.84), predicted to reflect activation of a cryptic GT dinucleotide at c.2490+5_6. Third, higher-resolution approaches such as next-generation sequencing (NGS) of RT-PCR products could provide more detailed qualitative and quantitative characterization of transcript species than gel-based analysis. However, such approaches would not resolve the inherent sequence-context limitations of minigene assays.
c.3717+2T>C (intron 22)
Under the experimental conditions tested, c.3717+2T>C does not generate residual WT transcript (Figure 3B), despite its intermediate SpliceAI ΔDL score (0.56). This variant is not listed in CFTR2, CFTR-France, CFTR1, or ClinVar (Figure 2B).
c.3873+2T>C (intron 23)
No WT transcript from the variant allele was observed in either minigene vector (Figure 3C). This finding is consistent with its high SpliceAI ΔDL score (0.98) and the designation of c.3873+2T>C as “CF-causing” in both CFTR2 and CFTR-France (Figure 2B). The variant is also classified as “pathogenic” in ClinVar (VCV000053828.18; review status: three stars), based on an expert panel submission from CFTR2. Moreover, consistent with its description in CFTR1, this variant is among the CFTR alleles recurrently identified in individuals with clinically diagnosed cystic fibrosis in Norway.54 Taken together, these data support c.3873+2T>C as a pathogenic CFTR variant causing classic cystic fibrosis. Any residual WT transcript generation from this allele, as previously reported,30 is therefore likely minimal.
c.4242+2T>C (intron 26)
No WT transcript from the variant allele was detected in either minigene vector (Figure 3D), in contrast to the earlier EMG report describing residual WT transcript generation and CFTR protein production30 and the very low SpliceAI ΔDL score for this variant (0.15). Notably, c.4242+2T>C is not listed in CFTR2 or CFTR-France but is reported in CFTR1 in two individuals with suspected cystic fibrosis (Figure 2B).
ClinVar lists c.4242+2T>C as “pathogenic/likely pathogenic” (VCV000035884.8; review status: two stars), based on three submissions from clinical genetic testing laboratories (one “pathogenic” and two “likely pathogenic). For two submissions, the collection method was described as “clinical testing” (SCV002630543.2 and SCV004675881.2), with affected status reported as unknown, whereas the remaining submission was described as “curation” (SCV000052191.3). In the absence of unequivocal clinical or functional evidence supporting pathogenicity, this variant is best considered as a VUS at present.
One methodological consideration is particularly relevant for interpreting these discrepant findings. In the previously reported EMG assay of c.4242+2T>C,30 the CFTR insert cloned into the pcDNA5/FRT vector comprised exons 1–25, intron 25, exon 26, intron 26, and exon 27, thereby preserving an extended native genomic context. This configuration differs markedly from the minigene systems used in the present study, in which shorter CFTR fragments were inserted into artificial sequence contexts. Such differences in local genomic context may account for the divergent splicing outcomes observed between the EMG and minigene assays. By contrast, all three vector systems are driven by the CMV promoter, and transfections were performed in either HEK293Flp or HEK293 cells, making contributions from promoter usage or cellular background unlikely to be significant.
In summary, locus-wide analysis of CFTR indicates that a subset of GT>GC variants retain the capacity to generate residual WT transcript. Functional testing of selected CFTR GT>GC variants revealed heterogeneous and context-dependent splicing outcomes, underscoring both inter-assay variability and the inherent limitations of minigene systems in fully recapitulating native splicing.22
Challenges of splicing assays: The case of BAP1 c.783+2T>C
As illustrated by the CFTR variants analyzed above, in vitro splicing assays may fail to yield clear or consistent results, thereby complicating variant classification. This challenge is further exemplified by BAP1 c.783+2T>C (GenBank: NM_004656.4, NG_031859.1). Because heterozygous germline LoF variants in BAP1 underlie tumor predisposition syndrome 1 (TPDS1; MIM: 614327) and BAP1 c.783+2T>C affects the canonical donor splice site of intron 9, this variant was initially classified in ClinVar as “likely pathogenic,” as noted by Goldberg et al.25 These authors subsequently reclassified the variant as a VUS based on the following considerations:
-
(1)
the variant was identified in seven families (six Jewish Iraqi and one non-Jewish) presenting with nonclassical TPDS1 tumor types;
-
(2)
the variant exhibited low penetrance;
-
(3)
the variant had a relatively high allele frequency in Jewish Iraqi individuals (3.6%);
-
(4)
the variant had a low SpliceAI score (0.08); and
-
(5)
no aberrantly spliced transcripts were detected by RT-PCR using RNA from blood cells of two probands heterozygous for BAP1 c.783+2T>C.
Consistent with this uncertainty, ClinVar currently lists c.783+2T>C under “conflicting classifications of pathogenicity” (VCV000422670.29; review status: one star), with five submissions reporting “uncertain significance” and one reporting “likely benign.”
A subsequent study by Sculco et al.26 identified a methodological limitation in the analysis by Goldberg et al.25 Specifically, the forward primer used by Goldberg et al. was located within exon 9, precluding detection of exon 9 skipping. Using a forward primer located in exon 8, Sculco et al. analyzed BAP1 c.783+2T>C heterozygotes and observed an alternative splicing pattern consistent with exon 9 skipping. On this basis, they proposed reinstating a “likely pathogenic” classification, while acknowledging that the effect of BAP1 c.783+2T>C on aberrant transcript generation might be incomplete.26 Indeed, any residual WT transcript derived from the c.783+2T>C allele could have been masked by transcript from the unaffected allele in heterozygous individuals. Importantly, skipping of exon 9 (c.660_783) would introduce a premature termination codon at c.661_663. Such a transcript would be predicted to undergo NMD47 or, if translated, to truncate the 729-amino-acid BAP1 protein after residue 220. In either scenario, meaningful residual function from exon 9-skipped transcripts is highly unlikely.
To further clarify this issue, we performed additional splicing assays (Figures 4 and S2). We first attempted FLGSA, as the BAP1 genomic sequence from start to stop codon (<8 kb) falls within the upper size limit of the insert cloned into the pcDNA3.1/V5-His-TOPO vector routinely used in this assay.12 However, the WT construct failed to yield clear RT-PCR bands following transfection into HEK293T cells (Figure 4A). Notably, in our previous study, only 33 of 119 (28%) WT full-length constructs produced a single or quasi-single band of the expected size after transfection.12 We therefore turned to minigene assays using the pET01 and pSPL3 vectors.
Figure 4.
Splicing analysis of the BAP1 c.783+2T>C variant in different assay contexts
(A) Full-length gene-splicing assay (FLGSA). The wild-type (WT) construct failed to generate clear reverse transcription-PCR (RT-PCR) bands.
(B) Minigene analysis of exons 8–10 (including exon 8, intron 8, exon 9, intron 9, and exon 10) with short flanking sequences from introns 7 and 10 cloned into the pET01 and pSPL3 vectors. No normally spliced transcripts were detected, in accordance with the BAP1 mRNA reference sequence GenBank: NM_004656.4, even from the WT constructs. For band annotations, see Figure S2.
(C) Minigene analysis of exon 9 with short flanking intronic sequences cloned into the pET01 and pSPL3 vectors. Arrows indicate bands corresponding to the normally spliced transcript, in accordance with the BAP1 mRNA reference sequence GenBank: NM_004656.4. Additional bands are annotated in Figure S2.
BAP1 genomic reference sequence GenBank: NG_031859.1.
WT constructs spanning genomic sequence from exon 8 to exon 10 (including exon 8, intron 8, exon 9, intron 9, and exon 10), with short flanking sequences from introns 7 and 10, produced multiple RT-PCR bands, none of which corresponded to the WT transcript defined by GenBank: NM_004656.4. The corresponding variant constructs also failed to generate a normally spliced transcript, although their banding patterns differed from those of the WT constructs (Figure 4B). In contrast, shorter constructs containing exon 9 with short flanking intronic sequences generated the expected WT band from the WT allele, whereas only transcripts lacking exon 9 were produced from the variant allele (Figures 4C and S2). While this latter result is consistent with the findings of Sculco et al.,26 the restricted sequence context of these minigene constructs precludes exclusion of residual WT transcript generation under native genomic conditions.
Given these inconclusive and context-dependent findings, definitive resolution of the splicing effect of BAP1 c.783+2T>C would likely require approaches capable of resolving allele-specific splicing effects, such as analyses enabled by cis-linked exonic variants or targeted knockin cell models.
Discussion
Functional and genetic effect spectrum of GT>GC variants
Building on our 2019 study,12 the present analysis further delineates the functional and genetic impact spectrum of GT>GC 5′ splice-site variants. Consistent with findings from FLGSA,12 we show that disease-associated GT>GC variants can generate a wide range of residual WT transcript levels, spanning from barely detectable to levels approaching normal. Importantly, our earlier postulate that “GT>GC variants in human disease genes may not invariably be pathogenic”12 has since been substantiated by subsequent studies, as exemplified by the DOCK8 c.3234+2T>C and NF1 c.888+2T>C variants.
Beyond an all-or-none view of splice disruption, our findings demonstrate that GT>GC variants generating substantial yet incomplete levels of WT transcript often fall within a functional gray zone, complicating their classification (Figure 5). Such intermediate effects expose the limitations of binary pathogenic versus benign frameworks and have direct clinical implications. A representative example is SPINK1 c.194+2T>C, which retains approximately 10% of normal WT transcript. In the simple heterozygous state, its effect is too weak to be classified as “pathogenic” yet too strong to be regarded as a mere “risk” allele. However, in homozygous, compound heterozygous, or transheterozygous states—or in the presence of environmental modifiers—the same variant confers a substantially greater pathogenic effect.8 This example underscores how variants with intermediate functional effects contribute to the genetic complexity of human disease and highlights the need for classification frameworks that can accommodate graded effects rather than binary outcomes.6,8
Figure 5.
Conceptual model of the functional gray zone in variant classification
Residual wild-type (WT) transcript levels and variant classification lie along continua rather than binary states. Variants with intermediate WT transcript levels often fall within a functional gray zone, where in silico prediction and experimental analysis become increasingly difficult. Such gray-zone variants may act through variant-variant (both intra- and intergenic) or variant-environment interactions, thereby contributing to the genetic complexity of human disease. Although illustrated here with splice-altering variants, this conceptual model also applies to other variant types and disease mechanisms shaped by graded levels of functional impact.
Limitations of in silico prediction for intermediate splice effects
Our findings further underscore the limited resolution of current in silico tools, such as SpliceAI,27 for predicting whether GT>GC variants generate residual WT transcript and, if so, to what extent.36 Whereas most variants retaining substantial residual WT transcript exhibit low SpliceAI ΔDL scores, a subset displays high values (>0.80), and no consistent quantitative relationship emerges between predicted scores and experimentally observed residual WT transcript levels.
This limitation partly reflects a fundamental biological feature of splicing, namely the intrinsic capacity of GC dinucleotides to function as natural donor sites in approximately 1% of U2-type introns.18,19,20 More broadly, it highlights the difficulty of computationally modeling splice-site behavior that deviates from all-or-none outcomes. Variants producing intermediate effects therefore tend to occupy a functional gray zone that current prediction algorithms cannot readily model or quantify.
Experimental challenges in assessing splice-altering variants
Our results also highlight the methodological challenges inherent in experimentally characterizing splice-altering variants. RNA from disease-relevant tissues provides the most reliable readout, yet such material is rarely available, and surrogate sources such as blood cells or lymphoblastoid lines are informative only if the gene of interest is expressed.55,56 In the absence of RNA from affected individuals, cell-based assays—such as FLGSA,12,57 EMG,30 maxigene58 or midigene,59 or conventional minigene60,61 systems—are often used, but each has intrinsic limitations. For example, not only is FLGSA impractical for large genes, but also not all successfully constructed full-length WT constructs yield the expected WT transcript. The most frequently used minigene assays are performed in a highly artificial background that lacks the natural genomic context required for proper splicing regulation.62,63 Such loss of context would be particularly problematic for variants with intermediate effects, for which classification depends not only on detecting aberrant transcripts but also on quantifying residual WT transcript from the variant allele.
Even when RNA from affected individuals is available, interpretation in heterozygotes can be confounded by WT transcript derived from the unaffected allele. For this reason, in our previous study,12 analyses of disease-associated GT>GC variants focused primarily on homozygotes, hemizygotes, and compound heterozygotes. Here, we expanded the analyses to individuals heterozygous for the variant and identified informative cases owing to the availability of cis-inherited variants that enabled allele-specific RT-PCR analyses41,51 or simultaneous RNA sequencing of index subjects and appropriately matched controls.45 However, RNA sequencing—although increasingly applied in clinical contexts64,65,66—faces persistent challenges related to tissue accessibility and allele-specific quantification. Collectively, these methodological constraints underscore the need for refined experimental systems, including approaches based on saturation genome editing67 and transactivation or transdifferentiation.68
Residual WT transcript as a generalizable phenomenon
GT>GC variants that retain residual WT transcript reflect the degeneracy and context dependence of splice-site recognition. As such, findings from this study can likely be extrapolated to many other splice-altering variants—excluding those at the invariant +1, −1, and −2 positions—because cis-acting motifs are degenerate, and deep intronic variants can generate novel splice sites that compete with physiological ones. Although a comprehensive survey is beyond the scope of this study, several observations support this view. In SpliceVarDB, approximately 50% of variants could not be assigned unambiguously to “splice-altering” or “non–splice-altering” categories but instead fell into a “low-frequency splice-altering” group, corresponding to weak or indeterminate evidence of spliceogenicity,69 suggesting that gray-zone variants are common. Among 12 FLGSA-analyzed SPINK1 coding sequence variants showing aberrant splicing, nine (including presumed missense and synonymous variants) also generated residual WT transcript.70 Likewise, deep intronic variants such as CFTR c.1584+18672A>G (GenBank: NM_000492.4, NG_016465.4)71 and CLRN1 (MIM: 606397) c.254−643G>T (GenBank: NM_174878.3, NG_009168.1)72 have been reported to yield both aberrant and WT transcripts. Collectively, these examples indicate that residual WT transcript generation is a recurrent and underappreciated phenomenon across genes and variant types, producing outcomes along a continuum rather than at binary extremes (Figure 5).
Because residual WT transcript cannot always be readily assessed in simple or compound heterozygotes, and splice-altering variants account for an estimated 10%–30% of all disease-associated alleles,11,24,69 we infer that variants generating residual WT transcript may represent a significant source of variable penetrance and expressivity.
Gene and context dependence of the functional gray zone
The boundaries of the functional gray zone (Figure 5) are inherently gene dependent and context dependent, reflecting differences in functional tolerance thresholds and modes of inheritance. Consequently, an identical level of residual WT transcript may be clinically inconsequential in one gene-disease context yet functionally significant in another. An illustrative example comes from a recent study reporting unexpectedly high levels of normally spliced transcripts in an individual with a recessive form of skeletal dysplasia, short stature, amelogenesis imperfecta, and skeletal dysplasia with scoliosis (MIM: 618363), who was compound heterozygous for two SLC10A7 (MIM: 611459) splice-altering variants, c.472−1G>T and c.722−16A>G (GenBank: NM_001029998.6, NC_000004.12). Approximately 45% of SLC10A7 transcripts in the affected individual were normally spliced.73 Such a level of residual WT transcript would be unlikely to cause cystic fibrosis if observed for CFTR, underscoring how gene-specific functional thresholds critically shape clinical outcomes.
Although developed here primarily in the context of LoF variants for conceptual simplicity, the gray-zone model may also extend to gain-of-function or dominant-negative variants,74 in which disease-associated alleles within a given gene are likewise unlikely to exert uniform effects. Notably, this conceptual framework is most readily applicable to genes in which complete or near-complete LoF alleles cause classical Mendelian diseases. As we have previously opined,8 some disease-associated genes may not necessarily harbor clearly “pathogenic” variants.
Implications for variant classification frameworks
Taken together, our study establishes GT>GC variants as a tractable model for interrogating the classification challenges posed by splice-altering variants with intermediate functional consequences. By using residual WT transcript as a measurable parameter—and integrating clinical and genetic data, RNA analyses of material derived from affected individuals, full-length and minigene assays, and computational predictions—we show that residual WT transcript generation is a recurrent and biologically meaningful phenomenon that frequently complicates variant classification.
Gradual functional effects are also observed for other variant types, as exemplified by missense variants in BRCA1 (MIM: 113705)67 and SOD1 (MIM: 147450),75 both analyzed using multiplexed assays of variant effect (MAVE). These observations underscore the need to move beyond binary pathogenic/benign frameworks toward models that explicitly accommodate a continuum of functional effects.6,8 Efforts focused solely on refining ACMG/AMP criteria within a binary framework (as reflected in several recent studies76,77,78,79) are therefore unlikely to resolve the fundamental challenge, which has been amplified by the field’s transition from a phenotype-first to a genotype-first approach.
Emerging conceptual, computational, and experimental strategies offer promising avenues forward. As highlighted by Whitcomb,80 our recently proposed “pathogenic-predisposing-risk-benign” framework8 represents an important step toward resolving the variant classification challenge. The integration of machine-learning models trained on large-scale population-based datasets9 with functional data generated from MAVE81,82 holds great promise for defining the full spectrum of variant effects revealed through analyses of increasingly large cohorts.83,84,85 Ultimately, gray-zone variants may constitute a substantial portion of those currently classified as VUS or excluded for exceeding the allele-frequency thresholds set for high-penetrance disease variants. Such variants could also account for part of the “missing heritability” observed in both rare and complex diseases.86,87,88 Recognizing the widespread presence of gray-zone variants and their contribution to genetic complexity will be essential for realizing the full potential of precision medicine.89
A workflow for evaluating GT>GC variants suspected to generate residual WT transcript
Although the boundaries of a functional gray zone are inherently gene dependent and context dependent, the analyses presented here support a structured framework for evaluating GT>GC variants—and, more broadly, splice-altering variants—suspected to generate residual WT transcript.
First, transcript-level evidence from disease-relevant tissues should be prioritized whenever available, ideally enabling separation of transcripts derived from the variant allele versus the WT allele. Such data provide the most direct and physiologically relevant assessment of residual splicing activity.
Second, in the absence of patient-derived material, assays should be applied hierarchically rather than interchangeably. Full-length gene assays best preserve native genomic architecture; maxigene or midigene systems offer a balance between contextual fidelity and tractability; and minigene assays, while accessible, require cautious interpretation due to their artificial sequence environment. Where discrepancies arise between assay systems or when qualitative assays cannot determine whether low-level residual WT transcript is present, higher-resolution approaches such as NGS may provide additional insight. These methods can detect complex or low-abundance splicing outcomes not readily resolved by gel-based analyses. However, NGS-based quantification remains constrained by construct design and experimental context and must be interpreted accordingly.
Third, qualitative detection of WT transcript must not be equated with quantitative sufficiency. Low-level residual splicing may escape detection depending on assay sensitivity and read depth, whereas modest differences in residual WT transcript abundance may translate into substantially different functional consequences across genes.
Finally, where gene-specific tolerance to reduced transcript dosage is known, residual WT transcript levels should be interpreted in light of those thresholds rather than through binary pathogenic versus benign classification alone.
Collectively, this framework provides a principled approach for integrating transcript-level evidence with existing classification systems and for evaluating splice-site variants that occupy an intermediate functional spectrum.
Study limitations and concluding remarks
This study has several limitations. First, the number of GT>GC 5′ splice-site variants for which substantial residual WT transcript has been documented remains limited. Second, residual WT transcript levels estimated by FLGSA or RT-PCR using material derived from affected individuals may not accurately reflect true in vivo splicing outcomes in pathologically relevant tissues. Third, a potential alternative impact of GT>GC variants—namely, a reduction in overall splicing efficiency without producing aberrant transcripts—could not be adequately assessed. Likewise, variants that generate both WT and aberrant transcripts, but in which the aberrant transcripts are degraded by NMD,90 could not be evaluated. These two scenarios could lead to underestimation of the prevalence of partial LoF GT>GC variants. Finally, the examples presented here illustrate recurring interpretive tension that arises when splice-site variants generate partial rather than complete functional disruption. While we highlighted the classification challenges posed by the variants discussed in detail, we did not systematically attempt to specify how each variant should be classified. This is primarily because, as we have previously noted, “defining precise thresholds between different variant categories remains challenging and must be tailored to each gene and disease context,”8 and because the available functional analytic data are not conclusive or fully reliable in some cases, and the associated phenotypes were not always well defined. Despite these limitations, our integrative analysis of GT>GC variants provides a coherent framework for understanding how residual WT transcript generation shapes the functional continuum of splice-altering variants and informs their broader biological and clinical interpretation.
In conclusion, viewing GT>GC variants through the lens of residual WT transcript generation offers valuable insight into how partial LoF splice-altering variants—and, more broadly, variants with intermediate functional effects—contribute to genetic complexity, complicate both in silico prediction and experimental analysis, and challenge variant classification (Figure 5).
Data and code availability
All data supporting this study are provided within the article and its supplemental information. This study did not generate any new datasets or code.
Acknowledgments
This study was funded by the National Natural Science Foundation of China (82000611 to J.-H.L.), the Shanghai Pujiang Program (2020PJD061 to J.-H.L.), and the National Key Research and Development Program of China (2025YFC2511900 to W.-B.Z.). Support for this study also came from the Institut National de la Santé et de la Recherche Médicale (Inserm), France. D.N.C. wishes to acknowledge financial support from Qiagen through a license agreement with Cardiff University. The funding bodies did not play any role in the study design, collection, analysis, and interpretation of data or the writing of the article and the decision to submit it for publication.
Author contributions
Conceptualization, J.-M.C.; study design, J.-H.L., H.W., W.-B.Z., and J.-M.C.; methodology, J.-H.L., H.W., W.-B.Z., and J.-M.C.; experiments, J.-H.L., H.W., X.-Y.T., and W.-B.Z.; literature search and variant collation, E.M. and J.-M.C.; variant database, P.D.S., A.D.P., and D.N.C.; writing – original draft, J.-M.C.; writing – review & editing, J.-H.L., H.W., X.-Y.T., E.M., P.D.S., A.D.P., D.N.C., C.F., Z.L., W.-B.Z., and J.-M.C.; project administration, W.-B.Z. and J.-M.C.; funding acquisition, J.-H.L., W.-B.Z., and J.-M.C. All authors have read and approved the final version of the manuscript and agree to be accountable for all aspects of the work, ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Declaration of interests
The authors declare no competing interests.
Declaration of generative AI and AI-assisted technologies in the writing process
During the preparation of this work, the authors used ChatGPT 5 in order to improve readability and language. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.xhgg.2026.100602.
Contributor Information
Wen-Bin Zou, Email: dr.wenbinzou@hotmail.com.
Jian-Min Chen, Email: jian-min.chen@univ-brest.fr.
Web resources
CFTR2, http://www.cftr2.org/
CFTR-France, https://cftr.chu-montpellier.fr/
ImageJ, https://imagej.net/ij/
OMIM, https://www.omim.org
SpliceAI Lookup, https://spliceailookup.broadinstitute.org/
Supplemental information
References
- 1.Fowler D.M., Rehm H.L. Will variants of uncertain significance still exist in 2030? Am. J. Hum. Genet. 2024;111:5–10. doi: 10.1016/j.ajhg.2023.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Forrest I.S., Huang K.L., Eggington J.M., Chung W.K., Jordan D.M., Do R. Using large-scale population-based data to improve disease risk assessment of clinical variants. Nat. Genet. 2025;57:1588–1597. doi: 10.1038/s41588-025-02212-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Wright C.F., Sharp L.N., Jackson L., Murray A., Ware J.S., MacArthur D.G., Rehm H.L., Patel K.A., Weedon M.N. Guidance for estimating penetrance of monogenic disease-causing variants in population cohorts. Nat. Genet. 2024;56:1772–1779. doi: 10.1038/s41588-024-01842-3. [DOI] [PubMed] [Google Scholar]
- 4.Torene R.I., Murphy K.M., Brandt T., Kelly M.A., Willard H.F., Retterer K. A scalable approach for genomic-first rare disorder detection in a healthcare-based population. Am. J. Hum. Genet. 2025;112:2565–2577. doi: 10.1016/j.ajhg.2025.09.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Richards S., Aziz N., Bale S., Bick D., Das S., Gastier-Foster J., Grody W.W., Hegde M., Lyon E., Spector E., et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med. 2015;17:405–424. doi: 10.1038/gim.2015.30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Masson E., Zou W.B., Génin E., Cooper D.N., Le Gac G., Fichou Y., Pu N., Rebours V., Férec C., Liao Z., Chen J.M. Expanding ACMG variant classification guidelines into a general framework. Hum. Genom. 2022;16:31. doi: 10.1186/s40246-022-00407-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Schmidt R.J., Steeves M., Bayrak-Toydemir P., Benson K.A., Coe B.P., Conlin L.K., Ganapathi M., Garcia J., Gollob M.H., Jobanputra V., et al. Recommendations for risk allele evidence curation, classification, and reporting from the ClinGen Low Penetrance/Risk Allele Working Group. Genet. Med. 2024;26 doi: 10.1016/j.gim.2023.101036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Wang Y.C., Masson E., Wang Q.W., Génin E., Le Gac G., Fichou Y., Cooper D.N., Liao Z., Férec C., Zou W.B., Chen J.M. SPINK1-related chronic pancreatitis: A model that encapsulates the spectrum of variant effects, genetic complexity, and classificatory challenges. Am. J. Hum. Genet. 2025;112:2043–2066. doi: 10.1016/j.ajhg.2025.07.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Forrest I.S., Vy H.M.T., Rocheleau G., Jordan D.M., Petrazzini B.O., Nadkarni G.N., Cho J.H., Ganapathi M., Huang K.L., Chung W.K., Do R. Machine learning-based penetrance of genetic variants. Science. 2025;389 doi: 10.1126/science.adm7066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Walker L.C., Hoya M.d.l., Wiggins G.A.R., Lindy A., Vincent L.M., Parsons M.T., Canson D.M., Bis-Brewer D., Cass A., Tchourbanov A., et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: Recommendations from the ClinGen SVI Splicing Subgroup. Am. J. Hum. Genet. 2023;110:1046–1067. doi: 10.1016/j.ajhg.2023.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sullivan P.J., Quinn J.M.W., Ajuyah P., Pinese M., Davis R.L., Cowley M.J. Data-driven insights to inform splice-altering variant assessment. Am. J. Hum. Genet. 2025;112:764–778. doi: 10.1016/j.ajhg.2025.02.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lin J.H., Tang X.Y., Boulling A., Zou W.B., Masson E., Fichou Y., Raud L., Le Tertre M., Deng S.J., Berlivet I., et al. First estimate of the scale of canonical 5′ splice site GT>GC variants capable of generating wild-type transcripts. Hum. Mutat. 2019;40:1856–1873. doi: 10.1002/humu.23821. [DOI] [PubMed] [Google Scholar]
- 13.Ramalho A.S., Beck S., Meyer M., Penque D., Cutting G.R., Amaral M.D. Five percent of normal cystic fibrosis transmembrane conductance regulator mRNA ameliorates the severity of pulmonary disease in cystic fibrosis. Am. J. Respir. Cell Mol. Biol. 2002;27:619–627. doi: 10.1165/rcmb.2001-0004OC. [DOI] [PubMed] [Google Scholar]
- 14.Raraigh K.S., Han S.T., Davis E., Evans T.A., Pellicore M.J., McCague A.F., Joynt A.T., Lu Z., Atalar M., Sharma N., et al. Functional assays are essential for interpretation of missense variants associated with variable expressivity. Am. J. Hum. Genet. 2018;102:1062–1077. doi: 10.1016/j.ajhg.2018.04.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Scalet D., Maestri I., Branchini A., Bernardi F., Pinotti M., Balestra D. Disease-causing variants of the conserved +2T of 5′ splice sites can be rescued by engineered U1snRNAs. Hum. Mutat. 2019;40:48–52. doi: 10.1002/humu.23680. [DOI] [PubMed] [Google Scholar]
- 16.Den Uijl I.E.M., Mauser Bunschoten E.P., Roosendaal G., Schutgens R.E.G., Biesma D.H., Grobbee D.E., Fischer K. Clinical severity of haemophilia A: does the classification of the 1950s still stand? Haemophilia. 2011;17:849–853. doi: 10.1111/j.1365-2516.2011.02539.x. [DOI] [PubMed] [Google Scholar]
- 17.Wong M.S., Kinney J.B., Krainer A.R. Quantitative activity profile and context dependence of all human 5′ splice sites. Mol. Cell. 2018;71:1012–1026.e3. doi: 10.1016/j.molcel.2018.07.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Burset M., Seledtsov I.A., Solovyev V.V. Analysis of canonical and non-canonical splice sites in mammalian genomes. Nucleic Acids Res. 2000;28:4364–4375. doi: 10.1093/nar/28.21.4364. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Abril J.F., Castelo R., Guigó R. Comparison of splice sites in mammals and chicken. Genome Res. 2005;15:111–119. doi: 10.1101/gr.3108805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Parada G.E., Munita R., Cerda C.A., Gysling K. A comprehensive survey of non-canonical splice sites in the human transcriptome. Nucleic Acids Res. 2014;42:10564–10578. doi: 10.1093/nar/gku744. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kralovicova J., Hwang G., Asplund A.C., Churbanov A., Smith C.I.E., Vorechovsky I. Compensatory signals associated with the activation of human GC 5′ splice sites. Nucleic Acids Res. 2011;39:7077–7091. doi: 10.1093/nar/gkr306. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Lin J.H., Wu H., Zou W.B., Masson E., Fichou Y., Le Gac G., Cooper D.N., Férec C., Liao Z., Chen J.M. Splicing outcomes of 5′ splice site GT>GC variants that generate wild-type transcripts differ significantly between full-length and minigene splicing assays. Front. Genet. 2021;12 doi: 10.3389/fgene.2021.701652. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Landrum M.J., Chitipiralla S., Kaur K., Brown G., Chen C., Hart J., Hoffman D., Jang W., Liu C., Maddipatla Z., et al. ClinVar: updates to support classifications of both germline and somatic variants. Nucleic Acids Res. 2025;53:D1313–D1321. doi: 10.1093/nar/gkae1090. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Stenson P.D., Mort M., Ball E.V., Chapman M., Evans K., Azevedo L., Hayden M., Heywood S., Millar D.S., Phillips A.D., Cooper D.N. The Human Gene Mutation Database (HGMD®): optimizing its use in a clinical diagnostic or research setting. Hum. Genet. 2020;139:1197–1207. doi: 10.1007/s00439-020-02199-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Goldberg Y., Laitman Y., Ben David M., Bazak L., Lidzbarsky G., Salmon L.B., Shkedi-Rafid S., Barshack I., Avivi C., Darawshe M., et al. Re-evaluating the pathogenicity of the c.783+2T>C BAP1 germline variant. Hum. Mutat. 2021;42:592–599. doi: 10.1002/humu.24189. [DOI] [PubMed] [Google Scholar]
- 26.Sculco M., La Vecchia M., Aspesi A., Clavenna M.G., Salvo M., Borgonovi G., Pittaro A., Witel G., Napoli F., Listì A., et al. Diagnostics of BAP1-tumor predisposition syndrome by a multitesting approach: A ten-year-long experience. Diagnostics. 2022;12:1710. doi: 10.3390/diagnostics12071710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Jaganathan K., Kyriazopoulou Panagiotopoulou S., McRae J.F., Darbandi S.F., Knowles D., Li Y.I., Kosmicki J.A., Arbelaez J., Cui W., Schwartz G.B., et al. Predicting splicing from primary sequence with deep learning. Cell. 2019;176:535–548.e24. doi: 10.1016/j.cell.2018.12.015. [DOI] [PubMed] [Google Scholar]
- 28.Castellani C., CFTR2 team CFTR2: How will it help care? Paediatr. Respir. Rev. 2013;14:2–5. doi: 10.1016/j.prrv.2013.01.006. [DOI] [PubMed] [Google Scholar]
- 29.Claustres M., Thèze C., des Georges M., Baux D., Girodon E., Bienvenu T., Audrézet M.P., Dugueperoux I., Férec C., Lalau G., et al. CFTR-France, a national relational patient database for sharing genetic and phenotypic data associated with rare CFTR variants. Hum. Mutat. 2017;38:1297–1315. doi: 10.1002/humu.23276. [DOI] [PubMed] [Google Scholar]
- 30.Joynt A.T., Evans T.A., Pellicore M.J., Davis-Marcisak E.F., Aksit M.A., Eastman A.C., Patel S.U., Paul K.C., Osorio D.L., Bowling A.D., et al. Evaluation of both exonic and intronic variants for effects on RNA splicing allows for accurate assessment of the effectiveness of precision therapies. PLoS Genet. 2020;16 doi: 10.1371/journal.pgen.1009100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Müller J.S., Piko H., Schoser B.G.H., Schlotter-Weigel B., Reilich P., Gürster S., Born C., Karcagi V., Pongratz D., Lochmüller H., Walter M.C. Novel splice site mutation in the caveolin-3 gene leading to autosomal recessive limb girdle muscular dystrophy. Neuromuscul. Disord. 2006;16:432–436. doi: 10.1016/j.nmd.2006.04.006. [DOI] [PubMed] [Google Scholar]
- 32.Seyama K., Nonoyama S., Gangsaas I., Hollenbaugh D., Pabst H.F., Aruffo A., Ochs H.D. Mutations of the CD40 ligand gene and its effect on CD40 ligand expression in patients with X-linked hyper IgM syndrome. Blood. 1998;92:2421–2434. [PubMed] [Google Scholar]
- 33.Bartolo C., Papp A.C., Snyder P.J., Sedra M.S., Burghes A.H., Hall C.D., Mendell J.R., Prior T.W. A novel splice site mutation in a Becker muscular dystrophy patient. J. Med. Genet. 1996;33:324–327. doi: 10.1136/jmg.33.4.324. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Frischknecht H., Dutly F., Walker L., Nakamura-Garrett L.M., Eng B., Waye J.S. Three new beta-thalassemia mutations with varying degrees of severity. Hemoglobin. 2009;33:220–225. doi: 10.1080/03630260903089060. [DOI] [PubMed] [Google Scholar]
- 35.Kume K., Masamune A., Kikuta K., Shimosegawa T. [-215G>A; IVS3+2T>C] mutation in the SPINK1 gene causes exon 3 skipping and loss of the trypsin binding site. Gut. 2006;55:1214. doi: 10.1136/gut.2006.095752. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Chen J.M., Lin J.H., Masson E., Liao Z., Férec C., Cooper D.N., Hayden M. The experimentally obtained functional impact assessments of 5′ splice site GT>GC variants differ markedly from those predicted. Curr. Genom. 2020;21:56–66. doi: 10.2174/1389202921666200210141701. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Jayasree D., Shaji R.V., George B., Mathews V., Srivastava A., Edison E.S. Clinical, hematological and molecular analysis of homozygous Hb E (HBB: c.79G > A) in the Indian population. Hemoglobin. 2016;40:16–19. doi: 10.3109/03630269.2015.1086880. [DOI] [PubMed] [Google Scholar]
- 38.Aebi M., Hornig H., Padgett R.A., Reiser J., Weissmann C. Sequence requirements for splicing of higher eukaryotic nuclear pre-mRNA. Cell. 1986;47:555–565. doi: 10.1016/0092-8674(86)90620-3. [DOI] [PubMed] [Google Scholar]
- 39.Aebi M., Hornig H., Weissmann C. 5′ cleavage site in eukaryotic pre-mRNA splicing is determined by the overall 5′ splice region, not by the conserved 5′ GU. Cell. 1987;50:237–246. doi: 10.1016/0092-8674(87)90219-4. [DOI] [PubMed] [Google Scholar]
- 40.Tesio N., Bauer D.E. Molecular basis and genetic modifiers of thalassemia. Hematol. Oncol. Clin. N. Am. 2023;37:273–299. doi: 10.1016/j.hoc.2022.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Nix P., Mundt E., Coffee B., Goossen E., Warf B.M., Brown K., Bowles K., Roa B. Interpretation of BRCA2 splicing variants: A case series of challenging variant interpretations and the importance of functional RNA analysis. Fam. Cancer. 2022;21:7–19. doi: 10.1007/s10689-020-00224-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Pruss D., Morris B., Hughes E., Eggington J.M., Esterling L., Robinson B.S., van Kan A., Fernandes P.H., Roa B.B., Gutin A., et al. Development and validation of a new algorithm for the reclassification of genetic variants identified in the BRCA1 and BRCA2 genes. Breast Cancer Res. Treat. 2014;147:119–132. doi: 10.1007/s10549-014-3065-9. [DOI] [PubMed] [Google Scholar]
- 43.Tubeuf H., Caputo S.M., Sullivan T., Rondeaux J., Krieger S., Caux-Moncoutier V., Hauchard J., Castelain G., Fiévet A., Meulemans L., et al. Calibration of pathogenicity due to variant-induced leaky splicing defects by using BRCA2 exon 3 as a model system. Cancer Res. 2020;80:3593–3605. doi: 10.1158/0008-5472.CAN-20-0895. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Huynh A., Gray P.E., Sullivan A., Mackie J., Guerin A., Rao G., Pathmanandavel K., Mina E.D., Hollway G., Hobbs M., et al. A Novel case of IFNAR1 deficiency identified a common canonical splice site variant in DOCK8 in Western Polynesia: The importance of validating variants of unknown significance in under-represented ancestries. J. Clin. Immunol. 2024;44:170. doi: 10.1007/s10875-024-01774-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Zimmermann H., Brannan T., Young C., Ramirez Castano J., Horton C., Richardson A., Molparia B., Richardson M.E. A plot twist: When RNA Yields unexpected findings in paired DNA-RNA germline genetic testing. Genes. 2025;16:1382. doi: 10.3390/genes16111382. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Kilic S., Sukruoglu Erdogan O., Tuncer S.B., Celik Demirbas B., Yalniz Kayim Z., Yazici H. RNA splicing aberrations in hereditary cancer: Insights from Turkish patients. Curr. Issues Mol. Biol. 2024;46:13252–13266. doi: 10.3390/cimb46110790. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Behera A., Panigrahi G.K., Sahoo A. Nonsense-mediated mRNA decay in human health and diseases: Current understanding, regulatory mechanisms and future perspectives. Mol. Biotechnol. 2025;67:3374–3390. doi: 10.1007/s12033-024-01267-7. [DOI] [PubMed] [Google Scholar]
- 48.Bardakov S.N., Deev R.V., Magomedova R.M., Umakhanova Z.R., Allamand V., Gartioux C., Zulfugarov K.Z., Akhmedova P.G., Tsargush V.A., Titova A.A., et al. Intrafamilial phenotypic variability of collagen VI-related myopathy due to a new mutation in the COL6A1 gene. J. Neuromuscul. Dis. 2021;8:273–285. doi: 10.3233/JND-200476. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Peretto L., Tonetto E., Maestri I., Bezzerri V., Valli R., Cipolli M., Pinotti M., Balestra D. Counteracting the common Shwachman-Diamond Syndrome-causing SBDS c.258+2T>C mutation by RNA therapeutics and base/prime editing. Int. J. Mol. Sci. 2023;24:4024. doi: 10.3390/ijms24044024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Schmaltz-Panneau B., Pagnier A., Clauin S., Buratti J., Marty C., Fenneteau O., Dieterich K., Beaupain B., Donadieu J., Plo I., Bellanné-Chantelot C. Identification of biallelic germline variants of SRP68 in a sporadic case with severe congenital neutropenia. Haematologica. 2021;106:1216–1219. doi: 10.3324/haematol.2020.247825. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Lao Q., Mallappa A., Rueda Faucz F., Joyal E., Veeraraghavan P., Chen W., Merke D.P. A TNXB splice donor site variant as a cause of hypermobility type Ehlers-Danlos syndrome in patients with congenital adrenal hyperplasia. Mol. Genet. Genom. Med. 2021;9 doi: 10.1002/mgg3.1556. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Riordan J.R., Rommens J.M., Kerem B., Alon N., Rozmahel R., Grzelczak Z., Zielenski J., Lok S., Plavsic N., Chou J.L., et al. Identification of the cystic fibrosis gene: cloning and characterization of complementary DNA. Science. 1989;245:1066–1073. doi: 10.1126/science.2475911. [DOI] [PubMed] [Google Scholar]
- 53.Kerem B., Rommens J.M., Buchanan J.A., Markiewicz D., Cox T.K., Chakravarti A., Buchwald M., Tsui L.C. Identification of the cystic fibrosis gene: genetic analysis. Science. 1989;245:1073–1080. doi: 10.1126/science.2570460. [DOI] [PubMed] [Google Scholar]
- 54.Lundman E., Gaup H.J., Bakkeheim E., Olafsdottir E.J., Rootwelt T., Storrøsten O.T., Pettersen R.D. Implementation of newborn screening for cystic fibrosis in Norway. Results from the first three years. J. Cyst. Fibros. 2016;15:318–324. doi: 10.1016/j.jcf.2015.12.017. [DOI] [PubMed] [Google Scholar]
- 55.Aicher J.K., Jewell P., Vaquero-Garcia J., Barash Y., Bhoj E.J. Mapping RNA splicing variations in clinically accessible and nonaccessible tissues to facilitate Mendelian disease diagnosis using RNA-seq. Genet. Med. 2020;22:1181–1190. doi: 10.1038/s41436-020-0780-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Wai H.A., Lord J., Lyon M., Gunning A., Kelly H., Cibin P., Seaby E.G., Spiers-Fitzgerald K., Lye J., Ellard S., et al. Blood RNA analysis can increase clinical diagnostic rate and resolve variants of uncertain significance. Genet. Med. 2020;22:1005–1014. doi: 10.1038/s41436-020-0766-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Wu H., Boulling A., Cooper D.N., Li Z.S., Liao Z., Chen J.M., Férec C. In vitro and in silico evidence against a significant effect of the SPINK1 c.194G>A variant on pre-mRNA splicing. Gut. 2017;66:2195–2196. doi: 10.1136/gutjnl-2017-313948. [DOI] [PubMed] [Google Scholar]
- 58.Karjosukarso D.W., Kiefmann J.F., Bukkems F., Duijkers L., Collin R.W.J. Context matters: The importance of a comprehensive genomic region when assessing the therapeutic potential of antisense oligonucleotides in splicing assays. Nucleic Acid Therapeut. 2025;35:241–245. doi: 10.1177/21593337251371582. [DOI] [PubMed] [Google Scholar]
- 59.Sangermano R., Khan M., Cornelis S.S., Richelle V., Albert S., Garanto A., Elmelik D., Qamar R., Lugtenberg D., van den Born L.I., et al. ABCA4 midigenes reveal the full splice spectrum of all reported noncanonical splice site variants in Stargardt disease. Genome Res. 2018;28:100–110. doi: 10.1101/gr.226621.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Buisine M.P., Bellanne-Chantelot C., Calmels N., Vaché C., Besnard T., Cogne B., Vitobello A., Piton A., Martins A., Gaildrat P., et al. RNA-based diagnostic studies in genetics: Review and guidance from a multidisciplinary French network. Eur. J. Hum. Genet. 2025;33:1219–1227. doi: 10.1038/s41431-025-01881-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Andreae H., Curcio M., Owrang D., Esmaeelpour S., Jahnke F., Benseler F., Brose N., Vona B. Protocol for a minigene splice assay using the pET01 vector. STAR Protoc. 2025;6 doi: 10.1016/j.xpro.2025.103908. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Fu X.D., Ares M., Jr. Context-dependent control of alternative splicing by RNA-binding proteins. Nat. Rev. Genet. 2014;15:689–701. doi: 10.1038/nrg3778. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Drexler H.L., Choquet K., Churchman L.S. Splicing kinetics and coordination revealed by direct nascent RNA sequencing through nanopores. Mol. Cell. 2020;77:985–998.e8. doi: 10.1016/j.molcel.2019.11.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Oh R.Y., AlMail A., Cheerie D., Guirguis G., Hou H., Yuki K.E., Haque B., Thiruvahindrapuram B., Marshall C.R., Mendoza-Londono R., et al. A systematic assessment of the impact of rare canonical splice site variants on splicing using functional and in silico methods. HGG Adv. 2024;5 doi: 10.1016/j.xhgg.2024.100299. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Zhao S., Macakova K., Sinson J.C., Dai H., Rosenfeld J., Zapata G.E., Li S., Ward P.A., Wang C., Qu C., et al. Clinical validation of RNA sequencing for Mendelian disorder diagnostics. Am. J. Hum. Genet. 2025;112:779–792. doi: 10.1016/j.ajhg.2025.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Sha Y., Ortiz J.B., Bristow S.L., Loranger K., Meng L., Zhao X., Xia F., Parmar S., ElNaggar A.C., Xu W. Refining the interpretation of variants of uncertain significance in hereditary cancer screening through integrated RNA sequencing. Genet. Med. Open. 2025;3 doi: 10.1016/j.gimo.2024.101914. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Findlay G.M., Daza R.M., Martin B., Zhang M.D., Leith A.P., Gasperini M., Janizek J.D., Huang X., Starita L.M., Shendure J. Accurate classification of BRCA1 variants with saturation genome editing. Nature. 2018;562:217–222. doi: 10.1038/s41586-018-0461-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Nicolas-Martinez E.C., Robinson O., Pflueger C., Gardner A., Corbett M.A., Ritchie T., Kroes T., van Eyk C.L., Scheffer I.E., Hildebrand M.S., et al. RNA variant assessment using transactivation and transdifferentiation. Am. J. Hum. Genet. 2024;111:1673–1699. doi: 10.1016/j.ajhg.2024.06.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Sullivan P.J., Quinn J.M.W., Wu W., Pinese M., Cowley M.J. SpliceVarDB: A comprehensive database of experimentally validated human splicing variants. Am. J. Hum. Genet. 2024;111:2164–2175. doi: 10.1016/j.ajhg.2024.08.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Wu H., Lin J.H., Tang X.Y., Marenne G., Zou W.B., Schutz S., Masson E., Génin E., Fichou Y., Le Gac G., et al. Combining full-length gene assay and SpliceAI to interpret the splicing impact of all possible SPINK1 coding variants. Hum. Genom. 2024;18:21. doi: 10.1186/s40246-024-00586-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Costantino L., Rusconi D., Soldà G., Seia M., Paracchini V., Porcaro L., Asselta R., Colombo C., Duga S. Fine characterization of the recurrent c.1584+18672A>G deep-intronic mutation in the cystic fibrosis transmembrane conductance regulator gene. Am. J. Respir. Cell Mol. Biol. 2013;48:619–625. doi: 10.1165/rcmb.2012-0371OC. [DOI] [PubMed] [Google Scholar]
- 72.Elasal M.A., Khateb S., Panneman D.M., Roosing S., Cremers F.P.M., Banin E., Sharon D., Sarma A.S. A leaky deep intronic splice variant in CLRN1 is associated with non-syndromic retinitis pigmentosa. Genes. 2024;15:1363. doi: 10.3390/genes15111363. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Zhao X.C., Zhang Z.C., Ye W.L., Ye Y.Y., Wang L.T., Zheng X.Q., Chang Y.B., Chen C. Unexpectedly high levels of normally spliced transcripts from the pathogenic SLC10A7 alleles in a recessive form of skeletal dysplasia. HGG Adv. 2026;7 doi: 10.1016/j.xhgg.2025.100545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Gerasimavicius L., Livesey B.J., Marsh J.A. Loss-of-function, gain-of-function and dominant-negative mutations have profoundly different effects on protein structure. Nat. Commun. 2022;13:3895. doi: 10.1038/s41467-022-31686-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Axakova A., Ding M., Cote A.G., Subramaniam R., Senguttuvan V., Zhang H., Weile J., Douville S.V., Gebbia M., Al-Chalabi A., et al. Landscapes of missense variant impact for human superoxide dismutase 1. Am. J. Hum. Genet. 2025;112:2295–2315. doi: 10.1016/j.ajhg.2025.08.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Richardson M.E., Bishop M.F.H., Holdren M.A., de la Hoya M., Spurdle A.B., Tavtigian S.V., Brannan T., Young C.C., Zec L., Hiraki S., et al. Specifications of the ACMG/AMP variant curation guidelines for the analysis of germline PALB2 sequence variants. Am. J. Hum. Genet. 2025;112:2266–2280. doi: 10.1016/j.ajhg.2025.08.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Rowlands C.F., Allen S., Garrett A., Durkie M., Burghel G.J., Robinson R., Callaway A., Field J., Frugtniet B., Palmer-Smith S., et al. Availability of benign missense variant “truthsets” for validation of functional assays: Current status and a systematic approach. Am. J. Hum. Genet. 2025;112:2281–2294. doi: 10.1016/j.ajhg.2025.08.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Liu S., Feng X., Wu Y., Bu F. Calibration and refinement of ACMG/AMP criteria for variant classification with BayesQuantify. J. Med. Genet. 2025 doi: 10.111136/jmg-112025-110863. [DOI] [PubMed] [Google Scholar]
- 79.Yuan X., Xia X., Li J., Zhao G. HCSeeker: A classification tool for human genetic variant hot and cold spots designed for PM1 and benign criteria in the ACMG-AMP guideline. Genet. Med. 2025;27 doi: 10.1016/j.gim.2025.101591. [DOI] [PubMed] [Google Scholar]
- 80.Whitcomb D.C. Adoption of a “Pathogenic, Predisposing, Risk and Benign” classification of DNA variants for precision medicine – a step forward. SMART-MD J. Precision Med. 2025;2:e.i–e.ii. doi: 10.69734/h7x6zm57. [DOI] [Google Scholar]
- 81.Rubin A.F., Stone J., Bianchi A.H., Capodanno B.J., Da E.Y., Dias M., Esposito D., Frazer J., Fu Y., Grindstaff S.B., et al. MaveDB 2024: A curated community database with over seven million variant effects from multiplexed functional assays. Genome Biol. 2025;26:13. doi: 10.1186/s13059-025-03476-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Allen S., Garrett A., Rowlands C.F., Durkie M., Burghel G.J., Robinson R., Callaway A., Field J., Frugtniet B., Palmer-Smith S., et al. Validating data from multiplex assays of variant effect: A CanVIG-UK national survey of NHS clinical scientists. Am. J. Hum. Genet. 2025;112:1479–1488. doi: 10.1016/j.ajhg.2025.04.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Heyne H.O., Karjalainen J., Karczewski K.J., Lemmelä S.M., Zhou W., Daly M.J., FinnGen. Havulinna A.S., Kurki M., Rehm H.L., Palotie A. Mono- and biallelic variant effects on disease at biobank scale. Nature. 2023;613:519–525. doi: 10.1038/s41586-022-05420-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Lee D.S.M., Cardone K.M., Zhang D.Y., Tsao N.L., Abramowitz S., Sharma P., DePaolo J.S., Conery M., Aragam K.G., Biddinger K., et al. Common-variant and rare-variant genetic architecture of heart failure across the allele-frequency spectrum. Nat. Genet. 2025;57:829–838. doi: 10.1038/s41588-025-02140-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.García Hernandez S., de la Higuera Romero L., Fernandez A., Luisa Peña Peña M., Mora-Ayestaran N., Basurte-Elorz M.T., Larrañaga-Moreira J.M., Cárdenas Reyes I., Villacorta E., Valverde-Gómez M., et al. Redefining the genetic architecture of hypertrophic cardiomyopathy: Role of Intermediate-effect variants. Circulation. 2025;152:1060–1075. doi: 10.1161/CIRCULATIONAHA.125.074529. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Manolio T.A., Collins F.S., Cox N.J., Goldstein D.B., Hindorff L.A., Hunter D.J., McCarthy M.I., Ramos E.M., Cardon L.R., Chakravarti A., et al. Finding the missing heritability of complex diseases. Nature. 2009;461:747–753. doi: 10.1038/nature08494. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Maroilley T., Tarailo-Graovac M. Uncovering missing heritability in rare diseases. Genes. 2019;10:275. doi: 10.3390/genes10040275. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Génin E. Missing heritability of complex diseases: case solved? Hum. Genet. 2020;139:103–113. doi: 10.1007/s00439-019-02034-4. [DOI] [PubMed] [Google Scholar]
- 89.Denny J.C., Collins F.S. Precision medicine in 2030-seven ways to transform healthcare. Cell. 2021;184:1415–1419. doi: 10.1016/j.cell.2021.01.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Torene R.I., Guillen Sacoto M.J., Millan F., Zhang Z., McGee S., Oetjens M., Heise E., Chong K., Sidlow R., O’Grady L., et al. Systematic analysis of variants escaping nonsense-mediated decay uncovers candidate Mendelian diseases. Am. J. Hum. Genet. 2024;111:70–81. doi: 10.1016/j.ajhg.2023.11.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data supporting this study are provided within the article and its supplemental information. This study did not generate any new datasets or code.





