Skip to main content
Cell Genomics logoLink to Cell Genomics
. 2025 Dec 2;6(2):101082. doi: 10.1016/j.xgen.2025.101082

Analysis of error profiles of indels and structural variants in deep-sequencing data

Ying Shao 1,10, Quang Tran 1,10, Yuan Feng 1,10, Pandurang Kolekar 1,10, Yanling Liu 1,10, Zhikai Liang 1, Li Fan 1, Andrea McBride 2, Tyler Jones 2, Alexis Cameron 2, Heather Mulder 1, Lingyun Ji 3, Benjamin J Huang 4, Jeffery M Klco 5, Soheil Meshinchi 6, Jinghui Zhang 1, William L Carroll 7, Mignon L Loh 8,, John Easton 1,∗∗, Patrick A Brown 9,∗∗∗, Xiaotu Ma 1,11,∗∗∗∗
PMCID: PMC12903382  PMID: 41338220

Summary

Despite extensive studies of the error profiles of SNVs, those of insertions/deletions (indels)/structural variants (SVs) remain elusive. Using ultra-deep sequencing, we show that the error rates of indel/SVs are >100-fold lower than those of SNVs, although repeat indels have high error rates of 1%. We validated this pattern in a cohort of 103 patients with relapsed B cell acute lymphoblastic leukemia (B-ALL). We analyzed repeat indels in 339 cancer driver genes and demonstrated that the number of repeat units is highly predictive of the error rate. We then analyzed minimal residual disease samples from 72 patients with relapsed B-ALL and demonstrated that our approach had positive detections in 61% of cases, outperforming clinical flow cytometry (51% detection). Overall, we established indel and SV error profiles in deep next-generation sequencing (NGS) data, enabling superior tumor detection at very low burdens, which has a significant impact on the clinical diagnosis and monitoring of human cancers and other diseases.

Keywords: sensitive detection, deep sequencing, cancer early detection, error profile, structural variants and indels, repeat indels

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • SVs and indels have much lower sequencing error rate than do SNVs

  • Repeat indels have an error rate predictable from their repeat length

  • SVs and indels provide much better biomarkers for tumor tracking


Shao et al. introduced a new approach for indel and structural variant genotyping, showing that these variants are far less error prone than SNVs, while repeat indel errors scale with repeat length. Both indels and structural variants provide more reliable biomarkers for cancer detection and monitoring.

Introduction

Genetic heterogeneity—caused by DNA alterations (or mutations), including base substitution (SNV), small insertion/deletion (indel), and structural alteration (structural variant [SV], which also results in copy number variations [CNVs])— are associated with human diseases such as cancer1,2,3,4,5,6,7,8 and infectious diseases including COVID-19.6,7 In these applications, genetic alterations acquired in a subpopulation of the parental organisms/cells can carry biological significance; however, accurate detection of such events is a great challenge, especially when the mutant subpopulation is rare (e.g., ≤1%). Deep DNA sequencing is a powerful tool to uncover genetic heterogeneity. For this purpose, sequencing fidelity (also known as error rate) is of pivotal importance to uncover rare genetic subpopulations. There have been intensive interests in sequencing error rates since the inception of the next-generation sequencing (NGS) technology. For example, the error rate of SNVs was reported to be ∼10−3 from 2011 to 2018.1,2,3,4 Our 2019 study9 indicated that SNV error rates can be computationally suppressed to between 10−4 and 10−5, which is 10- to 100-fold lower than previous reports. Furthermore, by addressing instrument fidelity, our 2021 study10 indicated that the SNV error rate can be reduced to as low as 10−6 by coupling computational error suppression with best-performing instruments.

Understanding of substitution (SNV) errors has progressed rapidly, with Stoler et al.11 focused on analyzing context-specific substitution errors and Costello et al.12 identifying oxidation artifacts (C>A substitution errors) originating from acoustic shearing in the presence of reactive contaminants during DNA extraction. However, neither work addressed indel or SV error profiles. Although indels and SVs typically occur less frequently than SNVs, these genetic alterations can play an important role in human diseases, especially cancers. For example, over >55% of childhood leukemias are driven by oncogenic fusions through SVs (also known as chromosomal rearrangements or translocations).8 These oncogenic fusions define cancer subtypes, can be predictive of outcome, typically remain intact throughout the disease course (in which SNVs can be eradicated by treatment or acquired de novo during treatment13), and can serve as excellent therapeutic targets (e.g., imatinib for BCR::ABL1/Philadelphia chromosome leukemia14). Accurate detection of such oncogenic fusions can be highly informative for disease management. Indels are also frequent key drivers for human cancer. For example, EGFR tyrosine kinase (TK)15 domain alterations by indels are frequently observed in lung cancers, for which inhibitors to TK are shown to benefit patients. Similarly, internal tandem duplication (ITD) of FLT3 exon 14 is frequently found in childhood acute myeloid leukemia, and FLT3 inhibitors are already in clinical use.16 Although some studies have examined optimal alignment of indels,17,18,19 a systematic understanding of sequencing errors of indels is lacking.

In this work, we aim to investigate the error profile of indels and SVs. We first proposed a theoretical procedure to analyze their error rates by using the technique of conditional probability, which indicated that error rate of indels and SVs can be much lower than that of SNVs. Because some indels and SVs can have complex mutation patterns and are commonly inaccurately notated,18 we developed a simple framework to ensure accurate allele representation of indels and SVs. We next performed ultra-deep sequencing (at ∼10,000,000×) by using a previously established dilution model involving a matched cancer/normal cell line pair established from a patient with melanoma (COLO829/COLO829BL).9 We then studied error profiles of 309 indels and 1,063 SVs detected from 103 patients with relapsed childhood leukemia13 by using whole-genome sequencing (WGS) data from 1,662 healthy donors.9,20 This study confirmed our hypothesis that indels and SVs typically have very low error rate (<2 × 10−7), although indels from repeat regions (repeat indel) can have error rates >1%.

We next studied the indels from 11,378 repeat regions of 339 cancer genes by using the WGS data of 1,662 heathy donors.9,20 We demonstrate that the number of repeat units is highly predictive of the error rate of repeat indels. We further discovered that deletion errors are ∼10-fold higher than insertion errors for microsatellite indels.

We next studied our ability to detect minimal residual disease (MRD) at day 28 post treatment for 72 patients with relapsed leukemia by using clinical MRD measurement21 as the benchmarking standard. We found that our ultra-deep-sequencing method (detection rate, 61%) outperformed clinical MRD measurement (detection rate, 51%). Interestingly, although our data indicate that the error rates of SNVs, indels, and SVs are much lower (upon proper computational error suppression) than those reported, the current challenge is instead signal loss, which significantly hinders the ability to amplify/capture and successfully sequence the mutant allele when the concentration is low (i.e., <0.1%).

Results

Terminology on signal and noise

As emphasized by Loeb and colleagues,4 “all measurements in science are limited by an assay’s signal-to-noise ratio,” a principle especially pertinent to studies of genetic heterogeneity where subclonal alterations can occur at extremely low allele fractions (AFs; AF < 1%, and <0.01% for clinical MRD). Detecting such rare events through bulk sequencing has been a major focus in recent years,9,10,11 leading to discoveries such as the 8-oxoG artifact caused by excessive acoustic shearing,12 which prompted widespread changes in sequencing protocols. Here, “signal” refers to true variants, while “noise” or “error” in NGS denotes sequencing artifacts—events that appear as mutations despite their absence in the true genome. Because proving the nonexistence of a mutation is impractical, we previously addressed this through dilution experiments using spike-in controls,9 where true mutations correlated with dilution concentration, while false positives showed consistently low AFs, defining the sequencing error rate. Thus, the distinction between signal and noise is inherently relative: if the error rate is ∼0.1%, an AF > 5% likely represents a real signal, whereas an AF near 0.1% is ambiguous. This framework underlies our analysis of sequencing error rates across mutation types, including SNVs,9,10,11 indels, and SVs, in this study.

Mathematical model of error rates of SNV, indel, and SV

Human cancer is driven by genetic alterations. The number of alterations per tumor, also known as the tumor mutation burden, can vary dramatically for SNVs, indels, and SVs. For example, in childhood acute lymphoblastic leukemia (ALL), there could be hundreds to thousands of SNVs (0.1–1 per Mb) but only 1–30 SVs (median 4–10) per tumor.13,22 These data indicate that DNA damage is more prone to result in SNVs than SVs, which can also be predicted by theoretical analysis using conditional probability (Figure 1A). Here, we first assume that only the damage type (SNV, indel, SV) is given but the exact alteration is not known, and the task is to predict the exact alteration. If the alteration is an SNV, then we can make a random prediction from the 12 possible substitutions with the chance of correct prediction being 1/12, which roughly equals 10%. For indels, the chance of a correct prediction also depends on indel length. At length ≤6, the chance is 0.01% (Figure 1A), much lower than that for SNV. This analysis equally applies to the scenario of recurrently observing the same alteration (including SV) in two different cells or two different PCR products during sequencing, based on which we predict that error rates of indels and SVs are much lower than those of SNVs, which is the subject of this study.

Figure 1.

Figure 1

Model of indel and SV errors

(A) Conditional probability analysis of error rate of substitution (SNV), insertion/deletion (indel), and rearrangement (SV). If only the information of damage types (SNV/indel/SV) is given, we can correctly guess the exact damage with a chance of nearly 10% for SNV. For indels with length ≤6 bp, the chance is ∼10−4. This reasoning also applies to SVs and indicated that the error rate of SVs and indels should be much lower than that of SNVs.

(B) Example of simple (an insertion of CCC) and complex (replacement of CTGAG by CCCCCGA) indels.

(C and D) (C) Methods for counting indel mutant reads. In addition to the exact mutant (orange) and reference (green) alleles, a small flanking buffer (8 bp each side in this work) is required to ensure perfect allele matches. A small percentage (1%) of mismatches in further flanking regions is allowed. To account for elevated error rates in read ends, partial matches in the mutant/reference region were not counted (N) unless the matched length exceeded certain cutoff values (8 bp in this work, corresponding to a chance of 4−8 = 2 × 10−5 for random matches) for reference (R) and mutant (M) alleles, respectively. SV reads were similarly counted (D).

(E) Criterion of counting high-confidence mutant and reference reads.

For indels, the mutation can be either “simple,” with a few reference base pairs inserted/deleted, or “complex,” with a few reference base pairs collectively replaced by a few mutant base pairs (Figure 1B). To fully account for such complex events, we developed a rigorous read-counting method (Figure 1C; STAR Methods). Here, a read can have one of five alignment statuses: (1) fully matching the reference (green), (2) fully matching the mutant (orange), (3) partially matching the reference, (4) partially matching the mutant, or (5) not matching either. Because the overhang in the partial-matching scenario can be potentially confounded by substitution errors that are enriched in the end of reads,9 we decided to count the partially matching reads only if the overhang has length ≥8 bp (STAR Methods). To account for indels that fall within long tandem repeats, we added a short flanking buffer around the core reference and mutant sequences. Outside the buffer zone, up to 1% of mismatches are allowed to account for substitution errors. A similar method allowing possible non-template insertion sequences has been developed for SVs (Figure 1D). Taken together, a read will be counted (Figure 1E) if (1) it matches the core of the mutant/reference, (2) the alignment score is above a threshold (e.g., >99% identity), and (3) the alignment score with the mutant allele is better than that with the reference allele.

Dilution experiment using an established cell model

To study the error rate of indels and SVs, we performed a cell-line dilution experiment using previously established melanoma cell model COLO829 and its matched non-cancer line COLO829BL (established from the same patient; ATCC # CRL-1974 and CRL-1980).9 We randomly selected 23 indels and 17 SVs (all somatic; 11 indels are from repeat regions [i.e., repeat indels]; Tables S1 and S2) detected from this cell model in published studies23,24,25 for ultra-deep amplicon sequencing. To determine the limit of detection, we designed four dilution ladders: 100%, 1%, 0.1%, and 0%, with the 100% and 0% ladders used collectively to validate the somatic status of selected variants. For ladders 1% and 0.1%, we aimed for 50,000× depth. We aimed for 5,000,000× depth for the 0% ladder so that we could report error rates as low as 10−7. To ensure reproducibility, all experiments were run in replicates. As shown in Figure 2A (see Table S3), the AF (a pseudo-count of 1 is added to both mutant and total counts throughout this work; see STAR Methods) of the designed indel (11 repeat indels are analyzed separately to account for their high error rate) and SV markers demonstrated high correlation with the designed ladders (R2 > 0.86, p < 10−39) for both replicates. Interestingly, although 90%–97% of designed markers are detectable (by comparing with the AF observed in pure non-cancer sample via Fisher’s exact test with false discovery rate [FDR] cutoff = 0.01) at dilution concentration of 1%, only ∼50% of the designed markers are detected at the dilution concentration of 0.1%, which is concordant with our previous study of SNV markers.9

Figure 2.

Figure 2

Limit of detection of indels and SVs using COLO829 dilution experiment

(A) Observed AFs of somatic alterations demonstrated a high correlation with designed dilution ladders across biological replicates (left and right). Gray shading indicates pure non-cancer cells.

(B) Error rates are highly concordant between replicates. (C–D) SNV error rate (y axis) for 12 substitution types (x axis) calculated from a previously published dataset (C; see main text) and the deep-sequencing dataset of this study (D). Blue horizontal bars indicate the median error rate, which is also indicated on the top of the panels.

(E) Error rate of spike-in indels and SVs in this study. Indels are further grouped into complex, simple, and repeat indels. Blue horizontal bars indicate the median error rate.

(F) Error rate of repeat indels is highly correlated with the number of repeat units.

In (A) and (B), symbols +, ×, and o represent SNV, SV, and indel, respectively. In (A), dark blue and red indicate detected and missed markers, respectively. Green markers with gray shading highlight error rate in non-cancer samples. In (A), 11 repeat indels are not included due to a high error rate as analyzed in (E) and (F). In (E), counts from both replicates are combined. In (F), all repeat indels have units with length of 1; thus, the homopolymer length equals the number of repeat units. Linear regression p values are shown in (A), (B), and (F). See also Tables S1, S2, S3, and S4.

Due to the high correlation between observed AF and dilution concentration ladders, we expect the AF of these markers to be 0 when the dilution concentration is 0% (pure non-cancer). Non-zero AFs observed in samples with pure non-cancer must arise from sequencing errors and can be estimated by ultra-deep-sequencing data (gray shading in Figure 2A). Due to prohibitive sequencing cost, we achieved a depth of ∼5,000,000× for the designed loci in both replicates (Table S3), which would allow us to investigate an error rate of ∼10−7. As seen in Figure 2B, the error rates of these markers cluster tightly around 10−6 with a high concordance (R2 = 0.9, p = 2 × 10−21) between the two biological replicates.

To provide a benchmark, we first investigated substitution error rates by using a previously published ultra-deep-sequencing dataset of Ma et al.9 and the ultra-deep-sequencing data of this study (non-cancer sample; STAR Methods). The dataset generated in this study demonstrated slightly better error rates than did those of Ma et al. (Figures 2C and 2D; Table S4). We next categorized our markers by mutation types (Figure 2E), observing that error rates of indels had a high dynamic range of 10−7–10−2, with a median of 2 × 10−6. By contrast, the SVs had a universally lower error rate of ∼10−7. Because the 11 repeat indels—for which the repeat units all have a length of 1 bp and thus the length of repeat equals that of the number of repeat units—have a higher dynamic range of error rates than complex and simple indels do, we studied their error rates with their repeat lengths and discovered a highly significant correlation (Figure 2F).

Overall, these data confirmed our theoretical prediction (Figure 1A) that SVs and complex indels have a very low error rate (<2 × 10−7). Higher rates were found in simple indels (∼3 × 10−7), SNVs (∼10−5), and repeat indels, for which the error rate (10−6–10−2) appears to be associated with the number of repeating units as further studied below.

Error rate of cancer-related SNV, indel, and SV mutations

To generalize the above observation, we next focused on 309 indels and 1,063 SVs (Tables S5 and S6) detected from primary diagnosis and relapse tumors from a cohort of 103 patients with ALL.13 With the observation that somatic alterations can be driver or passenger alterations (with driver alterations able to functionally promote tumor growth),22,26,27,28 we classified the indels (Table S7) and SVs (Table S8) into categories of pathogenic and others.28 Importantly, some SVs are known to result in oncogenic fusions that define cancer subtypes,8 and we further classified these pathogenic alterations as “fusion” to highlight their importance. Because it is too expensive to generate further ultra-deep-sequencing data to study the error profiles of these markers, we utilized WGS (30×) datasets generated from 1,662 healthy (cancer-free; Table S9) donors from a St. Jude LIFE (SJLIFE) cohort of childhood cancer survivors (≥5 years since initial diagnosis) as a control dataset (a total of ∼50,000× depth).9,20 The rationale is that the cancer-driving alterations are not expected from any healthy donors; therefore, mutant reads observed from these WGS data are likely to be sequencing errors; in turn, the error rate estimate from this dataset is theoretically the upper bound of the true error rate because some healthy donors may have pre-cancerous alterations.29 Indeed, SVs labeled as fusion or pathogenic did not have any matching mutant reads in the control data (Figure 3A; Table S8). However, a small fraction (1.8%; 15 of 846) of the SVs labeled as “other” demonstrated matching mutant reads in the SJLIFE WGS datasets. Among these 15, 14 (93%) were intra-chromosomal deletions; therefore, we also classified the SVs into four categories based on location and orientation of the breakpoints25: (1) inter-chromosomal translocations (CTXs), (2) inversions (ITXs), (3) deletions (DELs), and (4) insertions (INSs). SVs categorized as DELs (Figure S1A) and having small sizes (Figure S1B) are more error prone (p = 9 × 10−7), although we did not find statistically significant (p = 0.8) association between error rate and SV size (Figure S1C), possibly due to small sample size or the limited sequencing depth of ∼50,000×. Collectively, these data confirmed the low error rate for SVs discovered in COLO829 data (Figure 2), and our bigger pool of somatic SVs here indicated that SVs involving small deletions can have increased error rates. Further, these data also confirmed the very low error rate of pathogenic indels (Figure 3B). Interestingly, repeat indels demonstrated a clear association between error rates and the number of repeat units (Figure 3C; p = 4 × 10−32), which is confirmed when the length of repeat region is used (Figure S1D; Table S7).

Figure 3.

Figure 3

Error rate of indels and SVs detected from 103 patients with pediatric ALL

(A) Error rate of SVs as classified into categories of fusion, pathogenic, and other.

(B) Error rate of indels categorized as pathogenic and other.

(C) Error rate of repeat indels is highly correlated with the number of repeat units. The number of data points is indicated for each category. Blue horizontal bars indicate the median error rate, which is also indicated on the top of the bars. Linear regression p value is shown in (C). See also Figure S1 and Tables S5, S6, S7, S8, and S9.

Error rate of repeat indels in 339 cancer genes

The above data reinforced the observation that repeat indels are error prone (Figure 3C; Table S7). However, the limited number of markers still prevented us from generalizing the conclusion. To address this gap, we collected repeat indels from 339 (Table S10) cancer genes (the calculation of these markers already took more than 1 month, so a separate study will be needed to evaluate more regions), which resulted in 11,378 repeat regions (STAR Methods; Table S11) with sufficient data (total depth >20,000×). We then calculated error rates for insertion or deletion of 1 repeat unit for these indels in 1,662 SJLIFE healthy donors (Table S9). To account for the possibility that some healthy donors may have a polymorphism in some of these repeat indels, we used an established approach to exclude homozygous or heterozygous samples (STAR Methods). As seen in Figure 4A, a highly significant (R2 = 0.97, p = 1.9 × 10−6) correlation between error rate and the number of repeat units is observed. Of note, the error rate can exceed 1% when the number of repeating units is ≥ 9, which is consistent with observations in disease-causing repeat indels that can be averted by base editing to break the repeat length.30

Figure 4.

Figure 4

Error rate of repeat indels

(A) The error rate (y axis) of repeat indels (from 339 cancer genes) is well predicted by the number of repeat units (x axis).

(B) Deletion errors (blue) tend to be more enriched for short indels than insertion errors (red) are.

(C and D) By categorizing repeats into homopolymers (C) and microsatellites (D), we discovered that deletion errors are nearly 10-fold more prevalent in microsatellite indels (D).

Numbers of data points for each given number of repeat units are indicated. Linear regression p values are shown in each panel. See also Figures S2 and S3; Tables S9, S10, and S11.

Next, we classified the errors (Figure 4B) as deletion of 1 repeat unit or insertion of 1 repeat unit. Strikingly, for shorter repeats (<7 units), deletion of 1 unit was more prevalent. For longer repeats (>7 units), insertion of 1 unit was more prevalent. However, the deletion error rate distribution in Figure 4B still appears to have >1 component with a long tail. We, therefore, categorized the repeat indels into homopolymers (i.e., repeating units are a single base A/C/G/T; Figure 4C) and microsatellites (such as di- and trinucleotides; Figure 4D). Interestingly, the 1-unit deletion is more prevalent (∼10-fold higher) in microsatellites than in homopolymers. We studied the deletion or insertion of two repetitive units, and a similar trend was observed by using the number of repeat units (Figure S2) or the length of repeat regions (Figure S3), although the total sequencing depth (∼50,000×) was not well powered for this analysis.

Effect of library quality, sequencing instrument, and PCR enzymes

We also investigated how common sequencing parameters, including sample handling/quality, sequencing instrument, and PCR enzymes, may affect the error profile. First, we downloaded the deep-sequencing dataset (EGA accession EGAS00001003444) known to have variable library qualities due to sample handling/storage (as measured by C>A error rates; Figure 5 of Ma et al.9). Because of the extremely low error rate of complex indels and the majority of somatic SVs, we choose to use repeat indels to investigate these parameters (Table S12). The 47 samples were grouped according to average C>A error rate as follows: high error rate (n = 16), medium error rate (n = 15), and low error rate (n = 16). The error rate of repeat indels within each group was calculated (Figure 5A). The median error rates of these three quality groups did not show any statistical difference, although the median error rates demonstrated a subtle decrease with increasing sample quality. By analyzing the deep-sequencing dataset (Figure 4 of Ma et al.9) generated via HiSeq and NovaSeq (Figure 5B), we found that repeat indels have comparable error rates on these two platforms. Further, we used the deep-sequencing dataset generated by different PCR enzymes Kapa and Q5, for which NovaSeq machines were used at two sequencing centers (St. Jude and HudsonAlpha Institute of Biotechnology). As seen in Figure 5C, repeat indels demonstrated a significantly lower error rate when the DNA library was prepared using Q5. As such, our data clearly suggest that, by contrast with SNVs whose error rates can be affected by sample quality and sequencing instrument,9,10,11,12 repeat indels have comparable error rates under these parameters, although Q5 can lead to lower error rates.

Figure 5.

Figure 5

Effect of sample quality, sequencing instrument, and PCR enzymes on error rates of repeat indels

(A) Error rate (y axis) of repeat indels (with deletion of 1 unit) from samples known to have high (n = 16), medium (n = 15), or low (n = 16) quality as measured by C>A error rate. The length of the repeat unit (bp) is shown on the x axis, along with the number of markers.

(B) Effect of sequencing instrument (HiSeq and NovaSeq) on the error rate of repeat indels (with Kapa enzyme for PCR).

(C) Effect of PCR enzymes (Kapa and Q5) on the error rate of repeat indels. Two different NovaSeq instruments (one from HudsonAlpha Institute of Biotechnology [HAIB], another from St. Jude [SJ]) were used. Wilcoxon rank-sum test p values are indicated.

See also Table S12.

Indels and SVs provide superior detection to SNVs for leukemia MRD

We next asked whether our above insights on error rate of somatic mutations can be used to detect residual tumors in the end-of-induction (day 28) sample of patients with B-progenitor lymphoblastic leukemia, which are known to be one of the most informative predictors of treatment outcome.21 For this, we analyzed 72 cases of relapsed B-ALL from an ongoing Children’s Oncology Group study of relapsed ALL genomics (AALL1331, to be published elsewhere). Because this is a relapsed cohort, there must be cancer cells in the patients at the end of induction (day 28 of their initial treatment), although it is not guaranteed to have cancer cells detectable in the remission specimen due to sampling fluctuation (typically a few million cells). Therefore, methods with higher positive detection rates are superior because they enable more high-risk patients to be treated appropriately. Around five clonal (shared between diagnosis and relapse tumors) somatic markers were selected for each patient (total of 397 markers; Table S13) by minimizing error rates, for example, preferentially selecting SVs, complex indels, and low-error-rate SNVs (including A>C, C>G, and A>T substitutions). A custom capture panel was designed using these patient-specific markers to perform ultra-deep sequencing across 72 remission specimens (achieved average depth of ∼30,000×).

First, we validated the predicted low error rate of the selected markers. For each marker, samples from patients known to be negative for this marker (in both diagnosis and relapse tumors) are assumed to be wild-type and are pooled to calculate corresponding error rates. As seen in Figure 6A, SVs and indels have lower error rates (<3 × 10−7, only two of 83 markers achieved sampling saturation) than the best-performing SNVs (∼10−5, nearly all designed SNV markers achieved sampling saturation). By comparing index samples with pooled background control samples via binomial distribution with false-discovery rate control (STAR Methods), we first called the presence/absence status for each marker in the index sample, then an average AF was calculated by using pooled allele counts of all positive markers in the corresponding sample (STAR Methods). The average AF is termed the NGS MRD level. A sample without any positive markers detected is considered to have an NGS MRD of 0; 38 cases received a positive NGS MRD in this round of the experiment (round 1; Table S14).

Figure 6.

Figure 6

Detecting residual tumors from remission samples of 72 patients with relapsed B-ALL

(A) SVs and indels have a >10-fold lower error rate (<3 × 10−7) than the best-performing SNVs (∼10−5). Blue horizontal bars indicate median error rate, which is also indicated on top.

(B) Next-generation sequencing (NGS) approach had 44 positive detections (61%), and state-of-the-art clinical flow cytometry had 37 positive detections (51%). Linear regression p value is shown.

(C) NGS-based method detected 92% of designed markers for samples with MRD > 0.3%, and this detection rate dropped to 27% for MRD between 0.1% and 0.01%.

(D) Detection rate stratified by different somatic marker types: SV, indel, and SNV.

See also Tables S13, S14, and S15.

Among the 34 cases with a negative NGS MRD detection in our round 1 experiment, nine had a positive flow MRD. We hypothesized that the limited sequencing depth (∼30,000×) may have caused the false negatives. To test this hypothesis, we randomly selected 17 cases to generate additional sequencing data with ∼120,000× as round 2 experiments, using the same library. As it turned out, 11 cases (65% of 17) received a stronger signal than that in round 1 (Table S15). Therefore, higher sampling depth is required to achieve higher sensitivity for the NGS detection of MRD.

Using these data, we next compared the NGS MRD against the clinical flow MRD (Tables S13 and S14; Figure 6B). Collectively, our NGS MRD method detected 11 more positive patients and missed four patients who were positive for flow, which resulted in a detection yield of 61% for NGS MRD and 51% for flow MRD. Notably, a reasonably high correlation between NGS and flow MRD (R2 = 0.46; p = 1.6 × 10−5) was observed in patients with positive detection by both flow and NGS MRD method.

Next, we stratified patients into MRD level bins according to clinical flow MRD (>0.5%, >0.1%, >0.01%, <0.01%, and negative; Figure 6C) and studied the chance to have positive detection of designed somatic markers. Interestingly, a high detection rate (83%; 114 of 138) was achieved when the flow MRD was >0.1%. This rate decreased to 16% (41 of 259) when flow MRD level was <0.1% or negative (Table S13). These data are concordant with the COLO829 data (Figure 2) and indicate that the next challenge in ultra-deep sequencing is to overcome the signal-loss problem because the marker would have been detected given the low error rates. Collectively, we discovered that SVs and indels have higher detection rates (>45%) than SNVs (37%), leading to a preference of utilizing low-error-rate SV/indels for disease monitoring (Figure 6D).

Discussion

Although massively parallel NGS has been driving numerous major discoveries on the genetic underpinnings of human cancers, the error profiles of different types of genetic alterations remain poorly understood. For example, the error rate of substitutions remains ∼0.1% in reports from 2011 to 2018, and our 2019 study indicated that the error rate is between 10−5 and 10−4, or 10- to 100-fold lower than general reports with simple computational error suppression. This has enabled studies on cancer monitoring13 and on detecting clonal hematopoiesis in childhood cancer survivors,29 in whom accurate detection of low-frequency substitutions is of critical importance. However, error profiles of indels and SVs remain elusive in the literature.

In this work, we have rigorously assessed indel and SV errors using ultra-deep and aggregated WGS data. We discovered that error profiles of indels and SVs, except for repeat indels, tend to be much lower than those of SNVs. We revealed that the number of repeat units is highly predictive of error rate of repeat indels. By categorizing repeat indels as homopolymers and microsatellites, we further discovered that deletion of one repeat unit happens nearly 10-fold more often than insertion of one repeat unit for microsatellite indels. We then demonstrated an application of our findings in detecting cancer-relevant mutations from day-28 remission specimens of 72 patients in a relapsed B-ALL cohort, where our NGS method achieved a detection rate (61%) superior to that of a clinical flow-based MRD estimate (51%). Notably, we demonstrated that higher sequencing depth (here ∼120,000× relative to 30,000×) can result in higher detection sensitivity by using a second round of sequencing on 17 cases.

We expect our method to affect the assessment, monitoring, and improvement of cancer diagnosis and treatment. Of note, although SVs and indels can serve as reliable biomarkers with a lower error rate than substitutions, we observed that the mutant alleles are difficult to capture when their AFs are low, such as when <0.1%. We believe that this phenomenon warrants focused study in the future.

Limitations of the study

Our study has a few limitations. First, cost has constrained a deeper sequencing that in turn limited our ability to provide a precise estimate of SV/indel error rate when no mutant reads are observed. Second, the ultra-deep sequencing mandated focused targeted regions so that many genomic regions remain uncharacterized for SV and indel errors. Another limitation of our study is the lack of a comparison with a cohort of patients who did not experience relapse, which warrants a separate study.

We were not aware of any PCR-free methods for targeted sequencing applications that would meet the ultra-high depth requirement, which is a clear limitation of this study. However, it is well known that PCR efficiency, which measures the percentage of input molecules to be successfully amplified in each PCR cycle,4 can be low, and that justifies the need of additional PCR cycles. Innovative methods such as primary template-directed amplification31 have now been developed to overcome biases introduced during PCR and could be an interesting future direction of research.

Resource availability

Lead contact

Further information and requests for resources and questions should be directed to and will be fulfilled by the lead contact, Xiaotu Ma (xiaotu.ma@stjude.org).

Materials availability

Cell lines COLO829 and COLO829BL were purchased from ATCC with stock # CRL-1974 and # CRL-1980.

Data and code availability

Acknowledgments

This work was supported, in part, by NCI R01CA273326 (to X.M.), R01CA293587 (to M.L.L), the NCTN Operations Center Grant (U10CA180886), an NCTN Statistics and Data Center grant (U10CA180899), the Fund for Innovation in Cancer Informatics (to X.M. and J.M.K.), Cancer Center Support grant P30CA021765 (Developmental Fund to J.M.K. and X.M.), and the American Lebanese Syrian Associated Charities (ALSAC). S.M. is partly supported by Leukemia and Lymphoma Society (7025-21) and St. Baldrick's Foundation (SAT-21-064-01-SBF-ACS). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health or other funding agencies. We thank Cherise Guess for editing the manuscript.

Author contributions

X.M. conceived the research. Y.S. generated the ultra-deep-sequencing data for COLO829 and MRD experiment. Y.L. implemented the genotyping scripts. Q.T., Y.F., Y.S., P.K., Y.L., Z.L., L.F., A.M., T.J., A.C., H.M., L.J., B.J.H., J.M.K., S.M., J.Z., W.L.C., M.L.L., J.E., P.A.B., and X.M. performed the analysis. M.L.L., J.E., P.A.B., and X.M. supervised the study. All authors read and approved the final manuscript.

Declaration of interests

The authors declare no competing interests.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Biological samples

Melanoma cell line COLO829 ATCC # CRL-1974
Normal cell line COLO829BL ATCC # CRL-1980

Critical commercial assays

MagAttract HMW DNA Kit Qiagen N/A
Twist library preparation enzymatic fragmentation kit Twist Bioscience N/A
NovaSeq 6000 Illumina N/A
HiSeq X Ten sequencers Illumina N/A
DNA Blood Mini Kit Qiagen cat#51106
DNeasy Blood & Tissue Kit Qiagen cat#69506
Quant-iT dsDNA Assay Kit Life Technologies cat#Q33130
Agarose gel electrophoresis E-Gel, Life Technologies cat#G8008-01

Deposited data

COLO829 dilution data This paper https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1347745
Day 28 B-ALL remission ultra-deep sequencing data This paper https://dbgap.ncbi.nlm.nih.gov/beta/study/phs003409.v1.p1/#study

Software and algorithms

BWA Li et al.32 https://pubmed.ncbi.nlm.nih.gov/20080505/
SAMtools Li et al.33 https://pmc.ncbi.nlm.nih.gov/articles/PMC2723002/
SequencErr Eric Davis et al.10 https://pubmed.ncbi.nlm.nih.gov/33487172/
SVIndelGenotyper This paper34 https://github.com/stjude/SVIndelGenotyper
R N/A https://www.r-project.org/

Other

Supporting code This paper35 https://doi.org/10.5281/zenodo.16969317

Experimental model and study participant details

Cell lines COLO829 and COLO829BL were purchased from ATCC (stock # CRL-1974 and # CRL-1980). The WGS data from 1662 whole-genome samples from a previous St. Jude LIFE (SJLIFE) study (see detailed sample list in Table S9; data available at St Jude Cloud (https://platform.stjude.cloud/requests/cohorts; Study ID: SJC-DS-1002) are included in this work. Another WGS cohort of 103 relapsed B-ALL cases was used to study the Indel and SV error rates.

Method details

COLO829 dilution experiment

The melanoma cell line COLO829 (ATCC CRL-1974) and its matched normal cell line COLO829BL (ATCC CRL-1980, derived from peripheral blood of the same patient) have been extensively studied for somatic DNA variants and are proposed to serve as a reference sample for cancer genome sequencing.6,7 We have previously used this cell line pair to develop computational error suppression methods that achieved a 100- to 1000-fold lower error rate than those in common reports for substitutions. In this work, we randomly selected 41 somatic alterations (including 23 Indels and 17 SVs, Tables S1 and S2) from published studies6,7 in which to study the error rate via dilution experiment. Building on previous findings on substitution errors, we designed four dilution ladders: a) Cancer (100% of COLO829), b) Low (1% of COLO829), c) Very low (0.1% of COLO829), and d) Non-cancer (0% of COLO829).

COLO829BL and COLO829 DNA was extracted by using MagAttract HMW DNA Kit (Qiagen), then the mixture of DNA was generated by spiking 1% and 0.1% of COLO829 into COLO829BL. Two biological replicates were used for 1% and 0.1% dilution samples, and 100% of COLO829BL (non-cancer) and 100% of COLO829 (pure cancer). Altogether, 50 ng of each sample was fragmented by enzyme, ligated to adaptors, and amplified by NEB Q5 Hot Start HiFi DNA Polymerase to generate Illumina DNA library via the Twist library preparation enzymatic fragmentation kit. Then, 500 ng of each biological replicate from each sample was pooled and hybridized at 70°C for 16 h with a custom designed Twist panel. After a hybridization wash, each sample was split into 5 reactions for 16 cycles of PCR amplification with NEB Q5 Hot Start HiFi DNA Polymerase. Each capture library was diluted to 10 nM and pooled by ratio to generate 50,000X coverage for the 1% and 0.1% spike-in, and 5,000,000X coverage for COLO829BL on NovaSeq 6000 S2 paired-end 2 × 151 cycles sequencing. We achieved an average depth of ∼40,000X for 100%, 1%, and 0.1% concentrations and ∼3,000,000X for 0% concentration.

Read counting

As illustrated in Figures 1C–1E, A dedicated computer algorithm was written to ensure the accurate counting of sequencing reads. Taking Indel as an example, we first defined the core reference (green) and mutant (orange) sequences. Without loss of generality, the Indel is a simple insertion if the core reference has length zero; the Indel is a simple deletion if the core mutant has length zero; the Indel is a complex Indel if core reference and core mutant both have non-zero length.

To account for the scenario in which the Indel is in a tandem repeat region that can be very long, we added an 8-bp buffer sequence (blue in Figure 1C) to the core reference and core mutant sequences on both sides for the matching. A read is considered matched to the mutant allele if 1) the read fully covers the core mutant sequence and associated buffer sequences; 2) the read matches nearly perfectly (<1% mismatch) with the sequences outside the core/buffer sequences. To account for the scenario in which the read has only partial match to the core mutant sequences (in which case the matches fall into the read ends) and to simultaneously account for the known elevated error rate in read ends,9 we counted the read as supporting the mutant allele if the partial match to the core mutant sequence had length ≥8 bps, which corresponds to a random chance of 4−8 = 2 × 10−5. To account for the scenario in which a secondary substitution occurs on top of an Indel,36 we did not allow any additional mismatches/Indels within the core reference/mutant sequence. Instead, we left this problem to the initial curation algorithm to accurately curate the Indel for accurate error profiling, which in turn helps to determine p values to make a call. To account for the scenario in which an insert sequence is short so that both of its mate pairs support the same reference/mutation allele, our algorithm further checks the read names so that mate-pairs are counted only once, as we have done previously.9,10 This algorithm also applies to reference and mutant allele counting for Indels/SVs.

Whole-genome sequencing dataset

DNA was extracted from stored samples by using either the QIAamp DNA Blood Mini Kit (QIAGEN cat#51106) or the DNeasy Blood & Tissue Kit (cat# 69506). After extraction, the DNA concentration was fluorometrically measured by using the Quant-iT dsDNA Assay Kit (Life Technologies cat#Q33130), and DNA integrity was verified visually by agarose gel electrophoresis (E-Gel, Life Technologies, cat#G8008-01). WGS was performed at the HudsonAlpha Institute for Biotechnology Genomic Services Laboratory (Huntsville, AL, USA) by using Illumina HiSeq X Ten sequencers. A total of 1662 whole-genome samples from a previous St. Jude LIFE (SJLIFE) study (see detailed sample list in Table S9; data available at St Jude Cloud (https://platform.stjude.cloud/requests/cohorts; Study ID: SJC-DS-1002)) were included in this work.

WGS data were analyzed for error rate of repeat Indels or SVs. To account for polymorphisms, within each sample, only loci with ≥20X coverage and >95% (so that binomial p value of observing 1 non-reference alleles from 20 reads is 4 × 10−5 and binomial p value of observing 2 non-reference alleles from 40 reads is 1.5 × 10−9 given the locus is heterozygous) reads being reference allele were merged into a single-count file. Loci with heterozygous or homozygous alternative allele calls (i.e., no alleles with fraction > 95%) in any subject were excluded from analysis. We used only loci with ≥20,000X collapsed coverage in our error analysis for this dataset. A total of 13,264 repeat regions were (Table S11) collected from 339 cancer genes (Table S10). With the requirement of 20,000X total depth from homozygous reference donors, we analyzed repeat Indels from 11,378 regions in this work, including 22,653 deletions (11,314 being deletion of 1 repeat unit) and 22,733 insertions (11,341 being insertion of 1 repeat unit). The limited total sequencing depth of ∼50,000X did not provide sufficient power to study the error rate of insertion or deletion of 2 repeat units. For example, 14,527 Indels (64%; of 22,655 total) with insertion or deletion of 1 repeat unit have at least 1 mutant read; while only 1210 Indels (5%; of 22,731 total) with insertion or deletion of 2 repeat units have at least 1 mutant read observed. However, a similar trend was observed (Figure S2). The same analytical procedure is applied to the study of Indel and SV errors from 103 relapsed B-ALL cases13 (Tables S7 and S8).

Remission DNA dataset from patients with relapsed childhood ALL

The tumor burden measured from end-of-induction (day 28) samples is known to reflect responsiveness to therapy and is used as a highly prognostic marker for childhood ALL treatment stratification, with patients having undetectable or very low level (<0.01%) cancer cells typically having low risk of relapse while patients with a higher level (>1%) may have a refractory disease and high risk of relapse. To evaluate the clinical utility of our findings on Indel and SV error rates, we analyzed day 28 remission samples from 72 patients with newly diagnosed childhood B-ALL who subsequently experienced relapse. The diagnosis and relapse tumors of these cases have been sequenced by using whole-genome sequencing; identified somatic alterations will be published elsewhere. The study has been approved by the St. Jude Institutional Review Board. All patient samples have been de-identified by using St. Jude identifier or Children’s Oncology Group’s USI (Tables S13, S14, and S15).

Because of the relapse, the day 28 samples are expected to contain cancer cells and a higher detection rate would imply that the method is more sensitive, although the specimen may not contain cancer cells due to sampling fluctuation when the tumor burden is low (e.g., between 10−6 and 10−5), and we do not expect a 100% detection rate. To enable ultra-deep sequencing, we selected ∼5 somatic markers (range: 2–7; total n = 397; Table S13) for each patient, which included 316 SNVs, 34 Indels, and 47 SVs across 72 cases. These markers are selected based on two criteria: 1) the markers are present in diagnosis and relapse tumors so that they are expected to exist in remission; 2) the markers are known to have low error rate. For example, we only included substitutions A>C/A>T/G>C for SNV markers. For Indels, we excluded repeat Indels. For SVs, we prefer the subtype-defining rearrangements such as that for ETV6-RUNX1. The samples were sequenced to an average depth of 30,000X in the round 1 experiment.

For each specimen, 100 ng of DNA was fragmented by enzyme, adaptors were ligated to generate Illumina DNA library, and the Twist library preparation enzymatic fragmentation kit (Twist Bioscience) was used. Each 7 to 8 DNA libraries were pooled together and hybridized at 70°C for 16 h with a custom-designed SNV/Indel/SV capture panel. After a hybridization wash, each capture was split into 2 reactions for amplification using KAPA HiFi HotStart Ready Mix. Final libraries were sequenced on an Illumina NovaSeq 6000 S4 using paired-end 2 × 151 cycles sequencing to reach the designed coverage.

Upon sequencing, the paired-end FASTQ files were subject to adapter trimming followed by mapping to the human reference genome GRCh37 using BWA 0.7.12-r1039.32 For SNVs, we used SequencErr v2.0.910 to identify and filter flow cell tiles with high error rate (>100 per million), followed by allele counting for all the designed markers across all samples. For Indels and SVs, we performed allele counts by using custom scripts34 using the algorithm outlines in Figure 1. Similar to previous work,9,10 poor quality reads (>5% bases with Phred quality score <20) are filtered.

To detect variants, we utilized a previously established “rotation control” approach.13 Here, the remission sample is considered as an index sample for a given SNV/Indel/SV marker observed in corresponding diagnosis/relapse tumor samples. All remission samples from other cases without this marker are considered background control samples, and read counts of these background control samples are aggregated (to increase total depth) to estimate the sequencing error rate of this marker. Denote the background coverage as tb, the background mutant allele counts as mb, and the background error rate as rb = mb/tb. In the mutation-positive sample, denote the total (foreground) allele and mutant allele count as tf and mf, respectively. We then calculated the probability P of observing ≥ mf mutant reads out of tf reads by random chance rb(sequencing/library errors), using binomial distribution. With false-discovery rate control (using R package “qvalue”), a sample was called MRD-positive if the Q value < 0.001 (calculated across all sample/marker combinations) for any designed markers. The MRD level of an MRD-positive sample is estimated by using formula imif/itif, where i = 1,2,… denote designed markers and mif and tif denote the number of mutant and total reads for marker i, respectively.

For 17 cases, we performed a second round of sequencing with ∼120,000X depth by using the same TWIST DNA library as in the first-round experiment. Using a similar analysis as described above, 11 cases (Table S15) received stronger evidence of detection than in the first round, 4 cases remained negative, and 2 cases received weaker evidence of detection than in the first round, possibly due to PCR sampling fluctuation.

Power consideration and pseudo-counts

With more advanced sequencing technology and more rigorous sequencing methods (such as the avoidance of 8-oxoG damage12), the sequencing error rate can become a lot lower9,10 (e.g., <10−4-10−5) so that the sequencing data need higher depth (typically >105 depth) to sufficiently power the error rate estimate.

In general, there are two ways to estimate the sequencing error rate. When the sequencing depth is shallow, it is assumed that the sporadic sequencing errors observed in large genomic intervals are mutually comparable and can be “merged” together. However, our previous study has clearly indicated context effect (Figure 3 of Ma et al.;9) that indicates a simple merging could mask key biological insights. As such, in this work we aimed to investigate the locus-specific error rate estimate by generating sequencing data with millions of depths. However, even at this ultra-expensive sequencing depth, we still experience a limit-of-detection problem due to the extremely low error rate of SVs and Indels (excluding repeat Indels). We reasoned that it is impractical to add more sequencing depths. As such, we employed a “pseudo-count” method by assuming that, if we can sequence one more read, it will carry the mutation that we expect to detect (which clearly leads to an over-estimate of the sequencing error rate but here we consider the over-estimation will lead to conservative conclusions, which is a good practice). To avoid confusion, we applied the pseudo-count throughout the study. As such, please note that the over-estimation problem will be more severe when the overall sequencing depth is shallower, and it is less severe when the overall sequencing depth is much higher.

Defining repeat indels

Genomic regions of 339 well-established cancer genes were selected.28,36 A total of 11,378 loci with at least 3 consecutive repeat units within these regions were identified (https://doi.org/10.5281/zenodo.16969317).35 Six Indel markers were created for each locus, including deletion of 1-, 2-, and 3-unit losses, and insertion of 1-, 2-, and 3-units. Allele counts were calculated with custom scripts34 (see Code Availability) following the algorithm outlined in Figure 1. The same approach was applied to identify repeat regions in the 47 targeted sequencing samples from the Shanghai Children's Medical Center (SCMC) cohort (Figure 5A), yielding 1,512 loci for the analysis.

Quantification and statistical analysis

Quantitative and statistical analyses are described in detail within each section of the method details and in the main text. To assess the repeat Indel error rates across samples differing in quality (n ranges from 1200 to 2 across different repeat unit sizes), sequencing instrument (n = 118), and PCR enzymes (n = 118) (Figure 5), the Wilcoxon rank-sum test was used. Differences in error rate in repeat indels across varying repeat units and region length were compared by using ANOVA (Figures 4 and S3) (Sample sizes shown in Figure 4 for each group). The binomial test was used to estimate the likelihood of observing non-template polymorphism in sequence reads (see method details).

Published: December 2, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.xgen.2025.101082.

Contributor Information

Mignon L. Loh, Email: mignon.loh@seattlechildrens.org.

John Easton, Email: john.easton@stjude.org.

Patrick A. Brown, Email: pbrown.0314@gmail.com.

Xiaotu Ma, Email: xiaotu.ma@stjude.org.

Supplemental information

Document S1. Figures S1–S3
mmc1.pdf (955.6KB, pdf)
Table S1. Designed SNVs and indels for the COLO829 dilution experiment, related to Figure 2
mmc2.xlsx (11.1KB, xlsx)
Table S2. Designed SVs for the COLO829 dilution experiment, related to Figure 2
mmc3.xlsx (11.5KB, xlsx)
Table S3. Detecting SNV, indel, and SV markers from the COLO829 dilution experiment, related to Figure 2
mmc4.xlsx (20.7KB, xlsx)
Table S4. SNV error profiles using the current dataset, COLO829, and a previously published cohort from Ma et al. related to Figure 2
mmc5.xlsx (1.3MB, xlsx)
Table S5. Indels from 103 patients with relapsed ALL, related to Figure 3
mmc6.xlsx (23.9KB, xlsx)
Table S6. SVs from 103 patients with relapsed ALL, related to Figure 3
mmc7.xlsx (109.6KB, xlsx)
Table S7. Error rate of indels from 103 patients with relapsed ALL in WGS data from healthy SJLIFE donors from Table S9, related to Figure 3
mmc8.xlsx (61.1KB, xlsx)
Table S8. Error rate of SVs from 103 patients with relapsed ALL in WGS data from healthy SJLIFE donors from Table S9, related to Figure 3
mmc9.xlsx (138.9KB, xlsx)
Table S9. List of 1,662 SJLIFE samples, related to Figures 3 & 4
mmc10.xlsx (75.3KB, xlsx)
Table S10. List of 339 cancer driver genes with repeat indels studied in this work, related to Figure 4
mmc11.xlsx (13.4KB, xlsx)
Table S11. Error profile of repeat indels from 339 genes, related to Figure 4
mmc12.xlsx (4MB, xlsx)
Table S12. Indel markers identified from repeat regions by using the deep targeted sequencing data from SCMC cohort, EGA accession EGAS00001003444, related to Figure 5
mmc13.xlsx (99.9KB, xlsx)
Table S13. Detecting tumors from day-28 remission samples from a cohort of 72 patients with relapsed B-ALL, related to Figure 6
mmc14.xlsx (72.5KB, xlsx)
Table S14. Comparison of clinical-flow MRD and NGS MRD estimate from day-28 remission samples from a cohort of 72 patients with relapsed B-ALL from COG study AALL1331, related to Figure 6
mmc15.xlsx (14KB, xlsx)
Table S15. Deeper, i.e.,120,000×, sequencing in round 2 results in stronger signal than in round 1, i.e., ∼30,000×, related to Figure 6
mmc16.xlsx (10.8KB, xlsx)
Document S2. Transparent peer review records for Shao et al.
mmc17.pdf (1.2MB, pdf)
Document S3. Article plus supplemental information
mmc18.pdf (5MB, pdf)

References

  • 1.Glenn T.C. Field guide to next-generation DNA sequencers. Mol. Ecol. Resour. 2011;11:759–769. doi: 10.1111/j.1755-0998.2011.03024.x. [DOI] [PubMed] [Google Scholar]
  • 2.Mardis E.R. Next-generation sequencing platforms. Annu. Rev. Anal. Chem. 2013;6:287–303. doi: 10.1146/annurev-anchem-062012-092628. [DOI] [PubMed] [Google Scholar]
  • 3.Goodwin S., McPherson J.D., McCombie W.R. Coming of age: ten years of next-generation sequencing technologies. Nat. Rev. Genet. 2016;17:333–351. doi: 10.1038/nrg.2016.49. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Salk J.J., Schmitt M.W., Loeb L.A. Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations. Nat. Rev. Genet. 2018;19:269–285. doi: 10.1038/nrg.2017.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Rugbjerg P., Sommer M.O.A. Overcoming genetic heterogeneity in industrial fermentations. Nat. Biotechnol. 2019;37:869–876. doi: 10.1038/s41587-019-0171-6. [DOI] [PubMed] [Google Scholar]
  • 6.Neuzil K.M. Interplay between Emerging SARS-CoV-2 Variants and Pandemic Control. N. Engl. J. Med. 2021;384:1952–1954. doi: 10.1056/NEJMe2103931. [DOI] [PubMed] [Google Scholar]
  • 7.Rockett R., Basile K., Maddocks S., Fong W., Agius J.E., Johnson-Mackinnon J., Arnott A., Chandra S., Gall M., Draper J., et al. Resistance Mutations in SARS-CoV-2 Delta Variant after Sotrovimab Use. N. Engl. J. Med. 2022;386:1477–1479. doi: 10.1056/NEJMc2120219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Liu Y., Klein J., Bajpai R., Dong L., Tran Q., Kolekar P., Smith J.L., Ries R.E., Huang B.J., Wang Y.C., et al. Etiology of oncogenic fusions in 5,190 childhood cancers and its clinical and therapeutic implication. Nat. Commun. 2023;14:1739. doi: 10.1038/s41467-023-37438-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ma X., Shao Y., Tian L., Flasch D.A., Mulder H.L., Edmonson M.N., Liu Y., Chen X., Newman S., Nakitandwe J., et al. Analysis of error profiles in deep next-generation sequencing data. Genome Biol. 2019;20:50. doi: 10.1186/s13059-019-1659-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Davis E.M., Sun Y., Liu Y., Kolekar P., Shao Y., Szlachta K., Mulder H.L., Ren D., Rice S.V., Wang Z., et al. SequencErr: measuring and suppressing sequencer errors in next-generation sequencing data. Genome Biol. 2021;22:37. doi: 10.1186/s13059-020-02254-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Stoler N., Nekrutenko A. Sequencing error profiles of Illumina sequencing instruments. NAR Genom. Bioinform. 2021;3 doi: 10.1093/nargab/lqab019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Costello M., Pugh T.J., Fennell T.J., Stewart C., Lichtenstein L., Meldrim J.C., Fostel J.L., Friedrich D.C., Perrin D., Dionne D., et al. Discovery and characterization of artifactual mutations in deep coverage targeted capture sequencing data due to oxidative DNA damage during sample preparation. Nucleic Acids Res. 2013;41 doi: 10.1093/nar/gks1443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Li B., Brady S.W., Ma X., Shen S., Zhang Y., Li Y., Szlachta K., Dong L., Liu Y., Yang F., et al. Therapy-induced mutations drive the genomic landscape of relapsed acute lymphoblastic leukemia. Blood. 2020;135:41–55. doi: 10.1182/blood.2019002220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Druker B.J., Talpaz M., Resta D.J., Peng B., Buchdunger E., Ford J.M., Lydon N.B., Kantarjian H., Capdeville R., Ohno-Jones S., Sawyers C.L. Efficacy and safety of a specific inhibitor of the BCR-ABL tyrosine kinase in chronic myeloid leukemia. N. Engl. J. Med. 2001;344:1031–1037. doi: 10.1056/NEJM200104053441401. [DOI] [PubMed] [Google Scholar]
  • 15.Morgillo F., Della Corte C.M., Fasano M., Ciardiello F. Mechanisms of resistance to EGFR-targeted drugs: lung cancer. ESMO Open. 2016;1 doi: 10.1136/esmoopen-2016-000060. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Antar A.I., Otrock Z.K., Jabbour E., Mohty M., Bazarbachi A. FLT3 inhibitors in acute myeloid leukemia: ten frequently asked questions. Leukemia. 2020;34:682–696. doi: 10.1038/s41375-019-0694-3. [DOI] [PubMed] [Google Scholar]
  • 17.Tran Q., Gao S., Phan V. Analysis of optimal alignments unfolds aligners' bias in existing variant profiles. BMC Bioinf. 2016;17:349. doi: 10.1186/s12859-016-1216-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Ye K., Wang J., Jayasinghe R., Lameijer E.W., McMichael J.F., Ning J., McLellan M.D., Xie M., Cao S., Yellapantula V., et al. Systematic discovery of complex insertions and deletions in human cancers. Nat. Med. 2016;22:97–104. doi: 10.1038/nm.4002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Hagiwara K., Edmonson M.N., Wheeler D.A., Zhang J. indelPost: harmonizing ambiguities in simple and complex indel alignments. Bioinformatics. 2022;38:549–551. doi: 10.1093/bioinformatics/btab601. [DOI] [PubMed] [Google Scholar]
  • 20.Wang Z., Wilson C.L., Easton J., Thrasher A., Mulder H., Liu Q., Hedges D.J., Wang S., Rusch M.C., Edmonson M.N., et al. Genetic Risk for Subsequent Neoplasms Among Long-Term Survivors of Childhood Cancer. J. Clin. Oncol. 2018;36:2078–2087. doi: 10.1200/JCO.2018.77.8589. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Borowitz M.J., Wood B.L., Keeney M., Hedley B.D. Measurable Residual Disease Detection in B-Acute Lymphoblastic Leukemia: The Children's Oncology Group (COG) Method. Curr. Protoc. 2022;2 doi: 10.1002/cpz1.383. [DOI] [PubMed] [Google Scholar]
  • 22.Ma X., Liu Y., Liu Y., Alexandrov L.B., Edmonson M.N., Gawad C., Zhou X., Li Y., Rusch M.C., Easton J., et al. Pan-cancer genome and transcriptome analyses of 1,699 paediatric leukaemias and solid tumours. Nature. 2018;555:371–376. doi: 10.1038/nature25795. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Pleasance E.D., Cheetham R.K., Stephens P.J., McBride D.J., Humphray S.J., Greenman C.D., Varela I., Lin M.L., Ordóñez G.R., Bignell G.R., et al. A comprehensive catalogue of somatic mutations from a human cancer genome. Nature. 2010;463:191–196. doi: 10.1038/nature08658. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Craig D.W., Nasser S., Corbett R., Chan S.K., Murray L., Legendre C., Tembe W., Adkins J., Kim N., Wong S., et al. A somatic reference standard for cancer genome sequencing. Sci. Rep. 2016;6 doi: 10.1038/srep24607. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wang J., Mullighan C.G., Easton J., Roberts S., Heatley S.L., Ma J., Rusch M.C., Chen K., Harris C.C., Ding L., et al. CREST maps somatic structural variation in cancer genomes with base-pair resolution. Nat. Methods. 2011;8:652–654. doi: 10.1038/nmeth.1628. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Vogelstein B., Papadopoulos N., Velculescu V.E., Zhou S., Diaz L.A., Jr., Kinzler K.W. Cancer genome landscapes. Science. 2013;339:1546–1558. doi: 10.1126/science.1235122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Grobner S.N., Worst B.C., Weischenfeldt J., Buchhalter I., Kleinheinz K., Rudneva V.A., Johann P.D., Balasubramanian G.P., Segura-Wang M., Brabetz S., et al. The landscape of genomic alterations across childhood cancers. Nature. 2018;555:321–327. doi: 10.1038/nature25480. [DOI] [PubMed] [Google Scholar]
  • 28.Edmonson M.N., Patel A.N., Hedges D.J., Wang Z., Rampersaud E., Kesserwan C.A., Zhou X., Liu Y., Newman S., Rusch M.C., et al. Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): a cloud-based platform for curating and classifying germline variants. Genome Res. 2019;29:1555–1565. doi: 10.1101/gr.250357.119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Hagiwara K., Natarajan S., Wang Z., Zubair H., Mulder H.L., Dong L., Plyler E.M., Thimmaiah P., Ma X., Ness K.K., et al. Dynamics of Age- versus Therapy-Related Clonal Hematopoiesis in Long-term Survivors of Pediatric Cancer. Cancer Discov. 2023;13:844–857. doi: 10.1158/2159-8290.CD-22-0956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Matuszek Z., Arbab M., Kesavan M., Hsu A., Roy J.C.L., Zhao J., Yu T., Weisburd B., Newby G.A., Doherty N.J., et al. Base editing of trinucleotide repeats that cause Huntington's disease and Friedreich's ataxia reduces somatic repeat expansions in patient cells and in mice. Nat. Genet. 2025;57:1437–1451. doi: 10.1038/s41588-025-02172-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Gonzalez-Pena V., Natarajan S., Xia Y., Klein D., Carter R., Pang Y., Shaner B., Annu K., Putnam D., Chen W., et al. Accurate genomic variant detection in single cells with primary template-directed amplification. Proc. Natl. Acad. Sci. USA. 2021;118 doi: 10.1073/pnas.2024176118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Li H., Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25:1754–1760. doi: 10.1093/bioinformatics/btp324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Li H., Handsaker B., Wysoker A., Fennell T., Ruan J., Homer N., Marth G., Abecasis G., Durbin R., 1000 Genome Project Data Processing Subgroup The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–2079. doi: 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Liu Y., Ma X. SVIndelGenotyper. 2023. https://github.com/stjude/SVIndelGenotyper
  • 35.Ma X., Feng Y., Kolekar P. Analysis of error profiles of Indels and structural variants in deep sequencing data. Zenodo. 2025 doi: 10.1016/j.xgen.2025.101082. https://zenodo.org/records/16969317 [DOI] [PubMed] [Google Scholar]
  • 36.Kolekar P., Balagopal V., Dong L., Liu Y., Foy S., Tran Q., Mulder H., Huskey A.L.W., Plyler E., Liang Z., et al. SJPedPanel: A Pan-Cancer Gene Panel for Childhood Malignancies to Enhance Cancer Monitoring and Early Detection. Clin. Cancer Res. 2024;30:4100–4114. doi: 10.1158/1078-0432.CCR-24-1063. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S3
mmc1.pdf (955.6KB, pdf)
Table S1. Designed SNVs and indels for the COLO829 dilution experiment, related to Figure 2
mmc2.xlsx (11.1KB, xlsx)
Table S2. Designed SVs for the COLO829 dilution experiment, related to Figure 2
mmc3.xlsx (11.5KB, xlsx)
Table S3. Detecting SNV, indel, and SV markers from the COLO829 dilution experiment, related to Figure 2
mmc4.xlsx (20.7KB, xlsx)
Table S4. SNV error profiles using the current dataset, COLO829, and a previously published cohort from Ma et al. related to Figure 2
mmc5.xlsx (1.3MB, xlsx)
Table S5. Indels from 103 patients with relapsed ALL, related to Figure 3
mmc6.xlsx (23.9KB, xlsx)
Table S6. SVs from 103 patients with relapsed ALL, related to Figure 3
mmc7.xlsx (109.6KB, xlsx)
Table S7. Error rate of indels from 103 patients with relapsed ALL in WGS data from healthy SJLIFE donors from Table S9, related to Figure 3
mmc8.xlsx (61.1KB, xlsx)
Table S8. Error rate of SVs from 103 patients with relapsed ALL in WGS data from healthy SJLIFE donors from Table S9, related to Figure 3
mmc9.xlsx (138.9KB, xlsx)
Table S9. List of 1,662 SJLIFE samples, related to Figures 3 & 4
mmc10.xlsx (75.3KB, xlsx)
Table S10. List of 339 cancer driver genes with repeat indels studied in this work, related to Figure 4
mmc11.xlsx (13.4KB, xlsx)
Table S11. Error profile of repeat indels from 339 genes, related to Figure 4
mmc12.xlsx (4MB, xlsx)
Table S12. Indel markers identified from repeat regions by using the deep targeted sequencing data from SCMC cohort, EGA accession EGAS00001003444, related to Figure 5
mmc13.xlsx (99.9KB, xlsx)
Table S13. Detecting tumors from day-28 remission samples from a cohort of 72 patients with relapsed B-ALL, related to Figure 6
mmc14.xlsx (72.5KB, xlsx)
Table S14. Comparison of clinical-flow MRD and NGS MRD estimate from day-28 remission samples from a cohort of 72 patients with relapsed B-ALL from COG study AALL1331, related to Figure 6
mmc15.xlsx (14KB, xlsx)
Table S15. Deeper, i.e.,120,000×, sequencing in round 2 results in stronger signal than in round 1, i.e., ∼30,000×, related to Figure 6
mmc16.xlsx (10.8KB, xlsx)
Document S2. Transparent peer review records for Shao et al.
mmc17.pdf (1.2MB, pdf)
Document S3. Article plus supplemental information
mmc18.pdf (5MB, pdf)

Data Availability Statement


Articles from Cell Genomics are provided here courtesy of Elsevier

RESOURCES