Skip to main content
Nature Portfolio logoLink to Nature Portfolio
. 2026 May 19;58(6):1308–1319. doi: 10.1038/s41588-026-02592-0

Patterns and drivers of 43,617 mosaic chromosomal alterations in blood

David Tang 1,2,3,4,✉, Nolan Kamitaki 1,2,3,4, Ronen E Mukamel 2,3,4, Simone Rubinacci 2,3,4,5, Po-Ru Loh 2,3,4,✉
PMCID: PMC13263154  PMID: 42156563

Abstract

Clonal expansions of hematopoietic cells carrying mosaic chromosomal alterations (mCAs) are commonly detectable in elderly individuals. Here we studied 43,617 autosomal mCAs ascertained in 484,081 UK Biobank participants using new, high-resolution computational methods to analyze blood-derived, whole-genome sequencing data. Shorter mCAs (≤1 Mb) clustered at 53 genomic hotspots (46 previously undetected), several of which implicated chromosomal fragile sites as a recurrent source of somatic deletions. Chronic lymphocytic leukemia (CLL)-associated deletions at 13q14 were detectable in 1% of individuals aged 65–70 years, suggesting opportunities for incorporating this mosaic mutation into clinical screening and in genetic association studies of CLL. Rare protein-coding variants in 38 genes associated (P < 1.2 × 10−5; false recovery rate <0.01) with clonal expansions of copy-neutral loss-of-heterozygosity (CN-LOH) mutations that modified the allelic dosages of these variants, identifying likely targets of clonal selection via CN-LOH-induced allelic substitution. These results show that our blood genomes often accrue mCAs in predictable ways as they evolve with age.

Subject terms: Genome-wide association studies, Haematological cancer


High-resolution analyses of blood-derived whole-genome sequence data from UK Biobank detect new mosaic chromosomal alterations and identify rare protein-coding variants associated with clonal expansions of copy-neutral loss-of-heterozygosity mutations.

Main

Hematopoietic stem and progenitor cells typically acquire 500–1,500 somatic mutations across an individual’s lifespan1–3. Although most somatic mutations do not alter a cell’s clonal fitness, some mutations confer a proliferative advantage to the cell and its progeny, such that they clonally expand to constitute an outsized proportion of the individual’s blood cells3–5. This phenomenon, termed clonal hematopoiesis (CH), commonly occurs during aging, such that expansions with clonal fractions of several percent are frequently observed in elderly individuals6–9 and subtler clonal expansions are ubiquitous by middle age10. CH confers a tenfold increased risk of progression to hematological malignancies6–9 and is also a risk factor for several other diseases11–13.

CH can be classified by the types of mutations carried in expanded clones: CH of indeterminate potential (CHIP) involves somatic SNVs and insertions and/or deletions (indels) in leukemia driver genes8,9, whereas CH with mosaic chromosomal alterations (mCAs) involves megabase-scale to chromosome-scale mutations6,7. Most hematopoietic clones have unknown drivers (CH-UD)3,8,14,15. Inherited genetic variation modulates risk of all types of CH, exhibiting some shared but some distinct effects on CHIP16–20, autosomal20–23 and sex chromosome mCAs24,25 and CH-UD14,15. Similarly, different CH mutations increase the risk of different blood cancers26,27 and other diseases13. Here we focus on autosomal mCAs, which are more readily detectable at lower cell fractions compared to point mutations21–23,28,29, include common drivers of lymphoid malignancies26,30,31 and provide unique insight into genetic drivers of CH by their action on inherited and acquired variation21–23,32,33.

Despite remarkable progress over the past decade in understanding the causes and consequences of CH13, many questions remain, particularly regarding mCAs. Nearly all large-scale studies of mCAs have analyzed SNP-array allelic intensity data6,7,12,21–26,28, limiting detection of shorter mCAs (~1 Mb) and mCAs with low clonal fractions (<1%). In addition, the sample sizes of existing autosomal mCA call sets have not yet permitted precise natural history studies of specific mCAs. Recent work has demonstrated that analyzing whole-genome sequencing (WGS) data can increase mCA detection sensitivity29,34. Population-sequencing datasets also enable more comprehensive exploration of the influences of inherited rare variants on clonal expansion of mCAs, building on previous results20,22,29. Here we developed a new computational method optimized for detecting mCAs from WGS data and applied it to the UK Biobank (UKB)35,36, generating a rich catalog of mCAs and uncovering new insights into their genetic drivers and progression to malignancy.

Results

Analysis of UKB WGS data detects 43,617 autosomal mCAs

Our new method to detect mCAs from high-coverage (~30×) WGS data improves on the mosaic chromosomal alteration (MoChA) software pipeline21,22 by: (1) carefully modeling WGS depth of coverage and filtering kilobase-scale germline copy number variants (CNVs); (2) counting sequenced DNA fragments derived from each of an individual’s two haplotypes; and (3) calling mCAs by analyzing allelic imbalances and determining copy-number states from denoised read-depth data (Fig. 1a, Supplementary Figs. 1–3 and Methods). For computational efficiency, we performed mCA calling on a subset of chromosomes nominated by an initial run of MoChA with permissive parameter settings; this prefiltering procedure reduced detection sensitivity by only ~1.5% (Supplementary Note).

Fig. 1. Detection of 43,617 autosomal mCAs from WGS of 484,081 UKB participants.

Fig. 1

a, Overview of the computational pipeline for calling mCAs from WGS data. b, Genomic ‘pileup’ plot showing chromosomal coordinates of mCA calls. Each y coordinate above a chromosome corresponds to a single individual and shows all mCAs called on that chromosome. Most individuals have only a single mosaic loss (blue), CN-LOH (gold) or gain (red); individuals with multiple mCA calls on the chromosome are shown on top (for example, isochromosome 17 and second-hit 13q CN-LOH amplifying a 13q deletion). c, WGS read-depth deviation (that is, fractional increase or decrease relative to expectation assuming diploid copy number) and allelic imbalance within the boundaries of mCAs cluster along three curves corresponding to the relationship between these quantities for duplications (gain), CN-LOH mutations and deletions (loss). Each mCA was assigned a copy-number state based on linear separators, with mCAs in V(D)J recombination regions treated separately (Supplementary Note). The solid curves indicate a theoretical relationship between read-depth and allelic imbalance for each mCA class. d, Comparison of WGS versus SNP-array-based mCA detection across age groups. e, Comparison of WGS versus SNP-array-based mCA detection as a function of mCA allelic imbalance and size. HMM, hidden Markov model.

Applying this approach to blood-derived WGS data for 484,081 UKB participants (aged ~40–70 years at sample acquisition), we detected 43,617 autosomal mCAs in 35,033 individuals (Fig. 1b,c, Supplementary Figs. 4–25 and Supplementary Tables 1 and 2). This represented a twofold increase in detection sensitivity compared to our previous SNP-array-based analysis of the same cohort22 (Fig. 1d). Beyond improving sensitivity, the WGS-based pipeline assigned copy-number state to all mCAs (Fig. 1c and Supplementary Figs. 26 and 27), unlike SNP-array analysis (Fig. 1d), and mCA calls required minimal filtering for technical artifacts (Supplementary Note), unlike previous analyses of SNP-array data21,22 and WGS data29.

Multiple lines of evidence supported the validity of the mCA calls. First, the genomic distributions of loss, copy-neutral loss of heterozygosity (CN-LOH) and gain mCAs (Fig. 1b) were consistent with previous studies6,7,22,23,28,29. Second, mCA detections increased markedly with age, from 0.04 autosomal mCAs per individual aged 40–44 years to 0.14 mCAs per individual aged 65–69 years (Fig. 1d). Third, whole-exome sequencing (WES) data from the UKB37 replicated the allelic imbalances detected from WGS at the expected rates (Supplementary Fig. 28). Fourth, the WGS call set largely comprised a superset of the SNP-array-based calls22: among 16,977 individuals with an mCA called from the SNP-array data, 15,299 (90%) had a WGS-based mCA call on the same chromosome, typically with the same copy-number state (Supplementary Fig. 29).

WGS-based mCA analysis particularly improved detection sensitivity for two classes of mCAs: (1) large mCAs with low clonal fractions and (2) small mCAs with moderate-to-high clonal fractions (Fig. 1e). Specifically, WGS analysis detected 2.8-fold more mCAs with allelic imbalance <0.01 (that is, <51% representation of the overrepresented haplotype) and 10.5-fold more mCAs shorter than 1 Mb (Fig. 1e). The WGS analysis also identified 5.6-fold more whole-chromosome CN-LOH mutations (624 versus 111 from SNP-array analysis), reflecting the low clonal fractions of most somatic isodisomies.

WGS-based analysis identifies focal mCAs and mCA breakpoints

The increased power to detect short, interstitial mCAs from WGS data enabled deeper characterization of ‘hotspots’ (Methods) at which short mCAs are recurrently found in blood DNA. Several deletion hotspots have previously been observed, particularly at DLEU2 (13q14), DNMT3A and TET27,21,28. Here higher-resolution, WGS-based analysis identified 41 genomic loci (34 new relative to SNP-array analysis22) at which at least 10 UKB participants carried sub-megabase somatic deletions, along with 12 sites of recurrent short duplications (Fig. 2a,b and Supplementary Table 3). Duplication hotspots have not been observed from SNP-array data. To validate these new interstitial mCA hotspots, we searched for WGS read pairs that mapped ‘discordantly’ near the opposite ends of an mCA (Supplementary Note). Such read pairs are sometimes generated when a sequenced DNA fragment spans an mCA breakpoint (in a cell carrying the mutation) and are rarely observed by chance. For 48 of the 53 recurrent short mCAs, we observed discordant read support for >25% of mCA calls (Supplementary Table 3). The mCA calls at each hotspot exhibited varying breakpoint locations and varying levels of allelic imbalance, further supporting somatic origin (Fig. 2a and Supplementary Table 3).

Fig. 2. Hotspots of focal mosaic deletions and analysis of mCA breakpoints.

Fig. 2

a, Pileups of mCAs at four loci where focal deletions were newly detected from WGS-based analysis. Approximate breakpoints were determined from discordant read pairs and are accurate to ~1 kb. b, Properties of focal deletions observed in at least 10 UKB participants (precise n provided in Supplementary Table 3). The focal index of a deletion hotspot is a measure of the extent to which deletions at the locus overlap a single shared deletion region (Methods). Deletions at fragile sites (pink) tended to be less age enriched and less focal compared to deletions at CH driver genes (green). Deletion hotspots indicated as ‘new from WGS’ (filled circles) are those that were detected at least 5× more frequently from WGS data compared to SNP-array data. The error bars are 95% confidence intervals (CIs). c, Fraction of carriers of each of the 41 focal deletions who are male. The error bars are 95% CIs. *Two-sided binomial P < 0.05; **two-sided binomial P < 0.05, Bonferroni adjusted. d, Distribution of the number of basepairs of microhomology between the left and right breakpoints of mosaic deletions for which split reads resolved the breakpoints to basepair resolution. e, Two examples of mosaic complex SVs detected in the blood of healthy individuals by analyzing WGS read pairs. In the second example, the set of rearrangements and translocations include a deletion of DNMT3A, which presumably contributed to the clonal expansion.

Several newly identified deletion hotspots exhibited characteristics different from previously observed focal deletions, suggesting a distinct etiology. Whereas deletions at previously identified hotspots tended to vary in length but consistently overlap a single target gene, deletions at the new loci tended to be shorter and less focal, typically spanning 100–700 kb with breakpoints distributed throughout the locus and no consensus-deleted region (Fig. 2a,b and Supplementary Table 3). These less focal deletions tended to be less age enriched and to overlap fragile sites commonly deleted in cancers (for example, MACROD2 and RBFOX1)38 (Fig. 2b and Supplementary Table 3), suggesting that some might have originated from mutations at fragile sites early in development rather than rising in frequency due to clonal selection during aging. Among 19 autosomal fragile sites previously observed in cancer genomes38, 10 were mCA hotspots in blood. Three focal deletions exhibited Bonferroni significant sex biases (P < 0.05/41) relative to the cohort average: del(16p11.2) was more frequent in women21,22, whereas deletions at DLEU2 (13q14) and NRXN1 were more frequent in men (Fig. 2c and Supplementary Table 3). Mosaic deletions at NRXN1 observed in blood have previously been implicated in schizophrenia39.

To learn about the DNA-repair mechanisms that generate interstitial deletions in blood DNA, we analyzed read-level data to resolve breakpoints of a subset of mCAs. Discordant read pairs localized the breakpoints of 3,500 interstitial deletions (28% of deletions, including 48% of those with cell fraction >10%; Supplementary Figs. 30 and 31) to a resolution of ~1 kb, and chimeric (‘split’) reads fully resolved 2,274 deletions at basepair resolution.

Most fully resolved deletions (88%) had little or no microhomology at their breakpoints (at most 3 bp, with 0 bp being most common; Fig. 2d), suggesting that nonhomologous end-joining is the double-strand break-repair mechanism40 responsible for most mosaic deletions in blood. Other deletions exhibited intermediate amounts of microhomology (usually up to 8 bp; Fig. 2d), indicating that microhomology-mediated repair mechanisms also contribute to mCA formation.

WGS data also provided the opportunity to identify several complex structural variants (SVs) in the blood of cancer-free individuals: 159 individuals had discordant read pairs linking mCAs called on separate chromosomes to the same mutational event, and we fully reconstructed complex mosaic SVs for 22 individuals, revealing translocations, inversions and complex rearrangements (Fig. 2e).

Clonal expansions of CLL-associated 13q14 deletions commonly arise during aging

Our new mCA call set contained thousands of chromosomal alterations often seen in chronic lymphocytic leukemia (CLL), enabling deeper study of these mCAs and their progression to CLL. Deletion of 13q14 and trisomy 12 are the two most common mCAs in CLL (with most CLL genomes containing at least one of these two mCAs30) and our previous SNP-array analysis detected each alteration in 0.1% of UKB participants, conferring 100-fold to 200-fold increased risk of incident CLL22. Here WGS-based analysis detected twice as many mCAs of each type (Fig. 3a). However, given that many 13q14 deletions are short deletions of a focal 1-Mb region (spanning DLEU2, miR-15a, miR-16-1 and DLEU7; Fig. 1b and Supplementary Fig. 16), we reasoned that a targeted analysis of WGS read-depth in this region might uncover even more mosaic 13q14 deletions present at low cell fractions. This approach extended the lower limit of detectability to a cell fraction of 2% and increased the number of detected 13q14 deletions to 2,906—a 5-fold increase over the SNP-array analysis—with prevalence rising with age from 0.1% at age 40 years to 1.1% by age 70 years (Fig. 3a, Supplementary Figs. 32 and 33 and Supplementary Table 4).

Fig. 3. Causes and consequences of mosaic 13q14 deletion from 2,906 cases in UKB.

Fig. 3

a, Prevalence of mosaic 13q14 deletions detected in blood samples from UKB participants by analyzing SNP-array allelic imbalance data, WGS allelic imbalance data or WGS allelic imbalance and WGS read-depth data (union of calls from the two WGS-based approaches). b, Kaplan–Meier survival curves for rates of progression to CLL for individuals with del(13q14) clones of various sizes (shaded regions, 95% CIs). Individuals with 13q CN-LOH calls were excluded and del(13q14) clones were restricted to those with breakpoints localized by discordant read pairs (allowing more accurate estimation of the clonal fraction; Supplementary Note). Analyses were restricted to European-ancestry individuals without blood cancer diagnoses at assessment and controls were age and sex matched. c, Analogous CLL-free survival curves for individuals with 13q CN-LOH clones of various sizes. d, Comparison of GWAS on WGS-derived del(13q14) or trisomy 12 status versus CLL case status in the UKB (including up to 15 years of incident cases). Novel loci relative to ref. 45 are labeled on the Manhattan plot. P values from two-sided Firth logistic regression are as implemented in REGENIE81 (Methods); dashed lines indicate a genome-wide significance threshold (P < 5 × 10−8). e, Comparison of P values from the del(13q14) + tri(12) GWAS (x axis) and the CLL GWAS (y axis) at CLL-associated loci from ref. 45. For each index variant from ref. 45, the plotted P values indicate the most significant associations observed within 500 kb in each GWAS in the UKB. Dashed lines and shading correspond to the standard genome-wide significance threshold of 5 × 10−8. f, Hi-C contacts in naive B cells46 suggesting frequent physical proximity between the most common left and right breakpoints of 13q14 deletions. The intensity of each pixel reflects the frequency of physical contacts between genomic loci as previously measured using Hi-C. g, Pileups of 2,087 mosaic 13q14 deletions with breakpoints localized by discordant read pairs in 1,590 individuals with 13q14 deletions (Supplementary Note). Deletions are colored by estimated cell fraction. h, Breakpoint distribution of 13q14 deletions. Individuals with both 13q14 deletions and 13q CN-LOH mutations tend to have shorter 13q14 deletions, suggesting selective pressure in flanking regions that prevents longer deletions from becoming biallelic. Genes in the bottom decile of lymphoid and myeloid DepMap scores82 are annotated. cf, cell fraction.

Most 13q CN-LOH events, which are also associated with strongly increased CLL risk22, appeared to act on 13q14 deletions. Of the 149 individuals with a high cell-fraction (>5%) CN-LOH mutation overlapping 13q14, 109 had evidence of a 13q14 deletion made biallelic by the CN-LOH (Supplementary Fig. 34). However, a substantial minority (27%) had no evidence of reduced copy number at 13q14 (relative read-depth >0.99; Supplementary Fig. 34), suggesting that 13q CN-LOH mutations—nearly all of which overlap 13q14 (Fig. 1b and Supplementary Fig. 16)—may also commonly act on other genetic or epigenetic modifications that inhibit tumor-suppression function at 13q14 (ref. 41) (Supplementary Note).

To study the natural history of progression from 13q mCAs to CLL, we analyzed cancer outcomes across up to 15 years of follow-up. Deletions of 13q14 that were likely to be monoallelic (based on the absence of 13q CN-LOH) progressed to CLL at rates that increased with clone size: 10-year CLL-free survival rates ranged from 92% (95% CI 89–95%) for 13q14 deletions with cell fractions <0.05 to 72% (66–76%) for cell fractions >0.1 (Fig. 3b). The 13q CN-LOH clones progressed to CLL at similar rates (Fig. 3c) despite being likely to have inactivated both copies of the 13q14 tumor-suppressor region. This counterintuitive finding is consistent with epidemiological observations that patients with CLL and monoallelic versus biallelic loss of 13q14 have a similar prognosis42,43 and contrasts with the effects of CN-LOH mutations that amplify myeloid CHIP mutations44.

The >2-fold larger number of UKB participants with a CLL-associated mCA (3,679 individuals with 13q14 deletion or trisomy 12) compared to the number with prevalent or incident CLL diagnoses (1,502 individuals) suggested that mCAs might help discover germline variants that influence CLL risk. To evaluate the potential of this approach, we first verified that previously identified CLL risk variants45 exhibited broadly consistent associations with CLL-associated mCAs (Supplementary Fig. 35). We then performed genome-wide association analysis on CLL-associated mCA status in the UKB (that is, the presence of 13q14 deletion or trisomy 12) and on CLL case status in the UKB. The mCA association analysis was better powered, identifying 18 significant loci compared to 12, with 15 of the 18 loci falling within 500 kb of a CLL risk locus found by a previous meta-analysis of 6,200 CLL cases45 (Fig. 3d,e). This suggests that mCAs detectable from WGS of blood samples can potentially provide more information about genetic predisposition to CLL than CLL outcome data from over a decade of follow-up—although it is important to keep in mind that genome-wide association studies (GWAS) of CLL-associated mCAs may provide more information on susceptibility to monoclonal B cell lymphocytosis (a CLL precursor), and germline influences on progression of monoclonal B cell lymphocytosis to CLL could be missed.

Analyzing breakpoints of 13q14 deletions provided clues about the origins of these mutations and the varying strength of the clonal selection that they experience. Discordant read pairs localized the breakpoints of 2,087 deletions in 1,590 individuals to ~1 kb resolution. These breakpoints clustered in hotspots on the scale of 10–100 kb (Fig. 3f,g), suggesting that local genomic or epigenomic architecture might strongly influence generation of mCAs. The most common left and right breakpoints of 13q14 deletions appear to be physically proximal in B cells according to Hi-C contact maps46 (Fig. 3f), suggesting that chromatin looping in this region contributes to recurrent deletion of this 1-Mb segment. Factors influencing clone fitness also appeared to shape the distribution of observed 13q14 deletion breakpoints. Shorter deletions tended to reach high cell fractions more often than longer deletions, suggesting that deletion of larger regions of 13q14 may be detrimental to clone fitness (Fig. 3g). This effect, consistent with observations that mCA type is an incomplete determinant of clonal fitness47, remained after controlling for differential detection sensitivity for long versus short deletions by restricting to deletions with cell fractions >5% (Supplementary Fig. 36). Moreover, 13q14 deletions detected in conjunction with 13q CN-LOH tended to be shorter than 13q14 deletions without detectable CN-LOH (Fig. 3h), hinting at selective pressures against biallelic deletions of larger regions that might contain essential genes48.

CN-LOH mutations act on rare protein-altering variants in 38 genes

CN-LOH mutations do not alter gene dosage, yet are commonly observed in hematopoietic clones3,6,7 and comprised half of the mCAs that we identified (Fig. 1d). We and others have previously observed that some CN-LOH mutations appear to act on inherited or acquired protein-altering variants in several genes, conferring a proliferative advantage either by making a fitness-increasing allele homozygous or by removing a fitness-decreasing allele from a hematopoietic cell’s genome21–23,32. Here our largest-to-date call set of 21,050 CN-LOH mutations, together with WGS-derived genotype calls for rare protein-altering SNVs and indels, provided the opportunity to find many more driver genes underlying CN-LOH expansions.

Protein-altering variants in 38 genes are associated with strongly increased risk of CN-LOH mutations overlapping these genes (burden test P < 1.2 × 10−5, false recovery rate (FDR) <0.01; Fig. 4a and Supplementary Table 5). Most of these associations were newly identified; for eight genes, inherited coding variants had previously been implicated as targets of CN-LOH mutations in the UKB and BioBank Japan21–23 and, for three genes (DNMT3A, TET2 and JAK2), we previously observed evidence of CN-LOH mutations acting on acquired CHIP driver mutations22. Gene ontology enrichment analysis of the 38 genes implicated DNA double-strand break response (P = 1.0 × 10−7), apoptosis (P = 5.1 × 10−6), cytokine signaling (P = 1.3 × 10−6) and protein ubiquitination (P = 6.3 × 10−4) as cellular processes involved in clonal selection (Fig. 4b and Supplementary Table 6). Impaired DNA-damage response can contribute to proliferation by enabling further accumulation of proliferation-increasing mutations or by removing replication checkpoints that ensure genomic integrity; cytokine signaling has important regulatory roles in immune cell proliferation; and ubiquitination regulates protein turnover, such that reduced function of ubiquitin ligases targeting pro-proliferative proteins could stimulate proliferation.

Fig. 4. Rare protein-coding variants in 38 genes influence clonal expansion of CN-LOH mutations that modify their allelic dosage.

Fig. 4

a, Genes for which a burden of rare coding variants associated with the presence of an overlapping CN-LOH mutation (one-sided Fisher’s exact P < 1.2 × 10−5, FDR < 0.01). Genomic coverage by CN-LOH mutations is shown in yellow (indicating relative frequencies of CN-LOH mutations on different chromosomes). Genes with rare alleles carrying protein-altering variants that were more often made homozygous (respectively, more often removed) by CN-LOH are colored purple (respectively, green); genes for which a directional bias was unclear are colored black (Methods). b, CN-LOH-associated genes cluster in DNA-damage response, cytokine signaling, cell-cycle regulation and protein ubiquitination pathways. c, CN-LOH mutations can either amplify or remove alleles carrying protein-altering variants. For each gene, we determined whether or not CN-LOH mutations observed in the UKB more often removed LoF and high-impact (PrimateAI-3D score > 0.8) missense variants (based on a simple majority among carriers that could be phased). The proportions of genes for which coding variants were more often removed are shown in green for various categories of genes: DepMap.Q10, genes in the bottom decile of DepMap scores62,63 (n = 379); tumor-suppressor genes and oncogenes are from Memorial Sloan Kettering’s OncoKB83 (n = 84 and n = 79, respectively). The error bars are 95% Jeffreys intervals. d, Penetrance of LoF alleles (that is, fraction of carriers for whom the corresponding mosaic CN-LOH mutation was observed) for 33 CN-LOH-associated genes with at least 30 carriers (Supplementary Table 5) of rare (allele frequency (AF) < 0.001) LoF alleles (in any transcript and inclusive of LoF CNVs). Inset: penetrance as a function of age for TM2D3, MPL, ATM and all 33 genes combined. The error bars are 95% Jeffreys intervals. e, Fraction of individuals with CN-LOH detected on a given chromosome arm for whom the CN-LOH clonal expansion is potentially attributable to a rare coding variant (AF < 0.01) in one of the 38 CN-LOH-associated genes (specifically, an LoF variant or a missense variant with PrimateAI-3D score > 0.6 in any transcript). Each bar in the stacked bar plot represents 1 of the 38 genes associated with CN-LOH, sorted by the fraction of cis-CN-LOH clonal expansions potentially attributable to a rare coding variant in that gene.

The expanded list of putative targets of CN-LOH mutations also provided deeper insights into clonal selection on mutations in two previously identified target genes—TM2D3 and CTU2—which exhibit some of the strongest associations of protein-altering alleles with CN-LOH mutations in European and east Asian populations21–23. TM2D3 encodes one of three transmembrane 2 domain-containing proteins with unknown functions in human cells. Recent studies of model organisms have suggested that TM2D genes regulate Notch signaling49,50. Here frameshift mutations in the final exon of NOTCH1, which constitutively activate Notch signaling in hematopoietic malignancies51,52, and loss-of-function (LoF) variants in all three TM2D genes—TM2D1, TM2D2, and TM2D3—associated with CN-LOH mutations in cis (Fig. 4a,b and Supplementary Table 5), suggesting that these CN-LOH mutations may generate a proliferative advantage via dysregulation of NOTCH1, which is also a target of clonal expansion driver mutations (and CN-LOH) in esophageal tissues53,54. CTU2 encodes a subunit of the cytosolic thiouridylase complex involved in maintenance of genomic integrity55. Here we replicated the association of coding variation in CTU2 with CN-LOH mutations in cis23 and found a new association of coding variants in URM1 with CN-LOH mutations overlapping URM1 (Fig. 4b and Supplementary Table 5). URM1 encodes a ubiquitin-like protein that interacts with CTU2 in the thiolation of uridine in the wobble position of specific transfer RNAs56, suggesting a translation-related mechanism of clonal selection for these CN-LOH mutations.

Haplotype phasing analyses enabled determination of whether CN-LOH mutations more frequently made rare alleles homozygous or removed them from the genome (Fig. 4c). In our previous SNP-array-based analysis, CN-LOH mutations acted to make rare alleles carrying protein-altering variants homozygous for all but one target gene, MPL being the sole exception22. Here we again observed that, for most putative target genes, CN-LOH mutations appeared to provide a ‘second hit’57 by increasing the dosage of a proliferation-increasing allele (Fig. 4a–c). However, we also observed four genes for which CN-LOH mutations generated revertant mosaicism, appearing to perform ‘natural gene therapy’ on alleles that presumably reduce cell fitness, replacing alleles carrying coding variants with unmutated alleles (Fig. 4a). These genes included CFLAR, which encodes a protein (CASPER, also called c-FLIP) that inhibits CD95-mediated apoptosis58,59, and IL2RB, which encodes a component of a receptor, the activation of which increases T cell proliferation60. A fifth gene, ERG, also exhibited evidence of somatic genetic rescue in two carriers of phaseable variants, corroborating recent work61. This phenomenon appeared to extend to many other genes: among genes most essential for survival of hematopoietic cancer cell lines (lowest decile of DepMap gene dependency scores62,63), 225 experienced more CN-LOH mutations that removed damaging alleles than CN-LOH mutations that made such alleles homozygous, whereas only 154 had more CN-LOH mutations confer a second hit (two-sided binomial P = 3.1 × 10−4; Fig. 4c). CN-LOH mutations also tended to more often duplicate LoF alleles of tumor-suppressor genes and more often remove LoF alleles of oncogenes, as expected, although these differences were not statistically significant (Fig. 4c).

LoF alleles of several putative target genes of CN-LOH mutations exhibited high penetrance for these CN-LOH clonal expansions (restricted to genes for which most LoFs are germline mutations; Fig. 4d and Supplementary Fig. 37). More than half of the UKB participants who carried TM2D3 LoF alleles had detectable 15q CN-LOH mutations overlapping TM2D3 (ref. 21), and TM2D1 LoF alleles exhibited the next-highest penetrance at 27% (95% CI 17–39%) (Fig. 4d). The penetrance of LoF alleles associated with CN-LOH clonal expansions increased steadily with age (for example, from 30% (17–46%) to 69% (59–78%) for carriers of TM2D3 LoF alleles of age 40–44 years versus 65–69 years; Fig. 4d), reflecting the stochastic acquisition and slow clonal expansion of CN-LOH mutations over decades of aging47,64. Despite these high penetrances, the clonal fractions of most CN-LOH clonal expansions attributable to rare coding variants were modest (<5%; Supplementary Fig. 38).

The 38 identified target genes of CN-LOH mutations provided plausible explanations for some CN-LOH mutations observed on 22 chromosome arms (Fig. 4a,e). However, although the newly identified target genes provided insights into biological mechanisms that contribute to CH (Fig. 4b), they accounted for small fractions of the CN-LOH mutations observed on each arm—in contrast to the substantial fractions of mutations explained by target genes previously identified21,22 (Fig. 4e)—suggesting that drivers other than rare coding variants in cis contribute to the expansions of most blood clones with CN-LOH mutations.

Common variants contribute modestly to CN-LOH expansions

Common genetic variation is another source of allelic differences between homologous chromosomes on which CN-LOH mutations can act to generate a proliferative advantage. Previous studies have identified four loci at which common variants associate with CN-LOH mutations in cis22,23. Here after eliminating CN-LOH clonal expansions likely to be driven by rare coding variants, we observed just one previously reported common-variant association at the DLK1 locus23 (Fig. 5a). A genome-wide scan for common alleles associated with CN-LOH direction (that is, a tendency for CN-LOH mutations in cis to make an allele homozygous or, alternatively, a tendency to remove it from the genome) identified four additional loci (PRDM16, JAK2, ATM and SH2B3; Supplementary Table 7 and Supplementary Fig. 39). Dysregulation of PRDM16 has been implicated in myelodysplastic syndrome and acute myeloid leukemia involving t(1;3)(p36;q21) translocations65. The limited association signal from common variants contrasted strikingly with the many strong associations that we observed between CN-LOH mutations and rare coding variants in cis (Fig. 5a).

Fig. 5. Contribution of inherited common variants to CN-LOH clonal expansions.

Fig. 5

a, Miami plot comparing associations of rare coding variants (top: burden tests) and common variants (bottom) with the presence of an overlapping CN-LOH mutation. P values (y axis; log10 scale) are from one-sided Fisher’s exact test. The significance threshold plotted for the burden test is the FDR < 0.01 threshold of P < 1.2 × 10−5, whereas the significance threshold for the common-variant GWAS is the canonical threshold P < 5 × 10−8. b, Tendency of clonally expanded CN-LOH mutations in blood cells to amplify chromosome arms with proliferation-increasing PGSs for blood cell counts. The heatmap displays z-scores corresponding to the mean difference between PGSs on haplotypes made homozygous versus haplotypes removed by CN-LOH mutations, divided by the s.e.m. Positive differences (red) correspond to chromosome arms for which the retained haplotype tended to have a larger blood count PGS than the removed haplotype. *z-test FDR < 0.05; **z-test P < 0.05, Bonferroni adjusted. c, Tendency of alleles that associate with increased WBCs to be more often made homozygous by CN-LOH. The effect directions of common WBC-associated index variants were evaluated for consistency with the directions of CN-LOH mutations that overlapped them (Supplementary Note). The error bars are 95% Jeffreys intervals. d, Overall contribution of common inherited variants to determining the directionality of observed CN-LOH mutations (that is, which haplotype is amplified and which is removed) estimated using variance components analysis (Methods). SNP heritability estimates (defined in Methods) are shown for 31 chromosome arms with at least 200 CN-LOH mutations (Supplementary Table 2); the inverse-variance weighted average of these estimates is shown on the right. The error bars are 95% CIs.

Despite the paucity of common-variant associations with CN-LOH mutations, common variants could still influence cis-acting CN-LOH mutations at levels too subtle to discern at existing sample sizes (here <1,000 CN-LOH mutations on most chromosome arms). We previously observed that differences in polygenic scores (PGSs) for blood cell counts computed across alleles carried on haploid chromosome arms—a proxy for differences in proliferative potential between pairs of homologous arms—correlated with CN-LOH directionality (that is, which arm was duplicated and which arm was lost)22. This effect replicated and strengthened in significance here (Fig. 5b). Among specific variants associated with white blood counts (WBCs), CN-LOH mutations in heterozygous carriers of a variant more often duplicated the WBC-increasing allele for most variants, and the strength of this effect increased with the magnitude of variants’ effect sizes on the WBCs (Fig. 5c).

Our CN-LOH call set provided sufficient sample size to perform heritability analyses quantifying the extent to which the directions of CN-LOH mutations on a given chromosome arm are influenced by common variants spanned by these mutations. We defined a liability threshold model in which the observed direction of a CN-LOH mutation is determined by the sign of an unobserved liability composed of a genetic component (namely, the difference in total proliferative strengths of the alleles carried on the two homologous chromosome arms) and a random noninherited effect. We then used variance component analysis to estimate the fraction of variance in liability attributable to inherited common variants66–68 (Supplementary Fig. 40). Averaging across chromosome arms with at least 200 CN-LOH mutations, the average heritability attributable to common variants spanned by CN-LOH mutations was 0.08 (s.e.m. = 0.03) (Fig. 5d). This modest common-variant heritability, together with the modest fraction of CN-LOH mutations (9%) that appear to act on rare coding variants in putative target genes, suggests that, although allelic differences between the proliferative strengths of homologous chromosome arms play an important role in clonal expansions of many CN-LOH mutations, other proliferative influences must also drive expansion of clones with CN-LOH mutations. Further work will be required to determine whether CN-LOH mutations often represent passenger events on clones carrying other driver mutations, whether they act on epigenetic differences between homologous chromosomes or whether they obtain proliferative advantages for other reasons.

Discussion

In this study, we leveraged WGS of UKB participants to obtain new insights into the landscape of mCAs in blood. Improved detection of short mosaic deletions revealed new hotspots driven by genome fragility and a surprisingly high prevalence of 13q14 mutations, whereas improved detection of large mCAs present at low cell fractions powered genetic association analyses identifying numerous targets of CN-LOH mutations.

These results show that, because CN-LOH mutations commonly generate proliferative advantages by altering the dosages of inherited alleles, such mutations provide valuable clues that can point to fitness-influencing genetic variation. Consequently, increasingly large studies of CN-LOH mutations will provide increasing power to identify and fine-map further effects of inherited variants on clonal fitness, complementing analyses of somatic point mutations69.

The 38 genes that our burden tests identified as targets of CN-LOH mutations already provide intriguing insights into how these clones expand. For most of these genes, CN-LOH mutations provided a second hit, often to a gene in a proliferation-related pathway. However, we also observed revertant mosaicism of damaging alleles of several genes, and a broader analysis across DepMap-prioritized genes suggested that this phenomenon extends to dozens and perhaps hundreds of genes (based on the differential of 71 = 225 − 154 genes for which CN-LOH directionality favored reversion versus a second hit). This suggests that ‘natural gene therapy’ to remove deleterious alleles occurs much more frequently beyond previous observations in inherited bone marrow failure syndromes70, albeit to more modest extents in CH.

Our analyses of common-variant effects on CN-LOH mutations also demonstrate the promise of this approach: although we were not yet powered to identify many specific effects, we observed evidence of heritable, common-variant effects on CN-LOH directionality—suggesting the potential for future discovery from larger cohorts, similar to how earlier analyses of complex trait heritability68 preceded a wave of GWAS discoveries.

Taken together, these results underscore that CN-LOH mutations have unique fitness consequences determined more by the alleles that they carry than the chromosomes on which they fall47. Interpreting such mutations (for example, to quantify risk of progression to malignancy) will require nuanced approaches. This need for nuance extends beyond CN-LOH mutations: even for 13q14 deletions, deletion length appeared to influence clonal fitness.

The surprising frequency of 13q14 deletion (>1% of 70-year-old individuals; highest among autosomal mCAs) allowed us to perform a natural history study of its progression to CLL. Given the prevalence of this mutation, an important area for future work will be to determine whether or how cases of mosaic del(13q14) should be clinically managed71,72: as with other forms of CH, the factors that determine which individuals progress to hematological malignancy remain largely unknown. The high prevalence of 13q14 deletion also suggests the possibility that undetected 13q14 deletions and other mCAs could explain some cases of CH with unknown drivers, which is associated with the risk of lymphoid disorders15.

Our analysis of UKB WGS data had several limitations: we could not study longitudinal dynamics5 of mCAs or analyze specific cell populations of expanded clones73,74, nor could we study therapy-related CH75,76. We did not analyze mCAs of sex chromosomes, which have very different biology24,25. Available sequencing depth also limited our ability to detect short mCAs with low cell fractions. In addition, the UKB cohort had limited representation of non-European ancestries. Future WGS-based mCA analyses across multiple biobanks will provide opportunities to learn from other genetic ancestries, building on recent SNP-array-based efforts25,77,78, and analyses of mCAs in other tissues54,79,80 will deepen our knowledge of the genetic forces shaping clonal selection more broadly.

Methods

Ethics

This research complies with all relevant ethical regulations. The study protocol was determined not to be human participant research by the Broad Institute Office of Research Subject Protection and the Partners HealthCare Human Research Committee (because all data analyzed were previously collected and de-identified).

UK Biobank dataset

UKB is a prospective cohort of approximately half a million volunteer participants aged 40–70 years at recruitment between 2006 and 2010. Blood samples were acquired at initial assessment, from which aliquots were used for SNP-array genotyping, serum biochemistry assays and blood count assays, with any remaining blood preserved for future analyses. Health-related phenotypes were recorded from touchscreen questionnaires and nurse interviews and follow-up health outcome data have been accruing from linkage with UK national health registries.

WGS data were subsequently generated and made available for 490,414 UKB participants sequenced by the UKB Whole Genome Sequencing Consortium36. Blood-derived DNA was sequenced using Illumina NovaSeq 6000 machines to an average coverage of 32.5×, after which 151-bp paired-end reads were aligned to GRCh38 and SNP and indel calling was performed as previously described36. Blood samples used for WGS had been acquired at initial assessment for 99.6% of sequenced individuals (and for 99.7% of the 484,081 individuals in our primary dataset); for the remaining individuals, blood samples from a later visit to a UKB assessment center were used for sequencing. Given the small fraction of participants whose sequenced blood samples had been obtained after initial assessment, we included all individuals in our analyses (using age at initial assessment in analyses of age; averaged across the cohort, this underestimated age at blood draw by ~1 week).

UKB participants could withdraw from the study at any time and request that their data no longer be used. Consequently, some analyses reported in this manuscript have small discrepancies in participant counts, reflecting slight variability in the number of available (non-withdrawn) participants at the time of each analysis.

Overview of approach for detecting mCAs from WGS data

An effective approach for detecting mCAs from bulk genotyping or sequencing data is to search for long (megabase-scale) stretches of genome in which the allelic fractions of an individual’s two haplotypes exhibit imbalance21,84. This approach can be applied to either SNP-array genotyping intensity data or WGS data by appropriately modeling allelic imbalance information derived from either of these sources. Most large-scale studies of mCAs performed in the last few years have analyzed SNP-array data using the MoChA software package21,22 to model B allele frequency measurements computed from signal intensities of allele-specific genotyping probes. Recently, MoChA has also been used to perform WGS-based mCA calling by analyzing imbalances in ‘allelic depths’ computed from counts of haplotype-informative sequencing reads that span a heterozygous site29.

Compared to SNP-array data, high-coverage (~30×) WGS data can in theory provide more information about allelic imbalance—and enable higher-resolution mCA calls—because WGS data contain information about allelic balance at most heterozygous sites in a human genome (typically 2–3 million) rather than only those represented on an SNP array (typically 100,000–200,000 heterozygous sites per genome). This difference in the number of informative measurements is only partially offset by SNP-array intensity data tending to provide more information per heterozygous site (that is, lower s.d. (B allele frequency), typically 0.03–0.07) than WGS allelic depths.

A recent analysis of the TOPMed WGS dataset using MoChA demonstrated promising results of WGS-based mCA detection29: autosomal mCAs were detected at rates similar to the rates we previously observed in analyses of UKB SNP-array data (for individuals in the corresponding age ranges)22. This similar detection sensitivity presumably reflects two orthogonal considerations that counterbalanced each other: on the one hand, the WGS data in TOPMed provided more information about allelic imbalance than the SNP-array data in the UKB, but, on the other hand, the phased haplotypes in TOPMed (a smaller and more diverse cohort) were less accurate than the phasing in UKB.

Here we sought to further enhance mCA detection in the UKB 500,000 WGS cohort by performing an optimized analysis of this very large WGS dataset that addressed a few key limitations of MoChA’s WGS-based mCA analysis pipeline. First, we performed genome-wide, copy-number profiling on each WGS sample to (1) flag and filter individual specific genomic regions potentially containing inherited CNVs and (2) accurately calibrate WGS read-depth measurements. Germline SVs (particularly inherited duplications) are a major source of false positives in mCA analysis, such that stringent post-processing of mCA calls from existing pipelines is needed to filter false-positive mosaic calls29. In addition, technical noise in WGS read-depth measurement limits the extent to which copy-number states of mCAs can be confidently determined.

Second, we recomputed allelic depths at heterozygous sites in a way that avoids double-counting allelic observations derived from the same DNA fragment. As paired-end sequencing reads can overlap multiple nearby heterozygous sites, standard allelic depth measurements (which count the total number of sequencing reads supporting each allele) at nearby heterozygous sites do not provide independent measurements of allelic imbalance (Supplementary Fig. 1). MoChA circumvents this issue by imposing a minimum genomic distance between heterozygous sites considered in allelic imbalance analysis, dropping allelic depth measurements from sites that are deemed to be too near other sites. This approach loses some information, so here we instead ‘de-duped’ allelic depths by examining the read pair from which each observed allele is derived. This allowed us to generate allelic depths derived from nonredundant observations that we could subsequently model as independent. These analyses are described in detail in Supplementary Note and summarized in Supplementary Fig. 3.

Defining focal mCA regions

We defined focal regions by scanning the genome in 1-Mb windows and counting (separately) the numbers of short (<2 Mb) mosaic deletions and duplications that either partially or fully overlapped each 1-Mb window. We considered windows with >10 mCAs to be focal regions and merged adjacent windows to obtain our final list of 53 regions. For each region, we calculated a focal index as the maximum number of individuals with an mCA overlapping a 1-kb bin within the focal region (maximizing across all possible choices of this 1-kb bin) divided by the total number of individuals with an mCA overlapping any part of the focal region. We annotated focal regions as fragile sites in cancer based on the 19 autosomal fragile loci described in the somatic SV analysis of the Pan-Cancer Analysis of Whole Genomes Consortium38.

Calling mosaic 13q14 deletions from WGS read-depth

To increase power to detect del(13q14) mCAs, which are relatively common and usually span a specific ~1-Mb commonly deleted region on chromosome 13q starting at the DLEU2 long noncoding RNA, we implemented a separate detection method based on WGS read-depth in this region (chr. 13:49982549–50982549, GRCh38 coordinates). For each individual, we computed a 13q14 read-depth z-score by comparing the observed number of WGS reads aligned in this region (after applying the filters used in our genome-wide read-depth profiling analysis) to the expected number of aligned reads (according to the WGS sample’s genome-wide read-depth profile, assuming no mCA in this region). We computed a provisional z-score for the observed 13q14 read count assuming a Poisson distribution with a mean equal to the expected number of aligned reads (which we then approximated with a Gaussian). We then recalibrated the provisional z-scores to account for overdispersion from the Poisson distribution due to unmodeled technical influences on WGS read-depth. We performed recalibration based on the distribution of provisional z-scores observed among a subset of individuals expected to be minimally affected by del(13q14): specifically, individuals younger than 45 years with zprovisional < 4. We computed the mean and s.d. of the provisional z-scores observed among these individuals and then recalibrated all provisional z-scores by subtracting this mean and dividing by this s.d. We then called mosaic 13q14 deletions by setting a recalibrated z-score threshold of −3.5, such that individuals with recalibrated z < −3.5 were considered to have evidence of del(13q14). Among these individuals, we analyzed discordant read pairs to localize del(13q14) breakpoints for a subset of individuals (Supplementary Note).

We selected the threshold of z < −3.5 based on the following calculations indicating that it provides good control of FDR. First, assuming the recalibrated z-scores have a standard normal distribution under the null hypothesis (no 13q14 deletion), we determined that the expected false-positive rate (Φ(−3.5) = 0.00023) was 4.1% of the detection rate, suggesting an FDR of 4.1%. We corroborated this estimate with an orthogonal approach that utilized the principle that clonal events associate with age. We took the mean age of 61.5 years in individuals with z < −5 as the mean age for individuals with true mosaic 13q14 deletions and the mean age of 56.2 years in individuals with z > 0 as the mean age for individuals without mosaic events. Assuming that each z-score bin contains a mixture of true positive calls with a mean age 61.5 years and false-positive calls with a mean age 56.2 years, we then estimated the FDR in our 13q14 calls as the proportion α in the equation: (Mean age of individuals with calls) = (1 − α) × 61.5 + α × 56.2. Using this method, we estimated an FDR of 7.4%.

CLL survival analysis

We defined a CLL phenotype for the survival analysis based on histology and behavior codes (histology = 9823 and behavior = 3) reported in the National Health Service (NHS) cancer registry for participants residing in England or Wales and the NHS Central Register (NHSCR) for participants residing in Scotland (1,010 cases in total). We removed 2,232 individuals with a prevalent blood cancer diagnosis (histology code ≥9590 with a diagnosis date preceding the sample collection date). Among the remaining individuals (who constituted the ‘at-risk’ set at the start of the study), we restricted analysis to individuals with European genetic ancestry. We stratified those individuals with a del(13q14) and/or 13q CN-LOH mCA by mCA type and cell fraction and, among individuals with neither mCA, we identified an age-matched and sex-matched control subset by propensity score matching in a logistic regression of del(13q14) status on age and sex. We right-censored individuals with death dates reported in the NHS or NHSCR death registries. We also right-censored all cancer outcome data collected after 1 January 2018 as suggested by the UKB and restricted the survival analysis to 10 years of follow-up. We used the lifelines package85 to compute Kaplan–Meier curves and CIs.

In our analysis of del(13q14) progression to CLL, we removed individuals with 13q CN-LOH clones and restricted to del(13q14) calls with breakpoints that had been localized by discordant read pairs, because this allowed more accurate estimation of the mCA cell fraction (by using read-depth in the entire deleted region rather than just the most focal 1-Mb region). We chose the highest cell fraction clone for individuals with multiple del(13q14) or 13q CN-LOH clones. In our analysis of individuals with 13q CN-LOH clones, we included all such individuals regardless of del(13q14) status.

GWAS of del(13q14) and CLL

We performed genome-wide association analysis on both our WGS-derived del(13q14) + trisomy 12 phenotype and CLL diagnosis status of UKB participants, restricting to individuals with European genetic ancestry (as previously defined using principal components (PCs)86), TOPMed-imputed genotypes and <10% missingness of SNP-array genotypes. For the del(13q14) + trisomy 12 GWAS, we further restricted to individuals with available WGS data, comprising 3,470 cases (the union of 2,727 cases of del(13q14) based on read-depth z < −3.5 and 889 cases of mosaic trisomy 12) and 445,815 controls. To maximize power for the CLL GWAS, we merged all reports of the C91.1 code from the cancer registry (1,007 cases), hospital episode statistics (1,237 cases) and death registry (251 cases), comprising a total of 1,502 CLL cases. Restricting to individuals with European genetic ancestry, TOPMed-imputed genotypes and <10% SNP-array missingness left 1,393 CLL cases and 451,346 controls. We computed association tests using REGENIE’s approximate Firth logistic regression81 controlling for sex, age, age squared, smoking status and the first 20 genetic PCs as covariates. We fit REGENIE step 1 on the SNP-array genotypes and tested TOPMed-imputed variants with minor allele frequency >0.001 in step 2. To compare our GWAS results with results from meta-analysis of larger CLL cohorts, we downloaded the significant CLL GWAS hits in ref. 45 from the GWAS Catalog87.

Burden tests for association of rare coding variants with CN-LOH in cis

To define sets of protein-coding variants within genes to use for burden testing, we used variant annotations from the gnomAD v4.1.0 exomes release88, which included 416,555 exome-sequenced UKB participants. We extracted annotations for all protein-coding variants with high or moderate Variant Effect Predictor (VEP) consequence89, FILTER = PASS and a >90% call rate in UKB WES, from which we identified missense variants and high-confidence LoF variants with no LoF_filter or LoF_flags; we further annotated missense variants with PrimateAI-3D scores90. We extracted genotypes of these variants from the UKB population-level DRAGEN WGS bgen files and dropped variants with allele frequencies discordant with frequencies reported for non-UKB, non-Finnish Europeans in the gnomAD v4.1.0 dataset (AF disagreement >10-fold). Finally, we merged these genotypes with LoF CNV calls that we had previously generated from UKB WES data91.

We defined up to 64 burden masks per gene based on the following possible options for variant inclusion:

  • Maximum allele frequency of 0.01, 0.001, 0.0001 or singleton (4 possible choices)

  • Inclusion of LoF variants only or inclusion of missense variants with PrimateAI-3D score >0.6, >0.7 or >0.8 (4 possible choices)

  • Inclusion versus exclusion of LoF CNV calls (two possible choices)

  • Inclusion of only variants with coding consequences for the MANE_SELECT transcript92 versus inclusion of protein-coding variants for any transcript (two possible choices).

In total, we constructed 1,112,952 burden masks for 17,893 genes using REGENIE step 2 with flags --skip-test and --write-mask, which coded individuals with a genotype of 1 if they carried at least 1 variant in the mask and 0 otherwise.

We tested each burden mask for association with CN-LOH mutations in cis using Fisher’s exact test on the binary phenotype of whether or not an individual carried a CN-LOH mutation overlapping the gene being tested. As in our previous SNP-array-based analysis22, we removed related individuals (in ukb_rel.dat) by preferentially keeping individuals with CN-LOH calls and older controls. We applied the Benjamini–Hochberg procedure to control the FDR at 0.01, which corresponded to P < 1.2 × 10−5.

We proceeded to phase the burden mask genotypes onto our common-variant scaffold haplotypes to determine whether CN-LOH mutations that associated with rare coding variants in certain genes tended to make these alleles homozygous or remove them from the genome. We performed both statistical phasing using SHAPEIT5 phase_rare (filtering to confident phase calls, that is, phase probability (PP) >0.8) as well as read-based phasing using a customized pipeline to identify reads or read pairs that overlapped both a rare coding variant and a nearby common variant, thereby enabling the rare coding variant to be phased on to the common-variant scaffold. Specifically, we first used samtools mpileup (-Q 20) to find phase-informative reads (or read pairs) within 1 kb of a target variant to be phased. For each nearby heterozygous variant, we counted the number of fragments (reads or read pairs) supporting the variant with the same phase versus the opposite phase relative to the target variant. We restricted to nearby variants for which phase-informative fragments unanimously supported either the same or the opposite phasing of reference (REF) alleles. We then merged the resulting phase sets into our common-variant scaffold to determine the phase relative to CN-LOH, dropping phase sets containing variants with a relative phase that disagreed with the scaffold phase.

For each of the 38 genes with FDR-significant burden associations, we evaluated the directionality of CN-LOH mutations that were confidently phased relative to rare coding variants by the above pipeline. We used a Bayesian approach to decide whether the observed fraction of CN-LOH mutations that made rare coding variants homozygous reflected a confident skew away from 0.5. Specifically, we computed the Bayesian two-sided tail probability 2×min(P(a<0.5|phasecounts),P(a>0.5|phasecounts)), where a is the fraction of CN-LOH mutations that make rare coding variants in the gene homozygous, for which we assume an uninformative β(0.5,0.5) Jeffreys prior, and posterior probabilities of a<0.5 and a>0.5 are computed given the observed counts of CN-LOH mutations with each phase relative to rare coding variants. We considered the gene to have a confident skew in CN-LOH directionality if this two-sided tail probability was <0.1.

We performed gene ontology enrichment analysis on the 38 genes identified with the burden tests, using gseapy93 to test for enrichments of these genes in the GO_Biological_Process_2025 (ref. 94), KEGG_2021_Human95 and Reactome_Pathways_202496 gene sets.

Common-variant GWAS for CN-LOH in cis

We also tested the common variants (minor allele frequency >0.01) that we had imputed from the SHAPEIT5 200K reference panel for association with CN-LOH mutations in cis. We coded heterozygous genotypes as 1 and homozygous genotypes as 0 (under the model that a difference in proliferative potential of the two alleles should be observed in heterozygotes but unobserved in both hom-REF and hom-ALT individuals) and used Fisher’s exact test with the binary phenotype for whether or not an individual carried a CN-LOH mutation overlapping the variant being tested. We restricted analysis to unrelated individuals with European genetic ancestry, pruning related individuals as in the burden analysis. We further removed individuals with CN-LOH mutations potentially explained by rare (AF < 0.01) coding variants (including CNVs) in the 38 genes with FDR-significant burden associations (that is, we removed individuals with a CN-LOH mutation that overlapped an LoF variant or a missense variant with PrimateAI-3D score >0.6 in any transcript of one of the 38 genes); this mitigated the tendency for some common variants to tag associations driven by haplotypes containing large-effect rare variants. Finally, we also removed CN-LOH mutations that appeared likely to be driven by two rare splice region variants in MPL and ATM: chr. 1:43338725:G:C (Pangolin score of 0.68) and Chr. 11:108257471:T:G (Pangolin score of 0.82)97.

In addition to the above test for association between common variants and the presence of a CN-LOH in cis, we also tested the phase of common variants for association with the directionality of overlapping CN-LOH mutations (that is, whether or not CN-LOH mutations tended to make the alternative allele homozygous or remove it from the genome). We used a binomial test for deviation of the balance of directionality from 0.5.

Heritability of CN-LOH directionality

To estimate the extent to which common variants on each chromosome arm collectively influence the directions of CN-LOH mutations on that arm, we performed a heritability analysis under a liability threshold model. Our model assumes that each observed CN-LOH clonal expansion is driven by a differential underlying proliferative potential of a cell with the CN-LOH mutation versus other cells, yprolif, which we model as:

yprolif=(X1−X2)|CN−LOHβ+ε.

Here (X1−X2)|CN−LOH is the difference between normalized minor allele counts on the two haplotypes (coded as −12p(1−p), 0 and 12p(1−p) where p indicates allele frequency) restricted to common variants within the boundaries of the CN-LOH mutation (and equal to 0 elsewhere), β is a vector of variant effect sizes for proliferation and ε encapsulates all other effects on proliferation (for example, somatic mutations that confer proliferative potential to the cell distinct from other cells). The observed direction of the CN-LOH then depends on the sign of yprolif: when yprolif > 0, CN-LOH mutations that replace alleles on haplotype 2 with those on haplotype 1 have a proliferative advantage and expand and, when yprolif < 0, the opposite is true (under a parallel scenario in which the sign of ϵ is flipped). The common-variant heritability of the differential proliferative potential that determines CN-LOH direction is then:

hg2=Var((X1−X2)∣CN−LOHβ)Var(yprolif)

where the variances take into account randomness in the span of CN-LOH mutations that span different genomic intervals. We can estimate hg2 under the standard liability threshold model using restricted maximum likelihood (REML, as implemented in the Genome-wide Complex Trait Analysis (GCTA) software67) by randomly assigning haplotypes to X1 and X2 such that the ‘case prevalence’ is 0.5. We verified in simulations that REML produces unbiased estimates of hg2 across a wide range of parameter settings under this model of CN-LOH directionality (Supplementary Fig. 40).

We applied this model to estimate the common-variant heritability of CN-LOH directionality for 31 chromosome arms with at least 200 CN-LOH carriers (after excluding CN-LOH mutations containing rare coding variants in the 38 putative target genes). We performed linkage disequilibrium pruning on the common variants that we imputed for mCA detection, excluding variants with >5% missingness, using plink298 with a 500-kb window and r2 cutoff of 0.5. For each chromosome arm, we dropped CN-LOH events that overlapped the other arm of the chromosome (thereby dropping isodisomy events) and computed a ‘kinship matrix’ using the difference in haplotypes within CN-LOH boundaries after correcting switch errors identified during mCA calling. We normalized the kinship matrix by the mean of its diagonal to account for differences in lengths of various CN-LOH events and supplied the resulting covariance matrix to GCTA67 with the --grm flag and the -reml, --reml-no-constrain and --prevalence 0.5 flags to fit our liability threshold model with unconstrained REML. We reported the liability scale hg2 estimates from GCTA as our estimates for the heritability of CN-LOH directionality.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Online content

Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41588-026-02592-0.

Supplementary information

Supplementary Information (2.7MB, pdf)

Supplementary Note and Figs. 1–41.

Reporting Summary (73.9KB, pdf)
Peer Review File (4.2MB, pdf)
Supplementary Tables (100.3KB, xlsx)

Supplementary Tables 1–7.

Acknowledgements

We thank B. Handsaker, V. Sankaran, S. Raychaudhuri, A. Gusev, A. Sekar, G. Genovese and M. Hujoel for helpful discussions. This research was conducted using the UKB resource under application no. 40709. D.T. was supported by an NIH training grant (no. T32 HG002295). N.K. was supported by an NIH training grant (no. T32 HG002295) and fellowship (no. F31 DE034283). R.E.M. was supported by an NIH grant (no. K25 HL150334). S.R. was supported by a Swiss National Science Foundation Postdoc Mobility fellowship (no. P500PB_211106). P.-R.L. was supported by NIH grants (nos. R56 HG012698 and R01 HG013110) and a Burroughs Wellcome Fund Career Award at the Scientific Interfaces. This research was supported by the NIH Common Fund, through the Office of Strategic Coordination and/or Office of the NIH Director under award no. UM1 DA058230. The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript. Computational analyses were performed on the O2 High Performance Compute Cluster, supported by the Research Computing Group, at Harvard Medical School (http://rc.hms.harvard.edu) and the UKB Research Analysis Platform.

Author contributions

D.T. and P.-R.L. performed analyses and wrote the manuscript. N.K., R.E.M. and S.R. provided guidance on the analyses and interpretations of results.

Peer review

Peer review information

Nature Genetics thanks Paul Auer, Mitchell Machiela and George Vassiliou for their contribution to the peer review of this work. Peer reviewer reports are available.

Data availability

Anonymized mCA calls and full summary association statistics for the four GWAS in Supplementary Fig. 39 are available via Zenodo at 10.5281/zenodo.19186546 (ref. 99). WGS read-depth PCs and empirically estimated REF-bias in allelic depths at common SNPs and indels are provided at https://data.broadinstitute.org/lohlab/mCAs_WGS/. Individual-level mCA calls will be returned to the UKB for release on the UKB Research Analysis Platform when a mechanism for return becomes available. Access to the UKB (http://www.ukbiobank.ac.uk/) can be obtained by application.

Code availability

Customized code used to call mCAs and perform downstream analyses are available via GitHub at https://github.com/tangdavid/mCAs_WGS and via Zenodo at 10.5281/zenodo.19186546 (ref. 99). The following open-source software packages were also used: samtools v1.15.1 (https://github.com/samtools/samtools), bcftools v1.15.1 (https://github.com/samtools/bcftools), plink v1.9 (https://github.com/chrchang/plink-ng/tree/master/1.9), plink v2.0 (https://github.com/chrchang/plink-ng/tree/master/2.0), IMPUTE5 v2.0.0 (https://jmarchini.org/software/), XCFtools v5.0.0 (https://github.com/odelaneau/xcftools), SHAPEIT5 v5.1.1 (https://github.com/odelaneau/shapeit5), GCTA v1.94 (https://github.com/jianyangqt/gcta), REGENIE v3.1.1 (https://github.com/rgcgithub/regenie) and Python 3.8.10.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

David Tang, Email: davidtang@g.harvard.edu.

Po-Ru Loh, Email: poruloh@broadinstitute.org.

Supplementary information

The online version contains supplementary material available at 10.1038/s41588-026-02592-0.

References

  • 1.Lee-Six, H. et al. Population dynamics of normal human blood inferred from somatic mutations. Nature561, 473–478 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Osorio, F. G. et al. Somatic mutations reveal lineage relationships and age-related mutagenesis in human hematopoiesis. Cell Rep.25, 2308–2316 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Mitchell, E. et al. Clonal dynamics of haematopoiesis across the human lifespan. Nature606, 343–350 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Jaiswal, S. & Ebert, B. L. Clonal hematopoiesis in human aging and disease. Science366, eaan4673 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Fabre, M. A. et al. The longitudinal dynamics and natural history of clonal haematopoiesis. Nature606, 335–342 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Jacobs, K. B. et al. Detectable clonal mosaicism and its relationship to aging and cancer. Nat. Genet.44, 651–658 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Laurie, C. C. et al. Detectable clonal mosaicism from birth to old age and its relationship to cancer. Nat. Genet.44, 642–650 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Genovese, G. et al. Clonal hematopoiesis and blood-cancer risk inferred from blood DNA sequence. N. Engl. J. Med.371, 2477–2487 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Jaiswal, S. et al. Age-related clonal hematopoiesis associated with adverse outcomes. N. Engl. J. Med.371, 2488–2498 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Young, A. L., Challen, G. A., Birmann, B. M. & Druley, T. E. Clonal haematopoiesis harbouring AML-associated mutations is ubiquitous in healthy adults. Nat. Commun.7, 12484 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Jaiswal, S. et al. Clonal hematopoiesis and risk of atherosclerotic cardiovascular disease. N. Engl. J. Med.377, 111–121 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Zekavat, S. M. et al. Hematopoietic mosaic chromosomal alterations increase the risk for diverse types of infection. Nat. Med.27, 1012–1024 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Weeks, L. D. & Ebert, B. L. Causes and consequences of clonal hematopoiesis. Blood142, 2235–2246 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Zink, F. et al. Clonal hematopoiesis, with and without candidate driver mutations, is common in the elderly. Blood130, 742–752 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Stacey, S. N. et al. Genetics and epidemiology of mutational barcode-defined clonal hematopoiesis. Nat. Genet.55, 2149–2159 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Bick, A. G. et al. Inherited causes of clonal haematopoiesis in 97,691 whole genomes. Nature586, 763–768 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Kar, S. P. et al. Genome-wide analyses of 200,453 individuals yield new insights into the causes and consequences of clonal hematopoiesis. Nat. Genet.54, 1155–1166 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kessler, M. D. et al. Common and rare variant associations with clonal haematopoiesis phenotypes. Nature612, 301–309 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Weinstock, J. S. et al. Genetic determinants and genomic consequences of non-leukemogenic somatic point mutations. Nat. Commun.16, 9194 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Liu, J. et al. Germline genetic variation impacts clonal hematopoiesis landscape and progression to malignancy. Nat. Genet.57, 1872–1880 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Loh, P.-R. et al. Insights into clonal haematopoiesis from 8,342 mosaic chromosomal alterations. Nature559, 350–355 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Loh, P.-R., Genovese, G. & McCarroll, S. A. Monogenic and polygenic inheritance become instruments for clonal selection. Nature584, 136–141 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Terao, C. et al. Chromosomal alterations among age-related haematopoietic clones in Japan. Nature584, 130–135 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Thompson, D. J. et al. Genetic predisposition to mosaic Y chromosome loss in blood. Nature575, 652–657 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Liu, A. et al. Genetic drivers and cellular selection of female mosaic X chromosome loss. Nature631, 134–141 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Niroula, A. et al. Distinction of lymphoid and myeloid clonal hematopoiesis. Nat. Med.27, 1921–1927 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Gu, M. et al. Multiparameter prediction of myeloid neoplasia risk. Nat. Genet.55, 1523–1530 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Vattathil, S. & Scheet, P. Extensive hidden genomic mosaicism revealed in normal tissue. Am. J. Hum. Genet.98, 571–578 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Jakubek, Y. A. et al. Mosaic chromosomal alterations in blood across ancestries using whole-genome sequencing. Nat. Genet.55, 1912–1919 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Landau, D. A. et al. Mutations driving CLL and their evolution in progression and relapse. Nature526, 525–530 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Sekar, A. et al. Mosaic chromosomal alterations (mCAs) in individuals with monoclonal B-cell lymphocytosis (MBL). Blood Cancer J.14, 193 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Chase, A. et al. PRR14L mutations are associated with chromosome 22 acquired uniparental disomy, age-related clonal hematopoiesis and myeloid neoplasia. Leukemia33, 1184–1194 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.O’Keefe, C., McDevitt, M. A. & Maciejewski, J. P. Copy neutral loss of heterozygosity: a novel chromosomal lesion in myeloid malignancies. Blood115, 2731–2739 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Jakubek, Y. A. et al. Genomic and phenotypic correlates of mosaic loss of chromosome Y in blood. Am. J. Hum. Genet.112, 276–290 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Bycroft, C. et al. The UK Biobank resource with deep phenotyping and genomic data. Nature562, 203–209 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.The UK Biobank Whole Genome Sequencing Consortium. Whole-genome sequencing of 490,640 UK Biobank participants. Nature645, 692–701 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Backman, J. D. et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature599, 628–634 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Li, Y. et al. Patterns of somatic structural variation in human cancer genomes. Nature578, 112–121 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Maury, E. A. et al. Schizophrenia-associated somatic copy-number variants from 12,834 cases reveal recurrent NRXN1 and ABCB11 disruptions. Cell Genom.3, 100356 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Scully, R., Panday, A., Elango, R. & Willis, N. A. DNA double-strand break repair-pathway choice in somatic mammalian cells. Nat. Rev. Mol. Cell Biol.20, 698–714 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Puiggros, A., Blanco, G. & Espinet, B. Genetic abnormalities in chronic lymphocytic leukemia: where we are and where we go. BioMed. Res. Int.2014, 435983 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Garg, R. et al. The prognostic difference of monoallelic versus biallelic deletion of 13q in chronic lymphocytic leukemia. Cancer118, 3531–3537 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Puiggros, A. et al. Biallelic losses of 13q do not confer a poorer outcome in chronic lymphocytic leukaemia: analysis of 627 patients with isolated 13q deletion. Br. J. Haematol.163, 47–54 (2013). [DOI] [PubMed] [Google Scholar]
  • 44.Kishtagari, A. et al. Driver mutation zygosity is a critical factor in predicting clonal hematopoiesis transformation risk. Blood Cancer J.14, 6 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Law, P. J. et al. Genome-wide association analysis implicates dysregulation of immunity genes in chronic lymphocytic leukaemia. Nat. Commun.8, 14175 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Vilarrasa-Blasi, R. et al. Dynamics of genome architecture and chromatin function during human B cell differentiation and neoplastic transformation. Nat. Commun.12, 651 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Pershad, Y. et al. Determinants of mosaic chromosomal alteration fitness. Nat. Commun.15, 3800 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Pertesi, M. et al. Essential genes shape cancer genomes through linear limitation of homozygous deletions. Commun. Biol.2, 262 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Salazar, J. L. et al. TM2D genes regulate Notch signaling and neuronal function in Drosophila. PLoS Genet.17, e1009962 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Masuda, W. et al. TM2D3, a mammalian homologue of Drosophila neurogenic gene product Almondex, regulates surface presentation of Notch receptors. Sci. Rep.13, 20913 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Weng, A. P. et al. Activating mutations of NOTCH1 in human T cell acute lymphoblastic leukemia. Science306, 269–271 (2004). [DOI] [PubMed] [Google Scholar]
  • 52.Rossi, D. et al. Mutations of NOTCH1 are an independent predictor of survival in chronic lymphocytic leukemia. Blood119, 521–529 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Martincorena, I. et al. Somatic mutant clones colonize the human esophagus with age. Science362, 911–917 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Gao, T. et al. A pan-tissue survey of mosaic chromosomal alterations in 948 individuals. Nat. Genet.55, 1901–1911 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Dewez, M. et al. The conserved Wobble uridine tRNA thiolase Ctu1-Ctu2 is required to maintain genome integrity. Proc. Natl Acad. Sci. USA105, 5459–5464 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Van der Veen, A. G. et al. Role of the ubiquitin-like protein Urm1 as a noncanonical lysine-directed protein modifier. Proc. Natl Acad. Sci. USA108, 1763–1770 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Knudson, A. G. Mutation and cancer: statistical study of retinoblastoma. Proc. Natl Acad. Sci. USA68, 820–823 (1971). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Shu, H.-B., Halpin, D. R. & Goeddel, D. V. Casper is a FADD- and caspase-related inducer of apoptosis. Immunity6, 751–763 (1997). [DOI] [PubMed] [Google Scholar]
  • 59.Irmler, M. et al. Inhibition of death receptor signals by cellular FLIP. Nature388, 190–195 (1997). [DOI] [PubMed] [Google Scholar]
  • 60.Boyman, O. & Sprent, J. The role of interleukin-2 during homeostasis and activation of the immune system. Nat. Rev. Immunol.12, 180–190 (2012). [DOI] [PubMed] [Google Scholar]
  • 61.Zerella, J. R. et al. Germ line ERG haploinsufficiency defines a new syndrome with cytopenia and hematological malignancy predisposition. Blood144, 1765–1780 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.DepMap Public 25Q2 (DepMap, Broad, 2025).
  • 63.Arafeh, R., Shibue, T., Dempster, J. M., Hahn, W. C. & Vazquez, F. The present and future of the Cancer Dependency Map. Nat. Rev. Cancer25, 59–73 (2025). [DOI] [PubMed] [Google Scholar]
  • 64.Watson, C. J. & Blundell, J. R. Mutation rates and fitness consequences of mosaic chromosomal alterations in blood. Nat. Genet.55, 1677–1685 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Mochizuki, N. et al. A novel gene, MEL1, mapped to 1p36.3 is highly homologous to the MDS1/EVI1 gene and is transcriptionally activated in t(1;3)(p36;q21)-positive leukemia cells. Blood96, 3209–3214 (2000). [PubMed] [Google Scholar]
  • 66.Lee, S. H., Wray, N. R., Goddard, M. E. & Visscher, P. M. Estimating missing heritability for disease from genome-wide association studies. Am. J. Hum. Genet.88, 294–305 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Yang, J., Lee, S. H., Goddard, M. E. & Visscher, P. M. GCTA: a tool for genome-wide complex trait analysis. Am. J. Hum. Genet.88, 76–82 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Yang, J. et al. Common SNPs explain a large proportion of the heritability for human height. Nat. Genet.42, 565–569 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Bernstein, N. et al. Analysis of somatic mutations in whole blood from 200,618 individuals identifies pervasive positive selection and novel drivers of clonal hematopoiesis. Nat. Genet.56, 1147–1155 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Tsai, F. D. & Lindsley, R. C. Clonal hematopoiesis in the inherited bone marrow failure syndromes. Blood136, 1615–1622 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Steensma, D. P. & Bolton, K. L. What to tell your patient with clonal hematopoiesis and why: insights from 2 specialized clinics. Blood136, 1623–1631 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Bolton, K. L. et al. The clinical management of clonal hematopoiesis: creation of a clonal hematopoiesis clinic. Hematol. Oncol. Clin. North Am.34, 357–367 (2020). [DOI] [PubMed] [Google Scholar]
  • 73.Nam, A. S. et al. Single-cell multi-omics of human clonal hematopoiesis reveals that DNMT3A R882 mutations perturb early progenitor states through selective hypomethylation. Nat. Genet.54, 1514–1526 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Weng, C. et al. Deciphering cell states and genealogies of human haematopoiesis. Nature627, 389–398 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Bolton, K. L. et al. Cancer therapy shapes the fitness landscape of clonal hematopoiesis. Nat. Genet.52, 1219–1226 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Mitchell, E. et al. The long-term effects of chemotherapy on normal blood cells. Nat. Genet.57, 1684–1694 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Zhao, K. et al. Genetic drivers and clinical consequences of mosaic chromosomal alterations in 1 million individuals. Preprint at medRxiv10.1101/2025.03.05.25323443 (2025).
  • 78.Francis, M. et al. Multi-ancestry genome-wide association meta-analysis of mosaic loss of chromosome Y in the Million Veteran Program identifies 240 novel loci. Preprint at medRxiv10.1101/2024.04.24.24306301 (2025).
  • 79.Jakubek, Y. A. et al. Large-scale analysis of acquired chromosomal alterations in non-tumor samples from patients with cancer. Nat. Biotechnol.38, 90–96 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Coorens, T. H. H. et al. The somatic mosaicism across human tissues network. Nature643, 47–59 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Mbatchou, J. et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet.53, 1097–1103 (2021). [DOI] [PubMed] [Google Scholar]
  • 82.Tsherniak, A. et al. Defining a cancer dependency map. Cell170, 564–576 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Chakravarty, D. et al. OncoKB: a precision oncology knowledge base. JCO Precis. Oncol.2017, PO.17.00011 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Vattathil, S. & Scheet, P. Haplotype-based profiling of subtle allelic imbalance with SNP arrays. Genome Res.23, 152–158 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Davidson-Pilon, C. Lifelines, survival analysis in Python. Zenodo10.5281/zenodo.14007206 (2024).
  • 86.Hujoel, M. L. A. et al. Insights into DNA repeat expansions among 900,000 biobank participants. Nature650, 920–929 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Cerezo, M. et al. The NHGRI-EBI GWAS catalog: standards for reusability, sustainability and diversity. Nucleic Acids Res.53, D998–D1005 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Chen, S. et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature625, 92–100 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.McLaren, W. et al. The Ensembl variant effect predictor. Genome Biol.17, 122 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Gao, H. et al. The landscape of tolerated genetic variation in humans and primates. Science380, eabn8153 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Hujoel, M. L. A. et al. Protein-altering variants at copy number-variable regions influence diverse human phenotypes. Nat. Genet.56, 569–578 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Morales, J. et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature604, 310–315 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Fang, Z., Liu, X. & Peltz, G. GSEApy: a comprehensive package for performing gene set enrichment analysis in Python. Bioinformatics39, btac757 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.The Gene Ontology Consortium et al. The Gene Ontology knowledgebase in 2023. Genetics224, iyad031 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Kanehisa, M., Furumichi, M., Sato, Y., Matsuura, Y. & Ishiguro-Watanabe, M. KEGG: biological systems database as a model of the real world. Nucleic Acids Res.53, D672–D677 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Milacic, M. et al. The Reactome Pathway Knowledgebase 2024. Nucleic Acids Res.52, D672–D678 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Zeng, T. & Li, Y. I. Predicting RNA splicing from DNA sequence using Pangolin. Genome Biol.23, 103 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience4, 7 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Tang, D., Kamitaki, N., Mukamel, R., Rubinacci, S. & Loh, P.-R. Code and data from ‘Patterns and drivers of 43,617 mosaic chromosomal alterations in blood’. Zenodo10.5281/ZENODO.19186546 (2026). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (2.7MB, pdf)

Supplementary Note and Figs. 1–41.

Reporting Summary (73.9KB, pdf)
Peer Review File (4.2MB, pdf)
Supplementary Tables (100.3KB, xlsx)

Supplementary Tables 1–7.

Data Availability Statement

Anonymized mCA calls and full summary association statistics for the four GWAS in Supplementary Fig. 39 are available via Zenodo at 10.5281/zenodo.19186546 (ref. 99). WGS read-depth PCs and empirically estimated REF-bias in allelic depths at common SNPs and indels are provided at https://data.broadinstitute.org/lohlab/mCAs_WGS/. Individual-level mCA calls will be returned to the UKB for release on the UKB Research Analysis Platform when a mechanism for return becomes available. Access to the UKB (http://www.ukbiobank.ac.uk/) can be obtained by application.

Customized code used to call mCAs and perform downstream analyses are available via GitHub at https://github.com/tangdavid/mCAs_WGS and via Zenodo at 10.5281/zenodo.19186546 (ref. 99). The following open-source software packages were also used: samtools v1.15.1 (https://github.com/samtools/samtools), bcftools v1.15.1 (https://github.com/samtools/bcftools), plink v1.9 (https://github.com/chrchang/plink-ng/tree/master/1.9), plink v2.0 (https://github.com/chrchang/plink-ng/tree/master/2.0), IMPUTE5 v2.0.0 (https://jmarchini.org/software/), XCFtools v5.0.0 (https://github.com/odelaneau/xcftools), SHAPEIT5 v5.1.1 (https://github.com/odelaneau/shapeit5), GCTA v1.94 (https://github.com/jianyangqt/gcta), REGENIE v3.1.1 (https://github.com/rgcgithub/regenie) and Python 3.8.10.


Articles from Nature Genetics are provided here courtesy of Nature Publishing Group

RESOURCES