Abstract
Lung cancer is a highly heterogeneous disease primarily driven by tobacco smoking. About 20% of lung cancers occur among patients who have never smoked (LCINS) with differences in patient ancestry, sex, tumor histology, and clinical features. Our understanding of chromosomal instability in lung cancer, especially LCINS, is still limited. Here, we perform a comprehensive study of 182,429 somatic structural variations (SVs) detected in 1,209 whole-genome sequenced lung cancers, of which 864 LCINS. SVs are more abundant in tumors from patients who have smoked (LCSS); however, they are more complex and play more important roles in tumorigenesis in LCINS. EGFR mutations and KRAS mutations profoundly and independently shape the SV landscape. EGFR-mutant tumors have higher SV burden and more cancer-driving SVs. In contrast, KRAS mutations are associated with lower SV burden and less driver SVs. We decompose 16 SV signatures for both complex and simple SVs that likely represent divergent molecular mechanisms. The SV breakpoints have distinct distributions across the genome depending on the signatures due to mutagenic mechanisms and positive selection. Many established cancer-driving genes are recurrently rearranged by multiple SV signatures suggesting functional convergence of these genome instability mechanisms.
Introduction
Lung cancer is the leading cause of cancer death worldwide in both men and women1. Although most lung cancer diagnoses are attributed to tobacco smoking, approximately 15% of male patients and 50% of female patients have never smoked2 (defined as having smoked fewer than 100 cigarettes in their lifetime). Large geographic variations in incidence rates of lung cancers in patients who have never smoked (LCINS) have been documented3. LCINS are biologically and clinically distinct in many aspects. For example, all major histological types of lung cancers (small cell lung cancer, squamous cell carcinoma, adenocarcinoma and large cell carcinoma) are associated with smoking; however, LCINS are predominantly adenocarcinomas4,5. Genetically, EGFR mutations are enriched in LCINS6, while KRAS and TP53 mutations are more common in tumors from patients who have smoked7,8 (LCSS). Clinically, patients who have never smoked tend to have better survival9,10 and superior response to EGFR inhibitors11,12, but less durable response to immune checkpoint blockade13,14.
Genome instability is a well-established hallmark of cancer15,16 and somatic structural variations (SVs) are abundant in adult solid tumors17. SVs refer to alterations in the DNA sequence that involve large segments of a genome. They include deletions (Del), tandem duplications (TD), inversions (Inv), translocations (Tra) and other complex events. Somatic SVs are of clinical significance in LCINS due to the high frequency of oncogenic fusions caused by SVs18–20. Our knowledge of the underlying causes of somatic SVs and their contributions to genome instability is rather limited. Different mutational processes leave distinct footprints on genomic DNA which can be mathematically decomposed into mutational signatures. Somatic SV and somatic copy number alteration (SCNA) signatures have been deconvoluted in multiple pan-cancer studies21,22 without distinguishing complex and simple SVs. Complex SVs consist of multiple SV breakpoints produced all together as one-time events, such as chromothripsis23 and chromoplexy24. They arise through distinct molecular mechanisms. Some are clustered on one or a few chromosomes whereas others are scattered. Chromosomes in micronuclei - formed through lagging chromosomes during mitosis - and chromatin bridges -formed through dicentric chromosomes - can shatter and form chromothripsis25,26. Once chromosomes are shattered, some DNA fragments can be ligated into a circular form known as circular extrachromosomal DNA (ecDNA)25, which is often highly amplified in cancer27. ecDNA can sometimes integrate into linear chromosomes and form homogeneous staining regions (HSRs)28. Furthermore, chromoplexy was proposed to form through reciprocal translocations involving multiple partners simultaneously and the breakpoints are scattered across multiple chromosomes24. We call ‘clustered complex SVs’ the complex SVs that have breakpoints enriched in certain chromosomes more than expected by chance, such as chromothripsis. The complex SVs without breakpoint enrichment, for example, chromoplexy, are considered ‘non-clustered complex SVs’.
In this study, we performed a detailed analysis on somatic SVs in 1,209 whole-genome sequenced lung cancers, including 864 LCINS29. We separately characterized clustered complex SVs, non-clustered complex SVs, and simple SVs to more effectively elucidate the molecular mechanisms underlying SV formation and to better define their respective contributions to tumorigenesis.
Results
Study cohort and somatic SVs
To comprehensively study the somatic SV landscape in lung cancers, we collected a total of 1,209 cases, including 1,056 cases from the Sherlock-Lung Project (primarily LCINS)29–31 and 153 cases from published studies (mainly LCSS)32–37 (Supplementary Table S1). Patients were recruited from 30 different locations across four continents. Among these, 811 and 359 patients had European (EUR) and East Asian (EAS) genetic ancestries, respectively (Fig. 1a). The remaining 39 patients had South American and African genetic ancestries. Patients of European descent were almost exclusively recruited from Europe, Russia, USA and Canada, while patients of East Asian descent were recruited mainly from Hong Kong, Taiwan, South Korea and Canada (Supplementary Table S1). The dataset included 864 LCINS29 and 344 LCSS31. One subject had unknown smoking status. The majority (n = 1,032) of the tumors were adenocarcinomas, followed by 71 squamous cell carcinomas, 60 carcinoid tumors, 11 adenosquamous carcinomas and 35 of other histologies (Fig. 1a). Both tumor and matched normal samples were whole genome sequenced at an average depth of 80x and 40x coverage, respectively.
Figure 1. Landscape of somatic SVs in 1,209 whole-genome sequenced lung cancers.
a, Sankey diagram illustrates the relationships among ancestry, sex, self-reported smoking status and tumor histology type in our cohort of 1,209 lung cancers. The vertical bars represent four categories (ancestry, sex, smoking and tumor type). The height of each bar is proportional to the number of patients/tumors in each category. The connecting bands represent the flow of patients/tumors from one category to another. The thickness of each link is proportional to the number of patients/tumors transitioning between the connected nodes. b, SV burden and composition across tumor types. The black dots on the top panel depict the number of somatic SVs detected in each tumor grouped by tumor type. Red lines indicate the median numbers of SVs, while blue dashed lines indicate the interquartile range (IQR) boundaries (Q1 and Q3). Two-sided Wilcoxon rank-sum test with False Discovery Rates (FDR) correction is used to compare SV burdens across tumor types (excluding “Others”), with only FDRs < 0.05 shown. The lower panel shows the composition of SVs, including clustered complex SVs, non-clustered complex SVs and simple SVs. c, SV burden in LCINS and LCSS. Box plots within violins indicate the interquartile ranges and medians. A two-sided Wilcoxon rank-sum test is used to calculate the P value. d, e and f, Examples of clustered complex SVs (d), non-clustered complex SVs (e) and simple SVs (f). Colored arcs represent SVs of different types. Red bars beneath the arcs in (d) mark the regions of clustered complex SVs. The copy number profiles are shown as black bars above the chromosome models. Centromere positions are highlighted as red bars within the grey chromosome ideograms. Sample IDs are displayed next to the corresponding SV signatures, and specific SV signatures/types are labeled for each example (e.g., 1 ecDNA/HSR, chromoplexy, deletion, etc.).
A total of 182,429 somatic SVs were detected in 1,209 tumors. The number of SVs (ranging from 0 to 1,102) varied significantly across histology with medians of 180, 128 and 115 for adenosquamous carcinomas, squamous carcinomas and adenocarcinomas, respectively (Fig. 1b). Carcinoid tumors had the fewest SVs (median of 2). In addition, driver fusions were detected in 95 tumors (7.9%) (Supplementary Table S2), highlighting the impact of SVs in lung cancer. LCSS had significantly higher SV counts than LCINS (P=1.1 × 10−8, Fig. 1c). To gain further insights into smoking status, we assessed single-base substitution signature 4 (SBS4), which is known to be mostly driven by tobacco smoking38. Most tumors from self-reported never smokers did not carry SBS4 mutations and those from smokers carried more than 100 mutations of SBS4 as expected (Extended Data Fig. S1a). However, there were 55 LCINS (6.4%) with more than 100 mutations of SBS4, whereas 39 LCSS (11.3%) lacked SBS4. If mis-annotation or inaccurate self-reporting of smoking status were the primary causes of the inconsistency between smoking status and SBS4, the mutation spectrum would be expected to be similar in tumors defined by SBS4 but different in self-reported smoking status. Intriguingly, this was not the case. The primary type of EGFR mutation in LCINS with SBS4 was L858R, which was barely observed in LCSS with SBS429 (P=3.7 × 10−4, Extended Data Fig. 1b). Therefore, it is likely that at least a subset of patients who have never smoked were not misannotated, and the SBS4 mutations may be derived from passive smoking, air pollution, or other exposures29. Furthermore, driver fusions were found in 28.2% of LCSS without SBS4, which was much higher than that in LCINS without SBS4 (P=1.3 × 10−3, 9.6%) (Extended Data Fig. 1b). LCSS without SBS4 may have originated before patients began smoking. It is also possible that these tumors originated from cells not influenced by smoking, because it has been reported that a small fraction of pathologically normal bronchial epithelial cells in individuals with smoking history have mutation burdens similar to LCINS with very few SBS4 mutations39. Hence, tumors with SBS4 status different from self-reported smoking status may represent unique subtypes. In the remainder of this manuscript, SV signatures are decomposed for all tumors, but only 1,114 tumors with consistent self-reported smoking status and SBS4 were used to study the impacts of smoking on the repertoire of SVs. Furthermore, only EUR and EAS patients are used when genetic ancestry is studied as a factor.
Complex SV events
Out of 182,429 individual somatic SVs, 91,776 (50.3%) and 25,652 (14.1%) belonged to clustered (Fig. 1d) and non-clustered (Fig. 1e) complex SVs, respectively. The remaining 65,001 (35.6%) SVs were classified as simple SVs (Fig. 1f). A total of 1,457 clustered complex SV events from 928 (76.8%) tumors and 4,816 non-clustered complex SV events from 1,053 (86.6%) tumors were detected (Fig. 2a and Supplementary Tables S3, S4). We then used Starfish40 and ClusterSV22 to identify clustered (Fig. 1d) and non-clustered (Fig. 1e) complex SV signatures, respectively, which likely reflect their molecular mechanisms of formation. The most abundant signature of clustered complex SVs was Signature 2 chromatin bridge detected in 310 (25.6%) tumors (Fig. 2a). Signature 1 ecDNA/HSR, Signature 5 large gain, Signature 4 micronuclei, Signature 6 hourglass and Signature 3 large loss were found in 278 (23.0%), 222 (18.4%), 209 (17.3%), 170 (14.1%) and 142 (11.7%) tumors, respectively (Fig. 2a). We note that “ecDNA/HSR” here refers to a signature of clustered complex SVs with both circular and linear forms. The simple form of ecDNA, one or a few DNA fragments ligated in a circle, was not captured by this signature and is analyzed in detail in a companion paper41. Regarding non-clustered complex SVs, 487 (40.3%) and 587 (48.6%) tumors carried chromoplexy and cycle of templated insertions (possibly formed through a template-switching mechanism42), respectively (Fig. 2a). Non-clustered complex SVs not classified as chromoplexy or cycle of templated insertions were assigned as “complex unclear” which may form through multiple unknown mechanisms. Both clustered and non-clustered SVs were more complex in LCINS than in LCSS (P=1.2 × 10−4 and 0.035, Fig. 2b). Among LCINS, clustered SVs were more complex in tumors from EAS patients than those from EUR patients (P=0.012, Fig. 2b).
Figure 2. Clustered and non-clustered complex SVs.
a, Overview of complex SVs in 1,209 tumors. The thin bars along the x-axis of each track represent individual tumors. For the 3rd and 4th signature tracks, the numbers in parentheses indicate the number of tumors with a unique complex SV signature of that type, while the “Multiple” groups in clustered and non-clustered complex SVs represent tumors with multiple complex SVs of different signatures. Therefore, numbers in parentheses may be smaller than those reported in the text, which include all tumors carrying that signature regardless of whether they have other complex SV signatures. b, Comparison of the numbers of breakpoints per clustered (left) and non-clustered (right) complex SV events across ancestry and smoking status, where n indicates the number of tumors in each group. Box plots within violins indicate the interquartile ranges and medians. Statistical comparisons are performed using two-sided Wilcoxon rank-sum tests with FDR correction; only FDRs < 0.05 are shown. c, Factors associated with clustered and non-clustered complex SV signatures using logistic regression. Odds ratios (ORs) are presented for significant factors in corresponding signatures. Error bars represent 95% confidence intervals. SCNA subtypes, piano (few SCNAs), mezzo-forte (enriched with arm-level amplification) and forte (dominated by whole-genome duplication), are defined according to our previous study20. d, Distributions of clustered complex SV signatures across tumors in selected groups. Each vertical column in the middle represents a tumor group defined by smoking status and mutation status of TP53, EGFR and KRAS in the bottom blocks. Each horizontal bar within a column in the middle represents one tumor. The height of each tumor may vary depending on the number of tumors (n) of the group. Tumors are color-coded based on clustered complex SV signatures with tumors exhibiting multiple SV signatures shown as horizontally segmented bars with different colors. Pairwise comparisons are performed for all clustered complex SVs combined between groups differing by only one factor using Fisher’s exact test with FDR correction; only FDRs < 0.05 are shown.
Many clinical and genetic factors differ between lung cancer types. For example, LCSS are mainly in males, whereas LCINS in females; carcinoid tumors are almost all from patients who have never smoked (Fig. 1a); KRAS and TP53 mutations are enriched in LCSS, whereas EGFR mutations in LCSS29. To better assess factors that impact complex SVs, we performed a logistic regression using the presence or absence of complex SV signatures as the dependent variable and patient ancestry, sex, smoking status, tumor type, tumor purity, whole-genome duplication (WGD), SCNA subtype20, mutations of TP53, EGFR, KRAS and gene fusion status as independent variables. No strong multicollinearity was observed in the independent variables (Extended Data Fig. 2a and 2b). Tumor purity was positively associated with all complex SV signatures (Fig. 2c). Patient ancestry and sex had no effects on any complex SV signatures (Fig. 2c). Smoking was positively associated with chromoplexy (Fig. 2c). Carcinoid tumors had relatively quiet genomes with very few complex SVs (Fig. 1b and 2c). WGD was positively associated with some clustered complex SVs and all non-clustered complex SVs (Fig. 2c). TP53 mutations were associated with most of the complex SV signatures (Fig. 2c) as expected40,43. Intriguingly, EGFR mutations were positively associated with almost all clustered complex SV signatures, whereas KRAS mutations were negatively associated with ecDNA/HSR, large gain and hourglass signatures (Fig. 2c). Similar associations were found if only analyzing 602 high-clonality tumors, taking tumor purity into account (Methods) (Extended Data Fig. 2c). When LCSS and LCINS were considered separately, mutant TP53 was positively associated with various types of complex SVs in both groups (Extended Data Fig. 3). EGFR mutations were positively associated with nearly all complex SV signatures in LCINS but not in LCSS, while KRAS mutations were negatively associated with clustered complex SVs in LCINS (Extended Data Fig. 3). We further stratified lung adenocarcinomas (LUADs) into distinct groups based on smoking status and mutation status of TP53, EGFR and KRAS (Fig. 2d). Among 217 and 242 LCINS (the left two groups in Fig. 2d) with wildtype TP53 and KRAS but different EGFR status, clustered complex SVs were significantly more abundant in EGFR-mutant tumors (88.0%) than wildtype ones (76.5%). In LCINS with wildtype TP53 and EGFR, the ones with KRAS mutation, had the lowest number of complex SVs (33.3%, Fig. 2d). Similarly, in TP53-mutant LCSS, those with also KRAS-mutations were associated with less clustered complex SVs than in those with no KRAS (72.7% and 89.8%, respectively).
In summary, LCSS had more complex SVs; however, they were more complex (more SV breakpoints per event) in LCINS. EGFR mutations were positively associated with complex SVs, and this effect was independent from other factors. The effect of KRAS mutations was the opposite of that observed for TP53 and EGFR mutations.
Simple SVs
After excluding complex SVs, 65,001 simple SVs (53.8 per tumor on average) remained in our cohort (Supplementary Table S3). We performed de novo simple SV signature decomposition using SigProfilerExtractor44, a non-negative-matrix-factorization-based de novo signature inference algorithm, rather than fitting SVs to COSMIC SV signatures. This is because COSMIC SV signatures do not separate complex and simple SVs, and four out of ten COSMIC SV signatures involve clustered SVs45. Furthermore, we classified simple SVs into 49 categories based on SV type and size to capture the effect of SV size better than in the COSMIC SV signatures46. For example, deletions were classified into 18 size categories, whereas COSMIC used 5 size categories. As a result, eight simple SV signatures were extracted using our approach (Fig. 3a). There were three deletion signatures (namely “Del1”, “Del2” and “Del3”), while COSMIC can only capture two (SV5 and SV7). Moreover, we found two tandem duplication (TD) signatures (namely “TD1” and “TD2”), as well as three other signatures including foldback inversions (“Fb inv”), intra-chromosomal translocations (“Intra tra”) and inter-chromosomal translocations (“Inter tra”) (Fig. 3a). This suggested that there were likely eight distinct mutational processes generating simple SVs in lung cancer. Our TD1, TD2 and Fb inv signatures corresponded to COSMIC SV3, SV1 and SV8 signatures, respectively. Our Intra tra and Inter tra signatures together resembled COSMIC SV2 signature. In our cohort, the most common simple SV signature was Del2 (16.6%), followed by Intra tra (16.0%) and Del3 (14.3%). The least abundant simple SV signature was TD1 (8.4%).
Figure 3. Simple SV signatures.
a, Eight simple SV signatures detected in 1,209 tumors. The 49 SV categories used for signature extraction are shown on the x-axis. The eight simple SV signatures are shown on the left side of the y-axis. The height of each bar represents the proportion of SVs assigned to the corresponding signature. b, Five clusters of tumors based on the simple SV signatures (excluding tumors lacking simple SVs). Vertical lines represent individual tumors. c, Factors associated with simple SV signatures. ORs are presented for significant factors. Error bars represent 95% confidence intervals. d, Comparisons of simple SV signatures across selected groups. The violin plots illustrate the distributions of numbers of simple SVs for “Fb inv” signature on the left, “Intra tra” signature in the middle, and “Inter tra” signature on the right. Box plots within violins indicate the interquartile ranges and medians. Tumor group is defined by smoking status and mutation status of TP53, EGFR and KRAS in the bottom blocks. Statistical comparisons are performed between groups differing by only one factor using two-sided Wilcoxon rank-sum tests with FDR correction; only FDRs < 0.05 are shown.
We then performed unsupervised clustering on tumors based on the activities of simple SV signatures and found four major clusters and a minor cluster (Fig. 3b). Three of the four major clusters were enriched in TD1, TD2 and Del1,2,3 signatures respectively (Fig. 3b). The fourth major “Inv/Tra” cluster was the largest cluster composed of 590 (48.8%) tumors and was enriched in Fb inv, Intra tra and Inter tra signatures (Fig. 3b). The minor “Mixed” cluster, composed of three tumors, harbored different signatures and had no dominant signature. LCINS overall (P=2.1 × 10−18) and EGFR-mutated tumors (across all patients, P=7.4 × 10−25) were enriched in the Inv/Tra cluster (Fig. 3b). To investigate the effects of various clinical and genetic factors, we conducted linear regression analyses for the eight simple SV signatures. Tumor purity, smoking status and WGD were positively associated with all simple SV signatures; patient ancestry with Del1, Del3 and TD1; and sex with Del2 only (Fig. 3c). Carcinoid tumors had very stable genomes with fewer simple SVs (Fig. 3c). As expected, TP53 mutations were positively associated with all signatures (Fig. 3c). Intriguingly, EGFR mutations were positively associated with inversions, intra-chromosomal and inter-chromosomal translocations (Fb inv, Intra tra and Inter tra), but negatively associated with TD1 (Fig. 3c). KRAS mutations were negatively associated with all signatures (Fig. 3c). Similar associations were found when we restricted the analysis to 606 high-clonality tumors from EUR and EAS patients with consistent smoking status (Extended Data Fig. 4a). When tumors were stratified by smoking status and mutation status of TP53, EGFR and KRAS, the above relationships largely held (Extended Data Fig. 4b and Fig. 3d). Interestingly, LCINS with SBS4 mutations had fewer simple SVs as well as fewer Del1, Del2, TD1, TD2 and Fb inv than LCSS with SBS4 mutations; whereas LCSS without SBS4 mutations had more simple SVs as well as Del1, Del3 and TD1 than LCINS without SBS4 mutations (Extended Data Fig. 5).
In summary, smoking, WGD and TP53 mutations were positively associated with simple SVs. EGFR mutations were positively associated with foldback inversions, inter- and intra-chromosomal translocations. KRAS mutations were negatively associated with all simple SVs.
SV breakpoint microhomology
SVs form through erroneous DNA damage repair and aberrant DNA replication47. The sequence features at SV breakpoints can allow us to infer SV forming mechanisms. Most breakpoints had 0 to 3 bp microhomology (Extended Data Fig. 6). Although both non-homologous end-joining (NHEJ) and alternative end-joining (alt-EJ) primarily produce SV breakpoints with microhomology within this range, with tens of thousands of SVs detected in our cohort, there were clear distinctions in microhomology patterns across different SV signatures. Overall, complex SVs often had the most abundant 1 bp microhomology suggesting NHEJ, whereas most simple SVs had dominant 2 bp microhomology suggesting alt-EJ (Extended Data Fig. 6). Specifically, the breakpoints in the two translocation signatures (Intra tra and Inter tra) were most abundant in 0 and 1 bp homology suggesting that NHEJ is the primary mechanism (Extended Data Fig. 6). Two bp microhomology dominated breakpoints in deletion, tandem duplication and inversion signatures (Del2, Del3, TD1, TD2 and Fb inv) (Extended Data Fig. 6) suggesting alt-EJ is more prevalent. Del1 was the only signature with the most frequent breakpoints being blunt ends (0 bp microhomology) (Extended Data Fig. 6). Microhomology of 2 bp was also abundant in Del1. These suggested that both NHEJ and alt-EJ contribute to Del1 formation. Noticeable portions of breakpoints in ecDNA/HSR, large gain, cycle of templated insertions and TD2 signatures had insertions of more than 10 bp (Extended Data Fig. 6), suggesting possible involvement of microhomology mediated break induced repair (MMBIR). Both MMBIR and alt-EJ might be relevant in the cycle of templated insertions (Extended Data Fig. 6).
SV hotspots
The distribution of SV breakpoints across the genome is determined by a combined effect of DNA damage, repair and selection. SV breakpoint hotspots often represent regions with excessive DNA damage or SVs under positive selection. We split the genome into 1 Mb windows and calculated the percentage of tumors with SV breakpoints from individual SV signatures in each window. Different SV signatures displayed distinct breakpoint distributions (Fig. 4). Complex SVs were different from simple SVs. The hotspots in complex SVs primarily involved known oncogenes, whereas the hotspots in simple SVs were mostly fragile sites (i.e., genomic loci that are prone to DNA breaks48) as well as several oncogenes and tumor suppressors (Fig. 4). MYC and MDM2 are among the most frequently amplified oncogenes in cancer including lung cancer49 and were frequently rearranged by complex SVs including ecDNA/HSR, chromatin bridge, large gain and hourglass signatures (Fig. 4). Out of 54 EML4-ALK fusions, 32 (59.3%) were produced by complex SVs including micronuclei, hourglass, chromoplexy and cycle of templated insertions; 14 (25.9%) and 8 (14.8%) were formed through unbalanced inversions (classified as Intra tra) and reciprocal inversions (classified as Del2, Fig 3a), respectively (Supplementary Table S2). As the most commonly deleted tumor suppressor49, CDKN2A was disrupted by many complex and simple SV signatures including large loss, chromoplexy, Del2, Del3 and Intra tra (Fig. 4). These together suggest functional convergence on a handful of major cancer driver genes from distinct molecular mechanisms of SV formation. The SV breakpoint distributions of ecDNA/HSR, chromatin bridge and large gain signatures were similar (Fig. 4) since these signatures often result in DNA copy gains40. Although large loss, micronuclei, hourglass and chromoplexy are signatures that commonly lead to loss of DNA, the SV breakpoint distributions were different (Fig. 4). For example, large loss had a hotspot at MECOM; micronuclei had a hotspot at CCNE1; hourglass had a hotspot at MDM2; whereas chromoplexy had none of these hotspots (Fig. 4). Intriguingly, the three deletion signatures had different SV breakpoint distributions. Large deletions (Del3) had many hotspots at fragile sites, but small (Del1) and median deletions (Del2) had none (Fig. 4). Hotspots of small TDs (TD1) did not overlap with any known oncogenes or tumor suppressors but were all fragile sites very similar to those in Del3 (Fig. 4) suggesting possible shared SV forming mechanisms of TD1 and Del3. There were very few hotspots in inter-chromosomal translocations (Inter tra), and the two major ones were due to transductions of L1 retrotransposons including a very active germline L1 element30 on chromosome 22 in the first intron of TTC28 (Fig. 4). Interestingly, some SV hotspots were different between LCINS and LCSS. For example, MDM2 was frequently rearranged by ecDNA/HSR in LCINS, but less frequent in LCSS (Fig. 4). CDKN2A was deleted by large loss and chromoplexy exclusively in LCINS, whereas STK11 was frequently deleted by Del2 only in LCSS (Fig. 4).
Figure 4. Genome-wide SV breakpoint distribution for each SV signature.
Chromosomal structures are represented by grey bars at the bottom with red lines marking centromeres. Each vertical bar represents the percentage of tumors carrying SV breakpoints for a specific SV signature in the corresponding 1 Mb interval. Known oncogenes, tumor suppressors, fragile sites and L1 retrotransposition events are annotated in major peaks.
The breakpoint distribution across SV signatures could facilitate functional inferences of hotspot genes. PTPRD has been reported as a tumor suppressor50. However, it is a very large gene (2.3 Mb) and has been implicated as a fragile site similar to other large genes, such as MACROD2, FHIT and WWOX22,51. In our cohort, fragile site hotspots were shared between Del3 and TD1 but not present in any other signatures; PTPRD was not a hotspot in TD1 of which all hotspots were fragile sites; PTPRD was frequently rearranged in several SV signatures not involving fragile sites, such as chromoplexy, Del1 and Del2 (Fig. 4). These results combined indicate that PTPRD is likely a tumor suppressor instead of a fragile site. Interestingly, in some signatures, such as chromoplexy, Del2 and Del3, PTPRD was altered as frequently as CDKN2A (Fig. 4). However, it was the only frequently rearranged locus in Del1 in LCINS, whereas CDKN2A was the only major hotspot in intra-chromosomal translocations (Intra tra) in LCINS without frequent SVs at the PTPRD locus (Fig. 4).
In summary, SV breakpoints were not evenly distributed across the genome and differed across signatures. Distinct SV signatures shared several SV hotspots around well-known cancer driver genes, suggesting functional convergence. There were also noticeable differences between tumors from patients with different smoking statuses.
Cancer-driving events
To further study SVs under positive selection, we comprehensively identified a total of 1,847 focal copy number changes of known cancer-driving genes as well as 2,284 driver mutations/indels in known driver genes in 1,114 tumors with consistent smoking status (Fig. 5a and Supplementary Table 5). Chromosome-arm level copy number changes were excluded so that the focal copy number changes of known driver genes should all result from SVs and represent cancer-driving SVs. EGFR mutations were enriched in LCINS and KRAS and TP53 mutations in LCSS (Fig. 5a). Most driver genes were mutated more frequently in LCSS than LCINS (Fig. 5a). Some focal copy number changes, such as CDKN2A losses, MDM2 gains and EGFR gains, were enriched in LCINS, whereas others, such as STK11 losses, NKX2–1 gains, MCL1 gains and FGFR1 gains, were enriched in LCSS (Fig. 5a). Furthermore, among LCINS, focal gains of TERT, PTK6, EGFR, CTNND2 and BRAF as well as mutations in TP53 and RBM10 primarily occurred in tumors with EGFR mutations (Extended Data Fig. 7). Since EGFR mutation is the earliest genetic alteration in LCINS52 (ref companion manuscription on lung cancer evolution), these co-occurring genetic alterations suggest they may act as secondary oncogenic drivers specifically required for the oncogenic transformation of cells with EGFR mutations. We then performed linear regression analyses to model the number of driver events using patient ancestry, sex, smoking status, tumor histology, tumor purity, WGD status, SCNA subtype, mutation status of TP53, EGFR and KRAS as well as fusion status as independent variables. LCSS had more driver mutations (Fig. 5b). Although smoking was positively associated with SV burden, especially simple SVs (Fig. 3c), it was not associated with the number of driver SVs (Fig. 5b). Carcinoid tumors had fewer driver mutations and driver SVs (Fig. 5b). Interestingly, TP53, EGFR and KRAS mutations were all positively associated with overall driver mutations (Fig. 5b). In contrast, EGFR-mutant tumors had more driver SVs, whereas KRAS-mutant tumors had less driver SVs (Fig. 5b). Overall LCSS carried a higher number of driver mutations, a median of three, compared to one to two in LCINS (Fig. 5c left panel). Among LCINS, EGFR- and KRAS-mutant tumors had medians of two driver mutations, while the remaining tumors, including fusion-driven ones, had medians of one (Fig. 5c left panel). In contrast, when considering driver SVs, EGFR-mutant LCINS had a median of two such SVs, whereas all other groups had only one (Fig. 5c right panel). Among LCSS, KRAS-mutant tumors had the fewest driver SVs (median of one, Fig. 5c right panel). These results combined suggest that LCSS are mainly driven by mutations, whereas SVs play more important roles in tumorigenesis in EGFR-mutant LCINS. We further investigated the SV signatures in driver SVs. In LCINS, SV drivers were more likely to be complex events (Fig. 5d top left panel), such as ecDNA/HSR and chromatin bridge (Fig. 5d top right panel). In contrast, in LCSS, driver SVs were enriched in simple SVs (Fig. 5d bottom left panel), such as tandem duplications, inversions and translocations (Fig. 5d bottom right panel). Even though STK11 focal losses preferentially occurred in LCSS (Fig. 5a), STK11 in LCINS were more likely to be disrupted by complex SVs (Fig. 5d top left panel). In LCINS, no significant difference in SV signatures was observed between EGFR-mutant and wildtype tumors (Extended Data Fig. 8). These results collectively suggested that complex SVs play more important roles in tumorigenesis in LCINS.
Figure 5. Cancer-driving mutations/indels and SVs.
a, Proportions of tumors harboring mutations/indels and SVs in known cancer driver genes in LCINS and LCSS. “*”s that are in or above the green bars and in or below the pink bars represent alterations significantly enriched in LCINS and LCSS, respectively. Enrichment is assessed using Fisher’s exact test with FDR correction. b, Factors associated with the number of driver genes altered by mutations/indels (left) and by SVs (right). ORs are presented for significant factors. Error bars represent 95% confidence intervals. c, Numbers of driver genes altered by mutations/indels (left) and by SVs (right) across selected groups. Box plots within violins indicate the interquartile ranges and medians. Statistical comparisons are performed using two-sided Wilcoxon rank-sum tests with FDR correction; only FDRs < 0.05 are shown. d, SV signature composition of focal gains and losses in known cancer driver genes in LCINS (top) and LCSS (bottom). Colored bars represent SV types (left) and signatures (right). Enrichment of SV types or signatures comparing LCINS and LCSS is tested using Fisher’s exact test with FDR correction. Significantly enriched SV types and signatures are marked with black outlines.
As one of the commonly mutated genes in lung cancer, EGFR was mutated in 48.7% of the 809 LCINS. Among them, 86 (10.6%) had EGFR focal gains (driver SVs), which, in turn, were enriched with mutations/indels (80.2%, P=4.3 × 10−10, Fig. 6a). In contrast, 5.9% of LCSS had EGFR mutations/indels and a comparable fraction (6.3%, 1 out of 16) of tumors with focal gain of EGFR had mutations/indels (P=1, Fig. 6a). EGFR was frequently rearranged by several SV signatures, especially complex SVs (Fig. 5d), including ecDNA/HSR, chromatin bridge, large gain, cycle of templated insertions and Fb inv (Fig. 4 and Fig. 6b). In addition, in tumors with EGFR mutations/indels and focal gains, the mutant allele was always amplified according to tumor purity and mutant allele fraction (Fig. 6b). Furthermore, NKX2–1, encoding a transcription factor to regulate cellular identity53,54 and recently recognized as an oncogene in LUAD55, was also frequently rearranged by multiple SV signatures such as ecDNA/HSR, chromatin bridge, large gain, cycle of templated insertions, TD2 and Fb inv (Extended Data Fig. 9).
Figure 6. EGFR-associated SVs in LCINS and LCSS.
a, Oncoprint plot showing EGFR focal gains and mutations/indels. b, Examples of SV signatures driving EGFR focal gains. Colored arcs represent SVs of different types. Red bars beneath the arcs mark the regions of clustered complex SVs. The copy number profiles are shown as black bars above the chromosome models. Centromere positions and EGFR genes are highlighted as red bars and black boxes within the grey chromosome ideograms. Sample IDs are displayed next to the corresponding SV signatures. EGFR copy number (CN), tumor purity, and mutant allele fraction (MAF) are listed for the cases with EGFR point mutations/indels.
In light of the findings within this study, we propose an updated oncogenic model for lung cancer (Extended Data Fig. 10). Most lung cancers carry two to five driver events (Fig. 5c) including point mutations, indels, fusions and DNA copy number changes. In LCSS epithelial cells, KRAS mutation is the tumor-initiating event caused by tobacco smoking. Secondary oncogenic alterations are primarily point mutations. One or two SVs per tumor serve as secondary drivers and are often simple SVs. In contrast, in LCINS cells, EGFR mutation arises through aging or endogenous processes as the first hit. One or two SVs, more frequently complex ones, are secondary drivers. In both LCSS and LCINS, malignant transformation occurs via specific combinations of genes being altered depending on the tumor-initiating events.
Discussion
We performed a comprehensive study on somatic SVs in 1,209 lung cancers. We decomposed 16 signatures of complex and simple SVs. The distribution and microhomology of breakpoints revealed interesting insights into SV formation. Del3 and TD1 were associated with DNA breakage at fragile sites (Fig. 4). Complex SVs were more likely to form via NHEJ, whereas simple SVs were likely generated through alt-EJ (Extended Data Fig. 6).
We identified important factors associated with the SV panorama in lung cancer (Fig. 2c and Fig. 3c). They included smoking, WGD, and mutations in TP53, EGFR and KRAS. Patient ancestry and sex had little effect on genome instability. Intriguingly, although the kinase domain mutations in both EGFR and KRAS are well-known primary drivers of lung cancer, we found the opposite associations of these mutations with somatic SVs. Mutant EGFR was positively associated with complex SVs (Fig. 2c), inversions, intra- and inter-translocations (Fig. 3c), whereas mutant KRAS was not enriched with complex SVs, and was negatively associated with simple SVs (Fig. 3c). In a recent study, mutant KRAS was associated with more stable genomes in colorectal cancer56, which was consistent with our results in lung cancer. It is unlikely that the point mutations in EGFR and KRAS are the result of SVs, because EGFR and KRAS mutations are the earliest genetic alterations in lung cancer (ref companion manuscript on lung cancer evolution). Instead, it is more likely that EGFR and KRAS mutations shape the SV landscape, possibly involving different tumor cells of origin. The effects of these driver mutations can be through SV formation or selection. Indeed, EGFR-mutant tumors had more driver SVs, whereas KRAS-mutant tumors had less (Fig 5b). However, positive selection cannot entirely explain the SV patterns in EGFR-mutant tumors. Many known oncogenes, such as MYC, MDM2 and NKX2–1, were amplified by various SV signatures which were positively associated with EGFR mutations, including ecDNA/HSR, chromatin bridge and large gain. MDM2 and NKX2–1 were also amplified by TD2 signature (Fig. 4). Circular ecDNA was studied in details in our companion paper41. If copy number gains of these genes are under positive selection in EGFR-mutant cells and SV burden is driven by positive selection, we would expect EGFR mutations to be positively associated with TD2 as well. However, EGFR mutations were not associated with TD2 in the full cohort (Fig. 3c) and negatively associated with TD2 in the high-clonality samples (Extended Data Fig. 4a). Therefore, the mechanisms of EGFR mutations shaping SV landscape may be due to a combined effect of EGFR mutations affecting SV formation and SVs affecting positive selection of altered cancer-driving genes. The detailed molecular mechanisms still require further experimental study.
Through detailed analysis of driver SVs and driver mutations, we revealed distinct patterns of driver events in LCSS and LCINS. The key findings were that LCSS had more driver mutations and LCINS had more SV drivers, which were more likely to be complex events (Fig. 5c and 5d). There can be multiple possible reasons for these interesting patterns, including the positive selection of secondary drivers depending on the initial driver events or secondary drivers adapting to different microenvironments and immune modulation, which again warrant further study.
Methods
Ethics declarations
Since the National Cancer Institute only received de-identified samples and data from collaborating centers, had no direct contact or interaction with the study participants, and did not use or generate identifiable private information, Sherlock-Lung has been determined to constitute “Not Human Subject Research (NHSR)” based on the federal Common Rule (45 CFR 46; https://www.ecfr.gov/cgi-bin/ECFR?page=browse).
Sample collection and processing
This study comprised 1,209 patients with histologically confirmed lung cancer primarily from Sherlock-Lung Project sourced from multiple institutions29,30,34 as well as publicly available datasets we previously collected31–33,35–37. Fresh-frozen tumor tissues were obtained alongside matched whole-blood samples or, when feasible, normal lung tissues sampled at least 3 cm from the tumor sites. All samples originated from treatment-naïve patients. Stringent inclusion criteria were uniformly applied20, ensuring minimum sequencing coverage of >40x for tumors and >25x for normal samples, limiting cross-sample contamination to <1% (evaluated by Conpair57), and maintaining sample relatedness below 0.2 (assessed by Somalier58). Samples exhibiting abnormal copy number profiles in matched normal tissues (determined by Battenberg59), samples with SBS7 (ultraviolet exposure-associated) or SBS31 (platinum chemotherapy-associated) mutational signatures, and those with fewer than 100 genomic alterations or fewer than 1,000 alterations combined with the low number of reads per clonal copy (NRPCC <10)60 were excluded. To avoid redundancy, only the highest-purity sample was selected in the case of multi-region sequencing. Furthermore, high-clonality tumor samples30 (n = 677) were defined as: (1) tumor purity > 0.3 to ensure adequate tumor cell content; (2) NRPCC ≥10 to enable reliable subclone detection; and (3) availability of high-confidence copy number profiles generated by Battenberg, excluding any samples that underwent refitting, thereby reducing potential artifacts from SCNAs. Genetic ancestry was assessed using verifyBamID61 and inferred through comparison with reference populations from the 1000 Genomes Project.
All coordinates in this study were based on the hg38 genome assembly. Gene annotation data were sourced from ENSEMBL (GRCh38.p13) (https://useast.ensembl.org/index.html).
Variant calling and quality control
Genome-wide somatic alterations were identified using previously validated pipelines established in our Sherlock-Lung study20,29–31. Somatic mutations and small insertion-deletions (indels) were detected using a consensus approach involving four variant callers: Strelka (v2.9.10), MuTect, MuTect2, and TNscope (Sentieon Genomics v2023.08.03). Variants identified by at least three callers were retained. Germline contamination was excluded by filtering variants with allele frequency >0.001 in any of these germline databases (1000 Genomes Project, ExAC, Genome Aggregation Database). Functional annotations were assigned using ANNOVAR (v2020–06-08). Candidate driver genes were identified using the IntOGen pipeline19 (v2020.02.0123), incorporating seven complementary algorithms to highlight mutations under positive selective pressure. Allele-specific SCNAs were derived using the NGSpurity pipeline62, integrating Battenberg59 (v2.2.9), DPClust59,63 (v2.2.8), and ccube64 (v1.0). Joint estimation of tumor purity, ploidy, and clonal architecture was conducted20,31, followed by rigorous manual inspection and iterative refinement of purity and ploidy values.
Meerkat65 with suggested parameters and Manta66 with default parameters were used to call somatic SVs. To assess the sensitivity and specificity of SVs detection, thirty-five LUADs from the PCAWG study were used to benchmark SV calling strategies, since somatic SVs detected by PCAWG were robust with a 90% validation rate17. When comparing somatic SVs detected by PCAWG, Meerkat, and Manta, SVs were considered validated if the SV breakpoints were on the same chromosomes with the same orientations and within 50 bp. By the same token, SVs detected by Meerkat and Manta were merged if the SV breakpoints were on the same chromosomes with the same orientations and within 50 bp. Compared to individual algorithms, the union of SVs identified by the two algorithms showed the largest (2,884 somatic SVs) overlap with SVs detected by PCAWG (Extended Data Fig. 11a). In addition, there were 4,846 somatic SVs detected by Meerkat and Manta but not reported by PCAWG (Extended Data Fig. 11a). SCNA breakpoints were used to further evaluate the quality of SVs since unbalanced SVs should result in changes in DNA copy number. Three different distance cutoffs were used when comparing SV and SCNA breakpoints: 1 kb, 5 kb and 10 kb (Extended Data Fig. 11b). Among three groups of SVs, (a) only detected by Meerkat and Manta but not by PCAWG, (b) shared by Meerkat, Manta, and PCAWG and (c) only detected by PCAWG but not Meerkat or Manta, the fractions of SV breakpoints supported by SCNA breakpoints were comparable (Extended Data Fig. 11b). This suggests that the somatic SVs detected only by Meerkat and Manta but missed by PCAWG were as robust as those detected by PCAWG. Therefore, the union of somatic SVs detected by Meerkat and Manta was of high quality. Furthermore, some SV events were likely artifacts with limited support from SCNAs. To improve the reliability of SV calls in our dataset, we introduced additional filtering criteria tailored to our analysis: for Manta-specific SVs, at least 4 read pairs and 1 split read were required (Extended Data Fig. 11c). For Meerkat-specific duplications that are less than 1,000 bp, additional >10 read pairs and >10 split reads were required (Extended Data Fig. 11d). Microhomology length and insertion length around SV breakpoints were extracted from Meerkat and Manta.
SV breakpoints were annotated with protein-coding genes. SVs involving two distinct protein-coding genes with correct transcriptional orientations (5’-3’) were potential fusion genes. Fusions with 3’ partners being ALK, RET, ROS1, MET, NTRKs or with 5’ partners being FGFRs were considered driver fusions (Supplementary Table S2). To enhance reliability, only fusions with established partner genes67 were included in the analysis.
SV signatures
Six clustered complex SV signatures were identified by Starfish40 (https://github.com/yanglab-computationalgenomics/Starfish) according to SV and SCNA patterns. The clustered complex SV signature cannot be extracted for one sample that has been annotated as “Unassigned” in the rest analysis as the SCNA data were not available. Non-clustered complex SVs were identified by ClusterSV22 (https://github.com/cancerit/ClusterSV) after removing clustered complex SVs.
After removing both clustered and non-clustered complex SVs, the remaining SVs were simple SVs. They were assigned into four primary types based on breakpoint locations and orientations: deletions, tandem duplications, inversions and translocations. Deletions and tandem duplications with breakpoints located in fragile regions were designated as fragile site deletions and fragile site tandem duplications, respectively. The remaining deletions and tandem duplications were further classified into 18 subgroups based on size. For inversions and translocations, they were further divided into reciprocal inversions, fold-back inversions, unbalanced inversions, reciprocal translocations and unbalanced translocations. Unbalanced inversions and reciprocal inversions were further split into 3 and 5 subgroups based on their sizes, respectively. Meanwhile, fold-back inversions, unbalanced translocations and reciprocal translocations stood as three distinct subcategories. In total, all simple SVs were classified into 49 subcategories. SigProfilerExtractor44 (https://github.com/AlexandrovLab/SigProfilerExtractor) with default parameters was employed to extract simple SV signatures.
Simple SVs were classified into their respective simple SV signatures as follows. Del1 signature included 0 – 2.5 kb deletions; Del2 signature included 2.5 kb – 25 kb deletions and all reciprocal inversions; Del3 signature included 25 kb – 750 kb deletions and fragile site deletions; TD1 signature included 0 – 50 kb tandem duplications and fragile site tandem duplications; TD2 signature included 50 kb – 5 Mb tandem duplications; Fb inv signature included foldback inversions; Intra tra signature included > 750 kb deletions, > 5 Mb tandem duplications and all unbalanced inversions; Inter tra signature included unbalanced inversions and reciprocal translocations. Tumors were clustered by the R package ConsensusClusterPlus68 based on the abundance of simple SV signatures. The clustering was performed with the following parameters: 80% samples resampling (pItem), 100% signatures resampling (pFeature), 50 resamplings (reps), agglomerative hierarchical clustering algorithm (clusterAlg) upon 1 – Pearson correlation distances (distance).
Factors associated with SVs signatures and drivers
The list of cancer-driving genes was obtained from the Cancer Genome Atlas study37. Copy number alterations due to genome-level, chromosome level and chromosomal arm level alterations were discarded by considering the baseline copy numbers at genome, chromosome and chromosomal arm levels. The remainder of copy number alterations were focal events. Focal losses of tumor suppressors and focal gains of oncogenes were considered as SV drivers. Mutations and indels altering protein sequences were considered drivers. Logistic regression was used to identify associated factors for complex SV signatures, and linear regression was used to identify associated factors for simple SV signatures and the number of drivers. The presence and absence of complex SV signatures were used as dependent variables in logistic regression analyses. Similarly, the number of simple SVs and the number of drivers were used as dependent variables in linear regression. Tumors from African and South American patients were not analyzed because of their small sample size. All analyses were conducted in both the entire cohort and a high-clonality subset. The Fisher’s test assesses the differences in the distribution of tumors with and without clustered/non-clustered complex SVs across different groups; the two-sided Wilcoxon rank-sum test assesses the number of somatic SVs, simple SVs, breakpoints per complex event and signature-specific SVs in different groups; and the two-sided Student’s t-test evaluates the number of driver genes in different groups. Benjamini-Hochberg procedure was used for False Discovery Rate (FDR) correction. A cutoff of 0.05 was used for the significance level.
Clinical and genomic variables were encoded as numeric or binary values (e.g., smoking status, ancestry, WGD, driver mutations, fusion status), and a correlation matrix was generated for tumors from EUR and EAS patients with consistent smoking status. Pairwise Pearson correlations and Variance Inflation Factors (VIFs) were computed. Correlations > 0.8 or VIF > 5 were considered indicative of multicollinearity.
Hotspots analysis
The reference genome was partitioned into non-overlapping bins of 1 Mb each. Within each bin, the number of tumors containing SV breakpoints was tallied for each SV signature. If a tumor carried multiple SV breakpoints of the same signature within the same bin, it was counted only once.
Supplementary Material
Extended Data Figure 1. Self-reported smoking status and SBS4 mutations. a, Number of SBS4 mutations in LCINS and LCSS. Box plots within violins indicate the interquartile ranges and medians. b, Sex, ancestry, and mutation status of EGFR, KRAS, TP53, and fusion status are compared across four groups defined by self-reported smoking status and SBS4 mutations. “n” indicates the number of tumors in each group.
Extended Data Figure 2. Multicollinearity assessment and factors associated with complex SV signatures in 602 high-clonality tumors. a, Pairwise correlations of variables. Circle size represents the absolute value of correlation. Red and blue indicate positive and negative correlations, respectively. b, Variance inflation factor (VIF) values for each variable. Dashed red line indicates cutoff of 5. c, Forest plot for factors associated with complex SV signatures in 602 high-clonality tumors. Odds ratios (ORs) are presented for significant factors in the corresponding signatures. Error bars represent 95% confidence intervals. SCNA subtypes, piano (few SCNAs), mezzo-forte (enriched with arm-level amplification) and forte (dominated by whole-genome doubling), are defined according to our previous study20.
Extended Data Figure 3. Factors associated with complex SV signatures across selected subgroups. ORs are presented for significant factors in corresponding signatures. Error bars represent 95% confidence intervals.
Extended Data Figure 4. Factors associated with the simple SV signatures. a, Factors associated with simple SV signatures in 602 high-clonality tumors from EUR and EAS patients with consistent smoking status. b, Factors associated with simple SV signatures in tumors from selected subgroups. ORs are presented for significant factors in the corresponding signatures. Error bars represent 95% confidence intervals.
Extended Data Figure 5. Comparisons of SV burden across tumors stratified by self-reported smoking status and SBS4 mutations. Box plots within violins indicate the interquartile ranges and medians. Pairwise comparisons differing by only one factor are performed using Wilcoxon rank-sum tests with FDR correction. Only FDRs < 0.05 are shown.
Extended Data Figure 6. Microhomology and insertion at SV breakpoints. Bars represent the number of SV breakpoints with corresponding microhomology (right) or insertion size (left). Suggested repair mechanisms are indicated for each signature. NHEJ, nonhomologous end joining; alt-EJ, alternative end joining; MMBIR, microhomology mediated break induced repair.
Extended Data Figure 7. Driver alterations in EGFR-mutant and EGFR-wildtype LCINS. Bars represent the proportion of tumors carrying mutations/indels and SVs in known driver genes stratified by EGFR mutation status (y-axis). Enrichment is tested using Fisher’s exact test with FDR correction. Only FDRs < 0.05 are labeled by “*”s.
Extended Data Figure 8. SV signature composition of focal driver gains and losses in 809 EGFR-mutant and EGFR-wildtype LCINS. Each bar represents a driver gene, with color indicating the associated SV type (left) or SV signature (right) identified as the closest event to the focal alteration. Enrichment of SV types or signatures between LCINS and LCSS was tested using Fisher’s exact test with FDR correction. Significantly enriched SV types and signatures are marked with black outlines. No significant enrichment is detected.
Extended Data Figure 9. Examples of SVs of different signatures leading to NKX2–1 focal gains. Colored arcs represent SVs of different types. Red bars beneath the arcs mark the regions of clustered complex SVs. The copy number profiles are shown as black bars above the chromosome models. Centromere positions and NKX2–1 genes are highlighted as red bars and black boxes within the grey chromosome ideograms. Sample IDs are displayed next to the corresponding SV signatures.
Extended Data Figure 10. Proposed model of lung cancer progression in LCINS and LCSS. Schematic illustration shows how distinct tumor initiating events occur and follow by distinct secondary driver alterations, shaping downstream pathways and contributing to lung cancer development in LCINS and LCSS.
Extended Data Figure 11. Assessment of SV calling. a, SVs detected by Meerkat and Manta in 35 lung adenocarcinoma (LUAD) samples from the PCAWG consortium compared to SVs detected by the PCAWG consortium. The three bar plots represent SVs called by Manta (left), Meerkat (middle), and the union of both algorithms (right). The bar colors correspond to the Venn diagram on the right, which illustrates the relationships between SVs of the testing sets (Meerkat, Manta, or their union) and PCAWG SVs. b, Bar plots showing the percentage of SVs supported by SCNA breakpoints at varying distances (<1 kb, <5 kb, <10 kb) for three groups of SVs: those detected only by the union of Meerkat and Manta, those overlapping between Meerkat/Manta and PCAWG, and those detected only by PCAWG. c, Manta-called SVs across different combinations of split reads (y-axis) and read pairs (x-axis) in all Sherlock-Lung Project samples. The dot size represents the number of Manta-called SVs at each combination, while the color indicates the fraction of SVs supported by SCNAs within 1 kb for both break ends. Manta-specific SVs not detected by Meerkat with low read support, located within the red boxes, were filtered out. d, Distributions of read pairs and split reads supporting duplications (top), deletions (middle), and inversions (bottom) detected by Meerkat (left) and Manta (right). The SVs are categorized by size into greater than 1 kb (>1 kb) and less than or equal to 1 kb (≤1 kb). Red dashed lines indicate the 10 supporting-read cutoff, and the red arrows highlight regions identified as potential artifacts in smaller duplications (≤1 kb) detected by Meerkat with a large number of SVs having lower read support.
Supplementary Table S1. Clinical and genomic annotations of 1,209 lung cancer samples used in the study.
This table summarizes the clinical characteristics and genomic annotations of 1,209 whole-genome sequenced lung cancer tumor samples included in this study.
Supplementary Table S2. Driver gene fusions identified in 1,209 samples.
This table lists all driver gene fusions detected in the 1,209 whole-genome sequenced lung cancer samples analyzed in this study. For each fusion event, the genomic coordinates, strand orientations, and partner genes for both the 5′ and 3′ breakpoints are provided, along with the SV signature category.
Supplementary Table S3. SV frequencies across 1,209 samples.
This table reports the number of SVs detected in each tumor sample, including totals and counts stratified by SV subtype. Columns provide the total SV count, counts for complex and simple SVs, and signature-specific categories.
Supplementary Table S4. Clustered complex SV signatures.
This table lists all identified clustered complex SV events across the 1,209 whole-genome sequenced lung cancer samples. For each event, the sample ID, event ID, chromosome, genomic coordinates (start and end positions), and CGR (complex genome rearrangement) status are provided. Additional columns indicate the set of chromosomes involved in a specific event, as well as the assigned SV signature category.
Supplementary Table S5. Driver gene mutations and focal copy number alterations.
This table summarizes the presence and type of driver gene alterations identified in the 1,209 whole-genome sequenced lung cancer samples. For each sample, alterations are listed for a curated set of driver genes and focal copy number changes, such as focal gains or losses. Copy number events are indicated by “gain” or “loss,” and “NA” denotes no detected alteration in the corresponding gene.
Acknowledgments
This work was supported by the Intramural Research Program of the National Institutes of Health (NIH) (project ZIACP101231 to MTL), the Anne Wojcicki Foundation (Grant Number: LC009 to MTL), the NIH grant R01CA269977 (LY), American Lung Association grant LCD-821295 (LY) and University of Chicago Medicine Comprehensive Cancer Center (LY). MD-G was awarded a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública talent attraction (C005/24-ED CV1), funded by the European Union NextGenerationEU funds, through PRTR. The contributions of the NIH author(s) were made as part of their official duties as NIH federal employees, are in compliance with agency policy requirements, and are considered Works of the United States Government. However, the findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services. We thank the study participants and the staff at Westat for their assistance with collecting samples and corresponding clinical data. This work used the computational resources of the NIH HPC Biowulf cluster (http://hpc.nih.gov).
Footnotes
Disclosure
L.B.A. and M.D-G. declare a European patent application with application number EP25305077.7. All other authors have no competing interests to declare. L.B.A. is a co-founder, CSO, scientific advisory member, and consultant for io9, has equity and receives income. The terms of this arrangement have been reviewed and approved by the University of California, San Diego in accordance with its conflict of interest policies. L.B.A. is also a compensated member of the scientific advisory board of Inocras. L.B.A.’s spouse is an employee of Biotheranostics. E.N.B. and L.B.A. declare U.S. provisional patent application filed with UCSD with serial numbers 63/269,033. LBA also declares U.S. provisional applications filed with UCSD with serial numbers: 63/366,392; 63/289,601; 63/483,237; 63/412,835; and 63/492,348. L.B.A. is also an inventor of a US Patent 10,776,718 for source identification by non-negative matrix factorization. L.B.A. and M.D-G. further declare a European patent application with application number EP25305077.7. S.R.Y. has received consulting fees from AstraZeneca, Sanofi, Amgen, AbbVie, and Sanofi; received speaking fees from AstraZeneca, Medscape, PRIME Education, and Medical Learning Institute. All other authors declare that they have no competing interests.
Code availability
All scripts used for data analysis and visualization are available at https://github.com/yanglab-computationalgenomics/Sherlock.SV.
Data availability
Normal and tumor-paired CRAM files for the WGS subjects of the Sherlock-Lung study and the EAGLE study have been deposited in dbGaP under the accession numbers phs001697.v2.p1 and phs002992.v1.p1, respectively. Detailed access information for the publicly available datasets can be found in Supplementary Table 1. Human reference genome GRCh38 was downloaded from the GATK resources at https://github.com/broadinstitute/gatk/blob/master/src/test/resources/large/Homo_sapiens_assembly38.fasta.gz.
References
- 1.Siegel R. L., Miller K. D. & Jemal A. Cancer statistics, 2018. CA Cancer J Clin 68, 7–30 (2018). 10.3322/caac.21442 [DOI] [PubMed] [Google Scholar]
- 2.Couraud S., Zalcman G., Milleron B., Morin F. & Souquet P. J. Lung cancer in never smokers--a review. Eur J Cancer 48, 1299–1311 (2012). 10.1016/j.ejca.2012.03.007 [DOI] [PubMed] [Google Scholar]
- 3.Wakelee H. A. et al. Lung cancer incidence in never smokers. J Clin Oncol 25, 472–478 (2007). 10.1200/JCO.2006.07.2983 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Boffetta P. et al. Multicenter case-control study of exposure to environmental tobacco smoke and lung cancer in Europe. J Natl Cancer Inst 90, 1440–1450 (1998). 10.1093/jnci/90.19.1440 [DOI] [PubMed] [Google Scholar]
- 5.Brownson R. C., Alavanja M. C., Caporaso N., Simoes E. J. & Chang J. C. Epidemiology and prevention of lung cancer in nonsmokers. Epidemiol Rev 20, 218–236 (1998). 10.1093/oxfordjournals.epirev.a017982 [DOI] [PubMed] [Google Scholar]
- 6.Kosaka T. et al. Mutations of the epidermal growth factor receptor gene in lung cancer: biological and clinical implications. Cancer Res 64, 8919–8923 (2004). 10.1158/0008-5472.CAN-04-2818 [DOI] [PubMed] [Google Scholar]
- 7.Le Calvez F. et al. TP53 and KRAS mutation load and types in lung cancers in relation to tobacco smoke: distinct patterns in never, former, and current smokers. Cancer Res 65, 5076–5083 (2005). 10.1158/0008-5472.CAN-05-0551 [DOI] [PubMed] [Google Scholar]
- 8.Shigematsu H. et al. Clinical and biological features associated with epidermal growth factor receptor gene mutations in lung cancers. J Natl Cancer Inst 97, 339–346 (2005). 10.1093/jnci/dji055 [DOI] [PubMed] [Google Scholar]
- 9.Nordquist L. T., Simon G. R., Cantor A., Alberts W. M. & Bepler G. Improved survival in never-smokers vs current smokers with primary adenocarcinoma of the lung. Chest 126, 347–351 (2004). 10.1378/chest.126.2.347 [DOI] [PubMed] [Google Scholar]
- 10.Zell J. A., Ou S. H., Ziogas A. & Anton-Culver H. Epidemiology of bronchioloalveolar carcinoma: improvement in survival after release of the 1999 WHO classification of lung tumors. J Clin Oncol 23, 8396–8405 (2005). 10.1200/JCO.2005.03.0312 [DOI] [PubMed] [Google Scholar]
- 11.Thatcher N. et al. Gefitinib plus best supportive care in previously treated patients with refractory advanced non-small-cell lung cancer: results from a randomised, placebo-controlled, multicentre study (Iressa Survival Evaluation in Lung Cancer). Lancet 366, 1527–1537 (2005). 10.1016/S0140-6736(05)67625-8 [DOI] [PubMed] [Google Scholar]
- 12.Smrdel U. & Kovač V. Erlotinib in previously treated non-small-cell lung cancer. Radiology and Oncology 40 (2006). [Google Scholar]
- 13.Gainor J. F. et al. Clinical activity of programmed cell death 1 (PD-1) blockade in never, light, and heavy smokers with non-small-cell lung cancer and PD-L1 expression >/=50. Ann Oncol 31, 404–411 (2020). 10.1016/j.annonc.2019.11.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Mo J. et al. Smokers or non-smokers: who benefits more from immune checkpoint inhibitors in treatment of malignancies? An up-to-date meta-analysis. World J Surg Oncol 18, 15 (2020). 10.1186/s12957-020-1792-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Negrini S., Gorgoulis V. G. & Halazonetis T. D. Genomic instability--an evolving hallmark of cancer. Nat Rev Mol Cell Biol 11, 220–228 (2010). 10.1038/nrm2858 [DOI] [PubMed] [Google Scholar]
- 16.Hanahan D. & Weinberg R. A. Hallmarks of cancer: the next generation. Cell 144, 646–674 (2011). 10.1016/j.cell.2011.02.013 [DOI] [PubMed] [Google Scholar]
- 17.Consortium I. T. P.-C. A. o. W. G. Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020). 10.1038/s41586-020-1969-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Paik J. H. et al. Clinicopathologic implication of ALK rearrangement in surgically resected lung cancer: a proposal of diagnostic algorithm for ALK-rearranged adenocarcinoma. Lung Cancer 76, 403–409 (2012). 10.1016/j.lungcan.2011.11.008 [DOI] [PubMed] [Google Scholar]
- 19.Mescam-Mancini L. et al. On the relevance of a testing algorithm for the detection of ROS1-rearranged lung adenocarcinomas. Lung Cancer 83, 168–173 (2014). 10.1016/j.lungcan.2013.11.019 [DOI] [PubMed] [Google Scholar]
- 20.Zhang T. et al. Genomic and evolutionary classification of lung cancer in never smokers. Nat Genet 53, 1348–1359 (2021). 10.1038/s41588-021-00920-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Steele C. D. et al. Signatures of copy number alterations in human cancer. Nature 606, 984–991 (2022). 10.1038/s41586-022-04738-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Li Y. et al. Patterns of somatic structural variation in human cancer genomes. Nature 578, 112–121 (2020). 10.1038/s41586-019-1913-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Stephens P. J. et al. Massive genomic rearrangement acquired in a single catastrophic event during cancer development. Cell 144, 27–40 (2011). 10.1016/j.cell.2010.11.055 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Baca S. C. et al. Punctuated evolution of prostate cancer genomes. Cell 153, 666–677 (2013). 10.1016/j.cell.2013.03.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Zhang C. Z. et al. Chromothripsis from DNA damage in micronuclei. Nature 522, 179–184 (2015). 10.1038/nature14493 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Maciejowski J., Li Y., Bosco N., Campbell P. J. & de Lange T. Chromothripsis and Kataegis Induced by Telomere Crisis. Cell 163, 1641–1654 (2015). 10.1016/j.cell.2015.11.054 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Turner K. M. et al. Extrachromosomal oncogene amplification drives tumour evolution and genetic heterogeneity. Nature 543, 122–125 (2017). 10.1038/nature21356 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Storlazzi C. T. et al. Gene amplification as double minutes or homogeneously staining regions in solid tumors: origin and structure. Genome Res 20, 1198–1206 (2010). 10.1101/gr.106252.110 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Díaz-Gay M. et al. The mutagenic forces shaping the genomes of lung cancer in never smokers. Nature 644, 133–144 (2025). 10.1038/s41586-025-09219-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhang T. et al. Deciphering lung adenocarcinoma evolution and the role of LINE-1 retrotransposition. bioRxiv (2025). 10.1101/2025.03.14.643063 [DOI] [Google Scholar]
- 31.Zhang T. et al. APOBEC affects tumor evolution and age at onset of lung cancer in smokers. Nat Commun 16, 4711 (2025). 10.1038/s41467-025-59923-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Imielinski M. et al. Mapping the hallmarks of lung adenocarcinoma with massively parallel sequencing. Cell 150, 1107–1120 (2012). 10.1016/j.cell.2012.08.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Leong T. L. et al. Deep multi-region whole-genome sequencing reveals heterogeneity and gene-by-environment interactions in treatment-naive, metastatic lung cancer. Oncogene 38, 1661–1675 (2019). 10.1038/s41388-018-0536-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Landi M. T. et al. Environment And Genetics in Lung cancer Etiology (EAGLE) study: an integrative population-based case-control study of lung cancer. BMC Public Health 8, 203 (2008). 10.1186/1471-2458-8-203 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Lee J. K. et al. Clonal History and Genetic Predictors of Transformation Into Small-Cell Carcinomas From Lung Adenocarcinomas. J Clin Oncol 35, 3065–3074 (2017). 10.1200/JCO.2016.71.9096 [DOI] [PubMed] [Google Scholar]
- 36.Lee J. J. et al. Tracing Oncogene Rearrangements in the Mutational History of Lung Adenocarcinoma. Cell 177, 1842–1857 e1821 (2019). 10.1016/j.cell.2019.05.013 [DOI] [PubMed] [Google Scholar]
- 37.Cancer Genome Atlas Research, N. Comprehensive molecular profiling of lung adenocarcinoma. Nature 511, 543–550 (2014). 10.1038/nature13385 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Alexandrov L. B. et al. Mutational signatures associated with tobacco smoking in human cancer. Science 354, 618–622 (2016). 10.1126/science.aag0299 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Yoshida K. et al. Tobacco smoking and somatic mutations in human bronchial epithelium. Nature 578, 266–272 (2020). 10.1038/s41586-020-1961-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Bao L., Zhong X., Yang Y. & Yang L. Starfish infers signatures of complex genomic rearrangements across human cancers. Nat Cancer 3, 1247–1259 (2022). 10.1038/s43018-022-00404-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Khandekar A. et al. Examining the Role of Extrachromosomal DNA in 1,216 Lung Cancers. bioRxiv, 2025.2006.2003.657117 (2025). 10.1101/2025.06.03.657117 [DOI] [Google Scholar]
- 42.Zhang F. et al. The DNA replication FoSTeS/MMBIR mechanism can generate genomic, genic and exonic complex rearrangements in humans. Nat Genet 41, 849–853 (2009). 10.1038/ng.399 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Cortes-Ciriano I. et al. Comprehensive analysis of chromothripsis in 2,658 human cancers using whole-genome sequencing. Nat Genet 52, 331–341 (2020). 10.1038/s41588-019-0576-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Islam S. M. A. et al. Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor. Cell Genom 2, None (2022). 10.1016/j.xgen.2022.100179 [DOI] [Google Scholar]
- 45.Everall A. et al. Comprehensive repertoire of the chromosomal alteration and mutational signatures across 16 cancer types from 10,983 cancer patients. medRxiv, 2023.2006.2007.23290970 (2023). 10.1101/2023.06.07.23290970 [DOI] [Google Scholar]
- 46.Yang Y. et al. Transcription and DNA replication collisions lead to large tandem duplications and expose targetable therapeutic vulnerabilities in cancer. Nat Cancer 5, 1885–1901 (2024). 10.1038/s43018-024-00848-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Cosenza M. R., Rodriguez-Martin B. & Korbel J. O. Structural Variation in Cancer: Role, Prevalence, and Mechanisms. Annu Rev Genomics Hum Genet 23, 123–152 (2022). 10.1146/annurev-genom-120121-101149 [DOI] [PubMed] [Google Scholar]
- 48.Durkin S. G. & Glover T. W. Chromosome fragile sites. Annu Rev Genet 41, 169–192 (2007). 10.1146/annurev.genet.41.042007.165900 [DOI] [PubMed] [Google Scholar]
- 49.Beroukhim R. et al. The landscape of somatic copy-number alteration across human cancers. Nature 463, 899–905 (2010). 10.1038/nature08822 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Veeriah S. et al. The tyrosine phosphatase PTPRD is a tumor suppressor that is frequently inactivated and mutated in glioblastoma and other human cancers. Proc Natl Acad Sci U S A 106, 9435–9440 (2009). 10.1073/pnas.0900571106 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Bignell G. R. et al. Signatures of mutation and selection in the cancer genome. Nature 463, 893–898 (2010). 10.1038/nature08768 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Burr R. et al. Developmental mosaicism underlying EGFR-mutant lung cancer presenting with multiple primary tumors. Nat Cancer 5, 1681–1696 (2024). 10.1038/s43018-024-00840-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Winslow M. M. et al. Suppression of lung adenocarcinoma progression by Nkx2–1. Nature 473, 101–104 (2011). 10.1038/nature09881 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Orstad G. et al. FoxA1 and FoxA2 control growth and cellular identity in NKX2–1-positive lung adenocarcinoma. Dev Cell 57, 1866–1882 e1810 (2022). 10.1016/j.devcel.2022.06.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Pulice J. L. & Meyerson M. Amplified dosage of the NKX2–1 lineage transcription factor controls its oncogenic role in lung adenocarcinoma. Mol Cell 85, 1311–1329 e1316 (2025). 10.1016/j.molcel.2025.03.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Cornish A. J. et al. The genomic landscape of 2,023 colorectal cancers. Nature 633, 127–136 (2024). 10.1038/s41586-024-07747-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Bergmann E. A., Chen B. J., Arora K., Vacic V. & Zody M. C. Conpair: concordance and contamination estimator for matched tumor-normal pairs. Bioinformatics 32, 3196–3198 (2016). 10.1093/bioinformatics/btw389 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Pedersen B. S. et al. Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches. Genome Med 12, 62 (2020). 10.1186/s13073-020-00761-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Nik-Zainal S. et al. The life history of 21 breast cancers. Cell 149, 994–1007 (2012). 10.1016/j.cell.2012.04.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Dentro S. C. et al. Characterizing genetic intra-tumor heterogeneity across 2,658 human cancer genomes. Cell 184, 2239–2254 e2239 (2021). 10.1016/j.cell.2021.03.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Zhang F. et al. Ancestry-agnostic estimation of DNA sample contamination from sequence reads. Genome Res 30, 185–194 (2020). 10.1101/gr.246934.118 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Klein A., Zhong J., Landi M. T. & Zhang T. Sherlock-Genome: an R Shiny application for genomic analysis and visualization. BMC Genomics 26, 44 (2025). 10.1186/s12864-024-11147-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Zhu B. et al. The genomic and epigenomic evolutionary history of papillary renal cell carcinomas. Nat Commun 11, 3096 (2020). 10.1038/s41467-020-16546-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Yuan K., Macintyre G., Liu W., group P. w. & Markowetz F. Ccube: A fast and robust method for estimating cancer cell fractions. bioRxiv, 484402 (2018). 10.1101/484402 [DOI] [Google Scholar]
- 65.Yang L. et al. Diverse mechanisms of somatic structural variations in human cancer genomes. Cell 153, 919–929 (2013). 10.1016/j.cell.2013.04.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Chen X. et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 32, 1220–1222 (2016). 10.1093/bioinformatics/btv710 [DOI] [PubMed] [Google Scholar]
- 67.Kim P. & Zhou X. FusionGDB: fusion gene annotation DataBase. Nucleic Acids Res 47, D994–D1004 (2019). 10.1093/nar/gky1067 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Wilkerson M. D. & Hayes D. N. ConsensusClusterPlus: a class discovery tool with confidence assessments and item tracking. Bioinformatics 26, 1572–1573 (2010). 10.1093/bioinformatics/btq170 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Extended Data Figure 1. Self-reported smoking status and SBS4 mutations. a, Number of SBS4 mutations in LCINS and LCSS. Box plots within violins indicate the interquartile ranges and medians. b, Sex, ancestry, and mutation status of EGFR, KRAS, TP53, and fusion status are compared across four groups defined by self-reported smoking status and SBS4 mutations. “n” indicates the number of tumors in each group.
Extended Data Figure 2. Multicollinearity assessment and factors associated with complex SV signatures in 602 high-clonality tumors. a, Pairwise correlations of variables. Circle size represents the absolute value of correlation. Red and blue indicate positive and negative correlations, respectively. b, Variance inflation factor (VIF) values for each variable. Dashed red line indicates cutoff of 5. c, Forest plot for factors associated with complex SV signatures in 602 high-clonality tumors. Odds ratios (ORs) are presented for significant factors in the corresponding signatures. Error bars represent 95% confidence intervals. SCNA subtypes, piano (few SCNAs), mezzo-forte (enriched with arm-level amplification) and forte (dominated by whole-genome doubling), are defined according to our previous study20.
Extended Data Figure 3. Factors associated with complex SV signatures across selected subgroups. ORs are presented for significant factors in corresponding signatures. Error bars represent 95% confidence intervals.
Extended Data Figure 4. Factors associated with the simple SV signatures. a, Factors associated with simple SV signatures in 602 high-clonality tumors from EUR and EAS patients with consistent smoking status. b, Factors associated with simple SV signatures in tumors from selected subgroups. ORs are presented for significant factors in the corresponding signatures. Error bars represent 95% confidence intervals.
Extended Data Figure 5. Comparisons of SV burden across tumors stratified by self-reported smoking status and SBS4 mutations. Box plots within violins indicate the interquartile ranges and medians. Pairwise comparisons differing by only one factor are performed using Wilcoxon rank-sum tests with FDR correction. Only FDRs < 0.05 are shown.
Extended Data Figure 6. Microhomology and insertion at SV breakpoints. Bars represent the number of SV breakpoints with corresponding microhomology (right) or insertion size (left). Suggested repair mechanisms are indicated for each signature. NHEJ, nonhomologous end joining; alt-EJ, alternative end joining; MMBIR, microhomology mediated break induced repair.
Extended Data Figure 7. Driver alterations in EGFR-mutant and EGFR-wildtype LCINS. Bars represent the proportion of tumors carrying mutations/indels and SVs in known driver genes stratified by EGFR mutation status (y-axis). Enrichment is tested using Fisher’s exact test with FDR correction. Only FDRs < 0.05 are labeled by “*”s.
Extended Data Figure 8. SV signature composition of focal driver gains and losses in 809 EGFR-mutant and EGFR-wildtype LCINS. Each bar represents a driver gene, with color indicating the associated SV type (left) or SV signature (right) identified as the closest event to the focal alteration. Enrichment of SV types or signatures between LCINS and LCSS was tested using Fisher’s exact test with FDR correction. Significantly enriched SV types and signatures are marked with black outlines. No significant enrichment is detected.
Extended Data Figure 9. Examples of SVs of different signatures leading to NKX2–1 focal gains. Colored arcs represent SVs of different types. Red bars beneath the arcs mark the regions of clustered complex SVs. The copy number profiles are shown as black bars above the chromosome models. Centromere positions and NKX2–1 genes are highlighted as red bars and black boxes within the grey chromosome ideograms. Sample IDs are displayed next to the corresponding SV signatures.
Extended Data Figure 10. Proposed model of lung cancer progression in LCINS and LCSS. Schematic illustration shows how distinct tumor initiating events occur and follow by distinct secondary driver alterations, shaping downstream pathways and contributing to lung cancer development in LCINS and LCSS.
Extended Data Figure 11. Assessment of SV calling. a, SVs detected by Meerkat and Manta in 35 lung adenocarcinoma (LUAD) samples from the PCAWG consortium compared to SVs detected by the PCAWG consortium. The three bar plots represent SVs called by Manta (left), Meerkat (middle), and the union of both algorithms (right). The bar colors correspond to the Venn diagram on the right, which illustrates the relationships between SVs of the testing sets (Meerkat, Manta, or their union) and PCAWG SVs. b, Bar plots showing the percentage of SVs supported by SCNA breakpoints at varying distances (<1 kb, <5 kb, <10 kb) for three groups of SVs: those detected only by the union of Meerkat and Manta, those overlapping between Meerkat/Manta and PCAWG, and those detected only by PCAWG. c, Manta-called SVs across different combinations of split reads (y-axis) and read pairs (x-axis) in all Sherlock-Lung Project samples. The dot size represents the number of Manta-called SVs at each combination, while the color indicates the fraction of SVs supported by SCNAs within 1 kb for both break ends. Manta-specific SVs not detected by Meerkat with low read support, located within the red boxes, were filtered out. d, Distributions of read pairs and split reads supporting duplications (top), deletions (middle), and inversions (bottom) detected by Meerkat (left) and Manta (right). The SVs are categorized by size into greater than 1 kb (>1 kb) and less than or equal to 1 kb (≤1 kb). Red dashed lines indicate the 10 supporting-read cutoff, and the red arrows highlight regions identified as potential artifacts in smaller duplications (≤1 kb) detected by Meerkat with a large number of SVs having lower read support.
Supplementary Table S1. Clinical and genomic annotations of 1,209 lung cancer samples used in the study.
This table summarizes the clinical characteristics and genomic annotations of 1,209 whole-genome sequenced lung cancer tumor samples included in this study.
Supplementary Table S2. Driver gene fusions identified in 1,209 samples.
This table lists all driver gene fusions detected in the 1,209 whole-genome sequenced lung cancer samples analyzed in this study. For each fusion event, the genomic coordinates, strand orientations, and partner genes for both the 5′ and 3′ breakpoints are provided, along with the SV signature category.
Supplementary Table S3. SV frequencies across 1,209 samples.
This table reports the number of SVs detected in each tumor sample, including totals and counts stratified by SV subtype. Columns provide the total SV count, counts for complex and simple SVs, and signature-specific categories.
Supplementary Table S4. Clustered complex SV signatures.
This table lists all identified clustered complex SV events across the 1,209 whole-genome sequenced lung cancer samples. For each event, the sample ID, event ID, chromosome, genomic coordinates (start and end positions), and CGR (complex genome rearrangement) status are provided. Additional columns indicate the set of chromosomes involved in a specific event, as well as the assigned SV signature category.
Supplementary Table S5. Driver gene mutations and focal copy number alterations.
This table summarizes the presence and type of driver gene alterations identified in the 1,209 whole-genome sequenced lung cancer samples. For each sample, alterations are listed for a curated set of driver genes and focal copy number changes, such as focal gains or losses. Copy number events are indicated by “gain” or “loss,” and “NA” denotes no detected alteration in the corresponding gene.
Data Availability Statement
Normal and tumor-paired CRAM files for the WGS subjects of the Sherlock-Lung study and the EAGLE study have been deposited in dbGaP under the accession numbers phs001697.v2.p1 and phs002992.v1.p1, respectively. Detailed access information for the publicly available datasets can be found in Supplementary Table 1. Human reference genome GRCh38 was downloaded from the GATK resources at https://github.com/broadinstitute/gatk/blob/master/src/test/resources/large/Homo_sapiens_assembly38.fasta.gz.






