Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2025 Jul 24:2025.06.03.657117. Originally published 2025 Jun 7. [Version 2] doi: 10.1101/2025.06.03.657117

Examining the Role of Extrachromosomal DNA in 1,216 Lung Cancers

Azhar Khandekar 1,2,3,4, Phuc H Hoang 1, Jens Luebeck 5, Marcos Díaz-Gay 2,3,4,6, Wei Zhao 1, John P McElderry 1, Caleb Hartman 1, Mona Miraftab 1, Olivia W Lee 1, Kara M Barnao 1, Erik N Bergstrom 2,3,4, Yang Yang 7, Martin A Nowak 8,9, Nathaniel Rothman 1, Robert Homer 10, Soo-Ryum Yang 11, Qing Lan 1, David C Wedge 12,13, Lixing Yang 7,14,15, Stephen J Chanock 1, Tongwu Zhang 1,*, Ludmil B Alexandrov 2,3,4,16,*, Maria Teresa Landi 1,*
PMCID: PMC12157361  PMID: 40501879

Abstract

The role of extrachromosomal DNA (ecDNA) in lung cancer, particularly in subjects who never smoked (LCINS), remains unclear. Examination of 1,216 whole-genome-sequenced lung cancers identified ecDNA in 18.9% of patients. Recurrent amplification of MDM2 and other oncogenes via ecDNA possibly drives a LCINS subset. Tumors harboring ecDNA showed worse overall survival than tumors harboring other focal amplifications. A strong association with whole-genome doubling suggests most ecDNA reflects genomic instability in treatment-naïve lung cancer.


Extrachromosomal DNA (ecDNA) is a potent mechanism for oncogene amplification in human cancer1, exhibiting unique characteristics such as a circular structure2, non-Mendelian inheritance3,4, and an altered epigenetic and transcriptional landscape5. Prior studies have linked ecDNA to worse overall survival in pan-cancer analyses, highlighting its clinical significance6-8 and potential as a therapeutic target9. Recent advances in computational techniques10 have enabled large-scale detection of ecDNA from whole-genome sequencing (WGS) data. However, only three pan-cancer studies have examined ecDNA in primary lung cancer cohorts with over 100 samples, focusing predominantly on patients of European descent6-8. Two studies6,7 analyzed lung cancer samples from the United States, reporting ecDNA prevalence rates of about 21% in 36 lung adenocarcinomas (LUAD) and 27% in 47 lung squamous cell carcinomas (LUSC). Additionally, a UK lung cancer study8 analyzed 14,778 patients across 39 tumor types, including 718 LUAD tumors and 378 LUSC tumors, detecting ecDNA in 8.2% of LUAD and 22.4% of LUSC cases. All three studies6-8 reported lung cancer findings as part of broader pan-cancer analyses and did not distinguish between lung cancers in subjects who have never-smoked (LCINS) and subjects who have smoked (LCSS). Notably, based on the high prevalence of the tobacco-associated11 mutational signature SBS4 in these studies, more than 84% of the analyzed lung cancer cases were from individuals with a history of smoking.

Despite these prior studies, the clinical relevance of ecDNA in lung cancer remains unclear, including its prevalence across histologies, geographical areas, genetic ancestries, and biological sexes. Its role in LCINS, LCSS, and subjects who were exposed to passive smoking also requires further exploration. To address these questions, here, we analyzed WGS data from 1,216 treatment-naïve lung cancers from diverse patient populations (Fig. 1a; Supplementary Fig. 1). This large international cohort enabled a comprehensive evaluation of ecDNA’s association with genomic alterations, as well as its prevalence across smoking status, biological sexes, ancestries, and 18 areas defined primarily by country of residence. The dataset includes 871 LCINS, sourced from the Sherlock-Lung12 (n=817), the Environment And Genetics in Lung Cancer Etiology (EAGLE13, n=25), the Cancer Genome Atlas (TCGA14, n=15), and other publicly available cohorts (n=14)15-19, representing patients from 17 geographical areas (Supplementary Fig. 2a). Additionally, it comprises 345 LCSS, drawn from EAGLE13 (n=221) and TCGA14 (n=83) and other publicly available cohorts (n=41)15-19, spanning five geographical areas (Supplementary Fig. 2b). The majority of tumors were LUAD (n=1,023; 84.1%) or LUSC (n=67; 5.5%), with carcinoids (n=61; 5%, exclusively in LCINS) and other rarer histological lung cancer subtypes (n=65; 5.3%) also represented. Amongst Sherlock-Lung LCINS, 250 were exposed to secondhand tobacco smoke, while 208 had documented non-exposure (Fig. 1a).

Figure 1: Landscape of extrachromosomal DNA in lung cancer across histologies, sex, ancestry, countries, and smoking status.

Figure 1:

a) Epidemiological features of the LCINS (top) and LCSS (bottom) cohorts. Total number of samples in each category is indicated. b) The number and proportion of tumors harboring ecDNA among LCINS and LCSS, stratified by histology. Samples with 1 or more ecDNA were classified as ecDNA+, and samples with no ecDNA were classified as ecDNA. c) Forest plot of the full cohort (LCINS and LCSS) from a multivariate logistic regression analysis of key epidemiological and clinical features and the presence of ecDNA in the tumors. d) Volcano plot showing log2 odds ratio (x-axis) and log10 FDR corrected q-value (y-axis) for a one-vs-all statistical comparison of ecDNA prevalence by region. e) Forest plots for multivariate logistic regression models testing key epidemiological and clinical features with the presence of ecDNA in the tumors, stratified by smoking status. NA/EU: North America or Europe region, AS: Asia region EAS: East Asian genetic ancestry super-sample, EUR: European genetic ancestry super-sample, LUAD: lung adenocarcinoma, LUSC: lung squamous cell carcinoma.

To detect and reconstruct ecDNA across the cohort, a computational pipeline that accounts for both tumor purity and ploidy was developed (Methods). A total of 231 samples harbored at least one ecDNA (18.9%), with similar prevalence between LCINS (17%) and LCSS (23%; q-value: 0.21; Fig. 1b). No statistically significant differences were observed in the prevalence of ecDNA across various histological subtypes, biological sexes, ancestries or age at diagnosis (Fig. 1c&e), except for carcinoids, which showed a lower prevalence compared to LUAD in the overall analysis (OR=0.07; q=0.04). ecDNA was present at similar rates across all geographic locations (Fig. 1d; Supplementary Fig. 2c-d), as indicated by a one-versus-all multivariate logistic regression for each country. There was also no difference in ecDNA prevalence between LCSS and LCINS (Fig. 1c). ecDNA was not associated with tumor stage (Supplementary Fig. 3a&b), subsequent development of metastases (Supplementary Fig. 3c&d) or passive smoking status (Supplementary Fig. 3e). These findings remained consistent when analyses were restricted to only LUAD cases in either LCSS or LCINS (Supplementary Fig. 4).

To examine the association between ecDNA and genomic features previously linked to genome instability20-23, we evaluated the relationship of ecDNA with whole-genome doubling (WGD), chromothripsis, tumor mutational burden, mitochondrial copy number, overall genome instability index (wGII), and telomere length tumor/normal ratio. Amongst these, only WGD showed a significant association with ecDNA in both LCSS and LCINS (Fig. 2a), as determined by logistic regression adjusted for age, sex, ancestry, histology, and tumor purity. Specifically, ecDNA was present in 7.1% of LCINS and 8.9% of LCSS without WGD, but its prevalence increased markedly in WGD-positive tumors from LCINS (24.9%, OR=4.0; q=1.24 x 10−8) and LCSS (30.3%, OR=5.85; q=8.06 x 10−4). The same association was observed when only LUAD was examined (Supplementary Fig. 5a). Moreover, this association was further validated in a pan-cancer dataset of 974 samples with high-confidence WGD annotation14, using logistic regression adjusted for age, purity, sex, and cancer type (OR=2.81; q=1.63 x 10−6; Supplementary Fig. 5b).

Figure 2. Genomic and clinical characteristics associated with ecDNA.

Figure 2.

a) Volcano plot of logistic regression models of genomic features in association with ecDNA status for LCINS (left) and LCSS (right). b) Volcano plot of a logistic regression model of driver gene alterations in a sample in association with ecDNA status for LCINS (left) and LCSS (right). c) Volcano plot of a logistic regression model of hotspot driver mutations in a sample in association with ecDNA status for LCINS (left) and LCSS (right). In all volcano plots, the x-axes reflect the log2 odds ratio, and the y-axes correspond to the log10 FDR q-value. An FDR q-value threshold of 0.05 is indicated with the dashed red line, an FDR q-value threshold of 0.01 is indicated with the dashed blue line. d) Bar plot indicating the proportion of ecDNA containing a particular genomic element in LCINS vs LCSS. e) Kaplan–Meier survival curves for 5-year overall survival stratified by the mode of amplification status for lung cancers in LCINS (left) and LCSS (right). P-values for significance and HRs of the difference were calculated using two-sided Cox proportional-hazards regression with adjustment for age, sex, tumor stage, ancestry and histology and are indicated within each plot.

In addition to WGD, chromothripsis was the only other feature associated with ecDNA, but this was exclusively observed in LCINS (OR=1.76; q=0.048; Fig. 2a). Prior studies have indicated that chromothripsis may contribute to the formation of certain ecDNA23. In this study, 115 out of 151 LCINS with ecDNA (76.2%) and 60 out of 80 LCSS with ecDNA (75%) exhibited chromothripsis (Supplementary Tables 1-2). Nevertheless, in samples where both ecDNA and chromothripsis were detected, overlap between ecDNA segments and chromothripsis regions was rare, occurring more frequently in LCINS—observed in 31 of 115 LCINS cases compared to 4 of 60 LCSS (OR=5.01, p=0.02 comparing LCINS vs. LCSS; Supplementary Fig. 5c). This suggests that there could be potentially different mechanisms of ecDNA formation between LCINS and LCSS. However, while the majority of ecDNA emerges in genomically unstable lung cancers that also exhibit chromothripsis, the actual genomic overlap between ecDNA and chromothripsis regions is limited in both lung cancer types.

To evaluate the relationship between ecDNA presence and lung cancer driver genes, we analyzed driver genes altered in at least 10 samples as well as driver hotspot mutations occurring at the same genomic position in at least three samples. Using multivariate logistic regression, adjusted for age, sex, tumor purity, ancestry, and histology, no significant associations were found between ecDNA and individual driver gene alterations (Fig. 2b). However, hotspot analysis showed an enrichment of EGFR L858R mutations (OR=1.78; q=0.05; Fig. 2c; Supplementary Fig. 5d) in LCINS with ecDNA and a depletion of KRAS G12V in LCSS with ecDNA (OR=0.18; q=0.03; Fig. 2c).

Mutational signature analysis provides insights into both the endogenous and exogenous mutational processes that have been active throughout the lineage of a cancer genome11. We investigated the association between ecDNA presence and mutational signatures. As expected24, multiple copy-number and structural variant signatures were associated with ecDNA presence similarly in LCSS and LCINS (Supplementary Fig. 6a-b). Additionally, in LCINS, ecDNA-positive samples were enriched for clock-like signatures (SBS1, SBS5, ID1, ID2) and, as previously reported19,25, APOBEC-associated signatures (SBS2, SBS13; Supplementary Fig. 6a). In LCSS, enrichment was observed only for SBS3 (Supplementary Fig. 6b), a signature previously associated with homologous recombination deficiency26.

Previous studies have shown that ecDNA can harbor oncogenes, regulatory elements, and immunomodulatory genes in varying configurations3,8,27,28. In our cohort, 151 (17%) LCINS and 80 (23%) LCSS harbored ecDNA, with a total of 236 ecDNAs in LCINS and 123 in LCSS, reflecting multiple ecDNAs in some tumors. Among these, ecDNA in LCINS were enriched with oncogenes (p<0.001) and immunomodulatory genes (p=0.008) compared to LCSS (Fig. 2d; Supplementary Fig. 7a-b). Further, ecDNA containing oncogenes exhibited elevated copy numbers (CN) compared to other ecDNA without oncogenes in LCINS (p=1.5 x 10−3) and LCSS (5.8 x 10−3), suggesting greater selective pressure for these ecDNAs (Supplementary Fig. 7c), as previously reported in pan-cancer8. MDM2 was the most frequently amplified oncogene on ecDNA in both LCINS (n=36) and LCSS (n=5). Given its prominence, we examined the relationship between ecDNA-driven MDM2 amplification and TP53 alterations, as MDM2 negatively regulates TP53 through ubiquitination29. We found that MDM2 amplification on ecDNA was mutually exclusive with TP53 alterations (OR=7.51, p=0.01). In LCINS, the next four most prevalent oncogenes that did not co-occur with MDM2 were TERT (n=10), CCND1 (n=8), MYC (n=8), and EGFR (n=7; Supplementary Table 3). In LCSS, no oncogene besides MDM2 was amplified on ecDNA in more than two samples (Supplementary Table 4).

Prior pan-cancer studies have shown that cancers harboring ecDNA amplifications exhibit worse overall survival6-8. To evaluate this association in the lung cancer cohort, survival analysis was conducted on 693 LCINS and 333 LCSS with available stage and covariate data. Lung cancers were categorized as harboring ecDNA, other focal amplifications, or no focal amplifications. Multivariate Cox proportional-hazards regression analysis revealed that ecDNA-positive lung cancers had worse overall survival than those with only chromosomal amplifications; this association was statistically significant in LCINS (hazard ratio=2.17; p=0.025; Fig. 2e).

In summary, we examined the full spectrum of ecDNA in the largest dataset of whole-genome sequenced lung cancers to date, from many geographical regions and ancestry, and with unique information on tobacco smoking status. ecDNA was present in 17% of LCINS and 23% of LCSS. ecDNA in LCINS were enriched with oncogenes and immunomodulatory genes in comparison to LCSS. In LCINS, the MDM2 locus was often amplified through ecDNA, while the EGFR L858R mutation showed a weak association with ecDNA presence. No associations were identified between presence of ecDNA and histological subtypes, biological sexes, smoking status, or ancestries. However, tumors carrying ecDNA showed worse overall survival in comparison to those with focal chromosomal amplification, particularly in LCINS. A strong association was observed between ecDNA and WGD, and this observation was validated in an independent pan-cancer cohort, suggesting that most ecDNA in treatment-naïve lung cancer are likely to be a byproduct of genomic instability. Notably, ecDNA carrying oncogenes could contribute to driving LCINS and affecting prognosis.

METHODS

Criteria for sample inclusion

A total of 1217 samples were available for analysis, but one sample with unknown smoking status was excluded, resulting in a total of 1216 samples across the cohort.

Determination of sample purity and ploidy

The Battenberg algorithm (v.2.2.9)30 was used to determine the sample purity and ploidy. Any somatic copy number variations (CNV) profile determined to have low-quality after manual inspection underwent a refitting process. This process required new tumor purity and ploidy inputs, either estimated by Ccube (v.1.0)31 or recalculated from local copy number status. The Battenberg refitting procedures were iteratively executed until the final CNV profile was established and met the criteria of manual validation check.

ecDNA detection and characterization

CNVKit v.0.9.632 was run in tumor-normal mode to call somatic CNVs against the matched normal whole-genome sequenced samples for each patient, using the log ratio of tumor to normal read depths (logR) and the circular binary segmentation (CBS) method. Note that this method, by default, assumes a sample purity of 1 and a sample ploidy of 2 as it does not correct for purity and ploidy. This limitation can lead to high copy number segments being missed due to low purity33. Due to the wide range of tumor purities in our dataset, consistent with prior reports in lung cancer33, we rescaled the total copy number estimate using the purity and ploidy estimated from Battenberg according to the following equation:

TCN=(ψ(2logR))2(1ρ)ρ Equation 1.1.

where ψ is the tumor ploidy, ρ is the tumor purity, and logR represents the log ratio of the tumor read depth to the normal read depth (Supplementary Fig. 8a). These copy number calls were used to identify putative ‘seed’ regions. We then used AmpliconArchitect (v.1.3)10 to reconstruct the architecture of the amplicons and AmpliconClassifier (v.1.3.1) to classify amplicons according to the most probable mechanism of formation (ecDNA, BFB, complex non-cyclic, or linear). Visual inspection and verification were performed for all ecDNA amplicon regions. This pipeline allowed for samples with low purities to be utilized as it completely removed the effect of tumor purity in ecDNA detection (Supplementary Fig. 8b). As a result, more ecDNA+ samples were detected (Supplementary Fig. 8b) and all 1216 samples could be used for ecDNA analysis (Supplementary Fig. 8c). Of the additional ecDNA detected, eight of these were EGFR-ecDNA (seven in LCINS and one in LCSS) and each of these ecDNA except for one also harbored a driver mutation (Supplementary Table 5). Among the eight EGFR-ecDNA+ samples, five had RNA-seq data, which showed very high EGFR expression (Supplementary Fig. 8d).

Determination of whole genome doubling status

As previously done34, samples were considered to have a whole genome doubling event if the major copy number was ≥3 for more than 50% of copy number segments.

Chromothripsis detection and characterization

To detect chromothripsis regions in genomes, we applied the computational algorithm ShatterSeek35, which employs a set of statistical criteria to identify chromothripsis given an input of copy number and structural variants. We used the most stringent criteria corresponding to high confidence chromothripsis calls. Specifically, these criteria were: (i) a cluster of structural variants (>6 DUP/DEL/h2hINV/t2tINV), (ii) oscillating CNV between two states (>7 CNV events), (iii) chromosomal enrichment and distribution of DNA breakpoints (p-value<0.05), (iv) randomness of fragment joins (p-value>0.05) and/or ≥4 inter-chromosomal rearrangements between multiple chromosomes. The input CNV calls were generated using CNVkit in tumor-normal mode with default parameters, and the input structural variants calls were generated using a union of Meerkat (v.0.189) and Manta (v.1.6.0) calls with the recommended filtering.

Association of ecDNA with demographic and clinical features

Demographic and clinical features where information was available for all samples (age, sex, ancestry, and tumor histology and purity) served as input variables in a multivariate logistic regression model with ecDNA presence or absence as the dependent outcome variable, and these models were constructed separately for LCINS and LCSS. In addition, separate models for passive smoking, tumor stage, and presence of metastases were run along with the rest of the variables in both LCINS and LCSS.

Determination of telomere length tumor/normal lung tissue ratio

We estimated telomere length (TL) in kilobases using TelSeq (v.0.0.2)36. As was previously done37, we used seven as the threshold for the number of TTAGGG/CCCTAA repeats in a read for the read to be considered telomeric. The TelSeq calculation was done individually for each read group within a sample, and the total number of reads in each read group was used as weight to calculate the average TL for each sample. The telomere length ratio (log2 scale) between tumor and normal tissue was then calculated for each patient and served as the independent variable in multivariable logistic regressions.

Association of mutational signatures and ecDNA status

Mutational signature attributions were determined as was previously done, based on the methodology in Díaz-Gay et. al37. SigProfilerMatrixGenerator38,39 was used to create the mutational matrix for all types of somatic mutations. The deconvolution of mutational signatures was performed by SigProfilerExtractor (v.1.1.3)40 using the whole-genome sequencing setting and the COSMIC mutational signatures (v3.2) as reference. Signature extractions were performed separately for LCINS and LCSS. Signature attributions were input into a multivariate logistic regression model as either present or absent if the signature was present in less than 50% of the samples. If the signature was present in ≥50% samples ( e.g., SBS1 and SBS5), it was input as either above median or below median. For SBS1 and SBS5, the signature attributions were also input as continuous numeric variables to confirm the significant associations found with inputting them as above median or below median.

Association of driver genes and ecDNA status

The IntOGen pipeline (v.2020.02.0123)41, which combines seven state-of-the-art computational methods, was employed with default parameters to detect signals of positive selection in the mutational patterns of driver genes across the cohort. The status of each driver gene (altered or unaltered) was then input into a logistic regression model with covariates age, sex, and tumor purity, and ecDNA status was the output. These sets of genes were considered altered if there was a focal deletion overlapping the gene, an inactivating (stop or missense) mutation, or an activating gain of function mutation.

Association of driver mutations and ecDNA status

To identify driver mutations within the set of identified driver genes, we implemented a rigorous and multifaceted strategy, considering multiple criteria: (i) the presence of truncating mutations specifically in genes annotated as tumor suppressors, (ii) the recurrence of missense mutations in a minimum of 3 samples, (iii) mutations designated as "Likely drivers" with a boostDM score42 exceeding 0.5, (iv) mutations categorized as either "Oncogenic" or "Likely Oncogenic" based on the criteria established by OncoKB43, an expert-guided precision oncology knowledge base, (v) mutations previously recognized as drivers in the TCGA MC3 drivers paper44, and (vi) missense mutations characterized as "likely pathogenic" in genes that are annotated as tumor suppressors, as described in Cheng et al.45 Any mutation meeting one or more of these criteria was recognized as a potential driver mutation.

Calculation of genome instability index score

To calculate the genome instability index score (wGII), the method devised in Bailey et al.8 was used. The overall ploidy was calculated as the length-weighted average of all copy number segments. The proportion of segments that were aberrant (where total copy number was greater than the ploidy or less than the ploidy) was computed for each chromosome, these proportions were summed, and the total was divided by 22.

Identifying regulatory regions on ecDNA

A set of lung enhancer regions were downloaded from EnhancerAtlas46, and each ecDNA region was overlapped with these enhancer regions with the requirement that the entire enhancer had to be fully contained (100% overlap) on the ecDNA. A set of lung cancer promoter regions was extracted using a R script which downloaded regions 1000 base pairs upstream of transcription start sites for all genes.

Defining a set of immunomodulatory genes

ecDNAs were considered immunomodulatory if a gene from the significant gene sets mapped to an ecDNA that did not contain an oncogene, and that significant gene set had an immunomodulatory function (GO terms: 0006968, 0002228, 0042267, 0001906, 0001909, 0002698, 0001910, 0031341, 0002367, 0002695, 0050866, 0051250, 0050777), following the procedure in Bailey et. al8. This resulted in a set of 376 genes.

Mitochondrial copy number estimation

Mitochondrial copy number was estimated per sample using the ratio of mitochondrial read depth to nuclear genome depth, adjusted for tumor ploidy and purity, following the procedure in Reznik et al.47

Statistical analysis

All statistical analyses and graphic displays were performed using the R software v4.2.3 1011 (https://www.r-project.org/). Standard logistic regression with adjustment for covariates was used for the enrichment analyses of categorical variables. We applied Firth’s penalized logistic regression using the logistf package in R to model the association between ecDNA status and geographic regions, in order to reduce small-sample bias and address potential issues of data separation. For the comparison of numerical variables across groups, we used non-parametric Mann-Whitney (Wilcoxon rank sum) tests. P-values <0.05 were considered statistically significant. Survival analyses were performed using Cox proportional-hazards regression adjusted for age at diagnosis, sex, tumor stage, tumor histology, ancestry, and wGII. If multiple hypothesis testing was required, we used a false-discovery rate correction based on the Benjamini-Hochberg method and reported FDR. FDR <0.05 were considered statistically significant.

Supplementary Material

Supplement 1

Supplementary Table 1. Co-occurrence of ecDNA genetic cargo in LCINS

Supplementary Table 2. Co-occurrence of ecDNA genetic cargo in LCSS

Supplementary Table 3. Recurrently amplified oncogenes on ecDNA in LCINS across unique genomic loci

Supplementary Table 4. Recurrently amplified oncogenes on ecDNA in LCSS across unique genomic loci

Supplementary Table 5. Additional EGFR-ecDNA detected by modified pipeline

media-1.xlsx (46.8KB, xlsx)
Supplement 2

Supplementary Fig 1. Additional cohort information. Sankey diagrams showing the numbers of samples with clinical variables (stage, metastasis) as well as survival data available for a) LCINS and b) LCSS

Supplementary Fig 2. Distribution of ecDNA harboring lung cancers in LCINS and LCSS by country. Maps showing the worldwide prevalence of lung cancers harboring ecDNA (ecDNA+) by country for the a) LCINS cohort and b) LCSS cohort. Volcano plot of a multivariate one-vs-all logistic regression model for each country, with ecDNA status as the outcome, for c) LCINS and d) LCSS.

Supplementary Fig 3. Analysis of the association between stage, metastasis, and passive smoking with ecDNA status in all LCINS and LCSS. Forest plots of logistic regression model with ecDNA status as the outcome and the variables for stage in a) LCINS and b) LCSS; metastasis in c) LCINS and d) LCSS; and passive smoking in e) LCINS. The total number of samples with information for all variables is indicated above each plot.

Supplementary Fig 4. Analysis of the association between stage, metastasis, and passive smoking with ecDNA status in lung adenocarcinomas. Forest plots of logistic regression model with ecDNA status as the outcome and the variables for stage in a) LCINS and b) LCSS; metastasis in c) LCINS and d) LCSS; and passive smoking in e) LCINS. The total number of samples with information for all variables is indicated above each plot.

Supplementary Fig 5. Genomic features associated with the presence of ecDNA. a) Volcano plot of logistic regression models of genomic features in association with ecDNA status for lung adenocarcinoma in LCINS (left) and lung adenocarcinoma in LCSS (right). In all volcano plots, the x-axes reflect the log2 odds ratio, and the y-axes correspond to the log10 FDR q-value. An FDR q-value threshold of 0.05 is indicated with the dashed red line, an FDR q-value threshold of 0.01 is indicated with the dashed blue line. b) Forest plot of multivariate logistic regression modeling the presence or absence of ecDNA in relation to whole-genome doubling in a pan-cancer dataset comprising four different tumor types from the PCAWG project, adjusted for the covariates of age, purity, sex, and tissue type. All significant associations are indicated in red. c) Bar plot indicating the proportion of samples that contained overlapping chromothripsis and ecDNA regions (red) and the proportion of all samples that did not contain any overlap (gray) for all samples that had both an ecDNA and chromothripsis event detected. d) Bar plot showing the proportion of samples with an EGFR L858R mutation, stratified by ecDNA status.

Supplementary Fig 6. Association of mutational signatures with ecDNA status. Volcano plots of logistic regression models of the activities of Single base substitution (SBS), Indel (ID), copy number (CN) and structural variant (SV) signatures in association with ecDNA in lung cancers of a) LCINS and b) LCSS. The logistic regression models were adjusted for age, sex, ancestry, histology, and purity. In all volcano plots, the x-axes reflect the log2 odds ratio, and the y-axes correspond to the log10 FDR q-value. An FDR q-value threshold of 0.05 is indicated with the dashed red line, an FDR q-value threshold of 0.01 is indicated with the dashed blue line.

Supplementary Fig 7. Genomic content of ecDNA

a) Upset plot indicating the number of ecDNA with oncogenes, immunomodulatory genes, promoter or enhancer regions, other genes that are not oncogenes, or unidentified genomic elements (Unknown) or any combination thereof for LCINS and LCSS. b) Bar plot indicating the proportion of ecDNA containing a particular genomic element in LCINS vs LCSS. c) log2 median copy number (y-axis) for each class of ecDNA (x-axis) in LCINS and LCSS.

Supplementary Fig 8. Highly sensitive pipeline for ecDNA detection that accounts for tumor purity and ploidy.

a) Schematic depicting the methodology for the rescaled pipeline that considers tumor purity and ploidy. The logR track for the region containing the EGFR gene on chr7 and the purity and ploidy computed from Battenberg is indicated in one tumor as example (left). The exact location of the EGFR gene is shown by the dashed red lines. The equation that simultaneously accounts for sample purity and ploidy when computing the total copy number is shown, as well as the total copy number estimate when assuming a sample purity of 1.0 and ploidy of 2.0 (Uncorrected CN) vs. considering the purity and ploidy (Corrected CN). The Sashimi plot indicates the structural variants (colored arcs) and smoothed coverage around the EGFR region, which was annotated as ecDNA by AmpliconClassifier. b) Distribution of tumor purities for samples with an ecDNA annotation for both the default (left) and rescaled (right) pipelines. P-value for a Mann-Whitney test between ecDNA+ and ecDNA samples is indicated. The total number of samples in each group is also indicated. c) The ecDNA prevalence (y-axis), which is the proportion of ecDNA+ samples over the total number of samples, for all possible tumor purity cutoffs applied to samples (x-axis). The total number of samples available for analysis after each tumor purity cutoff is applied is also indicated. d) Gene expression (log2 CPM) of the EGFR gene for all samples with RNA-seq data available. LCINS samples with EGFR-ecDNA are indicated in red.

media-2.pdf (2.3MB, pdf)

ACKNOWLEDGEMENTS

This work was supported by the Intramural Research Program of the National Cancer Institute, US National Institute of Health (NIH) (project ZIACP101231 to MTL), and by the Anne Wojcicki Foundation (Grant Number: LC009). Part of this work was supported by the US National Institute of Health grants R01ES032547, R01CA269919, P01CA281819, and U01CA290479 to LBA as well as by LBA’s Packard Fellowship for Science and Engineering. MD-G was awarded with a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública, for talent attraction (C005/24-ED CV1), funded by the European Union NextGenerationEU funds, through PRTR. The research in this study was also supported by UC San Diego Sanford Stem Cell Institute. This work utilized the computational resources of the NIH HPC Biowulf cluster (https://hpc.nih.gov). The funders had no roles in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Footnotes

COMPETING INTERESTS

LBA is a co-founder, CSO, scientific advisory member, and consultant for io9, has equity and receives income. The terms of this arrangement have been reviewed and approved by the University of California, San Diego in accordance with its conflict of interest policies. LBA is a compensated member of the scientific advisory board of Inocras. LBA’s spouse is an employee of Hologic, Inc. LBA declares U.S. provisional applications with serial numbers: 63/289,601; 63/269,033; 63/366,392; 63/412,835 as well as international patent application PCT/US2023/010679. LBA is also an inventor of a US Patent 10,776,718 for source identification by non-negative matrix factorization. LBA and MD-G further declare a European patent application with application number EP25305077.7. Soo-Ryum Yang has received speaking fees from Medscape, OncLive, Cure Today, Medical Learning Institute, PRIME Education, AstraZeneca, Roche; and consulting fees from AstraZeneca, AbbVie, Merus, Eli Lilly, Boehringer Ingelheim, Roche, Amgen, Sanofi. All other authors declare no conflicts of interest.

Ethics declarations

Since the National Cancer Institute only received de-identified samples and data from collaborating centers, had no direct contact or interaction with the study participants, and did not use or generate identifiable private information, Sherlock-Lung has been determined to constitute “Not Human Subject Research (NHSR)” based on the federal Common Rule (45 CFR 46; https://www.ecfr.gov/cgi-bin/ECFR?page=browse).

DATA AVAILABILITY

Normal and tumor-paired CRAM files for the WGS subjects of the Sherlock-Lung study and the EAGLE study have been deposited in dbGaP under the accession numbers phs001697.v2.p1 and phs002992.v1.p1, respectively.

REFERENCES

  • 1.Benner S. E., Wahl G. M. & Von Hoff D. D. Double minute chromosomes and homogeneously staining regions in tumors taken directly from patients versus in human tumor cell lines. Anticancer Drugs 2, 11–25 (1991). 10.1097/00001813-199102000-00002 [DOI] [PubMed] [Google Scholar]
  • 2.Nielsen J., Walsh J. T., Degen D. R., Drabek S. M. & McGill J. R. Evidence of gene amplification in the form of double minute chromosomes is frequently observed in lung cancer. Cancer Genetics 65, 120–124 (1993). 10.1016/0165-4608(93)90219-C [DOI] [PubMed] [Google Scholar]
  • 3.Hung K. L. et al. Coordinated inheritance of extrachromosomal DNAs in cancer cells. Nature 635, 201–209 (2024). 10.1038/s41586-024-07861-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Lange J. T. et al. The evolutionary dynamics of extrachromosomal DNA in human cancers. Nat Genet 54, 1527–1533 (2022). 10.1038/s41588-022-01177-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wu S. et al. Circular ecDNA promotes accessible chromatin and high oncogene expression. Nature 575, 699–703 (2019). 10.1038/s41586-019-1763-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Kim H. et al. Mapping extrachromosomal DNA amplifications during cancer progression. Nat Genet 56, 2447–2454 (2024). 10.1038/s41588-024-01949-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Kim H. et al. Extrachromosomal DNA is associated with oncogene amplification and poor outcome across multiple cancers. Nat Genet 52, 891–897 (2020). 10.1038/s41588-020-0678-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Bailey C. et al. Origins and impact of extrachromosomal DNA. Nature 635, 193–200 (2024). 10.1038/s41586-024-08107-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Tang J. et al. Enhancing transcription-replication conflict targets ecDNA-positive cancers. Nature 635, 210–218 (2024). 10.1038/s41586-024-07802-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Deshpande V. et al. Exploring the landscape of focal amplifications in cancer using AmpliconArchitect. Nat Commun 10, 392 (2019). 10.1038/s41467-018-08200-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Alexandrov L. B. et al. The repertoire of mutational signatures in human cancer. Nature 578, 94–101 (2020). 10.1038/s41586-020-1943-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Landi M. T. et al. Tracing Lung Cancer Risk Factors Through Mutational Signatures in Never-Smokers. Am J Epidemiol 190, 962–976 (2021). 10.1093/aje/kwaa234 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Consonni D. et al. Outdoor particulate matter (PM10) exposure and lung cancer risk in the EAGLE study. PLoS One 13, e0203539 (2018). 10.1371/journal.pone.0203539 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Consortium, I. T. P.-C. A. o. W. G. Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020). 10.1038/s41586-020-1969-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lee J. J. et al. Tracing Oncogene Rearrangements in the Mutational History of Lung Adenocarcinoma. Cell 177, 1842–1857 e1821 (2019). 10.1016/j.cell.2019.05.013 [DOI] [PubMed] [Google Scholar]
  • 16.Imielinski M. et al. Mapping the hallmarks of lung adenocarcinoma with massively parallel sequencing. Cell 150, 1107–1120 (2012). 10.1016/j.cell.2012.08.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Lee J. K. et al. Clonal History and Genetic Predictors of Transformation Into Small-Cell Carcinomas From Lung Adenocarcinomas. J Clin Oncol 35, 3065–3074 (2017). 10.1200/JCO.2016.71.9096 [DOI] [PubMed] [Google Scholar]
  • 18.Leong T. L. et al. Deep multi-region whole-genome sequencing reveals heterogeneity and gene-by-environment interactions in treatment-naive, metastatic lung cancer. Oncogene 38, 1661–1675 (2019). 10.1038/s41388-018-0536-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Bergstrom E. N. et al. Mapping clustered mutations in cancer reveals APOBEC3 mutagenesis of ecDNA. Nature 602, 510–517 (2022). 10.1038/s41586-022-04398-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chen J. et al. MYC-driven increases in mitochondrial DNA copy number occur early and persist throughout prostatic cancer progression. JCI Insight 8 (2023). 10.1172/jci.insight.169868 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Bielski C. M. et al. Genome doubling shapes the evolution and prognosis of advanced cancers. Nat Genet 50, 1189–1195 (2018). 10.1038/s41588-018-0165-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.O'Sullivan R. J. & Karlseder J. Telomeres: protecting chromosomes against genome instability. Nat Rev Mol Cell Biol 11, 171–181 (2010). 10.1038/nrm2848 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Shoshani O. et al. Chromothripsis drives the evolution of gene amplification in cancer. Nature 591, 137–141 (2021). 10.1038/s41586-020-03064-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Steele C. D. et al. Signatures of copy number alterations in human cancer. Nature 606, 984–991 (2022). 10.1038/s41586-022-04738-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hadi K. et al. Distinct Classes of Complex Structural Variation Uncovered across Thousands of Cancer Genome Graphs. Cell 183, 197–210 e132 (2020). 10.1016/j.cell.2020.08.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Davies H. et al. HRDetect is a predictor of BRCA1 and BRCA2 deficiency based on mutational signatures. Nat Med 23, 517–525 (2017). 10.1038/nm.4292 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hung K. L. et al. ecDNA hubs drive cooperative intermolecular oncogene expression. Nature 600, 731–736 (2021). 10.1038/s41586-021-04116-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Zhu Y. et al. Oncogenic extrachromosomal DNA functions as mobile enhancers to globally amplify chromosomal transcription. Cancer Cell 39, 694–707.e697 (2021). 10.1016/j.ccell.2021.03.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Chene P. Inhibiting the p53-MDM2 interaction: an important target for cancer therapy. Nat Rev Cancer 3, 102–109 (2003). 10.1038/nrc991 [DOI] [PubMed] [Google Scholar]
  • 30.Nik-Zainal S. et al. The life history of 21 breast cancers. Cell 149, 994–1007 (2012). 10.1016/j.cell.2012.04.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Dharanipragada P. et al. Blocking Genomic Instability Prevents Acquired Resistance to MAPK Inhibitor Therapy in Melanoma. Cancer Discov 13, 880–909 (2023). 10.1158/2159-8290.CD-22-0787 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Talevich E., Shain A. H., Botton T. & Bastian B. C. CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12, e1004873 (2016). 10.1371/journal.pcbi.1004873 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Aran D., Sirota M. & Butte A. J. Systematic pan-cancer analysis of tumour purity. Nat Commun 6, 8971 (2015). 10.1038/ncomms9971 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang T. et al. Deciphering lung adenocarcinoma evolution and the role of LINE-1 retrotransposition. bioRxiv (2025). 10.1101/2025.03.14.643063 [DOI] [Google Scholar]
  • 35.Cortes-Ciriano I. et al. Comprehensive analysis of chromothripsis in 2,658 human cancers using whole-genome sequencing. Nat Genet 52, 331–341 (2020). 10.1038/s41588-019-0576-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Ding Z. et al. Estimating telomere length from whole genome sequence data. Nucleic Acids Res 42, e75 (2014). 10.1093/nar/gku181 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Diaz-Gay M. et al. The mutagenic forces shaping the genomic landscape of lung cancer in never smokers. medRxiv (2024). 10.1101/2024.05.15.24307318 [DOI] [Google Scholar]
  • 38.Khandekar A. et al. Visualizing and exploring patterns of large mutational events with SigProfilerMatrixGenerator. BMC Genomics 24, 469 (2023). 10.1186/s12864-023-09584-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Bergstrom E. N. et al. SigProfilerMatrixGenerator: a tool for visualizing and exploring patterns of small mutational events. BMC Genomics 20, 685 (2019). 10.1186/s12864-019-6041-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Islam S. M. A. et al. Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor. Cell Genom 2, 100179 (2022). 10.1016/j.xgen.2022.100179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Gonzalez-Perez A. et al. IntOGen-mutations identifies cancer drivers across tumor types. Nat Methods 10, 1081–1082 (2013). 10.1038/nmeth.2642 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Muinos F., Martinez-Jimenez F., Pich O., Gonzalez-Perez A. & Lopez-Bigas N. In silico saturation mutagenesis of cancer genes. Nature 596, 428–432 (2021). 10.1038/s41586-021-03771-1 [DOI] [PubMed] [Google Scholar]
  • 43.Chakravarty D. et al. OncoKB: A Precision Oncology Knowledge Base. JCO Precis Oncol 2017 (2017). 10.1200/PO.17.00011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Bailey M. H. et al. Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell 174, 1034–1035 (2018). 10.1016/j.cell.2018.07.034 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Cheng F. et al. Comprehensive characterization of protein-protein interactions perturbed by disease mutations. Nat Genet 53, 342–353 (2021). 10.1038/s41588-020-00774-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Gao T. & Qian J. EnhancerAtlas 2.0: an updated resource with enhancer annotation in 586 tissue/cell types across nine species. Nucleic Acids Res 48, D58–D64 (2020). 10.1093/nar/gkz980 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Reznik E. et al. Mitochondrial DNA copy number variation across human cancers. Elife 5 (2016). 10.7554/eLife.10769 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1

Supplementary Table 1. Co-occurrence of ecDNA genetic cargo in LCINS

Supplementary Table 2. Co-occurrence of ecDNA genetic cargo in LCSS

Supplementary Table 3. Recurrently amplified oncogenes on ecDNA in LCINS across unique genomic loci

Supplementary Table 4. Recurrently amplified oncogenes on ecDNA in LCSS across unique genomic loci

Supplementary Table 5. Additional EGFR-ecDNA detected by modified pipeline

media-1.xlsx (46.8KB, xlsx)
Supplement 2

Supplementary Fig 1. Additional cohort information. Sankey diagrams showing the numbers of samples with clinical variables (stage, metastasis) as well as survival data available for a) LCINS and b) LCSS

Supplementary Fig 2. Distribution of ecDNA harboring lung cancers in LCINS and LCSS by country. Maps showing the worldwide prevalence of lung cancers harboring ecDNA (ecDNA+) by country for the a) LCINS cohort and b) LCSS cohort. Volcano plot of a multivariate one-vs-all logistic regression model for each country, with ecDNA status as the outcome, for c) LCINS and d) LCSS.

Supplementary Fig 3. Analysis of the association between stage, metastasis, and passive smoking with ecDNA status in all LCINS and LCSS. Forest plots of logistic regression model with ecDNA status as the outcome and the variables for stage in a) LCINS and b) LCSS; metastasis in c) LCINS and d) LCSS; and passive smoking in e) LCINS. The total number of samples with information for all variables is indicated above each plot.

Supplementary Fig 4. Analysis of the association between stage, metastasis, and passive smoking with ecDNA status in lung adenocarcinomas. Forest plots of logistic regression model with ecDNA status as the outcome and the variables for stage in a) LCINS and b) LCSS; metastasis in c) LCINS and d) LCSS; and passive smoking in e) LCINS. The total number of samples with information for all variables is indicated above each plot.

Supplementary Fig 5. Genomic features associated with the presence of ecDNA. a) Volcano plot of logistic regression models of genomic features in association with ecDNA status for lung adenocarcinoma in LCINS (left) and lung adenocarcinoma in LCSS (right). In all volcano plots, the x-axes reflect the log2 odds ratio, and the y-axes correspond to the log10 FDR q-value. An FDR q-value threshold of 0.05 is indicated with the dashed red line, an FDR q-value threshold of 0.01 is indicated with the dashed blue line. b) Forest plot of multivariate logistic regression modeling the presence or absence of ecDNA in relation to whole-genome doubling in a pan-cancer dataset comprising four different tumor types from the PCAWG project, adjusted for the covariates of age, purity, sex, and tissue type. All significant associations are indicated in red. c) Bar plot indicating the proportion of samples that contained overlapping chromothripsis and ecDNA regions (red) and the proportion of all samples that did not contain any overlap (gray) for all samples that had both an ecDNA and chromothripsis event detected. d) Bar plot showing the proportion of samples with an EGFR L858R mutation, stratified by ecDNA status.

Supplementary Fig 6. Association of mutational signatures with ecDNA status. Volcano plots of logistic regression models of the activities of Single base substitution (SBS), Indel (ID), copy number (CN) and structural variant (SV) signatures in association with ecDNA in lung cancers of a) LCINS and b) LCSS. The logistic regression models were adjusted for age, sex, ancestry, histology, and purity. In all volcano plots, the x-axes reflect the log2 odds ratio, and the y-axes correspond to the log10 FDR q-value. An FDR q-value threshold of 0.05 is indicated with the dashed red line, an FDR q-value threshold of 0.01 is indicated with the dashed blue line.

Supplementary Fig 7. Genomic content of ecDNA

a) Upset plot indicating the number of ecDNA with oncogenes, immunomodulatory genes, promoter or enhancer regions, other genes that are not oncogenes, or unidentified genomic elements (Unknown) or any combination thereof for LCINS and LCSS. b) Bar plot indicating the proportion of ecDNA containing a particular genomic element in LCINS vs LCSS. c) log2 median copy number (y-axis) for each class of ecDNA (x-axis) in LCINS and LCSS.

Supplementary Fig 8. Highly sensitive pipeline for ecDNA detection that accounts for tumor purity and ploidy.

a) Schematic depicting the methodology for the rescaled pipeline that considers tumor purity and ploidy. The logR track for the region containing the EGFR gene on chr7 and the purity and ploidy computed from Battenberg is indicated in one tumor as example (left). The exact location of the EGFR gene is shown by the dashed red lines. The equation that simultaneously accounts for sample purity and ploidy when computing the total copy number is shown, as well as the total copy number estimate when assuming a sample purity of 1.0 and ploidy of 2.0 (Uncorrected CN) vs. considering the purity and ploidy (Corrected CN). The Sashimi plot indicates the structural variants (colored arcs) and smoothed coverage around the EGFR region, which was annotated as ecDNA by AmpliconClassifier. b) Distribution of tumor purities for samples with an ecDNA annotation for both the default (left) and rescaled (right) pipelines. P-value for a Mann-Whitney test between ecDNA+ and ecDNA samples is indicated. The total number of samples in each group is also indicated. c) The ecDNA prevalence (y-axis), which is the proportion of ecDNA+ samples over the total number of samples, for all possible tumor purity cutoffs applied to samples (x-axis). The total number of samples available for analysis after each tumor purity cutoff is applied is also indicated. d) Gene expression (log2 CPM) of the EGFR gene for all samples with RNA-seq data available. LCINS samples with EGFR-ecDNA are indicated in red.

media-2.pdf (2.3MB, pdf)

Data Availability Statement

Normal and tumor-paired CRAM files for the WGS subjects of the Sherlock-Lung study and the EAGLE study have been deposited in dbGaP under the accession numbers phs001697.v2.p1 and phs002992.v1.p1, respectively.


Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES