Skip to main content
Clinical Epigenetics logoLink to Clinical Epigenetics
. 2025 Jun 8;17:95. doi: 10.1186/s13148-025-01902-3

Mutational landscape and DNA methylation-based classification of squamous cell carcinoma and urothelial carcinoma

Min Ren 1,2,3,#, Midie Xu 1,2,3,#, Chen Chen 1,2,3,#, Ran Wei 1,2,3, Qianlan Yao 1,2,3, Liqing Jia 1,2,3, Peng Qi 1,2,3, Qifeng Wang 1,2,3, Qianming Bai 1,2,3, Xiaoli Zhu 1,2,3, Sheng Wu 4, Qinghua Xu 4, Xiaoyan Zhou 1,2,3,
PMCID: PMC12147257  PMID: 40484973

Abstract

Background

Identification of the tissue of origin is fundamental for cancer treatment. However, squamous cell carcinomas from different sites lack representative histological and immunohistochemical features. This study aimed to identify mutational profiles and further establish a DNA methylation-based classification for squamous cell carcinoma and urothelial carcinoma. Samples of unambiguous squamous cell carcinomas and urothelial carcinomas were collected for targeted next-generation sequencing and mutational landscape analysis. Moreover, using Illumina methylation BeadChip data from public datasets and a local cohort, we developed a DNA methylation-based classifier utilizing the CatBoost algorithm to identify four common types of squamous cell carcinoma (lung, head and neck, esophagus, and cervix) as well as urothelial carcinoma.

Results

The DNA mutational profiles of squamous cell carcinomas from different sites overlapped greatly, and there was no significant difference in tumor mutation burden or microsatellite status. On the basis of public datasets and analyses via various machine learning algorithms, a DNA methylation-based classification containing 106 features by the CatBoost algorithm was constructed and reached an accuracy of 98.79% (490/496) in the training set from PanCanAtlas datasets. The predictive accuracies of the methylation classification in the public validation set and local FUSCC validation set 1 with known primary were 86.96% (340/391) and 84.87% (101/119), respectively. The predictive accuracy for the primary samples (89.66%, 78/87) was obviously greater than that for the metastatic samples (71.88%, 23/32). FUSCC validation set 2 included ten complicated cancer of unknown primary (CUP) samples with squamous cell differentiation. When a well-established 90-gene expression assay was compared with the present classification, our methylation-based classification successfully classified two samples with no eligible RNA expression; the results for four sample were consistent with higher methylation prediction scores in three, and those for two samples were inconsistent. The methylation-based classification results of the remaining two samples were more compatible with the results of the clinical evaluation.

Conclusion

We successfully established a DNA methylation-based classification for squamous cell carcinomas (lung, head and neck, esophagus, and cervix) and urothelial carcinomas with outstanding diagnostic performance for the first time. This classification has high potential for clinical translation to address the dilemma of identifying the origin of squamous cell carcinoma of unknown primary.

Supplementary Information

The online version contains supplementary material available at 10.1186/s13148-025-01902-3.

Keywords: Cancer of unknown primary, Squamous cell carcinoma, DNA methylation, Machine learning, Mutational landscape

Introduction

Cancer of unknown primary (CUP) represents a biologically heterogeneous group of metastatic malignant tumors without identifiable primary sites. The biological behavior of CUP is highly aggressive, and the prognosis is extremely dismal, with a median overall survival of less than one year [1]. Several clinical trials have revealed that, compared with empirical therapy, molecular-guided site-specific treatment has superior therapeutic value [26], indicating the clinical significance of identifying the tissue-of-origin. For genuine CUP, in which the primary tumor is not detected via routine physical, hematological and radiologic examinations, pathological assessment is pivotal to provide diagnostic clues, including morphological and immunohistochemical (IHC) examinations. According to the tumor’s morphological characteristics, CUPs can be preliminarily classified as highly/moderately differentiated adenocarcinoma (50%), poorly differentiated carcinoma/adenocarcinoma (30%), squamous cell carcinoma (15%), or undifferentiated tumor (5%) [1]. Although the most common histological type is adenocarcinoma, there are morphological differences among adenocarcinomas from different sites, and many available tissue-specific IHC markers. However, squamous cell carcinomas from different origins share identical morphological features, and no effective tissue-specific IHC markers are available. In addition, the treatment for squamous cell carcinoma is principally based on the anatomical site; therefore, identification of the origin of metastatic/multiple squamous cell carcinomas is crucial in present-day clinical practice. The results of our previous study on CUP also revealed that the proportion of squamous cell carcinoma cases in China was relatively greater than that reported in the West [7], reaching 20%–30% [8, 9]. The tissue-of-origin classifications in our previous study and the literature also indicated that the predictive accuracy for squamous cell carcinoma was relatively inferior to that for adenocarcinoma [1012]. In addition, urothelial carcinomas often exhibit marked squamous cell differentiation during invasion and metastasis, which is difficult to distinguish via morphological and IHC analyses. Therefore, a classification system for the tissue-of-origin in common squamous cell carcinoma and urothelial carcinoma is of clinical importance.

DNA methylation, a critical epigenetic modification, is associated with the regulation of gene expression. Increasing evidence suggests that different tissues and cell types under various normal and pathological conditions exhibit distinct methylation patterns [13]. Previous studies have demonstrated the significant diagnostic value of methylation profiles in central nervous system tumors, bone and soft tissue tumors, and various other tumor types [11, 1416], suggesting that DNA methylation profiling, with its superior tissue-specific performance, could help address the challenge of distinguishing the origins of squamous cell carcinomas from different sites.

Therefore, in the present study, we first described and compared the DNA mutational landscape of squamous cell carcinomas from different sites and urothelial carcinomas using next-generation sequencing (NGS). Given the similar mutational profiles of squamous cell carcinomas and urothelial carcinomas reported in the literature, we further established a DNA methylation-based classification. An effective classification for squamous cell carcinoma could significantly enhance precise diagnosis, ultimately improving the treatment and prognosis of CUPs.

Materials and methods

Patients

Samples of primary and metastatic squamous cell carcinomas of the lung, vulva, head and neck, esophagus, and cervix and samples of urothelial carcinomas from patients diagnosed at the Fudan University Shanghai Cancer Center (FUSCC) between May 2016 and December 2023 were retrospectively collected as FUSCC validation set 1. Additionally, ten complicated CUP patients who had undergone a 90-gene expression assay [8, 10] and had available tumor tissue were included in FUSCC validation cohort 2. All patients were diagnosed by at least two professional pathologists and assessed for sufficient tumor tissue for subsequent next-generation sequencing (NGS) or methylation BeadChip testing. The study protocol was reviewed and approved by the Institutional Ethics Committee of FUSCC.

Next-generation sequencing

  1. Tissue DNA extraction: DNA was extracted from formalin-fixed paraffin-embedded (FFPE) tissue using the QIAamp DNA FFPE Tissue Kit (Qiagen, Hilden, Germany) according to the manufacturer’s instructions. The DNA concentration was measured via a Qubit dsDNA assay. The quantification of tissue DNA was performed using a Qubit 2.0 fluorometer and a double-stranded DNA HS assay kit (Life Technologies, USA).

  2. Capture-based targeted DNA sequencing: A minimum of 30 ng of DNA was required for NGS library construction. The tissue DNA was sonicated using an M220 ultrasonicator (Covaris, MA, USA), followed by end repair, phosphorylation, adaptor ligation, and purification of fragments with sizes between 200 and 400 base pairs. Target capture was performed using a commercial panel consisting of 520 cancer-related genes spanning 1.86 Mb of the human genome (OncoScreen Plus, Burning Rock Biotech) (gene list detailed in Additional file 1: Table S1). The quality and size of the fragments were assessed with a high-sensitivity DNA kit via a Bioanalyzer 2100 (Agilent Technologies, CA, USA). Indexed samples were sequenced on a NextSeq500 sequencer (Illumina, Inc., California, USA) with 150-base pair read lengths and a target sequencing depth of 1000 × for tissue samples.

  3. Sequence Data Analysis: Sequences were mapped to the reference human genome (hg19) via the Burrows–Wheeler Aligner (version 0.7.10). Local alignment optimization duplication marking and variant calling were performed via the Genome Analysis Tool Kit (version 3.2) and VarScan (version 2.4.3). Tissue samples were compared against paired white blood cells to eliminate most clonal haematopoiesis-related variants (detailed in Additional file 2). Somatic mutations were assigned to COSMIC v3 mutational signatures [17] using the Mutational Patterns framework [18].

DNA methylation-based classification

  1. DNA methylation microarray experiment: All the local samples were analyzed with commercially available Illumina Infinium EPIC or EPIC_v2.0 BeadChip arrays. All procedures were performed according to the Infinium HD Methylation Assay Protocol [19]. Briefly, at least 500 ng of sample DNA was subjected to bisulfite conversion (EZ DNA Methylation-GoldTM Kit-D5005). Using the Illumina Infinium Methylation Assay Kit, bisulfite-converted DNA was amplified, incubated, fragmented, precipitated, resuspended, and hybridized to a BeadChip. Finally, the hybridized BeadChip was further washed, extended, stained and scanned with a methylation microarray scanner (Illumina NextSeq™ 550).

  2. Establishment of a methylation classification
    • 2.1.
      Data source: The Pan-Cancer Atlas (PanCanAtlas) dataset includes data on molecular changes at the DNA, RNA, proteomic, and epigenetic levels for 33 human tumor types from The Cancer Genome Atlas (TCGA) Program (Download link: https://api.gdc.cancer.gov/data/99b0c493-9e94-4d99-af9f-151e46bab989) [20]. For this study, DNA methylation data for 1651 cases, including lung squamous cell carcinoma (LUSC, 364 cases), cervical squamous cell cancer (CESC, 195 cases), head and neck squamous cell carcinoma (HNSCC, 523 cases), esophageal squamous cell carcinoma (ESCC, 154 cases) and bladder urothelial carcinoma (BLCA, 415 cases), were downloaded from PanCanAtlas as training set. The validation sets included three parts: a public validation set, FUSCC validation set 1, and FUSCC validation set 2. The public validation set consisted of data on 107 cases of CESC (CGCI-HTMCP-CC), 108cases of LUSC (CPTAC-3), 39 cases of ESCC (GSE178212, 24 cases; GSE121930, 15 cases), 120 cases of HNSCC (GSE178216, 15 cases; GSE178218, 20 cases; GSE178219, 32 cases; E-MTAB-10576, 53 cases) and 17 cases of BLCA (GSE222933). The FUSCC validation sets included tumors from local Chinese patients at FUSCC, comprising set 1 which consisted of 119 cases with definite pathological diagnoses, and set 2 which included ten CUPs exhibiting squamous cell differentiation.
    • 2.2.
      Data preprocessing: All the data were analyzed according to beta signal values. For the training set from the PanCanAtlas dataset (based on 450 K BeadChip platform), the champ.filter method in the R package ChAMP (v2.29.1) was used for preprocessing to filter low-quality markers. The rejection criteria included the following: (1) markers with missing beta signal values; (2) markers with non-CpG check points; (3) markers associated with single-nucleotide polymorphisms (SNPs); (4) markers mapped to multiple locations; and (5) markers on sex chromosomes. Using the champ.norm method of the R package ChAMP (v2.29.1), the data were corrected and standardized by BMIQ (Beta Mixture Quantile dilation). Because that the ChAMP package showed limited compatibility with latest EPIC and EPIC v2.0 platforms, the openSesame flow in the R package sesame (v1.20.0) was used for preprocessing for the validation sets from public dataset and local FUSCC cohort.
    • 2.3
      Specific methylation marker screening: The five tumor types in the training set were divided into 15 groups according to two comparison methods: 1 vs. 1 (the comparison of two tumor types, a total of 10 groups) and 1 vs. all (the comparison of a single tumor type with other tumor types, a total of 5 groups). For these 15 cohorts, Limma differential analysis and receiver operating characteristic (ROC) curve analysis were performed. In these two analyses, specific methylation markers satisfying two conditions were screened: (1) markers with a false discovery rate (FDR) less than 0.05 in the Limma differential analysis which were defined as differentially methylated positions (DMPs) and (2) the top 20 markers with the highest area under the curve (AUC). Since the methylation markers contained in different datasets were different, only the markers common to these three Illumina BeadChip platforms including 450 K, EPIC and EPIC v2.0 were included in the screening. The overall number of screened markers was 396,065, and remained 324,873 markers after quality control. The DMP method by Limma differential analysis was achieved using the R package ChAMP (v2.29.1), and ROC analysis was performed via the ROC method in the R package pROC (v1.18.5).
    • 2.4.
      Machine learning model construction and evaluation: To streamline classifier construction, for specific methylation markers that satisfied the conditions, the 50% of the screened markers that contributed the most to the model were reserved. Feature selection was performed through the SelectFromModel method in the Python library scikit-learn (v1.2.2), and the Light Gradient Boosting Machine (LightGBM) classifier was used to estimate feature weights. Among the 16 machine learning models tested (detailed in Fig. 3A), the categorical boosting (CatBoost) algorithm outperformed the other classifiers and exhibited excellent predictive performance. Thus, in this study, the CatBoost model was used to classify the five tumor types in the training set. Classifier performance was evaluated via indicators such as accuracy, AUC, recall, precision, F1 score, Cohen’s kappa coefficient, and the Matthews correlation coefficient. The PanCanAtlas dataset was divided into a training set 1 (1155 samples) and set 2 (496 samples) at a 7:3 ratio. Specifically, 70% of the samples (training set 1) were used to train the model, while the remaining 30% (training set 2) served as an internal validation set to evaluate the model’s performance. The model was constructed using the create_model method and the default parameters from the Python library PyCaret (v3.0.4), and 10 × cross-validation was used for model training and evaluation (Fig. 1). The accuracy of the methylation classification was evaluated according to the clinicopathological diagnosis. Furthermore, DNA methylation data of the adjacent normal tissue from the TCGA were also downloaded and processed with above-mentioned methods, in order to evaluate the performance of the established methylation classification in corresponding normal tissues.

Fig. 3.

Fig. 3

Establishment of the DNA methylation-based classification: comparison of the performance of 16 machine learning algorithms A. The accuracy of the number of features included in the classification among the training set 1 (1155 samples) and set 2 (496 samples). Among total 212 specific methylation features, the top 106 features had the best predictive performance in both groups B

Fig. 1.

Fig. 1

The workflow for establishing and validating the DNA methylation-based classification

Results

DNA mutational landscape of squamous cell carcinomas and urothelial carcinomas

A total of 26 cases of CESC, 20 cases of HNSCC, 24 cases of vulvar squamous cell carcinoma (VSCC), 44 cases of LUSC, 61 cases of ESCC, and 15 cases of BLCA were subjected to targeted NGS testing which included 520 cancer-related genes. All the samples were eligible for analysis, and the DNA mutation heatmaps showing the top 15 genes of these six tumor types were detailed in Fig. 2 (the complete mutation heatmaps were shown in Additional file 3: Fig S1–S7). The mean tumor mutation burden (TMB) values in patients with CESC, HNSCC, VSCC, LUSC, ESCC and BLCA were 6.67, 4.59, 6.52, 10.07, 6.71, and 13.56 mutations/Mb, respectively, with no significant difference (P > 0.05). The results of microsatellite status (MSI) analysis revealed that only one case of ESCC, CESC or LUSC exhibited the microsatellite instability-high (MSI-H) phenotype, while the remaining cases were microsatellite stable (MSS) (Table 1).

Fig. 2.

Fig. 2

Landscape of genetic alterations including the top 15 genes across six types of squamous cell carcinoma: all samples A, CESC B, HNSCC C, ESCC D, LUSC E, VSCC F, and BLCA G. For the each of the heatmap, the bar chart above the heatmap represented the total number of mutations for each sample, while the right side listed gene names in descending order of mutation frequency. On the left, the mutation frequency of each gene was displayed. The different colors in the heatmap represented different mutation types (as indicated in the legend on the right side of the heatmap). Each column represented one sample, and samples were ordered by the frequency of mutation

Table 1.

Status of TMB, MSI, and COSMIC signatures enriched in different types of squamous cell carcinoma and BLCA

Tumor type Average TMB (mutations/Mb) Frequency of MSI-H COSMIC signature
CESC 6.67 3.8% (1/26)

APOBEC cytidine deaminase (C > T)

Defective DNA mismatch repair

HNSCC 4.59 0% (0/20) APOBEC cytidine deaminase (C > T)
VSCC 6.52 0% (0/24)

Defective DNA mismatch repair

APOBEC cytidine deaminase (C > T)

LUSC 10.07 2.3% (1/44) No biologically significant signature
ESCC 6.71 1.6% (1/61) No biologically significant signature
BLCA 13.56 0% (0/15) Not available

Due to the limited data on urothelial carcinoma, only the sequencing data from five different sites of squamous cell carcinomas were further compared with 30 mutation signatures from the COSMIC database, based on 96 distinct mutation spectra. The results revealed that APOBEC cytidine deaminase (C > T) and defective DNA mismatch repair signatures were identified in CESC and VSCC; APOBEC cytidine deaminase (C > T) was identified in HNSCC; and the SBS5 signature was identified in LUSC and ESCC, with no definite biological significance (Table 1). The results of the COSMIC signature analysis for different sites of squamous cell carcinomas were largely similar, primarily involving APOBEC cytidine deaminase (C > T) and defective DNA mismatch repair signatures, making differentiation between them difficult.

Overall, the mutational profiles of squamous cell carcinomas from different sites and urothelial carcinomas overlapped greatly, commonly involving genes in the cell cycle pathway (such as TP53, CDKN2 A/B, CCND1, and RB1), the RAS and AKT signaling pathways (such as PIK3CA, PTEN, and FGFR1/2/3), and the squamous differentiation pathway (such as NOTCH1/2 and TP63). Frequent mutation of KMT2 C/D, which is involved in the chromosome remodeling signaling pathway, was identified not only in squamous cell carcinomas but also in urothelial carcinomas, which was consistent with the results of the previous studies [21]. Furthermore, there was no significant difference in the TMB or COSMIC signature among the squamous cell carcinomas and urothelial carcinomas, and the overall frequency of MSI was extremely low.

Above all, these mutational signatures were unable to accurately distinguish different types of squamous cell carcinoma and urothelial carcinoma. In view of the excellent performance of DNA methylation in determining the tissue-of-origin, the present study further established a methylation-based classification for squamous cell carcinomas and urothelial carcinomas.

DNA methylation-based classification of squamous cell carcinomas and urothelial carcinomas

Data source and type

The training set was derived from the PanCanAtlas dataset, which contained the DNA methylation data of 195 CESCs, 364 LUSCs, 154 ESCCs, 523 HNSCCs, and 415 BLCAs on the Illumina HumanMethylation450 BeadChip chip platform. The public validation set was derived from the GEO, TCGA, and ArrayExpress public databases. The samples from the public databases were all primary tumors. The FUSCC validation set 1 was composed of 119 samples from local primary and metastatic samples with known primary, and FUSCC validation 2 included ten CUPs. The specific number, source and BeadChip platform of the training and three validation sets were detailed in Table 2.

Table 2.

Detailed information on the source, type of Illumina BeadChip and preprocessing method of the training set and three validation sets

Data group Tumor entity (N) Dataset ID (N) Data source Chip type Preprocessing method
Training set Five tumor types (1651) PanCanAtlas TCGA 450 K Champ.filter
Public validation set CESC (107) CGCI-HTMCP-CC (107) CGCI EPIC OpenSesame flow
LUSC (108) CPTAC-3 (108) CPTAC EPIC
ESCC (39) GSE178212 (24) GEO 450 K
GSE121930 (15) GEO
HNSCC (120) GSE178216 (15) GEO 450 K
GSE178218 (20) GEO
GSE178219 (32) GEO
E-MTAB-10576 (53) ArrayExpress
LUSC (17) GSE222933 (17) GEO EPIC
FUSCC validation set1 Five tumor types (119) FUSCC

EPIC (45)

EPIC v2.0 (74)

OpenSesame flow
FUSCC validation set2 CUP (10) FUSCC EPIC v2.0 OpenSesame flow

Establishment of a DNA methylation-based classification

With the goal of developing a DNA methylation-based classification with optimal predictive performance, the performances of 16 machine learning algorithms commonly used were compared in this study. The results showed that CatBoost outperformed the other algorithms, so it was ultimately used to establish the classification model (Fig. 3A). After screening the top 20 features with an FDR < 0.05 and the highest AUC values, a total of 212 specific features were obtained. When the number of features included in the classification was further evaluated, the top 106 features had the best predictive performance (Fig. 3B). The degree of methylation among the five tumors was significantly different (Fig. 4). This information can be used to further explore the biological function of highly informative hyper or hypomethylation markers for specific cancer types.

Fig. 4.

Fig. 4

Unsupervised hierarchical clustering and heatmaps of included 106 features in the training set A, public validation set B, and FUSCC validation set 1 C. The features in two validation sets were ordered according to clustering on the training set. The tumor types and sample types were displayed in the different colors indicated in the figure legends. The standardized methylation values were displayed from low (blue) to high (red). Each column represented one sample

The classification based on the CatBoost algorithm had excellent diagnostic performance in the training set 2 from PanCanAtlas dataset including 496 samples. On the basis of the referenced pathological diagnoses, the AUCs for BLCA, CESC, ESCC, HNSCC and LUSC were 0.995, 0.998, 0.969, 0.991 and 0.991, respectively. The overall predictive accuracy reached 98.79% (490/496) (Fig. 5A).

Fig. 5.

Fig. 5

The performance of the classification in the training set A, public validation set B, FUSCC validation set 1 (primary samples) C, and FUSCC validation set 1 (metastatic samples) D, including the receiver operating characteristic (ROC) curve, confusion matrix, and recall curve

Performance validation of the methylation classification

In the public validation set, the predictive accuracies of methylation classification in BLCA, CESC, ESCC, HNSCC, and LUSC were 94.12% (16/17), 92.52% (99/107), 89.74% (35/39), 73.33% (88/120), and 94.44% (102/108), respectively, with an overall accuracy of 86.96% (340/391) (Fig. 5B). In addition, in the public data of adjacent normal tissue, the overall predictive accuracy reached as high as 95.7% (111/116), with 50.0% (1/2) in CESC, 94.0% (47/50) in HNSCC, 100.0% (41/41) in LUSC, 50.0% (1/2) in ESCC, and 100.0% (21/21) in BLCA (Additional file 4: Figure S8). Because of the extremely limited adjacent normal tissues of ESCC and CESC in the TCGA, we will further investigate the diagnostic value of the present methylation-based classification in corresponding normal tissue.

Moreover, we assessed the performance of methylation classification in FUSCC validation set 1 from the local FUSCC cohort with confirmed origins. Compared with that of the histological diagnosis, the overall accuracy of our classification in FUSCC validation set 1 was 84.87% (101/119) (Table 3). The predictive accuracy for the primary samples (89.66%, 78/87) (Fig. 5C) was obviously greater than that for the metastatic samples (71.87%, 23/32) (Fig. 5D). The reasons for this discrepancy were likely that five of the nine mispredicted metastatic samples had a tumor cell content of ≤ 30%. In addition, the histological diagnosis of two patients (Sample ID: ESCC-21, ESCC-23) with squamous cell carcinoma in the lung with a history of ESCC tended toward metastatic ESCC, whereas the methylation classification indicated a diagnosis of LUSC. Given that the lung lesions in these two patients were solitary, it was also possible that they may constitute a second primary LUSC. Furthermore, the average prediction score for the matched samples was 0.76, which was significantly higher than that of the mismatched samples (0.51), suggesting that the results of the classification for samples with lower prediction scores should be interpreted with caution and evaluated in combination with clinical and pathological information. The detailed information and predictive results of the 119 samples in FUSCC validation set 1 were listed in Additional file 5: Table S2. In addition, we have uploaded the original methylation data of the local FUSCC validation cohort to the National Genomics Data Center (NGDC) public database.

Table 3.

The accuracy of the DNA methylation-based classification in the public validation set and FUSCC validation set 1

Group Tumor type Accuracy
Public validation set BLCA 94.12% (16/17)
CESC 92.52% (99/107)
ESCC 89.74% (35/39)
HNSCC 73.33% (88/120)
LUSC 94.44% (102/108)
Total 86.96% (340/391)
FUSCC validation set 1 BLCA 95.83% (23/24)
CESC 73.91% (17/23)
ESCC 92.59% (25/27)
HNSCC 81.81% (18/22)
LUSC 78.26% (18/23)
Total 84.87% (101/119)

Bolded data emphasize the overall accuracy across the two cohorts

Finally, the performance of the methylation classification was investigated with validation set 2, a real-life cohort of ten complicated CUPs with squamous cell differentiation (Table 4). These ten CUPs were selected from patients who had undergone a 90-gene expression assay for tissue-of-origin identification on the basis of RNA expression, but the primary tumor site remained unclear. The key clinicopathological features, and predictive results of the 90-gene expression assay and the present methylation classification of these ten samples were shown in Table 4. First, DNA methylation testing successfully classified two samples whose quality control were ineligible for detecting RNA expression (CUP-2 and CUP-3). Furthermore, when the results of the methylation classification were compared with those of the 90-gene expression assay, which has been validated with a large number of cases and in clinical trials [5, 22], the consistency in the predicted primary sites for four cases (CUP-1, CUP-5, CUP-6, and CUP-10) indicated the reliability of the present methylation classification. Notably, the methylation classifier provided stronger evidence with a higher prediction score in three cases (CUP-5, CUP-6, and CUP-10). However, the inconsistent results for CUP4 and CUP7 indicated the need for additional information for validation of the tissue-of-origin during patient follow-up. For the predicted primary sites of CUP-8 and CUP-9, the final diagnosis of CUP-8 was lung metastasis of BLCA given the history of BLCA and partial positivity for GATA3, which was consistent with the result of our methylation classification. Considering that the lesions were concentrated in the head and neck region and negative cervical biopsy, the suspected primary lesion of CUP-9 was ultimately identified as HNSCC, which was supported by the methylation classification. The results of CUP-8 and CUP-9 further validated the accuracy of the present methylation classification for primary identification in CUP.

Table 4.

Clinical and molecular testing information for ten CUP samples in FUSCC validation set 2

Sample ID Sex Biopsy site 90-gene assay Methylation classification Pivotal clinical information
CUP-1 Male Cervical lymph node HNSCC HNSCC No finding by nasopharyngeal and oropharyngeal examination
CUP-2 Male Inguinal lymph node Failure ESCC Head and neck, and gastroscopy were negative
CUP-3 Female Axillary lymph node Failure HNSCC Head and neck, and gastroscopy were negative
CUP-4 Female Abdominopelvic cavity CESC ESCC Cervical biopsy and gastroscopy were negative
CUP-5 Male Lumbar vertebrae LUSC (similarity score ≤ 45) LUSC PET scan shows multiple lesions in the lung
CUP-6 Male Cervical lymph node HNSCC (similarity score ≤ 45) HNSCC No finding in the head and neck region
CUP-7 Female Axillary lymph node CESC ESCC Cervical biopsy and gastroscopy were negative
CUP-8 Male Lung LUSC BLCA A history of BLCA
CUP-9 Female Supraclavicular lymph node CESC HNSCC Cervical biopsy was negative. Lesions were concentrated in the head and neck
CUP-10 Male Abdominal cavity BLCA (similarity score ≤ 45) BLCA Lesions were concentrated in the abdominal cavity

Bolded values emphasize the samples with consistent results between the two assays

Discussion

Although compared with empirical chemotherapy, molecular-guided site-specific treatment significantly improves the prognosis of patients with CUP [3, 23], identifying the source of CUP remains a technological challenge for modern cancer medicine. As reported in the literature and in our previous study, the proportion of CUPs with squamous cell carcinoma in China is greater than that in the Western population [9]. In terms of traditional histopathology and immunohistochemistry, squamous cell carcinomas from different sites are not different and cannot be accurately distinguished in clinical practice. Previous studies have established many multisite or targeted tissue-of-origin classifications based on different molecular platforms [24]. However, few studies have shown that the predictive accuracy of multisite classification for squamous cell carcinoma is significantly inferior to that for adenocarcinoma [1012]. To date, large-scale studies specifically targeting the origin of squamous cell carcinomas from different sites are still lacking. Previous studies have focused primarily on the origin of squamous cell carcinoma in the lungs of patients with a history of HNSCC [2530], whereas the included tumor types and the number of samples from China are very limited. Therefore, this study included the most common and clinically challenging-to-differentiate squamous cell carcinomas—those originating from the lung, head and neck, esophagus, and cervix—along with urothelial carcinoma, and provided a comprehensive description of the mutational profiles of these cancers. For the first time, a DNA methylation-based classification for squamous cell carcinoma was established and further validated in real CUP cases, aiming to address the clinical dilemma of the differential diagnosis of squamous cell carcinoma.

We initially expected to discover specific DNA mutational profiles of different squamous cell carcinomas and urothelial carcinomas. However, targeted NGS analysis indicated that the mutational profiles of squamous cell carcinomas and urothelial carcinomas overlap greatly [3135], which is consistent with the results of previous studies reporting that squamous cell carcinomas from different sites exhibit similar gene mutation profiles [20, 36]. The results of the COSMIC signature analysis also revealed that the signature types in squamous cell carcinomas were mostly APOBEC cytidine deaminase (C > T) and defective DNA mismatch repair, and there was no distinctive signature for differentiating tumor types. In addition, analysis of TMB and MSI across five tumors revealed that TMB was relatively high in LUSC and BLCA and relatively low in VUSC, but the differences were not statistically significant. Most of these tumors were MSS, with only a small proportion showing MSI-H, making MSI status an unreliable diagnostic marker.

Therefore, this study further aimed to establish a tissue-of-origin classification for squamous cell carcinoma and urothelial carcinoma on the basis of DNA methylation analysis. The application of DNA methylation analysis in clinical practice is becoming increasingly widespread and promising, with involvement in tumor screening, diagnosis, treatment, and prognosis evaluation [37]. In terms of tumor diagnosis and classification, the team at Heidelberg University Hospital in 2018 established 91 methylation classes for central nervous system tumors. This classification system was validated using over 1100 tumor samples, and demonstrated a high concordance rate of 88% [14]. In 2021, this team further established a methylation classification model for 62 types of soft tissue and bone sarcomas. After the results of methylation classification and histological diagnosis were compared, the diagnoses of 29 patients (29/428, 7%) were revised, suggesting the substantial diagnostic impact of methylation classification [15]. In an early pivotal study by Moran et al. in 2016, they established a tumor type classifier (EPICUP) based on DNA methylation profiles, which could predict the tissue-of-origin in 188 (87%) of 216 patients with CUP, and the application of type-specific therapy subsequently improved prognosis [4]. Recent studies have shown that various DNA methylation-based classifications achieve accuracies of 80%–90% in both known malignancies and CUP samples [12, 16, 3841]. The results of the present study further confirmed the high accuracy of DNA methylation-based classification in distinguishing different squamous cell carcinomas and urothelial carcinomas. This classification was established not only based on the public databases from Western countries but also on data from a large number of local primary and metastatic samples in China, with an accuracy of 83.50%. The accuracy of the primary samples was significantly greater than that of the metastatic samples (89.02% vs. 61.9%). This discrepancy may be due to the lower tumor content in some metastatic samples, which could be affected by surrounding tissues. Samples with lower tumor contents (≤ 30%) should be enriched or microdissected when scraping tumor tissues. In addition, there were two squamous cell carcinomas in the lungs of patients with a history of ESCC, and the methylation classification predicted these as primary LUSC. Since this classification was not used at the time of initial histological evaluation, the ultimate diagnoses of these two cases may have changed on the basis of histological, molecular and clinical assessments. Moreover, the analysis of the prediction scores revealed that the scores for correctly predicted samples were significantly higher than those for incorrectly predicted samples (0.80 vs. 0.51), suggesting that classifications with low prediction scores (≤ 0.50) should be evaluated in conjunction with additional clinical and pathological information for a more comprehensive assessment. Compared with DNA mutation- or RNA expression-based assays, DNA methylation testing is considered more applicable, reliable and targeted in clinical practice for identifying the tissue-of-origin in CUP patients with squamous cell differentiation. Given the limited number of CUP cases in the present study, our team will further validate the diagnostic value of this classification for CUP in future clinical practice.

This methylation-based classification was developed by evaluating the performance of various machine learning algorithms. The final results revealed that CatBoost outperformed the other classifiers, with an overall accuracy of 98.10%. This algorithm can accurately handle the interaction between features, minimize overfitting, and effectively improve the predictive performance. The present classification included only 106 markers that can be detected via quantitative PCR or targeted NGS platforms, making it potentially suitable for translation into clinical practice to assist in diagnosing squamous cell carcinomas of unknown primary.

Conclusions

In conclusion, the DNA mutational profiles of squamous cell carcinomas from different sites significantly overlapped, with no notable differences in TMB, MSI status, or COSMIC signatures. Most importantly, this study successfully established a DNA methylation-based classification for common squamous cell carcinomas and urothelial carcinomas for the first time. A comparison of various machine learning methods revealed that CatBoost outperformed the other classifiers. On the basis of data from public and local datasets, the DNA methylation-based classification could effectively distinguish these five tumors and CUPs with squamous cell differentiation. Therefore, the classification has significant clinical value in identifying the origin of squamous cell carcinomas of unknown primary, which will further improve the treatment and prognosis of these patients.

Electronic supplementary material

Table S1 (14.6KB, xlsx)

List of the 520 gene included in the next-generation sequencing panel.

Supplementary Material 1

Supplementary materials and methods (15.6KB, docx)

Data analysis of next-generation sequencing.

Supplementary Material 2

Complete mutation heatmaps across six tumor types (29.5MB, zip)

All samples, CESC, HNSCC, ESCC, LUSC, VSCC, and BLCA. The bar chart above the heatmap represented the total number of mutations for each sample, while the right side listed gene names in descending order of mutation frequency. On the left, the mutation frequency of each gene was displayed. The central heatmap showed the distribution of mutations across all samples, with different colors representing different mutation types. Each column represented one sample.

Supplementary Material 3

Figure S8 (419.7KB, jpg)

The performance of the classification in the adjacent normal tissues from TCGA, including the receiver operating characteristiccurve, confusion matrix, and recall curve.

Supplementary Material 4

Table S2 (14.6KB, xlsx)

Detailed information and predictive results of 119 samples in FUSCC validation set 1.

Supplementary Material 5

Acknowledgements

The authors thank the patients for their willingness to cooperate with our study.

Abbreviations

CUP

Cancer of unknown primary

CESC

Cervical squamous cell cancer

HNSCC

Head and neck squamous cell carcinoma

VSCC

Vulva squamous cell carcinoma

LUSC

Lung squamous cell carcinoma

ESCC

61 Cases of esophageal squamous cell carcinoma

BLCA

Bladder urothelial carcinoma

NGS

Next generation sequencing

ROC

Receiver operating characteristic

AUC

Area under the curve

TMB

Tumor mutation burden

MSI

Microsatellite instability

FUSCC

Fudan university shanghai cancer center

TCGA

The cancer genome atlas

Author contributions

M. R.: edited the manuscript and analyzed the data; M. X.: provided some samples; C. C. and L. J.: performed the experiments; R. W., Q. Y. and S.W.: data analysis; P. Q. and Q. W.: provided part of data; Q. B. and X. Z.: provided pivotal opinions about the study; Q. X.: gave valuable insight to the study concept; X. Z.: conceived and designed the study and revised the paper.

Funding

This work was supported by the Innovation Group Project of Shanghai Municipal Health Commission [Project No. 2019 CXJQ03], the Shanghai Science and technology development fund [Project No. 19MC1911000], the Shanghai Municipal Key Clinical Specialty [Project No.shslczdzk01301], and Innovation Program of Shanghai Science and Technology Committee [Project No.20Z11900300].

Availability of data and materials

The dataset used and analyzed in the present study are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

All methods were carried out in accordance with relevant guidelines and regulations, and all experimental protocols were approved by the Institutional Ethics Committee of Fudan University Shanghai Cancer Center. The study was reported in accordance with ARRIVE guidelines.

Consent for publication

Consent to publish has been obtained from the participants.

Competing interests

SW and QHX are employees of Canhelp Genomics. No other potential competing interests were disclosed by the author.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Min Ren, Midie Xu and Chen Chen have contributed equally to this work.

References

  • 1.Rassy E, Pavlidis N. Progress in refining the clinical management of cancer of unknown primary in the molecular era. Nature reviews Clin Oncol. 2020;17(9):541–54. [DOI] [PubMed] [Google Scholar]
  • 2.Hainsworth JD, Rubin MS, Spigel DR, Boccia RV, Raby S, Quinn R, et al. Molecular gene expression profiling to predict the tissue of origin and direct site-specific therapy in patients with carcinoma of unknown primary site: a prospective trial of the Sarah Cannon research institute. J Clin Oncol. 2013;31(2):217–23. [DOI] [PubMed] [Google Scholar]
  • 3.Rassy E, Bakouny Z, Choueiri TK, Van Allen EM, Fizazi K, Greco FA, et al. The role of site-specific therapy for cancers of unknown of primary: a meta-analysis. Eur J Cancer. 2020;127:118–22. [DOI] [PubMed] [Google Scholar]
  • 4.Moran S, Martínez-Cardús A, Sayols S, Musulén E, Balañá C, Estival-Gonzalez A, et al. Epigenetic profiling to classify cancer of unknown primary: a multicentre, retrospective analysis. Lancet Oncol. 2016;17(10):1386–95. [DOI] [PubMed] [Google Scholar]
  • 5.Liu X, Zhang X, Jiang S, Mo M, Wang Q, Wang Y, et al. Site-specific therapy guided by a 90-gene expression assay versus empirical chemotherapy in patients with cancer of unknown primary (Fudan CUP-001): a randomised controlled trial. Lancet Oncol. 2024;25(8):1092–102. [DOI] [PubMed] [Google Scholar]
  • 6.Krämer A, Bochtler T, Pauli C, Shiu K, Cook N, de Menezes JJ, et al. Molecularly guided therapy versus chemotherapy after disease control in unfavourable cancer of unknown primary (CUPISCO): an open-label, randomised, phase 2 study. The Lancet. 2024;404(10452):527–39. [DOI] [PubMed] [Google Scholar]
  • 7.Binder C, Matthes KL, Korol D, Rohrmann S, Moch H. Cancer of unknown primary: epidemiological trends and relevance of comprehensive genomic profiling. Cancer Med. 2018;7(9):4814–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ren M, Cai X, Jia L, Bai Q, Zhu X, Hu X, et al. Comprehensive analysis of cancer of unknown primary and recommendation of a histological and immunohistochemical diagnostic strategy from China. BMC Cancer. 2023;23(1):1175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Qi P, Sun Y, Liu X, Wu S, Wo Y, Xu Q, et al. Clinicopathological, molecular and prognostic characteristics of cancer of unknown primary in China: an analysis of 1420 cases. Cancer Med. 2022;12:13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ye Q, Wang Q, Qi P, Chen J, Sun Y, Jin S, et al. Development and clinical validation of a 90-gene expression assay for identifying tumor tissue origin. J Mol Diagn. 2020;22(9):1139–50. [DOI] [PubMed] [Google Scholar]
  • 11.Danilova L, Wrangle J, Herman JG, Cope L. DNA-methylation for the detection and distinction of 19 human malignancies. Epigenetics. 2022;17(2):191–201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Xia D, Leon AJ, Cabanero M, Pugh TJ, Tsao MS, Rath P, et al. Minimalist approaches to cancer tissue-of-origin classification by DNA methylation. Modern Pathol. 2020;33(10):1874–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Fernandez AF, Assenov Y, Martin-Subero JI, Balint B, Siebert R, Taniguchi H, et al. A DNA methylation fingerprint of 1628 human samples. Genome Res. 2012;22(2):407–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Capper D, Jones DTW, Sill M, Hovestadt V, Schrimpf D, Sturm D, et al. DNA methylation-based classification of central nervous system tumours. Nature. 2018;555(7697):469–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Koelsche C, Schrimpf D, Stichel D, Sill M, Sahm F, Reuss DE, et al. Sarcoma classification by DNA methylation profiling. Nat Commun. 2021;12(1):498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Modhukur V, Sharma S, Mondal M, Lawarde A, Kask K, Sharma R, et al. Machine learning approaches to classify primary and metastatic cancers using tissue of origin-based dna methylation profiles. Cancers. 2021;13(15):3768. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Alexandrov LB, Kim J, Haradhvala NJ, Huang MN, Tian Ng AW, Wu Y, et al. The repertoire of mutational signatures in human cancer. Nature. 2020;578(7793):94–101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Blokzijl F, Janssen R, van Boxtel R, Cuppen E. Mutational Patterns: comprehensive genome-wide analysis of mutational processes. Genome Medi. 2018;10(1):1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Noguera-Castells A, García-Prieto CA, Álvarez-Errico D, Esteller M. Validation of the new EPIC DNA methylation microarray (900K EPIC v2) for high-throughput profiling of the human DNA methylome. Epigenetics. 2023;18(1):2185742. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Sanchez-Vega F, Mina M, Armenia J, Chatila WK, Luna A, La KC, et al. Oncogenic signaling pathways in the cancer genome atlas. Cell. 2018;173(2):321-337.e10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Sanchez-Danes A, Blanpain C. Deciphering the cells of origin of squamous cell carcinomas. Nat Rev Cancer. 2018;18(9):549–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Qi P, Sun Y, Pang Y, Liu J, Cai X, Huang S, et al. Diagnostic utility of a 90-gene expression assay (Canhelp-Origin) for patients with metastatic cancer with an unclear or unknown diagnosis. Mol Diagn Ther. 2024;29(1):81–9. [DOI] [PubMed] [Google Scholar]
  • 23.van Mourik A, Tonkin-Hill G, O’Farrell J, Waller S, Tan L, Tothill RW, et al. Six-year experience of Australia’s first dedicated cancer of unknown primary clinic. Brit J Cancer. 2023;129(2):301–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Oróstica K, Mardones F, Bernal YA, Molina S, Orchard M, Verdugo RA, et al. Advances in machine learning for tumour classification in cancer of unknown primary: A mini-review. Cancer Lett. 2025;611: 217348. [DOI] [PubMed] [Google Scholar]
  • 25.Vachani A, Nebozhyn M, Singhal S, Alila L, Wakeam E, Muschel R, et al. A 10-gene classifier for distinguishing head and neck squamous cell carcinoma and lung squamous cell carcinoma. Clin Cancer Res. 2007;13(10):2905–15. [DOI] [PubMed] [Google Scholar]
  • 26.Muñoz-Largacha JA, Gower AC, Sridhar P, Deshpande A, O’Hara CJ, Yamada E, et al. miRNA profiling of primary lung and head and neck squamous cell carcinomas: addressing a diagnostic dilemma. J Thorac Cardiov Sur. 2017;154(2):714–27. [DOI] [PubMed] [Google Scholar]
  • 27.Jurmeister P, Bockmayr M, Seegerer P, Bockmayr T, Treue D, Montavon G, et al. Machine learning analysis of DNA methylation profiles distinguishes primary lung squamous cell carcinomas from head and neck metastases. Sci Transl Medi. 2019;11(509):eaaw8513. [DOI] [PubMed] [Google Scholar]
  • 28.Leitheiser M, Capper D, Seegerer P, Lehmann A, Schüller U, Müller KR, et al. Machine learning models predict the primary sites of head and neck squamous cell carcinoma metastases based on DNA methylation. J Pathol. 2022;256(4):378–87. [DOI] [PubMed] [Google Scholar]
  • 29.Bohnenberger H, Kaderali L, Ströbel P, Yepes D, Plessmann U, Dharia NV, et al. Comparative proteomics reveals a diagnostic signature for pulmonary head-and-neck cancer metastasis. EMBO Mol Med. 2018;10(9): e8428. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Richter A, Fichtner A, Joost J, Brockmeyer P, Kauffmann P, Schliephake H, et al. Quantitative proteomics identifies biomarkers to distinguish pulmonary from head and neck squamous cell carcinomas by immunohistochemistry. J Pathol Clin Res. 2022;8(1):33–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Gao Y, Chen Z, Li J, Hu X, Shi X, Sun Z, et al. Genetic landscape of esophageal squamous cell carcinoma. Nat Genet. 2014;46(10):1097–102. [DOI] [PubMed] [Google Scholar]
  • 32.The Cancer Genome Atlas N. Integrated genomic and molecular characterization of cervical cancer. Nature. 2017;543(7645):378–384 [DOI] [PMC free article] [PubMed]
  • 33.The Cancer Genome Atlas Research N. Comprehensive genomic characterization of squamous cell lung cancers. Nature. 2012;489(7417):519–525 [DOI] [PMC free article] [PubMed]
  • 34.Mountzios G, Rampias T, Psyrri A. The mutational spectrum of squamous-cell carcinoma of the head and neck: targetable genetic events and clinical impact. Ann Oncol. 2014;25(10):1889–900. [DOI] [PubMed] [Google Scholar]
  • 35.The Cancer Genome Atlas N. Comprehensive molecular characterization of urothelial bladder carcinoma. Nature. 2014;507(7492):315–22. [DOI] [PMC free article] [PubMed]
  • 36.Ti W, Wei T, Wang J, Cheng Y. Comparative analysis of mutation status and immune landscape for squamous cell carcinomas at different anatomical sites. Front Immunol. 2022;13:947712. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Hao X, Luo H, Krawczyk M, Wei W, Wang W, Wang J, et al. DNA methylation markers for diagnosis and prognosis of common cancers. P Natl Acad Sci. 2017;114(28):7414–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Zhang S, He S, Zhu X, Wang Y, Xie Q, Song X, et al. DNA methylation profiling to determine the primary sites of metastatic cancers using formalin-fixed paraffin-embedded tissues. Nat Commun. 2023;14(1):5686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Liu B, Liu Y, Pan X, Li M, Yang S, Li SC. DNA methylation markers for pan-cancer prediction by deep learning. Genes. 2019;10(10):778. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Hackeng WM, Dreijerink KMA, de Leng WWJ, Morsink FHM, Valk GD, Vriens MR, et al. Genome methylation accurately predicts neuroendocrine tumor origin: an online tool. Clin Cancer Res. 2021;27(5):1341–50. [DOI] [PubMed] [Google Scholar]
  • 41.Koelsche C, von Deimling A. Methylation classifiers: brain tumors, sarcomas, and what’s next. Gene Chromosome Canc. 2022;61(6):346–55. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Table S1 (14.6KB, xlsx)

List of the 520 gene included in the next-generation sequencing panel.

Supplementary Material 1

Supplementary materials and methods (15.6KB, docx)

Data analysis of next-generation sequencing.

Supplementary Material 2

Complete mutation heatmaps across six tumor types (29.5MB, zip)

All samples, CESC, HNSCC, ESCC, LUSC, VSCC, and BLCA. The bar chart above the heatmap represented the total number of mutations for each sample, while the right side listed gene names in descending order of mutation frequency. On the left, the mutation frequency of each gene was displayed. The central heatmap showed the distribution of mutations across all samples, with different colors representing different mutation types. Each column represented one sample.

Supplementary Material 3

Figure S8 (419.7KB, jpg)

The performance of the classification in the adjacent normal tissues from TCGA, including the receiver operating characteristiccurve, confusion matrix, and recall curve.

Supplementary Material 4

Table S2 (14.6KB, xlsx)

Detailed information and predictive results of 119 samples in FUSCC validation set 1.

Supplementary Material 5

Data Availability Statement

The dataset used and analyzed in the present study are available from the corresponding author on reasonable request.


Articles from Clinical Epigenetics are provided here courtesy of BMC

RESOURCES