Skip to main content
Cell Reports Medicine logoLink to Cell Reports Medicine
. 2025 Sep 4;6(9):102332. doi: 10.1016/j.xcrm.2025.102332

Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data

Jilei Liu 1,3, Meng Yang 1,3, Yajing Bi 1,3, Junqing Zhang 1,3, Yichen Yang 1, Yang Li 1, Hongru Shen 1, Kexin Chen 2,, Xiangchun Li 1,4,∗∗
PMCID: PMC12490231  PMID: 40912256

Summary

Tumor-type classification is critical for effective cancer treatment, yet current methods based on genomic alterations lack flexibility and have limited performance. Here, we introduce OncoChat, an artificial intelligence (AI) model designed to classify 69 tumor types by integrating diverse genomic alterations. Developed on genomic data from 158,836 tumors sequenced with targeted cancer gene panels, OncoChat demonstrates superior performance, achieving a micro-averaged precision-recall area under the curve (PRAUC) of 0.810 (95% confidence interval [CI], 0.803–0.816), accuracy of 0.774, and an F1 score of 0.756, outperforming baseline methods. In a cancer of unknown primary (CUP) dataset of 26 cases whose types were subsequently confirmed, OncoChat correctly identified 22 cases. In two larger CUP datasets (n = 719 and 158), tumor types predicted by OncoChat were associated with survival outcomes and mutation profiles consistent with those of known tumor types. OncoChat offers promising potential for clinical decision support, particularly in managing patients with CUP.

Keywords: cancer of unknown primary, tumor-type classification, artificial intelligence, large language model

Graphical abstract

graphic file with name fx1.jpg

Highlights

  • OncoChat, a large language model, classifies tumor types from genomic alterations

  • OncoChat classifies known and unknown primary cancers with superior accuracy

  • Structural variant integration greatly improves tumor classification accuracy

  • OncoChat predictions are prognostic for survival in cancers of unknown primary


Liu et al. develop OncoChat, a large language model that accurately classifies multiple cancer types, including cancers of unknown primary (CUP), from genomic data. OncoChat outperforms existing models, and its predictions for CUP cases correlate with clinical outcomes, offering a promising new tool for personalized cancer therapy.

Introduction

Accurate tumor-type localization is crucial for determining appropriate treatment strategies, optimizing clinical outcomes, and enabling access to targeted therapies and clinical trials.1,2 Identifying both the tumor type and its primary site allows oncologists to personalize treatment based on the tumor’s biological characteristics.3 Recent advances in molecular profiling have improved cancer classifications and supported the shift toward personalized medicine.4,5,6 However, despite these innovations, diagnosing complex cancers, particularly metastatic tumors, remains a major clinical challenge.7,8

Cancers of unknown primary (CUP) present one of the most challenging clinical scenarios in oncology.9,10 CUP refers to metastatic tumors where the primary sites remain unidentified despite extensive diagnostic efforts, including histopathology, imaging, and molecular testing.11,12 Accounting for 3%–5% of all cancers globally, CUP results in an estimated 60,000 to 100,000 new cases annually in the United States alone.1,13 These tumors are often aggressive, resistant to treatment, and associated with poor outcomes, with median survival ranging from 6 to 16 months.14 Due to the unknown primary site, treatment typically involves empirical chemotherapy, which is less effective than site-specific therapies.15 Identifying the tissue of origin is therefore crucial for expanding therapeutic options and improving outcomes. While pathology, including immunohistochemistry, tumor morphology, and clinical findings, plays a key role in identifying the primary tumor, conventional methods often fail in highly metastatic or poorly differentiated CUP cases.14,15,16 These limitations, exacerbated by intra- and inter-observer variability, highlight the urgent need for more reliable diagnostic approaches.17

To address these diagnostic challenges, recent efforts have focused on improving diagnostic precision through molecular profiling techniques such as whole-genome sequencing (WGS) and whole-exome sequencing (WES).18,19,20,21 These methods provide comprehensive insights into the genomic landscape of tumors, enabling the identification of tissue-specific mutations. However, WGS and WES are costly, resource-intensive, and not yet widely available in many clinical settings, limiting their scalability.22,23 In contrast, recent studies have shown that accurate primary tumor-type classifications can be achieved using next-generation sequencing (NGS) of targeted panels. These panels, now routinely used in many cancer centers, can be applied to hundreds of thousands of tumors, offering a more accessible and scalable alternative.24,25

Targeted genomic panels, which focus on mutations relevant to multiple cancer types, have emerged as practical alternatives for cancer classification.25,26 Various machine learning (ML) models have been developed to classify tumors based on these molecular features. Notable examples include Genome-Derived Diagnosis Ensemble (GDD-ENS) and OncoNPC, which demonstrate the potential of ML to enhance cancer diagnosis.1,13 GDD-ENS, developed on Memorial Sloan Kettering’s MSK-IMPACT gene panels,24 achieved high accuracy in predicting cancer types across 38 categories. Similarly, OncoNPC, trained on targeted NGS data, exhibited strong predictive performance for known tumors and proved clinically useful in classifying CUP. Patients whose CUP tumors were classified by OncoNPC received more precise targeted therapies, resulting in significantly better outcomes compared to those treated empirically. However, these models have limitations. While OncoNPC integrates tumor genomic data (single-nucleotide variants [SNVs]/copy number alterations [CNAs]) from multiple centers, its scope is limited to 22 cancer types. GDD-ENS covers 38 cancer types and incorporates SNVs, CNAs, and structural variants (SVs) but relies solely on MSK (Memorial Sloan Kettering) samples, potentially compromising generalizability.27,28 Integrating multi-institutional sequencing data across broader tumor spectra could enhance model generalizability—a critical requirement for clinical translation.29,30

Recent advances in artificial intelligence have highlighted the potential of large language models (LLMs) to revolutionize medical research.31,32,33,34 Models based on architectures such as Generative Pre-Training35 and Bidirectional Encoder Representations from Transformers29 have demonstrated proficiency in processing vast amounts of unstructured text, including clinical notes, pathology reports, and research publications. In clinical practice, LLMs are being applied to tasks such as medical coding, summarizing patient records, extracting data from electronic health records, and assisting with diagnosis.36,37,38,39,40 Despite these successes, the application of LLMs in genomic diagnostics remains largely unexplored.41

In this study, we introduce OncoChat, a diagnostic tool that leverages LLMs to integrate genomic data with clinical information for cancer-type prediction. OncoChat enhances existing molecular classifiers by addressing key limitations, incorporating mutations, copy number changes, and structure variation in a flexible manner. The model is developed on a dataset of 158,836 across 69 tumors, all sequenced using clinically targeted cancer gene panels. OncoChat demonstrates strong performance, particularly in classifying CUP cases.

Results

Overall study design

We developed OncoChat, a differential diagnostic tool that leverages LLMs to integrate genomic and clinical data for accurate tumor-type prediction. The study involved four key components: data collection, preprocessing, model development, and evaluation (Figure S1).

The dataset was sourced from the American Association for Cancer Research (AACR) Project Genomics Evidence Neoplasia Information Exchange (GENIE) and comprised 163,585 targeted panel sequencing samples from 19 institutions. Genomic data include SNVs, copy number variations, and SVs, while clinical data encompassed age, gender, tumor type, and follow-up information (Figure S1A). Of these, 158,836 samples represented cancers with a known primary (CKP), spanning 69 cancer types, including common malignancies (e.g., non-small cell lung cancer, colorectal cancer, and breast cancer) as well as rarer cancers (e.g., hepatobiliary and pancreatic cancers). The remaining 4,749 samples were categorized as CUP (Figure 1).

Figure 1.

Figure 1

Flowchart illustrating the development and evaluation of OncoChat

The data were preprocessed into a single-turn dialogue format suitable for instruction-tuning LLMs (Figure S1B). The CKP dataset was randomly split into training and testing sets, with the model trained on the former and evaluated on the latter. OncoChat’s performance was subsequently validated on two CUP datasets with follow-up information (n = 719 and 158, respectively), as well as an independent cohort of 26 CUP cases whose tumor types were subsequently confirmed (Figure S1D). Detailed evaluation criteria are provided in the STAR Methods.

Model performance comparison with baseline methods

OncoChat demonstrated superior predictive performance compared to existing models, including OncoNPC and GDD-ENS. On the testing set of 19,940 CKP cases consisting of 69 cancer types, OncoChat achieved an accuracy of 0.774 and an F1 score of 0.756 (Table S2; Figure 2C), outperforming OncoNPC (accuracy: 0.718, F1 score: 0.701) and GDD-ENS (accuracy: 0.616, F1 score: 0.595). The precision-recall area under the curve (PRAUC) for OncoChat was 0.81 (95% confidence interval [CI], 0.803–0.816), significantly higher than 0.789 (95% CI, 0.782–0.796) for OncoNPC and 0.689 (95% CI, 0.681–0.695) for GDD-ENS (p < 0.001) (Figure 2B).

Figure 2.

Figure 2

Performance of OncoChat compared with baseline models

(A) Normalized confusion matrix of OncoChat’s performance across 23 pre-selected cancer types in the held-out test set (n = 19,940). Precision is shown on the diagonal, recall is displayed below the matrix, and sample sizes are listed on the left.

(B) Precision-recall curves comparing OncoChat, GDD-ENS, and OncoNPC on the testing set. The statistical significance of the model-to-model differences was evaluated with permutation test (1,000 random permutations).

(C) Bar plot of overall accuracy and F1 scores for OncoChat, GDD-ENS, and OncoNPC.

(D) F1 scores across cancer types for OncoChat under four filtering thresholds, with dot size indicating the proportion of tumor samples retained.

(E) Scatterplot of filtering thresholds versus sample retention proportion, with F1 scores represented by dot size and colors distinguishing models.

(F) Comparison of OncoChat and baseline methods across 69 cancer types, with green, gray, and red bars indicating metrics better, equal, or worse than baseline methods, respectively. Figures 2A and 2D were plotted based on 23 representative tumor types. The remaining figures were generated using all 69 cancer types.

See Figure S14 and Table S3 for the complete list of tumor types included in OncoChat’s testing, along with their corresponding performance metrics, including precision, recall, and F1 score.

OncoChat’s performance was robust across a diverse range of tumor types, including traditionally challenging cases such as gliomas, prostate cancers, and soft tissue sarcomas, where it achieved superior precision and recall (Figures 2A, 2D, and S14; Table S3). At four distinct predictive thresholds, OncoChat maintained high F1 scores while retaining a substantial proportion of tumor samples. Additionally, OncoChat demonstrated greater classification stability under varying filtering thresholds compared to GDD-ENS and OncoNPC, as shown in Figure 2E and Table S3.

A detailed comparison in Figure 2F highlights OncoChat’s consistent superiority, particularly in cancers that are difficult to classify, such as endometrial cancers and head and neck cancers. OncoChat achieved higher precision, recall, and F1 scores across multiple cancer types, demonstrating exceptional reliability and accuracy (Table S3).

OncoChat demonstrated consistent and robust performance across diverse clinical and demographic settings, underscoring its broad applicability in cancer diagnosis. When stratified by cancer centers, OncoChat maintained stable predictive accuracy, achieving a PRAUC value of 0.809 at MSK (n = 9,964), 0.806 at DFCI (Dana-Farber Cancer Institute, n = 8,642), and 0.725 at DUKE (Duke University Medical Center, n = 513). These results underscore the model’s stability and its consistency compared to GDD-ENS and OncoNPC across institutions (Figure S2). To evaluate the impact of variability in gene panel coverage, predictions were further stratified by panel type. OncoChat exhibited stable performance across different panels, aligning with earlier findings13 (Figure S3). Similarly, the model maintained strong predictive accuracy across sample types, including primary (n = 11,605) and metastatic (n = 7405) tumors, as well as across patient ethnicities. OncoChat consistently outperformed OncoNPC and GDD-ENS among White (n = 15,679), Black (n = 1,122), Asian (n = 1,119), and other ethnicities (n = 463) (Figures S4 and S5; Table S2). In direct comparisons, OncoChat outperformed both OncoNPC and GDD-ENS across overlapping cancer types. For the 38 cancer types in GDD-ENS, OncoChat achieved superior accuracy (0.803 vs. 0.749 [OncoNPC] and 0.643 [GDD-ENS]) and F1 scores (0.796 vs. 0.747 and 0.638) (Figure S10B). This performance advantage persisted when restricted to OncoNPC’s 22 cancer types (accuracy: 0.827 vs. 0.774 and 0.658; F1: 0.828 vs. 0.779 and 0.657; Figure S10A).

As an ensemble of nine language models ranging in size from 100 Mb to 7 billion parameters (see STAR Methods), predictive performance positively correlated with model size. Accuracy increased from 0.59 for the smallest model to 0.753 for the largest, while F1 scores improved from 0.572 to 0.739 (Table S2). Notably, even a relatively small model achieved competitive results, with one model attaining an accuracy of 0.714 and an F1 score of 0.706, rivaling or surpassing the performance of OncoNPC and GDD-ENS. Cross-validation analysis revealed robust performance across CKP cohort partitions (accuracy: 0.744 ± 0.003; F1: 0.730 ± 0.005; Figure S12; Table S15), demonstrating model stability with minimal overfitting.

These results underscore OncoChat’s effectiveness in delivering high accuracy and robust performance across a broad spectrum of cancer types and clinical scenarios, making it a valuable tool for enhancing diagnostic workflows in oncology. Its precision in cancer classification has the potential to improve patient outcomes by enabling more accurate and tailored treatments.

Integration of SVs enhances OncoChat’s performance

Incorporating SV data significantly improved OncoChat’s predictive performance. The PRAUC increased from 0.802 (95% CI, 0.789–0.816) to 0.831 (95% CI, 0.819–0.843) (p < 0.001), highlighting the critical role of SVs in distinguishing complex cancer types (Figure 3B). Overall accuracy improved from 0.766 to 0.798, and F1 score rose from 0.749 to 0.781, reflecting a better precision-recall balance across diverse tumor types (Figure 3C).

Figure 3.

Figure 3

Enhanced performance of OncoChat with SV data

(A) Normalized confusion matrix of OncoChat’s performance on the held-out test set for 15 pre-selected cancer types using SV data (n = 5,799). Precision is shown on the diagonal, recall below, and sample size on the left.

(B) Precision-recall curves for OncoChat with and without SV data. The statistical significance of the difference was evaluated with permutation test (1,000 random permutations).

(C) Bar plot comparing accuracy and F1 scores for OncoChat with and without SV data.

(D) Classification metrics for OncoChat across cancer types with and without SV data.

As shown in the confusion matrix (Figure 3A), OncoChat achieved notably higher recall for challenging cancers such as gliomas and mature B cell neoplasms, where genomic complexity often hampers classification.42,43 Enhanced precision and recall were also observed for breast, colorectal, and hepatobiliary cancers, demonstrating the broad utility of SVs in refining predictive performance.

Further cancer-specific analysis (Figure 3D) revealed that SV inclusion significantly boosted precision, recall, and F1 scores for hard-to-classify tumors, including hepatobiliary and bone cancers. These enhancements reflect improved sensitivity and reduced false negatives, underscoring OncoChat’s heightened ability to accurately identify difficult cancer types.

Incorporating SV data improved not only the ensemble performance but also the classification performance of individual models. On average, accuracy increased by 2.7%, while the F1 score improved by 2.5% across individual models (Table S6). The magnitude of improvement correlated positively with model size, indicating that larger models or the full OncoChat framework are particularly well suited for predictions involving SV data (Figure S6).

Rare tumor classification performance

OncoChat demonstrated superior classification of rare cancers across prevalence thresholds in the full cohort. At the 15% prevalence threshold (see STAR Methods), OncoChat achieved higher accuracy (0.428 vs. 0.376 [OncoNPC] and 0.257 [GDD-ENS]) and F1 scores (0.531 vs. 0.474 and 0.358), with consistent advantages at stricter thresholds (<10%: F1 = 0.453; <5%: F1 = 0.352) (Figure S13; Table S14).

Performance remained robust for ultra-rare tumors (<200 samples): mesothelioma (F1 = 0.746 vs. 0.690 and 0.367), bone cancer (F1 = 0.560 vs. 0.428 and 0.331), and uterine sarcoma (F1 = 0.402 vs. 0.342 and 0.390). OncoChat correctly reclassified 21/64 and 66/133 mesothelioma cases misclassified by OncoNPC and GDD-ENS, respectively, demonstrating improved sensitivity.

While all models showed reduced accuracy for rare tumors, OncoChat’s consistent outperformance highlights its relative strength in this challenging domain, though further refinement remains necessary.44

Performance across broad cancer categories

OncoChat’s predictive performance was evaluated across broad categories, including gastrointestinal (GI), thoracic, breast, and brain/CNS tumors, reflecting the clinical grouping of cancers based on anatomical and biological characteristics (Figures 4A and S15). The model demonstrated robust performance across these categories. For instance, GI cancers—including colorectal, pancreatic, and esophagogastric cancers—were classified with an F1 score approaching 0.98, highlighting exceptional precision and recall.

Figure 4.

Figure 4

OncoChat’s performance across broad cancer categories

(A) Normalized confusion matrix showing classification performance for 11 broad cancer types in held-out test set (n = 19,940). Precision is shown on the diagonal, recall below, and sample size on the left.

(B) Bar plot of overall accuracy and F1 scores for OncoChat, GDD-ENS, and OncoNPC at the broad cancer-type level.

(C) F1 scores across broad cancer types for OncoChat at four different filter thresholds, with dot size indicating sample retention proportion.

(D) Detailed breakdown of prediction results for broad cancer types.

On the CKP test set (n = 19,940), OncoChat achieved an overall accuracy of 0.831 and an F1 score of 0.827 at the category level, outperforming baseline models. By comparison, OncoNPC achieved an accuracy of 0.776 and an F1 score of 0.772, while GDD-ENS lagged with an accuracy of 0.681 and an F1 score of 0.676 (Figure 4B). Detailed classification metrics were provided in Table S4.

As shown in Figure 4C, OncoChat maintained high predictive accuracy as classification threshold became more stringent, demonstrating its reliability across a range of clinical scenarios (Table S5). The model’s ability to retain a substantial proportion of samples with high confidence supports its versatility in varied diagnostic contexts. Figure 4D provides a detailed breakdown of these prediction results.

Performance of OncoChat in CUP samples

To explore the clinical utility of OncoChat, we analyzed whether its cancer-type predictions could stratify overall survival in a CUP cohort consisting of 719 patients. Predictions identified subgroups with distinct median survival outcomes (log rank test, p < 0.001; Figure 5A). Patients predicted to have pancreatic cancer (pancreatic adenocarcinoma) or esophagogastric cancer showed the shortest survival, reflecting the aggressive nature of these malignances. In contrast, CUP cases predicted as head and neck squamous cell carcinoma (HNSCC) or GI neuroendocrine tumors exhibited significantly longer survival. Notably, a strong correlation was observed between the median survival of CUP patients and that of their CKP counterparts (Spearman’s ρ = 0.75, p = 0.066; Figure 5B), suggesting that OncoChat captures prognostic features similar to those in cancers with known primaries. Further validation came from the high concordance of base substitution patterns between CUP and CKP samples (Spearman’s coefficient of 0.918, p < 0.001, Figure 5C), highlighting OncoChat’s reliability in detecting shared mutational signatures.

Figure 5.

Figure 5

OncoChat’s performance in CUP samples

(A) Kaplan-Meier survival curves stratified by OncoChat-predicted cancer types for CUP patients, with statistical significance assessed via the log rank test.

(B) Comparison of median survival between CUP patients and corresponding CKP patients, with dot size representing the log rank test p value.

(C) Analysis of single-base substitution types in CUP and CKP samples for corresponding cancer types, with point shapes denoting substitution types and colors indicating cancer types.

(D) Heatmap of cosine similarity between mutation signatures for CUP and CKP samples with the same predicted cancer type, normalized by rows and clustered by columns.

(E) Sankey plot comparing performance of OncoChat, OncoNPC, and GDD-ENS on an independent CUP cohort (n = 26). EGC/PAAD, esophagogastric adenocarcinoma/pancreatic adenocarcinoma; GI/PANET, gastrointestinal, pancreatic neuroendocrine tumor.

In an independent cohort of 158 CUP cases, treatment concordance with OncoChat’s predictions trended toward improved survival (log rank test, p = 0.065). Multivariable analysis confirmed this association (hazard ratio = 0.66, 95% CI: 0.436–1.00, p = 0.048), independent of gender, CNA burden, or histology (Table S13; Figure S7). These results align with NCCN CUP Guidelines (v2.2025) and prior evidence that molecular profiling improves outcomes, supporting OncoChat’s utility in treatment selection for unresolved primaries.16,45

Mutation analysis further underscored this consistency. For instance, CUP samples classified as non-small cell lung cancer displayed mutation signatures associated with smoking and DNA mismatch repair deficiency (e.g., SBS3, SBS4, SBS8, SBS14, and SBS29), while those predicted as melanoma exhibited UV-related patterns (e.g., SBS7) (Figure 5D).46 These findings demonstrate OncoChat’s ability to characterize the genomic landscape of various cancer types effectively.

Superior performance of OncoChat in confirmed CUP cohorts

In a cohort of 26 CUP samples with subsequently confirmed diagnosis, OncoChat accurately predicted the tumor type in 22 cases (Figure 5E; Table S7). This surpassed the performance of OncoNPC and GDD-ENS, which misclassified multiple samples. For example, GDD-ENS incorrectly identified prostate cancer as esophagogastric cancer and biliary cancer as sarcoma. OncoChat correctly classified challenging cases, such as three prostate cancers, two esophagogastric cancers, and one biliary cancer, demonstrating superior predictive accuracy.

Overall, OncoChat effectively predicts the tissue of origin for CUP cases with accuracy comparable to its performance in CKP samples. Its ability to classify aggressive and rare cancers and inform site-specific treatment strategies underscores its potential to enhance diagnostic workflows and improve patient outcomes, particularly in cases where conventional methods fall short.

Model interpretability reveals biological plausibility

Analysis of CKP samples identified tumor-type-specific driver mutations (Figure S8). For instance, TP53 dominated lung/ovarian/glioma classifications, while BRAF (melanoma/thyroid) and PIK3CA (breast/endometrial) showed subtype specificity. Driver genes like KDM6A (bladder), CTNNB1 (liver), and SRSF2 (leukemia) further validated biological alignment.

In 22/26 correctly classified CUP cases (Figure S9), feature contributions matched established biology—APC enrichment in colorectal cancer (CRC) predictions, VHL in renal cell carcinoma (RCC)—demonstrating decision transparency.

Attention mechanisms captured functional interactions: BRCA-PARP1 scores were elevated specifically in BRCA-mutated breast cancer (0.060 ± 0.122 vs. KRAS-EGFR: 0.010 ± 0.027; p < 1e−3), unlike control samples (0.008 ± 0.014; p = 0.260) (Figure S11). This unsupervised detection of synthetic lethality highlights OncoChat’s advantage over feature-agnostic models.47 Detailed analysis pipelines are provided in the STAR Methods.

Discussion

The development and validation of OncoChat, a LLM-based diagnostic tool, mark a significant advancement in cancer diagnosis, particularly in challenging cases such as CUP. OncoChat provides a robust mechanism for accurate tumor classification, surpassing the performance of existing ML models in both CKP and CUP cases. This dual capability underscores the model’s potential for broader applications in clinical oncology, where precise tumor classification is critical for determining the most effective therapeutic strategies.

OncoChat demonstrated impressive performance in the CKP test set, achieving an accuracy of 77.4%, an F1 score of 75.6%, and a PRAUC of 0.810. These metrics are particularly significant given the challenges in classifying rare and complex cancers, such as endometrial cancers and head and neck cancers, where OncoChat outperformed both OncoNPC and GDD-ENS. Furthermore, OncoChat offers a broader cancer classification spectrum, encompassing 69 tumor types, compared to models like OncoNPC and GDD-ENS, which focus on 20–30 types. Its ability to maintain strong predictive performance across a diverse array of tumor types highlights OncoChat’s generalizability and scalability, making it suitable for adoption in varied clinical settings. Besides, OncoChat demonstrated robust performance across diverse clinical and demographic contexts, reinforcing its potential broad applicability in cancer diagnosis.

OncoChat’s application to CUP cases is particularly noteworthy. CUPs are notoriously difficult to diagnose due to the lack of an identifiable primary site. In this study, OncoChat’s predictions for CUP samples showed performance comparable to its accuracy in CKP cases. The correlation between OncoChat’s predictions and survival outcomes in CUP patients further underscores its clinical relevance. Patients predicted to have aggressive cancers, such as pancreatic and esophagogastric cancers, exhibited shorter survival times, while those predicted to have less aggressive tumors, such as HNSCC, had better prognoses. These findings suggest that OncoChat not only captures tumor type but also aligns with clinically relevant prognostic indicators, enhancing its potential as a tool for guiding personalized treatment plans.

OncoChat demonstrates superior diagnostic performance in challenging clinical scenarios, particularly for CUP. The model correctly identified the tissue of origin in 84.6% (22/26) of confirmed CUP cases, outperforming existing approaches that frequently misclassified morphologically ambiguous tumors. More importantly, alignment between OncoChat’s predictions and empirical therapy selection was associated with improved survival outcomes, providing clinical validation of its potential utility in precision oncology. These findings strongly support current National Comprehensive Cancer Network (NCCN) guidelines advocating for molecular profiling in CUP management.

The clinical applicability of existing molecular classifiers has been limited by several factors: (1) restricted tumor-type coverage (22–38 types), (2) single-institution training biases, and (3) diminished performance in rare malignancies. OncoChat addresses these limitations through comprehensive training on 158,836 samples spanning 69 tumor types from 19 institutions. Notably, even when evaluated on the same tumor types as narrower scope models, OncoChat maintained superior performance (accuracy = 0.827 vs. 0.774 for OncoNPC), suggesting more robust feature learning. This advantage was particularly pronounced for rare tumors (≤15% prevalence: F1 = 0.531 vs. 0.474) and ultra-rare histologies (mesothelioma: F1 = 0.746 vs. 0.690), where existing models often fail.

The incorporation of SVs into OncoChat further boosted its performance, especially in classifying complex cancers. SVs, which are often implicated in cancer progression and tumor heterogeneity, enabled OncoChat to more accurately classify cancers that are difficult to diagnose using point mutations or CNAs alone.48 This improvement was especially evident in cancers such as gliomas and mature B cell neoplasms, where structural alterations play a crucial role.42,43 These results highlight the importance of comprehensive genomic profiling, as it allows OncoChat to capture the full spectrum of genomic alterations associated with various tumor types, thereby improving diagnostic accuracy.

Through systematic interpretability analyses, we demonstrate that our model’s predictions are anchored in established cancer biology.49 The framework identified driver mutations with distinct tissue-specific distributions: while TP53 mutations dominated epithelial malignancies (lung, ovarian, and esophageal) and gliomas, BRAF alterations were characteristic of melanocytic and thyroid lineages, and PIK3CA activation marked hormone-responsive cancers (breast and endometrial). Notably, the model’s classification confidence correlated with the strength of these biological signatures—correctly predicted CUPs exhibited mutation profiles mirroring their putative tissues of origin, including canonical APC alterations in colorectal cases and VHL mutations in renal carcinomas.50,51 A particularly compelling finding emerged from attention-based analysis, which revealed the model’s capacity to detect functional genomic interactions without prior biological annotation. In BRCA-mutant breast cancers, we observed significant co-attention between BRCA and PARP1, recapitulating the synthetic lethality relationship that underpins poly(ADP-ribose) polymerase (PARP) inhibitor sensitivity.47 This biologically grounded feature selection contrasts sharply with conventional ML approaches, which failed to capture this interaction due to their assumption of feature independence. These results not only validate the model’s biological plausibility but also suggest its potential to uncover novel therapeutic vulnerabilities through pattern recognition in genomic data.52

Conventional diagnostic workflows for CUP, reliant on sequential immunohistochemistry (IHC) and imaging, remain limited by marker specificity and interpretive subjectivity.14,15,16 Recent studies highlight the diagnostic potential of histomorphological approaches. Deep learning models applied to H&E slides can classify tumors with >80% accuracy in certain cancers.53,54 However, such methods may struggle with poorly differentiated or metastatic CUP cases, where morphological ambiguity is common.55,56 In contrast, OncoChat leverages genomic alterations preserved across tumor stages, enabling classification even in morphologically challenging cases. For instance, OncoChat correctly identified 22/26 CUP cases with confirmed diagnoses, including metastatic prostate and biliary cancers, which often lack distinct histopathological hallmarks. Molecular tools like CUP-AI-Dx and TOD-CUP using gene expression profiling achieve 72%–94% accuracy in retrospective studies but lack survival benefit verification.57,58 Transcriptomic classifiers similarly struggle with poorly differentiated tumors (e.g., pancreatic adenocarcinoma) that rapidly lose transcriptional identity.59,60 Emerging epigenomic methods such as DNA methylation profiling (EPICUP) and cell-free DNA-based assays (CUPiD) show promise but depend on specialized platforms, limiting accessibility.9,61 NCCN Guideline for Occult Primary notes that while DNA methylation and transcriptomic profiling classify 75% of CUPs, their clinical utility remains unproven. In contrast, targeted sequencing identifies pathogenic germline alterations in 14.5% of CUPs, enabling therapies with documented clinical benefit (e.g., 53.8% response rate).45 OncoChat bridges this gap by prioritizing mutations linked to approved therapies, offering a pragmatic balance between scalability and therapeutic relevance.

OncoChat addresses these challenges by leveraging stable genomic alterations, even in dedifferentiated tumors, resolving 37% of IHC-indeterminate cases in published cohorts.16 Its design exploits targeted DNA sequencing panels, widely adopted in clinical oncology due to compatibility with archival FFPE tissue and avoidance of labor-intensive assays (e.g., RNA sequencing, methylation profiling). In our CUP cohort, nearly half of samples harbored actionable mutations, highlighting the dual diagnostic-therapeutic utility of mutation-based profiling. Though OncoChat may not surpass transcriptomic or epigenomic classifiers under ideal conditions, its alignment with routine clinical infrastructure renders it an immediately deployable solution. Rather than replacing existing modalities, it complements them by unifying tumor classification and therapeutic guidance in a single assay. The NCCN Guideline for Occult Primary (CUP) version 2.2025 emphasizes molecular profiling as a tiered diagnostic tool alongside IHC, noting that genomic classifiers can resolve cases where histomorphology is inconclusive. Our results align with these recommendations: OncoChat’s predictions stratified CUP survival outcomes and mirrored mutational signatures of known primaries, underscoring its utility in guiding site-specific therapies when conventional methods fail. Future iterations of OncoChat will incorporate histomorphological data, building on frameworks like Orion and NCCN-endorsed multimodal strategies.62 For example, combining mutational profiles with spatially resolved immune infiltration patterns or peritumoral edema features could refine prognostic stratification.63,64

In conclusion, OncoChat represents a significant leap forward in cancer diagnosis, offering a powerful tool for accurately classifying both known and unknown tumors. By integrating genomic and clinical data, OncoChat provides superior performance in predicting cancer types from targeted genomic profiling, positioning it as a promising diagnostic tool in personalized medicine. As the model evolves, it has the potential to transform the clinical practices, especially for challenging cases like CUP, where precise tumor classification is critical for improving patient outcomes. Further validation and refinement will be essential for ensuring its widespread adoption.

Limitations of the study

While OncoChat demonstrates broad utility, key challenges remain. First, the model must be validated in substantially larger, well-annotated cohorts that include both cancers with a known-primary (CKP) tumors and CUP with definitive diagnostic labels, so that performance estimates reflect the full heterogeneity of real-world practice. Second, expanding the feature space to include multimodal data—such as bulk and single-cell RNA sequencing, epigenomic assays (e.g., DNA methylation and chromatin accessibility), circulating cell-free DNA fragmentomic patterns, and whole-slide pathology images—may uncover complementary biological signals and further enhance predictive power, particularly for challenging specimens with low tumor burden where DNA-based classifiers often underperform. Systematic evaluation of these transcriptomic, epigenetic, and histopathological features is essential to ensure robust performance across the full spectrum of clinical materials. Third, while OncoChat demonstrates robust performance across broad tumor types, we acknowledge that rare tumor entities remain markedly underrepresented in current training sets; enlarging the dataset to include more cases of uncommon malignancies—and subjecting predictions in these subgroups to rigorous prospective evaluation—will be essential before OncoChat can be confidently applied to the full breadth of oncologic diagnoses.

Resource availability

Lead contact

For further information and requests for resources, please contact the lead contact, Xiangchun Li (lixiangchun@tmu.edu.cn).

Materials availability

This study did not generate new unique reagents.

Data and code availability

Acknowledgments

This work was supported by the National Key Research and Development Program of China (grant no. 2021YFC2500400 to K.C.), the National Natural Science Foundation of China (grant nos. 32270688 and 31801117 to X.L. and 31900471 to M.Y.), the Program for Changjiang Scholars and Innovative Research Team in University in China (grant no. IRT_14R40 to K.C.), the China Postdoctoral Science Foundation (grant no. BX20240253 and 2024M762384 to H.S.), the Natural Science Foundation of Tianjin (grant no. 24JCQNJC01280 to H.S), and the Tianjin Key Medical Discipline (Specialty) Construction Project (YXZDXK-009A). We express our gratitude to the AACR Project GENIE team for their efforts in aggregating, managing, and sharing the data used in this study. We also extend our gratitude to Alibaba, Inc. for Qwen-1.5 and to the researchers from Carnegie Mellon University and Princeton University for Mamba. These open-source LLMs were integral to our research.

Author contributions

X.L. and K.C. designed and supervised the study. X.L. and J.L. performed data analysis and wrote the manuscript. X.L., J.L., M.Y., Y.B., and J.Z. developed the model. X.L., J.L., M.Y., Y.B., J.Z., Y.Y., Y.L., and H.S. collected the data. J.L., X.L., and K.C. revised the manuscript.

Declaration of interests

The authors declare no competing interests.

STAR★Methods

Key resources table

REAGENT or RESOURCE SOURCE IDENTIFIER
Deposited data

GENIE data AACR https://doi.org/10.7303/syn53210170
26 CUP with confirmed diagnoses Darmofal et al.1 https://aacr.silverchair-cdn.com/aacr/content_public/journal/cancerdiscovery/14/6/10.1158_2159-8290.cd-23-0996/5/cd-23-0996_table_s17_suppst17.xlsx
158 CUP with treatment information Moon et al.13 https://github.com/itmoon7/onconpc
719 CUP with follow-up information This paper https://github.com/deeplearningplus/OncoChat/tree/main/data

Software and algorithms

OncoNPC Moon et al.13 https://github.com/itmoon7/onconpc
GDD-ENS Darmofal et al.1 https://github.com/mmdarmofal/GDD_ENS
OncoChat This paper https://github.com/deeplearningplus/OncoChat
R (version 4.2.0) R Core Team RRID:SCR_001905
Pytorch (version 2.6.0) PyTorch Foundation RRID:SCR_018536
Qwen-1.5 Github https://github.com/QwenLM/Qwen
Mamba Github https://github.com/state-spaces/mamba
TransformerLens Github https://github.com/TransformerLensOrg/TransformerLens

Experimental model and study participant details

The study protocol received approval from the institutional review boards or ethics committees of the participating institutions. As this study is retrospective, patient consent was not required.

A total of 197,980 next-generation sequencing (NGS) panel samples were collected from 171,957 patients across 19 institutions participating in the AACR Project GENIE. Detailed demographic information, including age, gender, and racial distribution, is provided in Table S1. This publicly accessible international cancer registry aggregates real-world data through collaborative efforts among leading cancer centers. By harmonizing clinical-grade cancer genomic sequencing data with patient outcomes from routine medical care, the project provides a comprehensive dataset for cancer research.

Given the multi-institutional nature of the dataset, the targeted panels used varied, with differences in gene coverage. All panels detected single-nucleotide variants (SNVs) and copy number variations (CNVs), while some also identified structural variants (SVs). Additional details on panel specifications are provided in Table S10. An example lung cancer sample, illustrating its clinical and genomic data, is presented in Figure S1A. This includes patient demographics (e.g., age, gender), diagnostic information, follow-up data, and genomic features such as SNVs, CNVs, and SVs.

After quality control, 34,395 samples were excluded based on predefined criteria: 3,958 samples labeled as “UNKNOWN”, 731 with incomplete diagnostic information, 176 tumor types with fewer than 15 samples, and 29,530 samples with low mutation counts (<2 mutations). This resulted in a total of 163,585 eligible cancer samples, categorized as either Cancer of Known Primary (CKP; 158,836 samples across 69 cancer types) or Cancer of Unknown Primary (CUP; 4,749 samples).

The CKP samples were randomly split into a training set (80%, n = 127,069) and a testing set (20%, n = 31,767). Among the CUP samples, 26 cases were subsequently confirmed with specific cancer types after follow-up and designated as the CUP testing set (Tables S11 and S7). To enable direct comparison with baseline models (OncoNPC and GDD-ENS), which require both SNV and CNA inputs, we limited the test set to samples with complete genomic data (n = 19,940). This ensures consistent input features across all models during benchmarking.

Method details

Data preprocessing

To enable training of the large language model (LLM), unstructured raw data were transformed into a structured single-turn dialogue format. As illustrated in Figure S1B, the format consisted of three components: user input (X), user instruction (I), and model response (Y). User input (X) was created by concatenating clinical phenotypes (C) and genomic information (G) into a single string. For example, clinical phenotypes are expressed as:

C=Age:36,Gender:Female

Genomic information included SNVs, CNVs and SVs. SNVs and SVs captured detailed genetic alterations, such as protein changes or structural modification. For instance:

SNVs=ATMp.T2396S,CICp.A686V,NF1p.V34I,
SVs=INVERSION:NAB2STAT6,TRANSLATION:PRKAR1AUTRN,

CNVs, which indicate variations in gene copy number, were transformed into biologically interpretable terms. Numerical values from the raw data were converted as follows: −2 → deep loss; −1.5 → medium-level loss; −1 → single-copy loss; 1 → low-level gain; 1.5 → medium-level amplification; 2 → high-level amplification. For example, a CNV description might appear as:

CNVs=CDKN2A:deeploss,ELF3:highlevelamplification

User instruction (I) prompted the LLM to generate predictions. Instructions were simple queries, such as “What is the broad cancer origin?”. Model response (Y) corresponded to the predicted cancer type for the sample. For instance,

Y=nonsmallcelllungcancer

The raw data for all samples were processed using these predefined rules, ensuring consistency across the dataset. A detailed example of the dialogue structure is presented in Figure S1C, and the complete workflow for data preprocessing is illustrated in Figure S1B.

Model architectures

OncoChat is created through instruction-tuning of each individual model in the Qwen-series65 (qwen-1.5-0.5b, 1.8b, 4b, 7b) and Mamba-series66 models (mamba-130m, 370m, 790m, 1.4b, 2.8b). The results from each tuned Qwen and Mamba model are then aggregated using model ensembling to obtain the final OncoChat predictions. Detailed model parameters for each model are provided in Table S12.

Qwen is a decoder-only transformer that incorporates self-attention mechanisms with causal masking and feedforward neural networks (FFNs). In Qwen-1.5, Grouped Query Attention (GQA) replaces traditional multi-head attention (MHA), enhancing efficiency. The model uses the SwiGLU activation function and rotary positional embeddings (RoPE) to better capture positional information. Additionally, Qwen-1.5 uses QKV bias in the attention mechanism, improving its ability to handle longer sequences.

Mamba is a deep learning architecture tailored for sequence modeling, addressing the limitations of transformers when processing long sequences. It integrates the Structured State Space (S4) model, which combines continuous-time, recurrent, and convolutional mechanisms to effectively capture long-range dependencies. The S4 model allows Mamba to maintain an unbounded context while ensuring computational efficiency during both training and inference. Mamba introduces a dynamic mechanism that adjusts State Space Model (SSM) parameters based on input sequences. This adaptability enables the model to focus on relevant information and filter out less important data, transitioning from a time-invariant to a time-varying framework. The combination of SSM with multi-layer perceptron (MLP) blocks results in a streamlined architecture suitable for a wide range of sequence modeling tasks.

Instruction fine-tuning objective

Instruction-tuning involves fine-tuning a pre-trained LLM on a dataset D, where each training example includes a source (instruction and input, s) and a target (response, t).67 During training, the model parameters θare updated by minimizing the loss on the target tokens, conditioned on the source and previously generated target tokens. Formally, this can be defined as:

L(D,θ)=ijlogpθ(tij|si,ti,<j)

This approach ensures the model learns to follow instructions and generate accurate responses within a specific domain.

Model development and evaluation

Following data preprocessing, the structured sample data were used to fine-tune the Qwen-series and Mamba-series models for cancer classification. As depicted in Figure S1C, the model generates a tumor type prediction (denoted as Yp) based on the input data (X) and user instruction (I). During the training phase, OncoChat was fine-tuned by minimizing the discrepancy between the model’s prediction (Yp) and true label (Y) using the instruction fine-tuning method. We trained each model for 3 epochs with a learning rate of 2e−5 and batch size of 8. The AdamW optimizer was employed for parameter updates. Training was conducted using PyTorch (version 2.2.1) and the transformers library (version 4.21.1) on NVIDIA DGX A100 system equipped with 8 GPUs, each with 40 Gb of memory.

In the evaluation phase, the model parameters were frozen, and predictions (Ypi) were generated using only user input (X) and instruction prompt (I). These predictions were then compared to the corresponding ground truth labels (Yi) where i represents the sample index, to assess the model’s performance. Finally, as shown in Figure S1D, OncoChat was applied to infer the specific cancer types of CUP cases.

To enhance prediction reliability, we employed an ensemble strategy. For each sample, the nine models generate their own predicted label. Then, for each prediction label (e.g., “colorectal cancer”), we count how many models support that label. The confidence of the sample being predicted as a specific label is calculated as the proportion of models that predict that label, i.e., k9, where k is the number of models predicting that label. The label with the highest confidence is chosen as the final prediction for that sample. This approach treats the frequency of support for each label as an indicator of prediction confidence, ranging from 0 to 1. Model evaluation included 5-fold stratified cross-validation, with stratification by tumor type and institution to preserve distributional characteristics. Performance metrics were averaged across folds, with standard deviations reported to quantify variability. We also categorized tumors as “rare” based on three prevalence thresholds (<5%, <10%, and <15%), corresponding to tumor types with fewer than 1,000, 1,500, and 3,000 samples, respectively. Performance evaluations were conducted on the testing set according to the tumor types included under each threshold.

Benchmark methods

To evaluate the performance of OncoChat, we compared it against two state-of-the-art cancer type classification methods: OncoNPC13 and GDD-ENS.1 Both benchmark models were trained and evaluated using the same training and testing datasets as OncoChat, ensuring a consistent and fair comparison.

OncoNPC, a machine-learning classifier based on the eXtreme Gradient Boosting (XGBoost) algorithm, was trained on the training set using somatic alterations (mutations, mutational signatures, and copy number alterations) along with patient age and gender. Model optimization was performed through hyperparameter tuning with random search and 10-fold cross-validation. We used the original code and parameter settings provided by the authors for model training and evaluation.

GDD-ENS, a deep-learning system, was trained on the same training set. It utilized genomic features derived from the MSK-IMPACT panels, including mutations, indels, copy-number variations, gene fusions, and mutation signatures. The model architecture employed a hyperparameter ensemble of 10 multi-layer perceptrons. Training incorporated upsampling of smaller cancer types to balance class distribution and used a Gaussian process for hyperparameter optimization. Model optimization was conducted using random search and 10-fold cross-validation, followed by model ensembling. Final prediction results were obtained using the original code and default parameter settings provided by the authors.1

Evaluation metrics

To compare the performance of OncoChat, OncoNPC, and GDD-ENS, we used the following evaluation metrics.

Accuracy: The proportion of correctly predicted cancer types in the test set.

Accuracy=TruePositives+TrueNegativesTotalSamples

Precision: The proportion of true positive predictions among all positive predictions made by the model.

Recall: The proportion of true positive predictions among all actual positive cases.

Precision=TruePositivesTruePositives+FalsePositives
Recall=TruePositivesTruePositives+FalseNegatives

Weighted F1 Score: A harmonic mean of the precision and recall that accounts for class imbalance in the dataset, providing a more comprehensive measure of model performance.

WeightedF1Score=i=1N(2·Precisioni·RecalliPrecisioni+Recalli·weighti)

where Weightirepresents the proportion of class i in the dataset, and N is the total number of classes.

Risk stratification among patients with cancer of unknown primary (CUP)

To identify CUP subgroups with significant prognostic differences, we estimated survival functions for seven common subgroups predicted by OncoChat, each containing more than 35 CUP patients: non-small-cell lung cancer (NSCLC), pancreatic adenocarcinoma (PAAD), breast carcinoma (BRCA), head and neck squamous cell carcinoma (HNSCC), esophagogastric adenocarcinoma (EGC), gastrointestinal neuroendocrine tumors (GINET), pancreatic neuroendocrine tumor (PANET). Consistent with prior studies,13 subgroups with similar morphologies were merged for survival analysis: PAAD was combined with EGC, and GINET was combined with PANET.

CUP patients from DFCI center who were lost to follow-up at the time of sequencing were excluded, resulting in a final cohort of 719 samples for analysis (Table S8). Survival analysis and visualization were performed using the survival and survminer R packages.

Mutation profiles of CUP samples

To validate the reliability of CUP prediction, we analyzed the association between the SNVs of CUP samples and their corresponding CKP. This analysis included 9,272 CKP samples and 479 CUP samples from the DFCI Cancer Center, sequenced using the DFCI-ONCOPANEL-3.1 oncology panel, covering nine cancer types: Bladder Cancer, Breast Cancer, Colorectal Cancer, Esophagogastric Cancer, Head and Neck Cancer, Melanoma, Non-Small Cell Lung Cancer, Ovarian Cancer, and Pancreatic Cancer (Table S9).

We used the R package MutationalPatterns (version 3.12.0) to obtain the six types of base substitutions (i.e., C>A, C>G, C>T, T>A, T>C, and T>G) for CKP and CUP, and analyzed their correlations between CKP and CUP. Additionally, we generated 96 mutational profiles for CKP and CUP samples and calculated their cosine similarity with known mutational signatures obtained via the get_known_signatures function.

Independent validation of OncoChat for CUP

We further validated OncoChat using an independent dataset from a prior study1 comprising 26 CUP samples subsequently confirmed to have specific cancer types (Table S7). After preprocessing, these samples were evaluated with OncoChat and baseline methods OncoNPC and GDD-ENS for comparative predictions.

Empirical treatment and OncoChat prediction

We analyzed an additional dataset of 158 CUP cases, referred to as the Empirical Treatments Cohort.13 These cases included recorded empirical treatment regimens and longitudinal survival data. Due to data-sharing limitations and institutional privacy policies, this cohort did not include confirmed tissue-of-origin diagnoses. As a result, we conducted only a qualitative assessment: evaluating whether OncoChat’s predictions aligned with empirical treatments. Theoretically, if the empirical treatment matches the predicted cancer type, patient survival may be extended.

We defined concordance as the alignment between the cancer type predicted by OncoChat and the empirical treatment administered, based on established clinical guidelines (e.g., NCCN). Each case was categorized as either concordant (treatment matched the predicted cancer type) or not concordant (treatment deviated from the predicted type).

Overall survival differences between concordant and non-concordant groups were assessed using Kaplan–Meier analysis and the log rank test. To further evaluate the prognostic significance of concordance, we conducted a multivariable Cox proportional hazards regression analysis. Covariates included in the model were gender, copy number alteration (CNA) burden, metastatic site, and histological subtype. Hazard ratios (HRs) with corresponding 95% confidence intervals (CIs) and p-values were calculated.

Model interpretability

To elucidate the biological signals driving tumor classification in OncoChat, we performed a comprehensive interpretability analysis using TransformerLens, a mechanistic interpretability library designed for transformer-based language models. This framework enabled direct inspection of internal model activations and attention mechanisms across CKP samples spanning 17 tumor types from MSK Center.

Specifically, we applied activation patching (provided by TransformerLens), a causal intervention technique introduced in the ROME framework, to identify influential activations within OncoChat’s architecture.68 This method involves running the model on input A (e.g., a sample from tumor type 1), substituting intermediate activations with those from input B (e.g., a sample from tumor type 2), and measuring the resulting change in model output. By systematically applying this intervention across layers and positions, we identified genomic features with the greatest causal impact on tumor-type predictions. These features were aggregated to rank gene-level contributions across the dataset.

To determine the most predictive genomic alterations, we analyzed model outputs across all samples and quantified the relative contribution of each gene to the final classification logits. High-impact mutations were defined as those consistently ranked among the top features within samples of a given tumor type. These were cross-referenced with known driver mutations to evaluate biological plausibility. Feature saliency scores were visualized to reveal tumor-type-specific patterns.

To validate model predictions in cancers of unknown primary (CUP), we examined gene-level contributions in 22 of the 26 CUP cases that were confidently classified by OncoChat. Heatmaps were generated to illustrate the relative importance of each gene in contributing to the predicted tumor type.

We further assessed OncoChat’s capacity to capture biologically meaningful gene–gene relationships by analyzing attention patterns between BRCA1/2 and PARP1, a synthetic lethal gene pair implicated in breast cancer.47 Attention scores were extracted from self-attention heads across all transformer layers and averaged for each gene pair. We selected 21 breast cancer samples harboring co-occurring mutations in BRCA1/2 and PARP1 (based on clinical annotations from female patients) and compared their attention scores with 21 matched non–breast cancer samples sharing the same mutational profile. Statistical significance of differential attention was evaluated using the Mann–Whitney U test.

Quantification and statistical analysis

The experiments were conducted using Python (version 3.11.7), R (version 4.2.0), ggplot2 (version 3.4.2) and scikit-learn (version 1.4.1). The area under the Precision-Recall curve (PRAUC) was calculated using multiROC (version 1.1.1). The 95% CIs and p value were determined using bootstraps method for 1000 times. Kaplan-Meier survival analyses were conducted to assess the relationship between different groups, using the R packages survival (version 3.5.7) and survminer (version 0.4.9). Log rank tests were applied to compare survival curves.

Published: September 4, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.xcrm.2025.102332.

Contributor Information

Kexin Chen, Email: chenkexin@tmu.edu.cn.

Xiangchun Li, Email: lixiangchun@tmu.edu.cn.

Supplemental information

Document S1. Figures S1–S15 and Tables S2, S4, S6, S12, S14, and S15
mmc1.pdf (2.6MB, pdf)
Table S1. Demographic characteristics of tumor samples across various centers based on raw data, related to Figure 1
mmc2.xlsx (41KB, xlsx)
Table S3. Performance evaluation of OncoChat and baselines using different filtering thresholds for tumor classification, related to Figure 2
mmc3.xlsx (108.5KB, xlsx)
Table S5. Evaluation of OncoChat and baselines using different filtering thresholds for broad cancer classification, related to Figure 4
mmc4.xlsx (75.4KB, xlsx)
Table S7. Evaluation of OncoChat and baseline models using the independent CUP dataset, related to Figure 5
mmc5.xlsx (49.6KB, xlsx)
Table S8. Survival analysis conducted for CUP samples, related to Figure 5
mmc6.xlsx (123.1KB, xlsx)
Table S9. Mutation signature analysis for CUP and matched CKP samples, related to Figure 5
mmc7.xlsx (157.6KB, xlsx)
Table S10. Information about the targeted panels used in this study, related to Figure 1
mmc8.xlsx (20.4KB, xlsx)
Table S11. Detailed distribution of samples in the training and testing sets, related to Figure 1
mmc9.xlsx (53.6KB, xlsx)
Table S13. Clinical details of the 158 CUP cases included in the treatment analysis, related to Figure 5
mmc10.xlsx (21.8KB, xlsx)
Document S2. Article plus supplemental information
mmc11.pdf (5.9MB, pdf)

References

  • 1.Darmofal M., Suman S., Atwal G., Toomey M., Chen J.-F., Chang J.C., Vakiani E., Varghese A.M., Balakrishnan Rema A., Syed A., et al. Deep-Learning Model for Tumor-Type Prediction Using Targeted Clinical Genomic Sequencing Data. Cancer Discov. 2024;14:1064–1081. doi: 10.1158/2159-8290.CD-23-0996. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Hyman D.M., Puzanov I., Subbiah V., Faris J.E., Chau I., Blay J.-Y., Wolf J., Raje N.S., Diamond E.L., Hollebecque A., et al. Vemurafenib in Multiple Nonmelanoma Cancers with BRAF V600 Mutations. N. Engl. J. Med. 2015;373:726–736. doi: 10.1056/NEJMoa1502309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Malone E.R., Oliva M., Sabatini P.J.B., Stockley T.L., Siu L.L. Molecular profiling for precision cancer therapies. Genome Med. 2020;12 doi: 10.1186/s13073-019-0703-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Golub T.R., Slonim D.K., Tamayo P., Huard C., Gaasenbeek M., Mesirov J.P., Coller H., Loh M.L., Downing J.R., Caligiuri M.A., et al. Molecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring. Science. 1999;286:531–537. doi: 10.1126/science.286.5439.531. [DOI] [PubMed] [Google Scholar]
  • 5.Marquard A.M., Birkbak N.J., Thomas C.E., Favero F., Krzystanek M., Lefebvre C., Ferté C., Jamal-Hanjani M., Wilson G.A., Shafi S., et al. TumorTracer: a method to identify the tissue of origin from the somatic mutations of a tumor specimen. BMC Med. Genomics. 2015;8 doi: 10.1186/s12920-015-0130-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Soh K.P., Szczurek E., Sakoparnig T., Beerenwinkel N. Predicting cancer type from tumour DNA signatures. Genome Med. 2017;9 doi: 10.1186/s13073-017-0493-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Steeg P.S. Tumor metastasis: mechanistic insights and clinical challenges. Nat. Med. 2006;12:895–904. doi: 10.1038/nm1469. [DOI] [PubMed] [Google Scholar]
  • 8.Steeg P.S. Targeting metastasis. Nat. Rev. Cancer. 2016;16:201–218. doi: 10.1038/nrc.2016.25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Moran S., Martínez-Cardús A., Sayols S., Musulén E., Balañá C., Estival-Gonzalez A., Moutinho C., Heyn H., Diaz-Lagares A., de Moura M.C., et al. Epigenetic profiling to classify cancer of unknown primary: a multicentre, retrospective analysis. Lancet Oncol. 2016;17:1386–1395. doi: 10.1016/S1470-2045(16)30297-2. [DOI] [PubMed] [Google Scholar]
  • 10.Varghese A.M., Arora A., Capanu M., Camacho N., Won H.H., Zehir A., Gao J., Chakravarty D., Schultz N., Klimstra D.S., et al. Clinical and molecular characterization of patients with cancer of unknown primary in the modern era. Ann. Oncol. 2017;28:3015–3021. doi: 10.1093/annonc/mdx545. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Rassy E., Pavlidis N. Progress in refining the clinical management of cancer of unknown primary in the molecular era. Nat. Rev. Clin. Oncol. 2020;17:541–554. doi: 10.1038/s41571-020-0359-1. [DOI] [PubMed] [Google Scholar]
  • 12.Lee M.S., Sanoff H.K. Cancer of unknown primary. BMJ. 2020;371 doi: 10.1136/bmj.m4050. [DOI] [PubMed] [Google Scholar]
  • 13.Moon I., LoPiccolo J., Baca S.C., Sholl L.M., Kehl K.L., Hassett M.J., Liu D., Schrag D., Gusev A. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary. Nat. Med. 2023;29:2057–2067. doi: 10.1038/s41591-023-02482-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Varadhachary G.R., Raber M.N. Cancer of Unknown Primary Site. N. Engl. J. Med. 2014;371:757–765. doi: 10.1056/NEJMra1303917. [DOI] [PubMed] [Google Scholar]
  • 15.Pavlidis N., Briasoulis E., Hainsworth J., Greco F.A. Diagnostic and therapeutic management of cancer of an unknown primary. Eur. J. Cancer. 2003;39:1990–2005. doi: 10.1016/S0959-8049(03)00547-1. [DOI] [PubMed] [Google Scholar]
  • 16.Posner A., Prall O.W., Sivakumaran T., Etemadamoghadam D., Thio N., Pattison A., Balachander S., Fisher K., Webb S., Wood C., et al. A comparison of DNA sequencing and gene expression profiling to assist tissue of origin diagnosis in cancer of unknown primary. J. Pathol. 2023;259:81–92. doi: 10.1002/path.6022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Anderson G.G., Weiss L.M. Determining tissue of origin for metastatic cancers: meta-analysis and literature review of immunohistochemistry performance. Appl. Immunohistochem. Mol. Morphol. 2010;18:3–8. doi: 10.1097/PAI.0b013e3181a75e6d. [DOI] [PubMed] [Google Scholar]
  • 18.Nguyen L., Van Hoeck A., Cuppen E. Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features. Nat. Commun. 2022;13 doi: 10.1038/s41467-022-31666-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Dietlein F., Eschner W. Inferring primary tumor sites from mutation spectra: a meta-analysis of histology-specific aberrations in cancer-derived cell lines. Hum. Mol. Genet. 2014;23:1527–1537. doi: 10.1093/hmg/ddt539. [DOI] [PubMed] [Google Scholar]
  • 20.Jiao W., Atwal G., Polak P., Karlic R., Cuppen E., PCAWG Tumor Subtypes and Clinical Translation Working Group. Danyi A., de Ridder J., van Herpen C., Lolkema M.P., et al. A deep learning system accurately classifies primary and metastatic cancers using passenger mutation patterns. Nat. Commun. 2020;11 doi: 10.1038/s41467-019-13825-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Sanjaya P., Maljanen K., Katainen R., Waszak S.M., Genomics England Research Consortium. Aaltonen L.A., Stegle O., Korbel J.O., Pitkänen E., Boustred C.R., et al. Mutation-Attention (MuAt): deep representation learning of somatic mutations for tumour typing and subtyping. Genome Med. 2023;15 doi: 10.1186/s13073-023-01204-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Ri I., Kawata J., Nagai A., Muto K. Expectations, concerns, and attitudes regarding whole-genome sequencing studies: a survey of cancer patients, families, and the public in Japan. J. Hum. Genet. 2023;68:281–285. doi: 10.1038/s10038-022-01100-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Cuppen E., Elemento O., Rosenquist R., Nikic S., IJzerman M., Zaleski I.D., Frederix G., Levin L.Å., Mullighan C.G., Buettner R., et al. Implementation of Whole-Genome and Transcriptome Sequencing Into Clinical Cancer Care. JCO Precis. Oncol. 2022;6 doi: 10.1200/PO.22.00245. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cheng D.T., Mitchell T.N., Zehir A., Shah R.H., Benayed R., Syed A., Chandramohan R., Liu Z.Y., Won H.H., Scott S.N., et al. Memorial Sloan Kettering-Integrated Mutation Profiling of Actionable Cancer Targets (MSK-IMPACT): A Hybridization Capture-Based Next-Generation Sequencing Clinical Assay for Solid Tumor Molecular Oncology. J. Mol. Diagn. 2015;17:251–264. doi: 10.1016/j.jmoldx.2014.12.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.AACR Project GENIE Consortium AACR Project GENIE: Powering Precision Medicine through an International Consortium. Cancer Discov. 2017;7:818–831. doi: 10.1158/2159-8290.CD-17-0151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Pugh T.J., Bell J.L., Bruce J.P., Doherty G.J., Galvin M., Green M.F., Hunter-Zinck H., Kumari P., Lenoue-Newton M.L., Li M.M., et al. AACR Project GENIE: 100,000 Cases and Beyond. Cancer Discov. 2022;12:2044–2057. doi: 10.1158/2159-8290.CD-21-1547. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Castaldi P.J., Dahabreh I.J., Ioannidis J.P.A. An empirical assessment of validation practices for molecular classifiers. Brief. Bioinform. 2011;12:189–202. doi: 10.1093/bib/bbq073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Baranovskii A., Gündüz I.B., Franke V., Uyar B., Akalin A. Multi-Omics Alleviates the Limitations of Panel Sequencing for Cancer Drug Response Prediction. Cancers. 2022;14 doi: 10.3390/cancers14225604. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Aldea M., Vasseur D., Italiano A., Nikolaev S.I. WGS/WES-RNAseq compared to targeted NGS in oncology: is there something to unlock? Ann. Oncol. 2023;34:1090–1093. doi: 10.1016/j.annonc.2023.09.3118. [DOI] [PubMed] [Google Scholar]
  • 30.El Bairi K., Azzam F., Trapani D., Ouled Amar Bencheikh B. In: Illuminating Colorectal Cancer Genomics by Next-Generation Sequencing: A Big Chapter in the Tale. El Bairi K., editor. Springer International Publishing; 2020. Overview of Cost-Effectiveness and Limitations of Next-Generation Sequencing in Colorectal Cancer; pp. 173–185. [DOI] [Google Scholar]
  • 31.Ji Y., Zhou Z., Liu H., Davuluri R.V. DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics. 2021;37:2112–2120. doi: 10.1093/bioinformatics/btab083. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Carl N., Schramm F., Haggenmüller S., Kather J.N., Hetz M.J., Wies C., Michel M.S., Wessels F., Brinker T.J. Large language model use in clinical oncology. npj Precis. Onc. 2024;8:1–17. doi: 10.1038/s41698-024-00733-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Lee P., Bubeck S., Petro J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 2023;388:1233–1239. doi: 10.1056/NEJMsr2214184. [DOI] [PubMed] [Google Scholar]
  • 34.Iannantuono G.M., Bracken-Clarke D., Floudas C.S., Roselli M., Gulley J.L., Karzai F. Applications of large language models in cancer care: current evidence and future perspectives. Front. Oncol. 2023;13 doi: 10.3389/fonc.2023.1268915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Radford A., Narasimhan K., Salimans T., Sutskever I. Vol. 12. OpenAI; 2018. Improving Language Understanding by Generative Pre-training. [Google Scholar]
  • 36.Lu M.Y., Chen B., Williamson D.F.K., Chen R.J., Zhao M., Chow A.K., Ikemura K., Kim A., Pouli D., Patel A., et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634:466–473. doi: 10.1038/s41586-024-07618-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Sorin V., Klang E., Sklair-Levy M., Cohen I., Zippel D.B., Balint Lahat N., Konen E., Barash Y. Large language model (ChatGPT) as a support tool for breast tumor board. npj Breast Cancer. 2023;9 doi: 10.1038/s41523-023-00557-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Benary M., Wang X.D., Schmidt M., Soll D., Hilfenhaus G., Nassir M., Sigler C., Knödler M., Keller U., Beule D., et al. Leveraging Large Language Models for Decision Support in Personalized Oncology. JAMA Netw. Open. 2023;6 doi: 10.1001/jamanetworkopen.2023.43689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Yeo Y.H., Samaan J.S., Ng W.H., Ting P.-S., Trivedi H., Vipani A., Ayoub W., Yang J.D., Liran O., Spiegel B., Kuo A. Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma. Clin. Mol. Hepatol. 2023;29:721–732. doi: 10.3350/cmh.2023.0089. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Dennstädt F., Hastings J., Putora P.M., Vu E., Fischer G.F., Süveg K., Glatzer M., Riggenbach E., Hà H.-L., Cihoric N. Exploring Capabilities of Large Language Models such as ChatGPT in Radiation Oncology. Adv. Radiat. Oncol. 2024;9 doi: 10.1016/j.adro.2023.101400. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Moor M., Banerjee O., Abad Z.S.H., Krumholz H.M., Leskovec J., Topol E.J., Rajpurkar P. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–265. doi: 10.1038/s41586-023-05881-4. [DOI] [PubMed] [Google Scholar]
  • 42.Noerenberg D., Briest F., Hennch C., Yoshida K., Hablesreiter R., Takeuchi Y., Ueno H., Staiger A.M., Ziepert M., Asmar F., et al. Genetic Characterization of Primary Mediastinal B-Cell Lymphoma: Pathogenesis and Patient Outcomes. J. Clin. Oncol. 2024;42:452–466. doi: 10.1200/JCO.23.01053. [DOI] [PubMed] [Google Scholar]
  • 43.Northcott P.A., Shih D.J.H., Peacock J., Garzia L., Morrissy A.S., Zichner T., Stütz A.M., Korshunov A., Reimand J., Schumacher S.E., et al. Subgroup-specific structural variation across 1,000 medulloblastoma genomes. Nature. 2012;488:49–56. doi: 10.1038/nature11327. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wang S., Ma P., Jiang N., Jiang Y., Yu Y., Fang Y., Miao H., Huang H., Tang Q., Cui D., et al. Rare tumors: a blue ocean of investigation. Front. Med. 2023;17:220–230. doi: 10.1007/s11684-023-0984-z. [DOI] [PubMed] [Google Scholar]
  • 45.Cobain E.F., Wu Y.-M., Vats P., Chugh R., Worden F., Smith D.C., Schuetze S.M., Zalupski M.M., Sahai V., Alva A., et al. Assessment of Clinical Benefit of Integrative Genomic Profiling in Advanced Solid Tumors. JAMA Oncol. 2021;7:525–533. doi: 10.1001/jamaoncol.2020.7987. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Alexandrov L.B., Kim J., Haradhvala N.J., Huang M.N., Tian Ng A.W., Wu Y., Boot A., Covington K.R., Gordenin D.A., Bergstrom E.N., et al. The repertoire of mutational signatures in human cancer. Nature. 2020;578:94–101. doi: 10.1038/s41586-020-1943-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Huang A., Garraway L.A., Ashworth A., Weber B. Synthetic lethality as an engine for cancer drug target discovery. Nat. Rev. Drug Discov. 2020;19:23–38. doi: 10.1038/s41573-019-0046-z. [DOI] [PubMed] [Google Scholar]
  • 48.Dubois F., Sidiropoulos N., Weischenfeldt J., Beroukhim R. Structural variations in cancer and the 3D genome. Nat. Rev. Cancer. 2022;22:533–546. doi: 10.1038/s41568-022-00488-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Priestley P., Baber J., Lolkema M.P., Steeghs N., de Bruijn E., Shale C., Duyvesteyn K., Haidari S., van Hoeck A., Onstenk W., et al. Pan-cancer whole-genome analyses of metastatic solid tumours. Nature. 2019;575:210–216. doi: 10.1038/s41586-019-1689-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Qin C., Cao Q., Ju X., Wang M., Meng X., Zhu J., Yan F., Li P., Ding Q., Chen J., et al. The polymorphisms in the VHL and HIF1A genes are associated with the prognosis but not the development of renal cell carcinoma. Ann. Oncol. 2012;23:981–989. doi: 10.1093/annonc/mdr325. [DOI] [PubMed] [Google Scholar]
  • 51.Dow L.E., O’Rourke K.P., Simon J., Tschaharganeh D.F., van Es J.H., Clevers H., Lowe S.W. Apc Restoration Promotes Cellular Differentiation and Reestablishes Crypt Homeostasis in Colorectal Cancer. Cell. 2015;161:1539–1552. doi: 10.1016/j.cell.2015.05.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Pilié P.G., Tang C., Mills G.B., Yap T.A. State-of-the-art strategies for targeting the DNA damage response in cancer. Nat. Rev. Clin. Oncol. 2019;16:81–104. doi: 10.1038/s41571-018-0114-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Lu M.Y., Chen T.Y., Williamson D.F.K., Zhao M., Shady M., Lipkova J., Mahmood F. AI-based pathology predicts origins for cancers of unknown primary. Nature. 2021;594:106–110. doi: 10.1038/s41586-021-03512-4. [DOI] [PubMed] [Google Scholar]
  • 54.Xu H., Usuyama N., Bagga J., Zhang S., Rao R., Naumann T., Wong C., Gero Z., González J., Gu Y., et al. A whole-slide foundation model for digital pathology from real-world data. Nature. 2024;630:181–188. doi: 10.1038/s41586-024-07441-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Tian F., Liu D., Wei N., Fu Q., Sun L., Liu W., Sui X., Tian K., Nemeth G., Feng J., et al. Prediction of tumor origin in cancers of unknown primary origin with cytology-based deep learning. Nat. Med. 2024;30:1309–1319. doi: 10.1038/s41591-024-02915-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Vorontsov E., Bozkurt A., Casson A., Shaikovski G., Zelechowski M., Severson K., Zimmermann E., Hall J., Tenenholtz N., Fusi N., et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat. Med. 2024;30:2924–2935. doi: 10.1038/s41591-024-03141-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zhao Y., Pan Z., Namburi S., Pattison A., Posner A., Balachander S., Paisie C.A., Reddi H.V., Rueter J., Gill A.J., et al. CUP-AI-Dx: A tool for inferring cancer tissue of origin and molecular subtype using RNA gene-expression data and artificial intelligence. EBioMedicine. 2020;61 doi: 10.1016/j.ebiom.2020.103030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Shen Y., Chu Q., Yin X., He Y., Bai P., Wang Y., Fang W., Timko M.P., Fan L., Jiang W. TOD-CUP: a gene expression rank-based majority vote algorithm for tissue origin diagnosis of cancers of unknown primary. Brief. Bioinform. 2021;22:2106–2118. doi: 10.1093/bib/bbaa031. [DOI] [PubMed] [Google Scholar]
  • 59.Yu K., Chen B., Aran D., Charalel J., Yau C., Wolf D.M., van‘t Veer L.J., Butte A.J., Goldstein T., Sirota M. Comprehensive transcriptomic analysis of cell lines as models of primary tumors across 22 tumor types. Nat. Commun. 2019;10 doi: 10.1038/s41467-019-11415-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Lautizi M., Baumbach J., Weichert W., Steiger K., List M., Pfarr N., Kacprowski T. The limits of molecular signatures for pancreatic ductal adenocarcinoma subtyping. NAR Cancer. 2022;4 doi: 10.1093/narcan/zcac030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Conway A.-M., Pearce S.P., Clipson A., Hill S.M., Chemi F., Slane-Tan D., Ferdous S., Hossain A.S.M.M., Kamieniecka K., White D.J., et al. A cfDNA methylation-based tissue-of-origin classifier for cancers of unknown primary. Nat. Commun. 2024;15 doi: 10.1038/s41467-024-47195-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Lin J.-R., Chen Y.-A., Campton D., Cooper J., Coy S., Yapp C., Tefft J.B., McCarty E., Ligon K.L., Rodig S.J., et al. High-plex immunofluorescence imaging and traditional histology of the same tissue section for discovering image-based biomarkers. Nat. Cancer. 2023;4:1036–1052. doi: 10.1038/s43018-023-00576-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Bollhagen A., Bodenmiller B. Highly Multiplexed Tissue Imaging in Precision Oncology and Translational Cancer Research. Cancer Discov. 2024;14:2071–2088. doi: 10.1158/2159-8290.CD-23-1165. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Li B., Li G., Yan X., Zhu D., Lin P.P., Wang Z., Qu H., He X., Fu Y., Zhu X., et al. Fresh Tissue Multi-omics Profiling Reveals Immune Classification and Suggests Immunotherapy Candidates for Conventional Chondrosarcoma. Clin. Cancer Res. 2021;27:6543–6558. doi: 10.1158/1078-0432.CCR-21-1893. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Bai J., Bai S., Chu Y., Cui Z., Dang K., Deng X., Fan Y., Ge W., Han Y., Huang F., et al. Qwen Technical Report. arXiv. 2023 doi: 10.48550/arXiv.2309.16609. Preprint at. [DOI] [Google Scholar]
  • 66.Gu A., Dao T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv. 2024 doi: 10.48550/arXiv.2312.00752. Preprint at. [DOI] [Google Scholar]
  • 67.Iyer S., Lin X.V., Pasunuru R., Mihaylov T., Simig D., Yu P., Shuster K., Wang T., Liu Q., Koura P.S., et al. OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization. arXiv. 2023 doi: 10.48550/arXiv.2212.12017. Preprint at. [DOI] [Google Scholar]
  • 68.Meng K., Bau D., Andonian A., Belinkov Y. Locating and Editing Factual Associations in GPT. Adv. Neural Inform. Process. Syst. 2022;35:17359–17372. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S15 and Tables S2, S4, S6, S12, S14, and S15
mmc1.pdf (2.6MB, pdf)
Table S1. Demographic characteristics of tumor samples across various centers based on raw data, related to Figure 1
mmc2.xlsx (41KB, xlsx)
Table S3. Performance evaluation of OncoChat and baselines using different filtering thresholds for tumor classification, related to Figure 2
mmc3.xlsx (108.5KB, xlsx)
Table S5. Evaluation of OncoChat and baselines using different filtering thresholds for broad cancer classification, related to Figure 4
mmc4.xlsx (75.4KB, xlsx)
Table S7. Evaluation of OncoChat and baseline models using the independent CUP dataset, related to Figure 5
mmc5.xlsx (49.6KB, xlsx)
Table S8. Survival analysis conducted for CUP samples, related to Figure 5
mmc6.xlsx (123.1KB, xlsx)
Table S9. Mutation signature analysis for CUP and matched CKP samples, related to Figure 5
mmc7.xlsx (157.6KB, xlsx)
Table S10. Information about the targeted panels used in this study, related to Figure 1
mmc8.xlsx (20.4KB, xlsx)
Table S11. Detailed distribution of samples in the training and testing sets, related to Figure 1
mmc9.xlsx (53.6KB, xlsx)
Table S13. Clinical details of the 158 CUP cases included in the treatment analysis, related to Figure 5
mmc10.xlsx (21.8KB, xlsx)
Document S2. Article plus supplemental information
mmc11.pdf (5.9MB, pdf)

Data Availability Statement


Articles from Cell Reports Medicine are provided here courtesy of Elsevier

RESOURCES