Abstract
Artificial intelligence (AI)-systems can improve cancer diagnosis, yet their development often relies on subjective histological features as ground truth for training. Here, we developed an AI-model applied to histological whole-slide images (WSIs) using CDH1 bi-allelic mutations, pathognomonic for invasive lobular carcinoma (ILC) in breast neoplasms, as ground truth. The model accurately predicted CDH1 bi-allelic mutations (accuracy=0.95) and diagnosed ILC (accuracy=0.96). A total of 74% of samples classified by the AI-model as having CDH1 bi-allelic mutations but lacking these alterations displayed alternative CDH1 inactivating mechanisms, including a deleterious CDH1 fusion gene and non-coding CDH1 genetic alterations. Analysis of internal and external validation cohorts demonstrated 0.95 and 0.89 accuracy for ILC diagnosis, respectively. The latent features of the AI-model correlated with human-explainable histopathologic features. Taken together, this study reports the construction of an AI-algorithm trained using a genetic rather than histologic ground truth that can robustly classify ILCs and uncover CDH1 inactivating mechanisms, providing the basis for orthogonal ground truth utilization for development of diagnostic AI-models applied to WSI.
Keywords: Artificial intelligence, breast cancer, genomics
INTRODUCTION
Artificial intelligence (AI)-based systems for detection and classification of histologic specimens from human cancers have the potential to improve cancer outcomes by facilitating the detection of histologic entities and specific phenotypes with important clinical implications (1). Traditionally, development of AI-based cancer classification systems employs annotation of digital pathology slides (2) or histopathologic diagnostic labels (3) as ground truth for training. Albeit effective (2,3), these approaches are inevitably confounded by human subjectivity and may only replicate pathologists’ accuracy. A subset of breast cancers with biology and clinical behavior distinct from invasive ductal carcinoma of no-special type (IDC-NST), the common form of the disease, are characterized by pathognomonic genetic alterations (4), which can be considered essentially diagnostic (4).
Invasive lobular carcinoma (ILC), the second most common breast cancer type, is phenotypically characterized by tumor cell discohesion, caused by bi-allelic inactivation of CDH1, that encodes for the cell-cell adhesion protein E-cadherin. CDH1 pathogenic mutations coupled with loss of heterozygosity of the wild-type allele (LOH) are detected in ~80% of ILCs (5), but found in only <1% of non-ILC breast cancers (6), constituting a genotypic-phenotypic correlation in breast neoplasms. ILCs have a distinctive biology and clinical behavior, including lower response rates to chemotherapy and potentially distinct genetic vulnerabilities (7). Several studies have demonstrated only modest interobserver agreement (~0.7) in ILC diagnosis (8,9). Factors contributing to this low reproducibility include the inherent subjectivity of histopathologic examination, lack of diagnostic standardization (10), and imperfect correlation of E-cadherin loss and CDH1 alterations, exemplified by aberrant E-cadherin expression in 16%–24% of ILCs (11,12), and E-cadherin loss in non-ILC high-grade basal-like breast cancers (13).
Challenges in ILC diagnosis have hindered its incorporation in therapeutic decision-making and the development of ILC-specific clinical trials. Thus, a robust methodology for ILC diagnosis constitutes an unmet clinical need. The low interobserver agreement in ILC diagnosis poses challenges for the development of an AI-based classification of ILCs using a traditional histopathologic ground truth for training. The strong genotypic-phenotypic correlation in ILC presents a unique opportunity to integrate AI and genetics in a diagnostic system. To overcome the subjectivity/variability inherent to ILC histologic labels/features, we utilized CDH1 bi-allelic mutations as orthogonal ground truth to develop an AI-model for ILC diagnosis. This approach is expected to outperform histologic examination for ILC diagnosis, as we anticipate that a genetic alteration would offer a cleaner/less noisy ground truth than traditional histologic features/labels. Furthermore, the AI-model trained on CDH1 bi-allelic mutations would overcome issues related to aberrant/non-specific E-cadherin expression (11–13), and mitigate human subjectivity, and arbitrary region selection for analysis (14).
In this proof-of-concept study, we leveraged a genotypic-phenotypic correlation, utilizing CDH1 bi-allelic mutations, instead of histologic labels, as orthogonal genetic ground truth for the training, development, and external validation of an AI-model to diagnose ILC, mitigating the human subjectivity of diagnostic AI-models. Through the analysis of samples lacking CDH1 bi-allelic mutations but classified as ILC by the AI-algorithm, we uncovered novel biological mechanisms of CDH1 inactivation, supporting that ILC represents a convergent phenotype (4).
MATERIALS AND METHODS
Samples and study population
The study was approved by Memorial Sloan Kettering Cancer Center (MSK) Institutional Review Board. The AI-model was developed by Paige AI. 2,213 estrogen receptor-positive primary breast cancers were included in this study.
The 1,462 samples included in the MSK development (n=1,057) and validation cohorts (n=405) had been previously subjected to targeted sequencing using the MSK-Integrated Mutation Profiling of Actionable Targets (MSK-IMPACT) assay. 187 and 84 breast cancer samples corresponded to ILCs (Supplementary Table S1), and 870 and 321 breast cancer samples to non-ILCs in the MSK development and validation cohorts, respectively.
No access to images or labels of the MSK validation cohort was available to Paige prior to the completion of the final model developed in the MSK development cohort. 751 digital histologic whole-slide images retrieved from The Cancer Genome Atlas (TCGA) study together with their respective histologic diagnosis as reported in the TCGA landmark study (15) were employed as the external validation dataset. The TCGA validation cohort (n=751) included 201 ILCs and 550 non-ILCs.
Assessment of CDH1 mutational status in targeted sequencing data
MSK-IMPACT targeted sequencing data were analyzed to determine the CDH1 mutational status of the breast cancer samples in the MSK development and validation cohorts of this study. The list of CDH1 mutational variants was retrieved from cBioPortal (http://www.cbioportal.org). Raw MSK-IMPACT targeted sequencing data (FASTQ files) were reprocessed with a benchmarked bioinformatics pipeline to infer LOH, which was determined using FACETS (16), as previously described (5). CDH1 inactivating mutations associated with LOH or a second somatic loss-of-function mutation were considered as ‘CDH1 bi-allelic mutations’.
Development of an AI-based CDH1 bi-allelic mutation classifier in breast cancer using whole-slide images
We developed a machine learning-based system for detecting CDH1 bi-allelic mutations in digital whole-slide images of invasive breast cancer slides stained with hematoxylin-and-eosin (H&E).
The AI-system is based on a two-stage model architecture comprising i) a breast cancer feature extractor and ii) a CDH1 bi-allelic mutation classifier.
Breast cancer feature extractor
The model architecture, training strategy and code of the breast cancer feature extractor, that uses the same architecture as Paige Prostate (17) are described in Campanella et al (3). The breast cancer feature extractor consists of a convolutional neural network (CNN) based on a squeeze and excitation residual network (SE-ResNet-50) (18) architecture that computes a 512-dimensional embedding vector for each tile.
In brief, H&E whole-slide images of all cases were identified and digitized using an Aperio AT2 scanner scanned at 20x magnification (0.5 um/pixel). Foreground filtering was applied to identify the tissue regions in the whole-slide images. Each slide was broken up into 224×224-pixel tiles. Tiles that overlapped with the identified tissue regions were selected and processed by the breast cancer feature extractor to compute the feature embeddings for each tile. The SEResNET-50 computes a 512-dimensional embedding vector and a classification score for each tile. Classification groups include pre-malignant lesions, in situ carcinoma or invasive carcinoma. A tile was considered benign if all cancerous prediction scores were lower than the pre-defined threshold. The maximum tile score was used as final prediction score for the slide, and a decision threshold of 0.1 was applied to the slide score to determine the final classification. This threshold, determined during model tuning, was designed to favor high sensitivity. By applying a cutoff value of 0.1 to the continuous score, it generated a binary classification of invasive cancer for each tile.
CDH1 bi-allelic mutation classifier
In the second stage of the system, embedding vectors for tiles with scores surpassing the invasive cancer prediction score threshold (i.e., 0.1) were used to train an attention-based aggregation network for the detection of CDH1 bi-allelic mutations.
The CDH1 bi-allelic mutation classifier consists of two convolutional layers and global soft attention, followed by a linear layer per classification output head. The tile-selection constrained the system to only look at the invasive tissue determined by the breast cancer feature extractor. A feed-forward aggregation network was trained using a 10-fold cross-validation system and the inference probabilities for both tune and test sets were generated for each fold. The CDH1 bi-allelic mutation classifier aggregates all tile embeddings, and the outer layer employs a logistic sigmoid activation function, generating a continuous value ranging from 0 (indicating the absence of CDH1 bi-allelic mutation) to 1 (indicating the presence of CDH1 bi-allelic mutation).
Training of the CDH1 bi-allelic mutation classifier
Training of the 10 aggregator models on the extracted features (i.e., tile embeddings) was conducted on an AWS EC2 P4d instance [https://aws.amazon.com/ec2/instance-types/p4/], and the training time for each model was approximately 2.5 hours. On average, the training set for each aggregator comprised approximately 20.2 million tissue tiles, each with a size of 224×224 pixels at 0.5 microns per pixel (mpp). The aggregator model training procedure was parameterized as follows:
| Training Parameter Name | Training Parameter Value |
|---|---|
| Optimizer | Adam [https://arxiv.org/abs/1412.6980] |
| Learning rate | 0.001 |
| Weight decay | 0.0 |
| Learning rate schedule | Decay learning rate every 40 epochs |
| Learning rate decay rate | 0.5 |
| Number of Epochs | 128 |
| Batch size | 16 |
| Model dropout ratio | 0.2 |
| Training loss | Binary Cross-Entropy |
The 10-fold development cohort comprised a total of approximately 33.6 million tissue tiles, with an average of 17,500 tissue tiles per slide.
Computational requirements to run the CDH1 bi-allelic mutation classifier
The runtime metrics to run the trained system comprising the breast feature extractor and aggregator model in inference mode on an AWS EC2 P4d instance were as follows:
| Runtime metric | Runtime value |
|---|---|
| Tissue tiles per second | 150.7 tiles/sec (std: 7.17 tiles/sec) |
| Minutes per slide | 1.92 minutes/slide (std: 0.325 minutes/slide) |
Selection of operating threshold for CDH1-bi-allelic mutation classifier
By applying 10-fold cross-validation, the development cohort of 1,057 samples was split into ten partitions with a 6:3:1 ratio for the train, tune, and test sets, respectively. A model was trained for each fold, resulting in 10 individual aggregator models. The best checkpoint of each aggregator model was selected by maximizing the sum of the average precision and receiver operating characteristic (ROC) area under the curve (AUC). The best aggregator model checkpoint was then used to run inference on the test set of each fold. For inference on independent cohorts, the ten models were combined into a single ensemble. To select a universal operating threshold for the models trained from 10-fold cross-validation, the inference probabilities of the tune set from each fold were used for calibration utilizing the CalibratedClassifierCV of scikit-learn (RRID:SCR_002577; scikit-learn version 1.1.0) (19) with cv=“prefit” and LogisticRegression for base_estimator, and method=sigmoid as the default value. This corresponds to the logistic regression model based on Platt’s method, that scales the output of a classification model into probability values. This process entails learning a logistic regression model to transform prediction scores into calibrated probabilities. The fitted classifier from each fold was then used to calibrate the probabilities for the fold-specific test set. The model performance was evaluated on the combined calibrated test probabilities, and the universal operating threshold was determined by optimizing the F1-score. The universal operating threshold was used to generate sample-level binary predictions for the presence or absence of CDH1 bi-allelic mutation.
Performance metrics
The performance of the AI-model for CDH1 bi-allelic mutations detection, and for classification of ILC was assessed using the ROC AUC and precision-recall curve. Sensitivity, specificity, accuracy and precision were also determined. 95% confidence intervals (CIs) for these metrics were estimated using the Wilson method (20).
Assessment of false positive cases using targeted sequencing data
The genetic underpinning of BCs predicted to harbor CDH1 bi-allelic mutations by the AI-model, but lacking this genetic alteration (i.e., false positive samples) in the MSK development and validation cohorts, was interrogated using the MSK-IMPACT sequencing data for alternative genetic mechanisms of CDH1 bi-allelic inactivation. Homozygous deletions were identified using FACETS (16). Intragenic deletions were assessed using the structural variant caller Delly (21) and confirmed through manual inspection of the data using the Integrative Genomics Viewer (IGV) (22). The presence of non-coding genetic alterations in CDH1 was investigated by analyzing intronic regions with a sequence coverage depth of at least 200x. Mutations within 50 base pairs from exon/intron junctions were included. Candidate variants were filtered based on variant allele frequency (≥0.1) in their presence in Genome Aggregation Database (gnomAD; v2.1.1) (23). All candidate non-coding CDH1 variants were evaluated using IGV (22). The functional significance of the identified intronic CDH1 variants identified was reviewed using SpliceAI (24) and ClinVar (25) annotations. Their bi-allelic status was assessed through the examination of allele-specific copy number using FACETS (16).
CDH1 methylation assessment by digital droplet PCR
We assessed the CDH1 gene promoter methylation status by digital droplet PCR (ddPCR) of cases included in the MSK development and validation cohorts that were false positive for CDH1 bi-allelic mutation detection by the AI-system and lacked alternative genetic CDH1 inactivating mechanisms as detected by MSK-IMPACT. In brief, bisulfite treated genomic DNA was quantified using a PicoGreen system, and 0.2–9 ng of DNA were combined with locus-specific primers designed for the two CDH1 gene promoter CpG islands, as well as FAM-labeled and HEX-labeled probes, HaeIII, and digital PCR Supermix for probes (no dUTP).
CpG Methylated DNA (ThermoFisher, Waltham, MA) was used a positive control and Universal Unmethylated DNA (Millipore, Burlington, MA) as negative control. ddPCR reactions were conducted on a QX200 ddPCR system (Bio-Rad, Hercules, CA). Two technical duplicates were used per sample. Using the QX200 droplet generator, reactions were partitioned into ~41K droplets per well. Emulsified PCRs were then run on a 96-well thermal cycler as follows: 95°C for 10 minutes; 50 cycles at 94°C for 1 minute and 54°C for 2 minutes; 98°C for 10 minutes. The number of droplets positive for CDH1 gene promoter methylated, unmethylated, both, or neither were determined using a QuantaSoft software (Bio-Rad, Hercules, CA). Methylation Frequency (MF) was inferred as MF = 100 * Methylated / (Methylated + Unmethylated). CDH1 promoter methylation in a sample was defined as MF≥1% of accepted droplets for CpG island 1 and/or CpG island 2.
Validation on independent MSK and TCGA cohorts
To assess the model performance on an independent MSK validation cohort of 405 samples and on an external independent validation cohort of 751 samples from TCGA, the ten trained models from the development cohort were run on the two independent validation cohorts, separately. For each of the validation cohorts, ten inference probabilities were generated for each sample, and the corresponding ten calibrated classifiers were then applied to the inference probabilities. The universal optimal operating threshold determined from the MSK development cohort was applied to the average calibrated probabilities to generate a binary prediction for each sample.
Whole-genome sequencing
DNA from microdissected tumor and normal formalin-fixed paraffin-embedded (FFPE) samples from a classic ILC lacking CDH1 genetic alterations by targeted sequencing and lacking CDH1 gene promoter methylation by ddPCR was subjected to whole-genome sequencing at the MSKCC Integrated Genomics Operations (IGO) using validated protocols. Mean sequencing coverage depth of 69x and 37x genome-wide was obtained from tumor and normal samples, respectively. Whole-genome sequencing data were analyzed using a validated bioinformatics pipeline.
In brief, following the alignment of sequence reads to the reference human genome GRCh37/hg19 using the Burrows-Wheeler Aligner (BWA (RRID:SCR_010910) v0.7.15) (26), somatic single nucleotide variants (SNVs) were detected with MuTect (RRID:SCR_000559; v1.0) (27), and insertion and deletions (indels) using Strelka (RRID:SCR_005109; v2.0.15) (28), VarScan2 (v2.3.7) (29), Platypus (RRID:SCR_005389; v0.8.1) (30) and Scalpel (v0.5.3) (31). Copy number alterations and LOH were determined using FACETS (16). The cancer cell fraction (CCF) of each mutation was inferred using ABSOLUTE (32). Single base substitution (SBS) mutational signatures (COSMIC - Catalogue Of Somatic Mutations In Cancer (RRID:SCR_002260) v3.1) were determined using Sigprofiler (33) and Signal (34). Indel mutational signatures (COSMIC - Catalogue Of Somatic Mutations In Cancer (RRID:SCR_002260) v3.1) were inferred with Sigprofiler (33) using the SigProfilerExtractor Package. Structural variants (SVs) were detected using the combination of Manta (35), SvABA (36) and GRIDSS2 (37) for tumor and matched normal using WGS data, as previously described (38). The resulting SVs were combined based on the SV type, strand and constraining the genomic coordinates of breakpoint to 0.5 kbp. Circos plots depicting SV calls, SNVs, indels and copy number alterations were generated using the signature.tools.lib R package (34).
Central pathology review of samples in the MSK development and validation cohorts
Samples that were misclassified by the AI-model as false positive or false negative for ILC diagnosis were centrally reviewed by four pathologists (FP, FD, LCC and JSR-F). Samples classified as true positives for ILC diagnosis were centrally reviewed by two pathologists (FP and LCC). Histopathologic review and ILC variant classification were conducted following the current World Health Organization (WHO) criteria (39). Samples were subclassified, as classic ILC, ILC variants, or mixed ductal-lobular cases. Cases from the TCGA validation cohort were not subjected to central pathology review, and the original TCGA histology labels were used for the assessment of the AI-model’s performance.
Assessment of the human-explainability of latent features of the AI-based model
One of the ten aggregator network models (fold 0) was selected to investigate the association between the latent features learned by the AI-model and human-explainable histopathologic features. The embeddings (i.e., latent features) were extracted from the penultimate layer of the aggregator network, situated immediately before the classification heads, rendering them the most representative latent features for the prediction target. Using the t-distributed stochastic neighbor embedding (t-SNE) method (7), embedding vectors of 512-dimentional tiles (n=11,300) from whole-slide images (n=113) of 106 breast cancer cases were projected into a two-dimensional space. This projection was computed using the python package module sklearn.manifold.TSNE (scikit-learn version 1.3.2) [https://scikit-learn.org/stable/modules/generated/sklearn.manifold.TSNE.html]. The whole-slide images from which the tile-level embeddings were derived were sampled from the test dataset of the development cohort. Tile-level embeddings were subsampled by selecting the top 100 tiles, based on their isolated prediction scores for the presence of CDH1 bi-allelic mutation for each whole-slide image. By manually inspecting high-density areas in the t-SNE plot, seven regions of interest (ROIs) within this projection (ROI0 through ROI6) were identified. Three board certified pathologists (FP, FD and LCC) conducted the histologic review of the 1,428 tiles from the seven ROIs and identified distinctive histopathologic features for each ROI.
Statistical analysis
Statistical analyses were conducted using R (version 3.1.2). Mann-Whitney U-test was employed for comparison between continuous variables and Fisher exact test to compare categorical variables. All tests were two-sided; P values <0.05 were considered as statistically significant.
Data availability
The code of the CDH1 bi-allelic mutation classifier attention-based aggregator is available at https://github.com/Paige-AI/cdh1-cancer-res. The AI-model is freely accessible via https://paige.ai/cdh1. The TCGA validation cohort data were obtained from the TCGA Breast Invasive Carcinoma dataset using cBioPortal at http://www.cbioportal.org/study/summary?id=brca_tcga. The CDH1 mutational status of the MSK development and validation cohorts is included in Supplementary Table S1. Data generated in this study are available within the article and Supplementary Data or upon request from the corresponding authors.
RESULTS
Development on an AI-based method for diagnosis of invasive lobular carcinoma
We developed a deep learning model where instead of histologic labels, we utilized CDH1 bi-allelic mutations for training. Diagnostic clinical whole-slide images from an MSK development cohort of 1,057 primary estrogen receptor-positive breast cancers previously subjected to the FDA-approved clinical MSK-IMPACT targeted sequencing assay were used to train an AI-based model to detect CDH1 bi-allelic mutations (i.e. loss-of-function mutations with LOH, or two inactivating mutations (Fig. 1A)). Whole-slide images were pre-processed using a previously developed breast cancer feature extractor trained with >50,000 breast cancer whole-slide images (40) using multiple-instance learning (3), that classified breast lesions as pre-malignant, in situ or invasive carcinoma (Fig. 1B). Embedding vectors for tiles classified as invasive carcinoma were used to train the CDH1 bi-allelic mutation classifier, utilizing an attention-based aggregation network. A 10-fold cross-validation system and a dataset divided into ‘train’, ‘tune’ and ‘test’ partitions at a 6:3:1 ratio were used. A feed-forward aggregation network was trained on each fold, yielding 10 individual models, which were ensembled by calibrating the prediction output for each model. Ten prediction probabilities were generated and averaged into a single prediction for CDH1 bi-allelic mutations (Fig. 1B). The breast cancer histologic subtype was not utilized during algorithmic development as parameter or co-variate. This AI-classifier was employed to detect CDH1 bi-allelic mutations and to predict breast cancer classification as ILC or non-ILC first in the clinical MSK development (n=1,057) cohort. Subsequently, it was applied to an MSK validation cohort (n=405), and to an external independent validation dataset from TCGA (15) (n=751; Fig. 1A).
Figure 1. Artificial intelligence (AI)-model for prediction of CDH1 bi-allelic mutations and the invasive lobular carcinoma diagnosis.

A, The presence of CDH1 bi-allelic mutations was determined through the analysis of paired tumor/normal targeted sequencing data from 1,057 breast cancer patients previously subjected to the FDA-cleared MSK-IMPACT assay. Clinical diagnostic whole-slide images (WSI) derived from the 1,057 breast cancers were used as input for the model. The AI-model for diagnosis of invasive lobular carcinoma (ILC) from WSIs was trained using CDH1 bi-allelic mutations as ground truth. The performance of the AI-model to predict CDH1 bi-allelic mutations and ILC diagnosis was assessed in Memorial Sloan Kettering (MSK) developmental (n=1,057) and independent validation cohorts (n=405), and in an independent external dataset from The Cancer Genome Atlas (TCGA; n=751). B, Following WSI tiling, foreground filtering for the selection of tissue regions was conducted. WSIs were partitioned into 224 × 224 px tiles. Tiles overlapping with tissue regions identified during foreground filtering were selected and used as input for the previously described breast cancer feature extractor, a convolutional neural network (CNN) based on the SE-RESNET50 architecture. The breast cancer feature extractor outputs a 512-dimensional feature embedding vector, as well as cancer prediction scores for all tissue tiles in a WSI. Cancer predictions scores indicate the presence of a pre-malignant lesion, in situ or invasive carcinoma. Tile embedding vectors with invasive prediction scores ≥0.1 threshold were selected and used to train the CDH1 bi-allelic mutation classifier using a 10-fold cross-validation approach with a dataset divided into 10 folds with ‘train’, ‘tune’ and ‘test’ sets partitioned at a 6:3:1 ratio. A feed-forward aggregation network with attention trained on each fold, was applied to make a prediction upon receiving all filtered embedding vectors from a WSI. Inference probabilities for ‘tune’ set and ‘test’ sets were generated. The probabilities of the ‘tune’ set were used to calibrate those of the ‘test’ set. An optimal threshold computed based on calibrated probabilities was used to determine the presence or absence of CDH1 bi-allelic mutations.
Performance of the AI-based model for detection of CDH1 bi-allelic mutations and alternative mechanisms of CDH1 inactivation
The AI-classifier predicted CDH1 bi-allelic mutations in the MSK development cohort (n=1,057) with an AUC of 0.966±0.034, sensitivity=0.83±0.062, specificity=0.96±0.012, accuracy=0.95±0.014 and precision=0.77±0.066, detecting only 23 false negative and 34 false positive samples (Fig. 2A–2D, Supplementary Table S2).
Figure 2. Performance of the AI-model for detection of CDH1 bi-allelic mutations in the MSK development cohort and analysis of false positive and false negative cases.

A, Receiver operating characteristic (ROC) curve, B, precision-recall curve and C, sensitivity, specificity, accuracy, and precision for the prediction of CDH1 bi-allelic mutations in the MSK development cohort (n=1,057); error bars represent 95% confidence intervals. AUC, area under the curve; CV, cross validation D, Boxplot depicting prediction probability values for samples harboring CDH1 bi-allelic mutations (n=139) and lacking CDH1 bi-allelic mutations (n=918) in the development cohort, Mann Whitney U-test. AUC, area under the curve; CV, cross validation FN, false negative; FP, false positive, TN, true negative; TP, true positive. E, Chromosome 16 copy number plot in an invasive lobular carcinoma (ILC) harboring a CDH1 homozygous deletion (‘AI-CDH1–30’). The genomic position (X-axis), a chromosome 16 ideogram, and Log2 copy number ratios are shown. F, Alignment of sequence reads of chromosome 16 depicting CDH1 intragenic deletion (blue) spanning exons 3–7 in case ‘AI-CDH1–18’. G, Alignment of sequence reads of chromosome 16 depicting a non-coding likely pathogenic CDH1 c.1137G>A splice region variant in case ‘AI-CDH1–15. Transcripts predicted by Ensembl with corresponding grey density plots depicting base level coverage and chromosome ideograms are shown for (F) and (G). H, CDH1 gene promoter CpG islands and probes used for CDH1 gene promoter methylation assessment by digital droplet PCR (ddPCR) assay. TSS, transcription start site. I, 2D matrix of CDH1 promoter methylation assessment by ddPCR depicting FAM (Y-axis) and HEX (X-axis) fluorescence amplitude, and methylated (Met) and unmethylated (Unmet) DNA droplets of case ‘AI-CDH1–23’; blue (Met), red (Met + Unmet), green (Unmet) and grey (empty). J, Representative hematoxylin-and-eosin (H&E) photomicrographs (top left; bottom left), corresponding prediction heatmap (top right) and E-cadherin expression by immunohistochemistry (bottom right) of case ‘AI-CDH1–1’. Scale bars, 2 mm (top) and 200 microns (bottom). K, Circos plot of case ‘AI-CDH1–1’ whole-genome sequencing depicting (from outside to inside) inter-variant distance and type of single base substitution (SBS), indels, copy number alterations (CNAs) and structural variants. L, Schematic representation of a structural variant detected in case ‘AI-CDH1–1’ by whole-genome sequencing, depicting a fusion gene between an intergenic region in chromosome 13 and CDH1 in chromosome 16. Promoter, 3’ UTR, 5’UTR, exons and protein domains are shown. Vertical dashed lines represent the genomic breakpoints in chromosomes 13 and 16. The predicted deleterious fusion gene juxtaposed the intergenic region in chromosome 13 and intron 2 of CDH1, resulting in a deletion of the CDH1 transcription start site and no CDH1 transcription (bottom). CD, cytoplasmic domain; ED, extracellular domain; ITR, intergenic region; Pr, precursor domain; Sig, signal domain; TM, transmembrane domain; TSS, transcription start site; UTR, untranslated region. M, CDH1 genetic alterations as detected by targeted sequencing or whole-genome sequencing and CDH1 gene promoter methylation in false positive (FP) samples for prediction of CDH1 bi-allelic mutations by AI-model (top). CDH1 mutation type and bi-allelic inactivation in false negative (FN) samples for detection of CDH1 bi-allelic mutations (bottom). LOH, loss-of-heterozygosity of the wild type allele.
To elucidate the biology of breast cancers lacking CDH1 bi-allelic mutations by targeted sequencing but predicted to harbor them by the AI-algorithm (i.e., false positives), we conducted the re-analysis of targeted sequencing data for coding and non-coding genetic alterations, and CDH1 promoter methylation assessment. In these samples (n=34), we identified alternative mechanisms of CDH1 inactivation including CDH1 promoter methylation (n=18, 53%), CDH1 homozygous deletions (n=3, 9%), intragenic deletion with LOH (n=1, 3%), as well as likely pathogenic non-coding CDH1 splice region variants with LOH (n=2; 6%), including CDH1 c.1137G>A and CDH1 c.1320+5G>A, not previously reported in breast cancer (Fig. 2E–2I, Supplementary Fig. S1A–S1D, Supplementary Fig. S2; see Materials and Methods). As a hypothesis generating aim, we sought to explore if complex genetic alterations not captured by targeted sequencing could explain the remaining false positive samples. Whole-genome sequencing of a classic ILC lacking CDH1 coding/non-coding genetic alterations by targeted sequencing, CDH1 promoter methylation and E-cadherin expression, yet predicted to harbor a CDH1 bi-allelic mutation by the AI-model (case AI-CDH1–1; Fig. 2J, see Materials and Methods) revealed dominant aging mutational signatures, typical in ILC, no mutations in cell-cell adhesion genes, and a translocation t(13;16) predicted to result in a deleterious fusion gene affecting CDH1 (Fig. 2K–2L and Supplementary Fig. S3A–S3E). This rearrangement juxtaposed intron 2 of CDH1 with an intergenic region in chromosome 13, with ensuing loss of the exons 1–2 and regulatory regions of CDH1, including its transcription start site, associated with LOH (Fig. 2L). In summary, alternative mechanisms of CDH1 bi-allelic inactivation were identified in 25/34 (74%) initially considered false positive cases (Fig. 2M), and their prediction probabilities by the AI-model (n=25; median=0.97±0.03) were comparable to those of breast cancers harboring CDH1 bi-allelic mutations (n=139; median=0.98±0.05; P=0.602; Supplementary Fig. S4A). Re-analysis of targeted sequencing data from 23 samples harboring CDH1 bi-allelic mutations but misclassified by the algorithm as not (false negatives) confirmed the inactivating and bi-allelic nature of the CDH1 mutations (Fig. 2M). Taken together, the AI-model detected the phenotypic characteristics of CDH1 inactivation regardless of the molecular mechanism with sensitivity=0.86±0.053, specificity=0.99±0.007, accuracy=0.97±0.01 and precision=0.94±0.039 (Supplementary Fig. S4B).
The AI-based model trained with CDH1 bi-allelic mutations as ground truth robustly identified ILCs
We next sought to determine whether the AI-model that was trained to detect the presence of CDH1 bi-allelic mutations could be employed to diagnose ILC. Our analysis revealed that the AI-classifier identified ILCs, regardless of CDH1 status, with an AUC=0.981±0.024, sensitivity=0.79±0.058, specificity=1.0±0.004, accuracy=0.96±0.012 and precision=0.98±0.025 (Fig. 3A–3D, Supplementary Table S2). 3/1,057 breast cancers (0.28%) of other histologic types (a neuroendocrine tumor, an apocrine carcinoma and an IDC-NST) were misclassified by the algorithm as ILCs (false positives; Fig. 3E). Forty samples were classified as non-ILC by the AI-algorithm, yet were histologically ILCs (false negatives); these mostly corresponded to cases that phenotypically deviate from classic ILC, including mainly ILC variants/mixed ductal-lobular cases (95%; 38/40), such as pleomorphic ILC (n=20), mixed ductal-lobular (n=13), ILC with extracellular mucin, tubulobular, solid, alveolar and trabecular ILCs (n=1, each; Fig. 3E–3F). By contrast, only 35% (52/147) of true positive samples were ILC variants/mixed ductal-lobular cases, and most of them were classic ILCs (P=1.7e-12; Supplementary Tables S3–S4).
Figure 3. Detection of invasive lobular carcinoma diagnosis by the AI-model in the MSK development cohort.

A, Receiver operating characteristic (ROC) curve, B, precision-recall curves, and C, sensitivity, specificity, accuracy and precision for the AI-based detection of invasive lobular carcinoma (ILC) in the MSK development cohort (n=1,057); error bars represent 95% confidence intervals. AUC, area under the curve; CV, cross validation. D, Boxplot depicting prediction probability values for non-ILC (n=870) and ILC (n=187) breast cancers of the MSK development cohort, Mann Whitney U-test. FP, false positives; FN, false negatives; TP, true positives; TN, true negatives. E, Histologic type of false positive and false negative samples for the AI-based diagnosis of ILC in the MSK development cohort. F, Hematoxylin-and-eosin (H&E) micrographs (top) and prediction heatmaps (bottom) of false negative cases for the AI-based detection of ILC in the MSK development cohort including a pleomorphic ILC (left), a solid ILC (center left), a mixed ductal-lobular carcinoma (center right), and an ILC with extracellular mucin (right). Scale bars, 500 microns (left and center left), 200 microns (center right) and 1 mm (right).
Validation of the AI-model in an independent MSK validation cohort and an external TCGA validation dataset
For validation, the performance of the AI-model was evaluated in an independent MSK validation cohort (n=405) and an external validation cohort from the TCGA (n=751); see Materials and Methods. The 10 trained models from the MSK development cohort were run on the two validation cohorts separately, and 10 inference probabilities were generated for each sample. The 10 corresponding calibrated classifiers were applied on the inference probabilities. The universal optimal operating threshold, determined from the MSK development cohort, was applied to the averaged calibrated probabilities, and a binary prediction for each sample was generated (see Materials and Methods).
Six low-tumor-purity samples were excluded from the MSK validation cohort (n=405). In 399 breast cancers with available bi-allelic CDH1 mutational status, the AI-model predicted CDH1 bi-allelic mutations with AUC=0.959±0.029, sensitivity=0.78±0.107, specificity=0.95±0.023, accuracy=0.93±0.025 and precision=0.73±0.111 (Fig. 4A–4C). Of the 16 samples classified by the AI-model as CDH1 bi-allelically mutated but lacking these alterations (false positives), 58% (7/12) of cases interrogated displayed CDH1 promoter methylation and none displayed alternative genetic mechanisms of CDH1 inactivation by reanalysis of targeted sequencing data. Reanalysis of the 12 samples misclassified by the AI-model as lacking CDH1 bi-allelic mutations (false negatives) confirmed the presence of bi-allelic inactivating mutations affecting CDH1 (Fig. 4D–4E).
Figure 4. Performance of the AI-model in an independent MSK validation cohort, and explainability of AI-model latent features.

A, Receiver operating characteristic (ROC) curve, B, precision-recall curve and C, sensitivity, specificity, accuracy and precision for the AI-based detection of CDH1 bi-allelic mutations in the independent MSK validation cohort (n=399); error bars represent 95% confidence intervals. AUC, area under the curve. D, Boxplot depicting prediction probability values for cases harboring CDH1 bi-allelic mutations (n=55) and those lacking CDH1 bi-allelic mutations (n=344) in the independent MSK validation cohort, Mann-Whitney U-test. E, CDH1 gene promoter methylation status of false positive (FP) samples (top) and CDH1 mutation type and bi-allelic status of false negative (FN) samples (bottom) for the AI-based detection of CDH1 bi-allelic mutations in the independent MSK validation cohort. LOH, loss-of-heterozygosity. F, ROC curve, G, precision-recall curve and H, sensitivity, specificity, accuracy and precision for the AI-based diagnosis of invasive lobular carcinoma (ILC) in the independent MSK validation cohort (n=405); error bars represent 95% confidence intervals. AUC, area under the curve. I, Boxplot depicting prediction probability values for ILC (n=84) and non-ILC (n=321) breast cancers of the independent MSK validation cohort. J, Histologic type of false negative samples for the AI-based detection of ILC in the MSK validation cohort. K, t-distributed Stochastic Neighbor Embedding (t-SNE) plot depicting tile-level embeddings (n=11,300) sampled from the developmental cohort. Regions of interest (ROI) are indicated and sample class (TP, true positive; TN, true negative; FP, false positive, FN, false negative) is color-coded. L. Representative tiles of ROIs 0 through 6 are displayed. Proportion of tiles per sample class and according to histology are depicted in bars. The most salient histologic features of tiles in each ROI are listed.
In the MSK validation cohort, the AI-model robustly diagnosed ILCs with AUC=0.983±0.022, sensitivity=0.74±0.092, specificity=1.0±0.006, accuracy=0.95±0.22, and precision=1.0±0.029 (Fig. 4F–4H). No false positive samples were identified, and most ILCs misclassified as non-ILCs by the AI-model (false negatives) were ILC variants or mixed ductal-lobular cases (77%; 17/22; Fig. 4I–4J and Supplementary Tables S2–S3). Akin to the development cohort, in the MSK validation cohort, ILC variants and mixed ductal-lobular cases were enriched in the subset of ILCs misclassified by the AI-model (false negatives) as compared to the true positive samples (P=1e-03; Supplementary Table S4).
In an external independent validation dataset of whole-slide images and diagnostic labels from TCGA (n=751), the AI-model diagnosed ILC with AUC=0.934±0.013, sensitivity=0.65±0.065, specificity=0.97±0.014, accuracy=0.89±0.023 and precision=0.89±0.051 (Supplementary Figures S5A–S5F, Supplementary Table S5).
Human-explainable features of AI-model
To determine whether the latent features of the AI-model correlate with human-explainable histologic features, we projected 512-dimensional tile-level embeddings (n=11,300) from the penultimate layer of the network of one of the 10 ensemble model CDH1 aggregators, sampled from the developmental cohort, into a two-dimensional space using a t-SNE method (Fig. 4K and Materials and Methods). Manual inspection of t-SNE tile clusters revealed seven ROIs (ROI0 though ROI6) encompassing 1,428 tiles (Fig. 4K).
Most tiles in ROI0 (99%; 99/100), ROI1 (70.2%; 139/198) and ROI2 (76.7%; 217/283) were predicted to have CDH1 bi-allelic mutations and harbored those alterations (true positives). Accordingly, all ROI0 (n=100) and most ROI1 (72.2%; 143/198) and ROI2 (83.4%; 236/283) tiles exhibited an ILC phenotype (Fig. 4K–4L). In contrast, all ROI5 (n=468) and ROI6 (n=178) tiles were not predicted to harbor CDH1 bi-allelic mutations and lacked those alterations (true negatives), with all ROI5 (n=468) and all but two ROI6 (98.9%; 176/178) tiles displaying a non-ILC phenotype (Fig. 4K–4L). Although most ROI3 (57.3%; 43/75) and ROI4 (55.6%; 70/126) tiles were true negatives, 30.7% (23/75) and 43.7% (55/126) of ROI3 and ROI4 tiles, respectively, harbored CDH1 bi-allelic mutations that were not detected by the AI-model (false negatives), all of which displayed a mixed ductal-lobular phenotype (Fig. 4K–4L). Overall, tiles predicted by the AI-model to harbor CDH1 bi-allelic mutations were enriched in ROI0, ROI1 and ROI2.
Histopathologic review of the ROI0 through ROI6 tiles (n=1,428) by three pathologists revealed that features distinguishing ROI0, ROI1 and ROI2, predominantly predicted by the AI-model to be associated with CDH1 bi-allelic mutations, included discohesion, single cells/single cell files, trabeculae, monomorphism, low to intermediate nuclear grade, cytoplasmic clearing and intracytoplasmic inclusions. In contrast, features characterizing the other ROIs (ROI3-ROI6) encompassed cohesiveness, high nuclear grade, a high extent of tumor infiltrating lymphocytes (TILs) and stromal desmoplasia (Fig. 4L).
DISCUSSION
Here, through the integration of AI and genetics we have conceived and implemented a novel strategy for the development of AI-based cancer classification systems. In pan-cancer studies, Fu et al. (41) found a significant association between automatically learned computational histological features in breast cancer with CDH1 driver mutations with AUC=0.85. Saldanha et al. (42), through the integration of self-supervised feature extraction and attention-based multiple instance learning, predicted CDH1 alterations in breast cancer with AUC=0.68. In this study, we capitalized on the genotypic-phenotypic correlation in ILC by utilizing the pathognomonic genetic alteration (CDH1 bi-allelic mutations) as ground truth for training the algorithm, enabling the development of an AI-model that robustly (AUC>0.95) predicts this genetic alteration and detects a phenotypic/histologic entity, ILC. The AI-system trained using this approach not only detected ILCs harboring CDH1 bi-allelic mutations, but also ILCs driven by alternative CDH1 inactivating genetic or epigenetic alterations. This illustrates how the integration of genetics into AI-model training not only avoids human subjectivity, but also can yield information beyond the mere detection of a single genetic alteration.
Not uncommonly, histologic entities displaying a distinctive phenotype and characterized by pathognomonic genetic alterations, such as adenoid cystic carcinoma (MYB/MYBL1 rearrangements; MYB amplification) or secretory carcinoma (NTRK1/2/3 rearrangements) have distinctive biological features, clinical behavior and outcomes (4,43). Hence, the strategy proposed here could be implemented to develop AI-based systems for the diagnosis of cancer types harboring disease-defining genetic alterations.
ILCs have unique pathologic and biologic features, with important therapeutic implications. In addition to the modest responses to chemotherapy (7), ILC-specific genetic vulnerabilities that could be exploited therapeutically have been reported, including the synthetic lethal interaction between E-cadherin deficiency and ROS1 inhibition (7). Despite the biologic, clinicopathologic and therapeutic differences between ILCs and non-ILC breast cancers (7), the modest interobserver agreement for the diagnosis of ILC, ranging in kappa scores from 0.6–0.8 (8,9), poses challenges to its utilization in clinical decision-making and ILC-specific clinical trials (8,9). Although E-cadherin immunohistochemistry has been shown to increase ILC diagnostic reproducibility (9), its use as a defining feature has proven controversial. Hence, the implementation of digital tools to optimize ILC diagnosis has the potential of improving breast cancer outcomes.
Exploration of the explainability of the solution revealed an association between the latent features learned by the AI-model that predict the presence of CDH1 bi-allelic mutations and human-explainable features. These include tumor characteristics, primarily associated with cell discohesion and low histologic grade, as well as microenvironmental features, such as low TIL infiltration, and lack of stromal desmoplasia.
Investigation of the molecular basis of breast cancers predicted to harbor CDH1 bi-allelic mutations by the AI-model but lacking them by targeted sequencing allowed the identification of novel CDH1 inactivating genetic alterations, including a CDH1 deleterious fusion gene and non-coding CDH1 alterations. Non-coding alterations reported in breast cancer include alterations affecting PIK3CA, TP53, FOXA1 and DNA repair genes (RAD51B and GEN1) (44,45). We report the novel observation of CDH1 non-coding alterations in a subset of ILCs. Their burden in ILC, however, requires further study. Our findings also indicate that CDH1 gene promoter methylation is prevalent in ILCs lacking CDH1 pathogenic mutations. Our findings suggest that application of AI-tools trained to detect a genetic alteration could help decipher the biological underpinning of phenotypic/histologic entities and warrant further systematic analyses of ILCs lacking genetic/epigenetic CDH1 alterations, which may reveal alterations in other key cell-cell adhesion genes.
Our study has limitations, such as the modest number of cases per histologic ILC variant and the scarcity of mixed ductal-lobular cases included for training, given their rarity (46). Indeed, only 2.1% of breast cancers in the training set were mixed ductal-lobular, and only 55% of them harbored CDH1 bi-allelic mutations, as previously reported (47), resulting in their suboptimal detection. Previous genetic studies segregated mixed ductal-lobular cases into invasive ductal carcinoma (IDC)-like or ILC-like, challenging their classification as a distinct entity (47). Therefore, some mixed ductal-lobular cases in our cohort might be ILC-like, and truly misclassified by the AI-algorithm, while others might be IDC-like, potentially reducing the misclassification burden. Although the AI-system we developed, relying on a robust genotypic-phenotypic correlation, effectively detects most ILCs, it is constrained in breast cancer subtypes with low prevalence of CDH1 genetic alterations, such as mixed ductal-lobular cases or ILC variants. The performance of the model in the external validation dataset from TCGA was inferior to that observed in the MSK independent validation cohort; this is not surprising, given that the whole-slide images from TCGA constitute research, rather than clinical-grade diagnostic material. Despite these limitations, this proof-of-concept study demonstrated that by training an AI-model applied to whole-slide images using CDH1 bi-allelic mutations rather than histologic diagnosis labels as ground truth, we developed an AI-system that can precisely and accurately classify ILCs. These findings provide a new paradigm for AI-based cancer classification systems, indicating that disease-defining genetic alterations, in addition to pathology diagnoses, may be employed for their development.
Supplementary Material
STATEMENT OF SIGNIFICANCE.
Genetic alterations linked to strong genotypic-phenotypic correlations can be utilized to develop AI-systems applied to pathology that facilitate cancer diagnosis and biological discoveries.
Acknowledgements
This study was partially funded by the Breast Cancer Research Foundation. Research reported in this publication was partly funded by a Cancer Center Support Grant of the National Institutes of Health (NIH)/ National Cancer Institute (grant No P30CA008748). FP is funded in part by a NIH/NCI P50 CA247749 01 grant and by a Starr Cancer Consortium grant. BW is partially funded by a NIH/NCI P50 CA247749 01 grant, a Cycle for Survival and Breast Cancer Research Foundation grants. SC is funded in part by the Breast Cancer Research Foundation. JSRF was funded in part by an NIH/NCI P50 CA247749 01 grant, a Breast Cancer Research Foundation grant and a Susan G Komen Leadership grant. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Footnotes
Conflict of interest disclosures: FP reports membership of the scientific advisory board of MultiplexDX and membership of the diagnostic advisory board of AstraZeneca. Additionally, FP receives consultancy fees from AstraZeneca. YW, MB, JHB, JS, MCHL, RAG, AC, BR, JDK, JO, DSK and TJF receive salary and own stock from Paige. CK is a paid a consultant and equity holder at Paige. BW reports research support from Repare Therapeutics, outside the scope of the current study. SC receives grant support (to MSK) from Daiichi-Sankyo and AstraZeneca, financial interests in Totus Medicine and Odyssey Biosciences, and consulting fees from Lilly, Novartis, Paige.ai, AstraZeneca, SAGA, Boxer Capital, & Prelude Therapeutics. J.S.R.F reports receiving personal/consultancy fees from Goldman Sachs, Bain Capital, REPARE Therapeutics, Saga Diagnostics and Paige.AI, membership of the scientific advisory boards of VolitionRx, REPARE Therapeutics and Paige.AI, membership of the Board of Directors of Grupo Oncoclinicas, and ad hoc membership of the scientific advisory boards of Astrazeneca, Merck, Daiichi Sankyo, Roche Tissue Diagnostics and Personalis, outside the scope of this study. All other authors declare no potential conflicts of interest.
REFERENCES
- 1.Bera K, Schalper KA, Rimm DL, Velcheti V, Madabhushi A. Artificial intelligence in digital pathology - new tools for diagnosis and precision oncology. Nat Rev Clin Oncol 2019;16(11):703–15 doi 10.1038/s41571-019-0252-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Sandbank J, Bataillon G, Nudelman A, Krasnitsky I, Mikulinsky R, Bien L, et al. Validation and real-world clinical application of an artificial intelligence algorithm for breast cancer detection in biopsies. NPJ Breast Cancer 2022;8(1):129 doi 10.1038/s41523-022-00496-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Campanella G, Hanna MG, Geneslaw L, Miraflor A, Werneck Krauss Silva V, Busam KJ, et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat Med 2019;25(8):1301–9 doi 10.1038/s41591-019-0508-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ashworth A, Lord CJ, Reis-Filho JS. Genetic interactions in cancer progression and treatment. Cell 2011;145(1):30–8 doi 10.1016/j.cell.2011.03.020. [DOI] [PubMed] [Google Scholar]
- 5.Pareja F, Ferrando L, Lee SSK, Beca F, Selenica P, Brown DN, et al. The genomic landscape of metastatic histologic special types of invasive breast cancer. NPJ Breast Cancer 2020;6:53 doi 10.1038/s41523-020-00195-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Derakhshan F, Da Cruz Paula A, Selenica P, da Silva EM, Grabenstetter A, Jalali S, et al. Non-Lobular Invasive Breast Carcinomas with Bi-Allelic Pathogenic CDH1 Somatic Alterations: a Histologic, Immunophenotypic and Genomic Characterization. Mod Pathol 2023:100375 doi 10.1016/j.modpat.2023.100375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Van Baelen K, Geukens T, Maetens M, Tjan-Heijnen V, Lord CJ, Linn S, et al. Current and future diagnostic and treatment strategies for patients with invasive lobular breast cancer. Ann Oncol 2022;33(8):769–85 doi 10.1016/j.annonc.2022.05.006. [DOI] [PubMed] [Google Scholar]
- 8.Christgen M, Gluz O, Harbeck N, Kates RE, Raap M, Christgen H, et al. Differential impact of prognostic parameters in hormone receptor-positive lobular breast cancer. Cancer 2020;126(22):4847–58 doi 10.1002/cncr.33104. [DOI] [PubMed] [Google Scholar]
- 9.Christgen M, Kandt LD, Antonopoulos W, Bartels S, Van Bockstal MR, Bredt M, et al. Inter-observer agreement for the histological diagnosis of invasive lobular breast carcinoma. J Pathol Clin Res 2022;8(2):191–205 doi 10.1002/cjp2.253. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.De Schepper M, Vincent-Salomon A, Christgen M, Van Baelen K, Richard F, Tsuda H, et al. Results of a worldwide survey on the currently used histopathological diagnostic criteria for invasive lobular breast cancer. Mod Pathol 2022;35(12):1812–20 doi 10.1038/s41379-022-01135-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sarrio D, Moreno-Bueno G, Hardisson D, Sanchez-Estevez C, Guo M, Herman JG, et al. Epigenetic and genetic alterations of APC and CDH1 genes in lobular breast cancer: relationships with abnormal E-cadherin and catenin expression and microsatellite instability. Int J Cancer 2003;106(2):208–15 doi 10.1002/ijc.11197. [DOI] [PubMed] [Google Scholar]
- 12.Rakha EA, Patel A, Powe DG, Benhasouna A, Green AR, Lambros MB, et al. Clinical and biological significance of E-cadherin protein expression in invasive lobular carcinoma of the breast. Am J Surg Pathol 2010;34(10):1472–9 doi 10.1097/PAS.0b013e3181f01916. [DOI] [PubMed] [Google Scholar]
- 13.Mahler-Araujo B, Savage K, Parry S, Reis-Filho JS. Reduction of E-cadherin expression is associated with non-lobular breast carcinomas of basal-like and triple negative phenotype. J Clin Pathol 2008;61(5):615–20 doi 10.1136/jcp.2007.053991. [DOI] [PubMed] [Google Scholar]
- 14.Marini N, Marchesin S, Otalora S, Wodzinski M, Caputo A, van Rijthoven M, et al. Unleashing the potential of digital pathology data by training computer-aided diagnosis models without human annotations. NPJ Digit Med 2022;5(1):102 doi 10.1038/s41746-022-00635-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Cancer Genome Atlas N Comprehensive molecular portraits of human breast tumours. Nature 2012;490(7418):61–70 doi 10.1038/nature11412. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Shen R, Seshan VE. FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing. Nucleic Acids Res 2016;44(16):e131 doi 10.1093/nar/gkw520. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Raciti P, Sue J, Retamero JA, Ceballos R, Godrich R, Kunz JD, et al. Clinical Validation of Artificial Intelligence-Augmented Pathology Diagnosis Demonstrates Significant Gains in Diagnostic Accuracy in Prostate Cancer Detection. Arch Pathol Lab Med 2022. doi 10.5858/arpa.2022-0066-OA. [DOI] [PubMed] [Google Scholar]
- 18.Jie H, Li S, Gang S, Albanie S. Squeeze-and-excitation networks. 2018. [DOI] [PubMed] [Google Scholar]
- 19.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. J Mach Learn Res 2011;12:2825–30. [Google Scholar]
- 20.R.G. N, D.G. A. Proportions and their differences. In: Altman D DM, TN B MJG, editors. Statistics with Confidence: Confidence intervals and statistical guidelines. 2nd ed: BMJ Books; 2000. p 45–7. [Google Scholar]
- 21.Rausch T, Zichner T, Schlattl A, Stutz AM, Benes V, Korbel JO. DELLY: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics 2012;28(18):i333–i9 doi 10.1093/bioinformatics/bts378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Robinson JT, Thorvaldsdottir H, Winckler W, Guttman M, Lander ES, Getz G, et al. Integrative genomics viewer. Nat Biotechnol 2011;29(1):24–6 doi 10.1038/nbt.1754. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alfoldi J, Wang Q, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 2020;581(7809):434–43 doi 10.1038/s41586-020-2308-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, et al. Predicting Splicing from Primary Sequence with Deep Learning. Cell 2019;176(3):535–48 e24 doi 10.1016/j.cell.2018.12.015. [DOI] [PubMed] [Google Scholar]
- 25.Landrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, et al. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res 2014;42(Database issue):D980–5 doi 10.1093/nar/gkt1113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Li H, Durbin R. Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics 2010;26(5):589–95 doi 10.1093/bioinformatics/btp698. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Cibulskis K, Lawrence MS, Carter SL, Sivachenko A, Jaffe D, Sougnez C, et al. Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol 2013;31(3):213–9 doi 10.1038/nbt.2514. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Saunders CT, Wong WS, Swamy S, Becq J, Murray LJ, Cheetham RK. Strelka: accurate somatic small-variant calling from sequenced tumor-normal sample pairs. Bioinformatics 2012;28(14):1811–7 doi 10.1093/bioinformatics/bts271. [DOI] [PubMed] [Google Scholar]
- 29.Koboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, et al. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res 2012;22(3):568–76 doi 10.1101/gr.129684.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Rimmer A, Phan H, Mathieson I, Iqbal Z, Twigg SRF, Wilkie AOM, et al. Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications. Nat Genet 2014;46(8):912–8 doi 10.1038/ng.3036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Narzisi G, O’Rawe JA, Iossifov I, Fang H, Lee YH, Wang Z, et al. Accurate de novo and transmitted indel detection in exome-capture data using microassembly. Nat Methods 2014;11(10):1033–6 doi 10.1038/nmeth.3069. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Carter SL, Cibulskis K, Helman E, McKenna A, Shen H, Zack T, et al. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol 2012;30(5):413–21 doi 10.1038/nbt.2203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Islam SMA, Diaz-Gay M, Wu Y, Barnes M, Vangara R, Bergstrom EN, et al. Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor. Cell Genom 2022;2(11):None doi 10.1016/j.xgen.2022.100179. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Degasperi A, Amarante TD, Czarnecki J, Shooter S, Zou X, Glodzik D, et al. A practical framework and online tool for mutational signature analyses show inter-tissue variation and driver dependencies. Nat Cancer 2020;1(2):249–63 doi 10.1038/s43018-020-0027-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Chen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Källberg M, et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 2016;32(8):1220–2 doi 10.1093/bioinformatics/btv710. [DOI] [PubMed] [Google Scholar]
- 36.Wala JA, Bandopadhayay P, Greenwald NF, O’Rourke R, Sharpe T, Stewart C, et al. SvABA: genome-wide detection of structural variants and indels by local assembly. Genome Res 2018;28(4):581–91 doi 10.1101/gr.221028.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Cameron DL, Baber J, Shale C, Valle-Inclan JE, Besselink N, van Hoeck A, et al. GRIDSS2: comprehensive characterisation of somatic structural variation using single breakend variants and structural variant phasing. Genome Biol 2021;22(1):202 doi 10.1186/s13059-021-02423-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Selenica P, Marra A, Choudhury NJ, Gazzo A, Falcon CJ, Patel J, et al. APOBEC mutagenesis, kataegis, chromothripsis in EGFR-mutant osimertinib-resistant lung adenocarcinomas. Ann Oncol 2022;33(12):1284–95 doi 10.1016/j.annonc.2022.09.151. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.WHO Classification of Tumors Editorial Board. Breast tumours. WHO Classification of Tumors. 5th Edition. IARC: Lyon; 2019. [Google Scholar]
- 40.Hanna MG, Raciti P, Bozkurt A, Godrich R, Viret J, Lee D, et al. Subtyping invasive carcinomas and high-risk lesions for machine learning based breast pathology. Cancer Res 2022;82(4) doi 10.1158/1538-7445.Sabcs21-Pd11-02. [DOI] [Google Scholar]
- 41.Fu Y, Jung AW, Torne RV, Gonzalez S, Vohringer H, Shmatko A, et al. Pan-cancer computational histopathology reveals mutations, tumor composition and prognosis. Nat Cancer 2020;1(8):800–10 doi 10.1038/s43018-020-0085-8. [DOI] [PubMed] [Google Scholar]
- 42.Saldanha OL, Loeffler CML, Niehues JM, van Treeck M, Seraphin TP, Hewitt KJ, et al. Self-supervised attention-based deep learning for pan-cancer mutation prediction from histopathology. NPJ Precis Oncol 2023;7(1):35 doi 10.1038/s41698-023-00365-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Pareja F, Weigelt B, Reis-Filho JS. Problematic breast tumors reassessed in light of novel molecular data. Mod Pathol 2021;34(Suppl 1):38–47 doi 10.1038/s41379-020-00693-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Dietlein F, Wang AB, Fagre C, Tang A, Besselink NJM, Cuppen E, et al. Genome-wide analysis of somatic noncoding mutation patterns in cancer. Science 2022;376(6589):eabg5601 doi 10.1126/science.abg5601. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Elliott K, Larsson E. Non-coding driver mutations in human cancer. Nat Rev Cancer 2021;21(8):500–9 doi 10.1038/s41568-021-00371-z. [DOI] [PubMed] [Google Scholar]
- 46.Rakha EA, Gill MS, El-Sayed ME, Khan MM, Hodi Z, Blamey RW, et al. The biological and clinical characteristics of breast carcinoma with mixed ductal and lobular morphology. Breast Cancer Res Treat 2009;114(2):243–50 doi 10.1007/s10549-008-0007-4. [DOI] [PubMed] [Google Scholar]
- 47.Ciriello G, Gatza ML, Beck AH, Wilkerson MD, Rhie SK, Pastore A, et al. Comprehensive Molecular Portraits of Invasive Lobular Breast Cancer. Cell 2015;163(2):506–19 doi 10.1016/j.cell.2015.09.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The code of the CDH1 bi-allelic mutation classifier attention-based aggregator is available at https://github.com/Paige-AI/cdh1-cancer-res. The AI-model is freely accessible via https://paige.ai/cdh1. The TCGA validation cohort data were obtained from the TCGA Breast Invasive Carcinoma dataset using cBioPortal at http://www.cbioportal.org/study/summary?id=brca_tcga. The CDH1 mutational status of the MSK development and validation cohorts is included in Supplementary Table S1. Data generated in this study are available within the article and Supplementary Data or upon request from the corresponding authors.
