Skip to main content
Journal of Pathology Informatics logoLink to Journal of Pathology Informatics
. 2026 Aug 27;23:100711. doi: 10.1016/j.jpi.2026.100711

Interpretable prototype learning for EGFR amplification prediction from whole-slide images in glioblastoma

Homay Danaei Mehr a,b,⁎, Imran Noorani b,c,d, Cong Cong a,b, Mark Fabian e, Antonio Di Ieva b, Sidong Liu a,b
PMCID: PMC13627952  PMID: 42824774

Abstract

The epidermal growth factor receptor (EGFR) is one of the key biomarkers for diagnosis, treatment, and prognosis in glioblastoma (GBM). Current EGFR diagnostic methods, including immunohistochemistry, fluorescence in situ hybridization, and next-generation sequencing, are costly, time-consuming, and not uniformly accessible. A major challenge in whole-slide images (WSIs)-based prediction is to capture the full morphological heterogeneity of the tissue in a way that supports not only precise classification but also interpretability. To address these challenges, we proposed an automatic, interpretable prototype-learning framework that integrates a Variational Autoencoder with the Dirichlet Bayesian Gaussian Mixture Model to determine optimal tissue prototypes from morphological features of hematoxylin and eosin-stained WSI, which guide the classification of EGFR amplification. The classification performance is evaluated using internal 5-fold cross-validation on the Cancer Genome Atlas dataset and external validation on the Clinical Proteomic Tumor Analysis Consortium (CPTAC) dataset and the private GB-UK dataset. The proposed model achieved an area under the curve of 0.8087 ± 0.0153 for internal validation and 0.7740 and 0.7870 for external validations of the CPTAC and GB-UK cohorts, respectively, surpassing multi-instance learning models and conventional predefined clustering settings. Specific prototypes distinguish EGFR-amplified from EGFR-non-amplified cases, and histological analysis of these prototypes, which are consistent with expert-recognized morphological patterns, suggests an EGFR-amplified infiltration pattern. These findings demonstrate that the automatic prototype learning framework provides a rapid, cost-effective, interpretable, and scalable AI model for EGFR prediction in GBM, linking the learned prototypes to recognizable tissue patterns that can facilitate clinical decision-making.

Keywords: Glioblastoma, Prototype learning, Computational pathology, Whole-slide imaging, Multi-instance learning, EGFR biomarker, Histopathological image analysis

Highlights

  • •

    Prototype learning predicts glioblastoma-EGFR amplification from whole-slide images.

  • •

    Automatic prototype discovery captures meaningful tissue patterns.

  • •

    Automatic prototypes enhanced EGFR classification over predefined clustering.

  • •

    External validation supports generalization across external cohorts.

  • •

    Learned prototypes match expert-recognized morphological patterns.

Introduction

Glioblastoma (GBM), one of the most aggressive intrinsic brain tumors in adults, is notoriously resistant to treatment due to its intrinsic molecular heterogeneity. The epidermal growth factor receptor (EGFR) is a protein on the surface of certain cells that plays a pivotal role in controlling cell growth, survival, proliferation, and differentiation. EGFR amplification drives the development and progression of nearly 50% of human GBMs and has been shown to initiate GBM in mouse models, making it a crucial diagnostic and prognostic biomarker, as well as a therapeutic target in GBM.1, 2, 3, 4 The 2021 WHO Classification of Central Nervous System Tumors incorporates molecular features alongside histology to provide diagnoses. The EGFR amplification serves as a diagnostic marker in isocitrate dehydrogenase1(IDH1)-wildtype GBMs, particularly in gliomas that lack histological features of aggressiveness, including microvascular proliferation and pseudopalisading necrosis.5, 6 EGFR amplification can also serve as a prognostic marker of poorer overall survival in certain populations.7 Hence, these findings highlight the importance of accurate identification of the EGFR amplification.

The current diagnostic methods, such as immunohistochemistry (IHC), fluorescence in situ hybridization (FISH), digital PCR,7 and next-generation sequencing, are costly, require technical expertise, and are time-consuming. Hence, their limited availability poses practical challenges for clinical labs. Additionally, recent biological findings indicate that EGFR amplification in GBM is typically detected in extrachromosomal DNA (ecDNA), which produces striking spatial and temporal heterogeneity in tumors, making EGFR both a prominent driver and a challenging molecule to quantify accurately.8

Furthermore, in recent years, significant focus has been placed on clinical trials, including those targeting GBM-EGFR amplification9, underscoring the critical clinical need for early, rapid, accurate, widely accessible, and scalable identification of patients with GBM-EGFR amplification, which remains a major bottleneck in the neuro-oncological clinical practice.8

On the other hand, AI models, especially machine learning, deep learning (DL), and modeling techniques, have demonstrated outstanding achievements in solving real-world problems. Applications have now been broadened across different fields, such as agriculture, price forecasting,10, 11 energy demand,12 financial and socioeconomic modeling,13, 14 and various clinical diagnoses, for instance, heart disease and breast cancer.15, 16 For each problem, a broad range of methods, such as neural networks,17 Gaussian process regression,11 and graphical causal modeling techniques,18 alongside different hybrid strategies, which combine with the models to capture nonlinear, heterogeneous, and independent patterns,17 while improving the efficiency of the predictive models.

Advancements in digital pathology, such as computerized whole-slide imaging (WSI) have enabled DL models to obtain promising results in tumor diagnosis. Weakly supervised multiple instance learning (MIL) models stand out as the state-of-the-art DL methods for WSI analysis, learning from slide-level labels rather than relying on patch-level annotations, effectively addressing challenges related to the gigantic size of WSIs, label scarcity, and the sparse distribution of tumors.19 Various glioma pathology studies have shown that AI-based prediction models achieved outstanding results in biomarker diagnosis from WSIs. For instance, some recent studies applied a hybrid DL-based convolutional neural network (CNN)-transformer20 and also MIL21, 22 models to hematoxylin and eosin (H&E)- stained WSIs to predict IDH mutation status in gliomas. Consequently, these studies provide compelling evidence that applying AI-based models to WSIs can deliver precise, scalable, and cost-effective biomarker prediction in clinical oncology.

Given the promising diagnostic results of artificial intelligence (AI) using WSIs, several studies have investigated various AI methods for detecting GBM-EGFR amplification. For the diagnosis of EGFR amplification, an AI diagnostic study23 has demonstrated that weakly supervised MIL models surpassed the transformer-based model used in a prior study,24 emphasizing the feasibility of AI methodologies in GBM-EGFR prediction. However, these AI models face severe challenges in clinical translation. Because the weakly supervised MIL approaches depend on slide-level labels and do not fully utilize unlabeled histopathology images, limit their ability to generalize across different and heterogeneous cohorts. Additionally, because WSIs contain various tissue structures that do not contribute equally to the label decision, MIL-based predictions are made on a limited set of selected WSI patches rather than the overall tissue structure. This restricts the models' ability to represent the full morphological diversity of the tissue, prevents the capture of the global tumor pattern, and makes the prediction models vulnerable to bias, providing limited insight into how distinct tissue morphologies contribute to the final prediction.

To address these challenges, we introduced an automatic and optimal prototype-based framework, integrating variational auto encoder (VAE)25 with Dirichlet Bayesian Gaussian Mixture Model (DBGMM)26 for GBM-EGFR amplification prediction directly from WSIs. Firstly, applying a VAE, which compresses the high-dimensional features of WSI-patches into a compact latent space, enables more effective clustering of tissue morphologies. Our method then employs an unsupervised prototype-learning approach that requires no annotations to discover optimal tissue morphological patterns. In fact, based on feature similarities, the model groups tissue regions into distinct morphological clusters, each represented by a prototype, known as a cluster centroid. These prototypes capture semantically meaningful tissue patterns that then guide the EGFR amplification classification model. Moreover, the significant advantage of the proposed model is that, unlike conventional unsupervised models that require a predefined number of clusters, which may lead to over- or under-clustering and affect the performance of the downstream tasks, such as classification, this model automatically determines optimal, meaningful prototypes that are desirable for more precise classification. Therefore, we aimed to integrate these prototypes into the predictive model to enhance its diagnostic performance and interpretability.

We trained our proposed model using the H&E-stained WSIs from the Cancer Genome Atlas (TCGA) EGFR dataset to determine the morphological tissue prototypes automatically. We simultaneously applied a binary classifier for predicting GBM-EGFR amplification using 5-fold cross-validation. Validation on two unseen cohorts provided encouraging evidence of the model's generalizability, although further validation on larger, multi-institutional cohorts is needed to confirm robustness in broader clinical settings. We also carried out a comprehensive performance comparison against MIL baselines23, 24, 27 demonstrating the high performance of our proposed model in GBM-EGFR prediction. Importantly, in this proposed model, we integrated expert pathologist interpretation of learned WSI morphologies and quantified the diagnostic contribution of each learned prototype, thereby providing a clinically interpretable view of how different regions of a tissue are assigned to semantically meaningful patterns, offering a transparent prototype distribution underlying EGFR prediction.

In summary, this prediction model, based on interpretable EGFR tissue morphologies, aims to assist pathologists in integrating it into their diagnostic workflow, thereby supporting faster, more efficient decision-making.

Methods

The workflow of the proposed model is shown in Fig. 1. First, in the preprocessing step, we segmented the WSIs to discard the background, and then, using a tiling strategy, we divided each WSI into patches. In the next step, we extracted patch features using the UNI feature encoder, which produces 1024-dimensional patch embeddings. To compress the extracted features, we used a VAE, reducing them to 512 dimensions. These features were then processed through the DBGMM to automatically determine the optimal prototypes, which were used to cluster each WSI. The prototype-based WSI representations were then used as inputs for the slide-level classifier to discriminate between EGFR-amplified and EGFR-non-amplified cases. For fair evaluation, the compressed 512-dimensional features were fed to the MIL models.

Fig. 1.

Fig. 1

Overview of the proposed framework for EGFR classification. (a) H&E-stained WSIs. (b) WSI preprocessing, including tissue segmentation and patching. (c) Feature extraction using the UNI feature encoder which obtains 1024-dimensional patch embeddings. (d) Feature dimension compression by applying variational auto encoder (VAE), which compressed the original feature embeddings into 512-dimensional latent features. (e) Prototype learning via Dirichlet Bayesian Gaussian Mixture Model (DB-GMM), which determines the optimal clusters. The determined optimal prototypes are used to cluster each WSI, where each cluster represents the WSI's morphology. (f) Prototype assignment on a WSI, with each color corresponding to a specific prototype. (g) The prototype-based WSI representations are then utilized as inputs for the slide-level classifier to discriminate between EGFR-amplified and EGFR-non-amplified cases.

Dataset

We utilized two public and one private GBM H&E-stained WSI datasets with EGFR amplification. The two public datasets were TCGA dataset (https://www.cancer.gov/ccg/) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC) from the National Cancer Institute Centre for Cancer Genomics (https://www.cancer.gov/ccg/). The TCGA dataset comprises 329 samples, including 217 EGFR-amplified and 112 EGFR-non-amplified cases. The CPTAC dataset, used for external validation, comprises 217 samples, including 99 EGFR-amplified and 118 EGFR-non-amplified cases. EGFR amplification was defined as a copy number greater than or equal 4.28 Moreover, we used a private dataset (including 70 samples) from the GB-UK cohort—a single UK center,8 for external validation. It includes 20 WSIs with EGFR-amplified and 50 WSIs with EGFR-non-amplified cases.

WSI preprocessing

Following the CLAM model's29 WSI preprocessing pipeline, we initially performed tissue segmentation30 and tiled the segmented tissue sections into non-overlapped smaller patches of 224 × 224 pixels using 20× magnification (Fig. 1b). Furthermore, we utilized the UNI feature encoder,31 a foundation model trained on over 200 million pathology H&E-stained and IHC images, to extract patch features with 1024 dimensions (Fig. 1c), which were then used in the proposed model for GBM-EGFR prediction.

Feature compression using VAE

The VAE25 learns the latent representation of data using encoder and decoder components (Fig. 1d). Features can be sparse in high-dimensional data, such as WSIs, and the distances between them may become uniform. This can lead to weak clustering due to the reduction of meaningful features required for effective clustering. VAE reduces noise and data sparsity while preserving meaningful variation across tissues, minimizing the risk of overfitting in clustering.32, 33 Therefore, using the reduced and appropriate feature dimensionality results in accelerated computation tasks.33, 34

The VAE encoder learns a probabilistic mapping from high-dimensional input x to low-dimensional latent representations zapplying a variational posterior to transform high-dimensional inputs into low-dimensional representations (Eq. (1)).

q∅zx=Nzμ∅xlogσ∅2x (1)

x represents the high-dimensional original patch features of a WSI, obtained during the preprocessing step. z shows the sampled latent representations, which are used by decoders during training to reconstruct the input x. The encoder predicts two vectors of the distribution, including μ∅x as the mean of latent distribution, and logσθ2x the log variance of the learned latent distribution.

The second component of the VAE is the decoder, which mirrors the encoder to map the produced low-dimensional representations back to their original high-dimensional structure. Indeed, the VAE training process is designed to iteratively adjust and refine the encoder's learned latent representations over multiple epochs through the total reconstruction loss feedback, which measures how well the decoder reconstructs the original input features x. The VAE model is optimized by minimizing the following total loss function (Lx) (Eq. (2)), including reconstruction loss (Lrecon) and Kullback–Leibler divergence loss (lkl).

Lx=Lrecon+Lkl (2)
lrecon=x−x^2 (3)

lreconis a mean-squared error (MSE) (Eq. (3)), which calculates the loss between the input x and the reconstructed input x^ and the Lkl (Eq. (4)) measures how well the latent space aligns with the standard normal prior N0I, where I is the identity matrix, d is the dimension of the latent presentations, and μi represents ith element of the mean vector μix produced by the encoder. Moreover, σi2 indicates ith element predicted variance vector.

Lkl=−1d∑i=1d1+logσi2−μi2−σi2 (4)

Unsupervised optimal prototype discovery with DBGMM

In contrast to the majority of conventional clustering methods, which require predefined numbers of clusters, the DB-GMM takes advantage of the flexibility of the Dirichlet process (DP)35 and the Gaussian Mixture Models (GMMs)36 to determine the optimal clusters in a data-driven process.

The DP is a Bayesian nonparametric approach used to model distributions to learn the number of clusters based on infinite mixture of clusters. DP is defined in Eq. (5), where α indicates the concentration parameter that determines the likelihood of possible new clusters. As the value of α increases, it permits more clusters. G0 presents the base data distribution over cluster parameters such as the component mean.

DPαG0 (5)

In contrast to GMM,36 DBGMM does not rely on a fixed number of clusters and incorporates a non-parametric Bayesian prior in clustering. It fits the input data (latent vectors) z using the following (Eq. (6)):

Pzw,θ=∑k=1KwkNzμk∑k (6)

where Nzμk∑k represents the Gaussian component with mean (μk) and covariance (∑k). K is the number of components, and the DP assigns clusters. wk is the weight of cluster k derived from the DP.

The latent representations are assigned to Gaussian components with different probabilities according to their posterior responsibilities, which are derived from each Gaussian component's likelihood in the Gaussian distribution with its prior weight. Consequently, through a Bayesian prior, components receive low posterior weights if they contain a minimum number of assigned data (latent vectors). To refine the selection of meaningful clusters, we utilized a data-driven thresholding procedure based on the Elbow method33 applied to the posterior weights of the components to identify a weight threshold to prune low-weight components. To reduce the fluctuations in weight distribution and produce a stable curve for knee point detection, the Gaussian smoothing37 was applied to the posterior cluster weights. Considering w as the sorted posterior weights w=w1w2…wK), of the K components derived from DBGMM, the Gaussian smoothing of these weights was calculated as Eq. (7). GaussianFilter1D.σ denotes one-dimensional convolution with a Gaussian kernel of standard deviation (σ).

w∼=GaussianFilter1Dwσ (7)

Eq. (8) indicates how smoothed values are computed. i denotes the rank position being smoothed, whereas j ranges over all rank positions in w. wj shows unsmoothed/raw weights at rank j and Z ensures the sum of all j results in 1. The rank distance ∣i−j∣, which shows how many positions apart two components are in the sorted sequence. Components with smaller ∣i−j∣, ranked close to i, are assigned larger kernel weight and contribute more strongly to w∼i. This contribution decreases according to a Gaussian envelope whose width is set by σ.

w∼i=1Z∑jwje−i−j22σ2,Z=∑je−i−j22σ2 (8)

Here, instead of manually tuning a fixed σ, it is adaptively selected based on the variability in the weight distribution itself, as given in Eq. (9). In fact, σ denotes the degree of smoothing applied to the sorted component weight curve. Therefore, a higher value permits comparatively more smoothing. In contrast, the curve is moderately smoothed with a smaller value of Stdw.

σ=Stdw (9)

More generalized σ is presented by σ=c·Stdw, where c is the scale factor. Its variation tests the robustness of the σ in smoothing the weight curve. In this study, c is set to 1, and for the sensitivity analysis, we applied different values of it. Following the smoothing process, the Elbow method was applied to the cumulative sum of the smoothed weights. In this study, the Elbow method was not applied to directly select the number of clusters; rather, it was used to identify a weight threshold to prune low-weight components in DBGMM. Indeed, the Elbow method identifies the inflection point (knee point), which corresponds to the maximum curvature at rank position i∗. The unsmoothed weight corresponding to the inflection point (weight_threshold = wi∗) is the data-driven weight threshold. Components with weights (wk)exceeding this threshold were retained, whereas components below the threshold were discarded (Eq. (10)). The retained clusters and their corresponding centroids serve as the final WSI prototypes for downstream tasks (Fig. 1e).

Cluster=μkwk≥weight_threshold (10)

GBM-EGFR classification

The optimal prototypes identified by our proposed model, or the predefined prototypes used in competing models, were used for the downstream classification task. For classification purposes, each WSI using its patch features was mapped onto this prototype space (Fig. 1f). Then, a WSI-specific mixture proportion, mean vector, and covariance vector were calculated for each prototype, and then protype-specific statistics were combined to make a WSI-specific representation. In fact, for classification, these slide representations were separated into set of prototype-specific vectors. The count of vectors is equal to the count of prototypes. Therefore, for each vector, an independent MLP was applied, allowing the model to learn distinct morphological patterns. The produced prototype embeddings were then concatenated and passed through a final linear layer38 that generated two class logits for the EGFR-amplified and EGFR-non-amplified class labels.

Furthermore, because MIL-based weakly supervised learning models have gained popularity for their ability to leverage slide-level annotations,39 in this study, three state-of-the-art MIL models, including CLAM,29 TransMIL40, DTFD,41 and attention challenging MIL (ACMIL)42 were used for comparison purposes. CLAM is an attention-based MIL model, whereas DTFD is a two-stage strategy model that uses attention and feature distillation to classify WSIs. Moreover, TransMIL is a transformer-based model that utilizes self-attention to classify WSIs, whereas ACMIL is a multi-attention model that applies Top-k instance masking to hide the highest-attention instances, forcing the model to learn from more challenging patches. Therefore, we used slide-level classification rather than computationally expensive and labor-intensive patch-level predictions.

Evaluation metrics and statistical methods

The classification models were trained using 5-fold cross-validation on the TCGA dataset (329 samples), with a split of 80% for training, 10% for internal validation, and 10% for testing in each fold. The 5-fold cross-validation results are reported as the mean and standard deviation of the performance metrics. For external validation, we assigned 85% of the TCGA dataset for training and 15% for internal validation. Subsequently, we allocated the entire CPTAC (217 samples) and GB-UK (70 samples) datasets as unseen test sets. All splits were at the patient level. The area under the curve (AUC), sensitivity (true-positive rate), specificity (true-negative rate), and F1-score were used to assess performance. To quantify each prototype's contribution to the model's decision-making, we evaluated the distribution using Cohen's d relative to each prototype's probability. Cohen's d effect size was computed using the mean prototype probabilities based on the WSIs' ground-truth, EGFR-amplified and EGFR-non-amplified labels, scaled by the pooled standard deviation.

Implementation details

VAE

Following the preprocessing step of WSIs, the extracted features of each WSI goes through the VAE framework step. First, the encoder projects the 1024-dimensional input using two hidden layers with monotonically decreasing width (1024 → 768 → 640). Afterwards, it branches into two parallel linear output layers of 640 → 512, each independently providing the μϕxand logσϕ2x. Then, the decoder mirrors these steps and reconstructs the 1024-dimensional data with the final linear output layer. In each hidden layer, a linear transformation followed by normalization, ReLU nonlinearity, and a dropout rate = 0.2 was applied. Due to the computational simplicity of ReLU and its avoidance of vanishing gradients, which are prevalent issues in tanh and sigmoid, we selected ReLU for hidden layers. Kaiming43 initialization is used for hidden-layer weights, which is consistent with ReLU, whereas Xavier initialization44 was applied to the posterior parameter heads, (μϕx and logσϕ2x and the final reconstruction layer. Because gradual dimension reduction prevents the loss of information,45 we applied this dimension reduction strategy rather than a single direct reduction. For 1024 → 768 → 640 → 512, using a similar compression ratio at each step: ∼0.75, ∼0.83, ∼0.80, avoids one layer being disproportionately more compressed than others. The VAE was optimized by minimizing the total loss utilizing Adam with an initial learning rate of 0.001 and no additional weigh decay. The learning rate was reduced by a factor of 0.5. Training was conducted for a maximum of 50 epochs. Based on the devised early stopping, the training was stopped early if the monitored validation failed to improve by at least 0.0001 for 10 consecutive epochs. Moreover, a batch size of 64 was preferred. We utilized cloud-based computational resources, including a NVIDIA V100 GPU with 64 GB of GPU memory and 64 GB of system memory. Moreover, to ensure a fair evaluation, we applied VAE to all three datasets.

DBGMM

DBGMM was applied on whole TCGA cohort using scikit-learn's Bayesian Gaussian Mixture implementation with the following configurations. The K components were assigned to 50 with diagonal covariance (∑k). The DP concentration parameter was assigned to α=0.05. Other priors were kept as their standard bayes defaults. The default standard conjugates Normal-Wishart (Normal-Gamma under diagonal covariance) was taken for G0. Regarding convergence criteria, the model was fit up to 500 variational EM iterations alongside a convergence tolerance of 0.001. Before knee-point detection, the sorted component weights were smoothed using σ=Stdw with a scale factor (c=1). Robustness of the retained prototype set was evaluated through different scaling factors in the ablation study section.

WSI classification

The downstream classification model used an individual MLP for each prototype-specific representation, which was six optimal prototypes for our proposed framework and different numbers for competing models. Each MLP included a single fully connected layer followed by a ReLU activation. These resulting prototype embeddings were then concatenated and passed through a final linear layer to produce EGFR-amplified and EGFR-non-amplified class logits. The classifier was trained utilizing cross-entropy loss and the Adam optimizer with a learning rate of 0.0001 and dropout of 0.1. We trained the model for a maximum of 200 epochs, with early stopping when no improvement was observed for 30 consecutive epochs. Furthermore, the model was trained with a fixed random seed.

Results

Prototype learning discovers distinct morphological prototypes in glioblastoma

Following the EGFR WSI dataset preprocessing (Fig. 1b), 1024-dimensional patch features were extracted (Fig. 1c). We trained the classification model with different compressed dimensional sizes. Hence, the application of VAE-compressed 1024-dimensional features to 512 low-dimensional latent representations (Fig. 1d).

We applied DBGMM (Fig. 1e) to identify an optimal set of morphological prototypes from 512-dimensional features, yielding six distinct prototypes. These distinctive and optimal prototypes captured semantically meaningful tissue patterns without requiring predefined cluster numbers. Fig. 2(a1) and (b1) are examples EGFR status illustrating the original H&E-stained WSI, and the optimal prototype assignment maps over a representative WSI. Each patch is assigned to one of six defined clusters, showing spatial heterogeneity captured in the feature space. Furthermore, Fig. 2 (c1) (Proportion-Cluster) depicts the estimated prototype distribution within a WSI. In fact, it quantifies the proportion of patches per prototype, showing that the relative abundance of prototypes varies significantly across WSIs. Each bar corresponds to one of the six optimal clusters (C0–C5) determined by the proposed model. For instance, for a1 (Proportion-Cluster distribution) cluster C4 was the dominant tissue section, followed by cluster C5, and other clusters represent a lower proportion of the WSI. Moreover, as shown in Fig. 2c1, for b, the Proportion-Cluster distribution indicated that C3 was the dominant prototype, whereas C2 contributed a moderate proportion. Other clusters' proportion contribution remained at a minimal level. Fig. 2(c2) illustrates representative patch examples for each prototype in each WSI. The patches of each prototype were reviewed by a neuropathologist, who provided morphological interpretation. Therefore, these descriptions provide more clinically and morphologically interpretable prototype patterns.

Fig. 2.

Fig. 2

Visualization of the optimal prototypes derived by the automatic prototype learning framework on a representative WSI. (a1) and (b1) illustrate the original H&E-stained WSIs and the spatial distribution of patch-level cluster assignments on WSI tissues. (c1) Proportional visualization of patches assigned to each of the six optimal prototypes. Cluster 4 (C4) exhibits the highest mixture proportion, followed by a moderate proportion of patches by cluster 5 (C5) in a1. In b1, cluster 4 (C4) has the most extensive tissue coverage, followed by cluster 2 (C2) with a moderate proportion contribution. c2 panels show the representative patches related to each prototype, illustrating the distinct morphological patterns identified within each WSI. The reviewing neuropathologist added brief morphological descriptions after the model identified the prototypes.

Prototype-specific contributions to EGFR amplification prediction

We examined the distribution of prototypes based on the proportion of patches assigned to each of the six optimal prototypes across the test set (Fig. 4). Although the relative distribution of prototypes varied across WSIs (Fig. 4), the morphological prevalence of the prototypes did not reflect their diagnostic contribution. Therefore, to quantify the contribution of each learned morphological prototype to the EGFR amplification prediction, we evaluated the distribution using Cohen's d relative to prototype probabilities. To clarify how we derived each prototype-specific probability of each WSI, following the classification section, once the six optimal prototypes were identified, the mean and covariance vectors for each prototype were computed for each WSI. These prototype-specific statistics were combined to form a WSI-specific representation; subsequently, they were reorganized into six vectors, and each vector went through a prototype-specific MLP. The resulting prototype embeddings were concatenated and passed through a final linear layer, generating two class logits. To extract the prototype-specific label probability, the weight matrix of the final linear layer was divided into six weight blocks, each corresponding to one of the six concatenated prototype embeddings. Each weight block was applied to its corresponding prototype embedding to derive two prototype-specific logit contributions, representing support for the EGFR-amplified and EGFR-non-amplified classes. Softmax was then applied separately to each prototype-specific logit pair, producing prototype-specific class-support probabilities. As a result, for six prototypes, each WSI yielded six probability pairs, where each prototype-specific probability shows the class support provided by that prototype in the classification of the given WSI. These prototype probabilities, which are uncalibrated class supports specific to prototypes, were used only in this section for Cohen's d calculation. They are not six separate final WSI prediction results. The primary overall classification outcomes in this study are derived from the combined logit contributions of all prototypes, with the final linear layer generating two class logits. Therefore, Table 1 shows the mean prediction probability for each of the six optimal prototypes separately for WSIs with EGFR-amplified and for WSIs with EGFR-non-amplified labels. On average, across the test slides shown in Table 1, prototypes C4, C3, C2, and C5 had higher mean probabilities in WSIs with EGFR-amplified labels, indicating more support for an EGFR-amplified diagnosis. Conversely, prototypes C1 and C0 showed higher probabilities for EGFR-non-amplified labels. Furthermore, Cohen's d was computed considering the derived prototype-specific class support probabilities.

Fig. 4.

Fig. 4

Aggregated prototype mixture weights of EGFR status. The plot shows the mean proportion of patches assigned to each of the six optimal prototypes across EGFR-non-amplified and EGFR-amplified WSIs.

Table 1.

Prototype-specific contribution patterns for EGFR-amplified and EGFR-non-amplified.

Prototype Mean probability in WSIs with EGFR-amplified label Mean probability in WSIs with EGFR-non-amplified label Cohen's d
C0 0.5732 0.4507 0.7326
C1 0.5389 0.4903 0.3269
C2 0.6268 0.4507 1.0589
C3 0.6491 0.4394 1.0463
C4 0.7048 0.3616 2.0790
C5 0.6050 0.4064 1.1575

Cohen's d was calculated using the difference between the mean probabilities of these two groups, standardized by their pooled standard deviation, thereby measuring differences in prototype-specific correct-class confidence between EGFR-amplified and EGFR-non-amplified cases. Indeed, Cohen's d quantifies how strongly each prototype separates WSIs with EGFR-amplified labels from EGFR-non-amplified labels. Hence, Cohen's d was calculated as the difference between the mean of all prototype-specific probabilities supporting the WSIs' ground-truth for EGFR-amplified and EGFR-non-amplified labels, divided by the pooled standard deviation. According to Table 1, the effect size was large for prototype C4, with a very large Cohen's d (2.0790), indicating a robust separation between EGFR-amplified and EGFR-non-amplified. This was closely followed by prototypes C5 (Cohen's d = 1.1575), C2 (Cohen's d = 1.0589), and C3 (Cohen's d = 1.0463), which had a high discriminative impact on the model's EGFR-amplified decision. Conversely, C0 presented a medium effect size at Cohen's d = 0.78, whereas C1 had a small effect size (Cohen's d = 0.39).

This effect size quantifies separation in model-driven prototype scores and does not directly interpret the prototypes' corresponding morphological description quality specific to EGFR amplification prediction. Nevertheless, considering the strong separation of C4 and C3 and the expert's histopathological review of the slides revealed that C4 and C3 prototypes were characterized by low-cell-density regions with either minimally infiltrated or subtly abnormal brain tissue (Fig. 2), suggestive of a distinct EGFR-amplified infiltration pattern.

Optimal clusters outperform predefined clustering settings in EGFR amplification prediction

We conducted a comprehensive classification evaluation based on automatically determined clusters (six optimal clusters from the proposed automatic prototype learning) and manually selected predefined cluster numbers, followed the PANTHER38 framework (using k-means clustering strategy). Then, these prototypes were used in the classification stage. Table 2 summarizes MLP classification results for the proposed automatic prototype learning-based prototypes and predefined prototype numbers, obtained via 5-fold cross-validation on the TCGA cohort comprising 329 samples.

Table 2.

5-fold cross-validation classification results (on TCGA cohort) based on clustering methods and given prototype numbers.

Clustering Prototype AUC (Mean ± SD) Specificity (Mean ± SD) Sensitivity (Mean ± SD) F1-score (Mean ± SD) p-value (t-test) p-value (Wilxon)
Proposed automatic prototype learning 6 0.8087 ± 0.0153 0.7380 ± 0.1181 0.9385 ± 0.0422 0.9066 ± 0.0326 – –
Predefined cluster numbers 4 0.6914 ± 0.0402 0.4444 ± 0.0314 1.0000 ± 0.000 0.8780 ± 0.0229 0.0062 6.828E-07
5 0.7408 ± 0.0230 0.4775 ± 0.1407 0.9900 ± 0.0200 0.8597 ± 0.0844 0.0075 1.30E-06
6 0.7653 ± 0.0172 0.5654 ± 0.0560 0.9900 ± 0.0200 0.8760 ± 0.0706 0.0023 1.48E-06
7 0.7746 ± 0.0352 0.5360 ± 0.1171 0.9900 ± 0.0200 0.8696 ± 0.0846 0.0344 1.30E-06
8 0.7568 ± 0.0418 0.1600 ± 0.2585 0.9895 ± 0.0211 0.7954 ± 0.1240 0.067 3.51E-06
9 0.7544 ± 0.0218 0.5137 ± 0.0994 0.9900 ± 0.0200 0.8650 ± 0.0821 0.0038 1.30E-06
10 0.7508 ± 0.0143 0.4985 ± 0.1415 0.9900 ± 0.0200 0.8660 ± 0.0793 0.0045 1.30E-06
11 0.7467 ± 0.0251 0.5061 ± 0.1101 0.9835 ± 0.0209 0.8592 ± 0.0841 0.0022 1.20E-06
12 0.7452 ± 0.0277 0.4929 ± 0.1293 0.9900 ± 0.0200 0.8648 ± 0.0749 0.0093 1.45E-06
13 0.7385 ± 0.0207 0.4399 ± 0.1504 0.9900 ± 0.0200 0.8505 ± 0.0916 0.0036 1.16E-06
14 0.7370 ± 0.0202 0.4419 ± 0.1302 0.9900 ± 0.0200 0.8516 ± 0.0836 0.0028 1.20E-06
15 0.7289 ± 0.0196 0.4553 ± 0.1353 0.9900 ± 0.0200 0.8553 ± 0.0828 0.003 1.30E-06
16 0.7112 ± 0.0423 0.0133 ± 0.0267 1.0000 ± 0.000 0.7810 ± 0.1050 0.0026 9.43E-07
17 0.7074 ± 0.0217 0.4086 ± 0.0800 0.9900 ± 0.0200 0.8472 ± 0.0800 0.0008 9.13E-07
18 0.6817 ± 0.0425 0.0133 ± 0.0267 1.0000 ± 0.000 0.7810 ± 0.1050 0.0036 1.14E-07
32 0.6255 ± 0.0636 0.1133 ± 0.1950 1.0000 ± 0.000 0.7931 ± 0.1173 0.0001 8.41E-08

In Table 2, boldface highlights the proposed model results and the highest numerical values among the predefined-prototype competing models for each performance metric; bold p-values indicate statistically significant differences at p < 0.05. These highlighted values are interpreted in the Results section

The results demonstrated that the classifier using 6 optimal prototypes outperformed (AUC: 0.8087 ± 0.0153 and F1-score: 0.9066 ± 0.0326) all classification results based on the k-means clustering settings with various cluster numbers, which achieved the highest performance at seven clusters (AUC = 0.7746 ± 0.0352, F1-score = 0.8696). However, these results are still lower than the classifier using the optimal prototypes automatically derived from the proposed model. Furthermore, although several classification models using k-means-derived prototypes (e.g., 4–15 prototypes) achieved high F1-scores (>0.85), they consistently showed high sensitivity but comparatively low specificity, indicating a tendency to over-predict EGFR-amplified classes. In contrast, the classifier using optimal prototypes, achieved a higher AUC with a more balanced sensitivity and specificity. The receiver operating characteristic (ROC) curves (Fig. 5a) show that the classifier using the optimal clusters achieved a higher AUC for predicting EGFR cases. Fig. 5b illustrates the variation in classification performance (AUC) of the k-means-derived prototypes. The curve shows that AUC initially increases, reaches a peak at seven clusters (AUC of (0.7746 ± 0.0352)), and then declines, indicating diminished classification performance with large numbers of clusters. These findings underscore the diagnostic importance of automatically learning the representative tissue morphologies, achieved by our proposed model, in supporting more accurate GBM-EGFR classification from H&E-stained WSIs.

Fig. 5.

Fig. 5

EGFR classification performance evaluation across various prototype settings. (a) ROC curves were attained from a 5-fold cross-validation setting using the TCGA dataset. The MLP classifier employing the optimal six prototypes derived from the proposed model obtained a superior AUC of 0.8087 ± 0.0153, outperforming the MLP classifiers using predefined prototypes. (b) Illustration of the relationship between AUC of the EGFR classification and the number of prespecified clusters based on k-means. The highest performance was observed at seven prespecified clusters (AUC:0.7746 ± 0.0352), followed by a consistent AUC degradation, suggesting model's degradation related to over-clustering.

To investigate whether the proposed model's AUC improvements over 16 baselines were statistically significant, we conducted two statistical analyses (Table 2). First, a paired t-test was applied to fold-level AUC values from the 5-fold cross-validation splits comparing the proposed model against each predefined-prototype baseline. The results demonstrate that the improvement was statistically significant for 15 out of 16 models, with p-values <0.05. The baseline model with eight predefined prototypes showed a positive AUC difference; however, it did not reach statistical significance, suggesting that reliance on 5-fold cross-validation constrains statistical robustness, particularly when fold-level comparisons are smaller or less consistent across folds. Furthermore, we conducted a second statistical analysis: a one-sided paired Wilcoxon signed-rank test using the predicted probability assigned to the ground-truth labels. The results showed p-values ranging from 8.41×10−8 to 3.51×10−6, demonstrating the proposed model's significance over 16 baseline models (p-value <0.05). To further examine classification errors, representative false-negative (FN) and false-positive (FP) cases, along with their corresponding prototype distributions, are shown in Fig. 3. Across the 5-fold cross-validation on TCGA, 14 FP and 7 FN predictions were observed, yielding a mean overall error rate of 11.88% ± 2.95%.

Fig. 3.

Fig. 3

Visualization of misclassified cases on representative WSIs. (a1) and (a2) illustrate the original H&E-stained WSIs and the spatial distribution of patch-level cluster assignments on WSI tissues for FN and FP cases, respectively. (b1 and b2) Proportional visualization of patches assigned to each of the 6 optimal prototypes for FN and FP cases, respectively.

EGFR amplification prediction using optimal clusters generalizes across cohorts

Considering the importance of developing a GBM-EGFR diagnostic model that generalizes across various cohorts, we conducted a comprehensive evaluation of our proposed optimal clustering-based classification model using two unseen external cohorts: CPTAC (including 217 samples) and GB-UK (including 70 samples). The classifier utilized the TCGA dataset (including 329 samples) as the training set and the two unseen datasets as external test sets. Subsequently, the model was trained using the proposed automatic learning-derived optimal clusters, as well as a series of predefined cluster numbers (4–18 and 32), based on the k-means. Table 3 and Table 4 summarize the classification results of the CPTAC and GB-UK evaluations, respectively.

Table 3.

Classification performance evaluation using the TCGA trainset and the CPTAC unseen test set.

Clustering Prototype AUC Specificity Sensitivity F1-score
Proposed automatic prototype learning 6 0.7740 0.6525 0.8585 0.7555
Predefined cluster numbers 4 0.7119 0.5084 0.9494 0.7490
5 0.7165 0.5338 0.9494 0.7580
6 0.7262 0.5423 0.9494 0.7611
7 0.7312 0.5508 0.9494 0.7642
8 0.6970 0.1300 1.0000 0.786
9 0.6916 0.4661 0.9494 0.7343
10 0.6899 0.4745 0.9393 0.7322
11 0.6896 0.4576 0.9393 0.7265
12 0.6881 0.4745 0.9292 0.7272
13 0.6864 0.4491 0.9494 0.7286
14 0.6797 0.4745 0.9191 0.7222
15 0.6785 0.4661 0.9191 0.7193
16 0.6670 0.4290 1.0000 0.750
17 0.6644 0.4067 0.9393 0.7099
18 0.6610 0.5710 0.8780 0.7780
32 0.6400 0.355 1.0000 0.723

In Tables 3 and 4, boldface highlights the proposed model results and the highest numerical values among the predefined-prototype models for each metric, including individual high values that are specifically interpreted in the Results section in the context of the other performance measures. It is interpreted in the Results section.

Table 4.

Classification performance evaluation using the TCGA trainset and GB-UK unseen test set.

Clustering Prototype AUC Specificity Sensitivity F1-score
Proposed automatic prototype learning 6 0.7870 0.660 0.90 0.6546
Predefined cluster numbers 4 0.7241 0.5423 0.9494 0.7611
5 0.7372 0.5593 0.9494 0.7673
6 0.7488 0.5762 0.9494 0.7736
7 0.7551 0.5847 0.9494 0.7768
8 0.7430 0.540 0.950 0.6129
9 0.7410 0.5677 0.9494 0.7704
10 0.7312 0.5508 0.9494 0.7642
11 0.7217 0.5423 0.9494 0.7611
12 0.7179 0.5338 0.9494 0.7580
13 0.7158 0.5254 0.9494 0.7550
14 0.7138 0.5254 0.9393 0.750
15 0.7119 0.5084 0.9494 0.7490
16 0.7067 0.5169 0.9393 0.7469
17 0.6928 0.4745 0.9494 0.7372
18 0.690 0.360 1.0000 0.555
32 0.6350 0.2800 1.0000 0.5263

In Tables 3 and 4, boldface highlights the proposed model results and the highest numerical values among the predefined-prototype models for each metric, including individual high values that are specifically interpreted in the Results section in the context of the other performance measures. It is interpreted in the Results section.

Using optimal clusters, the classifier applied to the unseen CPTAC dataset (Table 3) yielded an AUC of 0.7740, specificity of 0.6525, sensitivity of 0.8585, and F1-score of 0.7555, surpassing all classifiers using predefined k-means-derived prototypes. The highest performance observed among prespecified cluster numbers was achieved by the classifier utilizing seven clusters with an AUC of 0.7312. However, classification performance diminished as the number of predefined clusters exceeded this peak. For example, when the cluster numbers increased to 32, overall performance decreased significantly, with AUC dropping to 0.6400, specificity to 0.355, and F1-score to 0.723, despite sensitivity remaining at 1. These findings highlight the negative impact of over-clustering, which results in an artificial increase in sensitivity at the cost of specificity and diminishes the classifier's performance in distinguishing between EGFR-amplified and EGFR-non-amplified cases.

Fig. 6a illustrates the ROC curves of the classifiers using clusters derived from our proposed model versus predefined settings on the CPTAC test set. The high performance of the classifier based on optimal clusters is visually noticeable, emphasizing the generalizability of using automatic prototype learning-derived clusters in classifying EGFR mutation while utilizing unseen and external datasets. Furthermore, Fig. 6b depicts the AUC trends for the classifiers across various prespecified cluster numbers. The classifier achieved its highest performance using seven clusters, after which the performance declined as the number of clusters increased.

Fig. 6.

Fig. 6

Performance evaluation using the CPTAC external validation cohort. (a) ROC curves evaluating the performance of the linear classifier applying the proposed optimal prototypes versus predefined prototypes. The classifier utilizing the proposed clustering approach obtained the highest AUC (0.7740), indicating its enhanced generalization capability while using unseen datasets. (b) AUC versus cluster numbers in predefined clusters from k-means. Whereas the highest AUC was achieved with seven prespecified clusters, performance consistently declined with an increasing number of clusters, reaffirming the potential issue of over-clustering and the robustness of our proposed approach.

The classifier trained on the TCGA dataset and tested on an unseen GB-UK dataset (Table 4) showed that the classification model utilizing six optimal prototypes obtained the highest AUC of 0.7870 with specificity of 0.660, sensitivity of 0.900, and F1-score of 0.6546. It also surpassed all classifiers using predefined clustering settings (Table 4 and Fig. 7a). Among predefined setting, the classification model based on seven prototypes reached achieved the highest performance (AUC of 0.7551) and after this peak point, performance declined as the number of predefined clusters increased. Moreover, as cluster numbers increased to 18 and 32, the classifier's performance noticeably diminished (Fig. 7b). Subsequently, when using 18 and 32 clusters, the AUC dropped to 0.690 and 0.6350, respectively, and specificity decreased to its lowest values of 0.36 and 0.28, despite maintaining sensitivity at 1. Consequently, the high sensitivity in several k-means-based models reflects a tendency toward over-prediction of EGFR-amplified, resulting in low specificity and indicating poor true-negative discrimination.

Fig. 7.

Fig. 7

Performance evaluation using the GB-UK external validation cohort. (a) ROC curves compare the performance of the MLP classifier using the proposed optimal prototypes versus prespecified prototypes. The MLP classifier utilizing the proposed clustering approach yielded the highest AUC (0.7870), surpassing all fixed clustering settings, demonstrating its robust generalizability on external datasets. (b) AUC versus cluster numbers in fixed cluster settings. The AUC peaked at seven clusters (AUC: 0.7551), followed by a consistent decline as the cluster numbers increased. The results demonstrate the effectiveness of using optimal prototypes on classification performance and the potential over-clustering issues associated with predefined clusters based on k-means.

Overall, these findings confirm that the optimal prototypes derived from the proposed model offer generalization, which enhances classifiers' discriminative capability on unseen data. Noting that the external cohorts evaluated here are relatively small and limited in diversity, further analysis in broader clinical settings is suggested.

Optimal clustering outperforms MIL approaches in GBM-EGFR amplification prediction

Given the clinical significance of a reliable GBM-EGFR amplification prediction, we compared the classification performance achieved by our proposed automatic clustering approach with established MIL models. The primary objective of this comparison is to evaluate whether our proposed approach can offer improvements and robustness in EGFR amplification prediction compared to MIL models. Table 5 summarizes the 5-fold cross-validation results using TCGA dataset, under three different setups, including classification results using our proposed optimal clusters, three state-of-the-art MIL models (i.e., CLAM,29 TransMIL40 and DTFD,41 and ACMIL42) using UNI feature encoder31 and also the same three MIL models of the previously published study14 using ResNet feature encoder.46 Table 5 also includes the results of a transformer-based model24 using different multi-institutional datasets (TCGA/BWH/UPenn). The classifier using optimal prototypes, trained and tested on the TCGA dataset, achieved the highest AUC (0.8087 ± 0.0153). In contrast, MIL models using the same TCGA dataset achieved lower AUC: CLAM (0.763 ± 0.085), DTFD (0.721 ± 0.062), and TransMIL (0.690 ± 0.161) and higher specificity values. For instance, DTFD achieved the highest specificity of 0.961 ± 0.056. Among the MIL models, ACMIL achieved the second highest AUC (0.7585 ± 0.0270) with a sensitivity of 0.9900 ± 0.0223 and the lowest specificity of 0.4883 ± 0.1137. The highest specificity of the MIL models could not elevate the overall classification performance and led to lower AUC and sensitivity in most EGFR cases, indicating the classifier's bias toward EGFR-amplified cases. The classification results of the study,23 which applied the same MIL models (reported an AUC of 0.736 ± 0.181, 0.706 ± 0.076, and 0.688 ± 0.172 for CLAM, DTFD, and TransMIL, respectively), and another study24 using a transformer-based model with an AUC of 0.71 ± 0.031 remained below the performance of our proposed approach. Consequently, the results suggest that the classification model using our proposed prototype strategy outperformed MIL-based models, enhancing the EGFR prediction efficiency and yielding higher AUC and sensitivity while maintaining an effective balance with specificity.

Table 5.

5-fold cross-validation performance (Mean ± SD) evaluation using MIL models.

Dataset Model type Feature encoder Ref Model AUC Specificity Sensitivity
TCGA Proposed prototype-based UNI This study MLP 0.8087 ± 0.0153 0.7380 ± 0.1181 0.9385 ± 0.0422
CLAM 0.7630 ± 0.085 0.741 ± 0.160 0.786 ± 0.128
TCGA MIL-based UNI This study DTFD 0.721 ± 0.062 0.961 ± 0.056 0.558 ± 0.103
TransMIL 0.69 ± 0.161 0.088 ± 0.018 1.0000 ± 0.000
ACMIL 0.7585 ± 0.0270 0.4883 ± 0.1137 0.9900 ± 0.0223
TCGA MIL-based ResNet 14 CLAM 0.736 ± 0.181 0.545 ± 0.351 0.741 ± 0.194
DTFD 0.706 ± 0.076 0.959 ± 0.067 0.512 ± 0.113
TransMIL 0.688 ± 0.172 0.627 ± 0.419 0.548 ± 0.604
TCGA/BWH/UPenn MIL-based Vision Transformers N/A 15 Transformer based-CHARM 0.71 ± 0.031 N/A N/A

In Table 5, boldface highlights the proposed model’s highest AUC and the highest values obtained among the MIL-based competing models for the corresponding metrics. The same convention is applied in Table 6 for the external-cohort comparisons. It is interpreted in the Results section.

Generalization capability of optimal-prototype-based classification over MIL models on external cohorts

We benchmarked classification performance beyond cross-validation settings. We conducted an external validation comparing optimal prototype models with MIL-based models (CLAM, TransMIL, DTFD, and ACMIL)42 utilizing two unseen cohorts: CPTAC and GB-UK (Table 6). All classifiers were trained on the TCGA dataset and tested using whole external datasets, separately. The classifier using optimal clusters and tested on the CPTAC outperformed all other models with an AUC of 0.7740, specificity of 0.6525, sensitivity of 0.8585, and F1-score of 0.7555. In contrast, MIL models did not show significant performance. The CLAM obtained the highest AUC (0.7698) among all MIL-based models tested on CPTAC, with remarkably low specificity (0.200), indicating the model's bias toward the prediction of positive cases (sensitivity: 0.800) at the cost of high prediction of FPs. The ACMIL obtained the second highest AUC of 0.7603 and the low specificity of 0.4915. TransMIL and DTFD exhibited low performance, with AUCs of 0.7581 and 0.7537, respectively, alongside notably lower F1-scores of 0.5192 and 0.5121.

Table 6.

Performance evaluation between optimal prototype-based classifiers and MIL models using unseen cohorts.

Test dataset Model type Model AUC Specificity Sensitivity F1-score
CPTAC Proposed prototype-based MLP 0.7740 0.6525 0.8585 0.7555
MIL-based CLAM 0.7698 0.2000 0.8000 0.6154
TransMIL 0.7580 0.5339 0.5454 0.5192
DTFD 0.7537 0.5339 0.53535 0.51208
ACMIL 0.7603 0.4915 0.8282 0.6804
GB-UK Proposed prototype-based MLP 0.7870 0.6600 0.9000 0.6546
MIL-based CLAM 0.7605 0.6800 0.800 0.6153
TransMIL 0.7380 0.020 0.8500 0.3953
DTFD 0.7348 0.5000 0.5353 0.5023
ACMIL 0.7520 0.6000 0.9000 0.6206

In Table 5, boldface highlights the proposed model’s highest AUC and the highest values obtained among the MIL-based competing models for the corresponding metrics. The same convention is applied in Table 6 for the external-cohort comparisons. It is interpreted in the Results section.

Comparing the models tested on the GB-UK external dataset, we observed a significant gap between prototype-based models and MILs. The classifier using optimal-prototype approach outperformed all other models. Among MIL models, CLAM achieved higher classification performance with an AUC of 0.7870, specificity of 0.660, sensitivity of 0.9000, and F1-score of 0.6154. The ACMIL gained the second highest AUC of 0.7520 with the highest sensitivity of 0.9000. The TransMIL obtained relatively low performance with an AUC of 0.738 and notably low specificity (0.020). Similarly, DTFD resulted in an AUC of 0.7348 and reduced sensitivity (0.5353).

Ablation study

Evaluation of how various dimensional features impact classification performance

To validate the selected latent dimension for the proposed framework, we conducted a sensitivity analysis by comparing various dimensions against our proposed 512-dimensional latent features. This evaluation includes both the discovery of prototypes and the performance of downstream classification (see Table 7). The classification results are based on the TCGA dataset using the same 5-fold cross-validation that was used in this study. Using full 1024-dimensional features without applying VAE to compress the features achieved seven prototypes and resulted in a moderate decline in EGGFR classification performance, with AUC of 0.7972 ± 0.0194, sensitivity of 0.8582 ± 0.0620, specificity of 0.5953 ± 0.1498, and F1-score of 0.8302 ± 0.0626. Hence, it shows a 1.42% relative reduction in AUC. Moreover, its computational cost (21 h) is substantially higher than that of our proposed 512-dimensional features (10 h).

Table 7.

Performance assessment using different dimensional features.

Metrics (mean ± SD) 512D (Current study) 1024D (Without VAE) 256D 128D
AUC 0.8087 ± 0.0153 0.7972 ± 0.0194 0.6426 ± 0.0337 0.5968 ± 0.0519
Sensitivity 0.9385 ± 0.0422 0.8582 ± 0.0620 0.9795 ± 0.0281 0.9589 ± 0.0655
Specificity 0.7380 ± 0.1181 0.5953 ± 0.1498 0.2264 ± 0.2039 0.1823 ± 0.2327
F1-score 0.9066 ± 0.0326 0.8302 ± 0.0626 0.8047 ± 0.1194 0.7825 ± 0.124
Cost (h) 10 21 8 7.5

In Table 7, boldface is additionally used to draw attention to notable comparative values; specifically, the bold training time of 21 h for the 1024-dimensional configuration indicates its substantially greater computational cost relative to the proposed 512-dimensional configuration (10 h) and therefore represents a disadvantage rather than superior performance. It is interpreted in the Results section.

Applying 256- and 128-dimensional representations reduced the number of prototypes to 4. Moreover, due to differences in the applied 256- and 128-dimensional representations, the retained prototypes are not identical across the two dimensions. Both dimensions showed a more severe decline in performance metrics. The AUC of the classification model utilizing 256-dimensional latent features declined to 0.6426 ± 0.0337, indicating a 20.5% performance reduction. More importantly, specificity decreased by 69.32%. Furthermore, the AUC of the classification model trained and tested on the 128-dimensional compressed features fell to 5968 ± 0.0519, demonstrating a 26.20% reduction compared to the classification model using our proposed 512-dimensional features. Utilizing both 256- and 128-dimensional features resulted in an imbalance of sensitivity and specificity metrics, with sensitivity values of 0.9795 ± 0.0281 and 0.9589 ± 0.0655, and specificity values of 0.2264 ± 0.2039 and 0.1823 ± 0.2327 for the 256- and 128-dimensional configurations, respectively. The computational cost was moderately reduced by employing 256- and 128-dimensional features, owing to the smaller pooled feature matrix. The classifier employing the proposed 512-dimensional configuration attained the best overall metrics within a moderate computational expense (10 h), thereby supporting the adoption of the proposed VAE-based feature compression within the framework.

DBGMM ablation experiments

α is the DP concentration parameter and controls how sparse the posterior weight distribution can become. Although it does not directly select the prototypes, it shapes the initial raw weight vectors that will go through the elbow thresholding procedure. To evaluate the practical impact of the concentration parameter (α), we compared the classification performance across models utilizing various α values (0.01, 1, and 5), with the results of our proposed model using α=0.05. This sensitivity analysis was conducted on the TCGA dataset using 5-fold cross-validation, fixing all parameters and varying only α. The comparison results (Table 8) presents that the final retained prototype set was identical (six prototypes) using α=0.05and0.01, and therefore gained the same AUC of 0.8087 ± 0.0153 suggesting stability in the neighborhood of the selected values. As the value of α increased to 1 and 5, the corresponding retained prototype count decreased to 8 at α=1 and 15 at α=5, and showed reduced AUC of 0.7206 ± 0.0106 and 0.6944 ± 0.0290, respectively. These outcomes support our selected small and sparsity-inducing concentration parameter.

Table 8.

Impact of α on retained prototypes and classification AUC.

α Final retained prototypes AUC (mean ± SD)
0.01 6 0.8087 ± 0.0153
0.05a 6 0.8087 ± 0.0153
1 8 0.7206 ± 0.0106
5 15 0.6944 ± 0.0290
a

In this study, α was assigned to 0.05.

Because the final threshold is determined by knee-detection applied to the cumulative weights using the elbow method, we applied Gaussian smoothing only to moderately smooth the weight curves before the elbow method. To investigate the impact of c on the classification results and the final retained prototypes, we used the same hyperparameters as the proposed model and subsequently selected various values of c on the TCGA dataset using 5-fold cross-validation (Table 9). The final retained prototypes obtained by the Elbow method and the AUC results (0.8087 ± 0.0153) were unchanged with c=136, confirming that the selected scale contributes only to mild denoising without altering the underlying structural threshold. In contrast, larger c values (10 and 30) showed decreased AUC with a high number of prototypes (7 prototypes for c=10 and 23 for c=30) suggesting that larger c spreads the weights far apart, which can introduce an arbitrary threshold and might dilute the Elbow method's ability to find the actual threshold.

Table 9.

Impact of con retained prototypes and classification AUC.

c Retained prototypes AUC (mean ± SD)
1a 6 0.8087 ± 0.0153
3 and 6 6 0.8087 ± 0.0153
10 7 0.7862 ± 0.0081
30 23 0.7510 ± 0.0105
a

In this study, c was assigned to 1.

Ablation analysis of the GB-UK external cohort using different imbalanced-aware loss functions

The TCGA training cohort contained a higher proportion of EGFR-amplified cases (217) than of EGFR-non-amplified cases (112), whereas the GB-UK external cohort had the opposite class distribution, comprising 20 EGFR-amplified and 50 EGFR-non-amplified WSIs. To investigate whether imbalance-aware training could improve the sensitivity–specificity balance on this external cohort, we retrained the TCGA classifier using focal loss and class-weighted cross-entropy and then evaluated the resulting models on the GB-UK cohort. Table 10 presents a performance comparison of the cross-entropy loss, used as the default loss function in the current study, and the imbalance-aware focal and class-weighted cross-entropy loss functions. With the default cross-entropy, the model obtained a sensitivity of 0.90 and specificity of 0.66, corresponding to 2 FNs and 17 FPs. Focal loss reduced FPs to 14, increasing specificity to 0.72, balanced accuracy to 0.81, and F1-score to 0.692 while preserving sensitivity at 0.90. Class-weighted cross-entropy reduced FPs to 12. It also increased specificity to 0.76, balanced accuracy to 0.83, and F1-score to 0.720, whereas sensitivity remained at 0.90. Moreover, AUC remained at 0.787 across all three loss functions indicating that the imbalance-aware loss functions improved the classification balance at the selected operating threshold rather than the overall ranking of cases.

Table 10.

Performance evaluation between the application of cross-entropy loss and imbalance-aware loss functions on classification results of our proposed model training on TCGA and testing on unseen GB-UK cohort.

Loss function AUC Specificity Sensitivity Balanced accuracy F1-score FP FN
aCross-entropy 0.787 0.660 0.900 0.780 0.655 17 2
Focal loss 0.787 0.720 0.900 0.810 0.692 14 2
Class-weighted cross-entropy 0.787 0.760 0.900 0.830 0.720 12 2
a

Our proposed model is trained using cross-entropy loss.

Performance analysis over a different feature encoder

To further investigate the proposed model's performance under a different feature encoder, we adopted Phikon-v2,47 a foundation model pre-trained on 450 million histology images. Following the same segmentation and tiling preprocessing pipeline, WSIs were divided into 224 × 224 patches, and Phikon-v2 was applied to extract 1024-dimensional feature embeddings. These features were then compressed to a 512-dimensional representation using the VAE, followed by DBGMM to derive the optimal prototypes for the TCGA cohort, resulting in seven prototypes. 5-fold cross-validation was performed using the same folds applied to the UNI feature embeddings, and classification results were compared against models using predefined prototype numbers. Table 11 presents the 5-fold cross-validation results on the TCGA cohort using the Phikon-v2 feature encoder. The proposed automatic prototype learning approach, which identified seven prototypes, outperformed both predefined prototype settings (8 and 16) across all four metrics, achieving the highest AUC (0.8294 ± 0.0205), specificity (0.7820 ± 0.0943), sensitivity (0.9478 ± 0.0514), and F1-score (0.9226 ± 0.0381). Performance also declined as the predefined prototype number increased from 8 to 16, suggesting that an excessive number of prototypes may introduce redundancy or noise into the classification process. These results are consistent with those obtained using the UNI encoder, further confirming that the proposed automatic prototype learning strategy generalizes effectively across different feature encoders.

Table 11.

5-fold cross-validation classification results of TCGA cohort (using Phikon-v2 feature encoder) based on clustering methods and given prototype numbers.

Clustering Prototype AUC (Mean ± SD) Specificity (Mean ± SD) Sensitivity (Mean ± SD) F1-score (Mean ± SD)
Proposed automatic prototype learning 7 0.8294 ± 0.0205 0.7820 ± 0.0943 0.9478 ± 0.0514 0.9226 ± 0.0381
Predefined number of prototypes 8 0.7703 ± 0.0383 0.7153 ± 0.1304 0.7352 ± 0.3310 0.7410 ± 0.2748
16 0.7355 ± 0.0161 0.5183 ± 0.3007 0.6261 ± 0.2406 0.6481 ± 0.1930

In Table 11, boldface highlights the performance of the proposed approach.

Discussion

EGFR amplification plays a pivotal role in the diagnosis of GBM and in therapeutic decision-making.

Current molecular diagnostic standards are costly and resource-intensive and are not accessible in all clinical environments. These challenges delay molecular testing, which hinders patients' treatment timelines. Therefore, there is a crucial need to develop a pathology-based computational model that provides a fast, precise, and widely accessible approach for predicting EGFR amplification from WSIs. In this study, we demonstrated that an automatic interpretable prototype-learning approach predicts GBM-EGFR amplification with high performance across internal and external validations. Although the external cohorts evaluated here are relatively limited in size. It also demonstrated a robust ability to link the prototypes to neuropathologists' recognizable tissue patterns.

A main challenge with prediction models using WSIs is to capture all morphological features of the tissue in a way that supports precise classification and interpretability. In this regard, MIL methods have achieved outstanding performance in classifying WSIs, particularly for predicting EGFR amplification. However, they rely only on slide-level labels and attention-based selected tissue patches, failing to capture global morphological tissue structures.

On the other hand, conventional unsupervised models rely on predefined, arbitrary cluster numbers, which can lead to over- or under-clustering. Over-clustering results from over-fragmenting the tissue's morphological features, whereas under-clustering results from restricted merging of diagnostically distinct tissue morphologies.

To address the challenges, we proposed an automatic prototype-learning framework for GBM-EGFR prediction that learns data-driven tissue morphological prototypes without requiring a predefined number of clusters or manual annotations.

The superior performance of the proposed model stems from its two major components. First, the application of VAE reduces feature sparsity while preserving informative morphological structure. Classification results using the 1024-dimensional features (without applying VAE) yielded a moderate AUC of 0.7972 ± 0.0194, but with lower specificity and sensitivity than those achieved by the model using a 512-dimensional embedding. Hence, it indicates that high-dimensional features increase sparsity. Furthermore, as the pairwise distances between patches become more concentrated with increasing dimensionality, clustering methods become less effective at distinguishing different tissue morphologies.48 Conversely, over-compressing the data to 256 and 128 dimensions resulted in a notably different and more severe failure mode. Instead of a moderate, consistent decline, the classification performance specifically deteriorated in discriminative ability, with AUC scores of 0.6426 and 0.5968, respectively. Therefore, results demonstrate that 512 dimensions represents the optimal trade-off between preserving discriminative morphological information and mitigating the instability associated with high-dimensional clustering, achieving the best classification performance at moderate computational cost.

Moreover, the DBGMM enables the model to more effectively explore the underlying data structure and determine the optimal clusters without relying on arbitrary, predefined cluster numbers.

Across internal and external validations, the proposed approach presented a more balanced sensitivity–specificity than the predefined clustering approach. As a result, it suggests that the automatic determination of morphological prototypes overcomes the over- and under-clustering effects that reduce discrimination between EGFR-amplified and EGFR-non-amplified cases.

Benchmarks against MIL models indicated that the proposed framework achieved higher classification performance. Although MIL models achieved competitive performance, they consistently showed imbalanced sensitivity and specificity on unseen validation cohorts, due to their attention-based aggregation strategy, which can lead to focus on non-discriminative tissue patterns, hindering their ability to overcome noisy instances and generalize across various cohorts.28, 29, 49, 50

A key finding of our model's quantitative prototype analysis based on Cohen's d revealed that prototype C4 exhibited the strongest ability to discriminate EGFR-amplified from EGFR-non-amplified cases, followed by C5, C2, and C3 prototypes. Neuropathology review indicated that C3 and C4 prototypes were characterized by low-cell-density regions with either minimally infiltrated or subtly abnormal brain tissue. As a result, it suggests that the learned prototypes were not only effective in the prediction model but also could present morphologically interpretable patterns. An experimental study using genetically engineered mouse models4 has shown that EGFR amplification or mutant EGFR signaling causes diffuse parenchymal invasion, tumor cell dispersal along neural structures, and invasion at tumor margins rather than bulk hypercellularity. Furthermore, a similar study51 on human GBM demonstrated that wild-type EGFR amplification enhances tumor growth without the need for angiogenic remodeling. These growth patterns are compatible with regions of deceptively preserved or only inconspicuously abnormal parenchyma surrounding invasive edges. A xenograft modeling study also demonstrated that patient-derived GBM cells with EGFR amplification transplanted into the brain and flanks of nude mice led to formation of invasive tumors that retained EGFR amplification52; in GBM slice models, inhibition of EGFR by gefitinib led to a reduction in GBM cell invasion for EGFR amplified tumors.53

These prior studies, showing EGFR-associated diffuse invasion, suggest the biological patterns that are highlighted by our infiltration-associated morphology findings. Therefore, the convergence of known tumor biology, neuropathological interpretation, quantification of the prototypes' effect size, and the automatic prototype learning provide a convincing explanation for why these prototypes play a significant role in the model's decision-making for EGFR-amplified cases.

From the pathology informatics standpoint, the results suggest that the prototype learning's contribution extends beyond solely improving the efficiency of the prediction model. As the framework represents WSIs as semantically meaningful tissue prototypes, it provides a more transparent and interpretable basis for analysis than relying only on a few selected patches. Therefore, this interpretable concept is particularly relevant to the digital pathology workflow, where molecular testing is crucial yet constrained by high cost, limited accessibility, and time-consuming nature. A morphology-driven AI prediction framework capable of identifying optimal prototypes of H&E-stained WSIs could be a strong assistive tool for molecular testing by facilitating case prioritization and providing rapid, early decision-support, especially in resource-constrained settings. Nevertheless, this approach should be considered complementary to molecular testing rather than a diagnostic replacement.

Limitations

Despite the promising outcomes of the proposed framework, it is important to acknowledge the study's limitations to inform future development. This study focused solely on GBM-EGFR amplification, and the generalizability of the framework to EGFR mutations (e.g., EGFRvIII) and broader GBM biomarkers (e.g., IDH, MGMT, and TERT54) remains to be evaluated.

Although the same preprocessing methods, including tissue segmentation, WSI tiling, and feature extraction, were applied across all three cohorts (i.e., TCGA, CPTAC, and GB-UK datasets), variations in fixation, staining intensity, tissue quality, and scanners may have introduced cohort-specific domain shifts. Hence, the use of an identical preprocessing pipeline reduces procedural variability; however, it cannot overcome preanalytical or scanner-related differences. Moreover, the retrospective inclusion of cases with available molecular annotations and WSIs of adequate quality may introduce selection bias and may not fully reflect the broader population encountered in routine clinical practice. Demographic and clinical factors, such as age, sex, ethnicity, institution, treatment history, and tumor sampling characteristics, were not formally adjusted. Therefore, it may represent residual confounders. In addition, the different class distributions and the relatively small size of external cohorts may affect the stability of sensitivity and specificity predictions.

Because EGFR-amplified cases dominated TCGA training cohort, it may lead the model to learn toward positive predictions when testing on the external GB-UK cohort, which includes the opposite class distribution. Consequently being biased toward positive cases. The model predicted 18 out of 20 EGFR-amplified cases in the GB-UK dataset, achieving high sensitivity (0.90), however, it misclassified 17 out of 50 EGFR non-amplified cases, resulting in lower specificity (0.66). Whereas sensitivity and specificity were calculated based on their respective classes the FP burden in the GB-UK cohort, which has predominantly EGFR-amplified cases, reduced the prevalence-dependent metrics such as F1-score. The imbalance-aware loss functions (focal loss and class-weighted cross-entropy) partially mitigated this effect, reducing FPs and increasing balanced accuracy from 0.78 to 0.81 and 0.83, while maintaining sensitivity at 0.90. Nonetheless, the observed performance might be influenced by differences in staining, tissue preprocessing, scanner settings, and other biological factors. Therefore, validation on broader, multi-institutional external cohorts is required to evaluate generalizability.

Finally, these findings suggest that the proposed framework has potential to yield interpretable and generalizable GBM-EGFR predictions, though broader validation is required to establish its clinical robustness. Hence, by effectively capturing tissue morphology, this approach can enhance digital pathology workflows and assist clinical decision-making.

CRediT authorship contribution statement

H.D.M: conceived the idea of this work, performed data curation, developed the methodology and software, provided resources, created visualizations, conducted formal analysis, supervised the project, and wrote the original draft, including review and editing. I·N: contributed to the review and editing, provided clinical investigation and contributed to the collection of the UK-Brain WSI Data. C.C: contributed to the review, editing and Investigation. M.F: contributed to the review and clinical investigation. A.D·I: Contributed to the review, editing, and provided clinical investigation. S.L: provided supervision, resources, review, and editing and contributed to the investigation. All authors reviewed and approved the final manuscript.

Ethics declaration

Images of the GB-UK dataset were obtained with Ethical approval from the UK Brain Archive Information Network (BRAIN UK) and the Research Ethics Committee South Central Hampshire B (REC reference number 14/SC/0098). Written informed consent was obtained from all patients in accordance with BRAIN UK protocols.8, 55

Code availability

The code used to develop and evaluate our model will be made publicly available on GitHub upon acceptance of the manuscript.

Declaration of generative AI and AI-assisted technologies in the writing process

During the preparation of this work, the authors used ChatGPT and Grammarly to improve the manuscript's readability. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

Funding

This research was supported by the Commonwealth through an Australian Government Research Training Program [https://doi.org/10.82133/C42F-K220] to H.D.M. The generation and curation of the GB-UK dataset were supported by funding to I·N from a National Institute of Health Research Clinical Lectureship, the University College London Biomedical Research Centre, and the Academy of Medical Sciences (grant number: SGL026\1035). NHMRC Ideas (ref no. GNT2020035) Computational Analysis and AI in Brain Tumor Imaging: Toward the Augmented Diagnostics of the Future to A.D, S.L and C·C.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this article.

Contributor Information

Homay Danaei Mehr, Email: homay.danaeimehr@hdr.mq.edu.au.

Imran Noorani, Email: imran.noorani@crick.ac.uk.

Cong Cong, Email: thomas.cong@mq.edu.au.

Mark Fabian, Email: mark.fabian@uhs.nhs.uk.

Antonio Di Ieva, Email: antonio.diieva@mq.edu.au.

Sidong Liu, Email: sidong.liu@mq.edu.au.

Data availability

The TCGA and CPTAC WSI datasets are publicly available through the National Cancer Institute Cancer Research Data Commons (CRDC) at https://www.cancer.gov/ccg/. The GB-UK dataset is not publicly accessible due to privacy and ethical restrictions; however, it may be made available upon reasonable request, subject to appropriate institutional review board approval.

References

  • 1.An Z., Aksoy O., Zheng T., et al. Epidermal growth factor receptor and EGFRvIII in glioblastoma: signaling pathways and targeted therapies. Oncogene. 2018;37(12):1561–1575. doi: 10.1038/s41388-017-0045-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Lassman A.B., Aldape K.D., Ansell P.J., et al. Epidermal growth factor receptor (EGFR) amplification rates observed in screening patients for randomized trials in glioblastoma. J Neuro-Oncol. 2019;144(1):205–210. doi: 10.1007/s11060-019-03222-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Weller M., Le Rhun E., Preusser M., et al. How we treat glioblastoma. ESMO Open. 2019;4 doi: 10.1136/esmoopen-2019-000520. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Noorani I., De La Rosa J., Choi Y., et al. PiggyBac mutagenesis and exome sequencing identify genetic driver landscapes and potential therapeutic targets of EGFR-mutant gliomas. Genome Biol. 2020;21(1) doi: 10.1186/s13059-020-02092-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jaunmuktane Z., Capper D., Jones D.T.W., et al. Methylation array profiling of adult brain tumours: diagnostic outcomes in a large, single centre. Acta Neuropathol Commun. 2019;7(1) doi: 10.1186/s40478-019-0668-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Noorani I., Mischel P.S., Swanton C. Leveraging extrachromosomal DNA to fine-tune trials of targeted therapy for glioblastoma: opportunities and challenges. Nat Rev Clin Oncol. 2022;19(11):733–743. doi: 10.1038/s41571-022-00679-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Fontanilles M., Marguet F., Ruminy P., et al. Simultaneous detection of EGFR amplification and EGFRvIII variant using digital PCR-based method in glioblastoma. Acta Neuropathol Commun. 2020;8(1):52. doi: 10.1186/s40478-020-00917-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Noorani I., Haughey M., Luebeck J., et al. Extrachromosomal DNA-driven oncogene spatial heterogeneity and evolution in glioblastoma. Cancer Discov. 2025;15(10):2078–2095. doi: 10.1158/2159-8290.CD-24-1555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Choi B.D., Gerstner E.R., Frigault M.J., et al. Intraventricular CARv3-TEAM-E T cells in recurrent glioblastoma. N Engl J Med. 2024;390(14):1290–1298. doi: 10.1056/NEJMoa2314390. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Jin B., Xu X. Predicting early indica rice prices utilizing Bayesian-optimized machine learning models employing Gaussian process regressions. Food Human. 2026;7 doi: 10.1016/j.foohum.2026.101294. [DOI] [Google Scholar]
  • 11.Jin B., Xu X. Machine learning-based forecasts of residential property prices in Hangzhou City, Zhejiang Province, China. Neural Comput Appl. 2025;37(6):4971–4988. doi: 10.1007/s00521-024-10726-w. [DOI] [Google Scholar]
  • 12.Jin B., Xu X. Chinese energy security index price forecasting through the neural network. Innov Emerg Technol. 2025;12 doi: 10.1142/S2737599425500367. [DOI] [Google Scholar]
  • 13.Xu X., Zhang Y. Individual time series and composite forecasting of the Chinese stock index. Mach Learn Appl. 2021;5 doi: 10.1016/j.mlwa.2021.100035. [DOI] [Google Scholar]
  • 14.Jin B., Xu X. Inter-city housing price causality in Fujian: contemporaneous analysis via vector error-correction and directed acyclic graph models. Qual Quant. 2026 doi: 10.1007/s11135-026-02740-y. Published online April. [DOI] [Google Scholar]
  • 15.Akan T., Alp S., Nobel Bhuiyan M.A. 2023 International Conference on Computational Science and Computational Intelligence (CSCI) IEEE; 2023. ECGformer: leveraging Transformer for ECG heartbeat arrhythmia classification; pp. 1412–1417. [DOI] [Google Scholar]
  • 16.Sahu A., Das P.K., Meher S. An efficient deep learning scheme to detect breast cancer using mammogram and ultrasound breast images. Biomed Signal Process Control. 2024;87 doi: 10.1016/j.bspc.2023.105377. [DOI] [Google Scholar]
  • 17.Xu X., Zhang Y. Steel price index forecasting through neural networks: the composite index, long products, flat products, and rolled products. Miner Econ. 2023;36(4):563–582. doi: 10.1007/s13563-022-00357-9. [DOI] [Google Scholar]
  • 18.Jin B., Xu X. Contemporaneous causal analysis of housing prices across Guangdong’s major cities: employing vector error-correction modeling and directed acyclic graphs. J Uncertain Syst. 2026 doi: 10.1142/S1752890926500042. Published online February 11. [DOI] [Google Scholar]
  • 19.Zheng Y., Wu K., Li J., et al. Partial-label contrastive representation learning for fine-grained biomarkers prediction from histopathology whole slide images. IEEE J Biomed Health Inform. 2025;29(1):396–408. doi: 10.1109/JBHI.2024.3429188. [DOI] [PubMed] [Google Scholar]
  • 20.Zhao Y., Wang W., Ji Y., et al. Computational pathology for prediction of isocitrate dehydrogenase gene mutation from whole slide images in adult patients with diffuse glioma. Am J Pathol. 2024;194(5):747–758. doi: 10.1016/j.ajpath.2024.01.009. [DOI] [PubMed] [Google Scholar]
  • 21.Zuo M., Xing X., Zheng L., et al. Weakly supervised deep learning-based classification for histopathology of gliomas: a single center experience. Sci Rep. 2025;15(1):265. doi: 10.1038/s41598-024-84238-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Innani S., Baheti B., Nasrallah M.P., et al. 2024 IEEE International Symposium on Biomedical Imaging (ISBI) IEEE; 2024. Weakly supervised IDH-status glioma classification from H&E-stained whole slide images; pp. 1–5. [DOI] [Google Scholar]
  • 23.Danaei Mehr H., Noorani I., Rana P., et al. In: Artificial Intelligence in Medicine: 22nd International Conference, AIME 2024, Proceedings, Part II. Finkelstein J., Moskovitch R., Parimbelli E., editors. Vol. 14845. Springer; Cham: 2024. AI in neuro-oncology: Predicting EGFR amplification in glioblastoma from whole slide images using weakly supervised deep learning; pp. 21–29. (Lecture Notes in Computer Science). [DOI] [Google Scholar]
  • 24.Nasrallah M.P., Zhao J., Tsai C.C., et al. Machine learning for cryosection pathology predicts the 2021 WHO classification of glioma. Med (Baltim) 2023;4(8):526–540.e4. doi: 10.1016/j.medj.2023.06.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kingma D.P., Welling M. Proceedings of the 2nd international conference on learning representations (ICLR 2014) Banff, AB; Canada: 2014. Auto-Encoding Variational Bayes. [Google Scholar]
  • 26.Yi D.H., Kim D.W., Park C.S. Prior selection method using likelihood confidence region and Dirichlet process Gaussian mixture model for Bayesian inference of building energy models. Energ Build. 2020;224 doi: 10.1016/j.enbuild.2020.110293. [DOI] [Google Scholar]
  • 27.Danaei Mehr H., Cong C., Noorani I., et al. In: Artificial Intelligence in Medicine: 23rd International Conference, AIME 2025, Proceedings, Part I. Bellazzi R., Juarez Herrero J.M., Sacchi L., Zupan B., editors. Vol. 15734. Springer; Cham: 2025. Adaptive clustering for EGFR amplification prediction in glioblastoma: a variational autoencoder-Dirichlet Bayesian Gaussian approach; pp. 88–97. (Lecture Notes in Artificial Intelligence). [DOI] [Google Scholar]
  • 28.Lopez-Gines C., Gil-Benso R., Ferrer-Luna R., et al. New pattern of EGFR amplification in glioblastoma and the relationship of gene copy number with gene expression profile. Mod Pathol. 2010;23(6):856–865. doi: 10.1038/modpathol.2010.62. [DOI] [PubMed] [Google Scholar]
  • 29.Lu M.Y., Williamson D.F.K., Chen T.Y., et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat Biomed Eng. 2021;5(6):555–570. doi: 10.1038/s41551-020-00682-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Otsu N. A threshold selection method from gray-level histograms. IEEE Trans Syst Man Cybern. 1979;9(1):62–66. doi: 10.1109/TSMC.1979.4310076. [DOI] [Google Scholar]
  • 31.Chen R.J., Ding T., Lu M.Y., et al. Towards a general-purpose foundation model for computational pathology. Nat Med. 2024;30(3):850–862. doi: 10.1038/s41591-024-02857-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zhou S., Xu H., Zheng Z., et al. A comprehensive survey on deep clustering: taxonomy, challenges, and future directions. ACM Comput Surv. 2025;57(3):1–38. doi: 10.1145/3689036. [DOI] [Google Scholar]
  • 33.Mukhdoomi M.A., Chachoo M.A. 2024 3rd Edition of IEEE Delhi Section Flagship Conference (DELCON) IEEE; 2024. Cancer subtyping through multi-omics feature selection: a review of methods and the superiority of variational autoencoders; pp. 1–7. [DOI] [Google Scholar]
  • 34.Sadat S., Buhmann J., Bradley D., et al. NIPS ‘24: Proceedings of the 38th International Conference on Neural Information Processing System. 2024. LiteVAE: lightweight and efficient variational autoencoders for latent diffusion models; pp. 3907–3936. 37. [Google Scholar]
  • 35.Bing Z., Yun Y., Huang K., et al. Context-based meta-reinforcement learning with Bayesian nonparametric models. IEEE Trans Pattern Anal Mach Intell. 2024;46(10):6948–6965. doi: 10.1109/TPAMI.2024.3386780. [DOI] [PubMed] [Google Scholar]
  • 36.Görür D., Rasmussen C.E. Dirichlet process Gaussian mixture models: choice of the base distribution. J Comput Sci Technol. 2010;25(4):653–664. doi: 10.1007/s11390-010-9355-8. [DOI] [Google Scholar]
  • 37.Deisenroth M.P., Ohlsson H. Proceedings of the 2011 American Control Conference. IEEE; 2011. A general perspective on Gaussian filtering and smoothing: explaining current and deriving new algorithms; pp. 1807–1812. [DOI] [Google Scholar]
  • 38.Song A.H., Chen R.J., Ding T., et al. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) IEEE; 2024. Morphological prototyping for unsupervised slide representation learning in computational pathology; pp. 11566–11578. [DOI] [Google Scholar]
  • 39.Zhao J., Zhao Z., Song X., et al. Multiple instance learning with hierarchical discrimination and smoothing attention for histopathological diagnosis. Appl Intell. 2025;55(6) doi: 10.1007/s10489-025-06300-z. [DOI] [Google Scholar]
  • 40.Shao Z., Bian H., Chen Y., et al. NIPS’21: Proceedings of the 35th International Conference on Neural Information Processing System. 2021. TransMIL: transformer based correlated multiple instance learning for whole slide image classification; pp. 2136–2147. 34. [Google Scholar]
  • 41.Zhang H., Meng Y., Zhao Y., et al. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) IEEE; 2022. DTFD-MIL: double-tier feature distillation multiple instance learning for histopathology whole slide image classification; pp. 18780–18790. [DOI] [Google Scholar]
  • 42.Zhang Y., Li H., Sun Y., et al. In: Computer Vision - ECCV 2024. Vedaldi A., Bischof H., Brox T., Frahm J.M., editors. Springer; Cham: 2024. Attention-challenging multiple instance learning for whole slide image classification; pp. 125–143. [DOI] [Google Scholar]
  • 43.He K., Zhang X., Ren S., et al. 2015 IEEE International Conference on Computer Vision (ICCV) IEEE; 2015. Delving deep into rectifiers: surpassing human-level performance on ImageNet classification; pp. 1026–1034. [DOI] [Google Scholar]
  • 44.Glorot X., Bengio Y. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings. Vol. 9. 2010. Understanding the difficulty of training deep feedforward neural networks; pp. 249–256. [Google Scholar]
  • 45.Hinton G.E., Salakhutdinov R.R. Reducing the dimensionality of data with neural networks. Science. 2006;313(5786):504–507. doi: 10.1126/science.1127647. [DOI] [PubMed] [Google Scholar]
  • 46.He K., Zhang X., Ren S., et al. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) IEEE; 2016. Deep residual learning for image recognition; pp. 770–778. [DOI] [Google Scholar]
  • 47.Filiot A., Jacob P., Mac Kain A., et al. Phikon-v2, a large and public feature extractor for biomarker prediction. arXiv [Preprint] 2024 doi: 10.48550/arXiv.2409.09173. arXiv:2409.09173. [DOI] [Google Scholar]
  • 48.Aggarwal C.C., Hinneburg A., Keim D.A. In: Database Theory - ICDT 2001. Van den Bussche J., Vianu V., editors. Vol. 1973. Springer; Berlin: 2001. On the surprising behavior of distance metrics in high dimensional space; pp. 420–434. (Lecture Notes in Computer Science). [DOI] [Google Scholar]
  • 49.Juyal D., Shingi S., Javed S.A., et al. 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) IEEE; 2024. SC-MIL: supervised contrastive multiple instance learning for imbalanced classification in pathology; pp. 7931–7940. [DOI] [Google Scholar]
  • 50.Ilse M., Tomczak J.M., Welling M. Proceedings of the 35th International Conference on Machine Learning. 2018. Attention-based deep multiple instance learning; pp. 2127–2136. [Google Scholar]
  • 51.Talasila K.M., Soentgerath A., Euskirchen P., et al. EGFR wild-type amplification and activation promote invasion and development of glioblastoma independent of angiogenesis. Acta Neuropathol. 2013;125(5):683–698. doi: 10.1007/s00401-013-1101-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Parker J.J., Dionne K.R., Massarwa R., et al. Gefitinib selectively inhibits tumor cell migration in EGFR-amplified human glioblastoma. Neuro-Oncology. 2013;15(8):1048–1057. doi: 10.1093/neuonc/not053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Giannini C., Sarkaria J.N., Saito A., et al. Patient tumor EGFR and PDGFRA gene amplifications retained in an invasive intracranial xenograft model of glioblastoma multiforme. Neuro-Oncology. 2005;7(2):164–176. doi: 10.1215/S1152851704000821. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Eckel-Passow J.E., Lachance D.H., Molinaro A.M., et al. Glioma groups based on 1p/19q, IDH, and TERT promoter mutations in tumors. N Engl J Med. 2015;372(26):2499–2508. doi: 10.1056/NEJMoa1407279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Noorani I., Sidlauskas K., Pellow S., et al. Clinical impact of anti-inflammatory microglia and macrophage phenotypes at glioblastoma margins. Brain Commun. 2023;5(3) doi: 10.1093/braincomms/fcad176. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The TCGA and CPTAC WSI datasets are publicly available through the National Cancer Institute Cancer Research Data Commons (CRDC) at https://www.cancer.gov/ccg/. The GB-UK dataset is not publicly accessible due to privacy and ethical restrictions; however, it may be made available upon reasonable request, subject to appropriate institutional review board approval.


Articles from Journal of Pathology Informatics are provided here courtesy of Elsevier

RESOURCES