Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 30;16:24854. doi: 10.1038/s41598-026-53395-6

Comparative evaluation of CNN models for nasopharyngeal carcinoma classification on pathology data

Muhammad Kabir Abdullahi 1, Sarina Mansor 1,✉, Wan Siti Halimatul Munirah Wan Ahmad 3, Md Serajun Nabi 1, Mohammad Faizal Ahmad Fauzi 2, Arbab Sufyan Wadood 1, Adam Malik Ismail 4
PMCID: PMC13458487  PMID: 42225742

Abstract

This study presents a systematic evaluation of deep learning models for the classification of nasopharyngeal carcinoma (NPC) using whole slide images (WSIs) obtained from Sarawak General Hospital (SGH) and Hospital Kuala Lumpur (HKL). NPC, a malignancy with high prevalence in Southeast Asia, presents diagnostic challenges due to the histological similarities between normal and pathological tissues. A dataset of 88,002 images, annotated by expert pathologists and categorized into four classes: normal, lymphoid hyperplasia (LHP), nasopharyngeal inflammation (NPI), and NPC, was utilized for model training and evaluation. Several convolutional neural network (CNN) architectures, including DenseNet201, MobileNet, EfficientNetB0, InceptionNet, XceptionNet, VGG16, and NASNetMobile were systematically assessed alongside hybrid architectures formed through intermediate-level feature fusion of top-performing backbones. All models were evaluated using accuracy, precision, F1-score, and training time to ensure a balanced assessment of predictive performance and computational efficiency. Among individual models, MobileNet achieved the highest accuracy (96.9%), while DenseNet201 demonstrated the most balanced classification performance with the highest F1-score (94.9%). The hybrid EfficientNetB0 + DenseNet201 model achieved the overall best accuracy (97.6%), indicating that combining complementary feature representations can further enhance predictive capability. The integration of data augmentation and class weighting effectively mitigated dataset imbalance, resulting in substantial improvements in generalization and minority class recognition. Overall, the findings highlight the strong potential of optimized CNN architectures and feature-level fusion strategies for robust multi-class NPC classification, supporting their applicability in computer-aided diagnosis and assisting pathologists in improving diagnostic accuracy.

Keywords: Nasopharyngeal carcinoma NPC, Lymphoid hyperplasia, LHP, Nasopharyngeal inflammation, NPI, Convolutional neural networks

Subject terms: Cancer, Computational biology and bioinformatics, Medical research

Introduction

In recent years, deep learning has revolutionized the field of computer vision, enabling significant advancements in image classification, object detection, and medical image analysis. Convolutional Neural Networks (CNN’s)1 have shown remarkable success in identifying intricate patterns in medical images, leading to breakthroughs in automated diagnostics. These models have been applied in various medical fields, including oncology, dermatology, and radiology, offering the potential to improve diagnostic accuracy and reduce the workload for medical professionals.

In the context of cancer diagnosis, especially for rare and geographically specific diseases like Nasopharyngeal Carcinoma (NPC)2,3, deep learning models hold promise for automating early detection. NPC is a malignant tumor that occurs in the nasopharynx, and it poses a significant health challenge, particularly in Southeast Asia, where the incidence rates are among the highest in the world4–6. Despite its prevalence in certain regions, NPC often goes undiagnosed until the advanced stage due to its non-specific symptoms and the complexity of interpreting histopathological images.

The primary diagnostic challenge in classifying NPC via histological images lies in the high degree of inter-observer subjectivity and the tumor’s “masked” morphological features. Under the World Health Organization (WHO) classification system, distinguishing between Type 2 (Non-keratinizing differentiated) and Type 3 (Undifferentiated) NPC is notoriously difficult due to overlapping cellular characteristics, leading to low inter-pathologist reproducibility7. This difficulty is compounded by the “lymphoepithelial” nature of the disease; malignant epithelial cells are often obscured by a dense infiltrate of non-cancerous lymphocytes, which can “hide” the tumor or cause it to be misidentified as lymphoma8. For deep learning models, these biological complexities translate into significant technical barriers, specifically “crush artifacts” during biopsy and the subtle textures of poorly differentiated squamous cells that are difficult for standard Convolutional Neural Networks to distinguish from surrounding inflamed stroma9. Consequently, while histopathology remains the gold standard, these intrinsic visual ambiguities necessitate the development of tailored and more complexed digital pathology models to ensure optimal diagnostic performance.

Given the urgency of improving NPC diagnosis, various deep learning architectures have been explored for image-based classification tasks. CNN architectures such as DenseNet20110, MobileNet11, EfficientNetBO12 and NASNetMobile13 have demonstrated their effectiveness in medical image classification. These models leverage large datasets and hierarchical learning to autonomously identify patterns, thus providing a potential solution for early and accurate NPC detection. Deep learning models outperform traditional markers in detecting Nasopharyngeal carcinoma recurrence via Tumor Infiltrating Lymphocytes (TIL’s) analysis14.

Although CNNs have demonstrated remarkable success in medical imaging tasks, including cancer detection, segmentation, and grading. However, most existing studies focus on binary classification (e.g., NPC vs. normal), overlooking the need for a multi-class approach that distinguishes NPC from other similar pathological conditions.

In this study, the performance of seven state-of-the-art CNN models was explored, DenseNet201, MobileNet, EfficientNetB0, InceptionV315, Xception16, VGG1617, and NASNetMobile on datasets of Nasopharyngeal images, as shown in Fig. 1. The motivation behind choosing these models stems from their established success in large-scale image classification tasks and their ability to process complex visual information with varying computational demands. The application of a range of pre-processing and augmentation techniques ensures robust training, including resizing the images to 224 × 224 pixels and using augmentation strategies like flipping and rotation to prevent overfitting.

Fig. 1.

Fig. 1

Whole-slide image sample.

This work aims to assess and compare the performance of these architectures in classifying NPC from medical images, ultimately aiming to determine the most effective deep learning algorithms for automating NPC detection and improving diagnostic accuracy.

While previous studies have investigated CNN-based classification of NPC, most focus primarily on binary classification (i.e., NPC vs. normal tissue) without considering other pathological states such as NPI and LHP. Furthermore, existing works often suffer from dataset limitations, where models trained on a single institution’s data fail to generalize well across different imaging environments. We can also notice that most studies that propose architectural fusion methods often do not expand their fusion techniques across different datasets to effectively asses it. Additionally, class imbalance remains a significant issue, as datasets typically contain far fewer samples of LHP and NPI compared to NPC, leading to biased model performance.

To address these challenges, this study makes the following contributions:

  • Comprehensive multi-class classification: Unlike prior works that focus on binary classification, our study evaluates multiple CNN architectures for distinguishing between four classes: Normal, LHP, NPI, and NPC to improve the robustness of automated NPC diagnosis.

  • Evaluating the generalization ability of state-of-the-art CNN models across multi-source datasets, done through the performance comparison between the pre-trained CNN models across datasets belonging to two different institutions.

  • Exploring hybrid models for improved accuracy, where best performing models were combined to form different combinations through a fusion technique, which proved helpful and slightly improving the accuracy.

Related work

Recent advancements in deep learning have significantly improved the automated classification of nasopharyngeal carcinoma (NPC) using histopathological and radiological imaging. Convolutional Neural Networks (CNNs) have been widely applied in tumor detection and classification, outperforming traditional machine learning approaches. This section reviews seven related studies that have explored deep learning-based NPC classification, highlighting their methodologies, findings, and contributions to the field.

CNNs, such as DenseNet, and MobileNet have achieved success in medical imaging by leveraging hierarchical feature extraction. EfficientNet optimizes performance through compound scaling, while Xception uses depth-wise separable convolutions for computational efficiency. Recent applications in pathology include cancer detection and subtyping5,6. Despite these advances, issues such as imbalanced datasets and generalizability across imaging conditions persist. This study builds on these works by evaluating CNN architectures in a clinically relevant NPC classification task, incorporating Area Under the Receiver Operating Characteristic curve (AUROC) and Weighted F1 metrics for better-imbalanced class evaluation.

18 explored the effectiveness of deep learning in NPC classification using biopsy images. Their study evaluated various CNN architectures, including ResNet and DenseNet, to determine the most effective model for differentiating between NPC and benign conditions. The methodology involved fine-tuning pre-trained CNN models using a dataset of high-resolution biopsy images. The findings indicated that deeper networks, such as DenseNet, performed better than shallow architectures due to their superior feature extraction capabilities. The study concluded that CNNs offer a promising solution for NPC classification, though the need for more diverse and larger datasets was highlighted as a limitation.

19 conducted research on the classification of NPC using deep learning applied to MRI scans. Their study compared CNN models trained on MRI images of patients with NPC and benign nasopharyngeal conditions. The methodology involved transfer learning, where pre-trained CNN models were fine-tuned on an specific datasets. The results demonstrated that CNNs could distinguish between NPC and benign lesions with high accuracy. The authors highlighted that integrating radiomics and deep learning could further enhance NPC diagnosis, but challenges related to inter-hospital data variability persisted.

20 proposed a learning-based segmentation method for NPC tumors in CT images. Their study focused on automating tumor segmentation to assist in treatment planning. The methodology included using a modified CNN model trained on CT scans labeled by expert radiologists. The model demonstrated high segmentation accuracy and was able to delineate tumor boundaries effectively. The study concluded that deep learning-based segmentation could significantly improve radiotherapy planning by providing precise tumor localization, although the authors emphasized the need for multi-center validation to ensure the model’s robustness.

21 investigated the application of deep learning models in histopathology images for tumor classification, with a focus on breast cancer but extending to nasopharyngeal carcinoma. Their study employed a CNN-based feature extraction model and compared different architectures such as ResNet, DenseNet, and EfficientNet for classifying pathological images. The results demonstrated that EfficientNet performed best, achieving an AUROC of 0.998, followed by DenseNet, while ResNet exhibited weaker generalization across datasets. They established that pre-trained models with transfer learning outperform those trained from scratch, reinforcing the importance of leveraging pre-existing features for medical image analysis.

22 developed a multi-magnification similarity learning approach for histopathological image classification. Their method combined deep learning and self-supervised learning to improve feature extraction across multiple scales. The study tested its approach on NPC datasets, showing that using multi-resolution inputs enhanced classification accuracy compared to single-resolution models. The findings suggested that deep learning models should incorporate multi-scale representations to capture complex morphological patterns in histopathology images. Their study also highlighted the importance of augmenting deep learning with domain-specific knowledge, such as tumor grading and feature interpretation.

23 provided a comprehensive review of radiomics and deep learning applications in NPC detection, emphasizing how CNN-based architectures have evolved in tumor segmentation and classification. The review covered segmentation based CNNs, object detection models, and hybrid AI approaches, concluding that deep learning models integrating clinical data with imaging features enhance diagnostic accuracy. Their findings indicated that while CNNs have revolutionized medical image analysis, challenges such as interpretability and datasets imbalance remain critical issues to address in future research.

24 developed a 3D DenseNet model for automatic NPC detection using MRI scans, integrating deep learning into radiological imaging. Unlike histopathological studies, this research focused on detecting tumor regions in MRI scans, demonstrating that deep learning could segment and classify NPC with an accuracy of 97.4%. The model’s self-constrained learning mechanism helped improve feature representation and tumor localization. Their work provided insights into how deep learning could complement histopathological analysis by integrating multi-modal imaging for better NPC classification.

25 explored the use of deep learning for classifying skull-base invasion (SBI) in NPC patients, leveraging whole-body bone scintigraphy data. Their CNN-based model successfully differentiated between NPC cases with and without skull-base invasion, achieving an F1-score of 91.3%. The study provided valuable insights into how deep learning could enhance staging and prognosis estimation, beyond just classification. This research underscored the potential of AI in personalized medicine, demonstrating how deep learning could refine NPC treatment strategies by predicting tumor invasiveness.

26 proposed a weakly supervised framework, WS-T2T-ViT, for nasopharyngeal carcinoma (NPC) classification using only slide-level labels. The method leverages a multi-resolution pyramid, Tokens-to-Token Vision Transformer (T2T-ViT), and a multi-scale attention module to capture both local and global features across different magnifications, mimicking the coarse-to-fine pathological analysis. WS-T2T-ViT demonstrated high performance on an 802-patient NPC dataset (AUC = 0.989) and showed strong generalizability on the CAMELYON16 dataset, reducing the reliance on labor-intensive manual annotations. This research suggested that CNN-based models might be limited in capturing long-range dependencies, which transformers can address more efficiently6.

27 explored deep learning for skull-base invasion detection in NPC using CT imaging, integrating CNN-based classifiers with radiomics features. Their study demonstrated that radiomics-enhanced deep learning models outperformed standalone CNNs, achieving an AUROC of 0.985. The findings emphasized that deep learning should incorporate clinical metadata and radiomics features to improve robustness. This research was crucial in showing how multi-modal AI approaches can enhance NPC diagnostics beyond histopathology alone.

28 Proposed CONV-FAN-POX, a deep learning architecture integrating a Fourier Analysis Network (FAN) as fully connected classifier layer and EfficientNetV2 feature extraction backbone to capture both spatial and frequency-domain features for medical image classification on the MSLD2.0 dataset in a six-class setting, achieving 98.81% accuracy and 98.56% F1-score demonstrating improved robustness and generalization across datasets alongside reduced parameter count.

29 Developed a deep learning model for mammogram classification based on BI-RADS categories using a private dataset covering all categories. The architecture combines a ConvNeXt backbone for low-level feature extraction with a randomly initialized naive inception module for higher-level feature learning. The model also incorporates ROI-based classification to better capture breast boundaries. Performance was evaluated through three experimental test setups. The achieved classification accuracy was between 82.08% and 86.27% across the experiments, demonstrating effective BI-RADS categorization from mammograms.

30 The study proposes a hybrid prognosis framework for cervical cancer using Hematoxylin and Eosin (H&E) whole slide images to predict metastasis and recurrence risk. The approach combines transfer learning-based deep feature extraction using Xception with a Random Forest classifier trained on patch-level representations, followed by majority voting to obtain patient-level predictions from whole-slide images. Experimental results demonstrate that the method achieves an AUC of 0.82, indicating that CNN-extracted histopathological features can effectively support non-invasive risk stratification and assist in guiding post-surgical adjuvant therapy decisions.

30 review discusses the growing use of bioprinting technologies for constructing in vitro tumor models of head and neck squamous cell carcinoma (HNSCC), a malignancy with increasing global incidence and significant clinical impact. It highlights the limitations of traditional tumor models in accurately replicating tumor microenvironment complexity and treatment response, as well as challenges in obtaining sufficient clinical specimens. Bioprinting enables precise spatial organization of cells and tissue-like structures, offering more physiologically relevant tumor models. The study summarizes recent advances in bioprinting techniques for HNSCC modeling and their applications in drug screening and personalized therapy, emphasizing their potential to improve experimental tumor research and support individualized treatment strategies.

31 study proposes a deep learning framework for predicting microsatellite instability (MSI) and tumor mutational burden (TMB) directly from H&E-stained whole-slide images in gastric and colorectal cancers, aiming to reduce reliance on costly sequencing-based assays. The method integrates image-level features extracted using CLAM with nuclear segmentation features obtained via HoVer-Net, and fuses them using Multimodal Compact Bilinear Pooling across multiple deep learning classifiers. Experimental results show that incorporating nuclear morphological information improves predictive performance, yielding gains of 1–3% in AUC and 5–11% in recall under cross-validation. External validation further demonstrates robustness, achieving an AUC of 0.81 for MSI prediction in colorectal cancer. The findings indicate that combining histopathological image features with cellular-level segmentation information enhances biomarker prediction accuracy and supports more cost-effective, image-based cancer stratification.

32 proposes a hybrid deep learning framework combining ResNet-18 and EfficientNet-B0 for automated HER2 immunohistochemistry (IHC) scoring in breast cancer. The model integrates ResNet-18’s spatial feature extraction with EfficientNet-B0’s multi-scale texture representation to improve classification of subtle HER2 expression levels (0, 1+, 2+, 3+), particularly borderline cases. Evaluated on the HER2-IHC-40x dataset, the method achieves 94% accuracy and a macro F1-score of 0.93, with a ROC-AUC of 0.994, while maintaining computational efficiency through frozen pretrained backbones. Interpretability is enhanced using Grad-CAM + + to highlight class-discriminative regions, and performance is validated through ablation studies against multiple baseline architectures. Overall, the results demonstrate improved consistency and reliability in automated HER2 scoring by effectively combining complementary feature representations from both architectures.

33 proposes an interpretable deep learning framework for automated HER2 immunohistochemistry (IHC) scoring in breast cancer, addressing challenges arising from subjective manual assessment, particularly in borderline categories (1 + and 2+). The model integrates EfficientNetB3 with a Convolutional Block Attention Module (CBAM) to enhance feature representation and emphasize diagnostically relevant regions. Interpretability is supported using Grad-CAM for visual explanation of class-discriminative areas. Evaluated on a HER2-IHC-40xWSI dataset comprising 10,997 patches across four classes (0, 1+, 2+, 3+), the method achieves 96% accuracy and a macro F1-score of 0.93, demonstrating strong performance with particular effectiveness in borderline cases. Overall, the results indicate that combining attention mechanisms with explainable AI improves both classification reliability and interpretability in automated HER2 scoring.

34 Study proposes a lightweight, real-time framework for progesterone receptor immunohistochemistry (PR-IHC) analysis that performs nucleus-level classification and automated Allred scoring for breast cancer assessment. The method uses auto-calibrated DAB intensity thresholds to categorize nuclei into Strong, Moderate, Weak, and Negative classes based on staining intensity, followed by aggregation of positive nuclei proportions to compute region-of-interest (ROI)–level Allred scores. Evaluated on a clinically validated dataset from the University of Malaya Medical Centre, the system achieves strong performance (macro-F1 = 0.948, IoU = 0.906) while maintaining low computational complexity.

35 study presents an automated instance segmentation framework for detecting progesterone receptor (PR)–expressing nuclei in high-resolution immunohistochemistry (IHC) breast cancer images, addressing challenges such as morphological variability, staining heterogeneity, and overlapping nuclei. The approach leverages the Cellpose deep learning model and introduces a novel ground truth dataset generated through a hybrid pipeline combining automated segmentation (StarDist), manual refinement, and multi-round pathological validation across 250 PR-IHC images. Experimental results demonstrate reliable segmentation performance, achieving an F1-score of 0.8535, precision of 0.8882, recall of 0.8215, and IoU of 0.7445 on the test set.

36 work investigates deep learning–based methods for automated detection and segmentation of breast cancer nuclei in digital pathology images, focusing on estrogen receptor (ER) assessment as a key prognostic indicator. A hybrid ensemble framework combining a customized U-Net and Cellpose was developed to leverage complementary strengths in handling irregular and overlapping nuclear morphologies across diverse tissue structures. Evaluated on a dataset of 220 regions of interest extracted from 44 whole slide images with pixel-level annotations, the model demonstrates improved segmentation performance and generalizability, achieving a mean F1-score of 0.7468 and mean IoU of 0.6018 on the test set.

37 developed deep learning models for automated melanoma detection from dermoscopic images to address limitations of conventional diagnostic methods, which are often subjective and time-consuming. Three architectures, CNN, DenseNet201, and U-Net were evaluated, incorporating attention mechanisms (SE and CBAM), data augmentation (Mixup), and explainable AI using Grad-CAM to enhance performance and interpretability. Among the models, U-Net achieved the highest classification accuracy (95%), followed by DenseNet201 (93%) and CNN (90%). The visual explanations generated by Grad-CAM improved model transparency and clinical interpretability.

38 proposes a deep learning framework for automated and interpretable HER2 IHC scoring in breast cancer to address the subjectivity and inconsistency of manual evaluation. The approach combines a custom CNN and a fine-tuned DenseNet121 model, supported by HSV-based preprocessing and expert-validated data selection. To enhance transparency, explainable AI techniques—Grad-CAM and SHAP—were integrated to provide pixel- and region-level visual explanations, particularly improving reliability in borderline HER2 cases (1 + and 2+). Both models achieved 93% accuracy with strong class-wise consistency, while the CNN showed better calibration through higher prediction confidence and lower training loss. Overall, the framework demonstrates robust predictive performance and improved interpretability, supporting trustworthy AI-assisted breast cancer diagnosis.

39 proposed KDH-Net, a hybrid deep learning framework for multiclass kidney disease characterization from CT images. The model integrates EfficientNetB0, ResNet50, and MobileNetV2 via feature-level fusion and employs a two-stage training strategy. Evaluated under a patient-level protocol to prevent data leakage, KDH-Net achieved 0.93 accuracy and a 0.91 macro F1-score, while calibration and Grad-CAM analyses confirmed reliable and interpretable predictions, supporting clinically trustworthy decision-making.

Several points can be derived from this review section:

  • Dominance of Pre-trained Architectures: Most studies18,20–22 rely heavily on established CNN.

  • backbones such as DenseNet, ResNet, and EfficientNet. These works consistently demonstrate that deeper.

  • architectures and transfer learning significantly outperform shallow models and those trained from scratch,

  • particularly in medical imaging tasks.Multi-Modal Exploration: Research has successfully moved beyond simple histopathology to include MRI18,23, CT scans19,26, and whole-body bone scintigraphy24, proving that deep learning can assist not only in tumor detection but also in staging and identifying invasion (e.g., skull-base invasion).

  • Multi-Modal Exploration: Research has extended beyond histopathological analysis to include MRI-based.

  • classification19,24, CT-based tumor segmentation20,27, and whole-body bone scintigraphy for skull-base invasion detection25. These studies highlight the growing role of deep learning in both diagnosis and staging of nasopharyngeal carcinoma.

  • Methodological Innovations: Recent works have introduced advanced techniques such as multi-magnification similarity learning22, weakly supervised transformer-based frameworks26, and Fourier Analysis Networks (FAN)28 to capture both spatial and frequency-domain representations, improving feature expressiveness and classification robustness.

  • Emergence of Hybrid Approaches: There is a clear trend toward hybrid and multi-modal AI systems. Studies combining deep learning with radiomics27 and feature-level fusion strategies32 demonstrate improved diagnostic accuracy and robustness compared to single-model approaches.

  • Predominance of Binary Classification: A significant portion of existing research18,19 focuses on binary classification (e.g., NPC vs. benign), with limited exploration of more clinically relevant multi-class scenarios involving conditions such as lymphoid hyperplasia (LHP) and nasopharyngeal inflammation (NPI).

  • Dataset Scarcity and Lack of Diversity: Many studies emphasize limitations in dataset size and diversity18,19, with models often failing to generalize across institutions due to variations in imaging protocols and acquisition conditions.Class Imbalance Issues: While some studies mention dataset imbalance22, few implement robust mechanical solutions like class weighting or specialized loss functions to protect the recognition of minority classes in a clinically relevant way.

  • Class Imbalance Issues: While dataset imbalance is acknowledged in several works22, only a few studies implement effective mitigation strategies such as class weighting or cost-sensitive learning to improve minority class detection.

  • Computational vs. Predictive Trade-offs: High-performance models such as 3D DenseNet24 and transformer-based approaches26 achieve strong accuracy but often at the expense of increased computational complexity, which may limit real-world deployment.

  • Limited Feature-Level Fusion: Although hybrid approaches exist26,32, the application of intermediate-level feature fusion between heterogeneous CNN architectures remains relatively underexplored, particularly for histopathological whole slide image (WSI) classification.

This study addresses these limitations by providing a systematic multi-class evaluation (four distinct classes) using a large-scale dataset of over 88,000 images from two major medical centers (SGH and HKL). By integrating class weighting and evaluating different multi-model combinations through intermediate-level feature fusion, this research moves beyond binary detection to offer a robust, computationally evaluated framework for assisting pathologists in complex diagnostic environments.

Methods

Dataset

The datasets used in this work includes 88,002 samples categorized into four classes: normal, lymphoid hyperplasia (LHP), nasopharyngeal inflammation (NPI), and nasopharyngeal carcinoma (NPC) as shown in Fig. 2. Data was collected in collaboration with Sarawak General Hospital (SGH) and Hospital Kuala Lumpur (HKL), with annotations made by three pathologists using the Cytomine platform.

Fig. 2.

Fig. 2

WSI patches from SGH and HKL hospitals.

An example of the original Whole-slide image samples collected from the hospitals is shown in Fig. 1.

The whole slide images (WSIs) from the two hospitals were scanned using different settings. At SGH, they were scanned with a 3DHistech Pannoramic SCAN 150 at 20× magnification, giving an image size of about 125,000 pixels wide and 295,000 pixels high.

At HKL, the WSIs were scanned with a 3DHistech Pannoramic MIDI at 40× magnification, producing larger images of about 250,000 pixels wide and 572,000 pixels high. An example of images extracted from both sources can be seen in Fig. 2 as well as examples from all four classes are shown in Fig. 3.

Fig. 3.

Fig. 3

Digital pathology samples from each of the four datasets classes.

The datasets is divided into 44,000 training samples, 22,000 for validation, and 22,000 for testing, as shown in Table 1.

Table 1.

Summary of Dataset.

Class Training Validation Test
NPC 34,093 17,047 17,046
NPI 5665 2832 2831
LHP 439 220 219
NORMAL 3805 1902 1903
Total 44,002 22,001 21,999
% Distribution 50% 25% 25%

A notable feature is the variation in imaging conditions: WSIs from SGH were scanned at 20x magnification (125,000 × 295,000 pixels), while HKL images were scanned at 40 × (250,000 × 572,000 pixels). This diversity in resolution and magnification adds complexity to model training and helps in ensuring performance generalization across different WSI collection methods conditions. .

Due to the severe imbalance in the datasets distribution, where the majority NPC class is extremely big compared to the rest of the classes, we opted for the use of data augmentation for the minority class (LHP), in which various techniques were used such as rotation (10° range), width and height shifting (0.1 range), zoom (0.1 range) and so on, as well as down-sampling the majority class (NPC) to level the distribution between the classes. Both of these steps were used only on the training set while keeping the testing data intact, thereby preserving the integrity of the model’s performance evaluation. The new dataset’s distribution is highlighted in Table 2.

Table 2.

Data distribution after pre-processing.

Class Training Validation Test
NPC 5000 17,047 17,046
NPI 5665 2832 2831
LHP 3000 220 219
NORMAL 3805 1902 1903
Total 17,470 22,001 21,999

Image pre-processing

To ensure consistent model performance, pre-processing steps were applied to the datasets. Each WSI underwent standard re-sizing to a uniform input size for compatibility with deep learning models. Given the differences in resolution between images from SGH and HKL, all images were re-sized to a fixed dimension of 224 × 224 pixels. The choice of this resolution is primarily dictated by the need for architectural consistency with the selected pre-trained backbones. Since models like EfficientNetB0, DenseNet201, and MobileNet were originally optimized on the ImageNet dataset using this standard dimension, maintaining this resolution preserves the integrity of their learned weights and spatial hierarchies without requiring aggressive interpolation. Additionally, this dimensionality offers a critical trade-off between clinical accuracy and resource management, providing sufficient morphological detail for histopathological analysis while significantly reducing the memory overhead and training time.

Initially, data augmentation techniques such as horizontal flips, random rotations, and shifts were employed to increase the diversity of the training set and helped the model generalize better. However, despite high training and validation accuracy, the testing accuracy remained unsatisfactory. As a result, data augmentation techniques were supported by applying class weight adjustments due to the class imbalance, particularly for the minority class LHP. In this experiment, class weighting was used to address class imbalance by assigning higher weights to underrepresented classes, which adjusted the model’s loss function to give errors from minority classes greater significance. The class weights were calculated using the formula (1):

graphic file with name d33e849.gif 1

where c is the minority class of the datasets.

This formula ensures that each class’s weight is directly related to its sample count. Since the denominator in the equation includes the number of samples in a given class, classes with a larger number of samples receive lower weights, while underrepresented classes are assigned relatively higher weights. The assigned weights to the minority classes will affect the loss computation in the back-propagation phase, where the loss will be higher if the models produces wrong predictions regarding the minority class, this will force the model to focus more correctly classifying them. Incorporating these weights into the categorical cross-entropy loss improved the model’s recall and F1-scores for minority classes without compromising accuracy for majority classes. Testing this weighting further optimized performance, showing that it was more effective than data augmentation alone for achieving balanced class sensitivity.

The methodology used moves beyond traditional binary classification approaches (NPC vs. normal tissue) by implementing a multi-class classification, furthermore, outlining a systematic evaluation and comparison between seven state-of-the-art CNN architectures, chosen for their established performance in image classification tasks. We Also tackle the challenge of class imbalance, a common issue in medical image datasets, by employing class weighting techniques to enhance the recognition of underrepresented classes. Finally, the study leverages a diverse dataset’s comprising WSIs from different hospitals with varying imaging conditions, contributing to a more robust assessment of model generalization across different clinical settings.

Baseline CNN pre-trained models

In this study, we picked seven pre-trained Convolutional Neural Network (CNN) architectures, each with distinct characteristics and strengths, the goal was to do a cross evaluation between them. The selected models were:

  1. DenseNet201: is a deep CNN architecture that introduces dense connections between layers, allowing feature reuse and improving gradient flow. It is known for its efficiency in feature extraction while requiring fewer parameters compared to traditional CNNs.

  2. MobileNet: Designed for mobile and resource-limited applications, it uses depth wise separable convolutions to reduce computational complexity while maintaining high accuracy. It is lightweight and efficient, making it ideal for real-time medical imaging.

  3. EfficientNetB0: introduces a compound scaling method that optimally balances model depth, width, and resolution, leading to improved performance with fewer parameters. It achieves high classification accuracy with minimal computational cost.

  4. InceptionV3: uses multiple convolutional filter sizes in parallel, enhancing feature extraction across different scales. It is particularly effective in recognizing intricate patterns in medical images.

  5. Xception: an extension of Inception, Xception replaces standard convolutions with depth wise separable convolutions, improving efficiency and performance. It is highly effective in capturing spatial correlations in histopathological images.

  6. VGG16: is a classic CNN model with a simple and uniform architecture composed of 16 layers. While computationally expensive, it provides high accuracy in image classification tasks due to its deep structure.

  7. NASNetMobile: is an automated neural architecture search (NAS)-based model designed for mobile applications. It optimizes accuracy and efficiency by dynamically selecting the most suitable convolutional operations.

Hybrid model fusion technique

The hybrid architectures developed in this study utilize an Intermediate-Level Feature Fusion (ILFF) strategy to synthesize diverse morphological representations. In this framework, the same input image is fed into a dual-stream pipeline where two distinct backbone networks (M1 and M2) process the data in parallel. Rather than relying on a single architectural bias, this approach extracts high-level feature vectors from the final Global Average Pooling (GAP) layers of both models. By capturing features at this bottleneck stage, the model preserves the abstract semantic information learned by each backbone, such as the complex spatial hierarchies in one and the efficient depth-wise patterns in another. These independent vectors are then concatenated to form a high-dimensional, unified feature representation. This fused vector serves as a comprehensive “feature map” that is subsequently passed through a dedicated classification head consisting of fully connected (Dense) layers. This head is responsible for learning the nonlinear correlations between the combined feature sets and the four diagnostic classes. The process concludes with a SoftMax activation layer, which maps the integrated features to a final probability distribution for multi-class NPC categorization. Detailed implementation is shown in Algorithm 1. Figure 4 presents the overall framework of this study. 

Fig. 4.

Fig. 4

Overall proposed framework.

Algorithm 1.

Algorithm 1

Intermediate feature-level fusion training pipeline.

Experimental setup

The experiments were conducted on the Kaggle platform, utilizing a dual NVIDIA T4 GPU setup to train all seven models. Keras transfer learning was applied to each model without fine-tuning, using pre-trained weights to streamline the training process for the specific task.

Threshold for classification and logits conversion

Classification rule and threshold selection

CNN models output logits (raw scores before applying SoftMax activation). These logits are converted into class probabilities using the SoftMax function, which is computed as in formula (2):

graphic file with name d33e948.gif 2

where:

  • P(yi) is the probability of class i,

  • Inline graphic is the logit for class

  • C is the number of classes (4 in this study).

For final classification, the model selects the class with the highest probability Inline graphic.

This method assumes equal misclassification costs across all classes and balanced class distributions. However, in real-world applications, particularly those involving class imbalance and clinical relevance such as our case, these assumptions may not hold. Therefore, relying solely on the highest SoftMax probability can lead to poor detection performance for underrepresented or minority classes.

Threshold fine-tuning with grid search

To ensure optimal classification performance, grid search was used to fine-tune threshold values for logits. The model was evaluated on multiple threshold configurations, optimizing for AUROC and F1-score. The final threshold values were selected based on the optimal trade-off between sensitivity and specificity, ensuring minority classes (LHP and NPI) were not underrepresented. This threshold fine-tuning approach essentially decouples the classification decision from the argmax operation, enabling more flexible and performance-driven decision boundaries. Additionally, it allows us to correct for the natural tendency of the model to be overconfident in majority classes due to imbalanced training data.

The final thresholds were chosen based on the grid configuration that maximized the macro-F1 score while ensuring there was no significant drop in specificity for the majority classes, alongside an improved recall for the underrepresented classes (LHP and NPI). This approach resulted in balanced decision boundaries that aligned with the real-world importance of each class, rather than being influenced solely by class frequency.

Experimental results and analysis

Performance metrics

Balanced Accuracy (3), precision (4), and F1Score (5) measure were employed to assess the classification performance of the model. The final reported accuracy represents the average result from three experimental runs. These metrics were calculated using the following formulas:

graphic file with name d33e1005.gif 3
graphic file with name d33e1009.gif 4
graphic file with name d33e1013.gif 5

Where:

  • TP is the number of true positives,

  • TN is the number of true negatives,

  • FP is the number of false positives, and.

  • FN is the number of false negatives.

Its worth noting that In the context of medical diagnostics, particularly for Nasopharyngeal Carcinoma, the F1-score serves as a more robust performance indicator than raw accuracy because it accounts for the class imbalance often found in clinical datasets. While accuracy can be misleadingly high if a model simply predicts the majority class (e.g., the more common Type 3 NPC), the F1-score provides the harmonic mean of precision and recall, effectively penalizing models that overlook minority classes or produce excessive false positives. Maintaining a high F1-score ensures that the model is both reliable in its positive identifications (precision) and comprehensive in its detection of all cancer subtypes (recall), which is vital for preventing life-altering misdiagnoses in a real-world pathological setting.

Experiments

Due to the Imbalanced nature of the datasets, we proposed the use of Augmentation and class weights to mitigate this issue. To highlight the benefit of implementing both of those techniques, experiments were carried through using augmentation only, class weight only and finally with using both of them.

Results and discussion

Table 3 presents the training and validation accuracy of the seven models. As shown in the table, all models achieved over 90% accuracy on the datasets in both training and validation.

Table 3.

Training & Validation accuracy of classification models.

Model Training Validation
DenseNet201 0.976 0.978
VGG16 1 0.96
MobileNet 0.954 0.953
Xception 0.996 0.995
InceptionV3 0.985 0.990
EfficientNetB0 0.996 0.996
NASNetMobile 0.985 0.986

Table 4 shows baseline results when not using either of augmentation or class weights, while Tables 5, 6 and 7 represent results when applying each at a time, and finally Table 5 represent the use of both techniques. Table 8 lists the obtained results when fusing the output features of best three models (DenseNet201, EfficientNet, MobileNet).

Table 4.

Testing results without augmentation nor class weights.

Model Accuracy Precision F1-Score Training Time (hrs)
DenseNet201 0.252 0.251 0.250 2.00
VGG16 0.248 0.247 0.245 2.71
MobileNet 0.245 0.249 0.236 1.16
Xception 0.249 0.248 0.242 5.25
InceptionV3 0.248 0.248 0.248 3.93
EfficientNetB0 0.250 0.193 0.218 2.27
NASNetMobile 0.252 0.252 0.252 1.77

Table 5.

Testing results when using augmentation only.

Model Accuracy Precision F1-Score Training Time (hrs)
DenseNet201 0.602 0.611 0.603 4.00
VGG16 0.552 0.570 0.566 4.71
MobileNet 0.568 0.580 0.532 3.16
Xception 0.580 0.592 0.600 9.25
InceptionV3 0.579 0.603 0.588 6.93
EfficientNetB0 0.601 0.623 0.593 4.27
NASNetMobile 0.567 0.580 0.570 3.77

Table 6.

Testing results when using class weights only.

Model Accuracy Precision F1-Score Training Time (hrs)
DenseNet201 0.369 0.373 0.372 2.00
VGG16 0.410 0.390 0.379 2.71
MobileNet 0.432 0.375 0.374 1.16
Xception 0.512 0.520 0.486 5.25
InceptionV3 0.513 0.524 0.480 3.93
EfficientNetB0 0.486 0.493 0.487 2.27
NASNetMobile 0.476 0.471 0.450 1.77

Table 7.

Testing results when using both augmentation and class weights.

Model Accuracy Precision F1-Score Training Time (hrs)
DenseNet201 0.963 0.940 0.949 4.00
VGG16 0.948 0.786 0.845 4.71
MobileNet 0.969 0.847 0.898 3.16
Xception 0.916 0.770 0.816 9.25
InceptionV3 0.917 0.669 0.749 6.93
EfficientNetB0 0.946 0.892 0.914 4.27
NASNetMobile 0.926 0.752 0.820 3.77

Table 8.

Testing results using different hybrid models.

Model Accuracy Precision F1-Score Training Time (hrs)
DenseNet201 + MobileNet 0.966 0.976 0.963 4.00
MobileNet + EfficientNetB0 0.931 0.802 0.848 3.16
DenseNet201 + EfficientNetB0 0.976 0.770 0.816 9.25

The results in Tables 5, 6 and 7 provide important context for understanding the improvements achieved with the combined strategy in Table 8. When using data augmentation only (Table 5), most models showed moderate performance, with accuracies ranging between 0.552 (VGG16) and 0.602 (DenseNet201, EfficientNetB0). Models such as Xception (0.580) and InceptionV3 (0.579) demonstrated relatively strong precision and F1-scores, but at the cost of very long training times (9.25 h and 6.93 h, respectively). On the other hand, MobileNet (0.568, 3.16 h) and NASNetMobile (0.567, 3.77 h) provided competitive accuracy with shorter training times, making them more efficient choices.

In contrast, when applying class weights only (Table 6), overall performance dropped significantly. The best-performing models in this setting, InceptionV3 (0.513 accuracy, 0.520 F1) and EfficientNetB0 (0.486 accuracy, 0.487 F1), still lagged behind their augmentation-only counterparts. Other models, such as DenseNet201 (0.369 accuracy) and VGG16 (0.410 accuracy), struggled to achieve reliable classification, despite the shorter training times observed across all models. This indicates that class weighting alone was insufficient for handling data imbalance and could not provide consistent generalization.

When both techniques were combined (Table 7), performance improved dramatically across all models, confirming the complementary effect of augmentation and class weights. Overall, MobileNet achieved the highest accuracy (0.969), indicating strong generalization capability with relatively low training time (3.16 h), making it both effective and computationally efficient. DenseNet201 also demonstrated excellent performance, with an accuracy of 0.963 and the highest F1-score (0.949), reflecting a well-balanced trade-off between precision and recall. In contrast, VGG16 and InceptionV3 produced comparatively lower precision and F1-scores, suggesting less robust classification performance under the same training conditions. Xception recorded the longest training time (9.25 h) while achieving only moderate performance (accuracy of 0.916), which may indicate higher computational complexity without proportional performance gains. Meanwhile, EfficientNetB0 showed competitive results (accuracy of 0.946 and F1-score of 0.914) with moderate training time, highlighting its efficiency. NASNetMobile achieved balanced but slightly lower metrics overall. These results suggest that combining augmentation with class weighting significantly benefits lightweight and well-optimized architectures such as MobileNet and DenseNet201, offering strong predictive performance with reduced computational cost.

To further explore potential performance gains, a fusion strategy was applied to the three best-performing architectures (MobileNet, DenseNet201 and EfficientNetB0) based on their reported accuracy and F1-Score in Table 7. As presented in Table 8. The intermediate-level fusion approach produced notable improvements. The DenseNet201 + MobileNet combination achieved strong and well-balanced results (0.966 accuracy, 0.963 F1-score) with moderate training time (4.00 h). The MobileNet + EfficientNetB0 configuration obtained 0.931 accuracy and 0.848 F1-score, reflecting stable but comparatively lower performance. Most notably, the EfficientNetB0 + DenseNet201 fusion achieved the highest overall accuracy (0.976), surpassing all individual models; however, this came at the expense of longer training time (9.25 h) and a comparatively lower F1-score (0.816), indicating less balanced class-wise performance. The confusion matrices for the models is shown in Fig. 5.

Fig. 5.

Fig. 5

Confusion matrices for fusion models (a) DenseNet201 + MobileNet (b) MobileNet+EfficientNetB0 (c) DenseNet201 + EfficientNetB0.

Moreover, the models demonstrated strong consistency by maintaining high performance across datasets from two different hospitals (SGH and HKL), which used different imaging magnifications and conditions. The successful application of transfer learning from pre-trained ImageNet weights further validated its advantage over training models from scratch, a finding supported by the broader literature on medical image analysis.

Limitations and future work

Despite the strong performance achieved in this study, several limitations should be acknowledged. Although the dataset contains a substantial number of annotated samples, it was derived from only two institutions, which may limit the generalizability of the proposed models to broader clinical settings with different staining protocols and patient populations. Additionally, the current framework operates at the patch level rather than full slide-level inference, which may restrict the ability to capture global tissue context relevant for clinical diagnosis. The intermediate feature fusion strategy, while improving predictive performance, introduces additional computational complexity that may affect deployment in resource-constrained environments. Furthermore, the models rely on transfer learning from natural image datasets, and domain-specific pre-training may further enhance robustness. Finally, the interpretability of deep learning models remains an important consideration for clinical adoption, highlighting the need for integrating explainable artificial intelligence techniques in future work.

Conclusion

This study has demonstrated the effectiveness of deep learning in the automated multi-class classification of nasopharyngeal carcinoma (NPC) using whole slide histopathological images. Among the seven convolutional neural network (CNN) models evaluated under the combined augmentation and class-weighting strategy, MobileNet achieved the highest individual accuracy (0.969), while DenseNet201 delivered the most balanced performance, achieving an accuracy of 0.963 and the highest F1-score (0.949). We attribute this performance superiority to DenseNet’s unique approach to feature propagation and gradient flow using dense connectivity, where each layer receives direct inputs from all preceding layers within a dense block. This architectural design ensures maximum feature reuse and significantly mitigates the vanishing gradient problem. For NPC histology, this is particularly advantageous, the dense connections preserve low-level morphological features (such as nuclear membranes and chromatin textures) that are often lost in deeper networks, while simultaneously building high-level semantic representations. These results highlight the strong capability of lightweight and well-optimized architectures to provide both high predictive performance and computational efficiency. Although Xception demonstrated acceptable classification ability, its significantly longer training time (9.25 h) without proportional performance gains makes it less practical for time-sensitive deployment scenarios.

To further enhance performance, intermediate-level fusion was applied to the top-performing architectures. The EfficientNetB0 + DenseNet201 fusion achieved the highest overall accuracy (0.976), surpassing all single-model configurations. The DenseNet201 + MobileNet combination also demonstrated strong and balanced results (0.966 accuracy, 0.963 F1-score) with moderate computational cost. The performance gains observed through intermediate-level feature fusion can be attributed to the complementary inductive biases inherent in the different architectures. While DenseNet201 utilizes a connectivity pattern that maximizes feature reuse and preserves low-level structural details, EfficientNetB0 employs compound scaling to capture high-level semantic abstractions with optimized depth and width. By concatenating both of the learned feature vectors, the unified classifier gains access to a heterogeneous feature space that individual models cannot construct alone, and allowing it to compensate for the representative weaknesses of one model with the strengths of another, ultimately leading to a more discriminative representation of complex NPC histological patterns. However, certain fusion configurations, despite achieving marginal accuracy improvements, exhibited lower F1-scores or increased training time, indicating that performance gains must be carefully weighed against computational overhead and class-wise balance.

The computational profile of the evaluated models reveals a clear trade-off between architectural depth and practical training efficiency. While complex architectures like Xception and the EfficientNetB0 + DenseNet201 fusion achieved high peak accuracies, their substantial training durations (up to 9.25 h) highlight a heavy reliance on high-performance hardware. In contrast, MobileNet and the DenseNet201 + MobileNet fusion emerged as the most viable candidates for clinical edge deployment, offering near-state-of-the-art accuracy with significantly reduced training times (ranging from 3.16 to 4.00 h). These findings suggest that for NPC histological classification, the diminishing returns in accuracy provided by ultra-heavyweight fusions may not always justify the 131% increase in computational overhead compared to optimized, lightweight alternatives like MobileNet.

The findings highlight that deep learning models can reliably distinguish between NPC, lymphoid hyperplasia (LHP), nasopharyngeal inflammation (NPI), and normal tissue, addressing a critical limitation in existing literature that often focuses on binary classification. The multi-class framework enhances clinical relevance by enabling more precise differentiation between malignant and non-malignant pathological conditions that may present with similar histological features.

Furthermore, the integration of data augmentation and cost-sensitive class weighting was instrumental in mitigating the inherent dataset imbalance. While augmentation expanded the representative diversity of minority classes to improve spatial invariance, class weighting adjusted the loss function to penalize misclassifications in underrepresented categories more heavily. This dual strategy prevented the models from developing a majority-class bias, thereby enhancing generalization and ensuring a balanced precision-recall trade-off across all histological subtypes. The consistently high F1-scores observed in top-performing models demonstrate improved recognition of minority classes such as LHP and NPI without sacrificing overall accuracy. Additionally, the use of transfer learning from ImageNet contributed positively to feature extraction quality and convergence stability, reaffirming the benefit of leveraging pre-trained representations in medical image analysis.

Despite these promising results, challenges remain. Datasets size and diversity continue to influence generalization capability, emphasizing the need for broader multi-institutional validation. Moreover, model interpretability remains a critical consideration for clinical deployment, necessitating the integration of explainable AI techniques to enhance transparency and trust among pathologists.

The findings of this study provide a framework for enhancing Computer-Aided Diagnosis (CAD) systems by reducing the inherent subjectivity of nasopharyngeal carcinoma (NPC) screening. By leveraging the EfficientNet + DenseNet201 fusion strategy, CAD systems can act as a high-precision “second opinion,” capable of identifying subtle morphological markers in undifferentiated NPC that may be overlooked during manual microscopic examination. Critically, the evaluation of these models across multi-center datasets from two different hospitals confirms their consistent performance despite variations in staining protocols and scanning hardware. This cross-dataset consistency suggests that the models have achieved high generalization capability, ensuring they remain reliable in diverse clinical environments. Furthermore, the high performance of MobileNet-based architectures suggests these models can be integrated into resource-constrained settings for real-time triage. This allows pathologists to prioritize complex cases flagged as high-risk, thereby streamlining the diagnostic workflow, reducing fatigue-related errors, and ultimately standardizing NPC grading across different healthcare institutions.

Overall, this study establishes a solid foundation for deep learning-assisted NPC diagnosis and highlights the potential of optimized single models and strategic model fusion for robust, clinically applicable classification systems.

Author contributions

Author Contributions: Conceptualization, M.K.A., S.B.M. and M.F.A.F.; Methodology, M.K.A.; Software, M.K.A.; Validation, M.K.A., A.S.W. and M.S.N.; Formal Analysis, M.K.A.; Investigation, M.K.A.; Data Curation, M.K.A., A.S.W. and A.M.I.; Writing Original Draft Preparation, M.K.A.; Writing Review & Editing, S.B.M., M.F.A.F. and W.S.H.M.W.A.; Visualization, M.K.A.; Supervision, S.B.M. and M.F.A.F.; Project Administration, S.B.M.; Funding Acquisition, M.F.A.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding and was not supported by any grant from funding agencies in the public, commercial, or not-for-profit sectors.

Data availability

The datasets used in this study are derived from private clinical data obtained from Sarawak General Hospital (SGH) and Hospital Kuala Lumpur (HKL). Due to ethical restrictions and patient confidentiality, the data are not publicly available. Access to the datasets may be considered upon reasonable request and is subject to approval by the respective institutions and relevant ethical review boards.

Declarations

Competing interests

The authors declare no competing interests.

Institutional ethical approval and compliance

All methods were carried out in accordance with relevant guidelines and regulations. The study protocol was reviewed and approved by the institutional ethics committees of Sarawak General Hospital (SGH) and Hospital Kuala Lumpur (HKL). Informed consent was obtained from all subjects and/or their legal guardians prior to the use of the data for research purposes.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.LeCun, Y. et al. Backpropagation applied to handwritten Zip Code recognition. Neural Comput.1(4), 541–551. 10.1162/neco.1989.1.4.541 (1989). [Google Scholar]
  • 2.Mienye, I. D., Swart, T. G., Obaido, G., Jordan, M. & Ilono, P. Deep convolutional neural networks in medical image analysis: A review. Information16(3), 195. 10.3390/info16030195 (2025). [Google Scholar]
  • 3.Wang, S.-X. et al. The detection of nasopharyngeal carcinomas using a neural network based on nasopharyngoscopic images. Laryngoscope134, 127–135. 10.1002/lary.30781 (2024). [DOI] [PubMed] [Google Scholar]
  • 4.Chan, J. Y., Paterson, I. C. & Yap, L. F. Nasopharyngeal carcinoma in Southeast Asia: Current landscape and future priorities. Br. J. Biomed. Sci.82, 15902. 10.3389/bjbs.2025.15902 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Su, Z. Y., Siak, P. Y., Leong, C.-O. & Cheah, S.-C. The role of Epstein–Barr virus in nasopharyngeal carcinoma. Front. Microbiol.14, 1116143. 10.3389/fmicb.2023.1116143 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Mohammed, M. A., Abd Ghani, M. K., Hamed, R. I. & Ibrahim, D. A. Analysis of an electronic methods for nasopharyngeal carcinoma: Prevalence, diagnosis, challenges and technologies. J. Comput. Sci.21, 241–254. 10.1016/j.jocs.2017.04.006 (2017). [Google Scholar]
  • 7.Cantù, G. Nasopharyngeal carcinoma. A “different” head and neck tumour. Part A: From histology to staging. Acta Otorhinolaryngol. Ital.43(2), 85–98. 10.14639/0392-100X-N2222 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Glastonbury, C. M. & Salzman, K. L. Pitfalls in the staging of cancer of nasopharyngeal carcinoma. Neuroimaging Clin. N. Am.23(1), 9–25. 10.1016/j.nic.2012.08.006 (2013). [DOI] [PubMed] [Google Scholar]
  • 9.Omoush, S. A., Alzyoud, J. A., El-Omari, N. K. T. & Alzyoud, A. J. The role of whole slide imaging in AI-based digital pathology: Current challenges and future directions—An updated literature review. J. Mol. Pathol.7(1), 2. 10.3390/jmp7010002 (2026). [Google Scholar]
  • 10.Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q. Densely connected convolutional networks, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pp. 4700–4708. (2017). 10.48550/arXiv.1608.06993
  • 11.Howard, A. G. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704 04861. 10.48550/arXiv.1704.04861 (2017). [Google Scholar]
  • 12.Tan, M. Efficientnet: Rethinking model scaling for convolutional neural networks, arXiv preprint arXiv:1905.11946, (2019).
  • 13.Zoph, B., Vasudevan, V., Shlens, J. & Le, Q. V. Learning transferable +architectures for scalable image recognition, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pp. 8697–8710. (2018). 10.48550/arXiv.1707.07012
  • 14.Wibawa, M. S. et al. AI-based risk score from tumour-infiltrating lymphocyte predicts locoregional-free survival in nasopharyngeal carcinoma. Cancers15(24), 5789. 10.3390/cancers15245789 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J. & Wojna, Z. Rethinking the inception architecture for computer vision, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pp. 2818–2826. (2016).
  • 16.Chollet, F. Xception: Deep learning with depthwise separable convolutions, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pp. 1251–1258 (2017). 10.48550/arXiv.1610.02357
  • 17.Simonyan, K. & Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint, (2014). arXiv:1409.1556.
  • 18.Chuang, W. Y., Chang, S. H., Yu, W. H., Yang, C. K. & Yeh, C. J. Successful identification of nasopharyngeal carcinoma in nasopharyngeal biopsies using deep learning. Cancers10.3390/cancers12020507 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Wong, L. M., King, A. D., Ai, Q. Y. H., Lam, W. K. J. & Poon, D. M. C. Convolutional neural network for discriminating nasopharyngeal carcinoma and benign hyperplasia on MRI. Eur. Radiol.10.1007/s00330-020-07451-y (2021). [DOI] [PubMed] [Google Scholar]
  • 20.Li, S. et al. The tumor target segmentation of nasopharyngeal cancer in CT images based on deep learning methods. Technol. Cancer Res. Treat.10.1177/1533033819884561 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Hu, D. et al. Application of deep learning in histopathology images of breast cancer: A review. Micromachines10.3390/mi13122197 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Diao, S., Luo, W., Hou, J. & Lambo, R. Deep multi-magnification similarity learning for histopathological image classification. IEEE J. Biomed. Health Inform.10.1109/JBHI.2023.3237137 (2023). [DOI] [PubMed] [Google Scholar]
  • 23.Li, S., Deng, Y., Zhu, Z., Hua, H. & Tao, Z. A comprehensive review on radiomics and deep learning for nasopharyngeal carcinoma imaging. Diagnostics10.3390/diagnostics11091523 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Ke, L. et al. Development of a self-constrained 3D DenseNet model in automatic detection and segmentation of nasopharyngeal carcinoma using magnetic resonance images. Oral Oncol.10.1016/j.oraloncology.2020.104862 (2020). [DOI] [PubMed] [Google Scholar]
  • 25.Mu, X. et al. Deep learning model using planar whole-body bone scintigraphy for diagnosis of skull base invasion in patients with nasopharyngeal carcinoma. J. Cancer Res. Clin. Oncol.10.1007/s00432-024-05969-y (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Hu, Z. et al. Weakly supervised classification for nasopharyngeal carcinoma with Transformer in whole slide images. IEEE J. Biomed. Health Inform.28(12), 7251–7262. 10.1109/JBHI.2024.3422874 (2024). [DOI] [PubMed] [Google Scholar]
  • 27.Nakagawa, J., Fujima, N., Hirata, K. & Harada, T. Diagnosis of skull-base invasion by nasopharyngeal tumors on CT with a deep-learning approach. Jpn. J. Radiol.10.1007/s11604-023-01527-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Tülü, Ç. N. & Kaya, Y. Convolutional Fourier Analysis Network (CONV-FAN-POX): A novel time–frequency approach for medical image analysis. Biomed. Signal Process. Control117, 109698. 10.1016/j.bspc.2026.109698 (2026). [Google Scholar]
  • 29.Huang, C., Liu, X., Wang, W. & Guo, Z. Exploring the mechanism of Centipeda minima in treating nasopharyngeal carcinoma based on network pharmacology. Curr. Comput. Aided Drug Des.21, 1–14. 10.2174/0115734099305631240930054417 (2025). [DOI] [PubMed] [Google Scholar]
  • 30.Ding, Y. et al. Bioprinting in tumor model construction for head and neck squamous cell carcinoma: A review. Int. J. Bioprint.11(2), 139–163. 10.36922/ijb.8100 (2025). [Google Scholar]
  • 31.Zhang, Y. et al. Deep learning-based fusion of nuclear segmentation features for microsatellite instability and tumor mutational burden prediction in digestive tract cancers: A multicenter validation study. Brief. Bioinform.26(6), bbaf580. 10.1093/bib/bbaf580 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Nabi, M. S. et al. Hybrid Deep Learning Framework for Multi-Class Breast Cancer Scoring Using Grad-CAM++, 2025 9th International Conference on Information Technology (InCIT), Phuket, Thailand, pp. 257–264, (2025). 10.1109/InCIT66780.2025.11276061
  • 33.Nabi, M. S. et al. Explainable AI for Breast Cancer Diagnosis Using EfficientNetB3 with Attention Mechanism, TENCON 2025–2025 IEEE Region 10 Conference (TENCON), Kota Kinabalu, Malaysia, pp. 1554–1558, (2025). 10.1109/TENCON66050.2025.11374955
  • 34.Bannah, H. 9th International Conference on Information Technology (InCIT), Phuket, Thailand, 2025, pp. 707–714, (2025). 10.1109/InCIT66780.2025.11276080
  • 35.Bannah, H. et al. Automated Nuclei Segmentation in PR-IHC Breast Cancer Images Using the Cellpose Deep Learning Model, TENCON 2025–2025 IEEE Region 10 Conference (TENCON), Kota Kinabalu, Malaysia, pp. 1709–1713, (2025). 10.1109/TENCON66050.2025.11375047
  • 36.Bannah, H. et al. Breast cancer nuclei segmentation in ER-IHC images using deep Learning, Multimedia University Engineering Conference (MECON), Cyberjaya, Malaysia, 2025, pp. 1–6, Cyberjaya, Malaysia, 2025, pp. 1–6, (2025). 10.1109/MECON67253.2025.11277170
  • 37.Nabi, M. S. et al. Skin Lesion Classification with Explainable AI to Enhancing Dermatological Diagnosis, Multimedia University Engineering Conference (MECON), Cyberjaya, Malaysia, 2025, pp. 1–6, Cyberjaya, Malaysia, 2025, pp. 1–6, (2025). 10.1109/MECON67253.2025.11277079
  • 38.Nabi, M. S. et al. Explainable deep learning models for HER2 IHC scoring in breast cancer diagnosis. Inform. Med. Unlocked 101700. 10.1016/j.imu.2025.101700 (2025). [Google Scholar]
  • 39.Nabi, M. S. et al. KDH-Net: Explainable medical AI for multiclass kidney disease characterization from CT images. J. Clin. Med.15(8), 3165. 10.3390/jcm15083165 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets used in this study are derived from private clinical data obtained from Sarawak General Hospital (SGH) and Hospital Kuala Lumpur (HKL). Due to ethical restrictions and patient confidentiality, the data are not publicly available. Access to the datasets may be considered upon reasonable request and is subject to approval by the respective institutions and relevant ethical review boards.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES