Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Feb 1;32(6):1556–1570. doi: 10.1111/odi.70210

SSDA_AOA: Stacked Sparse Denoising Autoencoder With Archimedes Optimization Algorithm Based Oral Cancer Detection on Histopathological Images

R Sathish Kumar 1,✉, M Govindarajan 1
PMCID: PMC13457694  PMID: 41622690

ABSTRACT

Objectives

This study aims to address the challenges in the diagnosis of oral cancer by proposing a novel computer‐aided diagnostic framework that leverages advanced deep learning (DL) and optimization techniques to enhance early detection and improve patient outcomes.

Materials and Methods

In the framework proposed, the histopathological images are subjected to a preprocessing technique, and then, the images are fed directly to the NASNet‐Large model for the extraction of high‐level discriminative texture features. The resultant vectors obtained from the features extracted act as input to the search space of Archimedes Optimization Algorithm that carries out dimensionality reduction and optimal hyperparameter tuning simultaneously. The optimized feature subset is fed to the final classifier, namely the Stacked Sparse Denoising Autoencoder that learns robust latent representations.

Results

The findings demonstrate that the proposed approach achieves superior performance, achieving an accuracy of 95.38%, a precision of 95.15%, a sensitivity of 91.78%, a specificity of 91.85%, and an F1‐score of 93.72%.

Conclusions

These findings underscore the potential of the SSDA‐AOA framework as an effective tool for the early detection and precise classification of oral cancer, paving the way for improved patient outcomes through timely intervention.

Clinical Relevance

This innovative approach may significantly enhance patient outcomes by facilitating earlier diagnosis and treatment, addressing the urgent need for more reliable diagnostic tools in oncology.

Keywords: Archimedes optimization algorithm, histopathological images, NASNet‐Large, oral cancer, stacked sparse denoising autoencoder

1. Introduction

Globally, oral cancer is a significant health concern, and its incidence is increasing annually (Pal et al. 2023). It can be classified into different types based on its origin, specifically carcinoma and sarcoma. Among these, oral squamous cell carcinoma (Oral SCC) stands out as the most prevalent form, primarily arising from oral potentially malignant disorders (OPMDs) (Rashid et al. 2024; Mira et al. 2024). This form of cancer has the potential to impact several regions of the oral cavity, such as the tongue, roof of the mouth, lips, cheeks, hard palate, soft palate, throat, and sinuses (Kumari et al. 2022). The prevalent risk factors associated with oral cancer include the unsanitary practice of chewing tobacco, the intake of alcohol, exposure to the human papillomavirus (HPV), hereditary conditions, and regional influences (Smędra and Berent 2023; Bouvard et al. 2022). Besides, oral cancer often exhibits a substantial mortality rate in the absence of timely detection. Despite advances in treatment modalities such as surgery, radiotherapy, and chemotherapy, prognosis remains poor, particularly in advanced stages, leading to high morbidity and reduced quality of life (Dar et al. 2022). Henceforth, early diagnosis, primarily through surgical biopsy and histopathologic evaluation, is crucial for improving survival rates, as prognosis declines significantly with disease progression.

At present, the histopathological examination of biopsy samples or tissue from affected oral cells, conducted by histopathologists, remains the gold standard for diagnosing oral cancer (Yang et al. 2022). Histopathologists primarily base their cancer detection on various criteria, including nuclear morphology, cellular distribution, cell size and shape, as well as patterns of cellular growth (Umapathy et al. 2022). These specialists must scrutinize cells with keen observation and interpret the visual data informed by their prior knowledge of medical science. This method of cancer diagnosis is labor‐intensive as well as prone to inaccuracies (Dixit et al. 2023). However, even for skilled pathologists, making a correct diagnosis can be difficult (Adhit et al. 2023). Consequently, there is a significant demand for an effective diagnostic tool that can aid histopathologists and medical professionals in more accurately and efficiently identifying oral malignancies. Happily, Artificial Intelligence (AI) can help make this process more precise and manageable by giving medical personnel the knowledge they need to act quickly to enhance patient outcomes (Al‐Rawi et al. 2022).

Recently, Deep learning (DL), a subset of AI, utilizes neural networks to automatically learn complex patterns from raw data, such as medical images, enabling advanced diagnostic and prognostic capabilities (Al‐Rawi et al. 2022; Varalakshmi et al. 2024). Specifically, Convolutional Neural Network (CNN) is widely adopted in the medical field for tasks like abnormality detection (Warin, Limprasert, Suebnukarn, Jinaporntham, Jantana, and Vicharueang 2022; Raval and Undavia 2023). DL has demonstrated its potential to match or exceed the diagnostic accuracy of healthcare professionals (Shah and Pareek 2024). In certain instances, a combination of DL models can be utilized to enhance model performance effectively. For instance, the deep neural network (DNN) first underwent training under the unsupervised learning paradigm during its initial learning phase. Subsequently, this was succeeded by a process of supervised fine‐tuning applied to the stacked network known as stacked autoencoders (SAEs). For instance, Stacked Sparse Autoencoders (SSAEs) have been effectively utilized for complex segmentation tasks, such as in vertebral analysis from CT scans (Qadri et al. 2022). The relevance of such DL techniques extends directly to histopathological image classification. Besides, researchers have employed patch‐based DL models for the classification of breast cancer from histopathological images, demonstrating the feasibility of this approach for cancer diagnostics (Hirra et al. 2021). Furthermore, the trend of combining autoencoders with optimization algorithms has yielded powerful results; the author in (Vaiyapuri et al. 2022) has integrated a stacked sparse denoising autoencoder (SSDA) with modified metaheuristic algorithms for cervical cancer classification, achieving superior performance. This paradigm of autoencoder‐assisted hybrid modeling is further validated in recent studies, such as the use of autoencoders within stacked ensemble frameworks for lymphoma subtype classification (Ogundokun et al. 2025). These studies collectively establish a strong foundation for using advanced, hybrid autoencoder‐based models for complex medical image classification tasks.

However, while these techniques have been explored for other cancers and imaging modalities, their specific application and optimization for oral cancer classification using histopathological images remain an area ripe for investigation. Therefore, this study aims to develop an oral cancer classification model by leveraging a hybrid DL framework using histopathological images. In the first stage, the input image is preprocessed with a median filter to enhance image quality. Then, a pretrained NasNet‐Large model is employed for robust feature extraction, with its hyperparameters fine‐tuned by the AOA to maximize feature relevance. Finally, the classification task is performed with the proposed Stacked Sparse Denoising Autoencoder (SSDA‐AOA) model. The model performance is validated with a histopathological image dataset and obtains superior accuracy and provides needful insights to the medical professional to take a suitable decision. Overall, this study provides the following research contribution and motivation,

  • To employ noise removal and median filter approach to enhance the contrast of the input image, hence improve the prediction accuracy of the model.

  • To develop the NasNet‐Large feature extraction model to extract the most important attributes and optimized with AOA, thereby emphasizing the robustness of the framework.

  • To propose SSDA‐AOA model for precise oral cancer prediction using the histopathological image.

  • To examine the proficiency of the proposed model through metrics such as accuracy, precision, recall, and F1‐score.

The paper is structured as, Section 1 provides an overview of oral cancer and its important to predict at an early stage. Section 2 reviews the existing studies in oral cancer classification using ML and DL models. Then, Section 3 deliberates the working principle of the proposed methodology. Section 4 depicts the key findings of the present study model. Finally, Section 5 ends with a conclusion and future direction of the research.

2. Related Works

Oral cancer remains a significant global health concern, with its incidence rising steadily and early detection being critical for improving patient outcomes. The traditional gold standard for oral cancer diagnosis is a tissue biopsy followed by histopathological assessment, a process that is invasive, requires specialized expertise, and can be time‐consuming and costly. As a result, several research studies have focused on developing and evaluating alternative diagnostic aids and screening methods to facilitate earlier, less invasive, and more accessible detection. This subsection provides the synthesized information of the prevailing studies that employed AI approaches to classify oral cancer stages.

Accordingly, Das et al. (2023) have aimed to develop automatic and early detection of oral SCC from histopathological images. It has proposed a CNN model against seven standard CNN models. The developed model has obtained remarkable cross‐validation accuracy and indicated its effectiveness for automatic classification of oral cancer data. Meanwhile, Myriam et al. (2023) have developed an oral cancer detection method using a CNN and an optimized Deep Belief Network (DBN). Additionally, a combination of Particle Swarm Optimisation (PSO) and Al‐Biruni Earth Radius Optimisation (BERO) approaches have been developed, referred to as PSOBER. The model yielded better accuracy of 97.35%. Similarly, Jubair et al. (2022) have utilized a pretrained EfficientNet‐B0 to develop a lightweight deep DCNN. The model has achieved an accuracy of 85.0%, specificity of 84.5%, and sensitivity of 86.7%. These results demonstrate that such deep CNNs can effectively support low‐cost, embedded diagnostic devices for early oral cancer detection. Another model by Warin, Limprasert, Suebnukarn, Jinaporntham, and Jantana (2022) has found that the use of DenseNet‐121 and ResNet‐50 yielded favorable diagnostic performance. The detection performance was the highest, with an AUC of 74.34%. Bansal et al. (2022) has exploed 5 deep TL models to detect lesions and malignancies of oral cancer using histopathological images. These models were used as feature extractors and optimized with Adam, SGD, and RMSprop. Among 5 TL model, DensNet has attained notable accuracy as 92.1 on standard and real‐world datasets. Another work by Sharma et al. (2022) has utilized Deep CNN for classifying Oral SCC and OPMDs. DenseNet‐169 has achieved near‐expert performance in classifying OSCC and OPMDs from 980 oral images, with a satisfactory level of performance. CNN‐based models outperformed general practitioners, highlighting their potential for early oral cancer detection.

The transfer learning (TL) strategy proposed by Rahman et al. (2022) has involved the use of AlexNet inside a CNN to extract rank features from OSCC biopsy images and then train the model. The model has achieved a moderate level of classification accuracy during training and testing. The study as suggested by Chang et al. (2023) has evaluated and compared five DCNN for identifying multiple types of oral cancer tissues using 16,200 Raman spectra from 180 tissue samples. Among them, ResNet50 has yielded better performance, with an overall accuracy of 92.81%, precision of 92.93%, and recall of 92.86%. Based on these results, the researchers have developed a prototype intelligent detection system to aid clinical diagnosis of oral cancer. Another prevailing work by Babu et al. (2024) has presented an automated DL system using TL to improve early detection and diagnosis of oral cancer, particularly in low‐ and middle‐income countries. By comparing different TL approaches, the authors have found that the Inception‐V3 algorithm has obtained a satisfactory level of accuracy for classifying oral cancer. Their preliminary results have showed that this method effectively supports early diagnosis and could enhance patient outcomes. Likewise, the study by Huang et al. (2023) has developed a DL‐based methodology enhanced by a metaheuristic approach for accurate oral cancer diagnosis. It has utilized gamma correction, noise reduction, and data augmentation for preprocessing, and optimized CNN weights using an improved squirrel search algorithm (ISSA). The framework has been applied to the Oral Cancer (Lips and Tongue) images dataset; the method demonstrated promising diagnostic accuracy compared to existing techniques. Similarly, Sampath et al. (2024) have explored an oral cancer detection model by leveraging a Feature Fusion DCNN combined with Stochastic Gradient‐based Logistic Regression (LR), using images of lips and tongue for diagnosis. The ResNet‐50 model has been utilized to extract multilayer convolutional features, followed by fusion and LR layers, achieving remarkable classification accuracy on a benchmark dataset.

Alternatively, the author Afify et al. (2023) in study has developed a deep TL model combining CNNs and Grad‐CAM to classify and localize lesions in oral squamous cell carcinoma (OSCC) histopathological images using a public dataset of 1224 images at two magnifications. Among 10 CNN models tested, ResNet‐101 has obtained notable accuracy at 100× magnification and EfficientNet‐b0 achieved 95.65% at 400×, outperforming other models. Grad‐CAM was used for lesion localization, enhancing the model's robustness and supporting early, accurate OSCC detection. Notably, Oya et al. (2023) has utilized histopathologic images from 90 tongue SCC patients, incorporating both cellular atypia (high magnification) and structural atypia (low magnification) for diagnosis. EfficientNet B0 was trained on huge image patches at various magnifications and input sizes with 512 × 512 pixel images. Then, Grad‐CAM analysis has confirmed that the AI model effectively focused on both cellular and structural atypia, particularly around the basal layer, demonstrating the suitability of this approach for diagnosing oral SCC. Another model by Ananthakrishnan et al. (2023) has implemented two classification methods: one based on handcrafted textural features with ML, and another using features from pretrained CNNs with a Random Forest (RF) classifier. The model has achieved test accuracies of 96.94% (AUC 0.976) for 400 images.

Recently, Ragab and Asar (2024) has introduced the SEHDL‐OSCCR technique, which applied a hybrid DL approach for recognizing Oral SCC in histopathological images. The method has involved bilateral filtering for noise removal, SE‐CapsNet for feature extraction optimized by an improved crayfish algorithm, and classification via a CNN‐BiLSTM model, obtained the improvised accuracy on benchmark datasets. Another work by Majeed et al. (2024) has integrated SMOTE and TL to address class imbalance issue in dataset. The model has yielded better performance with VGG16 and EfficientNetB3 models. This integrated approach further enhanced diagnostic performance, ensuring reliable detection even when cancerous samples are underrepresented. Similarly, Bharanidharan et al. (2024) has combined features extracted from VGG16, MobileNetV2, and ResNet50, followed by a variance‐based feature selection method. It has analyzed 439 oral cancer and 89 normal images, attained moderated level of performance and indicated improvisation in early oral cancer detection from histopathological images. Khan and Asif (2024) have incorporated a self‐attention block with TL models EfficientNetB0 and EfficientNetB1. This method effectively combined features from both models, enhancing predictive accuracy while reducing model complexity. Evaluated on the MOD dataset, the approach achieved an impressive accuracy. Another work by Panigrahi et al. (2022) has implemented a DL based capsule network method to classify oral cancer from histopathological images, utilizing dynamic routing to improve robustness against rotation and affine transformations. The approach has effectively identified Oral SCC by analyzing architectural differences in epithelial layers and the presence of keratin pearls. The model achieved moderate performance in classifying OSCC images. Sukegawa et al. (2023) has found that the VGG16 CNN model with the spectral angle mapper optimizer achieved the good accuracy and AUC for classifying histopathological images of oral SCC. When oral pathologists used the model's results as Supporting Information, their diagnostic performance improved significantly.

Recently, the author Saraswathi and Murali Bhaskaran (2025) have presented a hybrid DL model, Recurrent Deep Belief Network (RDBN), optimized using a Hybrid Beetle‐Barnacle Swarm Optimization (HBBSO) algorithm for oral cancer detection. Preprocessing was performed using median filtering and CLAHE to enhance image quality, and hyperparameters were fine‐tuned for optimal performance. Experimental results have showed the model outperformed existing methods, achieving up to 5.6% higher accuracy in classifying oral cancer. Srinivasulu et al. (2025) has developed a hybrid CNN‐RNN model that integrated feature extraction and temporal sequence learning to improve throat cancer prediction and detection using a dataset of patient symptoms, medical histories, and demographics. This approach has addressed limitations of traditional DL methods, such as overfitting and imbalanced data, resulting in enhanced diagnostic accuracy for throat cancer. The study as suggested by Meer et al. (2025) has evaluated a framework by integrating Self‐Attention MobileNet‐V2 and Self‐Attention DarkNet‐19, optimized using the Whale Optimization Algorithm and Quantum WOA for feature selection. Features from both models were fused via canonical correlation analysis and classified using wide neural networks. The method, has evaluated on augmented oral cancer datasets (100× and 400×), achieved promising accuracies outperforming state‐of‐the‐art techniques. Another work by Razmjouei et al. (2025) has presented fuzzy Non‐linear Fuzzy Rank‐based Ensemble DL (NFR‐EDL) and validated on multiple standard dataset and achieved satisfied level of performance comparing with other approach. Similarly, Shah et al. (2025) have introduced OCANet, a U‐Net‐based model with advanced attention mechanisms to improve tumor segmentation in oral cancer histopathology images. OCANet outperforms existing models in accuracy and generalizes well across datasets. Its integrated visual and textual interpretability enhances transparency and supports clinical decision‐making. Tanriver et al. (2021) proposed a patch‐based deep learning approach to extract the discriminative features from unlabeled data using a stacked sparse autoencoder (SSAE). SAE uses pixel intensities alone to learn high‐level features to recognize distinctive features from image patches. Each image is subjected to a sliding window operation to express image patches using autoencoder high‐level features, which are then fed into a sigmoid layer to classify whether each patch is a vertebra or not.

2.1. Problem Identification

The following aspects of research problems are identified by analyzing the existing studies, that are as follows:

  • The suggested study (Das et al. 2023) has trained on limited or non‐representative datasets; hence, the model is affected by the overfitting issue.

  • Though the considered studies (Huang et al. 2023; Sukegawa et al. 2023) performed well in detecting the oral SCC but yielded a moderate level of accuracy.

  • DL models (Srinivasulu et al. 2025) have faced a data imbalance issue that significantly diminished the performance of the model.

To address the insufficient accuracy and overfitting issue that exists in the prevailing work, this study utilizes the NasNet‐Large model to capture relevant features, and classification is performed by the proposed SSDA‐AOA model.

3. Methodology

Early detection of oral cancer could improve the quality of life of the affected patient. Thus, the motivation behind the proposed research is to accurately classify the oral cavity cancer by employing the proposed SSDA‐AOA model. It included four steps such as preprocessing, feature extraction, optimization, and classification. These steps are separately explained in the following section. Figure 1 depicts the overall process flow involved in the classification process.

FIGURE 1.

FIGURE 1

Workflow of proposed SSDA_AOA classification model.

Figure 1 illustrates the basic flow of the process involved in the proposed work. Accordingly, the process begins with raw histopathological images to detect oral cancer stages. These input images undergo preprocessing steps, including noise removal and contrast enhancement, to improve visual clarity and highlight diagnostically relevant features. Following preprocessing, the system employs NasNet‐Large model to extract the highly important attributes from the preprocessed image. Then, the Archimedes Optimization Algorithm (AOA) is used to fine‐tune model parameters, optimizing the subsequent classification process. The refined images are then classified using a Stacked Sparse Denoising Autoencoder (SSDA), a DL approach that effectively learns and distinguishes complex patterns within the data. Finally, the classification results are visualized graphically, allowing for clear comparison between different diagnostic categories, such as benign and malignant conditions, thus demonstrating a comprehensive integration of advanced AI techniques in medical image diagnostics. The integration of SSDA with AOA was designed specifically to address the diagnostic challenges inherent to oral cancer histopathological images, such as staining variability, artifacts, complex tissue morphology, and high intra‐class heterogeneity. This is because AOA identifies the most discriminative NasNet‐Large feature subset with simultaneous tuning of the SSDA hyperparameters.

3.1. Data Preprocessing

Commonly, preprocessing is an essential step in image analysis, which is applied on raw image to improve their quality and make diagnostically relevant features more discernible before further analysis. The input histopathological image typically has noise; in order to mitigate the noise present in unprocessed images, it is necessary to eliminate extraneous visual data. During the preprocessing phase, modifications are made to the contrast in order to enhance the distinction and demarcation between various structures, as well as between healthy and diseased structures. The median filter is a nonlinear filter that is effective in eliminating salt and pepper noise as well as Gaussian noise. Hence, the proposed model employs the approach; it aids in maintaining the image's clarity while eliminating noise. The effectiveness of the median filter is contingent upon the magnitude of the windowing.

In this context, considered image with dimensions of 255 × 255. The window size can be either 3 × 3 or 5 × 5. The first stage involves determining the median value. The typical median process involves organizing the pixel values in ascending order, then selecting the middle value as the median value, denoted as ni,j. The second stage involves determining the disparity between the median value ni,j and the current pixel value curri,j. In the third stage, the initial threshold value is established, and thereafter, the difference value is compared to the threshold value. The calculation of the difference between each pixel value in the window and the center pixel value is performed if the difference is less than the threshold value. The minimum value, denoted as a, is compared to the minimum threshold value. When the minimum value is lower than the threshold value, the current pixel value is deemed to be devoid of noise and is preserved in its original state. Alternatively, if the current pixel is deemed noisy, it is substituted with the median value.

3.1.

3.2. Feature Extraction With NasNet‐Large Model

After the completion of preprocessing, the preprocessed images are passed to NasNet‐Large model for feature extraction. The proposed methodology, derived on the existing NasNet‐Large concept, employs dense blocks as a means to enhance computational efficiency. Although the NasNet‐Large, which is based on CNN, has the potential to determine the optimal network design for medical image processing, it suffers from many accuracy‐related concerns. Dense blocks are added to the NasNet‐Large architecture to address the aforementioned issue and increase performance.

Furthermore, enhanced feature reuse facilitates the model's ability to extract a greater amount of relevant data from every input picture. The term “convolutional” refers to the process of doing a dot product between the values from the filters and the values from the input picture. In the context of 3‐dimensional images, it is essential that the depth of the filter be consistently equivalent to the depth of the input volume. Figure 2 demonstrates that the convolutional layer thoroughly examines the whole picture and extracts the most significant information.

FIGURE 2.

FIGURE 2

NasNet‐Large Architecture for feature extraction.

Before entering the pooling layer, the data undergo an activation function. The convolutional layer comprises additional parameters, namely the number of filters, strides, and padding. Strides refer to the distance between the input image and the convolutional layer, or the length of the receptive field. Padding is further divided into subcategories known as valid and same. Padding refers to the practice of maintaining same dimensions for both the input and output. The term “valid” refers to the absence of padding, whereas “padding” denotes the application of padding. The pooling technique used in this network serves the purpose of down‐sampling, reducing the dimensions of the input picture. This reduction facilitates the subsequent input of the image into the fully connected layers, enabling the classification operation to be executed more efficiently. The dataset also has pairs of reduction cells and normal cells. Reduction cells refer to cells that produce a feature map with a 2‐fold decrease in height and breadth, while normal cells refer to cells that produce a feature map with the same dimensions.

Assuming that the class label of the output image × corresponds to the preY label. The NasNet‐Large model is represented by a function fX that maps the input image to its matching class label:

Z=fx (1)

The parameters of this function are determined by a training procedure that minimizes a loss function LZ,Y′, which measures the difference between the predicted output Z′ and the actual output Z.

The loss function LZ,Z′ is generated using a logarithmic function, which imposes a higher penalty on faulty predictions compared to right ones. To provide more elaboration, the calculation of the difference between the observed label Zi and the expected label Zi′s is performed for every class label i in the output vector.

If the forecast is correct Zi=1, the difference is zero. Otherwise, it is positive and follows an inverse relationship with the prediction error. The aggregate disparity between Z and Z′ may be determined by summing these disparities over all class labels and applying the negative logarithm to the result. Mathematically, this may be expressed as.

LZ,Z′=∑mi=0ZilogZi′ (2)

This formula calculates the sum of all i class labels in the output vector, using Z as the actual label vector and Z′ as the predicted label vector. The logarithmic function improves feature prediction by increasing loss exponentially when predicted probabilities deviate from actual labels. The direct utilization of the full NasNet‐Large feature space may result in Over fitting and increased computational complexity during training. An Archimedes Optimization Algorithm is therefore incorporated for choosing the most informative subset of features while discarding redundant dimensions, hence allowing for more compact and disease‐specific representations.

3.3. Hyper Parameter Tuning With Archimedes Optimization Algorithm (AOA)

The proposed study employs the Archimedes Optimization Algorithm (AOA) to tune different parameters in feature extraction and also helps to select the relevant features as shown in Figure 3. The AOA draws its roots from the physical principle of buoyancy that guarantees a resilient foundation for optimization procedures on complex, multidimensional search spaces. With a strong balance of exploration, moving through varied solutions, and exploitation improving known solutions, AOA provides a general‐purpose methodology for traversing complex optimization problems. Inspired by the principle of Archimedes, AOA utilizes a fitness landscape‐directed mechanism for searching to ensure fast convergence to optimal solutions. Its inherent flexibility allows it to nimbly traverse local minima, and as a result, it proves especially capable within real‐world applications across many fields and domains. The ability of AOA in feature selection is one of its strongest aspects, whereby it can detect and separate crucial features from large datasets. Through dynamic adaptation of its search strategy to the individual properties of the dataset, AOA significantly decreases computational costs, also increases model interpretability and generalization performance. This dual advantage situation renders AOA an effective instrument for practical problem‐solving in areas ranging from ML to data mining and beyond.

FIGURE 3.

FIGURE 3

Flowchart of Archimedes Optimization Algorithm.

Figure 3 depicts the flows included in the AOA as follows:

Step 1—Initialise population location, volume, density, and acceleration using the following mathematical expression (Equations (3), (4), (5), (6)),

Xi=lowi+rand×uppi−lowi;i=1,2,…N (3)
ACCELi=lowi+rand×uppi−lowi;=1,2,…N (4)
DENSITYi=randA,B (5)
VOLUMEi=randA,B (6)

Here, A indicates population number and B represents search range dimension. The object Xi is the ith element in the N population. The search range is defined by two lower and higher limits, denoted as lowi and uppi, respectively. The notation randA,B represents A×B dimensional matrix that may be determined at random by the system function. The i th object's volume, density, and acceleration are, respectively, VOLUMEi, DENSITYi, and ACCELi. The individual Xbest with the highest fitness value is then selected, along with their corresponding ACCELbest, DENSITYbest, and VOLUMEbest.

Step 2—The density and volume of the t+1 th iteration of the i th object are to be enhanced as represented below:

DENSITYit+1=DENSITYit+rand×DENSITYbest−DENSITYit (7)
VOLUMEit+1=VOLUMEit+rand×VOLUMEbest−VOLUMEit (8)

The global optimal values of volume and density are represented by the terms DENSITYbest, and VOLUMEbest, respectively, which are expressed in Equations (7) and (8).

Step 3—Calculate the density decline factor DEC and the parameter DEC that allow the AOA technique to balance its capacity to converge globally and locally.

TF=expt−tmaxtmax (9)
Bt+1=expt−tmaxtmax−ttmax (10)

The current iterations and the maximum are indicated by tmax and t, respectively, in Equation (9). As the number of iterations increases, TF grows until TF=1. While Equation (10) demonstrates that as the number of iterations increases, B decreases and the search is transferred to the identified confined region.

Step 4—When TF ≤ 0.5 results in object exploration and collision. The following Equation (11) provides acceleration:

ACCELit+1=DENSITYmr+VOLUMEmr×ACCELmrDENSITYit+1+VOLUMEit+1 (11)

As per Equation (11), the acceleration, volume, and density of the i th individual at the (t + 1)th iteration are cit+1, VOLUMEit+1 and DENSITYit+1 correspondingly. The cit+1, VOLUMEit+1 and DENSITYit+1 of the random individuals are denoted by ACCELmr, DENSITYmr and VOLUMEmr accordingly.

On the other hand, when TF > 0.5, the exploitation stage occurs and there is no collision between the objects. The update of the acceleration is denoted below:

ACCELi,normt+1=u×ACCELit+1−minACCELmaxACCEL−minACCEL+l (12)

In Equation (12), the normalization range is upp and the fixed value is low at 0.9 and 0.1, respectively. The percentage change of each agent increment is denoted as ACCELi,normt+1. When object i deviates significantly from the global optima, the value of ACCELi,normt+1 could be elevated, indicating that the object is in the exploration stage.

Step 5—When the value of TF ≤ 0.5, the population X's position is adjusted using the calculation shown below:

Xit+1=Xit+C1×rand×ACCELi,normt+1×B×Xrand−Xit (13)

In Equation (13) represents the constant C1, which is equal to 2. Alternatively, in case where TF > 0.5, the population X's position is modified as expressed in Equation (14).

Xit+1=Xbestt+F×C2×rand×ACCELi,normt+1×B×T×Xbest−Xit (14)

Here, C1 is a constant that is equal to 6. T=C3×TF;T increases slowly over time. The parameter F alters the direction of the movement and is assessed using Equation (15).

P=2×rand−C4 (15)
F=+1,ifP≤0.5−1,ifP>0.5 (16)

where the variables C3 and C4 are used to alter the model's capabilities and avoid local optima by balancing the direction of the motions.

Step 6—Evaluation: The individual exhibiting the highest fitness level, as assessed through their acceleration, density, and volume, is selected according to the updated population data. This process is reiterated until the maximum iteration limit is reached. The Fitness Function (FF) within the AOA FS methodology considers both the classification outcomes and the quantity of features selected. This procedure effectively minimizes the number of features while simultaneously improving the precision of the classification results. Consequently, the FF serves as a metric for evaluating the individual solutions:

fitness=α.error_rate+1−β*selec_featall_feat (17)

Equation (17) deliberates the classifier error rate derived from designated features. This error rate quantifies the ratio of incorrect classifications to the overall number of classifications, falling within the range of [0,1]. The variable selec_feat indicates the count of selected features, while all_feat denotes the features of the original dataset. Furthermore, β influences the relevance of categorization quality and the length of the subset. Although AOA reduces the feature redundancy, the retained features from heterogeneous histopathological images, which often suffer from staining variations and scanner noise. Standard classifiers operate on fixed boundaries and may not fully capture the underlying nonlinear manifold structure of these optimized deep features. Therefore employed Stacked Sparse Denoising Autoencoder (SSDA) is employed which Learns hierarchical nonlinear representations, enabling better separation between morphologically similar tissue classes.

3.4. Proposed Stacked Sparse Denoising Autoencoder for Classification

After the successful completion of feature selection with optimized parameter tuning, the proposed Stacked Sparse Denoising Autoencoder (SSDA) classifier known with the aim to compress the size of the data by transforming the input from one form to another appropriate form. This study applies the SSDA model, which learns non‐linear transformations using many layers and a non‐linear activation function. In addition, it works well for eLearning several tiers using an autoencoder (Al‐Rawi et al. 2022). The AE is divided into two parts, a decoding component wφBφ that maps code CD to a rebuilt dataset β, and an encoding part wβBβ that maps input φ to code CD. It can be expressed as per the Equation (18):

φ→CD→β (18)

Since the output, β, is equal to the input, φ, the encoded portion has weight wφ and bias Bβ, while the decoded portion has weight wβ and bias Bβ.

CD=LOGsigwφφ+Bφ (19)
β=LOGsigwβCD+Bβ (20)

The log‐sigmoid function is indicated by LOGsig, and the output β is an approximation of the input φ derived from the equation:

LOGsigx=11+exp−x (21)

An AE version is the Sparse Denoising Autoencoder (SDA). To minimize the errors between the input vector φ and the output β, the following is assumed for the loss function of AE:

IAEwφwβBφBβ=1NUM_INSφ−β2 (22)

In Equation (22), the variable NUM_INS represents the quantity of training examples. The output β is derived by using the Equations (19) and (20):

β=Abs_AEβwβ,wφ,Bβ,Bφ (23)

In Equation (21) Abs_AE indicates the abstract of AE; Thus, Equation (22) is written as follows:

IAEwβwφBβBφ=1NUM_INSAbs_AEβwβ,wφ,Bβ,Bφ2 (24)

The L2 regularization term of the weight wβwφ and τs regularization terms of the sparsity constraints are chosen to avoid trivial or overcomplete mapping:

ISDAwβwφBβBφ=1NUM_INSABS_AEβwβ,wφ,Bβ,Bφ−β2+CS×τs+Cw×τω (25)

The sparsity regulation factor is denoted as CS in Equation (25), while the weighted regulation factors are represented as Cw. The determination of the sparsity regularization term is based on:

τs=∑j=1cdivekullμ−μ′=∑j=1cμlogμμ′+1−μlog1−μ1−μ′6 (26)

Equation (25) represents the Kullback–Leibler divergence function, where divekull represents the parameter |c|, which represents the element count of the internal code output c. The variable μ′ represents the average activation value on the j th sample of NUM_INS, while μ represents the selected value, known as the sparsity percentage factor. The determination of the weight regularization term τωis based on the following:

τω=12×wβwφ2 (27)

The flatten layer reduces the previous layer's output to one dimension. A single dropout layer with a 30% dropout mechanism after the flatten layer prevents overfitting. Finally, the output layer has a dense softmax activation layer. The model is trained using the “Adam” optimizer and category cross entropy loss function for binary classification. The following Algorithm 2 provides the steps involved in the oral cancer classification task.

3.4.

As shown in Algorithm 2, the proposed SSDA‐AOA model performs the classification task to identify specific oral cancer types such as lymphoma, mucosal melanoma, sarcomas, and squamous cell carcinoma. By using a set of selected features by employing the optimized NasNet‐Large model, the process begins by initializing the parameters (weights and biases) for both the encoding and decoding components of the neural network. The algorithm then computes the output of a logarithmic sigmoid (LOG_sig) activation function with respect to these weights and biases, updating the selected feature set accordingly. Through iterative optimization, the model minimizes the classification error across all instances, guided by a defined loss function, and incorporates a regularization term to prevent overfitting, as well as a sparsity regularization to encourage the use of only the most relevant features. If the number of instances matches the selected feature set, the process concludes; otherwise, the sparsity regularization is increased to further refine feature selection. The classified features are then passed through a flattening layer, and the final output layer uses a Softmax activation function to assign probabilities to each cancer type, enabling multi‐class classification. This approach leverages standard neural network training techniques such as weight and bias updates via backpropagation, error minimization, and regularization to achieve robust and accurate cancer type prediction.

4. Results and Discussion

This section presents the major findings of this research and discusses their meaning in the context of the research goals. Data are examined to point out significant trends, patterns, and irregularities. Comparisons are given with current literature in order to put the results into context. Implications of these findings are explored, offering suggestions for future research and applied practice.

4.1. Dataset Description

The histopathology dataset in this repository contains a total of 1224 photos (Rahman et al. 2020). There are two distinct sets of images, each characterized by a unique resolution. The first collection of images comprises 439 histological images depicting oral squamous cell carcinoma (OSCC) at a magnification of 100×, with 89 images illustrating the normal epithelium of the oral cavity. The second set consists of 495 histological images of OSCC at a magnification of 400×, as well as 201 pictures depicting the normal epithelium of the oral cavity. The Leica ICC50 HD microscope was used to capture pictures of H&E stained tissue slides obtained from a sample of 230 individuals. These slides were then analyzed and categorized by medical professionals. The dataset is split as Training: Validation: Testing (70:15:15). The dataset was initially labeled by a board‐certified oral pathologist with more than 10 years of diagnostic experience. Table 1 shows the number of images available in the dataset and Figure 4 indicates the few sample images from the dataset.

TABLE 1.

No. of images in the histopathology dataset.

Type Division No. of images
100× magnification Normal 89
OSCC 439
400× magnification Normal 201
OSCC 495

FIGURE 4.

FIGURE 4

Some images from the Dataset (a) 400× magnification (b) 100× magnification.

Dataset link: https://www.kaggle.com/datasets/ashenafifasilkebede/dataset.

Figure 4 presents representative histological micrographs illustrating tissue architecture and cellular morphology at two different magnifications. Figure 4a shows high‐magnification images that highlight detailed cellular features, including clearly defined cell boundaries, nuclear morphology, and patterns of cellular alignment. Polygonal cells with centrally located nuclei are evident in some regions, indicating preserved cellular integrity, while other areas display elongated, spindle‐shaped cells arranged in parallel bundles, suggestive of organized fibrous or connective tissue components. In contrast, Figure 4b presents lower magnification images that emphasize overall tissue architecture, including layering patterns and regional variations in cellular density, thereby providing a broader structural context.

4.2. Experimental Setup

The experimental setup of AOA and SSDA Training Parameters are given here:

  • Population size N = 30

  • Maximum iterations (tmax) = 100

  • Constant C 1 = 2,C 2 = 6,C 3 = 0.8,C 4 = 1s

  • Fitness weighting parameters: ∝=0.6,β=0.4

  • Random initialization randA,B implemented using uniform distribution

  • Number of stacked layers = 3

  • Hidden units per layer: 512 → 256 → 128

  • Learning rate = 0.001

  • Optimizer—Adam optimizer

  • Batch size = 32

  • Dropout rate after flatten layer = 0.3

  • Maximum training epochs = 50

  • Activation function: Log‐sigmoid, consistent with Equations ((19), (20), (21))

To ensure optimal model performance, first AOA was employed to optimize key training and architectural hyperparameters of the SSDA model. In this stage, each candidate solution in the AOA population encodes a set of hyperparameters, including the learning rate, number of hidden units per layer, dropout rate, and weighting coefficients of the fitness function. The population size was set to 30, and the optimization process was iterated for a maximum of 100 iterations to ensure sufficient exploration and exploitation of the search space. The AOA control constants C1=2,C2=6,C3=0.8,C4=1 were selected based on recommendations from the original AOA formulation to balance global exploration and local exploitation. Candidate solutions were randomly initialized using a uniform distribution to avoid bias in the initial search.

The fitness function, defined using weighted performance metrics ∝=0.6,β=0.4, guided the optimization toward configurations that balance accuracy and generalization. For each candidate solution, the SSDA model was trained and evaluated on the validation set, and the resulting performance was used to update the AOA population. The final network architecture (three stacked layers with 512, 256, and 128 hidden units), learning rate (0.001), Adam optimizer, batch size (32), dropout rate (0.3), and maximum of 50 training epochs were selected based on the best‐performing AOA solutions and further refined through empirical validation to ensure convergence stability and prevent overfitting. The log‐sigmoid activation function was adopted to maintain consistency. The model performance was evaluated with various metrics which was described in the following section.

4.3. Performance Metrics

The study's analysis results are assessed using the following metrics: F‐score, Accuracy, Specificity, and Sensitivity. The confusion matrix was utilized in the computation of these metrics. A confusion matrix is a performance measuring tool that is used to assess a model's accuracy in machine learning and classification jobs. This report provides an overview of the classification algorithm's performance by presenting the frequencies of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) predictions generated by the model on a given dataset.

Accuracy: The metric compares the proportion of accurately categorized samples (positives and negatives) to the total number of samples (number of classified samples).

Accuracy=TrP+TrNTrP+TrN+FP+FN (27)
Precision=TrPTrP+FP (28)

Sensitivity: The measurement is conducted to determine the proportion of positive samples that are correctly classified:

Sensitivity=TrPTrP+FN (29)

Specificity: The metric assesses the accuracy of accurately classifying negative samples.

Specificity=TrNTrN+FP (30)

F1‐score: It computes the disparity between the overall count of accurate positive results and the sum of accurate positive results plus the number of incorrect positive results:

F1−score=2TrP2TrP+FP+FN (31)

ROC curve: A graph which illustrates the performance of a classification model at various categorization levels.

where,

True positives TrP: In these circumstances, the diagnosis was made as to whether the patient had the illness, and it is confirmed.

True negatives TrN: we made a negative prediction, and they do not possess the illness.

False positives FP: although we made a prediction, they do not really possess the illness.

False negatives FN: The null hypothesis was made, nevertheless, the illness is indeed present.

4.4. Performance Analysis

The confusion matrix and the ROC curve of the proposed SSDA_AOA method are illustrated in Figures 6 and 7.

FIGURE 6.

FIGURE 6

Outcome of the ROC curve.

FIGURE 7.

FIGURE 7

Comparative analysis with the proposed SSDA‐AOA with existing SSDA Model.

Figure 5 depicts the results of the prediction made by the proposed model. It performs binary classification as either normal or oral SCC, which demonstrates strong diagnostic performance. The proposed framework correctly identified 280 cases of OSCC and 71 normal cases, while only misclassifying 13 normal cases as oral SCC and 4 oral SCC cases as a normal. These results yield high accuracy of 92.58%, precision of 95.6%, recall of 98.6%, and F1 score of 97.1, indicating the model is highly effective at both detecting cancer cases and minimizing missed diagnoses.

FIGURE 5.

FIGURE 5

Confusion matrix of SSDA_AOA model.

Figure 6 denotes the ROC curve of the SSDA‐AOA model, plots of TPR against FPR at various threshold levels. It perfectly distinguishes between the two classes' normal and Oral SCC. It has an archived AUC value of 0.9879 for both classes, which is close to the ideal value of 1. This indicates that the model can almost perfectly distinguish between normal and cancer cases across all classification thresholds, maintaining high sensitivity and a low false positive rate. Besides, high AUC values confirm the model's robustness and reliability, making it highly suitable for clinical applications where accurate and early detection is critical.

4.5. Comparative Analysis

The sub section compares the efficiency of the proposed model with the existing model to showcase the robustness and reliability.

Figure 7 and Table 2 depict the comparative study between the proposed and existing models. The existing system, employing the SSDA method, achieves commendable results with an accuracy of 87.83%, a precision of 88.92%, sensitivity of 88.92%, specificity of 85.31%, and an F‐score of 70.02%. However, due to the presence of the optimization method, the proposed system, SSDA_AOA method, showcases significant enhancements across all metrics. With an impressive accuracy of 95.38%, precision of 95.15%, sensitivity of 91.78%, specificity of 91.85%, and an F‐score of 93.72%, the proposed system outperforms the existing method.

TABLE 2.

Comparison analysis of the proposed model with existing SSDA models.

System Methods Accuracy (%) Precision (%) Sensitivity (%) Specificity (%) F1‐score (%)
Existing system SSDA 87.83 88.92 88.92 85.31 70.02
Proposed system SSDA‐AOA 95.38 95.15 91.78 91.85 93.72

The Table 3 shows the comparative analysis of the proposed methods with different DL methods and Figure 8 represents the graphical representation of comparison evaluation. Oral cancer screening plays a vital role in identifying early signs of oral SCC and its precursors, oral potentially malignant disorders (OPMDs). General practitioners are often the first point of contact for patients with oral lesions; however, their limited knowledge in this area can hinder timely referrals and treatment. Such delays can result in more invasive procedures and significantly impact patients' quality of life, as survival rates drop below 50% when diagnosed at advanced stages. While current screening primarily relies on visual examination, advancements in AI, particularly in image analysis, offer promising potential to enhance diagnostic accuracy and improve patient outcomes in oral cancer detection.

TABLE 3.

Comparative performance analysis with existing methods with optimizations.

References Methods Accuracy (%) Precision (%) Sensitivity (%) Specificity (%) F1‐score (%)
Welikala et al. (2020) ResNet‐101 81.37 84.77 89.51 62.29 87.07
Mira et al. (2024) HRNet‐W18 84.3% 85.47% 83% 96% 83.6%
Tanriver et al. (2021) DenseNet‐161 85.21%% 87.9% 84.1% 91.23% 84.4%
Das et al. (2023) Alexnet 88% 88% 89% 88% 89%
Warin et al. (2021) Faster RCNN 92.1% 76.67 82.14 79 79.31
Soni et al. (2024) EfficientNetB0 91.1 91.3 92.2 91.0 92.3
Proposed method SSDA‐AOA 95.38 95.15 91.78 91.85 93.72

FIGURE 8.

FIGURE 8

Comparative analysis with the proposed SSDA_AOA with existing models.

Accordingly, Welikala et al. (2020) utilized ResNet‐101, achieved an accuracy of 81.37%. Likewise, Mira et al. (2024) implemented HRNet‐W18, obtained an accuracy of 84.3%. Tanriver et al. (2021) employed DenseNet‐161, achieving an accuracy of 85.21%. Das et al. (2023) utilized Alexnet, attained a moderate accuracy of 88%. Additionally, Warin et al. (2021) achieved an accuracy of 92.1% by employing Fater RCNN model for classification. Similarly, Soni et al. (2024), developed a detection model with EfficientNetB0 and yielded the 91.1% of accuracy in classification task. While comparing the efficiency and robustness of the proposed SSDA_AOA method with existing state‐of art models, the proposed model is demonstrated superior accuracy than existing methods, highlighting its effectiveness in image classification tasks. The proposed method, SSDA‐AOA, showcased a remarkable accuracy of 95.38%, significantly outperforming other methods. It emphasized its importance in clinical diagnosis assistance in detecting Oral SCC and significantly enhance the well‐being of the patient. The current framework depends on expert‐annotated datasets from a few clinical centers, which may raise concerns over generalizability across institutions. Whole‐slide image analysis is not yet incorporated into the study instead, it focuses on patch‐level classification.

5. Conclusion and Future Direction

This research introduced a novel robust approach for the early detection of oral cancer by integrating advanced preprocessing methods with an improved NasNet‐Large‐based DL architecture. The use of median filtering effectively suppressed noise and highlighted lesion regions, thereby enhancing the quality of input data for subsequent analysis. The proposed methodology outperformed existing techniques, as evidenced by comprehensive comparative analysis, and demonstrated notable improvements in diagnostic accuracy for oral cancer. The proposed model yielded accuracy of 95.38, sensitivity of 91.78%, precision of 95.15%, specificity of 91.85%. These findings underscore the potential of the developed system to assist clinicians in making timely and reliable diagnoses, ultimately contributing to better patient outcomes and increased survival rates. Additionally, future research will focus on expanding the dataset to include a wider variety of clinical images representing all known oral lesion variants, thereby improving the model's generalizability and robustness. Additionally, the implementation of ensemble neural network strategies will be explored, leveraging the complementary strengths of multiple deep learning models to enhance feature extraction and classification performance.

Author Contributions

R. Sathish Kumar: methodology, data curation, validation, supervision. M. Govindarajan: methodology, data curation.

Funding

The authors have nothing to report.

Ethics Statement

The authors have nothing to report.

Conflicts of Interest

The authors declare no conflicts of interest.

Acknowledgements

The authors have nothing to report.

Data Availability Statement

Data sharing is not applicable to this article as no datasets were generated.

References

  1. Adhit, K. K. , Wanjari A., and Menon S.. 2023. “Liquid Biopsy: An Evolving Paradigm for Non‐Invasive Disease Diagnosis and Monitoring in Medicine.” Cureus 15, no. 12: e50176. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Afify, H. M. , Mohammed K. K., and Hassanien A. E.. 2023. “Novel Prediction Model on OSCC Histopathological Images via Deep Transfer Learning Combined With Grad‐CAM Interpretation.” Biomedical Signal Processing and Control 83: 104704. [Google Scholar]
  3. Al‐Rawi, N. , Sultan A., Rajai B., et al. 2022. “The Effectiveness of Artificial Intelligence in Detection of Oral Cancer.” International Dental Journal 72, no. 4: 436–447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Ananthakrishnan, B. , Shaik A., Kumar S., Narendran S., Mattu K., and Kavitha M. S.. 2023. “Automated Detection and Classification of Oral Squamous Cell Carcinoma Using Deep Neural Networks.” Diagnostics 13, no. 5: 918. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Babu, P. A. , Rai A. K., Ramesh J. V. N., et al. 2024. “An Explainable Deep Learning Approach for Oral Cancer Detection.” Journal of Electrical Engineering & Technology 19, no. 3: 1837–1848. [Google Scholar]
  6. Bansal, K. , Bathla R., and Kumar Y.. 2022. “Deep Transfer Learning Techniques With Hybrid Optimization in Early Prediction and Diagnosis of Different Types of Oral Cancer.” Soft Computing 26, no. 21: 11153–11184. [Google Scholar]
  7. Bharanidharan, N. , Abhinav K. K., Lathvik C. K., Chethan M., Reddy M. M. K., and Deepak K.. 2024. “Feature Extraction Using Hybridized Transfer Learning Approach for Oral Cancer Diagnosis.” In 2024 5th International Conference on Smart Electronics and Communication (ICOSEC), 975–979. IEEE. [Google Scholar]
  8. Bouvard, V. , Nethan S. T., Singh D., et al. 2022. “IARC Perspective on Oral Cancer Prevention.” New England Journal of Medicine 387, no. 21: 1999–2005. [DOI] [PubMed] [Google Scholar]
  9. Chang, X. , Yu M., Liu R., et al. 2023. “Deep Learning Methods for Oral Cancer Detection Using Raman Spectroscopy.” Vibrational Spectroscopy 126: 103522. [Google Scholar]
  10. Dar, G. M. , Agarwal S., Kumar A., et al. 2022. “A Non‐Invasive miRNA‐Based Approach in Early Diagnosis and Therapeutics of Oral Cancer.” Critical Reviews in Oncology/Hematology 180: 103850. [DOI] [PubMed] [Google Scholar]
  11. Das, M. , Dash R., and Mishra S. K.. 2023. “Automatic Detection of Oral Squamous Cell Carcinoma From Histopathological Images of Oral Mucosa Using Deep Convolutional Neural Network.” International Journal of Environmental Research and Public Health 20, no. 3: 2131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Dixit, S. , Kumar A., and Srinivasan K.. 2023. “A Current Review of Machine Learning and Deep Learning Models in Oral Cancer Diagnosis: Recent Technologies, Open Challenges, and Future Research Directions.” Diagnostics 13, no. 7: 1353. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Hirra, I. , Ahmad M., Hussain A., et al. 2021. “Breast Cancer Classification From Histopathological Images Using Patch‐Based Deep Learning Modeling.” IEEE Access 9: 24273–24287. [Google Scholar]
  14. Huang, Q. , Ding H., and Razmjooy N.. 2023. “Optimal Deep Learning Neural Network Using ISSA for Diagnosing the Oral Cancer.” Biomedical Signal Processing and Control 84: 104749. [Google Scholar]
  15. Jubair, F. , Al‐karadsheh O., Malamos D., Al Mahdi S., Saad Y., and Hassona Y.. 2022. “A Novel Lightweight Deep Convolutional Neural Network for Early Detection of Oral Cancer.” Oral Diseases 28, no. 4: 1123–1130. [DOI] [PubMed] [Google Scholar]
  16. Khan, S. U. R. , and Asif S.. 2024. “Oral Cancer Detection Using Feature‐Level Fusion and Novel Self‐Attention Mechanisms.” Biomedical Signal Processing and Control 95: 106437. [Google Scholar]
  17. Kumari, P. , Debta P., and Dixit A.. 2022. “Oral Potentially Malignant Disorders: Etiology, Pathogenesis, and Transformation Into Oral Cancer.” Frontiers in Pharmacology 13: 825266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Majeed, T. , Masoodi T. A., Macha M. A., Bhat M. R., Muzaffar K., and Assad A.. 2024. “Addressing Data Imbalance Challenges in Oral Cavity Histopathological Whole Slide Images With Advanced Deep Learning Techniques.” International Journal of System Assurance Engineering and Management 15: 1–9. [Google Scholar]
  19. Meer, M. , Khan M. A., Jabeen K., et al. 2025. “Deep Convolutional Neural Networks Information Fusion and Improved Whale Optimization Algorithm Based Smart Oral Squamous Cell Carcinoma Classification Framework Using Histopathological Images.” Expert Systems 42, no. 1: e13536. [Google Scholar]
  20. Mira, E. S. , Saaduddin Sapri A. M., Aljehanı R. F., et al. 2024. “Early Diagnosis of Oral Cancer Using Image Processing and Artificial Intelligence.” Fusion: Practice & Applications 14, no. 1: 293–308. [Google Scholar]
  21. Myriam, H. , A. Abdelhamid A., El‐Kenawy E.‐S. M., et al. 2023. “Advanced Meta‐Heuristic Algorithm Based on Particle Swarm and Al‐Biruni Earth Radius Optimization Methods for Oral Cancer Detection.” IEEE Access 11: 23681–23700. [Google Scholar]
  22. Ogundokun, R. O. , Owolawi P. A., Tu C., and van Wyk E.. 2025. “Autoencoder‐Assisted Stacked Ensemble Learning for Lymphoma Subtype Classification: A Hybrid Deep Learning and Machine Learning Approach.” Tomography 11, no. 8: 91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Oya, K. , Kokomoto K., Nozaki K., and Toyosawa S.. 2023. “Oral Squamous Cell Carcinoma Diagnosis in Digitized Histological Images Using Convolutional Neural Network.” Journal of Dental Sciences 18, no. 1: 322–329. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Pal, A. , Oshiro S., Verma P. K., et al. 2023. “Oral Cancer Detection at an Earlier Stage.” In International Conference on Computational Electronics for Wireless Communications (ICCWC), 375–384. Springer. [Google Scholar]
  25. Panigrahi, S. , Das J., and Swarnkar T.. 2022. “Capsule Network Based Analysis of Histopathological Images of Oral Squamous Cell Carcinoma.” Journal of King Saud University 34, no. 7: 4546–4553. [Google Scholar]
  26. Qadri, S. F. , Shen L., Ahmad M., Qadri S., Zareen S. S., and Akbar M. A.. 2022. “SVseg: Stacked Sparse Autoencoder‐Based Patch Classification Modeling for Vertebrae Segmentation.” Mathematics 10, no. 5: 796. [Google Scholar]
  27. Ragab, M. , and Asar T. O.. 2024. “Deep Transfer Learning With Improved Crayfish Optimization Algorithm for Oral Squamous Cell Carcinoma Cancer Recognition Using Histopathological Images.” Scientific Reports 14, no. 1: 25348. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Rahman, A.‐u. , Rahman A., Alqahtani A., et al. 2022. “Histopathologic Oral Cancer Prediction Using Oral Squamous Cell Carcinoma Biopsy Empowered With Transfer Learning.” Sensors 22, no. 10: 3833. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Rahman, T. Y. , Mahanta L. B., Das A. K., and Sarma J. D.. 2020. “Histopathological Imaging Database for Oral Cancer Analysis.” Data in Brief 29: 105114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Rashid, J. , Qaisar B. S., Faheem M., Akram A., Amin R. u., and Hamid M.. 2024. “Mouth and Oral Disease Classification Using InceptionResNetV2 Method.” Multimedia Tools and Applications 83, no. 11: 33903–33921. [Google Scholar]
  31. Raval, D. , and Undavia J. N.. 2023. “A Comprehensive Assessment of Convolutional Neural Networks for Skin and Oral Cancer Detection Using Medical Images.” Healthcare Analytics 3: 100199. [Google Scholar]
  32. Razmjouei, P. , Moharamkhani E., Aryanezhad S. S., Shokouhifar M., Hosseinzadeh M., and Zadmehr B.. 2025. “NFR‐EDL: Non‐Linear Fuzzy Rank‐Based Ensemble Deep Learning for Accurate Diagnosis of Oral and Dental Diseases Using RGB Color Photography.” Computers in Biology and Medicine 192: 110279. [DOI] [PubMed] [Google Scholar]
  33. Sampath, P. , Sasikaladevi N., Vimal S., and Kaliappan M.. 2024. “OralNet: Deep Learning Fusion for Oral Cancer Identification From Lips and Tongue Images Using Stochastic Gradient Based Logistic Regression.” Network Modeling Analysis in Health Informatics and Bioinformatics 13, no. 1: 24. [Google Scholar]
  34. Saraswathi, T. , and Murali Bhaskaran V.. 2025. “Heuristic Strategy Using Hybrid Deep Learning With Transfer Learning for Oral Cancer Detection.” International Journal of Wavelets, Multiresolution and Information Processing 23, no. 1: 2450039. [Google Scholar]
  35. Shah, R. , and Pareek J.. 2024. “Non‐Invasive Primary Screening of Oral Lesions Into Binary and Multi Class Using Convolutional Neural Network, Stratified K‐Fold Validation and Transfer Learning.” Indian Journal of Science and Technology 17, no. 7: 651–659. [Google Scholar]
  36. Shah, S. J. H. , Albishri A., Wang R., and Lee Y.. 2025. “Integrating Local and Global Attention Mechanisms for Enhanced Oral Cancer Detection and Explainability.” Computers in Biology and Medicine 189: 109841. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Sharma, D. , Kudva V., Patil V., Kudva A., and Bhat R. S.. 2022. “A Convolutional Neural Network Based Deep Learning Algorithm for Identification of Oral Precancerous and Cancerous Lesion and Differentiation From Normal Mucosa: A Retrospective Study.” Engineered Science 18, no. 15: 278–287. [Google Scholar]
  38. Smędra, A. , and Berent J.. 2023. “The Influence of the Oral Microbiome on Oral Cancer: A Literature Review and a New Approach.” Biomolecules 13, no. 5: 815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Soni, A. , Sethy P. K., Dewangan A. K., Nanthaamornphong A., Behera S. K., and Devi B.. 2024. “Enhancing Oral Squamous Cell Carcinoma Detection: A Novel Approach Using Improved EfficientNet Architecture.” BMC Oral Health 24, no. 1: 601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Srinivasulu, A. , Bohara B., and Kumar A. S.. 2025. “Enhancing Throat Cancer Prediction and Detection: A Hybrid CNN‐RNN Approach to Mitigate Common Deep Learning Model Drawbacks.” In Spatially Variable Genes in Cancer: Development, Progression, and Treatment Response, 293–314. IGI Global Scientific Publishing. [Google Scholar]
  41. Sukegawa, S. , Ono S., Tanaka F., et al. 2023. “Effectiveness of Deep Learning Classifiers in Histopathological Diagnosis of Oral Squamous Cell Carcinoma by Pathologists.” Scientific Reports 13, no. 1: 11676. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Tanriver, G. , Soluk Tekkesin M., and Ergen O.. 2021. “Automated Detection and Classification of Oral Lesions Using Deep Learning to Detect Oral Potentially Malignant Disorders.” Cancers 13, no. 11: 2766. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Umapathy, V. R. , Natarajan P. M., Swamikannu B., et al. 2022. “Emerging Biosensors for Oral Cancer Detection and Diagnosis—A Review Unravelling Their Role in Past and Present Advancements in the Field of Early Diagnosis.” Biosensors 12, no. 7: 498. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Vaiyapuri, T. , Alaskar H., Syed L., et al. 2022. “Modified Metaheuristics With Stacked Sparse Denoising Autoencoder Model for Cervical Cancer Classification.” Computers and Electrical Engineering 103: 108292. [Google Scholar]
  45. Varalakshmi, D. , Tharaheswari M., Anand T., and Saravanan K. M.. 2024. “Transforming Oral Cancer Care: The Promise of Deep Learning in Diagnosis.” Oral Oncology Reports 10: 100482. [Google Scholar]
  46. Warin, K. , Limprasert W., Suebnukarn S., Jinaporntham S., and Jantana P.. 2021. “Automatic Classification and Detection of Oral Cancer in Photographic Images Using Deep Learning Algorithms.” Journal of Oral Pathology & Medicine 50, no. 9: 911–918. [DOI] [PubMed] [Google Scholar]
  47. Warin, K. , Limprasert W., Suebnukarn S., Jinaporntham S., and Jantana P.. 2022. “Performance of Deep Convolutional Neural Network for Classification and Detection of Oral Potentially Malignant Disorders in Photographic Images.” International Journal of Oral and Maxillofacial Surgery 51, no. 5: 699–704. [DOI] [PubMed] [Google Scholar]
  48. Warin, K. , Limprasert W., Suebnukarn S., Jinaporntham S., Jantana P., and Vicharueang S.. 2022. “AI‐Based Analysis of Oral Lesions Using Novel Deep Convolutional Neural Networks for Early Detection of Oral Cancer.” PLoS One 17, no. 8: e0273508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Welikala, R. A. , Remagnino P., Lim J. H., et al. 2020. “Automated Detection and Classification of Oral Lesions Using Deep Learning for Early Detection of Oral Cancer.” IEEE Access 8: 132677–132693. [Google Scholar]
  50. Yang, G. , Wei L., Thong B. K. S., et al. 2022. “A Systematic Review of Oral Biopsies, Sample Types, and Detection Techniques Applied in Relation to Oral Cancer Detection.” Biotech 11, no. 1: 5. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing is not applicable to this article as no datasets were generated.


Articles from Oral Diseases are provided here courtesy of Wiley

RESOURCES