Abstract
A skin lesion, in medical terminology, refers to any abnormality on the skin’s surface. Skin lesions, both benign and malignant, can progress to skin cancer, highlighting the need of early identification. The scarcity of dermatological resources, particularly in rural regions, contributes to delayed diagnosis. Existing diagnostic techniques, which heavily rely on the expertise of dermatologists, are plagued by disparities in global healthcare access. Skin lesion classification plays a vital role in early intervention and timely treatment of various dermatological conditions. In this paper, a novel and optimal method to classify different type of skin lesions is proposed, by combining deep learning with optimization techniques. The proposedalgorithm uses the robust RegNetY032 model with a modified classification head as the backbone architecture to perform feature extraction and classification from the dermoscopy images, integrated with the Soft Attention Block to effectively discern and prioritize salient lesion features while disregarding artifacts such as hair and veins commonly present on the skin. Furthermore, Harris-Hawks Optimization is employed to perform hyperparame- ter optimization, enhancing the model’s classification performance. Experimental results on the HAM10000 benchmark dataset demonstrate the robustness of the proposed model, achieving superior classification accuracy of 99.27%. The proposed HHO-Reg-SA-Net holds promise for advancing automated dermatological diagnosis systems, contributing to an improved patient care and early intervention in skin cancer diagnosis and treatment.
Keywords: Skin lesion classification, Dermoscopy, Class imbalanced, RegNet, Soft attention, Harris-Hawks optimization
Subject terms: Computer science, Diseases
Introduction
Skin lesion is an area on the skin that differs from the surrounding tissue. Skin lesions are a potential cause of skin cancer. Benign skin lesions are non-cancerous growths on the skin that typically pose no harm. In many cases, these abnormalities, such as moles, birthmarks, cherry angiomas, skin tags, acne and freckles do not require treatment unless they cause discomfort or aesthetic concerns. Malignant skin lesions refer to abnormal growths on the skin that have the potential to become cancerous. The most commonly occurring skin cancer is Melanoma that arises from the pigment- producing cells (melanocytes) in the skin. Basal Cell Carcinoma and Squamous Cell Carcinoma typically develop in the outer layer of the skin. Malignant skin lesions often exhibit changes in colour, shape, size, or texture and may be accompanied by symptoms like itching, tenderness, or bleeding. Early detection is vital for effective treatment and a favourable prognosis. Research jointly conducted by the International Labour Organization (ILO) and the World Health Organization (WHO) reveals a concerning correlation between outdoor work and non-melanoma skin cancer deaths.
Approximately 1 out of every 3 fatalitiesfrom non-melanistic skin cancer is attributed to sun exposure on the job. The study, featured in Environment Interna- tional, underscores the escalating burden on outdoor workers, emphasizing the need for preventive measures to address this critical workplace hazard and prevent further loss of lives1. The combined data illustrates that around 1.5 billion people aged above 15, who were part of the workforce, encountered exposure to solar UV rays due to outdoor work activities in 2019. This represents 28% of the worldwide population within the working age bracket. Shockingly, nearly 19,000 people, from 183 countries around the globe were affected due to non-melanistic skin cancer in the same year due to occupational sun exposure, with males comprising 65% of the victims. These statistics position the exposure to solar ultraviolet radiation as the third-highest work-relatedrisk factor for cancer deaths worldwide. Notably,from 2000 to 2019, skin cancer related deaths attributed to on-the-job solar radiation exposure nearly doubled, experiencing a whooping 88% increase from 2000 to 2019.
Thesefindings underscoretheurgencyof implementing measuresto protect outdoor workers and mitigate the growing impact of occupational sun exposure on global health2. Skin cancer ranks among the most prevalent malignancies world- wide, experiencing a significant rise in incidence over the last decade3. Typically, dermatologists diagnose skin cancer through visual inspection, aided by dermoscopic imaging and confirmed by skin biopsy4. However, delayed skin cancer identification results in increased morbidities and melanoma mortality; these issues are exacerbated by the lack of dermatological resources and pathologylabs in rural clinics5. A dermatologist with advanced training is necessary to analyze potentially malignant lesions with an optical dermatoscope. Globally, there are 3,000,000 non-melanoma cases and 123,000 melanomacases recordedeach year. Annually,the numberof instances of skin cancer globally rises between 3 and 7% due to external factors such as prolonged sun exposure6. This uptick translates to an additional 4,500 melanoma cases each year. Morton and Machie7 highlight the complexity, noting that in melanoma cases, an aggressive skin cancer type, only 60–90% of malignant lesions are successfully detectedthrough visual inspection.Accuracy significantly fluctuates based on the dermatologist’s experience. According to research findings, dermatologists with over a decade of experience demonstrated an 80% accuracy in diagnosing skin lesions. Individuals with three to five years of professional experience correctly recognized 62% of the cases8.
The scarcity of expert dermatologists around the world, particularly those serving various ethnic and economic communities, plays a major role in widening health dis- parities globally. The advent of artificial intelligence, particularly deep learning (DL), has emerged as a pivotal advancement in computer-aided cancer diagnosis9. The application of deep learning algorithms in classification of skin lesions holds promise for automating screening and early detection, addressingthe challenges posed by the paucityof dermatologists and laboratory facilities in rural communities10. Globally, deep learning has gained widespread use in the field of medical imaging, particularly in the challenging task of accurately classifying skin lesions. However, this application comes with certain hurdles. One major obstacle lies in the necessity for high-quality data, a crucial factor for achieving reliable classification outcomes through deep learning methodologies.
Unlike many deep convolutional neural network models that benefit from training on millions of diverse images, the domain of skin lesion analysis presents a constraint with only a limited dataset of a few thousand images or even fewer11. This scarcity hampers the potential for achieving robust classification results. Another complicating factor is the substantial similarity among different lesion classes and the pronounced intra-class variations, posing additional challenges to accurate classification12. These intricacies not only make the visual inspection process more complex but also hinder the feature extraction process by introducing noise into dermoscopic images. Address- ing these issues is critical for increasing the efficiency of deep learning in accurately classifying skin lesions. The aim of the study is:
To develop and train an optimized neural network model that effectively classifies skin lesions despite imbalances in the dataset, leveraging advanced techniques such as soft attention and Harris Hawks optimization.
To integrate the trained model into clinical workflows, enabling dermatologists to efficiently and accurately classify skin lesions, facilitating timely interventions and treatments.
To address limitations in manual interpretation variability by providing a computer vision-based solution that ensures unbiased and consistent results in skin lesion classification.
To enhance classification accuracy for various types of skin lesions, contributing to more precise diagnoses and enabling early treatment, ultimately improving patient survival rates.
The contribution of the paper is as follows:
- A novel and optimal method for classifying skin lesions.
- Utilizing the RegNetY032 neural network as the backbone architecture.
- Integrating the Soft Attention Block for feature extraction and classification of skin lesions.
- EmployingtheHarris-HawksOptimization Algorithm(HHOA)for optimizing hyperparameters.
- The proposed model is named HHO-Reg-SA-Net (Harris-Hawks Optimized RegNet with Soft Attention).
Related work
Recent advances in research and development have focused on leveraging pre-trained models via transfer learning to train, enhance, and fine-tune deep convolutional neural network (DCNN) models, allowing for automatic feature extraction and classification with excellent performance. The various approaches towards skin lesion classification over the years, by various researchers around the world are summarized as follows:
Transfer learning
Pre-trained models that were initially developed for a certain job can be modified and refined for a different but similar task using the transfer learning approach. By utilizing the information that was gained during the first training phase, this method enables the model to generalize and function well in the new setting. Because the approach relies on pre-existing information from the source work, it is particularly useful when labeled data for the target task is limited. Transfer learning accelerates the training process, enhances model performance, and is widely employed across various domains, showcasing its versatility and efficiency in leveraging prior knowledge to tackle novel challenges. Iqbal et al. (2021) developed CSLNet, a Deep Convolutional Neural Network (DCNN) for automatic multi-class categorization of skin lesions using dermoscopic images13,14. Dermoscopic images from the International Skin Imaging Collaboration databases (ISIC-17, ISIC-18, and ISIC-19) are utilized in the research. The proposed CSLNet accurately classified skin lesions with an accuracy of 94%, sensitivity of 93%, and specificity of 91%. Calderonet al.15 proposed a bilinear convolutional neural network (CNN) technique for skin lesion categorization using ResNet50 and VGG16 architectures, including transfer learning and fine-tuning. The computational cost of the approach is also measured,with the training time reported as 238.6 min. Dermoscopic images from the HAM10000 dataset are used for training and testing the model. Muhammad Usama et al.16 worked with retraining of two CNN models, Darknet53 and InceptionV3, via transfer learning to perform feature extraction from the dermoscopic images. Augmentation techniques were used to perform class-balancing and to reduce overfitting of the model, resulting in a significant improvement in performance. Despite its performance, transfer learning in skin lesion classification encounters limitations such as domain mismatch, where the source and target tasks differ significantly, potentially affecting performance. Also, the optimal transferability of knowledge heavily depends on the choice of source and target domains, presenting challenges in achieving consistent improvements across diverse datasets.
Ensemble learning
Ensemble learning combines numerous separate models, known as base learners or weak learners, to produce a more robust and accurate prediction model. Ensemble Learning harnesses the diversity of the base models, often trained on different subsets of data or employing distinct algorithms, to collectively enhance the generalization capabilities and overall performance of the ensemble model. By aggregating the predic- tions or decisions of the constituent models, ensemble learning overcomes the problem of overfitting and minimizes the impact of individual model weaknesses. Using five deep neural network models as the foundation of the ensemble SeResNeXt, ResNet50, ResNeXt, Densenet—169 and Xception. Rahman et al.17 developed a weighted average ensemble learning-based model to categorize various types of skin lesions. Techniques for data augmentation, noise reduction, and class balance were also used. The influence of each base model on the average recall score was maximized by using the grid search approach to identify the optimal base model combinations in the ensemble. Three separate feature extractor modules are fused together to create a hybrid-CNN that Hasan et al.18 suggested to provide lesion feature maps with greater depth. After that, several fully connected layers are used to classify these feature maps, and the results are ensembled to predict a lesion class. In order to classify skin lesions, Mehmood et al.19 introduced a modified deep learning model called SBXception, which is based on the Xception network. The HAM10000 dataset, comprising 10,015 pictures of skin lesions, is used to train and assess the model. To better fit the categorization issue, the depth and breadth of the basic model, Xception, are increased and decreased, respectively. Ensemble learning has difficulties even if it performs better in skin lesion categorization than transfer learning. Choosing the diverse base models demands careful consideration, as not all combinations may yield synergistic improvements. Additionally, ensemble methods may face diminishing returns, where further model additions provide marginal benefits, potentially escalat- ing computational costs. The sensitivity of ensemble performance to the quality and quantity of data presents another hurdle, necessitating robust strategies to address data imbalances and ensure generalizability.
Feature extraction
Feature extraction refers to the process of identifying and selecting relevant information or distinctive characteristics from raw image data. This crucial step aims to transform high-dimensional input, such as pixel values, into a more compact and mean- ingful representation of the image. These extracted features serve as discriminative indicators that capture important patterns, textures, shapes, or structures within the image. This process not only reduces computational complexity but also facilitates a more focused analysis of the image content, enabling the classification algorithm to better differentiate between different classes or categories. An approach for multiclass skin lesion classification has been proposed by Afza et al.20. It consists of five primary phases: contrast enhancement and image acquisition; feature extraction via transfer learning algorithms; optimal feature selection via entropy-mutual infor- mation method and hybrid whale-optimization algorithm;feature fusion by means of a modified canonical correlation method; and classification using extreme learning machine (ELM). Multi-Hierarchy Contrastive Learning, along with Pareto Optimality (MHCL-PO) is a technique that Shuang et al.21 introduced to improve classification performance of the model. The four stages of the MHCL-PO framework are pareto optimality, contrastive learning, feature extraction, and classification. A new segmentation technique called EW-FCM and a deep learning structure called Wide- ShuffleNet were used by Hoang et al.22 to present a novel method for skin lesion categorization. By using Wide-ShuffleNet to perform segmentation and classification, this method aids in separating the skin lesion object from the surrounding environment. A portable healthcare system, for example, might benefit from the use of Wide-ShuffleNet, a lightweight deep learning architecture. Since skin lesions vary widely, it is difficult to design characteristics that are universally helpful. This makes it complex to develop a solution that works for everyone around the globe. Skin lesion categorization feature extraction is limited by the possibility of information loss in the process since reducing raw data to its most basic form may result in the omission of important but subtle nuances.
Deep convolutional neural networks
Withinthedomainof computer vision,DeepConvolutional NeuralNetworks (DCNNs) have become key architectures, exhibitingoutstanding performance in a range of image-related tasks as image segmentation, image classification, and picture captioning. Deep CNNs leverage multiple layers of convolutional operations to automatically extract hierarchical features from input data, allowing them to capture intricate patterns and representations in images. Transfer learning employs a different paradigm by utilizing models that are pre-trained on large datasets for a related task and refining them on a target task with limited data,whereas DCNN excels in end-to-end feature learning from enormous datasets. Thus, DCNN and transfer learning are not the same. The exploration of these nuances is crucial for optimizing model performance and advancing in image understanding and recognition. For the purpose of classifying skin, Zhang et al.23 suggested a unique DCNN architec- ture with a deep attention mechanism. An analysis of the network’s intrinsic self-awareness was conductedusing a deep convolutional neural network (DCNN). Using the attention maps, each ARL block develops residual learning mechanisms to enhance its capacity to categorize incoming data from the ISIC-2017, ISIC-2018 and the ISIC-2019 datasets. Dinga et al.24 built a Deep Attention Branch Network (DABN) model with an Entropy-guided Loss Weighting (ELW) technique to cate- gorize skin lesions. To further enhance the classification performance, the ELW technique raises the loss weights of samples that are prone to misclassification. A deep convolution neural network with weighted loss specific to each class and grouped multi-scale attention blocks (GMAB) was proposed by Qiana et al. in 202225. The model’s capacity to concentrate on the lesion region is improved by using the GMAB to extract multi-scale fine-grained features. A triplet-based network (TPN), which is based on deep-metriclearning and a mixed attention mechanismcomprise the deep-metricattention learning CNN (DeMAL-CNN) that Wang et al.26 sug- gested for skin lesion categorization. To increase the amount of training samples and learn embeddings that are resilient to both intra-class and inter-class changes, TPN takes an input of triplet of samples and employs the triplet loss to optimize the embeddings. The capacity of DeMAL-CNN to localize skin lesions is strengthened by the mixed attention mechanism, which focuses on both channel-wise and spatial-wise attention information. Naeem et. al. (2024) proposed DVFNet, by combining VGG19 architecture andtheHistogramof Oriented Gradients (HOG)featureextractor27. The model addresses key challenges in image-based skin cancer detection, such as image noise and class imbalance. The use of Synthetic Minority Over-sampling Technique with Tomek links (SMOTETomek) to resolve class imbalance issues in the ISIC 2019 dataset sets DVFNetapart, ensuring a more reliable classification performance in a multi-class setting. The authors evaluated their proposed technique using a dataset characterized by class imbalance. While they increased the number of images using borderline SMOTE, it would be beneficial to test the model on a more comprehensive dataset. Naeem et. al.28 proposed SNC Net, a hybrid model for skin cancer detectionthat combines handcrafted and deep learning-based features from dermoscopy images to classify eight types of skin cancer. The model was trained, validated and tested on the ISIC 2019 dataset. It uses a customized convolu- tional neural network SNC Net for classification and the SMOTE Tomek method to address dataset imbalance. The integration of handcrafted and deep learning features, along with Grad-CAM visualization, improves classification accuracy and provides interpretability. Naeem et. al. (2022) proposed SCDNet, combines VGG16 with con- volutional neural networks (CNN) for classification. The SCDNet model is trained, validated and tested against the ISIC 2019 dataset. The model’s strength lies in its ability to automatically extract prominent features from dermoscopy images, provid- ing faster, more accurate diagnosis compared to traditional methods. However, the study is limited by its inability to handle datasets of dark-skinned individuals. Deep Convolutional Neural Networks (DCNNs) have demonstrated great success in skin lesion segmentation and classification tasks; yet, they have limitations and problems. One notable limitation lies in the requirement for large-scale class-balanced datasets, which can be particularly scarce in the domain of dermatology. The intricacies of skin lesions demand extensive and diverse datasets to ensure model generalization and robustness, but the scarcity of such datasets may impede the performance of DCNNs. Additionally, the interpretability of DCNNs remains a concern, as the inherent com- plexity of these models makes it challenging to decipher the decision-making process, crucial for establishing trust in clinical settings. Furthermore, DCNNs may struggle with handling images with various artifacts, such as poor lighting conditions, pres- ence of hairs, wounds or other irrelevant artifacts on the skin, and the unconventional perspectives, which are common in real-world clinical scenarios.
Multimodal classifier models
Multimodal models for skin lesion classification leverage multiple types of data—such as images, patient metadata (age, gender), and clinical history—to improve diagnostic accuracy. These models integrate different modalities,like dermoscopic images and textual descriptions, using advanced deep learning architectures, such as convo- lutional neural networks (CNNs) combined with natural language processing (NLP) techniquesor transformer models. By capturing complementary information from various data sources, multimodal approaches can enhance the model’s ability to differ- entiate between similar-looking lesions, account for contextual factors, and ultimately provide more robust and reliable classification results compared to models that rely solely on image data.This approach is particularly beneficial in medical settings, where accurate diagnosis often depends on synthesizing diverse information. Maiti et al.29 conducted a study focusing on the classification of skin lesions, particularly melanoma, by leveraging various machine learning and deep learning frameworks. The research presents a novel approach to image preprocessing and segmentation, which enhances the dataset by extracting and refining 23 texture features and 10 shape features through feature engineering techniques. This refined dataset is then processed using deep neural networks, with the model employing binary cross-entropy for classification tasks. The study explores the effectiveness of different activation layers, optimization techniques, and model architectures. The proposed model demonstrates a high accuracy of 96.8% on 170 MED-NODE and 2000 ISIC images, outperforming 12 other machine learning models and 5 deep learning models. The study emphasizes the efficiency and time-saving nature of the selected model. Wang et al.30 devel- oped an interpretability-based multimodal convolutional neural network (IM-CNN) aimed at improving the accuracy and interpretability of skin lesion diagnosis, a critical step in skin cancer screening. Their IM-CNN model integrates skin lesion images and patient metadata for multiclass classification, addressing challenges in generalization and interpretability often faced by deep learning methods. The IM-CNN comprises three main pathways: one for processing patient metadata, another for domain- knowledge-based features from segmented skin lesions, and a thirdfor skin lesion images. To enhance interpretability, visual modules are included to provide explanations for both images and metadata. The model’s performance is evaluated using AUC, sensitivity, specificity, and a newly introduced metric, AUC SEN 80, which focuses on the AUC with a sensitivity greater than 80%. Extensive experiments on the HAM10000 dataset demonstrate that IM-CNN significantly outperforms popular deep learning models such as DenseNet and ResNet. Notably, the multimodal approach achieved a 72% improvement in sensitivity and a 21% improvement in AUC SEN 80 compared to single-modal models. The visual explanations provided by the model facilitatetrust from dermatologists, enabling effective man–machine collaboration andmitigating thelimitations of black-boxmodels in medical decision-making. Ahmed et al.31 developed a multimodal approach to improve the classification of skin lesions, specifically targeting melanoma detection using dermoscopic images. They addressedthe challenge of distinguishing between skin lesions in their early stages by employing hybrid techniques that integrate handcrafted features with deep learning models. The study utilized the HAM10000 and PH2 datasets and resolved the issue of data imbalance between them. Pre-trained MobileNet and ResNet101 models were employed, and hybrid techniquescombining SVM with these models (SVM-MobileNet, SVM-ResNet101, and SVM-MobileNet-ResNet101) were developed to enhance performance. The integration of handcrafted features, which capture color, texture, and shape characteristics, with the features extracted by MobileNet and ResNet101, led to the creation of highly accurate feature sets. These features were subsequently classified using an Artificial Neural Network (ANN). The hybrid approach achieved notable results: for the HAM10000 dataset, an AUC of 97.53%, accuracyof 98.4%, sensitivity of 94.46%, precision of 93.44%, and specificity of 99.43% were reported. The PH2 dataset showed even higher performance, achieving 100% across all metrics, demonstrating the effectiveness of combining handcrafted features with deep learning for early-stage skin lesion detection.
To successfully incorporate DCNNs into dermatological procedures, these obstacles must be overcome. This highlights the necessity for continued research to improve the DCNNs’ functionality, interpretability, and suitability for tasks involving the categorization of skin lesions. Analyzing the existing methodologies, algorithms and techniques, the limitations that are identified in the skin lesion classification tasks are as follows:
-
I.
Pre-trained models might not capture fine-grained features specific to skin lesions.
-
II.
Transferring knowledge from vast domains such as the ImageNet performs well, but is not the most optimal model for a specific use case.
-
III.
The success of ensemble methods depends on the diversity of base learners. It is a challenging task to ensure diversity and prevent correlation among the models that are being ensembled.
-
IV.
Traditional feature extraction methods such as handcrafted feature extraction, and simple neural network-based feature extraction might sometimes miss complex and non-linear relationships within the images. Also, the extracted features might not represent the intricate details of skin lesions.
-
V.
iv. Manual hyperparameter tuningfor image classification CNNs requires iterative adjustments and retraining, which can be time-consuming and computationally expensive.
The proposed architecture, HHO-Reg-SA-Net (“Harris-Hawks Optimized RegNet with Soft Attention”) aims to overcome the existing limitations, by integrating state-of-the-art techniques, such as Soft Attention Mechanism for feature selection, and the Harris Hawks Optimization (HHO) for model fine-tuning and hyper- parameter optimization, the proposed approach seeks to expedite the convergence process. The inclusion of the soft attention block enables the model to focus only on the relevant features. The use of HHO ensures that the model parameters are efficiently adjusted, allowing for a faster convergence towards an optimal solution. This not only enhances the accuracy of the model but also contributes to a more time-effective train- ing pipeline, making the proposed approach a robust and efficient solution for image classification tasks.
Proposed work
Figure 1 refers to the overall system design of the proposed workflow.
Fig. 1.
The system design of the proposed framework.
Dataset
HAM10000: The HAM10000 dermoscopic image dataset from the Harvard Dataverse archive32 is used to train the models. It includes 10,015 dermoscopic images of skin lesions in seven different types: melanoma (MEL), dermatofibroma (DF), melanocytic nevus (NV), benign keratosis lesion (BKL), vascular (VASC) and actinic keratosis (AK). In the HAM10000 dataset, each skin lesion is associated with specific charac- teristics. Actinic Keratosis (AK) stands out as the most prevalent precancer, often observed in skin which overexposed to ultraviolet rays over a prolonged period of time.
Basal Cell Carcinoma (BCC), on the other hand, is identified as a slowly pro- gressing, malignant, locally invasive and epidermal skin tumor. Among older adults, benign keratosis (BKL) is one of the most common benign skin neoplasms. Commonly occurring around the region of the lower leg skin, dermatofibroma (DF) is a benign fibrous lesion. Melanocyte cells which produce melanin are the source of melanoma (MEL), the most serious variant of skin cancer. Benign Neoplasms or Hamartomas made out of melanocytes, which are the pigment producing cells that are found in the epidermis are referred to as Melanocytic Nevus (NV). Last but not least, a vascular lesion (VASC) is a term used to describe relatively frequent skin and underlying tissue anomalies that are known as birthmarks.
Data preprocessing
The images from the HAM10000 dataset consists of seven classes, with a size of 700 × 460 pixels. A 224 × 224 pixel resizing is applied to the dermoscopic pictures from the HAM10000 collection. The image datasets (Fig. 2). from the ISIC-archive are highly imbalanced in nature, with a particular class present in a dominantly high num- ber, while another class is very low in number. The input images are augmented using the image augmentation strategies such as brightness and sharpness enhancement, random gaussian noise addition, contrast variations, random rotations, horizontal and vertical flipping and shear along X and Y-axes. Following the process of augmentation, the class-balanced collection of pictures is fed into the pipeline for model training, where 20% of the dataset is utilized for validation and testing of the models, and 80% of the dataset is used for training the models.
Fig. 2.
Sample images of different classes from the HAM10000 dataset.
RegNetY032
Regularized Networks33, also known as “RegNet” draws inspiration from the widely successful ResNet architecture but introduces several key innovations that enhance its performance and efficiency. Unlike ResNet, which relies on a fixed network archi- tecture, RegNet defines a design space characterized by a simple, quantized linear function. This function dictates the width and depth of each network stage, allowing for a wide range of potential networks to be generated. Neural architecture search (NAS) has become increasingly popular, even though it comes with a significant computational burden. One of the key differences between RegNet and ResNet lies in their parameter space exploration. ResNet utilizes manual design or computationally expensive Neural Architecture Search (NAS) techniques to search for optimal network configurations. Conversely, RegNet leverages the simplicity of its design space to efficiently explore a vast range of potential networks using a human-in-the-loop approach.
This approach involves human experts guiding the search process by providing feedback on performance and selecting promising network candidates for further exploration. Another crucial distinction lies in the resource efficiency of both architectures. RegNet’s design space allows for the creation of networks with significantly fewer parameters than ResNet while achieving equivalent or even better performance. This is especially beneficial for resource critical environments, where computational power and memory are limited. Furthermore, RegNet introduces a novel regularization technique that employs a differentiable group sparsity constraint. This constraint encourages groups of neurons within the network to become inactive, leading to a more compact and efficient representation of the learned features. Unlike the conventional method, this approach avoids arbitrary alterations to the network’s dimensions. While RegNet also incorporates NAS technology, its utilization differs from previous NAS systems like MobileNetV3 and EfficientNet. Instead of focusing solely on parameter combina- tions within a predefined search framework, RegNet delves into the organization of regions and specific network design principles.
Network
Figure 3 refers to the overall architecture of the RegNetY032 model. It defines the high-level structure of the different components and how they are connected. Within RegNet, there are multiple network configurations available, each with varying depth, width, and resource requirements. The specific network configuration chosen for a particular application depends on factors like the desired accuracy, available resources, and computational constraints.
Fig. 3.
Overall architecture of RegNetY032.
Stem
Figure 4 depicts the stem block of the RegNetY032 model. This is the initial block of the network, responsible for processing the input data. In RegNet, the stem typically consists of a convolution layer, which is followed by batch normalization layer and ReLU activation function. The stem extracts relevant features from the input data and prepares it for further processing in the subsequent layers.
Fig. 4.

Stem layer of RegNetY032.
Body
Figure 5 refers to the body, which is the core of the network, where much of the feature extraction and learning takes place. The body consists of four stages, with each stage consisting of many residual blocks, also known as the Y-blocks. Y-blocks are a key component of RegNet and are responsible for learning complex relationships between features. Figure 6 refers to the Y-Block of the RegNetY032 model. The design of the body, including the number of stages and residual blocks, plays a vital role in the accuracy and efficiency of the network. The RegNetY032 architecture consists of four stages, where each stage is a series of sequential Y-Blocks. Table 1 refers to the number of Y-blocks in each stage of the RegNetY032 model.
Fig. 5.

Body layer of RegNetY032.
Fig. 6.
Y-block of RegNetY032.
Table 1.
Number of Y-blocks in each stage of RegNetY032 model.
| Stage | Number of Y-blocks |
|---|---|
| 1 | 2 |
| 2 | 5 |
| 3 | 13 |
| 4 | 1 |
Stage
This is a sub-unit within the body of the RegNetY032 architecture as shown in Fig. 7. Each stage consists of a stack of Y-blocks, responsible for extracting features at a specific level of abstraction. As the network progresses through different stages, the feature representations become increasingly complex and capture higher-level infor- mation from the input data. The number of stages and the structure of each stage can be varied depending on the desired network configuration.
Fig. 7.

Original head of RegNetY032.
Head
The head block shown in Fig. 8 of the RegNetY032 network is responsible for making prediction, which is decided based on the features extracted by the body. It typically includes a few fully-connected layers followed by an output layer. The size and com- plexity of the head depend on the specific task the network is trained for. The head of the RegNetY032 is modified by introducing a dense layer after the global average pooling layer. The inclusion of a dense layer enhances the model’s flexibility, enabling it to capture nuanced relationships within the data, especially relevant in the diverse and complex domain of skin lesions.
Fig. 8.

Modified head of RegNetY032.
The parameter efficiency of dense layers, coupled with dropout regularization, aids in preventing overfitting and promotes a more compact representation of learned features, crucial for robust classification. RegNet is a neural network architectural design (Fig. 9), and does not fixate on a specific network or set of networks. RegNet comprises of two series: RegNetX and RegNetY. RegNetX is primarily fashioned after the conventional residual bottleneck block featuring grouped convolution. In contrast, the enhancement of RegNetY involves the strategic incorporation of SE (Squeeze- and-Excitation) modules into the existing RegNetX framework. Remarkably, RegNet surpasses the currently available benchmark networks which are trained on ImageNet in terms of accuracy while achieving a fivefold improvement in speed on the Graphics Processing Unit (GPU).
Fig. 9.
Architecture of proposed Reg-SA-Net model.
Soft attention block
Convolutional neural networks in particular have benefited greatly from the introduction of attention mechanisms, which allow models to choose focus on important information in the input. Better performance in a variety of activities, such as picture categorization, image captioning, and visual question responding, is achieved through this selective concentration, which also increases representation power. The model is able to assign distinct components of the input sequence different degrees of priority thanks to the calculation of attention weights using a softmax function, which allows for dynamic soft attention.
Figure 10. refers to the architecture of the Soft Attention Block, which is integrated within the RegNetY032. The mathematical formulation of the soft attention mechanism has been elaborated as follows. The attention mechanism assigns different importance weights to different regions of the feature map to focus on the most relevant areas. Let the input feature tensor be denoted as t ∈ Rh×w×d, where h, w and d represent the height, width, and depth (channels) of the feature map, respectively. The attention mechanism first applies a convolution operation to the feature tensor t, with weights Wm ∈ Rh×w×d×M , where M is the number of attention heads (in this case, M = 16). The convolutional output is then normalized using a softmax acti- vation function across the attention heads to produce the attention weights for each region of the feature map:
![]() |
1 |
Fig. 10.
Architecture of the proposed Soft Attention Block.
Here, Am ∈ Rh×w represents the attention map generated by the m th attention head. The softmax function ensures that the attention weights for each location sum to 1 across the attention heads. The attention maps are then linearly combined to form a weighted context vector µ ∈ Rh×w×d, which consolidates the information from all the attention heads:
![]() |
2 |
where ⊙ represents the element-wise product. This context vector µ contains the information from the feature map t, but weighted by the attention maps to highlight the most relevant regions. Next, the context vector µ from Eq. (2) is scaled by a learnable scalar δ which is a input feature map with shape C × H × W, where: C = number of channels, H and W = height and width of the feature map to further refine the feature representation:
![]() |
3 |
Finally, the output of the attention mechanism fas is concatenated with the original feature tensor t to form a residual connection, ensuring that the model retains both the original and the attention-refined feature representations. The final output of the soft attention block is:
![]() |
4 |
This formulation allows the model to dynamicallyfocus on the most clinically significant regions of the dermoscopic images, while still maintaining the broader contextual information from the original feature tensor.
Harris Hawks optimization algorithm
One of the notable swarm-based optimization algorithm developed by Heidari et al.34 is Harris Hawks Optimization Algorithm (HHOA). It is inspired through the cooperative hunting techniques found in nature, particularly by mimicking the movements of a hawk pack and their prey. With this method, the prey is the best answer to a particular single-objective problem, and the hawks’ chase activities operate as search agents. HHO demonstrates adaptability across a broad range of real-world optimization issues, including those in the fields of engineering design, manufacturing, geotechnical engineering, pattern recognition, feature selection, power quality analysis, segmentation of images, and medication design. Especially, HHO35 performs exceptionallywell in situations when the search space types are unknown, and it can handle discrete and continuous space problems with ease. Research indicates its capacity to deliver solutions of superior quality, achieve high accuracy in parameter optimization, and enhance predictive performance39,40. The Harris Hawks Opti- mization Algorithm consists of the following phases, to find an optimal solution:
Algorithm 1: Soft attention block.
Phase 1: exploration phase—unraveling levy flight and local dynamics
-
Levy flight exploration: In the quest for prey, hawks employ a Levy flight strategy, reflecting their unpredictable searching patterns. The mathematical expression governing this exploration is depicted as:

5 
6 where u and v denote random numbers drawn from normal distributions, and β, a constant, which is set to 1.5.
Here, Xi signifies the position of the ith hawk, d denotes the problem dimension, and Xa and Xbdenote the lower and upper bounds of the search space. The Levy distribution is harnessed through:
-
Surrounding Hawks’ Influence: Hawks further refine their positions based on the influence of the best hawk in their vicinity, fostering local exploration dynamics:

7 Here, Xgrepresents the best hawk in the neighborhoodof hawki, and rand() returns a random number between 0 and 1.
Phase 2: exploitation phase—strategic approaches to besiege and selection
-
Soft Besiege: During the prey’s evasive maneuvers, hawks execute a soft besiege, progressively closing in on the target:

8 Here, Xr denotes the rabbit’s position, symbolizing the best solution, and E denotes the energy of the rabbit which ranges between 0 and 1.
-
Hard Besiege: When the prey is cornered, hawks’ transition into a hard besiege, converging towards its location:

9 In this equation, Y represents a hawk which is selected from the population in a random manner, and Xavg is the average value of the position of all the hawks in the search space.
- Progressive Selection: Positions that enhance the fitness value of a hawk are selectively retained, fostering convergence towards the optimal solution. This selective approach ensures the algorithm dynamically refines its solutions iteratively.
where f represents the fitness function, which evaluates the quality of the solution represented by Xinew .
10
These phases, characterized by intricate mathematical formulations, iteratively unfold until a predetermined stopping criterion is met, such as reaching the maximum count of iterations. Harris Hawks Optimization’s dynamic equilibrium between explo- ration and exploitation positions it as a formidable optimization tool, showcasing its potential to tackle a diverse array of complex problems. Once the iterations are over, the best set of hyperparameters which tends to increase the test accuracy of the Reg- SA-Net model are returned by the Harris-Hawks Optimization Algorithm (HHOA), which are then used to train the HHO-Reg-SA-Net model. Table 2 refers to the hyperparameters that are tuned using the Harris-Hawks Optimization algorithm (HHOA), and their search space.
Table 2.
Hyperparameters and their search space for Harris-Hawks Optimization Algorithm (HHOA).
| S. no | Hyperparameter | Search space |
|---|---|---|
| 1 | Learning rate | Lower bound = 0.001, Upper bound = 0.01 |
| 2 | Optimizer | Adagrad, RMSProp, SGD, Adam |
| 3 | Batch size | 16, 32, 64 |
Algorithm 2: Harris Hawks optimization algorithm (HHOA).
Experimental results
Experimental environments
The RegNetY032 and the Reg-SA-Net models were trained on the Google Colabora- tory Platform, which consists of a Linux-based backend, with 12.7 gigabytes of memory and 15 gigabytes of NVIDIA T4 GPU. TensorFlow 2.15.0 and Keras 3.0.4 libraries were used to build the RegNetY032 model, the soft attention block. Harris-Hawks Optimization was implemented on the Amazon Web Services platform, on EC2 G5 instance with 16 gigabytes of memory, and 24 gigabytes of GPU memory.
Model training strategy
The input images are resized from a size of 600 × 450 pixels to 224 × 224 pixels. Then, the resized input images are augmented using image augmentation strategies, which include horizontal flipping and verticalflipping probability is 0.5, random rotations range is within ± 30 degrees, random cropping of 90%, random perspective with distortion scale set to 0.5, addition of random sharpness factor varied between.
0.5 (blur) and 2.0 (sharp) and addition of random gaussian noise with mean = 0 and standard deviation between 0.01 and 0.05.
Figure 11. represents the class-wise image count of the train dataset before performing image augmentation and after performing image augmentation respectively. Before training, the data set is separated into three subsets, with 80% of the data used for training the models, 10% of the data for validating the models and 10% for testing the models. This split data is used to train the RegNetY032, the Reg-SA-Net and the HHO-Reg-SA-Net models. Table 3 refers to the image counts present in the training, validation and test sets used to train and evaluate the model.
Fig. 11.
HAM10000 image count before image augmentation and after image augmentation.
Table 3.
Image count of the dataset across the training, validation and test splits in the HAM10000 dataset.
| Class | Train image count | Validation image count | Test image count |
|---|---|---|---|
| AKIEC | 5280 | 660 | 660 |
| BKL | 5680 | 710 | 710 |
| BCC | 6080 | 760 | 760 |
| DF | 4080 | 510 | 510 |
| MEL | 6000 | 750 | 750 |
| NV | 6400 | 800 | 800 |
| VASC | 3280 | 410 | 410 |
Evaluation metrics
Evaluation metrics are essential for evaluating how well deep learning classifier models function. They offer useful insights into the model’s advantages and disadvantages as well as aid in quantifying how well a model performs on a particular task. The following are the evaluation metricsused for the evaluation of the RegNetY032, Reg-SA-Net and the HHO-Reg-SA-Net models:
-
Accuracy: An overall gauge of the model’s correctness is provided by accuracy. It is the proportion of the sum of genuine positives and genuine negatives to the total number of test instances. Accuracy is calculated by using the equation:

11 Here, TPdenotesTruePositives,TN denotesTrueNegatives, FPdenotesFalse.
Positives, and FN denotes False Negatives.
- Precision: Precision is defined as the ratio of genuine positives to the sum of genuine and false positives, and is often referred to as positive-predictive value (PPV). It centers on how accurate positive predictions are. Precisionis calculated from the equation:

12 - Recall: The ratio of genuine positives to the sum of genuine positives and false negatives is known as recall or sensitivity. It assesses the model’s capacity to correctly identify every instance in the positive class. Recall or sensitivity is calculated from the equation:

13 - F1-Score: The harmonic mean of recall and precision is known as the F1- score. It comes handy in situations where there is an uneven distribution of classes, since it offers a balance between the two measures. F1-Score is calculated as follows:

14 - Specificity: The proportion of genuine negative results to all negative results in a diagnostic test is known as specificity. It gauges how well the test can identify those without the ailment as negative, reducing the number of false positive findings. Higher specificity denotes a decreased likelihood of misidentifyingnegative results. Specificity is calculated using the equation:

15 - Confusion Matrix: A confusion matrix is a matrix, which summarizes the counts of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) predictions for each class in the data. It is used to assess how well a classification model performs. A thorough examination of the model’s behavior across several classes is made possible by the confusion matrix. A thorough grasp of the model’s classification performance across several classes is obtained by calculating evaluation metrics for each class or overall performance, such as accuracy, precision, recall, F1 score, and specificity, by looking at the confusion matrix. The confusion matrix format, for multi- class classification is as follows:

16 ROC Curve:In binary and multi-class classification tasks, the AUC-ROC (Area Under the Receiver-Operator Characteristic Curve) is a graphicalrepresen- tation and assessment metric with application. It evaluates a classification model’s performance over a range of categorization thresholds. Plotting the sensitivity versus specificity for various threshold levels results in the ROC curve. The area under the ROC curve is known as the AUC. It offers a single figure that sums up the classifier’s performance over a range of thresholds. Higher AUC values correspond to stronger classifier performance; AUC values range from 0 to 1.
Results
RegNetY032: The resized and augmented images were used as the input images, for training process of the RegNetY032 model. Initially, the RegNetY032 model was trained without the integration of the Soft Attention Block. Using the Adam opti- mizer and a learning rate of value 0.01 along with early stopping criteria spanning over 100 epochs, the model was trained. Model converged on 65 epochs, exhibiting a training accuracy of 98% and a test accuracy of 96%. Figure 12 (a) depicts the accuracy versus epoch plot of the RegNetY032 model. The training loss was 12%, while the test loss was found to be 19%. Figure 12 (d) represents the loss versus epoch plot of the RegNetY032 model. Figure 13 represents the confusion matrix of the RegNetY032 model, which shows the fraction of correctly classified test images, out of the overall images used for testing the model. Figure 14 represents the AUC-ROC Curve of the RegNetY032 model, tested against the HAM10000 dataset. An area under curve of 1.0 is obtained for the classes AK, BCC, BKL, DF and VASC classes, area of 0.99 for the NV class and an area of 0.98 for the MEL class. This is due to the fact that the Melanoma (MEL) is the most data insufficient class in the dataset.
Reg-SA-Net:The resized and augmented images were used as the inputs for training the Reg-SA-Net model. The architecture of the RegNetY032 model is modified by insertinga Soft-Attention Block after the body of the RegNetY032 model.The model underwent over 100 epochs of training utilizing the Adam optimizer with a learning rate of 0.01 and early stopping conditions over 100 epochs. Model converged on 67 epochs, exhibiting a training accuracy of 99%, and a test accuracy of 98%. Figure 12 (b) represents the plot of the accuracy of the Reg-SA-Net model versus the number of epochs. The training loss was 11%, while the test loss was found to be 26%. Figure 12 (e) represents the graph of loss of the Reg-SA-Net model versus the number of epochs. Figure 15 represents the confusion matrix of the Reg-SA-Net model, which shows the part of accurately classified test instances, out of the overall images used for testing the model. An area of 1.0 for the classes AK, BCC, BKL, DF, NV and VASC classes, area of 0.98 for the MEL class. Figure 16 represents the AUC-ROC Curve of the Reg-SA-Net model, testedagainstthe HAM10000 dataset. An area under curve of 1.0 has been obtained for all the classes except Melanoma, and an area of 0.98 for the Melanoma class.
HHO-Reg-SA-Net: Harris-Hawks Optimization was implemented to per- form hyperparameter optimization on the proposed Reg-SA-Net model. The number of hawks, also known as the population size, is initialized to 10, and the Harris-Hawks optimization was run for 2 epochs. Table 4 refers to the best fit solution, or the best set of hyperparameters returned by the Harris-Hawks Optimization algorithm.
Fig. 12.
Accuracy and Loss vs Epoch plot for all the trained models.
Fig. 13.

Confusion matrix of RegNetY032 model.
Fig. 14.

Confusion matrix of Reg-SA-Net model.
Fig. 15.

Confusion matrix of HHO-Reg-SANet model.
Fig. 16.

ROC-Curve of RegNetY032 model.
Table 4.
Best-Fit values of the hyperparameters using the Harris-Hawks optimization algorithm.
| S. No | Hyperparameter | Best-Fit value |
|---|---|---|
| 1 | Learning rate | 0.00934439 |
| 2 | Optimizer | Adam |
| 3 | Batch size | 16 |
The Reg-SA-Net model was trained with the optimalset of hyperparameters returned by the Harris-HawksOptimization algorithm. The model was trained using the Adam optimizer, with a batch size of 16, learning rate of 0.00934439, and early stopping conditions spanning over 100 epochs. Model converged on 49 epochs, exhibiting a training accuracy of 97% and a test accuracy of 99%. Figure 12 (c) represents the plot accuracy of the model versus the number of epochs. The training loss was 17%, while the test loss was found to be 22%. Figure 12 (f) represents the plot of model loss versus the number of epochs. Figure 17 represents the confusion matrix of the HHO-Reg-SA-Net model, which shows the fraction of correctly classified test images, out of the overall images used for testing the model.
Fig. 17.

ROC-Curve of Reg-SA-Net model.
Figure 18 represents the AUC-ROC Curve of the HHO-Reg-SA-Net model, tested against the HAM10000 dataset. An area under curve of 1.0 has been obtained for the classes AK, BCC, DF, NV and VASC classes, area of 0.99 for the BKL and MEL classes respectively. The models trained, validated, and tested against the HAM10000 dataset are all compared and evaluated based on their performance, which is shown in Table 5 (Table 6).
Fig. 18.

ROC-curve of HHO-Reg-SANet model.
Table 5.
Comparison of evaluation metrics of all the models against the HAM10000 dataset.
| Model | RegNetY032 | Reg-SA-Net | HHO-Reg-SA-Net |
|---|---|---|---|
| Accuracy | 96.28 | 98.99 | 99.27 |
| Precision | 97.12 | 98.71 | 99.24 |
| Recall | 97.04 | 98.06 | 99.07 |
| F1-Score | 97.23 | 98.09 | 99.14 |
| Specificity | 96.84 | 98.17 | 99.09 |
Table 6.
Comparison of the proposed HHO-Reg-SA-Net model with the existing models against the HAM10000 dataset.
| Model | Accuracy | Precision | Recall | F1-Score | Specificity |
|---|---|---|---|---|---|
| Bilinear CNN15 | 93.21 | 92.92 | 93.00 | 93.21 | – |
| InceptionV3 + DarkNet53 + MFO16 | 95.80 | 95.11 | 91.75 | – | – |
| Ensemble Learning17 | 88.00 | 87.00 | 94.00 | 89.00 | – |
| SBXception Network19 | 96.97 | 85.34 | 95.43 | – | – |
| MHCL-PO21 | 92.20 | 90.20 | 84.72 | 86.41 | – |
| Fully Transformer Network36 | 92.70 | 85.70 | 62.10 | – | 93.60 |
| Reg-SA-Net [ours] | 98.99 | 98.71 | 98.06 | 98.09 | 98.17 |
| HHO-Reg-SA-Net [ours] | 99.27 | 99.24 | 99.07 | 99.14 | 99.09 |
Discussion
The classification accuracy was improved once the Soft-Attention Block was incorporated into the RegNetY032 network. Further improvement was achieved by training the model with the optimal hyperparameters returned by the Harris-Hawks Optimization Algorithm. Convergence was attained at early epochs by training the Reg-SA-Net model with the ideal hyperparameters, which decreased the model’s training time and increased its test accuracy. The performance comparison of the currently existing mod- els for skin lesion classification and the proposed HHO-Reg-SA-Net model is shown in Table 7.
Table 7.
Validation of the proposed HHO-Reg-SA-Net model with ISIC 2019 and PH2 datasets.
| Dataset | No. of Images | Skin Tone Range | Accuracy | Precision | Recall |
|---|---|---|---|---|---|
| HAM10000 | 10,015 | Light–Medium | 99.27 | 99.24 | 99.07 |
| ISIC 2019 | 25,331 | Medium–Dark | 98.84 | 98.61 | 98.45 |
| PH2 | 200 | Fair–Medium | 98.56 | 98.47 | 98.48 |
To assess the real-time applicability of HHO-Reg-SA-Net model, we evaluated its computational efficiency on different hardware configurations. The proposed model achieves an average inference time of 14.8 ms per image on GPU and 0.42 s per image on CPU, indicating suitability for near real-time analysis. The model contains 22.8 million parameters and requires 5.3 GFLOPs, maintaining a favorable balance between accuracy and computational cost. HHO-Reg-SA-Net achieved 0.82 s per image inference time and 99.27.
Table 4 represents the hyperparameters used for Harris-Hawks optimization algorithm. Tables 5, 6 highlights the computational efficiency gains achieved through the incorporation of the Soft-Attention Block and the application of the Harris- Hawks Optimization Algorithm. Figure 19 depictsthe comparisonof the proposed models with the existing models for skin lesion classification. The proposed Reg-SA- Net and HHO-Reg-SA-Net models demonstrates superior performance and exhibits enhanced scalability and convergence characteristics, rendering it a compelling choice for real-world deployment in dermatological diagnosis systems.
Fig. 19.
Attention maps generated by the Reg-SA-Net and HHO-Reg-SA-Net for various types of skin lesions.
To test the generalization and strength of the proposed HHO-Reg-SA-Net model beyond the HAM10000 dataset, we conducted additional tests on the ISIC 2019 and PH2 datasets. As shown in Table 7, the model performedconsistently well with high accuracy and F1-scores on all datasets, suggesting excellent generalization under different acquisition conditions and lesion types.
Figure 20 refers to the attention maps generated by the soft attention block for various types of skin lesions from the HAM10000 dataset. Regions which are red in colour represents the lesion area, while the outer blue maps represent the normal skin. The intensity of the attention map is high in regions where distinguishing features of the lesions are prominent, such as in melanoma (MEL), which often presents with asymmetrical borders, varied pigmentation, and irregular shapes. Similarly, attention intensity will be elevated for actinic keratosis (AK) and basal cell carcinoma (BCC), as these lesions typically exhibit unique textures and color variations that the model can readily identify. Conversely, the attention map intensity may be lower for lesions like benign keratosis (BKL) and melanocytic nevus (NV), which can often appear similar to surrounding skin or lack distinctive features, leading to a more generalized focus from the model. Dermato-fibromas (DF) and vascular lesions (VASC) might also yield lower attention intensities due to their less pronounced characteristics compared to more malignant forms of skin cancer. Figure 21 refers to the correlation of the intensity of the attention map with classification accuracy across different lesion types.
Fig. 20.

Correlation between attention map intensity and classification accuracy.
Fig. 21.
Comparison of the HHO-Reg-SA-Net model with the existing models for skin lesion classi- fication.
Ablation study
An ablation study was conducted to quantify the individual and combined contribu- tions of the Soft Attention Block (SA Block) and Harris Hawks Optimization (HHO) to the overall performance of the proposed model. To gain deeper insights, the study was extended to include comparisons with alternative attention architectures—namely the Squeeze-and-Excitation (SE) and Convolutional Block Attention Module (CBAM)—as well as other optimization algorithms, including Particle Swarm Optimization (PSO) and Grey Wolf Optimizer (GWO).
Results show that the inclusion of attention mechanisms generally improves model performance, with the Soft Attention Block yielding higher gains in F1 Score and speci- ficity compared to SE and CBAM. Additionally, the use of metaheuristic optimization further enhances both accuracy and convergence speed. Among the tested optimizers, HHO achieved the best balance between exploration and exploitation, resulting in the highest improvement in all performance metrics while reducing total training time by more than 70% compared to the baseline RegNetY032 model. Table 8 refers to the quantitative analysis of the performance boost of the models. All the improvements stated in the Table 8 are relative to the baseline RegNetY032 model.
Table 8.
Quantitative analysis of the performance boost of the models.
| Model | Attention type | Optimization | Accuracy improvement | Precison improvement | Recall improvement | F1-score improvement | Specificity improvement | Training time (h) |
|---|---|---|---|---|---|---|---|---|
| RegNetY032 | – | – | – | – | – | – | – | 4.11 |
| Reg-SE- Net | SE | – | 1.82 | 1.23 | 0.91 | 0.78 | 0.96 | 4.45 |
| Reg- CBAM- Net | CBAM | – | 2.16 | 1.44 | 1.05 | 0.84 | 1.12 | 4.37 |
| Reg-SA- Net | Soft Attention | – | 2.71 | 1.59 | 1.02 | 0.86 | 1.33 | 4.31 |
| PSO- Reg-SA- Net | Soft Attention | PSO | 2.84 | 1.88 | 1.47 | 1.31 | 1.74 | 3.95 |
| GWO- Reg-SA- Net | Soft Attention | GWO | 2.91 | 1.97 | 1.63 | 1.47 | 1.86 | 3.62 |
| HHO-Reg-SA- Net | Soft Attention | HHO | 2.99 | 2.12 | 2.03 | 1.91 | 2.25 | 1.14 |
Conclusion
This paper introduces a novel method for skin lesion classification using the HHO-Reg-SA-Net model. The model leverages a Soft Attention Block to highlight key features in dermoscopic images, improving classification accuracy by effectively mitigating artifacts like hairs and wounds.Data augmentation on the HAM10000 dataset addresses overfitting issues, and modifications to the RegNetY032 architecture enhance its performance. The Harris-Hawks Optimization algorithm optimizes hyper- parameters, leading to faster convergence and significantly reduced training time.Our model achieved a test accuracy of 98.99%, and with furtherhyperparameter tun- ing, it improved to 99.27%.Future research will explore lesion segmentation, validate the model on diverse datasets, and assess its applicability in other medical imaging domains.
Acknowledgements
We would like to extend our sincere gratitude to S. Harrish Ragavendar and S. Jayajith Roshan, a student of Mepco Schlenk Engineering College for their invaluable contribution to this research. Their dedication, keen insight and diligent work in data collection and analysis, and first draft preparation greatly enhanced the quality of this study. Their commitment to academic excellence and research integrity has been commendable throughout this project.
Author contributions
R. Naga Priyadarsini contributed to the conceptualization of the study, methodology design, and overall supervision of the research work. Bhawana Tyagi and M. Priyadharsini was involved in data analysis, interpretation of results, and critical revision of the manuscript. All authors reviewed and approved the final version of the manuscript.
Funding
Open access funding provided by Vellore Institute of Technology.
Data availability
The data is publicly available in the Harvard Dataverse. Dataset URL: 10.7910/DVN/DBW86T
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
R. Naga Priyadarsini, Email: nagapriyadarsini.r@vit.ac.in.
Bhawana Tyagi, Email: bhawana.tyagi@vit.ac.in.
References
- 1.(WHO), W. H. O. International agency for research on cancer (iarc) (2024). URL https://www.iarc.who.int.
- 2.Apalla, Z., Nashan, D., Weller, R. B. & Castellsagu´e, X. Epidemiology, disease burden, pathophysiology, diagnosis, and therapeutic approaches. Dermatol Ther (Heidelb) 7, 5–19 (2017). [DOI] [PMC free article] [PubMed]
- 3.Siegel, R. L., Miller, K. D. & Jemal, A. Cancer statistics, 2020. CA A, cancer J. Clinicians 70, 7–30 (2020). [DOI] [PubMed]
- 4.Vestergaard, M. E., Macaskill, P., Holt, P. E. & Menzies, S. W. Dermoscopy compared with naked eye examination for the diagnosis of primary melanoma: A meta-analysis of studies performed in a clinical setting. Brit. J. Dermatol.159, 669–676 (2008). [DOI] [PubMed] [Google Scholar]
- 5.Feng, H., Berk-Krauss, J., Feng, P. W. & Stein, J. A. Comparison of dermatologist density between urban and rural counties in the united states. JAMA Dermatol.154, 1265–1271 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Leiter, U., Eigentler, T. & Garbe, C.Epidemiology of skin cancer. Sunlight, Vitamin D and Skin Cancer (2015).
- 7.Morton, C. & Mackie, R. Clinical accuracyof the diagnosis of cutaneous malignant melanoma. Br. J. Dermatol.138, 283–287 (1998). [DOI] [PubMed] [Google Scholar]
- 8.Chaturvedi, S. S., Tembhurne, J. V. & Diwan, T.A multi-classskin cancer classification using deep convolutional neural networks. Multimedia Tools and Applications 79, 28477–28498 (2020).
- 9.Litjens, G. et al. A survey on deep learning in medical image analysis. Med. Image Anal.42, 60–88 (2017). [DOI] [PubMed] [Google Scholar]
- 10.Adegun, A. & Viriri, S. Deep learning techniques for skin lesion analysis and melanoma cancer detection: A survey of state-of-the-art. Artif. Intell. Rev.54, 811–841 (2020). [Google Scholar]
- 11.Maiti, A., Chatterjee, B., Ashour, A. S. & Dey, N. Computer-aided diagnosis of melanoma: a review of existing knowledge and strategies. Current Med. Imag.16, 835–854 (2020). [DOI] [PubMed] [Google Scholar]
- 12.Song, Y. et al. Large margin local estimate with applications to medical image classification. IEEE Trans. Med. Imag.34, 1362–1377 (2015). [DOI] [PubMed] [Google Scholar]
- 13.Iqbal, I., Younus, M., Walayat, K., Kakar, M. U. & Ma, J.Automated multi- class classification of skin lesions through deep convolutional neural network with dermoscopic images. Comput. Med. Imag. Graph. 88 (2021). [DOI] [PubMed]
- 14.Veeramani, N., Jayaraman, P., Krishankumar, R., Ravichandran, K. S. & Gan- domi, A. H. Ddcnn-f: double decker convolutional neural network’f’feature fusion as a medical image classification framework. Sci. Reports 14, 676 (2024). [DOI] [PMC free article] [PubMed]
- 15.Calderon, C., Sanchez, K., Castillo, S. & Arguello, H. Bilsk: A bilinear convolu- tional neural network approach for skin lesion classification. Comput. Methods Programs Biomed Update 1 (2021).
- 16.Usama, M., Naeem, M. A. & Mirza, F. Multi-class skin lesions classification using deep features. Sensors, MDPI 22 (2022). [DOI] [PMC free article] [PubMed]
- 17.Rahman, Z., Hossain, M. S., Islam, M. R., Hasan, M. M. & Hridhee, R. A. An approach for multiclass skin lesion classification based on ensemble learning. Inform. Med Unlocked 25 (2021).
- 18.Hasan,M. K., Elahi,M. T. E., Alam, M. A., Jawad,M. T. & Mart´ı, R. Dermoexpert: Skin lesion classification using a hybrid convolutional neural net- work through segmentation, transfer learning, and augmentation. Inform. Med. Unlocked 28 (2022).
- 19.Mehmood, A. et al. Sbxception: A shallower and broader xception architecture for efficient classification of skin lesions. Cancers, MDPI 15 (2023). [DOI] [PMC free article] [PubMed]
- 20.Farhat, A. et al. Multiclass skin lesion classification using hybrid deep features selection and extreme learning machine. Sensors, MDPI 22 (2022). [DOI] [PMC free article] [PubMed]
- 21.Liang, S. et al.Skin lesion classification base on multi-hierarchy contrastive learning with pareto optimality. Biomed. Signal Process. Control 86 (2023).
- 22.Hoang, L., Lee, S.-H., Lee, E.-J. & Kwon, K.-R.Multiclass skin lesion classifi- cation using a novel lightweight deep learning framework for smart healthcare. Appl. Sci. MDPI 12 (2022).
- 23.Zhang, J., Xie, Y., Xia, Y. & Shen, C. Attention residual learning for skin lesion classification. IEEE Trans. Med. Imaging38, 2092–2103 (2019). [DOI] [PubMed] [Google Scholar]
- 24.Ding, S. et al.Deep attention branch networks for skin lesion classification. Comput. Methods Programs Biomed. 212 (2021). [DOI] [PubMed]
- 25.Qian, S., Ren, K., Zhang, W. & Ning, H. Skin lesion classification using cnns with grouping of multi-scale attention and class-specific loss weighting. Comput. Methods Programs Biomed. 226 (2022). [DOI] [PubMed]
- 26.He, X., Wang, Y., Zhao, S. & Yao, C. Deep metric attention learning for skin lesion classification in dermoscopy images. Complex & Intell. Syst. 8, 1487–1504 (2022).
- 27.Naeem, A. & Anees, T. Dvfnet: A deep feature fusion-based model for the multi- classification of skin cancer utilizing dermoscopy images. PLoS ONE19, e0297667 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Naeem, A. et al. Snc net: Skin cancer detection by integrating handcrafted and deep learning-based features using dermoscopy images 12 (2024).
- 29.Maiti, A. & Chatterjee, B. Improving detection of melanoma and naevus with deep neuralnetworks. Multimed. Appl.79, 15635–15654 (2020). [Google Scholar]
- 30.Wang, S., Yin, Y., Wang, D., Wang, Y. & Jin, Y. Interpretability-based multi- modal convolutional neural networks for skin lesion diagnosis. IEEE Trans. Cybernet.52, 12623–12637 (2021). [DOI] [PubMed] [Google Scholar]
- 31.Ahmed, I. A., Senan, E. M., Shatnawi, H. S. A., Alkhraisha, Z. M. & Al-Azzam, M. M. A. Multi-models of analyzing dermoscopy images for early detection of multi-class skin lesions based on fused features. Processes11, 910 (2023). [Google Scholar]
- 32.Tschandl, P., Rosendahl, C. & Kittler, H. The ham10000 dataset, a large col- lection of multi-source dermatoscopic images of common pigmented skin lesions. Sci. Data5, 1–9 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Xu, J. et al.Regnet: Self-regulatednetwork for image classification. IEEE Trans. Neural Networks Learn. Syst. 34, 9562–9567 (2023). [DOI] [PubMed]
- 34.Heidari, A. A. et al. Harris hawks optimization: Algorithm and applications. Futur. Gener. Comput. Syst.97, 849–872 (2019). [Google Scholar]
- 35.Yang, N., Tang, Z., Cai, X., Chen, L. & Hu, Q. Cooperative multi-population Harris Hawks optimization for many-objective optimization. Complex & Intell. Syst.8, 3299–3332 (2022). [Google Scholar]
- 36.He, X. et al. Fully transformer network for skin lesion analysis. Med. Image Anal. 77, 102357 (2022). [DOI] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data is publicly available in the Harvard Dataverse. Dataset URL: 10.7910/DVN/DBW86T
















