Abstract
In light of the ongoing battle against COVID-19, while the pandemic may eventually subside, sporadic cases may still emerge, underscoring the need for accurate detection from radiological images. However, the limited explainability of current deep learning models restricts clinician acceptance. To address this issue, our research integrates multiple CNN models with explainable AI techniques, ensuring model interpretability before ensemble construction. Our approach enhances both accuracy and interpretability by evaluating advanced CNN models on the largest publicly available X-ray dataset, COVIDx CXR-3, which includes 29,986 images, and the CT scan dataset for SARS-CoV-2 from Kaggle, which includes a total of 2,482 images. We also employed additional public datasets for cross-dataset evaluation, ensuring a thorough assessment of model performance across various imaging conditions. By leveraging methods including LIME, SHAP, Grad-CAM, and Grad-CAM++, we provide transparent insights into model decisions. Our ensemble model, which includes DenseNet169, ResNet50, and VGG16, demonstrates strong performance. For the X-ray image dataset, sensitivity, specificity, accuracy, F1-score, and AUC are recorded at 99.00%, 99.00%, 99.00%, 0.99, and 0.99, respectively. For the CT image dataset, these metrics are 96.18%, 96.18%, 96.18%, 0.9618, and 0.96, respectively. Our methodology bridges the gap between precision and interpretability in clinical settings by combining model diversity with explainability, promising enhanced disease diagnosis and greater clinician acceptance.
Keywords: Deep learning, Ensemble, Convolutional neural networks (CNN), Explainable AI, Chest X-ray images, CT scan images, COVID-19
Subject terms: Classification and taxonomy, Image processing
Introduction
The ongoing C-19 (COVID-19) pandemic, though diminishing, continues to present significant challenges with periodic outbreaks1. Early and accurate identification of infections remains crucial in controlling the spread of the virus. While real-time reverse transcription-polymerase chain reaction (RT-PCR) has been the standard diagnostic procedure, it has notable limitations, including reliability issues, sensitivity concerns, and a troubling rate of false-negative results2, 3. These challenges highlight the need for alternative diagnostic approaches.
CT (Computed Tomography) scans and chest X-rays (CXR) play a crucial role in diagnosing C-19, delivering a more dependable, efficient, and precise diagnostic method, particularly useful during the initial stages of the disease4, 5. Nevertheless, interpreting these imaging techniques can be challenging, especially when distinguishing C-19 from other types of viral pneumonia. Timely and accurate diagnosis is essential for controlling the virus’s spread by enabling prompt isolation and treatment of infected individuals. To enhance the diagnostic accuracy of these imaging modalities, deep learning, particularly convolutional neural networks (CNNs), has been extensively adopted. These models have significantly advanced the field of medical imaging, improving disease detection across various modalities, including CT scans and chest X-rays6–11. However, in healthcare, where decisions are critical, accuracy alone is insufficient. The “black-box” nature of deep learning models raises concerns about their transparency and trustworthiness. Explainable artificial intelligence (XAI) techniques are essential in addressing these concerns by offering an understanding of how these models arrive at their decisions. This is especially crucial in diagnostic imaging, where healthcare providers need to comprehend and have confidence in the AI’s results for them to be effectively used in clinical practice.
In this study, we focus on evaluating the effectiveness of five widely used pre-trained deep learning models—VGG16, ResNet50, DenseNet169, EfficientNetB3, and Xception—for C-19 classification using radiological images12–16. Unlike previous studies that frequently incorporate explainability into newly developed models without considering how these techniques influence model selection, our approach distinctly prioritizes the explainability of pre-trained models. We specifically use various XAI techniques, including LIME, SHAP, Grad-CAM, and Grad-CAM + +17–20, to analyze and visualize the crucial regions in radiological images that impact model predictions. This method not only depends on our understanding of the decision-making processes of pre-trained models but also builds greater trust in their application for C-19 diagnosis by ensuring that model selection is guided by the interpretability of the results.
To further improve classification accuracy, we conducted a comprehensive evaluation and interpretability analysis across multiple models. Based on quantitative measures of both performance and interpretability, we identify the top-performing models and combine them to develop an ensemble model. This ensemble approach leverages the strengths of each model, resulting in improved C-19 classification accuracy and a more robust and generalized model performance across different datasets. By integrating explainability into the diagnostic process, our study not only advances the technical performance of AI models but also addresses the ethical and practical considerations necessary for their adoption in real-world healthcare settings. Our findings underscore the importance of transparency in AI-driven diagnostics and assist in the ongoing initiatives to enhance the public health response to pandemics. The primary goal of this research is to evaluate how well pre-trained deep neural networks perform in classifying C-19 cases with radiological image datasets. Specifically, we aim to:
Evaluate the accuracy and performance of various pre-trained neural networks (e.g., VGG16, DenseNet169, ResNet50, EfficientNetB3, Xception) in C-19 classification, utilizing multiple datasets to ensure robustness.
Apply state-of-the-art interpretability techniques (LIME, SHAP, GradCAM, GradCAM++) to highlight and understand the image key regions that contribute most to the models’ classification decisions, thereby identifying key features and enhancing the transparency of the models.
Perform a comprehensive quantitative analysis to identify top-performing models based on their accuracy and interpretability scores. We will then combine these models into an ensemble model to further improve classification performance and generalizability across different datasets.
Ensure the interpretability of the ensemble model by analyzing the combined predictions and understanding how the ensemble approach enhances the overall decision-making process, thereby building trust in the model’s clinical applicability.
Section 2 presents an overview of relevant studies, while Section 3 addresses the pre-trained networks, interpretability methods, and the introduced ensemble model. Section 4 discusses data description, preprocessing of input images, model tuning, and a comparative evaluation of model performance and explanations using XAI. Finally, Section 5 presents the conclusions and results.
Related work
The use of computer-aided diagnostic systems powered by deep learning with radiological images is becoming more popular for accurately detecting and classifying lung diseases. Recently, researchers have become more interested in developing CAD systems that use radiographic images for the automatic identification of C-1921–23. Abubakar et al.24 introduced a novel approach for detecting C-19 from the CT scan images, which integrated deep learning models with the Histogram of Oriented Gradient (HOG). Using SVM, they achieved a 99.4% overall accuracy on the CT scan dataset. This proved the substantial boost in diagnostic efficacy that their innovative strategy made achievable. Ragab et al.25 used the CT dataset and came up with a novel compression technique that allowed them to correctly classify C-19 and NC-19 CT images with a 99.3% overall accuracy. Their DeepCSFusion model worked better and used less computing time when compared to other pipelines. In Haynes et al.26, the issue of model generalization was explored using C-19 detection from chest X-rays. By showing how pre-training on pertinent chest x-ray imagery can stabilize model performance and reduce feature dependency, their work highlighted the need for new techniques for trustworthy model learning with minimal datasets. Suhartanto et al.27 presented SCOV-CNN, a Convolution Neural Network architecture particularly created for C-19 classification from CT images. Achieving an accuracy of 0.96, a precision of 0.98, and an F1 score of 0.95, the model demonstrated superior performance compared to previous deep learning algorithms. Their integration of the model into an online platform shows how practical it is for use in actual medical settings. For C-19 detection, Zhao et al.28 introduced MLF-AttNet, a multi-level feature attention network. Their method achieved accuracies of 93.66% on the X-ray dataset and 87.08% on the CT dataset. Designed for deep feature extraction from datasets of chest X-rays and computed tomography scans, CNNs were employed to develop an automated method for C-19 detection29. The Hybrid Whale Optimization (HWO) algorithm was used to achieve a 99.43% accuracy in classification trials. Additionally, canonical correlation-based fusion contributed to surpassing traditional CNNs in performance for augmented images. Hoffer et al.30 built a smartphone application that employs thermal imaging to detect C-19 infection. Deep learning analysis by them attained a sensitivity of 88.7% and specificity of 92.3%, which exhibits the potential of noninvasive thermal imaging in primary C-19 screening, pneumonia identification as well as acting as a pragmatic alternative to hospital-based diagnostic methods. To exactly discover chronic lung diseases in C-19 patient’s lungs, Sanampudi et al.31 recommended Local Search Enhanced AHO-based Inception-ResNet-v2 Model (LSAIRM). The model was helpful in disease classification because its notable performance metrics stood at precision of 98.95% and accuracy of 98.97%.
A two-phase defense framework was presented by Sheikh and Aasim32 to strengthen the resilience of deep learning-based diagnostic models against adversarial attacks in radiology image-based C-19 diagnosis. Their method demonstrated enhanced model resilience, preserving a high degree of precision in C-19 diagnosis from uncontaminated images while successfully reducing the negative effects of hostile assaults. Using the MobileNetV2, InceptionV3, and VGG19 architectures, Türk and Kökver33 provide a three-channel fusion CNN model for successful lung opacity detection. With their approach, lung opacity regions can be monitored over an extended period with substantial ease and advantages for clinicians. On a combined dataset, their approach achieves high accuracy values of 92.52% for two classes, 92.44% for three classes, 87.12% for four classes, and 91.71% for five classes. Saheb et al.34 presented the automated deep learning-based C-19 identification framework, which also showed that it is accurate in detecting C-19 from CT dataset. Achieving a accuracy of 98.49%, their system surpasses existing methods, showcasing the effectiveness of their innovative approach to automated diagnosis of C-19. Introducing AI into clinical workflows has posed a major challenge in terms of explainability, as noted in previous studies35.
Koul et al.36 investigated advanced convolutional neural networks combined with Explainable Artificial Intelligence for detecting and classifying respiratory diseases from medical images. Among the many deep learning models investigated, InceptionResNetV2 obtained the highest accuracy (99.15%), precision (98%), recall (99%), and F1 score (0.0275 loss) in their study. These are noteworthy performance measures. The C-19 diagnosis decision support system developed by Chadaga et al.37 achieves a strong 96% accuracy using patient data from Manipal hospitals. Their method combines explainable artificial intelligence, deep learning, and machine learning approaches, providing useful tools for preliminary screening and reducing the load on the medical infrastructure.
The current review of the literature reveals several constraints, as outlined in Table 1. First, many surveyed studies lack sufficient emphasis on interpretability, potentially hindering the understanding of the inner workings and decision-making processes of the deep learning models employed. Second, the literature predominantly focuses on performance metrics when forming ensembles, neglecting other crucial factors that participate in the overall reliability and strength of the models in real-world applications. Therefore, in this work, we have selected models based on both performance and explanation through quantitative evaluation to create a robust ensemble model. Additionally, we have applied explainability techniques to the ensemble model itself to increase trust in its decision-making process. This approach not only enhances the accuracy of C-19 classification but also addresses the critical need for interpretability and trustworthiness in clinical applications.
Table 1.
Survey of DL Models’ Performance Across different datasets for C-19 detection.
| Author(s) & Year | Architecture/Models | Image Type | Performance Parameters (%) | Task | Datasets Used |
|---|---|---|---|---|---|
| Prinzi et al. (2024)21 | Support Vector Machine and Random Forest classifiers | CXR images | AUC: 81.9%, Accuracy: 73.3%, Specificity: 70.5%, Sensitivity: 76.1% |
Prognosis Prediction |
Multi-Centric Dataset from Covid CXR Hackathon (1589 patients) |
| Soda et al. (2021)22 | Convolutional Neural Networks (CNNs) and handcrafted features |
CXR images |
10-fold CV: AUC: 81.3%, Sensitivity: 77.2%, Specificity: 75.1% |
Prognosis Prediction |
Multi-Centric Dataset from 6 Italian Hospitals (820 patients) |
| Abubakar et al. (2024)24 | HOG and deep learning | CT scans | Accuracy: 99.4% | Classification | Three Public Datasets |
| Suhartanto et al. (2024)27 | SCOV-CNN | CT images | Accuracy:96%, Precision: 98%, F1-Score: 95% | Classification | CT images of 120 patients from hospitals in Brazil. |
| Zhao et al. (2024)28 | MLF-AttNet | Chest CT and CXR images | CT: Accuracy: 93.66%, X-ray: Accuracy: 87.08% | Classification | COVID-CT dataset, Chestx-ray8 |
| Abdellatef and Allah (2024)29 | Optimized CNNs with HWO method and fusion based on Canonical Correlation | CT scans and CXR images | Accuracy: 99.43% | Classification | C-19 CT and CXR Datasets |
| Hoffer et al. (2024)30 | Thermal imaging based approach | Thermal images | Sensitivity:88.7%, Specificity: 92.3% | Detection | Private Dataset |
| Türk and Kökver (2023)33 | Three-channel fusion CNN model (MobileNetV2, InceptionV3, VGG19) | Radiological images | Accuracy: 92.52% (2 classes), 92.44% (3 classes), 87.12% 4 classes), 91.71% (5 classes) | Classification | C-19 Radiological Image Dataset |
| Saheb et al. (2023)34 | Automated DL-based Framework for C-19 Detection | CT images | Accuracy: 98.49% | Detection | C-19 CT Dataset |
| Chadaga et al. (2023)37 | Decision support system integrating ML, DL, and XAI tech- niques | Chest X-ray images | Accuracy: 96% | Classification | C-19 CXR images |
Proposed methodology
This section describes propose approach to improving the diagnosis of C-19 through the use of radiological images. Our approach leverages state-of-the-art convolutional neural networks (CNNs), advanced explainability techniques, and various strategies to address the potential risks of small datasets and overfitting. The methodology consists of a systematic, multi-step process designed to optimize model performance, ensure generalizability, and enhance the understanding of prediction mechanisms. Figure 1 illustrates the main steps in building the model, including data processing, model training, ensemble creation, and interpretability enhancement which are discussed below.
Fig. 1.
Major steps in model development.
- Step 1: Model Evaluation
- Fine-tune and assess the capabilities of five models—VGG16, DenseNet169, ResNet50, EfficientNetB3, and Xception—using X-ray and CT scan image datasets, focusing on their performance and the explanations provided.
- Apply data augmentation methods, including rotation, zooming, and flipping, to enhance the diversity of the training dataset and decrease the likelihood of overfitting.
- Step 2: Techniques for Understanding
- Use LIME, SHAP, Grad-CAM, and Grad-CAM + + methods to visualize results.
- Improve understanding by gaining insights into how the models make decisions.
- Step 3: Model Selection
- Choose the top three models based on the combined criteria of performance and interpretability.
- Step 4: Ensemble Model Creation
- Combine the strengths of the best-performing models to create an ensemble model.
- Leverage ensemble diversity to enhance predictive performance and reduce overfitting.
- Optimize the ensemble model using weighted averaging to combine the predictions of the selected models effectively.
- Step 5: Classification and Validation
- Classify test images and assess the model’s robustness through cross-validation to confirm its generalizability across various data subsets.
- Evaluate model effectiveness using metrics including accuracy, sensitivity, specificity, and the geometric mean score (Gmean).
Pretrained CNNs and XAI techniques
Pretrained models have become essential tools in leveraging prelearned features and patterns for improving C-19 diagnosis from radiological images. To mitigate the risks associated with small datasets, we fine-tune these pretrained models, which have been trained on large, diverse datasets, enabling efficient transfer learning and reducing the likelihood of overfitting. Table 2 summarizes key architectural details of several widely used pretrained CNN models, including their depth, input layer size, and parameter count. In addition to utilizing pretrained models, our approach integrates XAI methods to improve AI systems’ clarity, transparency, and reliability. XAI offers understandable insights into how these models make decisions, ensuring that the model not only achieves high performance but also delivers explanations that are accessible to clinicians.
Table 2.
Pretrained CNN models’ architectural comparison.
Proposed ensemble model
The proposed model employs a structured, multi-step approach to significantly enhance classification performance for radiological images. Initially, Convolutional Neural Networks (CNNs) such as VGG16, ResNet50, and DenseNet169 are fine-tuned for binary classification targeting C-19 and NC-19 categories. These CNNs are initialized with weights from pretrained models, which have been trained on extensive and diverse datasets. This initialization leverages learned features from these broad datasets, enabling efficient transfer learning. Fine-tuning involves modifying the upper layers of the CNNs (fully connected layers or classification heads) where several new layers are added to better adapt the models to the specific classification task. The newly added layers include a layer Conv2D with filter size of 256 and kernel size of [3, 3], with a ReLU activation function. To mitigate overfitting, a dropout rate of 0.315 is applied, and batch normalization is used to stabilize training. Another layer of Conv2D with filter size of 512 and a kernel size of [3, 3] is added, along with another ReLU activation function. Following this, another layer of dropout with a 0.05 rate and batch normalization are included. A GlobalAveragePooling2D layer aggregates spatial information, and a FullyConnected layer with 2 units is added for binary classification as shown in Table 4. These additions help in extracting more relevant features and ensure better specialization for the binary classification task. Techniques such as data augmentation and adjusted learning rates are utilized to expedite convergence and enhance model performance.
In the subsequent step, a high-dimensional feature dataset is constructed based on the outputs from these fine-tuned CNNs. Each image in the training set is processed through the CNNs to obtain class-specific prediction probabilities—one for the C-19 class and one for the NC-19 class. These predictions are concatenated into a multi-dimensional feature vector for each image, creating a robust feature representation that integrates diverse model perspectives for the next phase. The third step involves training a softmax classifier, which functions as a meta-classifier, on the generated feature vectors. This meta-model synthesizes the multi-dimensional outputs from the CNNs into a unified prediction, enhancing classification accuracy. The softmax classifier is optimized using a weighted averaging technique to combine the strengths of the individual CNNs effectively while addressing their weaknesses. Additionally, cross-validation is used to improve the model’s generalization to unseen data and reduce the risk of overfitting.
Performance evaluation is a crucial component of the model validation process. Each CNN model is rigorously assessed on a separate test set, with metrics including classification accuracy, sensitivity, specificity, and geometric mean score (Gmean) being calculated. Such performance measures provide a thorough assessment of the model’s performance in handling class imbalances and ensuring reliable predictions across different classes. The entire methodology, from initial image input through CNN training, feature vector generation and meta-classifier training, is visually summarized in Fig. 2 and discussed in Algorithm 1, offering a clear and detailed schematic representation of the proposed approach.
| Algorithm 1 Ensemble of CNN models for image classification |
| Input: chest X-ray and CT image datasets. |
| Output: Classification output. |
| 1: Fine-tune pretrained CNN models using transfer learning on X-ray or CT image datasets: |
| 2: Substitute the upper layers of pretrained CNNs with newly designed top layers. |
| 3: Train the pretrained CNNs on X-ray or CT image samples from the training set: |
| 4: DenseNet169 = Train(DenseNet169, image, label) |
| 5: ResNet50 = Train(ResNet50, image, label) |
| 6: VGG16 = Train(VGG16, image, label) |
| 7: Obtain predictions from the individual trained CNNs and generate feature vectors from the softmax classifier: |
| 8: for i = 1 to N do |
| 9: [Pc−d,Pnc−v] = DenseNet169predict(imagei) |
| 10: [Pc−r, Pnc−r] = ResNet50.predict(imagei) |
| 11: [Pc−v, Pnc−d] = VGG16.predict(imagei) |
| 12: end for |
| 13: F = concatenation of [Pc−d, Pnc−d, Pc−r, Pnc−r, Pc−v, Pnc−v] |
| 14: Train softmax classifier on the feature vector F: |
| 15: Ensemble model = Train(F, label) |
| 16: Classify test images: |
| 17: predicted label = classify(Ensemble model, test images) |
| 18: return predicted labels |
| *where N indicates the total count of images in the training set. |
Fig. 2.
Diagram depicting the schematic of the proposed ensemble CNN model.
Experiments
This section covers the specifics of the datasets employed in the study, the evaluation criteria for assessing the model’s performance, and the necessary model tuning and hyperparameter optimization.
Datasets
For this research, we employed publicly available chest radiographs and CT scans from Kaggle and other repositories to create and validate our models for C-19 classification. The datasets were meticulously selected to provide a balanced representation of both C-19 and NC-19 cases. Several preprocessing steps were applied to enhance image quality and maintain consistency across datasets. Figure 3 illustrates examples of C-19 and NC-19 cases obtained from these radiographic and CT imaging datasets. Recognizing the importance of a rigorous dataset preparation process that directly impacts model reliability and accuracy, we further incorporated two external datasets—one with chest radiographs and another with CT scans—to evaluate our models’ generalization ability in diverse real-world situations and across various data sources.
Fig. 3.
Illustrative examples of COVID-19 (C-19) and non-COVID-19 (NC-19) cases from X-ray and CT image datasets.
The public datasets of chest radiographs and CT scans used in this work consist of confirmed C-19 cases, obtained from various public sources. The collected images were preprocessed to remove noise, normalize pixel values, and resize them to a consistent shape. Additionally, augmentation techniques (e.g., rotations, flips, and crops) were employed to expand the dataset size, improve model generalization, and correct class imbalances.
Internal Dataset 1: SARS-CoV-2 CT Scan Slice Dataset40
Out of 2,482 total CT scan images, 1,252 CT images were classified as positive for C-19, while the remaining 1,230 images represented patients without the virus.
Internal Dataset 2: COVIDx CXR-3 Dataset41
This dataset consists of a total of 30,386 images obtained from 16,648 patients. Among these, 16,194 images were identified as C-19 positive, while 14,192 images were categorized as C-19 negative.
External Dataset 1: Large COVID-19 CT Scan Slice Dataset42
This dataset encompasses 16,752 CT scan slices sourced from 7 different public repositories, featuring 7,593 scans marked as C-19 positive and 9,159 as NC-19. It is continually updated to include diverse images from multiple origins.
External Dataset 2: COVID-19 Radiography Dataset43
This dataset comprises 3,616 images positive for C-19 and 10,192 images of NC-19 cases, making a total of 13,808 images. The dataset is updated periodically and collected from various sources.
Data preprocessing
All images were adjusted to a size of 224 × 224 pixels to ensure uniform processing. The CT scan dataset utilized for this study consisted of preprocessed 2D slices, which were extracted from original 3D volumetric CT scans by the dataset providers. Therefore, our analysis was based solely on these preprocessed 2D slices without any additional modifications or selections from the original 3D data. For the CT scan dataset, the data was segmented into training, validation, and test subsets with a distribution ratio of 70% for training, 10% for validation, and 20% for testing. The training subset was used to build the model, the validation subset for evaluating performance and fine-tuning hyperparameters, and the test subset for final evaluation. We ensured that the distribution of positive and negative samples remained balanced across these subsets. Similarly, the CXR dataset was divided into training and validation subsets in an 80:20 ratio, with testing images provided separately. Table 3 shows the distribution of images within both the X-ray and CT scan datasets, detailing the allocation of images for training, validation, and testing.
Table 3.
Image distribution across different subsets in the chest X-ray and CT scan datasets.
| Dataset Type | Label | Chest X-ray | CT scan | ||||
|---|---|---|---|---|---|---|---|
| Train | Val | Test | Train | Val | Test | ||
| Internal | C-19 | 12,795 | 3,199 | 200 | 877 | 125 | 250 |
| NC-19 | 11,194 | 2,798 | 200 | 860 | 122 | 248 | |
| External | C-19 | 2,891 | 725 | 200 | 5,315 | 759 | 1,519 |
| NC-19 | 3,154 | 1,038 | 200 | 4,804 | 686 | 1,374 | |
| Total (CXR) | 30,034 | 7,760 | 800 | - | - | - | |
| Total (CT) | - | - | - | 11,856 | 1,692 | 3,391 | |
To improve the model’s generalization and address class imbalance, various data augmentation techniques were applied exclusively to the training sets of all datasets. These methods, including rotation, flipping, and cropping, help simulate variations that may occur in real-world scenarios, thereby enhancing the robustness of the model.
Evaluation metrics
Model’s performance evaluation metrics
To evaluate the significance of the CNN models and the introduced ensemble model, we have employed a range of evaluation metrics. These metrics are represented by Eqs. (1), (2), (3), (4), and (5). The exact mathematical expressions for each of these evaluation criteria are detailed below:
![]() |
![]() |
![]() |
![]() |
![]() |
To measure the importance and accuracy of our model in classifying C-19 case, we examine the following factors: true positive (correctly identified C-19 cases, TPc), false positive (incorrectly identified C-19 case, FPc), true negative (correctly identified non-C-19 case, TNn), and false negative (incorrectly identified non-C-19 case, FNn). True positive are instances accurately classified as C-19, while false positives are cases mistakenly labeled as C-19. True negatives represent correctly identified non-C-19 case, whereas false negatives indicate cases incorrectly labeled as non-C-19. Assessing our model’s ability to differentiate between C-19 and non-C-19 situations relies heavily on these factors.
Model’s interpretations evaluation metric
To assess the effectiveness of interpretation methods, we use the following metrics44:
Impact on Decision Making: This metric quantifies the change in classification decisions when the key regions identified by the XAI methods are excluded. Let f(z) represents the function used by the deep learning model to produce the classification outcome for an input image z, referred to as the decision function. The Decision Impact Ratio (DIR) is evaluated as:
![]() |
6 |
where
is an indicator function that equals 1 if the decision changes when the critical region rj is omitted from the image
, and 0 otherwise.
Impact on Confidence: This metric assesses the decrease in confidence scores when the key regions recognized by the XAI methods are omitted. Let γ(z) denote the confidence function of the deep learning model that calculates the probability of classification for an input image z, known as the confidence function. The Confidence Impact Ratio (CIR) is given by:
![]() |
7 |
where
represents the confidence score for the j-th image, and
represents the confidence score when the critical region
is excluded.
Model tuning and hyperparameter optimization
To make the pre-trained models work well with both CT and X-ray image datasets, we adjusted the top layer(fully connected layers or classification heads) by adding more layers. We designed the network architecture with several layers to improve the training process. These additional layers are displayed in Table 4. To achieve optimal results, the experiment utilized a learning rate and a batch size of 0.003 and 32 respectively, both of which were determined to be the most effective hyperparameters. Table 5 provides an overview of the hyperparameters used in this research for the pre-trained models. These parameters were chosen as the most effective settings for all models. Although the number of training epochs varied—100 for the CT scan images and 10 for the X-ray images—the other hyperparameter settings were uniform across both datasets. The training process aimed to minimize the loss function based on categorical cross-entropy, with optimization performed using the Adam algorithm.
Table 4.
Details of the newly added layers in the pretrained CNN models.
| Name of layer | Filters/parameters |
|---|---|
| Conv2D_1 | 256, [3,3] |
| ReLU | - |
| Dropout | 0.315 |
| BatchNormalization | - |
| Conv2D_2 | 512, [3,3] |
| ReLU | - |
| BatchNormalization | - |
| Dropout | 0.05 |
| GlobalAveragePooling2D | - |
| FullyConnected | 2 |
Table 5.
Details of the hyperparameter settings for the training of Pretrained Models.
| Hyperparameter | Optimized Value |
|---|---|
| Image size | 224X224 |
| Pooling | Global Average Pooling |
| Dropout | 0.5 |
| Activation Function | ReLU |
| Learning Rate | 0.003 |
| Batch Size | 32 |
| Optimizer | Adam |
| Epochs (CT scan) | 100 |
| Epochs (CXR-2) | 10 |
Results and discussion
This section delves into the explanations generated by various techniques, and a performance comparison among different models. The information provided contributes to a thorough understanding of the methodology and results of our research.
Performance evaluation of pretrained CNN models
The performance results in classifying C-19 and NC-19 samples using both datasets are shown in Tables 6 and 7, respectively. To showcase the performance, we have also included Figs. 4 and 5, which depict the corresponding confusion matrices. To further illustrate our findings, we have included Figs. 6 and 7, which demonstrate the ROC curves of the CNN models on both datasets.
Table 6.
Pretrained CNN model performance on the X-ray dataset.
| Model | F1-score | AUC | Sensitivity (%) | Specificity (%) | Accuracy (%) | G-mean (%) |
|---|---|---|---|---|---|---|
| VGG16 | 0.9800 | 0.9800 | 98.00% | 98.00% | 98.00% | 98.00% |
| ResNet50 | 0.9750 | 0.9700 | 97.50% | 97.50% | 97.50% | 97.49% |
| Xception | 0.9424 | 0.9500 | 94.50% | 94.50% | 94.25% | 94.44% |
| DenseNet169 | 0.9800 | 0.9800 | 98.08% | 98.08% | 98.00% | 98.06% |
| EfficientNetB3 | 0.9524 | 0.9600 | 95.51% | 95.51% | 95.25% | 95.44% |
Table 7.
Pretrained CNN model performance on chest CT image dataset.
| Model | F1-score | AUC | Sensitivity (%) | Specificity (%) | Accuracy (%) | G-mean (%) |
|---|---|---|---|---|---|---|
| VGG16 | 0.9377 | 0.9300 | 92.77% | 92.77% | 92.77% | 92.77% |
| ResNet50 | 0.9500 | 0.9600 | 95.80% | 95.80% | 95.78% | 95.79% |
| Xception | 0.9297 | 0.9300 | 92.99% | 92.99% | 92.97% | 92.99% |
| DenseNet169 | 0.9397 | 0.9400 | 94.11% | 94.11% | 93.98% | 94.07% |
| EfficientNetB3 | 0.9215 | 0.9300 | 92.52% | 92.52% | 92.17% | 92.43% |
Fig. 4.
Assessment of pre-trained CNN model performance shown through the confusion matrix for the X-ray images.
Fig. 5.
Assessment of pre-trained CNN model performance shown through the confusion matrix for the CT images.
Fig. 6.
ROC curve illustrating the outcome of the pre-trained CNN models on the X-ray images.
Fig. 7.
ROC curve illustrating the outcome of the pre-trained CNN models on the CT images.
Quantitative evaluation of explanation generated using XAI techniques
In this research, we assessed the impact of decision-making and confidence metrics for four interpretation techniques—Grad- CAM, Grad-CAM++, SHAP, and LIME—across five pretrained deep learning models: VGG16, ResNet50, DenseNet169, EfficientNetB3, and Xception. Each of these interpretation methods was used with the same set of deep learning models, which were assembled by integrating the respective pretrained image classification networks with fully connected layers. For consistency, a uniform final convolutional layer was selected as the focus for all interpretation techniques (Table 8).
Table 8.
Decision impact and confidence impact ratios for different Pretrained models and the Ensemble Model.
| XAI Techniques | VGG16 | ResNet50 | DenseNet169 | EfficientNetB3 | Xception | Ensemble Model |
|---|---|---|---|---|---|---|
| Decision Impact | ||||||
| Grad-CAM | 0.85 | 0.84 | 0.87 | 0.80 | 0.83 | 0.88 |
| Grad-CAM++ | 0.88 | 0.85 | 0.89 | 0.82 | 0.86 | 0.91 |
| SHAP | 0.82 | 0.80 | 0.85 | 0.78 | 0.83 | 0.86 |
| LIME | 0.79 | 0.77 | 0.81 | 0.75 | 0.80 | 0.82 |
| Confidence Impact | ||||||
| Grad-CAM | 0.68 | 0.66 | 0.70 | 0.65 | 0.67 | 0.71 |
| Grad-CAM++ | 0.72 | 0.69 | 0.74 | 0.68 | 0.71 | 0.76 |
| SHAP | 0.64 | 0.62 | 0.66 | 0.60 | 0.65 | 0.68 |
| LIME | 0.61 | 0.59 | 0.63 | 0.58 | 0.62 | 0.64 |
We utilized radiographic images from patients, specifically those that were correctly classified by the models. Heat maps depicting the significant regions were generated using the four interpretation techniques, with these visualizations presented in Tables 9 and 10. To measure the influence of the identified critical regions on model predictions, we computed prediction scores for each image under two conditions: the original image and a modified version where the identified critical regions were omitted. These prediction scores are detailed in Table 8. As shown in the table, the prediction scores changed significantly when the corresponding critical areas were removed, underscoring the importance of these regions in the models’ decision-making processes. This experiment was designed to measure how effectively different methods interpret model predictions in a controlled setting, where the same networks were evaluated using identical data.
Table 9.
Explanation of Pretrained CNN models on the X-ray dataset using LIME, SHAP, GRAD-CAM, and GRAD-CAM++.
Table 10.
Explanation of pretrained CNN models on the CT Scan dataset using LIME, SHAP, GRAD-CAM, and GRAD-CAM++.
To maximize the robustness and reliability of our model interpretations, we selected the top three models and interpretation methods to form an ensemble based on their performance in decision and confidence impact assessments. DenseNet169 paired with Grad-CAM + + exhibited the highest Decision Impact ratio of 0.89, indicating that Grad-CAM + + effectively identified critical regions that significantly influenced the model’s decisions. The Confidence Impact ratio was also notably high at 0.70, suggesting that the identified regions were crucial for the model’s confidence in its predictions. ResNet50, when interpreted with Grad-CAM++, demonstrated a strong Decision Impact ratio of 0.84, making it a reliable model for capturing essential image features. The Confidence Impact ratio of 0.66 confirmed that the regions identified by Grad-CAM + + were critical for the model’s confidence in its predictions. VGG16 paired with Grad-CAM + + showed a significant Decision Impact ratio of 0.88, indicating its effectiveness in identifying key regions affecting the model’s decisions, with a Confidence Impact ratio of 0.72, further validating the reliability of this model and interpretation technique.
By selecting DenseNet169, ResNet50, and VGG16, we formed an ensemble that leverages the strengths of these models and interpretation methods. This ensemble approach ensures that the most critical image regions are consistently identified, enhancing the reliability and accuracy of model predictions. The diversity in model architectures and interpretation techniques contributes to the ensemble’s robustness, making it well-suited for clinical decision-making scenarios. To thoroughly evaluate the ensemble model, we applied all four interpretation techniques—Grad-CAM, Grad-CAM++, SHAP, and LIME. Each technique provided unique insights into the critical regions influencing model predictions, with Grad-CAM + + emerging as particularly effective based on the results discussed. Consequently, we used Grad-CAM + + to illustrate the ensemble model’s results in more detail, as it offered the most comprehensive visualizations of the key regions impacting predictions.
Performance evaluation of the proposed ensemble model
We utilized 5-fold cross-validation to assess the effectiveness of our ensemble model in identifying COVID-19 from radiological images. This method ensures a more robust evaluation by dividing the dataset into five subsets, training the model on four subsets, and validating it on the remaining one. This process is repeated five times, each time with a different subset as the validation set. The performance measures, detailed in Table 11, illustrate the model’s ability to accurately classify CT and X-ray images, confirming its effectiveness in distinguishing between C-19 positive and negative cases.
Table 11.
Assessment of the Ensemble Model’s performance on internal X-ray and CT scan datasets (5-Fold evaluation).
| Dataset | Fold | Sensitivity (%) | Specificity (%) | Accuracy (%) | G-mean (%) | F1-scor | e | AUC |
|---|---|---|---|---|---|---|---|---|
| Fold 1 | 99.00 | 99.00 | 99.00 | 99.00 | 0.990 | 0.990 | ||
| Fold 2 | 99.20 | 99.00 | 99.10 | 99.10 | 0.990 | 0.990 | ||
| COVIDx CXR-3 | Fold 3 | 99.10 | 98.90 | 99.00 | 98.90 | 0.989 | 0.989 | |
| Fold 4 | 98.90 | 98.70 | 98.80 | 98.80 | 0.988 | 0.988 | ||
| Fold 5 | 99.00 | 98.80 | 98.90 | 98.90 | 0.989 | 0.989 | ||
| Mean | 99.04 ± 0.13 | 99.08 ± 0.12 | 99.02 ± 0.11 | 99.10 ± 0.11 | 0.989 ± | 0.001 | 0.989 ± 0.001 | |
| Fold 1 | 96.00 | 95.90 | 95.95 | 96.00 | 0.9600 | 0.960 | ||
| Fold 2 | 96.10 | 96.00 | 96.05 | 96.10 | 0.9610 | 0.965 | ||
| SARS-COV-2 | Fold 3 | 96.18 | 96.18 | 96.18 | 96.18 | 0.9618 | 0.970 | |
| Fold 4 | 95.90 | 95.70 | 95.80 | 95.90 | 0.9590 | 0.955 | ||
| Fold 5 | 96.00 | 96.00 | 96.00 | 96.00 | 0.9600 | 0.960 | ||
| Mean | 96.04 ± 0.12 | 95.96 ± 0.14 | 96.00 ± 0.12 | 96.05 ± 0.10 |
0.9605 0.002 |
± | 0.962 ± 0.005 |
Visual performance comparison
A simple visual comparison of our ensemble model to pre-trained CNN models on the CXR and computed tomography datasets is displayed in Figs. 8 and 9. Furthermore, comprehensive confusion matrices and ROC are provided in Fig. 10 to offer an additional understanding of the performance of our model.
Fig. 8.
Contrasting the Performance of Ensemble Model with Pre-trained CNNs on the X-ray Dataset.
Fig. 9.
Contrasting the Performance of Ensemble Model with Pre-trained CNNs on Chest CT Dataset.
Fig. 10.
Visualization of the confusion matrix and ROC curve depicting the performance of the proposed ensemble model on CT and X-ray image datasets.
Comprehensive analysis and comparison
An extensive comparison of our model’s efficiency metrics with previous studies for CXR and computed tomography image datasets is presented in Tables 12 and 13. To further demonstrate the robustness of the proposed ensemble model, we performed a cross-dataset evaluation. In this evaluation, the model was trained on internal datasets and then tested on external datasets to assess its generalizability in classifying CXR and computed tomography images into C-19 and NC-19 categories. The results of this cross-dataset evaluation are summarized in Table 14.
Table 12.
Comparison of Model Performance Metrics for X-ray dataset.
| Author(s) & Year | Architecture / Models | Image Type | Performance Parameters (%) |
|---|---|---|---|
| Liu and Shen45 | CECT: Controllable ensemble CNN and transformer | C-19 radiography dataset, COVIDx CXR-3 dataset | Accuracy: 98.1% (intra-dataset), 90.9% (unseen) |
| Dey et al.46 | VGG19,InceptionV3, MobileNet | chest-xray- pneumonia, COVIDx CXR-3, covid-chest x-ray-dataset | Accuracy: 98.62%, F1-Score: 98.67% |
| Eshraghi et al.47 | COV-MobNets Ensemble Model | COVIDCXR-3, CXR Images (Pneumonia) | Accuracy: 97.75% |
| Abad et al.48 | Ensemble: ResNet50, DenseNet121, Inception-ResNet-v2 | COVIDxCXR-3, COVIDGR, Labeled OpticalCoherence Tomography (OCT) and CXR Images | Accuracy: 97.38 (Internal), 81.18% (External) |
| Proposed | Ensemblemodel: VGG16, ResNet50, DenseNet169 | COVIDx CXR-3 | Accuracy: 99% (Internal Dataset) |
Table 13.
Comparison of Model Performance Metrics for CT scan dataset.
| Author (s) & Year | Architecture\Models | Image type | Performance parameters (%) |
|---|---|---|---|
| Panwar et al.49 | VGG19 | COVID-chest X- ray, SARS-COV-2 CT-scan, and Chest X-Ray Images (Pneumonia) | Accuracy: 95.61% |
| Silva et al.50 | EfficientNet with Voting-based approach | COVID-CT dataset, SARS- CoV-2 CT-scan | Accuracy: 87.68% (Single Dataset), 56.16% Cross-Dataset) |
| Yang et al.51 | VGG16, DenseNet121, ResNet50, ResNet152 | SARS-CoV-2 CT Scan Dataset | Accuracy: CT-Scan 96% |
| Proposed Model | Ensemble model: VGG16, ResNet50, DenseNet169 | SARS-CoV-2 CT Scan Dataset | Accuracy: 96% (Internal Dataset) |
Table 14.
Cross-dataset evaluation results: accuracy for C-19 and NC-19 classification.
Limitations
Our study demonstrates promising results for C-19 detection using radiological images, but there are several limitations to consider. First, the datasets used primarily consist of images from specific geographic regions and demographic groups, which may not represent a global population. This limitation could introduce biases in the model’s performance when applied to images from different populations or healthcare settings. Future work should consider using a more diverse dataset to enhance the model’s generalizability.
Second, the interpretability methods used in our study, such as Grad-CAM, Grad-CAM++, SHAP, and LIME, present certain challenges. These methods can be computationally expensive, especially when applied to ensemble models, leading to increased processing times. Additionally, the clarity of the explanations provided by these methods may vary depending on the model’s complexity, potentially making it difficult for clinicians to fully understand the model’s predictions.
Conclusion and future work
In our investigation, we highlighted the importance of providing detailed explanations for pre-trained models before integrating them into ensemble methods, particularly for identifying C-19 from radiological images. Our proposed approach combines a softmax classifier with pre-trained models—ResNet50, DenseNet169, and VGG16—within a deep CNN ensemble framework. This method effectively extracts diverse types of information from chest radiographs and unique visual features. Our evaluation on two datasets showed that the ensemble model achieved classification accuracies of 96.18% for CT images and 99.00% for chest X-rays. These results, obtained using transfer learning and data synthesized from the COVID-Net dataset, demonstrate significant improvements over existing methods that rely solely on X-ray imaging for C-19 diagnosis. This underscores the importance of using interpretable pre-trained models to enhance the performance of C-19 image classification.
Despite these advancements, further investigation is required. Firstly, our study’s findings need additional medical validation to ensure that the explanations provided by our models are consistent with clinical interpretations and practices. Moreover, the generalizability of our approach to other patient populations and imaging sites needs to be evaluated. Testing across diverse clinical settings and varying image qualities is essential to assess the model’s robustness and real-world applicability.
Author contributions
R.R. and M.G. conceived the study, designed the methodology, and conducted the primary research efforts data preprocessing, model implementation, and performance evaluation. S.J. and V.B.S. contributed to reviewing the methodology, providing feedback on experimental design, and assisting with data interpretation. All authors participated in the discussion of results and contributed to the preparation of the manuscript.
Funding
Funding was not provided to the authors for this work.
Data availability
The datasets generated and/or analyzed during the current study are available in the SARS-COV-2 CT-Scan Kaggle Dataset repository at [https://www.kaggle.com/datasets/plameneduardo/sarscov2-ctscan-dataset] and the COVIDx CXR-3 Kaggle Dataset repository at [https://www.kaggle.com/datasets/mahy143/covid19xraydataset? rvi=1]. Researchers can access and utilize these public datasets for further analysis and investigation related to COVID-19 and its impacts on medical imaging.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.World Health Organization. Coronavirus disease (covid-19): Post-covid-19 condition. https://www.who.int/health-topics/coronavirus. Accessed: 2024-02-07.
- 2.Fang, Y. et al. Sensitivity of chest CT for COVID-19: comparison to RT-PCR. Radiology 296(2), E115–E117 (2020). [DOI] [PMC free article] [PubMed]
- 3.Ai, T. et al. Correlation of chest CT and RT-PCR testing in coronavirus disease 2019 (COVID-19) in China: a report of 1014 cases. Radiology 296(2), E32–E40 (2020). [DOI] [PMC free article] [PubMed]
- 4.Ng, M. Y. et al. Imaging profile of the covid-19 infection: radiologic findings and literature review. Radiol. Cardiothorac. Imaging2(1), e200034 (2020). [DOI] [PMC free article] [PubMed]
- 5.Kanne, J. P., Little, B. P., Chung, J. H., Elicker, B. M. & Ketai, L. H. Essentials for radiologists on covid-19: an update—radiology scientific expert panel. Radiology 296(2), E113–E114 (2020). [DOI] [PMC free article] [PubMed]
- 6.Zhang, J. et al. Recent developments in segmentation of covid-19 ct images using deep-learning: an overview of models, techniques and challenges. Biomed. Signal. Process. Control. 91. 10.1016/j.bspc.2024.105970 (2024).
- 7.Dai, H., Yang, Y., Yue, X. & Chen, S. Improving retinal oct image classification accuracy using medical pre-training and sample replication methods. Biomed. Signal. Process. Control. 91. 10.1016/j.bspc.2024.106019 (2024).
- 8.Zhang, L. et al. Deep learning model based on primary tumor to predict lymph node status in clinical stage ia lung adenocarcinoma: a multicenter study. J. Natl. Cancer Cent.. 10.1016/j.jncc.2024.01.005 (2024). [DOI] [PMC free article] [PubMed]
- 9.Mary, A. R. & Kavitha, P. Diabetic retinopathy disease detection using shapley additive ensembled densenet-121 resnet-50 model. Multimed Tools Appl. 1–28. 10.1007/s11042-024-18309-6 (2024).
- 10.Nehru, V. & Prabhu, V. Automated multimodal brain tumor segmentation and localization in mri images using hybrid res2-unext. J. Electr. Eng. Technol. 1–13. 10.1007/s42835-023-01779-3 (2024).
- 11.Pathan, S., Kumar, P., Pai, R. M. & Bhandary, S. V. An automated classification framework for glaucoma detection in fundus images using ensemble of dynamic selection methods. Prog Artif. Intell.12, 287–301. 10.1007/s13748-023-00304-x (2023).
- 12.Simonyan, K. & Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).
- 13.He, K., Zhang, X., Ren, S. & Sun, J. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell.37, 1904–1916. 10.1109/TPAMI.2015.2389824 (2015). [DOI] [PubMed]
- 14.Huang, G., Liu, Z., Pleiss, G., Van Der Maaten, L. & Weinberger, K. Q. Convolutional networks with dense connectivity. IEEE Trans. Pattern Anal. Mach. Intell.44, 8704–8716. 10.1109/TPAMI.2019.2918284 (2019). [DOI] [PubMed]
- 15.Tan, M., Le, Q. & Efficientnet Rethinking model scaling for convolutional neural networks. In International conference on machine learning, 6105–6114. 10.48550/arXiv.1905.11946 (2019).
- 16.Chollet, F. & Xception Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1251–1258. 10.48550/arXiv.1610.02357 (2017).
- 17.Ribeiro, M. T., Singh, S. & Guestrin, C. Why should i trust you? explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, 1135–1144. 10.1145/2939672.2939778 (2016).
- 18.Lundberg, S. M. & Lee, S. I. A unified approach to interpreting model predictions. Adv. Neural Inform. Process. Syst.30. 10.48550/arXiv.1705.07874 (2017).
- 19.Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, 618–626. 10.1007/s11263-019-01228-7 (2017).
- 20.Chattopadhay, A., Sarkar, A., Howlader, P. & Balasubramanian, V. N. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In IEEE winter conference on applications of computer vision (WACV), 839–847. 10.1109/WACV.2018.00097 (IEEE, 2018).
- 21.Prinzi, F., Militello, C., Scichilone, N., Gaglio, S. & Vitabile, S. Explainable machine-learning models for covid-19 prognosis prediction using clinical, laboratory and radiomic features. IEEE Access11, 121492–121510. 10.1109/ACCESS.2023.3327808 (2023). [Google Scholar]
- 22.Soda, P. et al. Aiforcovid: Predicting the clinical outcomes in patients with covid-19 applying ai to chest-x-rays. An Italian multicentre study. Med. Image Anal.74. 10.1016/j.media.2021.102216 (2021). [DOI] [PMC free article] [PubMed]
- 23.Sun, Y. et al. Use of machine learning to assess the prognostic utility of radiomic features for in-hospital covid-19 mortality. Sci. Rep.13. 10.1038/s41598-023-34559-0 (2023). [DOI] [PMC free article] [PubMed]
- 24.Abubakar, H., Al-Turjman, F., Ameen, Z. S., Mubarak, A. S. & Alturjman, C. A hybridized feature extraction for covid-19 multi-class classification on computed tomography images. Heliyon. 10.1016/j.heliyon.2024.e26939 (2024). [DOI] [PMC free article] [PubMed]
- 25.Ragab, D. A., Fayed, S., Ghatwary, N. & Deepcsfusion Deep compressive sensing fusion for efficient covid-19 classification. J. Imaging Inf. Med.1–13. 10.1007/s10278-024-01011-2 (2024). [DOI] [PMC free article] [PubMed]
- 26.Haynes, S. C., Johnston, P. & Elyan, E. Generalisation challenges in deep learning models for medical imagery: insights from external validation of covid-19 classifiers. Multimed Tools Appl. 1–20. 10.1007/s11042-024-18543-y (2024).
- 27.Suhartanto, H. et al. Scov-cnn: a simple cnn architecture for covid-19 identification based on the ct images. JOIV: Int. J. Inf. Vis.8. 10.62527/joiv.8.1.1750 (2024).
- 28.Zhao, A., Wu, H., Chen, M. & Wang, N. A multi-level feature attention network for covid-19 detection based on multi-source medical images. Multimed Tools Appl. 1–32. 10.1007/s11042-023-18014-w (2024).
- 29.Abdellatef, E. & Allah, M. F. Hybrid whale optimization and canonical correlation based covid-19 classification approach. Multimed Tools Appl. 1–22. 10.1007/s11042-024-18153-8 (2024).
- 30.Hoffer, O. et al. Smartphone-based detection of covid-19 and associated pneumonia using thermal imaging and a transfer learning algorithm. J. Biophotonics. e202300486. 10.1002/jbio.202300486 (2024). [DOI] [PubMed]
- 31.Sanampudi, A. & Srinivasan, S. Local search enhanced optimal inception-resnet-v2 for classification of long-term lung diseases in post-covid-19 patients. Automatika. 65, 473–482. 10.1080/00051144.2023.2295142 (2024).
- 32.Zafar, A. et al. Robust medical diagnosis: a novel two-phase deep learning framework for adversarial proof disease detection in radiology images. J. Imaging Inf. Med.1–31. 10.1007/s10278-023-00916-8 (2024). [DOI] [PMC free article] [PubMed]
- 33.Türk, F. & Kökver, Y. Detection of lung opacity and treatment planning with three-channel fusion cnn model. Arab. J. Sci. Eng. 1–13. 10.1007/s13369-023-07843-4 (2023). [DOI] [PMC free article] [PubMed]
- 34.Saheb, S. K., Narayanan, B. & Rao, T. V. N. Adl-cdf: a deep learning framework for covid-19 detection from ct scans towards an automated clinical decision support system. Arab. J. Sci. Eng.48, 9661–9673. 10.1007/ s13369-022-07271-w (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Holzinger, A., Biemann, C., Pattichis, C. S. & Kell, D. B. What do we need to build explainable ai systems for the medical domain? arXiv Preprint arXiv:1712 09923. 10.48550/arXiv.1712.09923 (2017).
- 36.Koul, A., Bawa, R. K. & Kumar, Y. Enhancing the detection of airway disease by applying deep learning and explainable artificial intelligence. Multimed Tools Appl. 1–33. 10.1007/s11042-024-18381-y (2024).
- 37.Chadaga, K. et al. A decision support system for diagnosis of covid-19 from non-covid-19 influenza-like illness using explainable artificial intelligence. Bioengineering. 10, 439. 10.3390/bioengineering10040439 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778. 10.48550/arXiv.1512.03385 (2016).
- 39.Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4700–4708. 10.48550/arXiv.1608.06993 (2017).
- 40.Soares, E., Angelov, P., Biaso, S., Froes, M. H. & Abe, D. K. Sars-cov-2 ct-scan dataset: a large dataset of real patients ct scans for sars-cov-2 identification. MedRxiv. 10.1101/2020.04.24.20078584 (2020).33173903 [Google Scholar]
- 41.Pavlova, M. et al. Covid-net cxr-2: an enhanced deep convolutional neural network design for detection of covid-19 cases from chest x-ray images. Front. Med.9. 10.3389/fmed.2022.861680 (2022). [DOI] [PMC free article] [PubMed]
- 42.Maftouni, M. et al. A robust ensemble-deep learning model for covid-19 diagnosis based on an integrated ct scan images database. In IIE annual conference. Proceedings, 632–637 (Institute of Industrial and Systems Engineers (IISE), (2021).
- 43.Rahman, T. & collaborators. COVID-19 Chest X-Ray Database. (2022). https://www.kaggle.com/tawsifurrahman/covid19-radiography-database Accessed: 2024-08-11.
- 44.Zou, L. et al. Ensemble image explainable ai (xai) algorithm for severe community-acquired pneumonia and covid-19 respiratory infections. IEEE Trans. Artif. Intell.4, 242–254. 10.1109/TAI.2022.3153754 (2022). [Google Scholar]
- 45.Liu, Z., Shen, L. & Cect Controllable ensemble cnn and transformer for covid-19 image classification. Comput. Biol. Med.173. 10.1016/j.compbiomed.2024.108388 (2024). [DOI] [PubMed]
- 46.Dey, S. et al. A fuzzy ensemble model for covid-19 detection from chest x-rays. Expert Syst. Appl.206. 10.1016/j.eswa.2022.117812 (2022). [DOI] [PMC free article] [PubMed]
- 47.Eshraghi, M. A., Ayatollahi, A. & Shokouhi, S. B. Cov-mobnets: a mobile networks ensemble model for diagnosis of covid-19 based on chest x-ray images. BMC Med. Imaging. 23, 83. 10.1186/s12880-023-01039-w (2023). [DOI] [PMC free article] [PubMed]
- 48.Abad, M., Casas-Roma, J. & Prados, F. Generalizable disease detection using model ensemble on chest x-ray images. Sci. Rep.14. 10.1038/s41598-024-56171-6 (2024). [DOI] [PMC free article] [PubMed]
- 49.Panwar, H. et al. A deep learning and grad-cam based color visualization approach for fast detection of covid-19 cases using chest x-ray and ct-scan images. Chaos Solitons Fractals. 140, 110190. 10.1016/j.chaos.2020.110190 (2020). [DOI] [PMC free article] [PubMed]
- 50.Silva, P. et al. Covid-19 detection in ct images with deep learning: a voting-based scheme and cross-datasets analysis. Inf. Med. Unlocked. 20. 10.1016/j.imu.2020.100427 (2020). [DOI] [PMC free article] [PubMed]
- 51.Yang, D. et al. Detection and analysis of covid-19 in medical images using deep learning techniques. Sci. Rep.11. 10.1038/s41598-021-99015-3 (2021). [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets generated and/or analyzed during the current study are available in the SARS-COV-2 CT-Scan Kaggle Dataset repository at [https://www.kaggle.com/datasets/plameneduardo/sarscov2-ctscan-dataset] and the COVIDx CXR-3 Kaggle Dataset repository at [https://www.kaggle.com/datasets/mahy143/covid19xraydataset? rvi=1]. Researchers can access and utilize these public datasets for further analysis and investigation related to COVID-19 and its impacts on medical imaging.



















