Abstract
Saliency maps are popularly used to “explain” decisions made by modern machine learning models, including deep convolutional neural networks (DCNNs). While the resulting heatmaps purportedly indicate important image features, their “trustworthiness,” i.e., utility and robustness, has not been evaluated for musculoskeletal imaging. The purpose of this study was to systematically evaluate the trustworthiness of saliency maps used in disease diagnosis on upper extremity X-ray images. The underlying DCNNs were trained using the Stanford MURA dataset. We studied four trustworthiness criteria—(1) localization accuracy of abnormalities, (2) repeatability, (3) reproducibility, and (4) sensitivity to underlying DCNN weights—across six different gradient-based saliency methods (Grad-CAM (GCAM), gradient explanation (GRAD), integrated gradients (IG), Smoothgrad (SG), smooth IG (SIG), and XRAI). Ground-truth was defined by the consensus of three fellowship-trained musculoskeletal radiologists who each placed bounding boxes around abnormalities on a holdout saliency test set. Compared to radiologists, all saliency methods showed inferior localization (AUPRCs: 0.438 (SG)–0.590 (XRAI); average radiologist AUPRC: 0.816), repeatability (IoUs: 0.427 (SG)–0.551 (IG); average radiologist IOU: 0.613), and reproducibility (IoUs: 0.250 (SG)–0.502 (XRAI); average radiologist IOU: 0.613) on abnormalities such as fractures, orthopedic hardware insertions, and arthritis. Five methods (GCAM, GRAD, IG, SG, XRAI) passed the sensitivity test. Ultimately, no saliency method met all four trustworthiness criteria; therefore, we recommend caution and rigorous evaluation of saliency maps prior to their clinical use.
Supplementary Information
The online version contains supplementary material available at 10.1007/s10278-024-01136-4.
Keywords: Saliency methods, Trustworthy AI, Musculoskeletal radiology, Deep learning
Introduction
Deep learning algorithms are being increasingly explored for their potential to assist radiologists in disease diagnosis from medical images [1]. A class of these algorithms known as deep convolutional neural networks (DCNNs) map input images to output classifications by using convolutional filters that spatially process image information [2]. Different DCNNs are characterized by architectural differences such as the different kinds of layers used (e.g., residual layers, auxiliary networks), the depth of the network, the order of the layers, among other architectural choices. DCNNs have had success on a variety of radiology tasks, ranging from abnormality classification on chest X-rays [3, 4] to brain tumor segmentation [5].
As deep learning algorithms transition from academic innovations to clinically implemented tools, the capacity to interpret model predictions (i.e., understand why and how a decision was made) is increasingly important in order to build physician trust in an algorithm [6]. This is an important challenge in the AI research community given the “black-box” nature of many deep learning algorithms. For example, DCNNs trained to diagnose disease from X-ray images taken at one hospital have not necessarily generalized to other hospital settings because models can incorrectly focus on confounding signals (e.g., hardware meta-data) rather than clinical features [7]. Model interpretability can help engineers and physicians debug these issues before reaching clinical deployment, or even reduce the amount of samples required for training [8].
Saliency methods are the current standard to produce visual explanations of a DCNN’s decision-making. These methods generate post hoc—i.e., after training the DCNN—heatmaps that quantify at the pixel level which regions were considered significant during classification of an image. These were initially, and continue to be, employed to supplement DCNN predictions with visual intuitions of “why” these decisions were made [3, 9–12]. However, recent studies have underscored issues with saliency maps as reliable explanation tools. Adebayo et al. and Zhang et al. established that common saliency methods are not resistant to perturbations in the input image [13, 14]. Arun et al. proposed a framework for quantifying the “trustworthiness” of saliency maps in the context of chest X-ray diagnosis. The constituent criteria included localization of disease abnormalities, a fundamental measurement that probed whether saliency heatmaps aligned with clinical reasoning [15].
One remaining question from these studies is how notions of saliency map trustworthiness apply to different clinical data types, especially since the prevailing analysis has focused on chest X-rays [14, 16]. Jin et al. approached this from a perspective of multi-model imaging: the authors analyzed if saliency maps carry information relevant to a particular MRI modality [17]. The results suggested that many saliency methods failed to correlate for modality-specific features. However, the alternative perspective remains to be explored: how saliency maps compare between different clinical presentations within the same modality and task.
The purpose of our study was to systematically evaluate the trustworthiness of saliency maps as visual explanations for deep learning algorithms that identify abnormalities on upper extremity radiographs. To our knowledge, this kind of analysis has not yet been explored in the context of musculoskeletal X-rays. We additionally probed how saliency trustworthiness differed across disease subtypes. Taken with the insights from Jin et al. [17], these experiments can shape a more comprehensive picture of how saliency maps vary across clinical presentations.
Materials and Methods
Data
Our experiments used the Stanford MURA dataset [18]: a large open-source collection of musculoskeletal radiograph studies made available for research use by Rajpurkar et al. in 2018. All radiographs were obtained from the Picture Archive and Communications System at the Stanford Hospital between 2001 and 2012 and were de-identified prior to public release. At the time of initial interpretation, board-certified Stanford radiologists annotated each radiographic study with a binary label indicating whether an abnormality was present. Rajpurkar et al. retrospectively collected these studies and divided the overall dataset into a training set (13,457 studies, 36,808 images), validation set (1199 studies, 3197 images), and test set (207 studies, 556 images), ensuring no patient overlap. However, only the training set and validation set were publicly released; the original test set was hidden for online competition purposes.
We only used the publicly available MURA data—i.e., the original training and validation sets—to train predictive DCNN models. Our training set was taken identically from the original dataset. We divided the originally provided validation set using a 60/40% split at the patient level to create our own validation and testing sets. A patient level split was done to avoid data leakage between the splits. Table 1 provides a complete description of our training, validation, and testing sets.
Table 1.
Description of (1) in-house training/validation/testing splits by number of images, studies, patients, and distribution of anatomical regions and (2) the saliency test set by number of images, studies, patients, and distribution of anatomical regions and abnormality subgroups
| Training set | Validation set | Testing set | Saliency test set | ||
|---|---|---|---|---|---|
| Number of images | 36,808 | 1886 | 1311 | 588 | |
| Number of studies | 13,457 | 703 | 496 | 217 | |
| Number of patients | 11,184 | 469 | 314 | 189 | |
| Region, n (%) | Elbow | 4931 (13.4) | 274 (14.5) | 191 (14.6) | 95 (16.2) |
| Finger | 5106 (13.9) | 278 (14.7) | 183 (14.0) | 105 (17.9) | |
| Forearm | 1825 (5.0) | 176 (9.3) | 125 (9.5) | 61 (10.4) | |
| Hand | 5543 (15.1) | 273 (14.5) | 187 (14.3) | 63 (10.7) | |
| Humerus | 1272 (3.5) | 173 (9.2) | 115 (8.8) | 56 (9.5) | |
| Shoulder | 8379 (22.8) | 332 (17.6) | 231 (17.6) | 108 (18.4) | |
| Wrist | 9752 (26.5) | 380 (20.1) | 279 (21.3) | 100 (17.0) | |
| Subgroup, n (%) | Arthritis | n/a | n/a | n/a | 197 (33.5) |
| Hardware/fracture | n/a | n/a | n/a | 367 (62.4) | |
| Other | n/a | n/a | n/a | 24 (4.1) |
Our institutional review board acknowledged our study as non-human subjects research because our work used publicly available data only.
Radiologist Annotations
We curated spatial annotations for a subset of MURA radiographs to establish a clinical baseline for abnormality identification. Three board-certified and fellowship-trained musculoskeletal radiologists annotated where abnormalities were present in each test set image that was in fact labeled for presence of an abnormality (638 positive label images out of 1311 test set images). Radiologists knew a priori each image was positively labeled and were instructed to place rectangles (and as needed, free-form polygons) surrounding regions where they observed an abnormality. The ground-truth annotation for each image was obtained by taking a majority vote at the pixel level across the three radiologists’ annotations (Fig. 1a). After making these spatial annotations, radiologists were finally instructed to provide a succinct textual description of the kind of abnormality. These descriptions were not already available in the public MURA release; we used this information in subgroup analysis of saliency map trustworthiness.
Fig. 1.
Overview of trustworthiness framework. Images were annotated (a) by a team of musculoskeletal-trained radiologists and a consensus annotation was derived using a pixel-level majority vote. Four trustworthiness criteria—localization accuracy (b), heatmap consistency (includes both repeatability and reproducibility) (c), and heatmap sensitivity (d)—were quantitatively proposed and evaluated
For 50 of the 638 images, the derived ground-truth annotation was empty. These were excluded from trustworthiness evaluations due to ambiguity; therefore, the size of the annotated subset—i.e., “saliency test set”—was 588 images (Table 1). The remaining images were tagged with one of three musculoskeletal abnormality subgroups: arthritis, hardware/fracture, or “other.” Tagging was based on one radiologist’s textual annotation, and ambiguities were clarified based on annotations from the other two radiologists. Hardware and fracture cases were combined due to frequent co-occurrence. The “other” category included tumors, lacerations, effusions, metaphyseal lucent lines, and amputations. This curated saliency test set was used in trustworthiness experiments.
DCNN Training and Evaluation
We trained DCNNs to predict the presence of an abnormality from individual radiographs. Our base classification models were DCNNs fine-tuned from ImageNet weights. We used the InceptionV3 [19] and DenseNet169 [20] architectures due to their prior success on medical imaging tasks [3, 4, 18, 21] and previous use in saliency map trustworthiness evaluations for chest X-ray diagnosis [15]. Standard data augmentation and preprocessing routines were consistently applied for input images (see Supplemental S.1 for details). Images were resized to 320 × 320 pixels to match prior experiments with the MURA dataset [18]. Hyperparameters were selected to maximize area under the receiver operating characteristic curve (AUROC) on our validation set (Supplemental S.1). Once optimal hyperparameters were identified, four final models—three InceptionV3s and one DenseNet169—were trained on the union of our training and validation sets.
All DCNNs were evaluated on the testing set, which consisted of radiographic studies distinct from those in the training and validation sets. This lack of image overlap ensures that the following test statistics were unbiased and independent (from, for example, hyperparameter tuning on the validation set). For each anatomical region, the AUROC, accuracy, precision, recall, and F1 scores were computed at the threshold yielding maximum accuracy. An overall weighted AUROC (wAUROC) was reported as the weighted average of individual anatomical region AUROCs. The weights were the proportion of abnormal test set images from each anatomical region. Other metrics were micro-averaged to yield overall performance metrics. Bootstrapping with 1000 resamples was implemented to estimate 95% confidence intervals for each metric. Pairwise Tukey tests were used to statistically compare DCNN classification performance. Model training and statistical analyses were done using Python.
Trustworthiness Evaluation
We measured the “trustworthiness” of saliency maps in visualizing DCNN abnormality predictions on a musculoskeletal radiograph. Our operative definition of “trustworthiness” was adapted from Arun et al. [15]. We examined four criteria (described in the following subsections): (1) localization accuracy, (2) repeatability, (3) reproducibility, and (4) sensitivity (Fig. 1). We additionally contribute a novel formulation of weak and strong versions for the first three criteria that represent sanity checks and radiologist-level comparisons, respectively. All trustworthiness experiments employed only images from the saliency test set, i.e., the annotated subset of positive label images in the test set.
Six saliency gradient-based methods were tested: gradient-weighted class activation mapping (GCAM) [22], gradient explanation (GRAD) [23], integrated gradients (IG) [24], Smoothgrad (SG) [25], smooth integrated gradients (SIG) [24, 25], and XRAI (XRAI) [26]. (Guided BackProp [27] and Guided GradCAM [22, 28] were excluded because Adebayo et al. [13] showed these methods lacked sensitivity to higher-level network weights. Hooker et al. [29] additionally showed that Guided BackProp had worse feature importance estimations than random assignment.) To ensure fair comparisons, saliency maps were preprocessed before conducting experiments to match the radiologists’ annotation styles (Supplementary S.2).
Localization Accuracy
Localization measured the spatial concordance between heatmaps and ground-truth annotations. Saliency maps were derived from the best-performing InceptionV3 classifier (as determined by wAUROC). We quantified localization accuracy using the pixel-level AUPRC distribution between ground-truth annotations and saliency maps (Fig. 1b). To fairly evaluate the clinical utility of saliency localizations, we only considered radiographs in the saliency test set that were correctly classified by the InceptionV3 DCNN. Indeed, if the underlying prediction was incorrect, there is a confounding factor that can limit saliency map localization accuracy.
In the weak version of this criteria, we compared the localization accuracy of saliency maps to that of raw images passed through a Canny edge detector [30] (with additional preprocessing similar to the saliency maps). The rationale was that saliency maps should be at least as accurate as an edge detector that can only identify edge-patterned abnormalities without any learning. In the strong version, we compared against the average of each radiologist’s localization accuracy with respect to the consensus annotation. Pairwise Tukey tests were performed to determine if saliency map AUPRCs were statistically higher than or comparable to that of its associated comparison in the weak and strong versions, respectively. Additionally, tests were performed at the abnormality subgroup level to identify potential hidden stratification.
Heatmap Consistency
Heatmap consistency included both repeatability and reproducibility scores. Repeatability tests evaluated the consistency of saliency maps between two InceptionV3 DCNNs. In contrast, reproducibility scores measured the consistency of saliency maps between an InceptionV3 and DenseNet169 DCNN. The two best-performing InceptionV3 DCNNs were chosen for repeatability experiments, and the best InceptionV3 was chosen to compare against the DenseNet169 for reproducibility experiments. In both cases, the compared models had similar weighted AUROCs. We quantified consistency using the structural similarity index measure (SSIM), intersection-over-union (IoU), and binarized pixel error. This combination of metrics evaluated the similarity of saliency maps from three distinct perspectives: spatial structural information, regional overlap, and naïve per-pixel accuracy (Fig. 1c). In the weak formulation, we compared each metric’s mean to a fixed threshold of 0.5. A saliency method passed the weak criteria only if every metric was significantly greater (in the case of SSIM and IoU) or lesser (in the case of pixel error) than this threshold. In the strong version, we set the threshold to be the mean radiologist concordance with ground-truth annotations. A saliency method passed the strong criteria if each metric was not significantly lesser (in the case of SSIM and IoU) or greater (in the case of pixel error) than its radiologist-derived threshold. One-sample one-sided t-tests were performed to make each comparison.
Heatmap Sensitivity
Heatmap sensitivity examined how saliency maps changed in response to randomization of the underlying model weights. Trustworthy saliency methods should demonstrate variance with respect to the trained DCNN: if successive DCNN layers are reinitialized to random values, then the associated saliency maps should degrade in quality. This process of accumulated randomization is known as cascading randomization [13, 15]. We evaluated this condition on the InceptionV3 architecture over a small subset of 100 images from the saliency test set. Three InceptionV3s were each divided into 17 layers and successively randomized from top (i.e., the logit layer) to the bottom (i.e., the first convolutional layer). At each step, saliency maps from a partially randomized model were compared using SSIM to those of its non-randomized fully trained parent model (Fig. 1d). The sequence of SSIM distributions was plotted to visualize the effects of cascading randomization.
The sensitivity criterion was obtained by comparing the similarity scores of the fully randomized model to a saliency method-specific degradation threshold. This threshold was based on the average pairwise similarity of saliency maps derived from the fully trained models. A saliency method passed if its similarity scores were significantly lesser than the degradation threshold; otherwise, the saliency methods demonstrated unwanted invariance to model training. One-sample one-sided t-tests were used to make these comparisons.
Results
The best-performing InceptionV3 model and the DenseNet169 model reported wAUROCs of 0.904 (LCI = 0.887, UCI = 0.921) and 0.905 (LCI = 0.887, UCI = 0.922), respectively, on the testing set (Table 2). No statistical differences were observed between these wAUROCs (p > 0.1). These scores are comparable to those reported in the original MURA study [18] and more recent experiments [31].
Table 2.
DCNN classification performances of the four DCNNs (three InceptionV3s and one DenseNet169) used in trustworthiness experiments, as measured by anatomical region-weighted AUROC (wAUROC) and Cohen’s kappa statistic. One thousand-sample bootstrapping was implemented to obtain 95% confidence intervals (lower-bound, LCI; upper-bound, UCI). Superscript letters indicate statistical comparisons within the same metric using pairwise Tukey tests at significance level
| DCNN architecture | Weighted AUROC (wAUROC) | Cohen’s kappa statistic | ||||
|---|---|---|---|---|---|---|
| Mean | LCI | UCI | Mean | LCI | UCI | |
| InceptionV3 | 0.904a | 0.887 | 0.921 | 0.680c | 0.641 | 0.719 |
| 0.902b | 0.884 | 0.918 | 0.689b | 0.649 | 0.726 | |
| 0.898c | 0.880 | 0.914 | 0.680c | 0.641 | 0.717 | |
| DenseNet169 | 0.905a | 0.887 | 0.922 | 0.701a | 0.658 | 0.736 |
All saliency methods passed the weak localization test and failed the strong localization test (Fig. 2a). XRAI scored the highest localization AUPRC (mean = 0.590, SD = 0.154), and SG scored the lowest (mean = 0.438, SD = 0.203). Comparisons against the edge detector baseline (mean = 0.191, SD = 0.159) and radiologist annotations (mean = 0.816, SD = 0.100) were significant (p < 0.001). Stratified analysis by the kind of abnormality revealed that hardware/fracture cases had better heatmap localizations compared to arthritis cases (Fig. 2b). These differences were prominent for IG, where the difference in mean AUPRC between hardware and arthritis cases was 0.174 (hardware, mean = 0.557, SD = 0.207; arthritis, mean = 0.383, SD = 0.248). Nonetheless, despite this observed hidden stratification, saliency methods continued to pass the weak localization test and fail the strong localization test at the subgroup level. (All localization scores are available in table form in Supplementary S.3.)
Fig. 2.
Localization scores reported for saliency methods at the a aggregate level and b abnormality subgroup level (color legend: blue, hardware/fractures; orange, arthritis; green, other). Saliency methods (including edge detector and inter-radiologist baselines) are plotted along the horizontal axis, and the vertical axis plots pixel-level area under the precision-recall curve (AUPRC). Box-plots display median, interquartile range, and outliers. (a) Superscript letters indicate statistical comparisons between (a) saliency methods (including edge-detector and inter-radiologist baseline) or (b) subgroups within each saliency method using pairwise Tukey tests at significance level
Heatmap consistency demonstrated a dichotomy between repeatability (Fig. 3a) and reproducibility (Fig. 3b). Four saliency methods passed the weak repeatability criterion: GRAD, IG, SIG, and XRAI. The other two methods failed only because their mean IoU scores fell below the 0.5 threshold (GCAM, IoU: mean = 0.459, SD = 0.228; SG, IoU: mean = 0.427, SD = 0.218). However, no saliency method passed the weak reproducibility criteria; each was only constrained by the IoU condition. For example, XRAI was the closest to passing (SSIM: mean = 0.784, SD = 0.176; IoU: mean = 0.502, SD = 0.284; pixel error: mean = 0.183, SD = 0.165) but the IoU score distribution was not significantly greater than 0.5 (p > 0.05). No saliency method passed either the strong repeatability or reproducibility criteria. Indeed, the strong thresholds ( = 0.951, = 0.613, = 0.034) together posed a stringent requirement on heatmap consistency. (All repeatability and reproducibility scores are available in table form in Supplementary S.3.)
Fig. 3.
Saliency method a repeatability and b reproducibility, as measured by SSIM (left panel), IoU (middle panel), and pixel error (right panel). Superscript letters indicate statistical comparisons between saliency methods using pairwise Tukey tests at significance level
Saliency methods generally demonstrated sensitivity to model randomization (Fig. 4). Five out of six saliency methods (all except SIG) had statistically lesser SSIM scores for the fully randomized model compared to the degradation threshold. This difference was largest for GRAD ( = 0.768, mean = 0.647, SD = 0.068) and GCAM ( = 0.579, mean = 0.469, SD = 0.097). Although SIG failed the test, SSIM scores did fall below the degradation threshold ( = 0.721, mean = 0.720, SD = 0.092), suggesting that SIG sensitivity could have benefited from more InceptionV3 iterations. The impact on cascading randomization on heatmap quality was observed alongside the change in weighted AUROC (Fig. 5). This provides a useful context for the cascading randomization plots: non-monotonicity and non-uniformity in layerwise changes in SSIM were coupled with similar patterns in classification performance. (All sensitivity scores are available in table form in Supplementary S.3.)
Fig. 4.
Saliency method sensitivity to model weights as measured by cascading randomization (color legend: red, GCAM; dark blue, GRAD; yellow, IG; green, SG; light blue/teal, SIG; purple, XRAI). Dotted horizontal lines indicate the associated saliency method’s degradation threshold. Sequential randomization layers are plotted along the horizontal axis (beginning with the non-randomized model), and the SSIM relative to non-randomized heatmaps is plotted along the vertical axis
Fig. 5.
GradCAM sensitivity (red, left vertical axis) and DCNN classifier wAUC (blue, right vertical axis) as a function of cascading randomization layer (horizontal axis)
Discussion
Although saliency maps have been used widely for explainability and disease localization in AI for radiology, their trustworthiness has been called into question [13, 15, 16]. In this study, we evaluated the trustworthiness of six state-of-the-art saliency methods for the diagnosis of abnormalities on upper extremity radiographs. Trustworthiness was defined as the combination of saliency map localization, repeatability, reproducibility, and sensitivity. For the first three criteria, we defined a weak version, which served as a sanity check, and a strong version, which enforced stricter conditions that correlate to human performance. Finally, we identified hidden stratification in localization accuracies by subgrouping the annotated abnormalities into arthritis, hardware/fracture, and extraneous cases. Our trustworthiness results are summarized in Table 3.
Table 3.
Overall trustworthiness results. Four criteria were evaluated (along columns: localization, repeatability, reproducibility, and sensitivity) for each of six saliency methods (along rows: gradient-weighted class activation mapping (GCAM), gradient explanation (GRAD), integrated gradients (IG), Smoothgrad (SG), smooth integrated gradients (SIG), and XRAI). Pass/fail indicates whether the saliency method satisfied the statistical requirements of each trustworthiness criteria
| Saliency method | Localization | Repeatability | Reproducibility | Sensitivity | |||
|---|---|---|---|---|---|---|---|
| Weak | Strong | Weak | Strong | Weak | Strong | ||
| GCAM | Pass | Fail | Fail | Fail | Fail | Fail | Pass |
| GRAD | Pass | Fail | Pass | Fail | Fail | Fail | Pass |
| IG | Pass | Fail | Pass | Fail | Fail | Fail | Pass |
| SG | Pass | Fail | Fail | Fail | Fail | Fail | Pass |
| SIG | Pass | Fail | Pass | Fail | Fail | Fail | Fail |
| XRAI | Pass | Fail | Pass | Fail | Fail | Fail | Pass |
All six saliency methods passed the weak localization criterion but failed the strong version, suggesting they possessed non-trivial yet sub-radiologist disease localization capacity. This finding is consistent with studies of saliency map localization on chest X-rays. Saporta et al. [16] found that localization of chest X-ray pathologies present in CheXpert [32] was significantly worse than human benchmarks, and Arun et al. [15] found that localization of both pneumothorax and pneumonia was inferior to specifically trained segmentation or object-detecting DCNNs, respectively. We also found that XRAI was the best localizing saliency method; this agrees with the results from Arun et al. [15] and is likely due to the recursive scheme of the XRAI method [26]. It is important to note that our analysis of localization excludes instances of inaccurate DCNN predictions—that is, only if the image was correctly classified do we evaluate localization accuracy. This mitigated the potential for DCNN errors to confound our analysis of saliency methods.
We also identified disparities in localization performance based on the kind of abnormality. These results echoed the take-home points from Oakden-Rayner et al. [33]: schema completion—i.e., assigning more complete subclass labels beneath the superclass—often reveals label imbalances that correlate to uneven AI performance across these labels. In our study, we extended these insights to saliency map efficacy. Hardware/fracture cases had higher saliency localization AUROCs than arthritis cases. We posit this was likely due to the sharper contours of fractures and clearer boundaries for metalwork relative to arthritis presentation. Similar findings of hidden stratification in saliency localization based on the disease pathology have been shown by Saporta et al. [16]; our study builds on this by clearly highlighting an instance of hidden stratification that be dangerous in AI visualization.
We showed that saliency methods differentiate between repeatability and reproducibility. While most saliency methods passed the weak repeatability test, all failed the weak reproducibility test. In other words, saliency maps for a single image notably differed based on the underlying DCNN architecture (i.e., InceptionV3 vs. DenseNet121), despite similar classification performance. GCAM and SG struggled on both kinds of heatmap consistency evaluations. This partially concords with the results presented by Arun et al. [15], where SG exhibited similar struggles on repeatability and reproducibility. Differences for other methods might be due to domain-specific phenomena, i.e., chest versus musculoskeletal X-rays. However, in concordance with the overall conclusion from Arun et al. [15], we found that no saliency method passed a strong repeatability or reproducibility test. The inability to pass strong repeatability and reproducibility criteria indicates that radiologist annotations were significantly more consistent than deep learning-derived saliency maps.
Every saliency method except for SIG demonstrated sensitivity to DCNN training. Arun et al. [15] similarly found that most saliency methods performed well on the weight randomization task. The only discrepancy between the results was with SIG, and this difference might be due to distribution-dependent effects on the randomization task, as critiqued by Yona and Greenfeld [34].
There are a few limitations in this study. First, the MURA dataset proved occasionally difficult to annotate for disease abnormality. Since the dataset is structured into studies, i.e., groups of images, one image alone may not have elucidated the abnormality apparent in the patient. Therefore, the annotated test set likely contained some views of a patient that did not highlight the abnormality that was known to exist. We accommodated this by removing those views from the saliency test set. Second, we only studied two DCNN architectures, InceptionV3 and DenseNet, in our reproducibility tests. The chosen architectures have been widely used in medical imaging studies with state-of-the-art performance; prior work by Arun et al. [15] similarly used these two base architectures for evaluating saliency maps on chest X-rays. Finally, we only studied the properties of gradient-based saliency methods, since these are the current state-of-art in visual explanation. Newer approaches such as SHAP [35] and h-SHAP [36] (which have more precise theoretical foundations [37]), as well as directly trained saliency maps [38], can be analyzed under this framework in future work.
Conclusions
Evaluating the trustworthiness of saliency methods used for diagnostic deep learning algorithms on upper extremity radiographs demonstrated that saliency methods critically suffer in localization capacity, repeatability, and reproducibility, as compared to radiologist annotations of disease. Saliency methods also demonstrate hidden stratification wherein abnormalities with clearer contours, such as hardware inserts and fractures, have better localization accuracy as opposed to less localized pathologies, such as arthritis. These results suggest that saliency methods may not be fully ready for clinical use, and more work must be done to make deep learning interpretations truly “trustworthy” before they are deployed in clinical diagnostic settings.
Supplementary Information
Below is the link to the electronic supplementary material.
Author Contribution
All authors contributed to the study conception and design. Material preparation, data collection, and analysis were performed by K.V. The first draft of the manuscript was written by K.V. and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.
Funding
This work was supported by the Johns Hopkins Biomedical Engineering Leong Undergraduate Research Fund. J.S. is supported by NSF CAREER Award CCF 2239787.
Code and Data Availability
We make our code repository available upon publication at https://github.com/kvenkatesh5/saliency-trustworthiness. All imaging data used is in the public domain and available in the references provided in the manuscript text.
Declarations
Ethics Approval
This study used publicly available data and was considered not human subjects research. Therefore, institutional review board review was waived.
Competing Interests
The authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.P. Rajpurkar and M. P. Lungren, “The Current and Future State of AI Interpretation of Medical Images,” N. Engl. J. Med., vol. 388, no. 21, pp. 1981–1990, May 2023. [DOI] [PubMed] [Google Scholar]
- 2.A. Khan, A. Sohail, U. Zahoora, and A. S. Qureshi, “A survey of the recent architectures of deep convolutional neural networks,” Artif. Intell. Rev., vol. 53, no. 8, pp. 5455–5516, Dec. 2020. [Google Scholar]
- 3.P. Rajpurkar et al., “CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning,” arXiv [cs.CV], 14-Nov-2017.
- 4.P. Rajpurkar et al., “Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists,” PLoS Med., vol. 15, no. 11, p. e1002686, Nov. 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.R. Ranjbarzadeh, A. Bagherian Kasgari, S. Jafarzadeh Ghoushchi, S. Anari, M. Naseri, and M. Bendechache, “Brain tumor segmentation based on deep learning and an attention mechanism using MRI multi-modalities brain images,” Sci. Rep., vol. 11, no. 1, p. 10930, May 2021. [DOI] [PMC free article] [PubMed]
- 6.L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. Specter, and L. Kagal, “Explaining Explanations: An Overview of Interpretability of Machine Learning,” arXiv [cs.AI], 31-May-2018.
- 7.J. R. Zech, M. A. Badgeley, M. Liu, A. B. Costa, J. J. Titano, and E. K. Oermann, “Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study,” PLoS Med., vol. 15, no. 11, p. e1002683, Nov. 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.J. Teneggi, P. H. Yi, and J. Sulam, “Examination-level Supervision for Deep Learning–based Intracranial Hemorrhage Detection at Head CT,” Radiology: Artificial Intelligence, p. e230159, Dec. 2023. [DOI] [PMC free article] [PubMed]
- 9.N. Bien et al., “Deep-learning-assisted diagnosis for knee magnetic resonance imaging: Development and retrospective validation of MRNet,” PLoS Med., vol. 15, no. 11, p. e1002699, Nov. 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.A. Mitani et al., “Detection of anaemia from retinal fundus images via deep learning,” Nat Biomed Eng, vol. 4, no. 1, pp. 18–27, Jan. 2020. [DOI] [PubMed] [Google Scholar]
- 11.Z. Kang, E. Xiao, Z. Li, and L. Wang, “Deep Learning Based on ResNet-18 for Classification of Prostate Imaging-Reporting and Data System Category 3 Lesions,” Acad. Radiol., Jan. 2024. [DOI] [PubMed]
- 12.L. Alzubaidi et al., “Trustworthy deep learning framework for the detection of abnormalities in X-ray shoulder images,” PLoS One, vol. 19, no. 3, p. e0299545, Mar. 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim, “Sanity Checks for Saliency Maps,” arXiv [cs.CV], 08-Oct-2018.
- 14.J. Zhang, H. Chao, G. Dasegowda, G. Wang, M. K. Kalra, and P. Yan, “Revisiting the Trustworthiness of Saliency Methods in Radiology AI,” Radiol Artif Intell, vol. 6, no. 1, p. e220221, Jan. 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.N. Arun et al., “Assessing the Trustworthiness of Saliency Maps for Localizing Abnormalities in Medical Imaging,” Radiol Artif Intell, vol. 3, no. 6, p. e200267, Nov. 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.A. Saporta et al., “Benchmarking saliency methods for chest X-ray interpretation,” Nature Machine Intelligence, vol. 4, no. 10, pp. 867–878, Oct. 2022. [Google Scholar]
- 17.W. Jin, X. Li, and G. Hamarneh, “One Map Does Not Fit All: Evaluating Saliency Map Explanation on Multi-Modal Medical Images,” arXiv [cs.CV], 11-Jul-2021.
- 18.P. Rajpurkar et al., “MURA: Large Dataset for Abnormality Detection in Musculoskeletal Radiographs,” arXiv [physics.med-ph], 11-Dec-2017.
- 19.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” arXiv [cs.CV], 02-Dec-2015.
- 20.G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely Connected Convolutional Networks,” arXiv [cs.CV], 25-Aug-2016.
- 21.S. S. Halabi et al., “The RSNA Pediatric Bone Age Machine Learning Challenge,” Radiology, vol. 290, no. 2, pp. 498–503, Feb. 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” arXiv [cs.CV], 07-Oct-2016.
- 23.K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps,” arXiv [cs.CV], 20-Dec-2013.
- 24.M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70, Sydney, NSW, Australia, 2017, pp. 3319–3328.
- 25.D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg, “SmoothGrad: removing noise by adding noise,” arXiv [cs.LG], 12-Jun-2017.
- 26.A. Kapishnikov, T. Bolukbasi, F. Viégas, and M. Terry, “XRAI: Better Attributions Through Regions,” arXiv [cs.CV], 06-Jun-2019.
- 27.J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller, “Striving for Simplicity: The All Convolutional Net,” arXiv [cs.LG], 21-Dec-2014.
- 28.R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra, “Grad-CAM: Why did you say that?,” arXiv [stat.ML], 22-Nov-2016.
- 29.S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, “A benchmark for interpretability methods in deep neural networks,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA: Curran Associates Inc., 2019, pp. 9737–9748.
- 30.J. Canny, “A computational approach to edge detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 8, no. 6, pp. 679–698, Jun. 1986. [PubMed] [Google Scholar]
- 31.M. He, X. Wang, and Y. Zhao, “A calibrated deep learning ensemble for abnormality detection in musculoskeletal radiographs,” Sci. Rep., vol. 11, no. 1, p. 9097, Apr. 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.J. Irvin et al., “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison,” arXiv [cs.CV], 21-Jan-2019.
- 33.L. Oakden-Rayner, J. Dunnmon, G. Carneiro, and C. Re, “Hidden stratification causes clinically meaningful failures in machine learning for medical imaging,” in Proceedings of the ACM Conference on Health, Inference, and Learning, Toronto, Ontario, Canada, 2020, pp. 151–159. [DOI] [PMC free article] [PubMed]
- 34.G. Yona and D. Greenfeld, “Revisiting Sanity Checks for Saliency Maps,” arXiv [cs.LG], 27-Oct-2021.
- 35.S. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” arXiv [cs.AI], 22-May-2017.
- 36.J. Teneggi, A. Luster, and J. Sulam, “Fast Hierarchical Games for Image Explanations,” arXiv [cs.CV], 13-Apr-2021. [DOI] [PubMed]
- 37.J. Teneggi, B. Bharti, Y. Romano, and J. Sulam, “SHAP-XRT: The Shapley Value Meets Conditional Independence Testing,” Transactions on Machine Learning Research, 11-Jul-2023.
- 38.Z. Liu, E. Adeli, K. M. Pohl, and Q. Zhao, “Going Beyond Saliency Maps: Training Deep Models to Interpret Deep Models,” Inf. Process. Med. Imaging, vol. 12729, pp. 71–82, Jun. 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
We make our code repository available upon publication at https://github.com/kvenkatesh5/saliency-trustworthiness. All imaging data used is in the public domain and available in the references provided in the manuscript text.





