Skip to main content
iScience logoLink to iScience
. 2026 Feb 16;29(3):115026. doi: 10.1016/j.isci.2026.115026

A comprehensive review of explainable artificial intelligence in healthcare methods, evaluation, and clinical integration

Kai Zhang 1,5, Dongqi Wang 2,4,5, Fuxin Lin 2, Jue Xie 3, Weihua Zhou 2,4,
PMCID: PMC12996819  PMID: 41858885

Abstract

Explainable artificial intelligence (XAI) is essential for healthcare trust, yet a substantial gap persists between XAI techniques and actual clinical adoption. This review addresses this gap by framing clinical integration through three complementary lenses. First, we introduce a three-dimensional XAI classification framework—property, dependency, and scope—that moves beyond descriptive cataloging and serves as a practical guide for matching XAI approaches to specific clinical tasks. Second, we propose an integrated evaluation system that balances technical robustness, including fidelity, with measures of clinical utility such as workflow alignment and clinician confidence. Third, we analyze the divergent and often competing needs of key stakeholder groups to produce a role-characteristic mapping that clarifies what constitutes meaningful explainability in different clinical contexts. By positioning clinical integration as the center, this review outlines a pathway for translating XAI from methodological innovation to a dependable component of clinical decision support.

Subject areas: health informatics, health sciences, medical specialty, medicine


Health informatics; Health sciences; Medical specialty; Medicine

Introduction

The rapid uptake of artificial intelligence (AI) across medicine has enabled substantial advances in data interpretation, diagnostic support, treatment planning, and population-level health management. Yet many AI models, particularly deep learning systems, remain opaque because of their complex architectures and limited interpretability. This lack of transparency is especially problematic in clinical settings, where decision-making requires not only high accuracy but also clear justification and traceability to meet the expectations of clinicians, patients, and regulators alike.1,2,3,4,5,6,7,8,9

In response to these challenges, explainable artificial intelligence (XAI) has emerged as a central avenue for addressing the opacity problem.10,11,12 At its core, XAI seeks to render model behavior more transparent by articulating the reasoning processes that connect inputs to outputs, thereby enhancing accountability and supporting informed clinical judgment.13,14,15,16 Beyond improving interpretability, XAI is widely viewed as essential for aligning AI systems with existing clinical norms, documentation practices, and medico-legal expectations.1,10 As such, it is increasingly regarded as a prerequisite for the safe and trustworthy deployment of AI in healthcare.17

Beyond its role in ensuring safety and accountability, XAI offers broad utility across clinical applications, including diagnosis, rehabilitation management, and surgical planning.18 Its value is reflected in two main domains. First, XAI can help clinicians understand how models generate their outputs, for example, through visualization techniques, thereby strengthening confidence in AI-assisted decisions.19,20 Second, XAI can reveal latent structure in high-dimensional medical data, offering new data-driven insights that advance biomedical research.21,22 Finally, explainability supports system-level benefits, such as enhancing patient communication, facilitating cross-institutional quality assurance, and enabling more transparent model governance as healthcare AI adoption expands.23,24,25,26

Despite this progress, a substantial gap persists: the limited and frequently unsuccessful integration of XAI into large-scale clinical practice. Translating XAI methods from controlled research settings to the high-stakes, time-constrained, and heterogeneous workflows of routine care remains profoundly challenging. Existing reviews have largely emphasized methodological taxonomies or domain-specific applications without confronting the core issue of clinical integration. They offer neither a prescriptive framework for aligning XAI methods with real clinical tasks nor a coherent standard for evaluating practical utility, offering only limited guidance for managing the competing priorities of diverse stakeholder groups. Thus, we synthesize the landscape of XAI in healthcare through a systematic review of 170 key publications. We shift the emphasis from the descriptive question of what XAI is to the prescriptive question of how XAI can be integrated by advancing three contributions.

  • 1.

    Clinically oriented classification framework. We introduce a three-dimensional taxonomy—model property, dependency, and scope—that functions as a decision-support tool rather than a descriptive catalog. This framework enables precise alignment of XAI method characteristics with the requirements of specific clinical workflows.

  • 2.

    Practice-focused evaluation system. We propose an evaluation schema that extends beyond technical metrics alone. By integrating measures of technical reliability with indicators of clinical utility, including clinician trust and workflow compatibility, this dual-dimensional system offers a pragmatic standard for assessing whether an XAI tool is suitable for routine clinical use.

  • 3.

    Stakeholder-centered integration model. We analyze the needs and expectations of clinicians, patients, developers, and regulators to construct a role-XAI characteristic mapping. This model supports human-centered design and offers a structured approach for navigating competing priorities during implementation.

Given the complexity and high-stakes nature of healthcare, clear terminology around model comprehension is essential. In this review, we define transparency as system-level disclosure of data sources, model structure, and design principles; interpretability as the inherent clarity of models whose internal logic can be directly understood; and explainability as the generation of post hoc rationales for complex models whose internal mechanisms are not readily accessible. We assess the quality of explanations through fidelity, referring to their technical accuracy with respect to the underlying model, and credibility, referring to their perceived trustworthiness to human users. Traceability ensures that decisions can be audited across the AI life cycle. Together, these components form the foundation for the rigorous evaluation of clinical applicability and organizational accountability, and they constitute the core requirements for a trustworthy medical AI system.

The paper is organized as follows. We first review mainstream XAI methods and existing evaluation frameworks. We then survey current XAI tools, followed by an examination of implementation requirements from the perspectives of multiple stakeholder groups. Next, we summarize representative clinical application cases. We conclude by outlining key challenges and future directions and providing closing remarks.

Methodology

We conducted this review in accordance with PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines to ensure methodological rigor and reproducibility (Figure 1).27 Core references were identified through systematic searches of ScienceDirect, PubMed, and Google Scholar. In February 2025, we performed initial searches using the terms “explainable AI and healthcare,” “explainable artificial intelligence and healthcare,” and “explainable AI.” The search was restricted to studies published between 2019 and 2025 to capture recent developments. For each database, we screened the top 200 returned articles. We then applied a snowballing strategy by reviewing the reference lists of these articles to identify additional relevant and foundational studies.

Figure 1.

Figure 1

PRISMA diagram

The diagram summarizes the identification, screening, eligibility assessment, and final inclusion of studies reviewed in this work. Records were retrieved from multiple databases, duplicates were removed, and articles were excluded based on predefined inclusion and exclusion criteria, resulting in 170 core publications included for analysis.

The final selection process involved a two-level screening procedure to ensure that the included studies aligned closely with the scope of this review.

  • 1.
    Level 1 Screening (Title and Abstract): Duplicate records were removed, and non-English articles were excluded. Articles were further excluded if the title or abstract indicated:
    • Domain Irrelevance: The study focused outside healthcare, medicine, or biomedical engineering (e.g., finance, manufacturing).
    • Methodological Irrelevance: The work primarily addressed traditional statistical approaches or non-AI computational models, without involving machine learning or a specific XAI technique.
  • 2.

    Level 2 Screening (Full-Text Review-Core Inclusion/Exclusion Criteria): Remaining articles underwent a rigorous full-text review guided by the central themes of this systematic analysis. Studies were assessed according to the following criteria:

Inclusion Criteria (I)

The study proposes, validates, or systematically compares at least one contemporary XAI method (e.g., LIME, SHAP, attention mechanisms, or intrinsically interpretable models) within a healthcare context.

The study introduces or applies a framework or metric to evaluate XAI outputs (e.g., fidelity, human-subject trust, clinical utility, or fairness).

The study demonstrates a concrete application of XAI in a clinical or medical domain and addresses implementation challenges or stakeholder perspectives.

Exclusion Criteria (E)

General perspectives, news items, or commentaries lacking new empirical data or structured review insights.

Studies with insufficient methodological detail regarding the XAI approach or evaluation, rendering findings non-reproducible.

Articles focused solely on general machine learning performance without the discussion of interpretability, transparency, or explainability.

Following this process, a total of 170 core references were included. The yearly distribution of publications from 2016 to 2025 was: 2, 3, 9, 12, 17, 16, 33, 35, 25, and 8 articles, respectively (Figure 2). Figure 3 illustrates the focus points of existing medical XAI reviews, showing that the application of XAI techniques in healthcare receives the most attention. However, there is a lack of systematic discussion on how XAI can be better integrated into clinical practice/applications.

Figure 2.

Figure 2

Distribution of references by publication year

The figure shows the annual number of studies on XAI in healthcare included in this review, highlighting the rapid growth of the field in recent years and the increasing research focus on XAI methodologies and clinical applications.

Figure 3.

Figure 3

Main focuses of existing reviews

The figure categorizes prior reviews according to their primary emphasis, including XAI methods, evaluation metrics, application domains, and clinical integration. It illustrates that while methodological and application-oriented discussions dominate the literature, systematic analysis of real-world clinical integration remains comparatively limited.

Explainable artificial intelligence methods and their evaluation approaches

Explainable artificial intelligence methods

The field of XAI has seen rapid growth, giving rise to a variety of classification systems. Most studies categorize methods along three principal dimensions: (1) functional characteristics, including feature-importance analysis, decision-path visualization, and concept-based explanation methods28,29,30,31,32; (2) model dependency, distinguishing intrinsically interpretable models29 from post-hoc explanation approaches that rely on external tools1,20,33,34; and (3) explanation scope (i.e., granularity), encompassing global explanations35 and local explanations.36,37 Additionally, the distinction between perceptual and structural interpretability19 provides a cognitive science perspective that enriches the theoretical foundation for method classification.

To better accommodate the diverse task requirements encountered in clinical settings, we propose a three-dimensional XAI classification framework tailored for medical applications (Figure 4). This framework systematically integrates three key dimensions.

  • 1.

    Model properties: distinguishing self-explanatory models from post-hoc explanation methods.

  • 2.

    Model dependency: differentiating model-specific approaches from model-agnostic methods.

  • 3.

    Explanation granularity: distinguishing global from local explanations.

Figure 4.

Figure 4

Classification of XAI methods

The framework organizes XAI approaches along three orthogonal dimensions: model properties (self-explanatory versus post-hoc), model dependency (model-specific versus model-agnostic), and explanation granularity (global versus local). This structure is designed to support prescriptive matching between XAI methods and specific clinical tasks, workflows, and ethical constraints.

Our framework deliberately combines the widely adopted dimensions of model properties and explanation granularity with model dependency to establish a medical-specific taxonomy. Notably, we exclude functional characteristics of explanations as a primary axis. This distinction marks a departure from prior descriptive reviews: rather than merely cataloging explanation types, our framework emphasizes the practical constraints of clinical integration, including model compatibility, explanation granularity, and implementation pathways.

This approach transforms the taxonomy from a descriptive catalog into a prescriptive decision-making tool, creating a direct and actionable link between an XAI method’s technical features and the specific demands of a clinical task. Beyond its practical utility, the framework functions as a first-class ethical filter: each choice along the three dimensions carries direct ethical implications that must be explicitly justified. For example, explanation granularity: local explanations, targeting individual patients, may pose privacy risks if reverse-engineered, whereas global explanations, summarizing population-level trends, risk reinforcing stereotypes or exposing systemic biases. Model properties: Self-explanatory (interpretable) models are more amenable to fairness audits, whereas post-hoc explanations for black-box models may serve as manipulable “justifications” that conceal potentially biased internal logic.31

By requiring development teams to explicitly debate and document their choices—for instance, “we selected a local, self-explanatory approach to prioritize privacy and auditability”—the framework embeds ethical-by-design principles at the outset of the development life cycle, rather than relegating them to a post-deployment checklist.

The innovative value of this framework is reflected in four key aspects, systematically distinguishing it from prior descriptive reviews: 1. Clinical Problem Adaptation (Prescriptive Matching): The framework moves beyond generalized categorization to align technical features precisely with clinical requirements. For example, it directly links the high-fidelity lesion localization needed in medical imaging with visual, model-specific explanation methods such as Grad-CAM, transforming the framework into a prescriptive matching tool. 2. Workflow Integration and Timeliness (Implementation Pathway): Suitable explanation methods are recommended according to the temporal demands of clinical scenarios. Local, computationally efficient methods are favored for low-latency settings, such as real-time decision-making in emergency departments, whereas global, higher-cost methods are more appropriate for post-hoc analyses in research or chronic care planning. 3. Full-Cycle Governance and Coverage: The framework spans the entire patient care trajectory, from initial disease screening to continuous prognosis prediction and risk adjustment. This systematic coverage provides a holistic methodological perspective on XAI deployment. 4. Embedded Ethical and Auditability Filter: By requiring development teams to proactively select and justify paths based on model properties and explanation granularity, the framework embeds ethical-by-design principles. For instance, choosing an intrinsically interpretable model facilitates fairness audits, whereas selecting a local explanation mitigates privacy risks by avoiding disclosure of population-level biases. This ensures responsible AI governance from the outset.

From a technical perspective, self-explanatory models leverage the inherent interpretability of their structure. For example, decision trees38 provide explicit rule-based explanations via branching conditions, while linear regression39 assigns transparent causal weights to features through regression coefficients. These approaches excel in structured data contexts, such as EHR analysis. However, for high-dimensional and complex data—such as medical imaging or genomics—their performance often falls short of deep learning models,40 underscoring the need for robust post-hoc explanation methods.

Post-hoc explanations can be further divided into model-specific and model-agnostic approaches based on their dependence on model architecture. Model-specific methods, such as Grad-CAM in CNNs, generate heatmaps that enable precise lesion localization in conditions such as pneumonia and brain tumors.29,41,42 LRP decomposes the contribution of each layer to the final prediction, demonstrating unique value in precision oncology.21 In contrast, model-agnostic methods are broadly applicable across model types. LIME approximates black-box decision logic by constructing local linear surrogate models around input samples, a strategy widely employed in emergency triage.34,43 SHAP assigns fair, quantitative contributions to individual features, supporting both global feature rankings and case-specific explanations—making it particularly popular in EHR-based predictive modeling.35,44 Additional approaches include attention mechanisms, which visualize model focus through attention weights,45,46 and adversarial example analysis, which probes model robustness and sensitivity by introducing minimal perturbations.47

The choice of explanation granularity is critical in clinical applications. Global explanation methods aim to reveal overarching model behavior. For instance, SHAP value analysis has been used to identify key clinical indicators influencing breast cancer prediction48 and to discover potential imaging biomarkers for Alzheimer’s disease.49 While these global approaches provide important theoretical foundations for model optimization and feature engineering, their direct utility in patient-level clinical decision-making is limited. Local explanation methods, by contrast, focus on the decision logic underlying individual predictions. Case-specific rules generated by LIME, for example, can help physicians understand an AI system’s assessment of complication risks for individual patients with diabetes.50 In practice, global and local methods are often employed complementarily. In Alzheimer’s disease diagnosis, researchers first use global SHAP analysis to identify hippocampal atrophy as the most predictive feature, then apply local LIME explanations to illustrate its manifestation and diagnostic weight in specific patients.49,51 This multi-level explanation strategy significantly enhances clinician trust and adoption, thereby strengthening support for clinical decision-making.

To clarify the applicability of XAI methods across medical tasks, we establish a method-task-scenario mapping (Table 1), linking mainstream methods with representative diseases, clinical task types, and technical characteristics. In medical image analysis, model-specific local visualization approaches are widely used for lesion localization due to their intuitive interpretability.29,41 For EHR-based risk prediction, methods offering both global and local explanations demonstrate distinct advantages.52,53 In time-sensitive emergency scenarios, lightweight local explanation methods have been successfully applied to streptococcal pharyngitis triage54 and methanol poisoning response.55 This refined mapping mechanism provides both theoretical guidance and practical engineering support, promoting the generalizability and clinical utility of medical XAI systems.

Table 1.

Example mapping between XAI Methods and application diseases

Disease/application area XAI methods Reference
Pneumonia diagnosis Grad-CAM, LRP 29,56
ECG signal detection of myocardial infarction Grad-CAM 57
Cataract Grad-CAM 58
Thyroid cancer diagnosis Grad-CAM 59
Alzheimer’s disease Grad-CAM, SHAP, SVM, Random Forest, Saliency Maps, Naive Bayes, KNN, LRP, CAM, Linear Regression, GBP, LIME 49,51,60,61,62,63,64
Glioma Grad-CAM 65
COVID-19 Grad-CAM, Grad-CAM++, LIME, SHAP, XGBoost, SVM, LRP, Logistic Regression, CAM, BP, GBP, Integrated gradients 56,66,67,68,69,70,71,72,73
Brain tumor Grad-CAM, GBP, NeuroXAI framework 41,74,75
MRI Grad-CAM 76
Chest diseases Grad-CAM, Grad-CAM++, DeepLIFT, Occlusion, Integrated gradients 14,60
Skin cancer classification Grad-CAM, Kernel SHAP 63,77
Breast cancer risk prediction Knowledge-based, SHAP, Decision Tree, Explanation ensembles 48,78,79,80
Chronic kidney disease prediction Example-based 50
Neurological decision support system Example-based, Counterfactual Reasoning, Template Language, Probability Score, Crowdsourced Score, Feature Importance, Decision Tree 81
Streptococcal pharyngitis Example-based, Feature Importance 54
Early prediction of heart disease XGBoost, SVM, Random Forest 55,82,83
Heart failure detection XGBoost 53
Stroke prediction XGBoost, SVM, LIME, KNN 83,84
Prediction of intubation needs in methanol poisoning patients XGBoost, SVM, Random Forest, Decision Tree 55
Real-time prediction of disease deterioration XGBoost, Random Forest, Decision Tree 85
Lung cancer diagnosis XGBoost, SVM, Linear Logistic, LGBM 52,79,86
Tumor treatment SVM, Random Forest, GBM 87
Computer-aided diagnosis AdaBoost 88
Diagnosis of iron deficiency anemia and aplastic anemia LIME, SHAP 89
Dopamine imaging techniques for Parkinson’s disease LIME 28
Cancer diagnosis LIME 90
X-ray image classification LIME 91
Detection of acoustic biomarkers and pulmonary diseases LIME, Feature Importance 92
Nasopharyngeal carcinoma LIME, SHAP 93
Readmission of sepsis patients in the intensive care unit LIME, SHAP 94
Precision oncology LRP 21
Time series analysis of clinical gait LRP 95
Prediction of all-cause mortality SHAP, Decision Tree 96
Identification of gait biomechanical parameters associated with anterior cruciate ligament injury SHAP 97
Diagnosis of ophthalmic diseases SHAP, GBP 98,99,99
Non-melanoma skin cancer SHAP 63
Detection and diagnosis of heart diseases SHAP 100
Glioma grading SmoothGrad 101
Fetal growth assessment PCBM 102
Diabetic retinopathy grading Integrated gradients 103
Age-related macular degeneration Expressive gradients 104
Classification of estrogen receptor status by breast MRI Integrated gradients attribution method, smooth-grad noise reduction algorithm 105
Estimation of biventricular volumes of the heart Shape attentive U-Net 106
Diagnosis of autism spectrum disorder Auto-ASD-Network 107
Salivary gland cancer densMAP 108
Diagnosis of colorectal cancer and heart disease Fuzzy Classifier 109
Segmentation of colorectal polyps GBP 110

A critical analysis of the methods summarized in Table 2 highlights substantial and persistent obstacles to clinical integration. Perturbation-based, model-agnostic approaches remain prominent in the literature, yet their known instability,30,111 substantial computational demands,30,112 and susceptibility to generating misleading or non-faithful explanations31 limit their suitability for time-sensitive clinical settings. Intrinsically interpretable models, although transparent by design, frequently underperform when applied to high-dimensional modalities such as imaging or genomics.40 Popular visualization techniques offer intuitive appeal, but they are post hoc and architecture-specific, largely confined to CNNs, with no assurance that the displayed heatmaps faithfully represent the model’s underlying decision pathway.113,114 Rather than a set of mature, deployment-ready tools, the current landscape often appears a collection of specialized techniques whose fragility becomes evident once evaluated against clinical constraints. A frequently cited barrier to real-world adoption is the mismatch between the substantial computational and financial costs of these systems and the still-limited evidence of clinical benefit.

Table 2.

Summary of main XAI methods

Types XAI methods Strength Weakness
LRP LRP28,29,30,50,113 Provides high-resolution, detailed relevance heatmaps that are more stable and suitable for deep layers than many alternatives. The choice of relevance propagation rule can significantly impact the result, potentially requiring specific tuning for different models.
LRP-CNN29
LRP-DNN29
LRP-BiLRP29
LRP-Deeplight LRP19,29
CAM-based CAM10,19,28,37,113 Widely applicable to various CNN models without retraining, providing visually intuitive, coarse localization maps for the predicted class. The relevance maps are low-resolution and typically only highlight the most discriminative parts, often missing fine-grained details or broader context.
Grad-CAM10,19,28,29,30,37,109,113,114,115
Grad-CAM++10,28,33
Guided Grad-CAM29
Respond-CAM29
Multi-layer CAM29
LIME LIME4,28,29,30,37,50,111,113,116 Highly model-agnostic and locally faithful, it can explain any classifier’s individual predictions by building a simple, interpretable local surrogate model. The explanation’s stability and robustness can be low, as the local sampling and perturbation process relies heavily on hyperparameters and dataset specifics.
SHAP SHAP4,20,28,39,50,112,117 Provides a unified, theoretically sound measure of feature importance based on Shapley values, ensuring consistency and local accuracy. The exact calculation of Shapley values is computationally intensive and slow, requiring approximate sampling for most real-world, high-dimensional problems.
Tree SHAP37
Deep SHAP20,30
Kernel SHAP30,37
Gradient-based GBP10,113 Simple, computationally efficient, and widely applicable across different neural network architectures since they only require one backward pass. Standard gradient methods often suffer from high noise or saturation, leading to unintuitive or sparse explanations for confident predictions.
Integrated Gradient10
BP33
Full-Gradient118
Expressive Gradient113
Smooth Gradient10,37,119
Vanilla Gradient113
DeepLIFT10,33,111,113,115
KNN KNN1 Provides a direct, instance-based explanation by identifying the most similar training data points responsible for a given prediction. The explanation’s quality is highly sensitive to the curse of dimensionality and the chosen distance metric, making it impractical for high-feature spaces.
Rules sets Knowledge-based28,29,78 Generates highly transparent and human-readable explanations in the form of logical IF-THEN conditions that directly lead to a prediction. The complexity and interpretability of the model significantly degrade as the number of rules or conditions within the rules increases.
Template Language28,29
Feature Importance28,29
Explanation Ensembles28,29
Fuzzy Rules29,50
Example-based Example-based10 Provides intuitive, concrete explanations by showing the user the most influential or representative training data points for a specific outcome. These methods can be computationally demanding, especially for large datasets, and their effectiveness relies on the quality and representativeness of the training examples.
Linear regression Linear Regression29 Coefficients provide a simple and direct measure of each feature’s magnitude and direction of impact on the predicted output. The model is highly limited, as it can only accurately explain relationships that are strictly linear, failing on complex or non-linear data.
Linear Logistic Regression1,29
Naive Bayes1,29
GLM1,29
Decision tree Decision Tree1,10 The explanation is highly transparent, following a clear, easy-to-read sequence of nested feature conditions (rules) leading to the final prediction. The interpretability is rapidly lost as the tree deepens, making large or complex trees functionally equivalent to a black box.
Random Forest10
GBM10
LGBM10
XGBoost10
AdaBoost10
EBM50
GAM10
Saliency Maps Saliency Maps28,31,115 They are simple and fast to compute, only requiring a single backward pass to visually highlight important input pixels for the prediction. The resulting maps are often noisy, unstable, and suffer from the gradient saturation problem, leading to poor visual quality and reliability.
GAM1,19,50
SVM SVM1 In the linear case, the separating hyperplane and margin provide an intuitive global explanation of class separation, with support vectors acting as key boundary examples. The use of non-linear kernels transforms the data into high-dimensional space, making the resulting decision boundary and feature relationships opaque and difficult to interpret.
LinearSVM29
PolynomialSVM29
Occlusion Occlusion111,113,120,121 High model-agnostic and intuitive, as it directly shows the causal effect of removing input regions on the final prediction score. The method is computationally very expensive because it requires multiple forward passes of the model for every single input region being tested.
Deep Taylor Decomposition Deep Taylor Decomposition37,115 Provides a theoretically grounded, high-resolution relevance decomposition that effectively handles complex, non-linear neural network architectures. The method can be sensitive to the choice of the reference point (root point) around which the Taylor expansion is performed, potentially affecting the final map.
PDP PDP29,37 Provides an intuitive, global view of how a feature, on average, influences the model’s prediction across the entire dataset. The plot can be misleading if the feature being analyzed strongly interacts with other features, as it assumes feature independence.
Attention maps Attention Maps29,113 Provides a native, structurally meaningful explanation by showing how the model’s internal attention mechanism weights different input parts. Visualized attention weights do not always correlate perfectly with the true causal importance of a feature for the final prediction.
Feature dimensional reduction PCA116 Simplifies high-dimensional data, enabling visual inspection and cluster identification of complex feature relationships or internal model representations. The transformed low-dimensional representation can distort the original local and global feature distances and relationships.
t-SNE116
UMAP116
CAV CAV28 Quantifies a model’s reliance on high-level, human-friendly concepts rather than just individual features or pixels, providing concept-based explanations. The method requires the user to manually define and provides sufficient example sets for each concept they wish to test, adding significant overhead.

Within this context, this review focuses on four widely used XAI approaches in medical AI: SHAP, LIME, CAM, and LRP. SHAP provides predominantly model-specific, local feature attributions; LIME enables model-agnostic local explanations through surrogate modeling; and CAM and LRP serve as deep-learning-specific techniques that generate local visual explanations by interrogating internal activation patterns within neural networks.

SHAP

SHAP is a game theory-based explainability method for machine learning models, designed to explain model predictions. SHAP is grounded in the theory of Shapley values, treating model prediction as a cooperative game in which each feature is an important player. By calculating the marginal contribution of each feature to the final prediction, SHAP quantifies the importance of individual features and provides consistent and fair explanations. The calculation formula for the Shapley value is as follows:

i(v)=SN{i}|S|!(|N||S|1)!|N|!(v(S{i})v(S)) (1)

Among them, N denotes the set of all features; N∖{i}represents the subset of N excluding feature i; v(S) indicates the contribution value of feature subset S; and ∅i(v) denotes the Shapley value.

Due to the high computational complexity of enumerating all possible subsets z', the SHAP method typically employs approximation techniques to calculate SHAP values. Common approximations and algorithms include Kernel SHAP, a kernel-based approximation algorithm that estimates SHAP values through weighted linear regression, and Tree SHAP, a fast computation algorithm specifically designed for tree-based models that efficiently calculates SHAP values by leveraging the structural characteristics of decision trees.

The SHAP method is widely used for interpretation tasks across various machine learning models. Its main applications include: (1) the interpretation of single-sample predictions, which explains the model’s prediction for an individual instance and helps users understand the contribution of each feature to the prediction; (2) global pattern interpretation, which analyzes the overall predictive behavior of the model and reveals the model’s dependence on different features; and (3) feature importance evaluation, which quantifies the importance of features and helps identify those that have the greatest impact on model predictions.

LIME

LIME is a model-agnostic local explanation method designed to interpret the predictions of complex models by constructing locally interpretable models. The core idea of LIME is to generate perturbed samples in the vicinity of the target instance and use a simple model to approximate the behavior of the complex model in this local region. The specific implementation steps of LIME are as follows.

  • 1.

    Generating perturbed samples: A set of perturbed samples is generated in the vicinity of the target sample x {x1,x2,,xn}

  • 2.

    Obtaining predictions from the complex model: The complex model f is used to make predictions on the perturbed samples, resulting in the corresponding prediction outcomes {f(x1),f(x2),,f(xn)}

  • 3.

    Training a locally interpretable model: Using the perturbed samples and their corresponding prediction results, a simple and interpretable model g is trained.

  • 4.

    Explaining the target sample: The interpretable model g is used to explain the prediction result of the target sample x.

Given a complex model f and a target sample x, the goal of LIME is to find an interpretable model g such that g can approximate the predictive behavior of f in the vicinity of x. The optimization problem of LIME can be formalized as follows:

minL(f,g,πx)+Ω(g),gG (2)

Here, G denotes the set of interpretable models; L(f,g,πx) represents the prediction difference between the complex model f and the interpretable model g in the vicinity of the target instance x; πx denotes the weighting function around the target instance x, usually an exponential kernel function πx(z) = exp⁡(-D(x,z)2/σ2), where D(x,z) is the distance between instances x and z, and σ is the bandwidth parameter; Ω(g) represents the complexity penalty term for the interpretable model g, used to control the model’s complexity.

LIME’s primary application scenarios include single-instance prediction explanation, model debugging and validation, feature importance analysis, and model comparison and selection.

CAM

CAM is a technique used to visualize the decision-making process of convolutional neural networks, aiming to highlight image regions that contribute most significantly to the prediction of a specific category by generating heatmaps. The core idea of CAM is to utilize the feature maps of the last convolutional layer in a CNN and their corresponding weights to generate a heatmap of the same size as the input image, where brighter regions indicate areas with the greatest contribution to the classification result.

The mathematical formulation of CAM is as follows: Given a CNN model, assume its last convolutional layer outputs are fk(x,y), where k denotes the feature map index and (x,y) represents the spatial position. Assume the weights of the fully connected layer are wkc, where c denotes the class index, and k denotes the feature channel index. The CAM heatmap Mc(x,y) can be calculated using the following formula:

Mc(x,y)=kwkcfk(x,y) (3)

Here, Mc(x,y) represents the heatmap for class c; wkc represents the weight of feature map channel k for class c in the fully connected layer, and fk(x,y) represents the feature map from the last convolutional layer.

The specific steps are as follows.

  • 1.

    Train CNN Model: First, train a CNN model, ensuring that its last convolutional layer is connected to a global average pooling layer and a fully connected layer.

  • 2.

    Extract Feature Maps: For the input image, extract the feature maps fk(x,y) from the last convolutional layer.

  • 3.

    Compute Heatmap: Based on the weights wkc of the fully connected layer, compute the heatmap Mc(x,y) for class c.

  • 4.

    Upsample Heatmap: Upsample the heatmap Mc(x,y) to the size of the input image, generating the final heatmap.

CAM visually highlights the image regions of interest to the model through heatmaps, without requiring retraining or modification of the original CNN model. It has demonstrated wide application value in visualizing model attention regions, model debugging and validation, medical image analysis, and autonomous driving, providing an important tool for research into the explainability of deep learning models.

LRP

In the explainability research of deep learning models, LRP is a backpropagation-based explanation method that aims to explain model prediction results by propagating relevance scores layer by layer. The core idea of LRP is to decompose the model’s output prediction value into contributions from the input features. Specifically, LRP starts from the model’s output layer and propagates relevance scores layer by layer toward the input layer, thereby generating relevance scores for the input features. The key to LRP lies in defining the relevance propagation rules for each layer, ensuring that the relevance scores are conserved during propagation.

The mathematical formulation of LRP is as follows: Given a deep learning model, assume its output is f(x), where x denotes the input features. The goal of LRP is to decompose the output f(x)into relevance scores for the input features, such that:

f(x)=iRi (4)

where Ri represents the relevance score of input feature xi.

LRP achieves this goal by propagating relevance scores layer by layer. Assuming the neuron activation value at layer l is ajl, and the neuron activation value at layer l+1 is akl+1, then the relevance score Rjl at layer l can be calculated using the following formula:

Rjl=kajlwjkjajlwjkRkl+1 (5)

Where wjk denotes the relevance score transferred from neuron j in layer l to neuron k in layer l+1.

The algorithmic steps for LRP are as follows.

  • 1.

    Initialization: Start from the model’s output layer. Set the output prediction value as the initial relevance score, where L denotes the output layer.

  • 2.

    Layer-wise Relevance Propagation: Propagate relevance scores layer by layer from the output layer toward the input layer, calculating relevance scores according to the propagation rules for each layer.

  • 3.

    Generate Input Feature Relevance Scores: This step yields the relevance score for each input feature, representing its contribution to the output prediction value.

Evaluation methods

In the evaluation of explainability within medical AI systems, a rigorous, clinically grounded assessment framework is essential for closing the gap between algorithmic performance and real-world clinical utility. A synthesis of the existing XAI evaluation literature, as organized in Table 3, reveals a highly fragmented landscape defined by siloed metrics and minimal linkage to clinical outcomes. Current practice relies on three largely independent categories: (1) computer-centered metrics that quantify the technical fidelity of an explanation to a model’s internal reasoning; (2) human-centered metrics that capture users’ subjective perceptions, trust, or usability; and (3) task-specific metrics tailored to particular clinical contexts. As reflected in Table 3, most studies privilege a single category—typically computer-centered measures—because they are convenient, quantifiable, and inexpensive to compute. This pattern exposes a central shortcoming in contemporary evaluation methodology: technical performance is frequently assessed in isolation from clinical relevance.

Table 3.

Summary of XAI evaluation methods

Reference Evaluation methods Details Basis for assessment
Chaddad et al.20 Human-centered, Computer-centered Human-centered XAI is costly, requiring domain expert (e.g., clinician) evaluation and feedback analysis. Computer-centered XAI is cheaper, using algorithms to assess explanation quality. Ability to provide accurate and understandable explanations for its decisions.
Kong et al.10 Human-centered XAI evaluation Match user demands and goals with perception-based evaluation measures. Mental model, Trust, Usefulness, Improved performance, Ethics, Satisfaction
Rong et al.17 Human-centered Understanding: A user’s internal grasp of how the ML model works; Usability: How successfully, efficiently, and satisfactorily a user can achieve goals with the product; Trust: A combination of confidence in model accuracy, comfort with its use, and willingness to delegate decisions; Huma-AI Collaboration: Performance in scenarios where AI predicts, but humans decide. Understanding, Usability, Trust, Human-AI Collaboration
Kerz et al.112 Clinical and technical perspectives Truthfulness: Explanations should truthfully reflect the AI model's decision process. Informative plausibility: User’s judgment on the explanation plausibility may inform users about AI decision quality, including potential flaws or biases. Truthfulness, Informative plausibility
Yeh et al.122 To saliency explanations Many explanations are unified by optimizing infidelity via perturbation. Theoretically and empirically, there is no necessary trade-off between sensitivity and infidelity; the right amount of smoothing can improve both metrics. (In)fidelity, Sensitivity
Ghnemat et al.123 To causal interpretability Human subject-based evaluation metrics; non-human based evaluation metrics; counterfactual explanations evaluation metrics; model-based evaluation metrics; causal fairness evaluation. Goodness

Such fragmentation creates a substantial barrier to translation, generating a persistent “validation gap.” Although computer-centered metrics are easily reproducible, they remain imperfect surrogates for clinical benefit; they have not been shown to correlate with reduced diagnostic error, improved workflow efficiency, or measurable gains in patient safety. Human-centered assessments, while vital for understanding trust and usability, are often inconsistent, rely on non-standardized instruments, and rarely meet the evidentiary standards of clinical research, such as randomized or controlled evaluations. As a result, many XAI methods are deemed “successful” on technical grounds without demonstrating any empirically verified advantage in real clinical settings.

This gap is further widened by the near-absence of economic and implementation-oriented metrics. Current literature offers no standardized frameworks to assess cost-effectiveness, resource demands, or return on investment for XAI deployment. Without such data, healthcare administrators lack the evidence required to justify integration, even when technical performance appears promising. Collectively, these limitations underscore a systemic failure to couple computational rigor with clinical, operational, and economic utility. Addressing this deficiency demands a unified, multi-dimensional evaluation model capable of governing the safe, effective, and accountable deployment of XAI in healthcare.

To address this validation gap and the deficiencies revealed in Table 3, we propose a dual-dimensional XAI evaluation framework (Table 4) that integrates “technical reliability” with “clinical utility.” The framework advances the field in three key ways. First, it introduces methodological integration, unifying established clinical evaluation procedures112 with XAI performance metrics124 for the first time. This linkage closes the long-standing divide between algorithmic validation and clinically grounded assessment. Second, it offers context-aware adaptability, enabling the weighting and prioritization of metrics according to the clinical task and deployment environment. For example, imaging-based diagnostic systems require a stringent assessment of anatomical plausibility,114,125 whereas real-time use cases such as emergency triage necessitate optimization for computational efficiency. Third, the framework incorporates embedded ethical governance, elevating fairness auditing, privacy protection, and accountability from peripheral considerations to core, measurable components of clinical utility. By positioning ethical and legal safeguards as essential validation criteria rather than post-hoc additions, the framework establishes a more robust foundation for the safe and trustworthy deployment of XAI in healthcare. Ethical and legal considerations are integral to both the technical methodology and clinical application domains. Consequently, relevant ethical and legal metrics are encompassed within both corresponding evaluation columns.

Table 4.

Two-dimensional framework for medical XAI evaluation

Evaluation dimensions Technical reliability Clinical utility
Core metrics Explanation fidelity122,124 Effectiveness of policy support78,101
Local/global explanation accuracy Clinical workflow improvement
Computational efficiency112,118 User acceptance126,127
Healthcare specificity Anatomical plausibility in medical imaging59,114,125 Patient-understandable explanation format54,128
Ethical and legal compliance Fairness and bias audits129,130
Privacy preservation
Accountability and traceability131

This dual-dimensional structure represents a necessary evolution beyond existing evaluation taxonomies. Prior frameworks have focused almost exclusively on technical performance, often relying on algorithmic proxies for fidelity.124,132 The innovation of our approach lies in the systematic incorporation of a dedicated clinical utility dimension. By elevating user-centric and workflow-oriented metrics—such as decision-support effectiveness, workflow efficiency, and clinician acceptance—to a status coequal with technical reliability, the framework provides a direct bridge between algorithmic capability and real-world clinical adoption. It offers, for the first time, a unified structure for determining an XAI system’s readiness for deployment, shifting the field’s central question from “can it work?” to “does it work in practice?.” This dimension also embeds ethical and legal compliance as measurable, non-negotiable components of utility, reframing fairness, privacy, and accountability as essential properties rather than post-hoc obligations. Within this dual-dimensional framework, technical reliability encompasses three core categories of metrics. The first is explanation fidelity, captured through assessments of both local and global faithfulness124 as well as stability under input perturbations.122 The second is computational efficiency, measured through explanation generation time112 and resource consumption profiles.118 The third is medical plausibility, which requires that explanations remain anatomically coherent and clinically interpretable, maintaining spatial alignment with relevant structures or biomarkers.59,114 Together, these metrics establish a rigorous foundation for evaluating whether an XAI method is not only technically sound but also viable for use within the constraints of clinical practice.

The clinical utility dimension captures user-centered performance in real medical environments. It encompasses decision-support effectiveness, including improvements in diagnostic accuracy101 and gains in workflow efficiency,85 as well as user acceptance, reflected in clinician trust78 and patient comprehension.54 It also incorporates ethical and legal compliance, such as fairness audits to ensure equitable performance across demographic groups,129,130 privacy-preservation checks, and traceability mechanisms that support institutional and regulatory accountability.131,133

Current evaluation practices in medical XAI fall largely into three paradigms. The first is technical-validation-oriented evaluation, which emphasizes algorithmic robustness and typically relies on controlled experiments, such as pixel-level perturbation testing to assess stability.118 The second is clinical-validation-oriented evaluation, which measures real-world utility by observing clinician behavior (for example, diagnostic consistency101) or soliciting structured patient feedback. The third is a hybrid evaluation model that combines quantitative testing with expert qualitative.

Despite rapid methodological progress, three persistent challenges limit the maturity of XAI evaluation. First, there is a dimensional imbalance, with a disproportionate emphasis on algorithmic performance and inadequate attention to user experience and ethical considerations.50 Second, there is a lack of standardized criteria, illustrated by the inappropriate application of uniform evaluation standards across disparate disease categories. Third, validation depth remains insufficient, evidenced by the scarcity of longitudinal, real-world deployment studies.85

To address these limitations, we recommend three strategies. The first is the adoption of a human-AI collaborative evaluation paradigm, in which clinicians participate in the early co-design of evaluation criteria to ensure alignment with workflow realities.134 The second is the development of disease-specific evaluation standards, such as mandatory pathological confirmation in oncology-related XAI assessments.79,125 The third is the implementation of multi-center, long-term tracking studies conducted in accordance with FDA real-world evidence guidelines.133 Together, these measures provide a path toward more rigorous, clinically grounded, and ethically robust evaluation of XAI systems.

Explainable artificial intelligence tool library

Current efforts in medical XAI have produced a diverse ecosystem of open-source tool libraries, each designed to support distinct components of the explanation pipeline. These tools can be broadly categorized according to their core functions and technical characteristics (Figure 5).

Figure 5.

Figure 5

Categories of XAI toolboxes

The figure groups existing XAI tools according to their primary functions, including general-purpose explanation libraries, deep learning-specific interpretability tools, fairness and bias auditing frameworks, interactive analysis platforms, and high-performance computing systems. This overview highlights the diversity of the current XAI ecosystem and its fragmented support for clinical deployment.

General-purpose explanation tools

Kater135 is an open-source Python library designed to elucidate model behavior across a range of machine-learning pipelines. It supports both global and local explanations, producing feature-importance summaries, partial-dependence plots, and structured explanation graphs.

ELI522,83,136 offers similarly broad coverage, with compatibility across scikit-learn, XGBoost, LightGBM, and Keras models. Its visualizations make it straightforward to examine the contribution of key clinical variables to a model’s predictions.

InterpretML,137 developed by Microsoft, provides a comprehensive suite of interpretability algorithms with support for both global and local explanations. Its interactive visualization interface further lowers the entry barrier for clinical researchers, facilitating exploratory model assessment and hypothesis generation.

Deep learning specific tools

PyTorch Captum,138 developed by the PyTorch team, provides a unified interface for interpreting deep learning models across architectures. It implements a broad set of attribution methods—including Integrated Gradients, DeepLIFT, and feature-occlusion techniques, making it applicable to CNN, RNN, and Transformer-based systems.

TensorFlow tf-explain139 is tailored to the interpretability needs of vision models; its Grad-CAM implementation, validated across multiple medical imaging studies,29,114 reliably highlights disease-relevant regions. Its seamless integration within the TensorFlow ecosystem has made it a preferred tool in clinical imaging research workflows.

iNNvestigate140 offers particularly strong support for LRP, including variants such as LRP-ϵ. The fine-grained heatmaps it produces have proven especially valuable in tasks that require anatomical precision, such as brain tumor segmentation.65

Specialized function tools

AI Fairness 360 (AIF360)129,130 addresses core ethical and regulatory requirements by providing more than twenty fairness metrics and a suite of bias-mitigation algorithms. In applications such as disease-risk prediction, it enables systematic detection and correction of demographic biases, supporting the development of more equitable clinical models.

The What-If Tool (WIT)141 offers interactive, code-free model exploration, allowing clinical experts to probe decision boundaries, examine counterfactual scenarios, and compare individual cases. Its case-level analysis capabilities make it particularly useful for generating personalized, patient-specific treatment explanations.

High-performance computing platforms

H2O’s automated machine-learning platform streamlines the development of clinical prediction models by automating model selection, hyperparameter optimization, and performance comparison.142 Its distributed architecture enables efficient processing of large-scale datasets, including cohorts containing millions of patient records. The platform’s integrated explainability module provides a comprehensive suite of tools—from global feature-importance summaries to case-level attribution analyses—allowing clinicians and data scientists to audit both population-level patterns and individual predictions within a unified environment.

Specialized analysis tools

Neuroscope143 is designed for the in-depth exploration of the internal structure and functional pathways of neural networks. It enables researchers to analyze neuron-activation patterns, information-flow dynamics, and layer-level interactions across diverse architectures. The tool’s visualization interface reveals how representations evolve within complex deep learning models, making it particularly valuable for interpreting high-dimensional systems used in medical imaging and sequence analysis.

QLattice144,145,146 is a supervised learning framework for symbolic regression that constructs compact, mathematically explicit models. By representing learned hypotheses as intuitive computational graphs, it provides inherently interpretable structures that reveal how input variables interact to produce predictions. QLattice performs well on small, noisy clinical datasets while remaining scalable to larger cohorts. Its algorithmic design, inspired by Feynman’s path-integral formalism, allows rapid convergence through localized search and evolutionary refinement. All computations are executed locally, offering strong privacy assurances for sensitive medical data.

Taken together, these tools constitute key components of the technical infrastructure for medical XAI. Their successful clinical deployment, however, still requires careful coordination between computational demands, depth of explanation, and the cognitive load placed on clinical users.

For different stakeholders

XAI is increasingly deployed in healthcare because it addresses the heterogeneous priorities of multiple stakeholder groups.1,147 Prior research indicates that the core function of XAI is to meet the professional expectations and societal responsibilities attached to AI systems within specific clinical and operational contexts.7,148 In healthcare, the primary stakeholders include patients, clinicians, hospital administrators, medical researchers, AI developers, and regulatory agencies, each with distinct requirements for explanation quality, system behavior, and evaluation standards (Figure 6).10,81,111,149

Figure 6.

Figure 6

Stakeholders based on functional classification

The figure illustrates major stakeholders—including patients, clinicians, hospital administrators, medical researchers, AI developers, and regulatory bodies—and their respective roles and expectations regarding explainability. It emphasizes the heterogeneity of stakeholder requirements and the need for role-sensitive, context-aware XAI design.

Patients

As the direct recipients of care, patients require explanations that are intuitive, accessible, and aligned with their individual values and treatment goals.54,150 Their involvement in shared decision-making is expanding, yet current XAI research demonstrates a persistent lack of meaningful patient participation during system design. This often results in value misalignment, where the optimization target embedded in the model diverges from patient priorities. For example, a predictive model may emphasize survival endpoints while a patient receiving palliative care may prioritize comfort or symptom relief. When an XAI system cannot articulate these trade-offs, it risks undermining trust and compromising adherence to treatment plans.78,151

Clinicians

Clinicians require explanations that balance technical clarity with actionable clinical relevance. They need insight into which features or image regions drive model predictions (for instance, the localization of abnormalities in CT imaging41,125) and whether the model’s reasoning is consistent with established medical knowledge.134,152 These explanatory needs become particularly acute in multidisciplinary discussions where decisions depend on reconciling diverse expert perspectives.54,81

Hospital Management

Hospital leadership evaluates XAI through the lens of operational performance and institutional accountability.85,153 Their priorities include measurable improvements in diagnostic accuracy,101 cost-effectiveness, and workflow efficiency. Explanation systems also carry legal relevance, providing traceability mechanisms that support risk management and help mitigate the likelihood of diagnostic disputes.131,133

Medical Researchers

Researchers use XAI to interrogate the scientific validity of models and to support hypothesis generation.146,154 They rely on fine-grained analytical capabilities such as feature traceability (for example, temporal evolution of SHAP values96), bias detection, and model-behavior decomposition. These tools are essential for studying disease mechanisms and informing the design of clinical trials.21,155

AI Developers

Developers employ XAI to support model debugging, optimization, and robustness evaluation.40,138 Their focus lies in technical diagnostics, including feature attribution verification,35,44 hyperparameter sensitivity, and adversarial robustness testing.47 For this group, XAI serves as both a performance-monitoring tool and a structural guide for model refinement.

Regulatory Bodies

Regulatory authorities prioritize transparency, accountability, and safety in the oversight of medical AI. They require XAI systems to support fairness evaluation (such as demographic parity in diagnostic performance129,130) and to provide traceable decision pathways that meet legal and documentation standards for high-risk clinical applications.131,143,150,156 These requirements play a central role in approval processes and post-market surveillance.

The heterogeneity of stakeholder expectations introduces three foundational challenges for the development and deployment of XAI in healthcare. First, at the technical level, developers must navigate a persistent tension between producing explanations that are sufficiently detailed to be meaningful yet sufficiently concise to remain comprehensible to end users.149,157 Second, at the evaluative stage, the lack of standardized, role-sensitive metrics complicates efforts to assess user satisfaction and interpretability performance in a consistent and comparable manner across different professional groups.109,112 Third, at the implementation level, existing workflows rarely include systematic mechanisms for coordinating the objectives of clinicians, patients, administrators, and regulatory bodies, resulting in fragmented adoption pathways and inconsistent integration practices.133,134

Addressing these challenges requires a unified strategy centered on mental model alignment, a concept grounded in cognitive science. This approach underscores that system usability depends on the degree to which its observable behavior corresponds to the user’s internal representation of how it should function.7 For XAI, alignment is particularly critical. Clinicians reason through diagnostic heuristics and causal structures; patients interpret information through perceptions of trust, risk, and personal agency; and administrators operate within frameworks focused on value, efficiency, and institutional accountability. When these mental models are not aligned with the system’s explanatory outputs, technically robust XAI tools may appear opaque, counterintuitive, or misaligned with clinical priorities, limiting their acceptance and practical utility.

Several harmonization strategies have emerged in recent literature. For instance, one human-centered XAI framework enables configurable explanation granularity to accommodate users with different informational needs and levels of technical fluency.10 Role-sensitive tailoring of explanation content has also been emphasized, ensuring that each stakeholder receives context-relevant information.149 In parallel, disease-specific evaluation standards have been proposed for high-stakes domains, such as oncology, where pathological verification and anatomical plausibility are essential.79 Regulatory scholarship further highlights the importance of integrating privacy protection and compliance audits early in the deployment pipeline to support trustworthy and legally sound implementation.133

Future research should expand the adaptability of clinician-AI co-design models across diverse care settings, as suggested by134; develop standardized methodologies for real-world evidence generation in line with emerging regulatory guidance133; and establish validated metrics for human-centered evaluation constructs such as comprehensibility and trust.10 Progress in these areas will be essential for aligning XAI capabilities with the practical, ethical, and evidentiary needs of modern healthcare, thereby advancing the responsible and clinically meaningful deployment of AI systems.131,150,152

Practical application of frameworks

Bridging theory and practice: an application of the integration frameworks

The three integration frameworks proposed in this review are intended not only as conceptual models but as prescriptive tools for clinical deployment. To illustrate their practical function, this section provides a stepwise guide for a clinical-technical team tasked with integrating XAI into a high-stakes workflow: the deployment of a real-time sepsis early warning system (EWS).158,159,160

Step 1: Identify Stakeholder Priorities and Conflicts (Stakeholder Framework)

The integration effort begins by applying the stakeholder framework to map requirements and uncover potential tensions among stakeholder groups. Clinicians, who operate in environments already burdened by alert fatigue, require high-fidelity, causally meaningful explanations that can be interpreted rapidly at the bedside. Hospital administrators prioritize seamless EHR integration and demand quantifiable operational gains, such as reductions in ICU length-of-stay. Regulatory and ethics stakeholders require evidence of fairness, including formal audits demonstrating that the EWS does not underperform in specific racial or socioeconomic groups, an established vulnerability in sepsis prediction models.

Step 2: Select an Appropriate Explanation Method

The three-dimensional XAI classification framework is then used to prescribe an XAI method aligned with these constraints. Because the EWS must generate real-time, patient-specific alerts, the explanation granularity must be local; global methods are excluded. Achieving clinically competitive predictive performance requires a complex temporal model (such as an RNN or Transformer), making self-explanatory models unsuitable. The choice therefore lies between post-hoc model-specific methods (e.g., LRP) and model-agnostic methods. Model-agnostic methods are rejected due to their computational burden, which conflicts with real-time implementation needs, and their documented instability, which undermines clinician trust. A local, model-specific method (e.g., LRP) can be selected to satisfy these constraints, given its computational efficiency, established stability on temporal physiological data, and the ability to directly link alerts to evolving vital sign trends.

Step 3: Construct a Validation Strategy (Dual-Dimensional Evaluation Framework)

Guided by the dual-dimensional framework, the team develops a validation plan that addresses both technical reliability and clinical utility. Technical reliability focuses on fidelity testing (whether the explanations reflect the model’s internal logic) and stability testing (whether explanations remain consistent under minor input perturbations). Clinical utility is assessed through three components.

  • 1.

    Workflow improvement: A randomized controlled trial evaluating the EWS’s effect on clinically meaningful endpoints such as time-to-antibiotics and ICU admission rates.

  • 2.

    User acceptance: A qualitative study measuring trust calibration, perceived usefulness, and effects on alert fatigue among frontline clinicians.

  • 3.

    Ethical and legal compliance: A mandatory bias audit assessing performance across demographic groups, meeting regulatory requirements for high-risk predictive systems.

Together, these steps demonstrate how the stakeholder, 3D classification, and dual-dimensional evaluation frameworks function as an integrated governance pathway. Applied collectively, they guide clinical-technical teams from the identification of stakeholder tensions to the selection of appropriate XAI methods and, ultimately, to a validated, ethically grounded, and clinically deployable system.

Explainable artificial intelligence application (research) examples

Integrating XAI into healthcare has the potential to materially improve the quality, safety, and reliability of clinical decision-making.18,20,50,161 An examination of representative application studies yields several cross-cutting insights that shape real-world deployment.

Context specificity

XAI requirements vary substantially across clinical domains. Systems designed for time-critical infectious disease management demand low-latency, high-salience explanations,67,72 whereas chronic disease management often requires longitudinal, higher-granularity interpretability to support risk monitoring and shared decision-making.49,51

Workflow integration

The mode of integration is as consequential as the explanation method itself. Evidence shows that XAI tools embedded directly into existing clinical workflows—rather than appended as standalone explanatory overlays—achieve greater clinician uptake, reduce cognitive burden, and deliver more meaningful decision support.85

Human-AI collaboration

The utility of XAI is strongly mediated by the user’s interaction with the system. Studies indicate that multimodal explanation strategies, such as coupling visual heatmaps with concise textual rationales,114 can enhance comprehension, trust calibration, and diagnostic consistency among clinicians.101,127

To consolidate these observations, the current application landscape can be organized into five principal domains, as summarized in Table 5.

Table 5.

Synthesis of XAI application domains, functions, and key integration insights

Application domain Primary clinical function Key insights and integration challenges
Medical imaging diagnosis and interpretation Validation, localization, and discovery Insight: Moves beyond localization41,65,162 to scientific discovery(e.g., new prognostic markers).21,86,93
Challenge: Most critical function is validation by exposing model flaws, such as reliance on spurious features (e.g., articles in COVID-19 images).66
Clinical decision support Enhancing model transparency for structured and multimodal data Insight: Effectiveness is highly user-dependent; XAI can improve novice performance but disrupt expert workflow.81
Challenge: Poor or misaligned explanations directly erode trust, leading to increased costs (e.g., 35% rise in unnecessary confirmatory tests).54
Real-time monitoring Dynamic risk prediction from high-frequency data streams Insight: Focuses on dynamic, longitudinal risk assessment (e.g., clinical alarms),85,96 not static diagnosis.
Challenge: Requires computationally efficient methods to process continuous data streams (e.g., EEG and EMG) and provide immediate, actionable alerts.83,101
Process optimization Augmenting team performance and workflow efficiency Insight: XAI’s greatest value is in collaborative performance, where XAI + physician teams outperform individuals or AI alone.163,164
Challenge: Must address system-level tasks (e.g., drug repurposing and165 early detection rates92) and data governance challenges (e.g., sensitive HER data access).111
Human-Computer Interaction Designing for trust and cognitive alignment Insight: Reveals the core “perception gap”: developers prioritize technical fidelity while clinicians demand clinical rationality.134
Challenge: The primary barrier to integration is not technical, but cognitive. Solutions must be personalized to the user’s mental model.1,5

Medical imaging diagnosis and interpretation

In medical imaging, the role of XAI has progressed from basic transparency tools33,91,115,121 to fulfilling three distinct clinical functions: localization, validation, and discovery. Localization remains the most widely adopted application. Techniques such as LIME and Grad-CAM have been shown to enhance lesion identification and diagnostic performance across radiology domains76,113,166,167,168 and a range of oncologic indications, including brain, thyroid, and breast cancers.41,65,78,125,162,169,170 Comparative evaluations demonstrate that deep learning models augmented with XAI often surpass intrinsically interpretable, shallow models in diagnostic accuracy.171

The most consequential function for clinical integration, however, has been validation. During the COVID-19 pandemic, XAI played a critical role in revealing model failure modes that were not apparent from performance metrics alone. Explanation analyses showed that several high-performing systems were relying on spurious correlates—such as institution-specific text markers, acquisition artifacts, or scanner-dependent noise—rather than pathological findings, preventing inappropriate deployment in clinical settings.66,67,68,69,116,172,173,174,175

Beyond diagnostic support, XAI is emerging as a tool for scientific discovery by shifting from lesion characterization to prognostic inference. Studies combining XGBoost with SHAP have identified clinically relevant prognostic factors,86,93 while pan-cancer analyses using LRP have uncovered new interactions among molecular markers with potential biological significance.21 Evidence that XAI-derived annotations can strengthen the confidence of non-specialist radiologists176 further underscores its dual value as both a clinical and research instrument.

Clinical decision support

For CDS systems, which typically operate on structured or multimodal inputs such as ECG signals integrated with laboratory data,53,177,178,179 the value of XAI is defined less by visual appeal than by its impact on trust calibration and workflow fit. Adoption has accelerated in cardiovascular applications, including myocardial infarction classification and cardiac risk prediction,82,180,181 as well as in neurological domains such as Alzheimer’s disease and stroke.49,51,60,61

Across these settings, empirical studies consistently show that the effectiveness of XAI is strongly user-dependent and not uniformly beneficial. This variability constitutes a central integration barrier. A randomized controlled trial in neurological decision support demonstrated that explanations improved diagnostic accuracy for less experienced clinicians but introduced friction for expert users by interrupting established reasoning processes.81 Misaligned or overly intrusive explanations can therefore produce measurable harm. In a telehealth setting, insufficiently tailored explanations were shown to undermine clinician trust, increasing unnecessary confirmatory testing by 35 percent and diminishing workflow efficiency.54 Conversely, explanations that convey only superficial or low-value information—such as simplistic rule sets—have demonstrated limited benefit in complex tasks including glioma grading, where clinicians require more nuanced, pathology-aligned reasoning to support decision-making.101,128

Real-time monitoring

In contrast to static diagnostic applications, XAI for real-time monitoring operates on high-frequency, continuous data streams and is evaluated primarily by its capacity to deliver timely, clinically meaningful justifications for dynamic risk changes. Its central function is to support interpretable alerts that clarify why a patient’s risk trajectory is shifting, thereby improving both responsiveness and clinician trust. Representative applications include ward-level deterioration alerts,85 longitudinal all-cause mortality prediction,96 and real-time identification of evolving stroke risk factors such as age or triglyceride patterns.182

Integration in this domain is constrained simultaneously by computational and cognitive demands. Explanations must be generated with minimal latency to remain actionable in time-critical contexts. This requirement is illustrated in systems that analyze continuous EEG streams for stroke prediction,83 deliver operator-sensitive feedback in fetal ultrasound monitoring,102 and interpret EMG patterns to explain gesture classification in prosthetic control.183 In these settings, explanations must be concise, immediately interpretable, and tightly coupled to the triggering signal so that high-priority alerts do not add cognitive load or delay clinical action.

Process optimization

In this domain, XAI extends beyond individual patient diagnosis to system-level workflow optimization, where value is assessed in terms of efficiency and collaborative performance. Evidence indicates that XAI-assisted physician teams often achieve higher diagnostic accuracy than either individual clinicians or the AI system operating alone, providing a robust, evidence-based rationale for integrating XAI into clinical workflows.163,164

Beyond diagnostic support, XAI is increasingly applied to optimize broader healthcare processes, including improving early detection rates of lung diseases by combining sound biomarkers with interpretable explanations,103 accelerating drug discovery and repurposing,165,184 and enhancing precision pathology through bias detection and uncovering novel molecular mechanisms.152

The primary integration challenges in this domain are often non-technical, encompassing data governance, patient privacy, and practical difficulties associated with accessing and interpreting models trained on large-scale, sensitive electronic health records.111 Addressing these barriers is essential to realizing the full potential of XAI in optimizing healthcare systems.

Human-computer interaction

The successful integration of XAI into healthcare ultimately hinges on human-computer interaction (HCI). Rather than a standalone application, HCI serves as a cross-cutting methodology that emphasizes cognitive alignment and trust. Research in this area is critical for designing and validating explanations, highlighting a core “perception gap”: developers tend to prioritize technical interpretability, whereas clinicians require clinical rationality and workflow compatibility.134

HCI efforts, therefore, focus on creating user-centric, personalized interfaces—such as conversational AI systems185 or role-based interaction patterns99—that align explanations with the user’s cognitive background and mental model.1,5,17 By doing so, these designs foster calibrated trust rather than blind reliance on AI systems.153,186,187

Challenges and future outlook

Despite substantial advances in XAI—ranging from sophisticated interpretability methods and structured evaluation frameworks to explicit the consideration of diverse stakeholder needs—translating these approaches from laboratory proofs-of-concept to robust, scalable clinical deployment remains a formidable challenge. Successful integration into high-stakes medical decision-making requires addressing gaps spanning technical, cognitive, and ethical domains. We synthesize these unresolved issues into seven critical challenges, each paired with a prospective research agenda that demands interdisciplinary collaboration across AI, cognitive science, and clinical medicine.

Reconciling model performance with interpretability

The classic trade-off between predictive accuracy and interpretability persists: complex models excel in high-dimensional tasks but are opaque, whereas simpler models are transparent but underperform. A paradigm shift toward inherently interpretable hybrid architectures is needed. Neurosymbolic AI, which chains deep learning perceptual capacity with symbolic reasoning engines (e.g., clinical knowledge graphs), offers one pathway to ensure decisions are both accurate and medically coherent.188 Extending concept bottleneck models (CBMs) with unsupervised or self-supervised discovery of clinical concepts can further enable transparency without manual feature engineering, maintaining high fidelity to ground truth.189

Adaptive explainable artificial intelligence for heterogeneous users and clinical scenarios

Static explanations fail to meet the diverse requirements of stakeholders, including clinicians, patients, and regulators. Future XAI must be context-aware and interactive, employing Dynamic Role-Based Adaptation where output format (e.g., saliency maps, counterfactuals) is tailored to user role and clinical context. Robust Interactive Explanatory Query Interfaces will enable clinicians to perform contrastive interrogation (“Why this diagnosis and not the differential?”) and sensitivity analysis (“Which input change would flip the recommendation?”), balancing exploration and exploitation based on case uncertainty and user expertise.

Standardizing validation through human-grounded evaluation

Current evaluation methods—often reliant on technical proxies or expert judgment—lack standardization, limiting cross-study comparability and long-term validation. Evidence-based evaluation protocols, including human-centered iterative evaluation and longitudinal in-situ studies, are needed to measure the impact of explanations on clinician performance, diagnostic errors, workflow efficiency, and alert fatigue. Large-scale Medical XAI Benchmarks with multi-modal ground truths and integrated health economic analyses will provide objective, quantitative validation for both clinical efficacy and cost-benefit considerations.

Ensuring equity via causal inference

Systematic biases in medical datasets can be amplified or hidden by XAI, risking algorithmic discrimination. Integrating Causal XAI allows the identification and mitigation of such biases, constraining learned mechanisms with known medical causal graphs to prevent reliance on spurious correlations (e.g., patient race as a proxy for kidney function). Explainability-driven auditing should quantify disparate impacts across intersectional subgroups, ensuring fairness is embedded in model design and evaluation.

Personalization and dynamics for longitudinal care

Static explanations are insufficient for longitudinal healthcare, where patient histories, genetic profiles, and treatment responses evolve over time. Dynamic Causal XAI must provide evolving explanations, forecasting how reasoning changes as new data accrue. This capability is essential for truly personalized medicine,190,191 enabling clinicians to understand adaptive treatment plans and evolving risk predictions over the patient trajectory.

Cognitive alignment and trust calibrating

Misalignment between a user’s mental model and the model’s internal reasoning can result in misleading trust or over-reliance. Research should develop metrics for Human-XAI Mental Model Alignment and methods for detecting manipulative explanations. Techniques such as Explanatory Red Teaming and “counter-explanations” can identify locally faithful but globally misleading rationales, promoting calibrated trust while addressing value misalignments—such as when patient goals (e.g., palliative care) differ from AI-optimized endpoints.

Accountability, privacy, and ethical governance

Explanation artifacts introduce new ethical, legal, and social risks. XAI must integrate privacy-preserving mechanisms (e.g., differential privacy and secure federated learning) and transparent Auditing Frameworks to prevent unintended exposure of sensitive information or algorithmic discrimination. Clear Accountability Proxies are needed to delineate responsibilities among developers, clinicians, and institutions. Ethical and legal safeguards must be embedded throughout the AI lifecycle—from design and data sourcing to deployment and post-market monitoring—rather than applied as post hoc checklists.

Together, these seven challenges define a research agenda for advancing medical XAI from technically sophisticated prototypes to trusted, ethically aligned, and clinically integrated systems. Meeting them requires coordinated effort across AI research, clinical practice, cognitive science, and health policy.

Conclusion

XAI will not become clinically consequential by accumulating more explanation techniques alone. Our synthesis of 170 studies indicates that the dominant translation barrier is not methodological scarcity but a persistent clinical integration gap: explanations are rarely treated as verifiable, auditable components of clinical decision support that must satisfy workflow constraints, evidentiary standards, and governance requirements. In practice, much of the support provided by XAI remains “persuasive” rather than “actionable.” Such support may look intuitive, yet lack demonstrated faithfulness, stability, and operational value when deployed in real clinical settings.

To address this gap, we position clinical integration as the organizing principle and introduce three interoperable instruments: a clinically oriented three-dimensional taxonomy that prescribes method selection based on clinical task demands rather than descriptive categories; a dual-dimensional evaluation framework that couples technical reliability with clinical utility; and a stakeholder-centered integration model that operationalizes “meaningful explainability” as role and context-dependent. Together, these frameworks transcend descriptive cataloging to offer actionable guidance for aligning XAI methods with clinical workflows, validating their utility, and addressing the diverse needs of clinicians, patients, administrators, and regulators.

Our analysis underscores that the principal barrier to widespread XAI adoption is not a lack of technical methods, but a persistent mismatch between algorithmically generated explanations and clinical reasoning processes. This mismatch is shaped by trade-offs among explanation fidelity, interpretability, and clinical usability: higher-fidelity explanations may increase cognitive burden or latency in time-critical settings, whereas overly simplified explanations can induce over-trust and mask model brittleness. Accordingly, explainability in healthcare should be optimized rather than universally maximized, and clinical readiness should be judged by deployment-level evidence rather than perceived plausibility alone. At minimum, an XAI-enabled clinical decision support tool should demonstrate: (1) faithfulness and stability under clinically plausible perturbations; (2) clinical plausibility, such as anatomically coherent attributions in imaging or physiologically consistent temporal rationales in monitoring; (3) workflow compatibility, with demonstrated impact on decision time, alert fatigue, and downstream testing; (4) trust calibration, showing improvements in decision quality without inappropriate reliance; and (5) governance properties, including traceability, bias auditing, and privacy-preserving handling of explanation artifacts.

Ultimately, the next phase of XAI in clinical deployment calls for a paradigmatic shift from proxy-based explanation quality to evidence-based explanation effects. Explanations should be conceptualized and evaluated as clinical interventions, validated by their demonstrable causal impact on patient-relevant outcomes and real-world workflow performance, rather than by technical surrogates alone. Our three integration frameworks operationalize this shift by linking stakeholder needs to method choice and a dual-dimensional evidence package, thereby advancing XAI from methodological novelty to deployable, monitored clinical decision support. Only under this evidence-first paradigm can XAI support calibrated trust and meet the accountability standards required in high-stakes care.

Acknowledgments

This work is supported by the National Natural Science Foundation of China (Grant 72192823).

Author contributions

Conceptualization: K. Z., D. W., and W. Z.; methodology: K. Z. and D. W.; investigation: K. Z., D. W., and J. X.; writing – original draft: K. Z., D. W., and W. Z.; writing – review and editing: K. Z., D. W., F. L., and W. Z.; funding acquisition: W. Z.; resources: W. Z.; supervision: D. W., J. X., and W. Z.

Declaration of interests

The authors declare no competing interests.

References

  • 1.Barredo Arrieta A., Díaz-Rodríguez N., Del Ser J., Bennetot A., Tabik S., Barbado A., Garcia S., Gil-Lopez S., Molina D., Benjamins R., et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion. 2020;58:82–115. [Google Scholar]
  • 2.Gunning D., Stefik M., Choi J., Miller T., Yang G.Z., Stumpf S. XAI—explainable artificial intelligence. Sci. Robot. 2019;4 doi: 10.1126/scirobotics.aay7120. [DOI] [PubMed] [Google Scholar]
  • 3.DW G.D.A. DARPA’s explainable artificial intelligence program. AI Mag. 2019;40:44. [Google Scholar]
  • 4.Ali S., Akhlaq F., Imran A.S., Kastrati Z., Daudpota S.M., Moosa M. The enlightening role of explainable artificial intelligence in medical & healthcare domains: a systematic literature review. Comput. Biol. Med. 2023;166 doi: 10.1016/j.compbiomed.2023.107555. [DOI] [PubMed] [Google Scholar]
  • 5.Escalante H.J., Escalera S., Guyon I., Baro X., Gtclitirk Y., Guicli U., van Gerven M. Explainable and Interpretable Models in Computer Vision and Machine Learning. Springer International Publishing; Cham, Switzerland: 2018. [Google Scholar]
  • 6.Castelvecchi D. Can we open the black box of AI? Nature. 2016;538:20–23. doi: 10.1038/538020a. [DOI] [PubMed] [Google Scholar]
  • 7.Wang D., Yang Q., Abdul A., Lim B.Y. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. Glasgow, Scotland. 2019. Designing theory-driven user-centric explainable AI; pp. 1–15. [Google Scholar]
  • 8.Shin D. The effects of explainability and causability on perception, trust, and acceptance: implications for explainable AI. Int. J. Hum. Comput. Stud. 2021;146 [Google Scholar]
  • 9.Lapuschkin S., Wäldchen S., Binder A., Montavon G., Samek W., Müller K.R. Unmasking clever hans predictors and assessing what machines really learn. Nat. Commun. 2019;10:1096. doi: 10.1038/s41467-019-08987-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kong X., Liu S., Zhu L. Toward human-centered XAI in practice: a survey. Mach. Intell. Res. 2024;21:740–770. [Google Scholar]
  • 11.Jonathan D., Sean P., Andrew A., Margaret M.B. International Conference on Intelligent User Interfaces. 2018. What should be in an XAI explanation? What IFT reveals. [Google Scholar]
  • 12.Kim T.W. Explainable artificial intelligence (XAI), the goodness criteria and the grasp-ability test. arXiv. 2018 doi: 10.48550/arXiv.1810.09598. Preprint at. [DOI] [Google Scholar]
  • 13.Yee C.M., Maimó M.F.R., Ramis S., Sansó R.M. Handbook of Artificial Intelligence in Healthcare: Vol 2: Practicalities and Prospects. Springer International Publishing; Cham: 2021. Advances in XAI: Explanation Interfaces in healthcare; pp. 357–369. [Google Scholar]
  • 14.Gunning D., Vorm E., Wang Y., Turek M. Authorea Preprints; 2021. DARPA’s Explainable AI (XAI) Program: A Retrospective. [Google Scholar]
  • 15.Islam M.R., Ahmed M.U., Barua S., Begum S. A systematic review of explainable artificial intelligence in terms of different application domains and tasks. Appl. Sci. 2022;12:1353. [Google Scholar]
  • 16.Veit R., Josefine R., Gregor K., Georg S., Stefan E., Christian B. Taming the chaos?! Using eXplainable artificial intelligence (XAI) to tackle the complexity in mental health research. Eur. Child Adolesc. Psychiatr. 2021;30:1143–1146. doi: 10.1007/s00787-021-01836-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Rong Y., Leemann T., Nguyen T.T., Fiedler L., Qian P., Unhelkar V., Seidel T., Kasneci G., Kasneci E. Towards human-centered explainable ai: A survey of user studies for model explanations. IEEE Trans. Pattern Anal. Mach. Intell. 2024;46:2104–2122. doi: 10.1109/TPAMI.2023.3331846. [DOI] [PubMed] [Google Scholar]
  • 18.Secinaro S., Calandra D., Secinaro A., Muthurangu V., Biancone P. The role of artificial intelligence in healthcare: a structured literature review. BMC Med. Inform. Decis. Mak. 2021;21:125. doi: 10.1186/s12911-021-01488-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tjoa E., Guan C. A survey on explainable artificial intelligence (XAI): toward medical XAI. IEEE Trans. Neural Netw. Learn. Syst. 2021;32:4793–4813. doi: 10.1109/TNNLS.2020.3027314. [DOI] [PubMed] [Google Scholar]
  • 20.Chaddad A., Peng J., Xu J., Bouridane A. Survey of explainable AI techniques in healthcare. Sensors. 2023;23:634. doi: 10.3390/s23020634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Keyl J., Keyl P., Montavon G., Hosch R., Brehmer A., Mochmann L., Jurmeister P., Dernbach G., Kim M., Koitka S., et al. Decoding pan-cancer treatment outcomes using multimodal real-world data and explainable artificial intelligence. Nat. Cancer. 2025;6:307–322. doi: 10.1038/s43018-024-00891-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Chen V., Yang M., Cui W., Kim J.S., Talwalkar A., Ma J. Applying interpretable machine learning in computational biology—pitfalls, recommendations and opportunities for new developments. Nat. Methods. 2024;21:1454–1461. doi: 10.1038/s41592-024-02359-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Johannssen A., Chukhrova N. The crucial role of explainable artificial intelligence (XAI) in improving health care management. Health Care Manag. Sci. 2025;28:565–570. doi: 10.1007/s10729-025-09720-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Adeniran A.A., Onebunne A.P., William P. Explainable AI (XAI) in healthcare: Enhancing trust and transparency in critical decision-making. World J. Adv. Res. Rev. 2024;23:2647–2658. [Google Scholar]
  • 25.Liu Y., Liu C., Zheng J., Xu C., Wang D. Improving explainability and integrability of medical AI to promote health care professional acceptance and use: mixed systematic review. J. Med. Internet Res. 2025;27 doi: 10.2196/73374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Abbas Q., Jeong W., Lee S.W. Explainable AI in clinical decision support systems: a meta-analysis of methods, applications, and usability challenges. Healthcare MDPI. 2025;13:2154. doi: 10.3390/healthcare13172154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Liberati A., Altman D.G., Tetzlaff J., Mulrow C., Gøtzsche P.C., Ioannidis J.P.A., Clarke M., Devereaux P.J., Kleijnen J., Moher D. The PRISMA statement for reporting systematic reviews and meta-analyses of studies that evaluate healthcare interventions: explanation and elaboration. Bmj. 2009;339:b2700. doi: 10.1136/bmj.b2700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Sadeghi Z., Alizadehsani R., Cifci M.A., Kausar S., Rehman R., Mahanta P., Bora P.K., Almasri A., Alkhawaldeh R.S., Hussain S., et al. A review of explainable artificial intelligence in healthcare. Comput. Electr. Eng. 2024;118 [Google Scholar]
  • 29.Sheu R.K., Pardeshi M.S. A survey on medical explainable AI (XAI): recent progress, explainability approach, human interaction and scoring system. Sensors. 2022;22:8068. doi: 10.3390/s22208068. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Saranya A., Subhashini R. A systematic review of explainable artificial intelligence models and applications: Recent developments and future trends. Decis. Anal. J. 2023;7 [Google Scholar]
  • 31.Moraffah R., Karami M., Guo R., Raglin A., Liu H. Causal interpretability for machine learning-problems, methods and evaluation. SIGKDD Explor. Newsl. 2020;22:18–33. [Google Scholar]
  • 32.Mothilal R.K., Sharma A., Tan C. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 2020. Explaining machine learning classifiers through diverse counterfactual explanations; pp. 607–617. [Google Scholar]
  • 33.Nazir S., Dickson D.M., Akram M.U. Survey of explainable artificial intelligence techniques for biomedical imaging with deep neural networks. Comput. Biol. Med. 2023;156 doi: 10.1016/j.compbiomed.2023.106668. [DOI] [PubMed] [Google Scholar]
  • 34.Ribeiro M.T., Singh S., Guestrin C. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, USA. 2016. “Why should i trust you?” Explaining the predictions of any classifier; pp. 1135–1144. [Google Scholar]
  • 35.Lundberg S.M., Lee S.I. A unified approach to interpreting model predictions. 2017. A unified approach to interpreting model predictions; p. 30. [Google Scholar]
  • 36.Rai A. Explainable AI: From black box to glass box. J. Acad. Mark. Sci. 2020;48:137–141. [Google Scholar]
  • 37.Linardatos P., Papastefanopoulos V., Kotsiantis S. Explainable AI: A review of machine learning interpretability methods. Entropy. 2020;23:18. doi: 10.3390/e23010018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Quinlan J.R. Induction of decision trees. Mach. Learn. 1986;1:81–106. [Google Scholar]
  • 39.Draper N.R., Smith H. John Wiley & Sons; Hoboken, America: 1998. Applied Regression Analysis. [Google Scholar]
  • 40.Montavon G., Samek W., Müller K.R. Methods for interpreting and understanding deep neural networks. Digit. Signal Process. 2018;73:1–15. [Google Scholar]
  • 41.Esmaeili M., Vettukattil R., Banitalebi H., Krogh N.R., Geitung J.T. Explainable artificial intelligence for human-machine interaction in brain tumor localization. J. Pers. Med. 2021;11:1213. doi: 10.3390/jpm11111213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Selvaraju R.R., Cogswell M., Das A., Vedantam R., Parikh D., Batra D. Proceedings of the IEEE International Conference on Computer Vision. Venice, Italy. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization; pp. 618–626. [Google Scholar]
  • 43.Rajkomar A., Oren E., Chen K., Dai A.M., Hajaj N., Hardt M., Liu P.J., Liu X., Marcus J., Sun M., et al. Scalable and accurate deep learning with electronic health records. npj Digit. Med. 2018;1:18. doi: 10.1038/s41746-018-0029-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Lundberg S.M., Erion G., Chen H., DeGrave A., Prutkin J.M., Nair B., Katz R., Himmelfarb J., Bansal N., Lee S.I. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell. 2020;2:56–67. doi: 10.1038/s42256-019-0138-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Lim B., Arık S.Ö., Loeff N., Pfister T. Temporal fusion transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 2021;37:1748–1764. [Google Scholar]
  • 46.Bahdanau D., Cho K., Bengio Y. Neural machine translation by jointly learning to align and translate. arXiv. 2014 doi: 10.48550/arXiv.1409.0473. Preprint at. [DOI] [Google Scholar]
  • 47.Goodfellow I.J., Shlens J., Szegedy C. Explaining and harnessing adversarial examples. arXiv. 2014 doi: 10.48550/arXiv.1412.6572. Preprint at. [DOI] [Google Scholar]
  • 48.Islam T., Sheakh M.A., Tahosin M.S., Hena M.H., Akash S., Bin Jardan Y.A., FentahunWondmie G., Nafidi H.A., Bourhia M. Predictive modeling for breast cancer classification in the context of Bangladeshi patients by use of machine learning approach with explainable AI. Sci. Rep. 2024;14:8487. doi: 10.1038/s41598-024-57740-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Mahmud T., Barua K., Habiba S.U., Sharmen N., Hossain M.S., Andersson K. An explainable ai paradigm for alzheimer’s diagnosis using deep transfer learning. Diagnostics. 2024;14:345. doi: 10.3390/diagnostics14030345. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Loh H.W., Ooi C.P., Seoni S., Barua P.D., Molinari F., Acharya U.R. Application of explainable artificial intelligence for healthcare: a systematic review of the last decade (2011–2022) Comput. Methods Programs Biomed. 2022;226 doi: 10.1016/j.cmpb.2022.107161. [DOI] [PubMed] [Google Scholar]
  • 51.Fania A., Monaco A., Amoroso N., Bellantuono L., Cazzolla Gatti R., Firza N., Lacalamita A., Pantaleo E., Tangaro S., Velichevskaya A., Bellotti R. Machine learning and XAI approaches highlight the strong connection between O3 and NO2 pollutants and Alzheimer’s disease. Sci. Rep. 2024;14:5385. doi: 10.1038/s41598-024-55439-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Wani N.A., Kumar R., Bedi J. DeepXplainer: An interpretable deep learning based approach for lung cancer detection using explainable artificial intelligence. Comput. Methods Programs Biomed. 2024;243 doi: 10.1016/j.cmpb.2023.107879. [DOI] [PubMed] [Google Scholar]
  • 53.Botros J., Mourad-Chehade F., Laplanche D. Explainable multimodal data fusion framework for heart failure detection: Integrating CNN and XGBoost. Biomed. Signal Process Control. 2025;100 [Google Scholar]
  • 54.Gomez C., Smith B.L., Zayas A., Unberath M., Canares T. Explainable AI decision support improves accuracy during telehealth strep throat screening. Commun. Med. 2024;4:149. doi: 10.1038/s43856-024-00568-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Moulaei K., Afrash M.R., Parvin M., Shadnia S., Rahimi M., Mostafazadeh B., Evini P.E.T., Sabet B., Vahabi S.M., Soheili A., et al. Explainable artificial intelligence (XAI) for predicting the need for intubation in methanol-poisoned patients: a study comparing deep and machine learning models. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-66481-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Huang S.C., Kothari T., Banerjee I., Chute C., Ball R.L., Borus N., Huang A., Patel B.N., Rajpurkar P., Irvin J., et al. PENet—a scalable deep-learning model for automated diagnosis of pulmonary embolism using volumetric CT imaging. npj Digit. Med. 2020;3:61. doi: 10.1038/s41746-020-0266-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Jahmunah V., Ng E.Y.K., Tan R.S., Oh S.L., Acharya U.R. Explainable detection of myocardial infarction using deep learning models with Grad-CAM technique on ECG signals. Comput. Biol. Med. 2022;146 doi: 10.1016/j.compbiomed.2022.105550. [DOI] [PubMed] [Google Scholar]
  • 58.Elsawy A., Keenan T.D.L., Chen Q., Thavikulwat A.T., Bhandari S., Quek T.C., Goh J.H.L., Tham Y.C., Cheng C.Y., Chew E.Y., Lu Z. A deep network DeepOpacityNet for detection of cataracts from color fundus photographs. Commun. Med. 2023;3:184. doi: 10.1038/s43856-023-00410-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Song D., Yao J., Jiang Y., Shi S., Cui C., Wang L., Wang L., Wu H., Tian H., Ye X., et al. A new xAI framework with feature explainability for tumors decision-making in ultrasound data: comparing with Grad-CAM. Comput. Methods Programs Biomed. 2023;235 doi: 10.1016/j.cmpb.2023.107527. [DOI] [PubMed] [Google Scholar]
  • 60.Alatrany A.S., Khan W., Hussain A., Kolivand H., Al-Jumeily D. An explainable machine learning approach for Alzheimer’s disease classification. Sci. Rep. 2024;14:2637. doi: 10.1038/s41598-024-51985-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Adarsh V., Gangadharan G.R., Fiore U., Zanetti P. Multimodal classification of Alzheimer's disease and mild cognitive impairment using custom MKSCDDL kernel over CNN with transparent decision-making for explainable diagnosis. Sci. Rep. 2024;14:1774. doi: 10.1038/s41598-024-52185-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Eitel F., Ritter K. Interpretability of Machine Intelligence in Medical Image Computing and Multimodal Learning for Clinical Decision Support: Second International Workshop, iMIMIC 2019, and 9th International Workshop, ML-CDS 2019, Held in Conjunction with MICCAI 2019, Shenzhen, China, October 17, 2019, Proceedings 9. Springer International Publishing; 2019. Alzheimer’s disease neuroimaging initiative (ADNI). Testing the robustness of attribution methods for convolutional neural networks in MRI-based Alzheimer’s disease classification; pp. 3–11. [Google Scholar]
  • 63.Meena J., Hasija Y. Application of explainable artificial intelligence in the identification of squamous cell carcinoma biomarkers. Comput. Biol. Med. 2022;146 doi: 10.1016/j.compbiomed.2022.105505. [DOI] [PubMed] [Google Scholar]
  • 64.Lombardi A., Diacono D., Amoroso N., Biecek P., Monaco A., Bellantuono L., Pantaleo E., Logroscino G., De Blasi R., Tangaro S., Bellotti R. A robust framework to investigate the reliability and stability of explainable artificial intelligence markers of mild cognitive impairment and Alzheimer’s disease. Brain Inform. 2022;9:17. doi: 10.1186/s40708-022-00165-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Zeineldin R.A., Karar M.E., Elshaer Z., Coburger J., Wirtz C.R., Burgert O., Mathis-Ullrich F. Explainable hybrid vision transformers and convolutional network for multimodal glioma segmentation in brain MRI. Sci. Rep. 2024;14:3713. doi: 10.1038/s41598-024-54186-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.DeGrave A.J., Janizek J.D., Lee S.I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat. Mach. Intell. 2021;3:610–619. [Google Scholar]
  • 67.Sarp S., Catak F.O., Kuzlu M., Cali U., Kusetogullari H., Zhao Y., Ates G., Guler O. An XAI approach for COVID-19 detection using transfer learning with X-ray images. Heliyon. 2023;9 doi: 10.1016/j.heliyon.2023.e15137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Olar A., Biricz A., Bedőházi Z., Sulyok B., Pollner P., Csabai I. Automated prediction of COVID-19 severity upon admission by chest X-ray images and clinical metadata aiming at accuracy and explainability. Sci. Rep. 2023;13:4226. doi: 10.1038/s41598-023-30505-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Bhandari M., Shahi T.B., Siku B., Neupane A. Explanatory classification of CXR images into COVID-19, pneumonia and tuberculosis using deep learning and XAI. Comput. Biol. Med. 2022;150 doi: 10.1016/j.compbiomed.2022.106156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Sun J., Shi W., Giuste F.O., Vaghani Y.S., Tang L., Wang M.D. Improving explainable ai with patch perturbation-based evaluation pipeline: a covid-19 x-ray image analysis case study. Sci. Rep. 2023;13 doi: 10.1038/s41598-023-46493-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Chung J., Kim D., Choi J., Yune S., Song K.D., Kim S., Chua M., Succi M.D., Conklin J., Longo M.G.F., et al. Prediction of oxygen requirement in patients with COVID-19 using a pre-trained chest radiograph xAI model: efficient development of auditable risk prediction models via a fine-tuning approach. Sci. Rep. 2022;12 doi: 10.1038/s41598-022-24721-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Rajpoot R., Gour M., Jain S., Semwal V.B. Integrated ensemble CNN and explainable AI for COVID-19 diagnosis from CT scan and X-ray images. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-75915-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Hou S., Han J. COVID-19 detection via a 6-layer deep convolutional neural network. Comput. Model. Eng. Sci. 2022;130:855–869. [Google Scholar]
  • 74.Pereira S., Meier R., Alves V., Reyes M., Silva C.A. In: Understanding and Interpreting Machine Learning in Medical Image Computing Applications. Stoyanov D., Taylor Z., Kia S.M., Oguz I., Reyes M., Martel A., Maier-Hein L., Marquand A.F., Duchesnay E., Löfstedtm T., editors. Springer; Cham, Switzerland: 2018. Automatic brain tumor grading from MRI data using convolutional neural networks and quality assessment. [Google Scholar]
  • 75.Zeineldin R.A., Karar M.E., Elshaer Z., Coburger J., Wirtz C.R., Burgert O., Mathis-Ullrich F. Explainability of deep neural networks for MRI analysis of brain tumors. Int. J. Comput. Assist. Radiol. Surg. 2022;17:1673–1683. doi: 10.1007/s11548-022-02619-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Qian J., Li H., Wang J., He L. Recent advances in explainable artificial intelligence for magnetic resonance imaging. Diagnostics. 2023;13:1571. doi: 10.3390/diagnostics13091571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Springer Nature; 2019. Interpretability of Machine Intelligence in Medical Image Computing and Multimodal Learning for Clinical Decision Support: Second International Workshop, IMIMIC 2019, and 9th International Workshop, ML-CDS 2019, Held in Conjunction with MICCAI 2019, Shenzhen, China, October 17, 2019, Proceedings. [Google Scholar]
  • 78.Yan L., Liang Z., Zhang H., Zhang G., Zheng W., Han C., Yu D., Zhang H., Xie X., Liu C., et al. A domain knowledge-based interpretable deep learning system for improving clinical breast ultrasound diagnosis. Commun. Med. 2024;4:90. doi: 10.1038/s43856-024-00518-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Ukwuoma C.C., Cai D., Eziefuna E.O., Oluwasanmi A., Abdi S.F., Muoka G.W., Thomas D., Sarpong K. Enhancing histopathological medical image classification for Early cancer diagnosis using deep learning and explainable AI–LIME & SHAP. Biomed. Signal Process Control. 2025;100 [Google Scholar]
  • 80.Watson M., Awwad Shiekh Hasan B., Al Moubayed N. Using model explanations to guide deep learning models towards consistent explanations for EHR data. Sci. Rep. 2022;12 doi: 10.1038/s41598-022-24356-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Gombolay G.Y., Silva A., Schrum M., Gopalan N., Hallman-Cooper J., Dutt M., Gombolay M. Effects of explainable artificial intelligence in neurology decision support. Ann. Clin. Transl. Neurol. 2024;11:1224–1235. doi: 10.1002/acn3.52036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.El-Sofany H., Bouallegue B., El-Latif Y.M.A. A proposed technique for predicting heart disease using machine learning algorithms and an explainable AI method. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-74656-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Islam M.S., Hussain I., Rahman M.M., Park S.J., Hossain M.A. Explainable artificial intelligence model for stroke prediction using EEG signal. Sensors. 2022;22:9859. doi: 10.3390/s22249859. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Leonardi G., Montani S., Striani M. Explainable process trace classification: An application to stroke. J. Biomed. Inform. 2022;126 doi: 10.1016/j.jbi.2021.103981. [DOI] [PubMed] [Google Scholar]
  • 85.De T., Giri P., Mevawala A., Nemani R., Deo A. Explainable AI: a hybrid approach to generate human-interpretable explanation for deep learning prediction. Procedia Comput. Sci. 2020;168:40–48. [Google Scholar]
  • 86.Flyckt R.N.H., Sjodsholm L., Henriksen M.H.B., Brasen C.L., Ebrahimi A., Hilberg O., Hansen T.F., Wiil U.K., Jensen L.H., Peimankar A. Pulmonologists-Level lung cancer detection based on standard blood test results and smoking status using an explainable machine learning approach. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-82093-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Orcutt X., Chen K., Mamtani R., Long Q., Parikh R.B. Evaluating generalizability of oncology trial results to real-world patients using machine learning-based trial emulations. Nat. Med. 2025;31:457–465. doi: 10.1038/s41591-024-03352-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.El-Sappagh S., Alonso J.M., Ali F., Ali A., Jang J.H., Kwak K.S. An ontology-based interpretable fuzzy decision support system for diabetes diagnosis. IEEE Access. 2018;6:37371–37394. [Google Scholar]
  • 89.Racoceanu D. Explainable AI and its instantiation in computational pathology for a better understanding of alzheimer’s disease. Alzheimers Dementia. 2023;19 [Google Scholar]
  • 90.Amoroso N., Pomarico D., Fanizzi A., Didonna V., Giotta F., La Forgia D., Latorre A., Monaco A., Pantaleo E., Petruzzellis N., et al. A roadmap towards breast cancer therapies supported by explainable artificial intelligence. Applied Sciences. 2021;11:4881. [Google Scholar]
  • 91.Van der Velden B.H.M., Kuijf H.J., Gilhuijs K.G.A., Viergever M.A. Explainable artificial intelligence (XAI) in deep learning-based medical image analysis. Med. Image Anal. 2022;79 doi: 10.1016/j.media.2022.102470. [DOI] [PubMed] [Google Scholar]
  • 92.Chen Z., Liang N., Li H., Zhang H., Li H., Yan L., Hu Z., Chen Y., Zhang Y., Wang Y., et al. Exploring explainable AI features in the vocal biomarkers of lung disease. Comput. Biol. Med. 2024;179 doi: 10.1016/j.compbiomed.2024.108844. [DOI] [PubMed] [Google Scholar]
  • 93.Alabi R.O., Elmusrati M., Leivo I., Almangush A., Mäkitie A.A. Machine learning explainability in nasopharyngeal cancer survival using LIME and SHAP. Sci. Rep. 2023;13:8984. doi: 10.1038/s41598-023-35795-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Hu C., Li L., Li Y., Wang F., Hu B., Peng Z. Explainable machine-learning model for prediction of in-hospital mortality in septic patients requiring intensive care unit readmission. Infect. Dis. Ther. 2022;11:1695–1713. doi: 10.1007/s40121-022-00671-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Slijepcevic D., Horst F., Lapuschkin S., Horsak B., Raberger A.M., Kranzl A., Samek W., Breiteneder C., Schöllhorn W.I., Zeppelzauer M. Explaining machine learning models for clinical gait analysis. ACM Trans. Comput. Healthc. 2021;3:1–27. [Google Scholar]
  • 96.Qiu W., Chen H., Dincer A.B., Lundberg S., Kaeberlein M., Lee S.I. Interpretable machine learning prediction of all-cause mortality. Commun. Med. 2022;2:125. doi: 10.1038/s43856-022-00180-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Kokkotis C., Moustakidis S., Tsatalas T., Ntakolia C., Chalatsis G., Konstadakos S., Hantes M.E., Giakas G., Tsaopoulos D. Leveraging explainable machine learning to identify gait biomechanical parameters associated with anterior cruciate ligament injury. Sci. Rep. 2022;12:6647. doi: 10.1038/s41598-022-10666-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Quintana-Quintana O.J., Aceves-Fernández M.A., Pedraza-Ortega J.C., Alfonso-Francia G., Tovar-Arriaga S. Deep learning techniques for retinal layer segmentation to aid ocular disease diagnosis: A review. Computers. 2025;14:298. [Google Scholar]
  • 99.Tjeerd A.J.S., Wiard J., Mark A.N., Karel V.D.B. Human-centered XAI: Developing design patterns for explanations of clinical decision support systems. Int. J. Hum. Comput. Stud. 2021;154:102684. [Google Scholar]
  • 100.Anand A., Kadian T., Shetty M.K., Gupta A. Explainable AI decision model for ECG data of cardiac disorders. Biomed. Signal Process Control. 2022;75 [Google Scholar]
  • 101.Jin W., Fatehi M., Guo R., Hamarneh G. Evaluating the clinical utility of artificial intelligence assistance and its explanation on the glioma grading task. Artif. Intell. Med. 2024;148 doi: 10.1016/j.artmed.2023.102751. [DOI] [PubMed] [Google Scholar]
  • 102.Bashir Z., Lin M., Feragen A., Mikolaj K., Taksøe-Vester C., Christensen A.N., Svendsen M.B.S., Fabricius M.H., Andreasen L., Nielsen M., Tolsgaard M.G. Clinical validation of explainable AI for fetal growth scans through multi-level, cross-institutional prospective end-user evaluation. Sci. Rep. 2025;15:2074. doi: 10.1038/s41598-025-86536-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Sayres R., Taly A., Rahimy E., Blumer K., Coz D., Hammel N., Krause J., Narayanaswamy A., Rastegar Z., Wu D., et al. Using a deep learning algorithm and integrated gradients explanation to assist grading for diabetic retinopathy. Ophthalmology. 2019;126:552–564. doi: 10.1016/j.ophtha.2018.11.016. [DOI] [PubMed] [Google Scholar]
  • 104.Yang H.L., Kim J.J., Kim J.H., Kang Y.K., Park D.H., Park H.S., Kim H.K., Kim M.S. Weakly supervised lesion localization for age-related macular degeneration detection using optical coherence tomography images. PLoS One. 2019;14 doi: 10.1371/journal.pone.0215076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Papanastasopoulos Z., Samala R.K., Chan H.P., Hadjiiski L., Paramagul C., Helvie M.A., Neal C.H. Explainable AI for medical imaging: deep-learning CNN ensemble for classification of estrogen receptor status from breast MRI. Medical Imaging 2020: Computer-Aided Diagnosis. SPIE. 2020;11314:228–235. [Google Scholar]
  • 106.Sun J., Darbehani F., Zaidi M., Wang B. International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer International Publishing; Cham Switzerland: 2020. Saunet: Shape attentive u-net for interpretable medical image segmentation; pp. 797–806. [Google Scholar]
  • 107.Eslami T., Raiker J.S., Saeed F. In: Neural Engineering Techniques for Autism Spectrum Disorder. El-Baz A.S., Suri J.S., editors. Academic Press; Nevada, USA: 2021. Explainable and scalable machine learning algorithms for detection of autism spectrum disorder using fmri data; pp. 39–54. [Google Scholar]
  • 108.Pertzborn D., Arolt C., Ernst G., Lechtenfeld O.J., Kaesler J., Pelzel D., Guntinas-Lichius O., von Eggeling F., Hoffmann F. Multi-class cancer subtyping in salivary gland carcinomas with maldi imaging and deep learning. Cancers. 2022;14:4342. doi: 10.3390/cancers14174342. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Caroprese L., Vocaturo E., Zumpano E. Argumentation approaches for explanaible AI in medical informatics. Intell. Syst. Appl. 2022;16 [Google Scholar]
  • 110.Wickstrøm K., Kampffmeyer M., Jenssen R. Uncertainty and interpretability in convolutional neural networks for semantic segmentation of colorectal polyps. Med. Image Anal. 2020;60 doi: 10.1016/j.media.2019.101619. [DOI] [PubMed] [Google Scholar]
  • 111.Di Martino F., Delmastro F. Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artif. Intell. Rev. 2023;56:5261–5315. doi: 10.1007/s10462-022-10304-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Kerz E., Zanwar S., Qiao Y., Wiechmann D. Toward explainable AI (XAI) for mental health detection based on language behavior. Front. Psychiatry. 2023;14 doi: 10.3389/fpsyt.2023.1219479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.De Vries B.M., Zwezerijnen G.J.C., Burchell G.L., van Velden F.H.P., Menke-van der Houven van Oordt C.W., Boellaard R. Explainable artificial intelligence (XAI) in radiology and nuclear medicine: a literature review. Front. Med. 2023;10 doi: 10.3389/fmed.2023.1180773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Saporta A., Gui X., Agrawal A., Pareek A., Truong S.Q.H., Nguyen C.D.T., Ngo V.D., Seekins J., Blankenberg F.G., Ng A.Y., et al. Benchmarking saliency methods for chest X-ray interpretation. Nat. Mach. Intell. 2022;4:867–878. [Google Scholar]
  • 115.Singh A., Sengupta S., Lakshminarayanan V. Explainable deep learning models in medical image analysis. J. Imaging. 2020;6:52. doi: 10.3390/jimaging6060052. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Fuhrman J.D., Gorre N., Hu Q., Li H., El Naqa I., Giger M.L. A review of explainable and interpretable AI with applications in COVID-19 imaging. Med. Phys. 2022;49:1–14. doi: 10.1002/mp.15359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Karim M.R., Islam T., Shajalal M., Beyan O., Lange C., Cochez M., Rebholz-Schuhmann D., Decker S. Explainable AI for bioinformatics: methods, tools and applications. Brief. Bioinform. 2023;24 doi: 10.1093/bib/bbad236. [DOI] [PubMed] [Google Scholar]
  • 118.Srinivas S., Fleuret F. Advances in neural information processing systems. 2019. Full-gradient representation for neural network visualization; p. 32. [Google Scholar]
  • 119.Smilkov D., Thorat N., Kim B., Viégas F., Wattenberg M. Smoothgrad: Removing noise by adding noise. arXiv. 2017 doi: 10.48550/arXiv.1706.03825. Preprint at. [DOI] [Google Scholar]
  • 120.van Leersum C.M., Maathuis C. Human centred explainable AI decision-making in healthcare. J. Respon. Technol. 2025;21 [Google Scholar]
  • 121.Borys K., Schmitt Y.A., Nauta M., Seifert C., Krämer N., Friedrich C.M., Nensa F. Explainable AI in medical imaging: An overview for clinical practitioners–Beyond saliency-based XAI approaches. Eur. J. Radiol. 2023;162 doi: 10.1016/j.ejrad.2023.110786. [DOI] [PubMed] [Google Scholar]
  • 122.Yeh C.K., Hsieh C.Y., Suggala A., Inouye D.I., Ravikumar P.K. On the (in)fidelity and sensitivity of explanations. Adv. Neural Inf. Process. Syst. 2019;32:10967–10978. [Google Scholar]
  • 123.Ghnemat R., Alodibat S., Abu Al-Haija Q. Explainable artificial intelligence (XAI) for deep learning based medical imaging classification. J. Imaging. 2023;9:177. doi: 10.3390/jimaging9090177. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Molnar C., Casalicchio G., Bischl B. Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer International Publishing; Cham, Switzerland: 2020. Interpretable machine learning–a brief history, state-of-the-art and challenges; pp. 417–431. [Google Scholar]
  • 125.Liu T., Guo Q., Lian C., Ren X., Liang S., Yu J., Niu L., Sun W., Shen D. Automated detection and classification of thyroid nodules in ultrasound images using clinical-knowledge-guided convolutional neural networks. Med. Image Anal. 2019;58 doi: 10.1016/j.media.2019.101555. [DOI] [PubMed] [Google Scholar]
  • 126.Hoogestraat A.T., Wulff A. A Vision on user-centered implementation and evaluation of explainable AI for predicting hospital-onset bacteremia. Stud. Health Technol. Inform. 2024;316:766–770. doi: 10.3233/SHTI240525. [DOI] [PubMed] [Google Scholar]
  • 127.Alkhalaf S., Alturise F., Bahaddad A.A., Elnaim B.M.E., Shabana S., Abdel-Khalek S., Mansour R.F. Adaptive aquila optimizer with explainable artificial intelligence-enabled cancer diagnosis on medical imaging. Cancers. 2023;15:1492. doi: 10.3390/cancers15051492. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Jasper V.D.W., Elisabeth N., Anita C., Mark N. Evaluating XAI: A comparison of rule-based and example-based explanations. Artif. Intell. 2020;291:103404. [Google Scholar]
  • 129.IBM Research AI Fairness 360 (AIF360) [Computer software] 2025. https://ai-fairness-360.org/
  • 130.IBM Research AI Fairness 360 GitHub repository [Computer software] 2025. https://github.com/Trusted-AI/AIF360
  • 131.Voigt P., Von dem Bussche A. A Practical Guide. 1st edition. Springer International Publishing; Cham, Switzerland: 2017. The eu general data protection regulation (gdpr) [Google Scholar]
  • 132.Molnar C., Casalicchio G., Bischl B. Ulmer Informatik-Berichte; 2019. Quantifying interpretability of arbitrary machine learning models through functional decomposition; p. 41. [Google Scholar]
  • 133.Yu K.H., Beam A.L., Kohane I.S. Artificial intelligence in healthcare. Nat. Biomed. Eng. 2018;2:719–731. doi: 10.1038/s41551-018-0305-z. [DOI] [PubMed] [Google Scholar]
  • 134.Bienefeld N., Boss J.M., Lüthy R., Brodbeck D., Azzati J., Blaser M., Willms J., Keller E. Solving the explainable AI conundrum by bridging clinicians’ needs and developers’ goals. NPJ Digit. Med. 2023;6:94. doi: 10.1038/s41746-023-00837-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Skater Model Interpretation Library for Python [Computer software] 2025. https://github.com/GapData/skater
  • 136.ELI5: Explain Like I'm 5 [Computer software] 2025. https://eli5.readthedocs.io
  • 137.Microsoft Research InterpretML: A toolkit for training interpretable models [Computer software] 2025. https://interpret.ml/
  • 138.Captum Model interpretability for PyTorch [Computer software] 2025. https://captum.ai/
  • 139.Sicara tf-explain: Interpretability Methods for tf.keras models with TensorFlow 2.x [Computer software] 2025. https://github.com/sicara/tf-explain
  • 140.Alber M. iNNvestigate: A toolbox for analyzing neural networks [Computer software] 2025. https://github.com/albermax/innvestigate
  • 141.Google PAIR What-If Tool [Computer software] 2025. https://pair-code.github.io/what-if-tool/
  • 142.H2O.ai H2O: Open Source Machine Learning Platform [Computer software] 2025. https://h2o.ai/
  • 143.Hazan L. Z.M. Neurosuite: Klusters, Neuroscope and NDManager [Computer software] 2025. https://neurosuite.sourceforge.net/
  • 144.Abzu The QLattice: Explainable AI by Abzu [Computer software] 2025. https://www.abzu.ai/qlattice/
  • 145.Garralda E., Beaulieu M.E., Moreno V., Casacuberta-Serra S., Martínez-Martín S., Foradada L., Alonso G., Massó-Vallés D., López-Estévez S., Jauset T., et al. MYC targeting by OMO-103 in solid tumors: a phase 1 trial. Nat. Med. 2024;30:762–771. doi: 10.1038/s41591-024-02805-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Li K., Desai R., Scott R.T., Steele J.R., Machado M., Demharter S., Hoarfrost A., Braun J.L., Fajardo V.A., Sanders L.M., Costes S.V. Explainable machine learning identifies multi-omics signatures of muscle response to spaceflight in mice. NPJ Microgravity. 2023;9:90. doi: 10.1038/s41526-023-00337-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Samek W., Montavon G., Vedaldi A., Hansen L.K., Müller K.R. Springer Nature; 2019. Explainable AI: interpreting, explaining and visualizing deep learning. [Google Scholar]
  • 148.Preece A., Harborne D., Braines D., Tomsett R., Chakraborty S. Stakeholders in explainable AI. arXiv. 2018 doi: 10.48550/arXiv.1810.00184. Preprint at. [DOI] [Google Scholar]
  • 149.Vermeire T., Laugel T., Renard X., Martens D., Detyniecki M. Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer International Publishing; Cham, Switzerland: 2021. How to choose an explainability method? Towards a methodical implementation of XAI in practice; pp. 521–533. [Google Scholar]
  • 150.Amann J., Blasimme A., Vayena E., Frey D., Madai V.I., Precise4Q consortium Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med. Inform. Decis. Mak. 2020;20:310–319. doi: 10.1186/s12911-020-01332-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Joyce D.W., Kormilitzin A., Smith K.A., Cipriani A. Explainable artificial intelligence for mental health through transparency and interpretability for understandability. NPJ Digit. Med. 2023;6:6. doi: 10.1038/s41746-023-00751-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Klauschen F., Dippel J., Keyl P., Jurmeister P., Bockmayr M., Mock A., Buchstab O., Alber M., Ruff L., Montavon G., Müller K.R. Toward explainable artificial intelligence for precision pathology. Annu. Rev. Pathol. 2024;19:541–570. doi: 10.1146/annurev-pathmechdis-051222-113147. [DOI] [PubMed] [Google Scholar]
  • 153.Nagendran M., Festor P., Komorowski M., Gordon A.C., Faisal A.A. Quantifying the impact of AI recommendations with explanations on prescription decision making. NPJ Digit. Med. 2023;6:206. doi: 10.1038/s41746-023-00955-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.Nguyen D.Q., Vo N.Q., Nguyen T.T., Nguyen-An K., Nguyen Q.H., Tran D.N., Quan T.T. BeCaked: an explainable artificial intelligence model for COVID-19 forecasting. Sci. Rep. 2022;12:7969. doi: 10.1038/s41598-022-11693-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Lipkova J., Chen R.J., Chen B., Lu M.Y., Barbieri M., Shao D., Vaidya A.J., Chen C., Zhuang L., Williamson D.F.K., et al. Artificial intelligence for multimodal data integration in oncology. Cancer Cell. 2022;40:1095–1110. doi: 10.1016/j.ccell.2022.09.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.Neri E., Aghakhanyan G., Zerunian M., Gandolfo N., Grassi R., Miele V., Giovagnoni A., Laghi A., SIRM expert group on Artificial Intelligence Explainable AI in radiology: a white paper of the Italian Society of Medical and Interventional Radiology. Radiol. Med. 2023;128:755–764. doi: 10.1007/s11547-023-01634-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157.Langer M., Oster D., Speith T., Hermanns H., Kästner L., Schmidt E., Sesing A., Baum K. What do we want from Explainable Artificial Intelligence (XAI)?–A stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI research. Artif. Intell. 2021;296 [Google Scholar]
  • 158.Adams R., Henry K.E., Sridharan A., Soleimani H., Zhan A., Rawat N., Johnson L., Hager D.N., Cosgrove S.E., Markowski A., et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat. Med. 2022;28:1455–1460. doi: 10.1038/s41591-022-01894-0. [DOI] [PubMed] [Google Scholar]
  • 159.Henry K.E., Hager D.N., Pronovost P.J., Saria S. A targeted real-time early warning score (TREWScore) for septic shock. Sci. Transl. Med. 2015;7 doi: 10.1126/scitranslmed.aab3719. [DOI] [PubMed] [Google Scholar]
  • 160.Henry K.E., Adams R., Parent C., Soleimani H., Sridharan A., Johnson L., Hager D.N., Cosgrove S.E., Markowski A., Klein E.Y., et al. Factors driving provider adoption of the TREWS machine learning-based early warning system and its effects on sepsis treatment timing. Nat. Med. 2022;28:1447–1454. doi: 10.1038/s41591-022-01895-z. [DOI] [PubMed] [Google Scholar]
  • 161.Jin W., Li X., Fatehi M., Hamarneh G. Guidelines and evaluation of clinical explainable AI in medical image analysis. Med. Image Anal. 2023;84 doi: 10.1016/j.media.2022.102684. [DOI] [PubMed] [Google Scholar]
  • 162.Lakshmi K., Amaran S., Subbulakshmi G., Padmini S., Joshi G.P., Cho W. Explainable artificial intelligence with UNet based segmentation and Bayesian machine learning for classification of brain tumors using MRI images. Sci. Rep. 2025;15:690. doi: 10.1038/s41598-024-84692-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 163.Das N., Happaerts S., Gyselinck I., Staes M., Derom E., Brusselle G., Burgos F., Contoli M., Dinh-Xuan A.T., Franssen F.M.E., et al. Collaboration between explainable artificial intelligence and pulmonologists improves the accuracy of pulmonary function test interpretation. Eur. Respir. J. 2023;61 doi: 10.1183/13993003.01720-2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 164.Das N., Happaerts S., Dinh-Xuan A.T., Vanderelst E., Haenebalcke C., Brusselle G., Derom E., Burgos F., Franssen F., Watz H., et al. Pulmonologists collaborate with explainable artificial intelligence for superior interpretation of pulmonary function tests. Allergy. 2022 [Google Scholar]
  • 165.Wang Q., Huang K., Chandak P., Zitnik M., Gehlenborg N. Extending the nested model for user-centric xai: A design study on gnn-based drug repurposing. IEEE Trans. Vis. Comput. Graph. 2023;29:1266–1276. doi: 10.1109/TVCG.2022.3209435. [DOI] [PubMed] [Google Scholar]
  • 166.Farahani F.V., Fiok K., Lahijanian B., Karwowski W., Douglas P.K. Explainable AI: A review of applications to neuroimaging data. Front. Neurosci. 2022;16 doi: 10.3389/fnins.2022.906290. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 167.Yang G., Ye Q., Xia J. Unbox the black-box for the medical explainable AI via multi-modal and multi-centre data fusion: A mini-review, two showcases and beyond. Inf. Fusion. 2022;77:29–52. doi: 10.1016/j.inffus.2021.07.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 168.Papandrianos N.I., Feleki A., Moustakidis S., Papageorgiou E.I., Apostolopoulos I.D., Apostolopoulos D.J. An explainable classification method of SPECT myocardial perfusion images in nuclear cardiology using deep learning and grad-CAM. Appl. Sci. 2022;12:7592. [Google Scholar]
  • 169.Katja H., Alexander K., Sarah H., Roman C.M., Christof V.K., Jochen S.U., Friedegund M., Sarah H., Frank F.G., Mildred S., et al. Explainable artificial intelligence in skin cancer recognition: A systematic review. Eur. J. Cancer. 2022;167:54–69. doi: 10.1016/j.ejca.2022.02.025. [DOI] [PubMed] [Google Scholar]
  • 170.Jungyo S., Dongwon K., Bumjin L., Cheryn S., Dalsan Y., In G.J., Jun H.H., Bumsik H., Hanjong A., Choung-Soo K., et al. MP47-01 Development of individual level inference available explainable artificial intelligence model for cancer-specific survival after nephrectomy in renal cell carcinoma patients. J. Urol. 2022;207 [Google Scholar]
  • 171.Prinzi F., Currieri T., Gaglio S., Vitabile S. Shallow and deep learning classifiers in medical image analysis. Eur. Radiol. Exp. 2024;8:26. doi: 10.1186/s41747-024-00428-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 172.Hou J., Gao T. Explainable DCNN based chest X-ray image analysis and classification for COVID-19 pneumonia detection. Sci. Rep. 2021;11 doi: 10.1038/s41598-021-95680-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 173.Giuste F., Shi W., Zhu Y., Naren T., Isgut M., Sha Y., Tong L., Gupte M., Wang M.D. Explainable artificial intelligence methods in combating pandemics: A systematic review. IEEE Rev. Biomed. Eng. 2023;16:5–21. doi: 10.1109/RBME.2022.3185953. [DOI] [PubMed] [Google Scholar]
  • 174.Nambiar A., Harikrishnaa S., Sharanprasath S. Model-agnostic explainable artificial intelligence tools for severity prediction and symptom analysis on Indian COVID-19 data. Front. Artif. Intell. 2023;6 doi: 10.3389/frai.2023.1272506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 175.Malinverno L., Barros V., Ghisoni F., Visonà G., Kern R., Nickel P.J., Ventura B.E., Šimić I., Stryeck S., Manni F., et al. A historical perspective of biomedical explainable AI research. Patterns. 2023;4 doi: 10.1016/j.patter.2023.100830. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 176.Gaube S., Suresh H., Raue M., Lermer E., Koch T.K., Hudecek M.F.C., Ackery A.D., Grover S.C., Coughlin J.F., Frey D., et al. Non-task expert physicians benefit from correct explainable AI advice when reviewing X-rays. Sci. Rep. 2023;13:1383. doi: 10.1038/s41598-023-28633-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 177.Antoniadi A.M., Du Y., Guendouz Y., Wei L., Mazo C., Becker B.A., Mooney C. Current challenges and future opportunities for XAI in machine learning-based clinical decision support systems: a systematic review. Appl. Sci. 2021;11:5088. [Google Scholar]
  • 178.Hatwell J., Gaber M.M., Atif Azad R.M. Ada-WHIPS: explaining AdaBoost classification with applications in the health sciences. BMC Med. Inform. Decis. Mak. 2020;20:250. doi: 10.1186/s12911-020-01201-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 179.Van Molle P., De Strooper M., Verbelen T., Vankeirsbilck B., Simoens P., Dhoedt B. Understanding and Interpreting Machine Learning in Medical Image Computing Applications: First International Workshops, MLCN 2018, DLF 2018, and iMIMIC 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16-20, 2018, Proceedings 1. Springer International Publishing; 2018. Visualizing convolutional neural networks to improve decision support for skin lesion classification; pp. 115–123. [Google Scholar]
  • 180.Salih A., Boscolo Galazzo I., Gkontra P., Lee A.M., Lekadir K., Raisi-Estabragh Z., Petersen S.E. Explainable artificial intelligence and cardiac imaging: toward more interpretable models. Circ. Cardiovasc. Imaging. 2023;16 doi: 10.1161/CIRCIMAGING.122.014519. [DOI] [PubMed] [Google Scholar]
  • 181.Moreno-Sánchez P.A. Improvement of a prediction model for heart failure survival through explainable artificial intelligence. Front. Cardiovasc. Med. 2023;10 doi: 10.3389/fcvm.2023.1219586. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 182.Moulaei K., Afshari L., Moulaei R., Sabet B., Mousavi S.M., Afrash M.R. Explainable artificial intelligence for stroke prediction through comparison of deep learning and machine learning models. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-82931-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 183.Noemi G., Lorenzo M., Fabio M., Alessandra P. XAI for Myo-Controlled Prosthesis: Explaining EMG Data for Hand Gesture Classification. Knowl. Base Syst. 2022;240:108053. [Google Scholar]
  • 184.Vo T.H., Nguyen N.T.K., Kha Q.H., Le N.Q.K. On the road to explainable AI in drug-drug interactions prediction: A systematic review. Comput. Struct. Biotechnol. J. 2022;20:2112–2123. doi: 10.1016/j.csbj.2022.04.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 185.Slack D., Krishna S., Lakkaraju H., Singh S. Explaining machine learning models with interactive natural language conversations using TalkToModel. Nat. Mach. Intell. 2023;5:873–883. [Google Scholar]
  • 186.Nagendran M., Festor P., Komorowski M., Gordon A.C., Faisal A.A. Eye tracking insights into physician behaviour with safe and unsafe explainable AI recommendations. npj Digit. Med. 2024;7:202. doi: 10.1038/s41746-024-01200-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 187.Cálem J., Moreira C., Jorge J. Intelligent systems in healthcare: A systematic survey of explainable user interfaces. Comput. Biol. Med. 2024;180 doi: 10.1016/j.compbiomed.2024.108908. [DOI] [PubMed] [Google Scholar]
  • 188.Nawaz U., Anees-ur-Rahaman M., Saeed Z. A review of neuro-symbolic AI integrating reasoning and learning for advanced cognitive systems. Intell. Syst. Appl. 2025;26 [Google Scholar]
  • 189.Koh P.W., Nguyen T., Tang Y.S., Mussmann S., Pierson E., Kim B., Liang P. PMLR; 2020. Concept Bottleneck Models. International Conference on Machine Learning; pp. 5338–5348. [Google Scholar]
  • 190.Mazumdar H., Khondakar K.R., Das S., Halder A., Kaushik A. Artificial intelligence for personalized nanomedicine; from material selection to patient outcomes. Expert Opin. Drug Deliv. 2025;22:85–108. doi: 10.1080/17425247.2024.2440618. [DOI] [PubMed] [Google Scholar]
  • 191.Njei B., Dranoff J., Lim J. S1366 explainable artificial intelligence accurately predicts high-risk nonalcoholic steatohepatitis and identifies new subphenotypes in lean individuals. Am. J. Gastroenterol. 2023;118:S1047–S1048. [Google Scholar]

Articles from iScience are provided here courtesy of Elsevier

RESOURCES