Abstract
Artificial intelligence (AI) is increasingly used in drug repurposing to integrate chemical, biological, and clinical data and to prioritize candidate drug–disease associations. Yet current evaluation practices remain dominated by benchmark metrics such as AUROC, AUPRC, F1, or top- retrieval, which do not by themselves establish pharmacological credibility. A useful prediction is not merely one that scores highly on retrospective datasets, but one that can be connected to a biologically plausible mechanism, safety-relevant context, and a traceable evidentiary basis for experimental follow-up. Mechanistic explainability should therefore be treated as a core objective for translationally oriented AI-enabled repurposing, while also emphasizing that explainability complements rather than replaces experimental validation. Specifically, explanations should connect drugs, targets, pathways, phenotypes, and clinical outcomes in ways that are intelligible to pharmacologists and compatible with experimental prioritization, translational decision-making, and emerging regulatory expectations. We identify a central gap in the literature: many explainability methods emphasize feature attribution or local model transparency, but rarely produce pharmacology-aligned evidence chains that also communicate uncertainty, robustness, and provenance. We introduce biomedical knowledge graphs and neurosymbolic approaches as a promising foundation for more credible repurposing systems because they can support relational inference, structured mechanistic reasoning, and provenance-aware explanation. We further argue that future evaluation should extend beyond predictive accuracy to include mechanistic coherence, uncertainty, robustness, and provenance (MURP) for experimental pharmacology, ideally within workflows that integrate computational prediction with laboratory and, where feasible, real-world validation. Reframing success in this way would improve rigor, translational relevance, and regulatory readiness in AI-driven drug repurposing.
Keywords: biomedical knowledge graphs, drug repurposing, explainable artificial intelligence (XAI), graph neural networks (GNNs), mechanistic interpretability, translational pharmacology, uncertainty quantification
1. Introduction
Artificial intelligence (AI) has increasingly become a catalyst for accelerating drug repurposing by integrating vast chemical, biological, pharmacological, and clinical data. Models that leverage these heterogeneous sources can rapidly propose drug–disease hypotheses and identify mechanistic leads that traditionally require substantial time and human labor (Pushpakom et al., 2019; Zhang et al., 2025). Yet, despite this promise, the field remains dominated by an “accuracy-first” mindset in which performance metrics such as area under the receiver operating characteristic (AUROC), area under the precision–recall curve (AUPRC), or top- retrieval accuracy are treated as primary indicators of scientific value. These metrics are useful for benchmarking algorithmic progress, but they do not equate to pharmacological credibility. A link predicted with high probability on a retrospective dataset is not necessarily biologically meaningful, clinically actionable, or safe to advance.
This misalignment has become increasingly salient as regulatory bodies sharpen their expectations for transparency and trustworthy AI. In January 2026, the European Medicines Agency (EMA) and United States Food and Drug Administration (FDA) introduced joint Guiding Principles of Good AI Practice in Drug Development 1 , emphasizing transparency, risk-based lifecycle management, clear communication of model capabilities and limitations, and the need for traceable evidence to support AI-informed decision-making. These principles are not specific to drug repurposing; rather, they address AI across the broader medicines lifecycle. Nevertheless, they are directly relevant to repurposing contexts in which biological plausibility, safety, traceability, and human oversight are essential. The EMA–FDA principles explicitly call for clear and relevant information about model behavior, uncertainty, and limitations; repurposing predictions that lack mechanistic justification fail this requirement by design.
Mechanistic explainability in AI-enabled drug repurposing increasingly functions as a form of scientific and regulatory infrastructure rather than a purely interpretive aid. The EMA–FDA principles emphasize lifecycle governance, including transparency, traceability, change management, and post-deployment monitoring of model behavior. Under this framing, explainability is important not only for scientific understanding but also for enabling auditability and defensible decision-making as models evolve over time. For repurposing systems, this implies that predictions should be accompanied by mechanistic rationales that remain intelligible when models, data sources, or knowledge graph (KG) snapshots are updated. Structured, provenance-aware evidence chains directly support these expectations by making explicit the biological assumptions underlying a prediction. They also enable reviewers to assess how changes in data, model parameters, or graph structure affect mechanistic conclusions.
In practical drug repurposing workflows, these requirements translate into a need for predictions that clarify why a drug might modulate disease biology, through what mechanism, in which tissues or cell types, and with what anticipated liabilities. In the absence of this mechanistic visibility, even highly ranked suggestions pose challenges for experimental prioritization, trial design, and regulatory engagement.
We therefore propose a pharmacology-centered standard. First, current explainability approaches are often misaligned with the epistemic needs of pharmacology because they rarely produce biologically intelligible evidence chains. Second, KGs, graph-based learning, and neurosymbolic approaches provide a promising substrate for pharmacology-aligned explanation. Third, future evaluation should treat mechanistic coherence, uncertainty, robustness, provenance, and experimental actionability as central dimensions of model quality. We collectively refer to these dimensions as the MURP framework (mechanistic coherence, uncertainty, robustness, and provenance). As a Perspective article, this paper does not present a new benchmark, dataset, or executable pipeline, nor does it follow a systematic review methodology. Instead, it synthesizes representative methodological and translational examples from the literature to argue for a more rigorous standard of evaluation in AI-enabled drug repurposing.
2. Why benchmark accuracy is not enough for pharmacology
Although statistical performance metrics remain necessary, they are insufficient for judging whether an AI-generated repurposing hypothesis is scientifically viable. A top- hit in a retrospective benchmark does not guarantee that the underlying biological rationale is sound. Many models learn shortcuts tied to network topology, database incompleteness, label leakage, or biased edge distributions (Kondratyeva et al., 2022; Singh et al., 2023; Brière et al., 2025). Without insight into the mechanistic chain linking a drug to a disease, researchers must rely on manual curation, which is slow, subjective, and prone to confirmation bias. Moreover, unexplained predictions impede translation: pharmacologists, clinicians, and regulatory reviewers require clarity on the molecular and physiological basis of proposed interventions.
Two factors amplify the need for mechanistic explainability. First, repurposing inherently draws on prior clinical knowledge about safety, off-target effects, pharmacokinetics, and contraindications. A candidate drug’s historical profile can either support or weaken confidence in a new indication, but only if the proposed rationale can be assessed against known biology. Second, clinical and translational decision-making is risk-sensitive. Understanding potential failure points or safety concerns early allows improved experimental design and more disciplined resource allocation. For these reasons, benchmark metrics should be viewed as one component of evaluation rather than a surrogate for pharmacological value.
It is important, however, to distinguish several related concepts that are often conflated in the literature. Interpretability typically refers to the extent to which a model’s internal workings are understandable to humans. Explainability refers more broadly to the extent to which a model’s outputs can be made understandable, whether intrinsically or by post-hoc methods. Mechanistic plausibility is a stronger requirement: it asks whether a computational explanation is consistent with known biological directionality, tissue or cell-type context, and pharmacological logic. Experimental validation is stronger still, requiring empirical confirmation through wet-lab studies, clinical data, or real-world evidence.
The central problem is therefore not simply that many current explainability methods are imperfect, but that they are often poorly matched to the epistemic needs of pharmacology. Generic post-hoc tools such as SHAP (Lundberg and Lee, 2017) or LIME (Ribeiro et al., 2016) may identify influential features, yet they rarely explain whether a predicted drug–disease association is consistent with known mechanism of action, pathway biology, tissue context, adverse-effect liabilities, or contraindication patterns. In drug repurposing, however, these are precisely the questions that determine whether a computational signal becomes a credible pharmacological hypothesis. A useful explanation should therefore resemble an evidence chain rather than a feature ranking: it should connect drug, target, pathway, phenotype, and clinical consequence in a traceable and biologically intelligible way, while also communicating uncertainty and provenance.
BOX 1. Definition: pharmacological evidence chain.
We define a pharmacological evidence chain as an ordered, semantically constrained sequence of biomedical relations that provides a mechanistic rationale for a drug–disease association. In its canonical form, an evidence chain links:
Each edge in the chain may be annotated with uncertainty estimates and provenance metadata, and extended with safety-relevant relations such as adverse drug reactions, drug–drug interactions, or contraindications. Evidence chains thus encode not only connectivity, but structured, inspectable mechanistic reasoning suitable for pharmacological interpretation and experimental prioritization.
In practice, evidence chains can be represented as weighted paths or explanatory subgraphs, where each relation is associated with a confidence score reflecting uncertainty, robustness, or evidentiary strength. Canonical chains will often span four or five biologically meaningful hops, although shorter or longer chains may be acceptable when justified by the disease context. What matters is not path length alone, but whether the chain respects biological directionality, ontology-consistent relation types, disease-relevant tissue context, and the balance of supporting versus conflicting evidence.
3. Knowledge graphs as foundations for evidence chains
Biomedical knowledge graphs provide a natural substrate for mechanistic reasoning. By organizing entities such as drugs, targets, genes, pathways, diseases, phenotypes, adverse events, and clinical observations into a semantically governed network, they enable multi-layered inference and transparent provenance tracking. Figure 1 illustrates how reasoning over an explanatory knowledge graph can be conceptualized. In the life sciences, Hetionet established one of the first comprehensive biomedical KGs designed for drug repurposing, constructed from 29 public databases and providing over 47,000 nodes and 2.25 million relationships (Himmelstein et al., 2017). Using Hetionet, Himmelstein et al., 2017 demonstrated that path-based features in such a heterogeneous network could prioritize drug–disease pairs while simultaneously producing interpretable metapaths that aligned with known biology.
FIGURE 1.
Illustrative biomedical knowledge graph for explainability-centered drug repurposing. The highlighted path shows how a candidate drug–disease association can be justified through a mechanistic evidence chain linking drug, target, pathway, phenotype, and disease, while additional nodes encode safety-relevant information such as adverse drug reactions and drug–drug interactions. This representation exemplifies how knowledge graphs support reasoning that is more transparent and pharmacologically interpretable than a prediction score alone.
Subsequent KGs expanded scale and semantic depth. SPOKE integrates data from more than 40 sources, spanning molecular, physiological, and clinical domains, resulting in over 27 million nodes and 53 million edges governed by 11 ontologies (Morris et al., 2023). Its molecular-to-clinical coverage facilitates translational inference at a scale unattainable by manual curation.
PharMeBINet offers further expansion by integrating extensive adverse drug reaction (ADR), drug–drug interaction (DDI), and genomic variant information into a multi-million-node Neo4j graph, making it particularly valuable for safety-aware reasoning in repurposing workflows (Königs et al., 2022).
Beyond their structural and semantic properties, biomedical networks have already demonstrated translational impact. During the COVID-19 pandemic, network proximity analysis identified 16 repurposable candidates and three drug combinations for SARS-CoV-2, several of which were subsequently supported by experimental or clinical evidence (Zhou et al., 2020). Notably, that approach relied on topological proximity between drug targets and disease-associated proteins rather than on explicit mechanistic paths. This distinction underscores the opportunity that path-based and evidence-chain approaches offer: not only identifying that a drug’s targets are close to a disease module, but articulating the mechanistic rationale and safety context underlying the association through multi-hop biological relationships that can be interpreted as mechanistic hypotheses.
Recent path-based repurposing methods already demonstrate that such mechanistic explanations can be operationalized in practice: explanatory rationales can be expressed as biologically interpretable paths linking compounds to diseases through intermediate genes, pathways, phenotypes, and adverse effects (Jiménez et al., 2024; Nunes et al., 2025). We therefore do not claim novelty for the mere idea of path-based explanation. Rather, our argument is that such explanations should become a primary evaluation target and should be judged not only by path plausibility, but also by uncertainty, robustness, provenance, and usefulness for downstream validation. The significance is broader than any single method: these approaches demonstrate that explainability in repurposing need not be limited to generic feature attribution, but can instead be operationalized as structured evidence chains that domain experts can inspect, challenge, and experimentally prioritize.
4. From neural prediction to pharmacological reasoning
Knowledge graphs provide structured biomedical context, whereas neural models provide the capacity to learn complex relational patterns from incomplete and heterogeneous data. Graph neural networks (GNNs) are therefore attractive for drug repurposing (Dai et al., 2024), but their internal reasoning remains difficult to interpret in pharmacologically meaningful terms. Model-aligned explainers such as GNNExplainer (Ying et al., 2019) and PGExplainer (Luo et al., 2020) identify prediction-relevant subgraphs, while SubgraphX (Yuan et al., 2021) explains GNN outputs by explicitly searching for important substructures using Monte Carlo tree search and Shapley-value-based scoring. Yet influence alone is not equivalent to mechanism: an influential subgraph may help explain a prediction computationally while still failing basic pharmacological expectations regarding causal direction, tissue relevance, safety, or known mechanism of action.
Beyond such neural explainers, several approaches have demonstrated that relational reasoning over knowledge graphs can produce explicitly interpretable paths or rules. Rule learning systems such as AnyBURL (Meilicke et al., 2019) derive multi-hop logical Horn rules in a bottom-up manner and expose reusable reasoning patterns underlying predictions, while extensions such as SAFRAN (Ott et al., 2021) improve their applicability by identifying redundant rules and aggregating them via non-redundant probabilistic schemes. Path-oriented approaches take a complementary perspective: methods such as MINERVA (Das et al., 2018) formulate reasoning as a reinforcement learning problem over the knowledge graph, learning to traverse multi-hop paths conditioned on a query relation, while optimization-based frameworks such as PoLo (Liu et al., 2021) explicitly search for high-quality explanatory paths connecting drugs, targets, and diseases. These approaches collectively demonstrate that explainability in relational models can be operationalized as structured multi-hop reasoning rather than simple feature attribution. However, they primarily optimize for predictive relevance, path quality, or rule aggregation and therefore address only parts of pharmacologically meaningful explainability; in the terminology of this Perspective, they do not yet jointly satisfy the full MURP requirements of mechanistic coherence, uncertainty, robustness, and provenance, which are outlined in Section 5.
Although these methods improve transparency, they do not by themselves guarantee pharmacological coherence or mechanistic plausibility. Neurosymbolic AI could bridge that gap. Neurosymbolic AI refers to approaches that integrate neural learning from large, heterogeneous datasets with symbolic knowledge representations (such as ontologies, logical rules, or curated biological constraints) in a unified framework where symbolic knowledge actively constrains, guides, or explains model predictions (Garcez and Lamb, 2023; Bhuyan et al., 2024). In this sense, neurosymbolic systems differ from purely neural or purely rule- or path-based approaches by coupling statistical learning with explicit, interpretable reasoning constraints. Recent work such as MARS (DeLong et al., 2024) illustrates both the promise and the current limitations of this paradigm: by combining learned representations with logical rule-based reasoning over knowledge graphs, it produces more interpretable mechanism-of-action predictions, but also demonstrates that even neurosymbolic systems can remain susceptible to topological shortcuts such as degree bias if biological constraints are not carefully enforced (DeLong et al., 2025). For drug repurposing, this integration is particularly attractive. Neural explainers and rule- or path-based approaches can identify which subgraphs or reasoning chains influenced a prediction, but they do not by themselves ensure that the resulting rationale satisfies all four MURP dimensions. Figure 2 demonstrates this neurosymbolic AI pipeline.
FIGURE 2.

Neurosymbolic pipeline for explainability-centered drug repurposing. Structured biomedical knowledge graphs provide the biological and safety context for predictive learning. Neural models generate candidate drug–disease associations, while explanatory subgraphs articulate mechanistic rationales. Symbolic pharmacological constraints enforce biological plausibility and safety relevance, yielding evidence chains that are uncertainty-aware, provenance-rich, and suitable for experimental prioritization.
Neurosymbolic models can address this gap by constraining or re-ranking candidate explanations according to pharmacological knowledge, such as pathway directionality, tissue-specific target expression, contraindication patterns, or ontology-consistent relations (DeLong et al., 2024; Hossain and Chen, 2025). For example, a GNN may propose a drug–disease association on the basis of multi-hop connectivity, while a symbolic layer retains only those explanatory paths that remain consistent with known mechanism-of-action and safety knowledge. In this way, neurosymbolic AI offers not only greater transparency, but also a stronger bridge between statistical prediction and pharmacological reasoning (Garcez and Lamb, 2023; Drancé, 2022).
Yet mechanistic plausibility alone is insufficient: if explanations are unstable, overconfident, or weakly supported, they cannot reliably guide experimental follow-up. This motivates the evaluation framework introduced in the next section.
5. Mechanistic coherence, uncertainty, robustness, and provenance as essential components
BOX 2. Definition: Pharmacologically Meaningful Explainability.
We define explainability in AI-driven drug repurposing as a multi-dimensional assessment profile:
where denotes a candidate repurposing hypothesis and:
M (Mechanistic coherence): the presence of an explicit, biologically plausible evidence chain linking a drug to a clinical outcome (e.g., drug target pathway phenotype outcome).
U (Uncertainty): quantified epistemic and data-driven uncertainty associated with both predictions and explanatory links.
R (Robustness): stability of predictions and explanations under perturbations of model parameters, graph structure, or data provenance.
P (Provenance): traceability of each explanatory element to its underlying data sources, curation level, and evidentiary support.
Each dimension is assessed qualitatively or semi-quantitatively in a domain-appropriate manner. The tuple representation is deliberate: we do not propose collapsing these dimensions into a single scalar score, because their relative importance varies by context of use (e.g., early exploration vs. regulatory-facing decision-making). An explanation that is deficient in any single dimension should be treated with caution, regardless of predictive performance. For brevity, we refer to these four dimensions collectively as MURP.
For AI-based drug repurposing, explainability is only useful if the proposed rationale is accompanied by an assessment of how reliable, stable, and traceable that rationale actually is (Perdomo-Quinteiro and Belmonte-Hernández, 2024). A mechanistic explanation may appear biologically plausible, yet still rest on weak evidence, unstable graph structure, or overconfident predictions. In this sense, these properties are not secondary technical considerations, but essential conditions for deciding whether a computational hypothesis is sufficiently trustworthy to justify experimental follow-up.
To make this framework more operational, each component should be assessed in practice. Mechanistic coherence can be evaluated by checking whether an extracted chain respects ontology-consistent relation types, biological directionality, and tissue or cell-type relevance, and whether domain experts judge the chain to be pharmacologically plausible. Uncertainty should be reported at both the prediction level and the explanatory-edge level, so that users can distinguish broadly confident hypotheses from those dependent on fragile links. Robustness should be tested by graph perturbation, ablation of low-provenance edges, alternative graph snapshots, or variation across model seeds, asking whether the same core rationale persists. Provenance should identify the source database, evidence type, curation status, and supporting literature for each explanatory edge.
Uncertainty estimation is particularly important because high predictive scores can mask substantial epistemic or data-driven uncertainty (Guo et al., 2017). In drug repurposing, this distinction matters directly for prioritization: a candidate supported by a coherent mechanistic explanation but associated with high uncertainty should be interpreted differently from one supported by similarly plausible evidence with stronger calibration and greater stability. Methods such as deep ensembles have shown strong performance in producing better-calibrated uncertainty estimates and in detecting distributional shift (Lakshminarayanan et al., 2017), while Monte Carlo dropout offers a practical approximation of model uncertainty within standard neural architectures (Gal and Ghahramani, 2016). Rather than treating uncertainty as an abstract property of the model, repurposing workflows should expose it at the level of explanatory edges and paths, thereby allowing researchers to distinguish relatively robust evidence chains from predictions that remain mechanistically fragile.
Robustness is equally important because even an apparently interpretable explanation may depend on unstable graph topology, low-quality relations, or isolated shortcuts learned by the model. This is especially relevant in biomedical KGs, where the completeness and quality of edges vary substantially across domains, diseases, and data sources. Robustness analysis should therefore test whether an explanation remains stable when low-provenance relations are removed, when subgraphs are perturbed, or when alternative KG snapshots are used (Ying et al., 2019; Luo et al., 2020; Agarwal et al., 2023). From a pharmacological perspective, this matters because a rationale that collapses under minor perturbation is unlikely to provide a reliable basis for mechanistic interpretation or experimental design.
Provenance adds a further layer of credibility by clarifying where each element of an explanation originates and how much evidentiary weight it should carry. In biomedical KGs, explanatory paths may combine manually curated relations, ontology-grounded mappings, inferred links, and literature-derived associations, all of which differ in reliability and interpretability (Königs et al., 2022; Himmelstein et al., 2017). Provenance tracking is therefore not only a matter of technical reproducibility, but also of scientific judgment: it allows researchers to distinguish well-supported mechanistic links from weaker associations that require additional scrutiny. Conflicting biological evidence should likewise be made explicit rather than hidden. When supporting and contradicting chains coexist, the prediction should be flagged for expert adjudication instead of being presented as a singular mechanistic truth.
Taken together, these four dimensions determine whether an explanation is merely computationally plausible or credible enough to support pharmacological follow-up and should accordingly be treated as integral to model evaluation rather than as optional post-hoc diagnostics.
6. Toward a pharmacology-centered standard for AI repurposing
If explainability is to function as a criterion of pharmacological credibility rather than a post-hoc reporting accessory, the field needs a clearer standard for model development, evaluation, and communication. We argue that AI-based repurposing systems should be designed to produce, alongside each predicted drug–disease association, a MURP-compliant evidence chain: a mechanistically interpretable path annotated with uncertainty, robustness, and provenance information. Under such a standard, explainability is not a supplementary visualization layer but part of what constitutes a usable repurposing hypothesis.
The starting point is a carefully selected and versioned KG substrate, chosen for its coverage of drugs, targets, pathways, phenotypes, clinical signals, and safety-relevant relations. Heterogeneous GNNs or hybrid neurosymbolic models then learn relational patterns for link prediction while being instrumented for explanation extraction. Explanatory subgraphs are produced via GNNExplainer or PGExplainer and subsequently refined through symbolic constraints and uncertainty estimation. The final product for each predicted drug–disease pair is a structured evidence chain: a serialized, human-interpretable subgraph that highlights mechanistic steps, indicates provenance of each edge, communicates uncertainty, and incorporates safety considerations such as adverse drug reactions, drug–drug interactions, and contraindications – a capability especially well supported by ADR-enriched KGs such as PharMeBINet (Königs et al., 2022).
Evaluation should therefore combine conventional predictive metrics with the MURP dimensions introduced above, plus explicit assessment of safety relevance and experimental actionability. A model that performs well statistically but fails these tests may remain computationally impressive while being weak as a pharmacological decision-support system.
The importance of these criteria also varies by context of use. In early exploratory prioritization, a statistical signal may still be useful for hypothesis generation even before a detailed mechanism is available. Mechanistic explainability becomes more important when predictions are used to prioritize costly experiments, assess safety risks, support translational decisions, or inform regulatory-facing deliberation. In other words, the stronger the downstream decision consequences, the stronger the requirement for structured, uncertainty-aware, provenance-rich explanation.
This integrated approach not only enables rigorous scientific assessment but also supports a more collaborative, human-in-the-loop research process. Pharmacologists and domain experts can review, prune, or augment evidence chains; their feedback can then be incorporated as constraints, ranking signals, or adjudication criteria, helping align model outputs with the forms of reasoning that experimental pharmacology actually requires. In practical terms, such a framework would allow repurposing predictions to be prioritized not only by score, but by the quality of their mechanistic rationale, safety context, evidentiary robustness, and readiness for downstream validation.
7. Illustrative vignettes
To clarify how the proposed framework relates to practice, we distinguish between literature-based examples and hypothetical scenarios.
A useful counterexample is the discovery of halicin by deep learning (Stokes et al., 2020). In that case, a model identified a structurally novel antibiotic candidate with limited mechanistic explainability at the time of prediction, yet the finding gained strong credibility through extensive experimental validation, including activity against multiple pathogens and efficacy in murine infection models. This case is important because it shows that mechanistic explainability is not the only route to a valuable discovery in early-stage exploration. At the same time, subsequent work has shown that explainable deep learning can help identify structural rationales underlying antibacterial activity, suggesting that explainability and experimental validation are complementary rather than competing goals (Wong et al., 2024).
A second literature-based example is provided by recent network-medicine work in Alzheimer’s disease (Bykova et al., 2026). There, network-based prediction identified doxycycline and irbesartan as repurposing candidates, and large-scale real-world patient data were then used to assess whether prescription exposure was associated with reduced disease incidence. This paradigm is instructive because it combines computational prioritization with translational corroboration rather than treating either one as sufficient on its own. It therefore exemplifies the broader argument advanced here: pharmacological credibility is strongest when mechanistic reasoning, data integration, and empirical follow-up reinforce one another.
The following three scenarios are illustrative rather than empirical findings, and they are intended only to show how explainability-centered workflows can alter decision-making.
In the first hypothetical scenario, a black-box GNN ranks a COX-2 inhibitor as a promising candidate for a neurodegenerative condition. Lacking mechanistic visibility, the suggestion appears arbitrary. An explainable KG-based model, by contrast, reveals a multi-hop chain linking COX-2 inhibition to microglial PGE2 modulation, synaptic pruning pathways, and neuroinflammation signals observed in clinical cohorts. It also surfaces known gastrointestinal adverse drug reactions and a potential drug–drug interaction with anticoagulants. This enriched rationale informs dose selection, safety monitoring strategies, and downstream experimental planning, which would be inaccessible from a probability score alone.
In the second hypothetical scenario, a neurosymbolic model proposes a JAK/STAT-modulating compound for fibrotic disease. The explanation identifies a critical protein–protein interaction that connects the compound’s primary target to TGF- signaling. However, uncertainty estimation highlights this edge as low-confidence due to limited provenance. Researchers may then prioritize validating this molecular interaction before committing resources to in vivo studies. Again, mechanistic transparency and uncertainty together guide more efficient and disciplined experimentation.
In the third hypothetical scenario, a GNN achieves high AUROC for a predicted drug–disease association driven by dense network connectivity. Inspection of the extracted evidence chain reveals a missing mechanistic link: the proposed target is not expressed in the disease-relevant tissue, and the inferred pathway direction contradicts established disease biology. In addition, safety annotations surface a known adverse effect that directly exacerbates a core clinical feature of the disease. Despite strong benchmark performance, the prediction is rejected as pharmacologically non-viable due to the absence of a coherent mechanism and unresolved safety conflicts.
8. Discussion
The central message of this Perspective is not that benchmark metrics are irrelevant, nor that every useful discovery must begin with a complete mechanistic account. Rather, it is that pharmacological credibility cannot be inferred from predictive performance alone. Drug repurposing decisions require more than ranked predictions: they require hypotheses that can be interrogated in terms of mechanism of action, pathway relevance, tissue context, safety liabilities, and translational plausibility. From this perspective, the limitation of many current AI approaches is not simply opacity in the abstract, but the absence of explanations that are usable within pharmacological reasoning and experimental decision-making.
This point should be calibrated across stages of the translational pipeline. In early exploratory discovery, a statistically strong signal may still be valuable as a hypothesis-generating starting point even before a detailed mechanistic rationale is available. The halicin case underscores this possibility (Stokes et al., 2020). By contrast, when predictions are used to prioritize costly experiments, justify biological plausibility, assess safety risks, support translational investment, or inform regulatory-facing use, mechanistic explainability becomes substantially more important. In these contexts, opaque rankings provide too little guidance about why a candidate might work, how it might fail, and which liabilities should be monitored.
Network medicine approaches (Barabási et al., 2011) illustrate this complementarity: network proximity–based screening identified sildenafil as a candidate for Alzheimer’s disease, later supported by real-world patient data and iPSC experiments (Fang et al., 2021), and multiple repurposable drugs for SARS-CoV-2 under pandemic time pressure (Zhou et al., 2020). In both cases, translational credibility emerged not from computational prediction alone, but from the reinforcement of mechanistic rationale by empirical follow-up.
The value of a computational rationale is not only that it is interpretable after the fact, but that it helps determine which assays to run, which liabilities to monitor, which mechanistic links to validate first, and which predictions are too weakly supported to justify immediate follow-up. In this sense, explainability is inseparable from resource allocation in translational pharmacology: it helps convert broad computational possibility into disciplined experimental prioritization. But the ultimate test of pharmacological credibility remains empirical, whether through in vitro target engagement, phenotypic assays, in vivo studies, clinical observation, or real-world evidence.
Path-based and KG-based explainability for drug repurposing already exists, and important reviews have highlighted the opportunities and limitations of explainable AI (XAI) in drug discovery more broadly (Jiménez et al., 2024; Nunes et al., 2025; Lavecchia, 2025; Tanoli et al., 2025). Our intended contribution is therefore not to claim invention of evidence-chain reasoning itself, but to argue that pharmacology-aligned explanation – enriched by the above-mentioned dimensions – should become a central evaluation target for AI-enabled repurposing systems, especially when they are used for translationally consequential decisions.
This Perspective has several limitations. First, it does not present a new benchmark, executable pipeline, or prospective experimental study; its purpose is conceptual and normative rather than empirical. Second, the literature coverage is representative rather than systematic. Third, the proposed framework will require adaptation across disease areas, data modalities, and contexts of use. These limitations also define an agenda for future work: prospective studies should test whether evidence-chain-based evaluation improves experimental hit rates, resource allocation, safety assessment, or translational decision quality compared with accuracy-first baselines.
9. Conclusion
AI-enabled drug repurposing should not equate high predictive performance with pharmacological credibility. Predictions become more useful for translation when accompanied by inspectable rationales linking drugs to outcomes through targets, pathways, phenotypes, safety context, uncertainty, robustness, and provenance. Knowledge graphs, graph-based learning, and neurosymbolic approaches provide a practical foundation for such evidence-chain reasoning. Their value, however, depends on evaluation standards that reward not only benchmark accuracy, but also mechanistic coherence and experimental usefulness. Prediction without rationale may support exploration, but it is often insufficient for translation. Explainability without uncertainty invites overconfidence. Progress should therefore be measured by how effectively models generate hypotheses that scientists can interrogate, experiments can validate, and decision-makers can trust.
Acknowledgements
The authors thank colleagues who provided feedback on earlier conceptual discussions related to AI-enabled drug repurposing.
Funding Statement
The author(s) declared that financial support was not received for this work and/or its publication.
Edited by: Terry Kenakin, University of North Carolina at Chapel Hill, United States
Reviewed by: Fabio Ferreira, Federal University of Amazonas, Brazil
Data availability statement
The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.
Author contributions
MB: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review and editing. SS: Conceptualization, Supervision, Validation, Writing – review and editing.
Conflict of interest
Authors MB and SS were employed by Unisys.
Generative AI statement
The author(s) declared that generative AI was used in the creation of this manuscript. During the preparation of this work, the authors used Microsoft Copilot in order to improve structure, language and readability. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
- Agarwal C., Queen O., Lakkaraju H., Zitnik M. (2023). Evaluating explainability for graph neural networks. Sci. Data 10, 144. 10.1038/s41597-023-01974-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Barabási A.-L., Gulbahce N., Loscalzo J. (2011). Network medicine: a network-based approach to human disease. Nat. Rev. Genet. 12, 56–68. 10.1038/nrg2918 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bhuyan B. P., Ramdane-Cherif A., Tomar R., Singh T. (2024). Neuro-symbolic artificial intelligence: a survey. Neural Comput. Appl. 36, 12809–12844. 10.1007/s00521-024-09960-z [DOI] [Google Scholar]
- Brière G., Stosskopf T., Loire B., Baudot A. (Forthcoming 2025). Benchmarking the impact of data leakage on the performance of knowledge graph embedding models for biomedical link prediction. 10.1101/2025.01.23.634511 [DOI] [PMC free article] [PubMed]
- Bykova M., Karavani E., Danziger M., Tonegawa-Kuji R., Martin W., Sha Z., et al. (2026). Network-based prediction and real-world patient data observation identify doxycycline as a repurposable drug in Alzheimer’s disease. Neurotherapeutics 23, e00836. 10.1016/j.neurot.2026.e00836 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dai E., Zhao T., Zhu H., Xu J., Guo Z., Liu H., et al. (2024). A comprehensive survey on trustworthy graph neural networks: privacy, robustness, fairness, and explainability. Mach. Intell. Res. 21, 1011–1061. 10.1007/s11633-024-1510-8 [DOI] [Google Scholar]
- Das R., Dhuliawala S., Zaheer M., Vilnis L., Durugkar I., Krishnamurthy A., et al. (2018). “Go for a walk and arrive at the answer: reasoning over paths in knowledge bases using reinforcement learning,” in International Conference on Learning Representations (ICLR). [Google Scholar]
- DeLong L. N., Gadiya Y., Galdi P., Fleuriot J. D., Domingo-Fernández D. (Forthcoming 2024). Mars: A Neurosymbolic Approach for Interpretable Drug Discovery. 10.48550/arXiv.2410.05289 [DOI] [Google Scholar]
- DeLong L. N., Mir R. F., Fleuriot J. D. (2025). Neurosymbolic AI for reasoning over knowledge graphs: a survey. IEEE Trans. Neural Netw. Learn. Syst. 36 (5), 7822–7842. 10.1109/TNNLS.2024.3420218 [DOI] [PubMed] [Google Scholar]
- Drancé M. (2022). “Neuro-symbolic xai: application to drug repurposing for rare diseases,” in International Conference on Database Systems for Advanced Applications (Springer; ), 539–543. [Google Scholar]
- Fang J., Zhang P., Zhou Y., Chiang C.-W., Tan J., Hou Y., et al. (2021). Endophenotype-based in silico network medicine discovery combined with insurance record data mining identifies sildenafil as a candidate drug for alzheimer’s disease. Nat. Aging 1, 1175–1188. 10.1038/s43587-021-00138-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gal Y., Ghahramani Z. (2016). “Dropout as a bayesian approximation: representing model uncertainty in deep learning,” in International Conference on Machine Learning (New York, NY: Proceedings of Machine Learning Research; ), 1050–1059. [Google Scholar]
- Garcez A. d., Lamb L. C. (2023). Neurosymbolic AI: the 3rd wave. Artif. Intell. Rev. 56, 12387–12406. 10.1007/s10462-023-10448-w [DOI] [Google Scholar]
- Guo C., Pleiss G., Sun Y., Weinberger K. Q. (2017). “On calibration of modern neural networks,” in International Conference on Machine Learning (PMLR), 1321–1330. [Google Scholar]
- Himmelstein D. S., Lizee A., Hessler C., Brueggeman L., Chen S. L., Hadley D., et al. (2017). Systematic integration of biomedical knowledge prioritizes drugs for repurposing. elife 6, e26726. 10.7554/eLife.26726 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hossain D., Chen J. Y. (2025). A Study on neuro-symbolic Artificial Intelligence: Healthcare Perspectives. [Google Scholar]
- Jiménez A., Merino M. J., Parras J., Zazo S. (2024). Explainable drug repurposing via path based knowledge graph completion. Sci. Rep. 14, 16587. 10.1038/s41598-024-67163-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kondratyeva L., Alekseenko I., Chernov I., Sverdlov E. (2022). Data incompleteness may form a hard-to-overcome barrier to decoding life’s mechanism. Biology 11, 1208. 10.3390/biology11081208 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Königs C., Friedrichs M., Dietrich T. (2022). The heterogeneous pharmacological medical biochemical network pharmebinet. Sci. Data 9, 393. 10.1038/s41597-022-01510-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lakshminarayanan B., Pritzel A., Blundell C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. Adv. Neural Inf. Process. Syst. 30, 6402–6413. 10.5555/3295222.3295387 [DOI] [Google Scholar]
- Lavecchia A. (2025). Explainable artificial intelligence in drug discovery: bridging predictive power and mechanistic insight. Wiley Interdiscip. Rev. Comput. Mol. Sci. 15, e70049. 10.1002/wcms.70049 [DOI] [Google Scholar]
- Liu Y., Hildebrandt M., Joblin M., Ringsquandl M., Raissouni R., Tresp V. (2021). “Neural multi-hop reasoning with logical rules on biomedical knowledge graphs,” in The Semantic Web – 18th International Conference, ESWC 2021 (Springer; ), 375–391. 10.1007/978-3-030-77385-4_22 [DOI] [Google Scholar]
- Lundberg S. M., Lee S.-I. (2017). A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 30, 4765–4774. [Google Scholar]
- Luo D., Cheng W., Xu D., Yu W., Zong B., Chen H., et al. (2020). Parameterized explainer for graph neural network. Adv. Neural Inf. Process. Syst. 33, 19620–19631. 10.5555/3495724.3497370 [DOI] [Google Scholar]
- Meilicke C., Chekol M. W., Ruffinelli D., Stuckenschmidt H. (2019). “Anytime bottom-up rule learning for knowledge graph completion,” in Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI 2019), 3137–3143. [Google Scholar]
- Morris J. H., Soman K., Akbas R. E., Zhou X., Smith B., Meng E. C., et al. (2023). The scalable precision medicine open knowledge engine (spoke): a massive knowledge graph of biomedical information. Bioinformatics 39, btad080. 10.1093/bioinformatics/btad080 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nunes S., Badreddine S., Pesquita C. (2025). “Rewarding explainability in drug repurposing with knowledge graphs,” in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI 2025), 4624–4632. 10.24963/ijcai.2025/515 [DOI] [Google Scholar]
- Ott S., Meilicke C., Samwald M. (2021). “Safran: an interpretable, rule-based link prediction method outperforming embedding models,” in Proceedings of the 3rd Conference on Automated Knowledge Base Construction (AKBC 2021). [Google Scholar]
- Perdomo-Quinteiro P., Belmonte-Hernández A. (2024). Knowledge graphs for drug repurposing: a review of databases and methods. Brief. Bioinform. 25, bbae461. 10.1093/bib/bbae461 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pushpakom S., Iorio F., Eyers P. A., Escott K. J., Hopper S., Wells A., et al. (2019). Drug repurposing: progress, challenges and recommendations. Nat. Rev. Drug Discov. 18, 41–58. 10.1038/nrd.2018.168 [DOI] [PubMed] [Google Scholar]
- Ribeiro M. T., Singh S., Guestrin C. (2016). “ Why should I trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144. [Google Scholar]
- Singh S., Kumar R., Payra S., Singh S. K. (2023). Artificial intelligence and machine learning in pharmacological research: bridging the gap between data and drug discovery. Cureus 15, e44359. 10.7759/cureus.44359 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stokes J. M., Yang K., Swanson K., Jin W., Cubillos-Ruiz A., Donghia N. M., et al. (2020). A deep learning approach to antibiotic discovery. Cell 180, 688–702.e13. 10.1016/j.cell.2020.01.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tanoli Z., Fernández-Torras A., Özcan U. O., Kushnir A., Nader K. M., Gadiya Y., et al. (2025). Computational drug repurposing: approaches, evaluation of in silico resources and case studies. Nat. Rev. Drug Discov. 24, 521–542. 10.1038/s41573-025-01164-x [DOI] [PubMed] [Google Scholar]
- Wong F., Zheng E. J., Valeri J. A., Donghia N. M., Anahtar M. N., Omori S., et al. (2024). Discovery of a structural class of antibiotics with explainable deep learning. Nature 626, 177–185. 10.1038/s41586-023-06887-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ying R., Bourgeois D., You J., Zitnik M., Leskovec J. (2019). GNNExplainer: generating explanations for graph neural networks. Adv. Neural Inf. Process. Syst. 32, 9240–9251. 10.5555/3454287.3455116 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yuan H., Yu H., Wang J., Li K., Ji S. (2021). “On explainability of graph neural networks via subgraph explorations,” in Proceedings of the 38th International Conference on Machine Learning (ICML), PMLR, 139, 12241–12252. [Google Scholar]
- Zhang K., Yang X., Wang Y., Yu Y., Huang N., Li G., et al. (2025). Artificial intelligence in drug development. Nat. Med. 31, 45–59. 10.1038/s41591-024-03434-4 [DOI] [PubMed] [Google Scholar]
- Zhou Y., Hou Y., Shen J., Huang Y., Martin W., Cheng F. (2020). Network-based drug repurposing for novel coronavirus 2019-ncov/sars-cov-2. Cell Discovery 6, 14. 10.1038/s41421-020-0153-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

