Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 May 11:2026.04.29.721775. [Version 2] doi: 10.64898/2026.04.29.721775

A Scalable Sign-Aware Multi-Omics Knowledge Graph Foundation Model for Mechanistic Drug Action and Clinical Response Predictions

Mohammadsadeq Mottaqi 1, Shuo Zhang 2, Ian Adoremos 3, Pengyue Zhang 4, Lei Xie 5,6,*
PMCID: PMC13192585  PMID: 42182201

Abstract

Mechanistically predicting drug action requires distinguishing activating from inhibitory interactions across broad chemical space, yet most biomedical knowledge graphs and graph neural networks (GNNs) rely on unsigned associations that obscure regulatory logic and have a limited chemical coverage. Here we present SIGMA-KG (SIGned Multi-omics Atlas Knowledge Graph) and FLASH (Fast Lightweight Architecture for Signed Heterogeneous GNN), a graph foundation model pretrained through self-supervised learning on SIGMA-KG. SIGMA-KG integrates chemogenomic perturbation, transcriptomic, proteomic, and clinical data while explicitly encoding biological polarity and directionality. FLASH preserves polarity composition across multi-hop pathways through structural balance principles, enabling scalable mechanistic reasoning. Without task-specific fine-tuning, FLASH consistently outperforms or matches nine state-of-the-art unsigned, relational, and signed graph baselines across drug mode-of-action prediction, clinical response modeling, and drug–drug interaction prediction, while substantially improving computational efficiency. FLASH further enables explainable inductive drug repurposing, achieving a 69.6% external clinical validation success rate across four complex diseases.

Keywords: machine learning, deep learning, graph neural network, graph foundation model, drug discovery, drug repurposing, drug-drug interaction, signed multi-omics knowledge graph

1. Introduction

Understanding the complex interactions between drugs, genes, proteins, and diseases is fundamental to drug discovery, precision medicine, and chemical safety assessments [1]. Recent advances in graph neural networks (GNNs) have enabled powerful models that capture the topological and relational structure of biomedical knowledge graphs (KGs), learning embeddings that jointly encode biological semantics and network connectivity [2]. These representations show promise for predicting drug-target interactions (DTIs) [3], drug-drug interactions (DDIs) [4], and identifying new therapeutic uses for existing drugs (drug repurposing) [5]. However, existing frameworks remain limited in their ability to capture the mechanistic logic underlying drug action.

The first limitation is chemical coverage. Many widely used biomedical KGs, including Hetionet (47,031 nodes; 2.2 million edges) [6], PrimeKG (129,375 nodes; 4 million edges) [7], and BioPathNet(32,000 edges) [8], focus primarily on only a few thousand approved or well-annotated compounds, thereby under-representing the much larger chemical space relevant to early-stage discovery. The second limitation is that all of these KGs encode interactions as unsigned relations without explicit effect polarity, thereby conflating activation with inhibition. Signed biological databases address polarity in more restricted settings: SIGNOR 2.0 encodes ~33,000 activation or inhibition edges across proteins, RNA, and metabolites [9], whereas OmniPath [10] and TRRUST v2 [11] cover additional single-modality regulatory layers. Although these resources preserve directional signs, +1 for activation/up-regulation and −1 for inhibition/down-regulation, they are not integrated into a unified, multi-omics atlas that simultaneously connects chemical perturbations, molecular responses, and clinical phenotypes in a form suitable for translational modeling (see Supplemental Information).

This missing polarity information is consequential because many biological signaling and regulatory processes are inherently sign-dependent. Within these signaling networks, the functional polarity of a multi-hop pathway is governed by the composition of constituent interactions, where the net regulatory effect is determined by the multiplicative product of signs along the path [12, 13]. Consistent with this view, empirical studies of gene-gene interaction networks have shown that balanced triads are significantly overrepresented, whereas unbalanced triads are underrepresented relative to randomized controls; this suggests that interaction signs in biological networks are organized in a manner consistent with structural balance principles [14, 15]. Together, these observations support the use of sign-consistent path composition as a biologically grounded inductive bias for modeling net mechanistic effects in biological systems. The practical importance of considering polarity is illustrated by vemurafenib, a BRAF inhibitor approved for melanoma. In unsigned KGs, its proximity to the ERK signaling cascade could suggest therapeutic benefit; yet in KRAS-mutant cells, the same drug can paradoxically hyperactivate ERK and promote tumor growth [16]. Such context-dependent sign reversals are difficult to capture with unsigned proximity-based models.

While relational (heterogeneous) GNNs (rGNNs) assign distinct embeddings to different edge types, they still represent an inhibition or activation interaction as a categorical label rather than an algebraic operator, and therefore cannot model signal transduction along multi-hop paths (e.g., A→−B→−C⇒A→+C [13]. We define this conceptual failure as polarity drift: the systematic divergence between a model’s predicted functional direction and the ground-truth sign-product of a regulatory pathway. This divergence compounds as path length increases, necessitating signed Graph Neural Networks (sGNNs) that explicitly integrate edge polarity into the message-passing framework. By grounding representation learning in structural balance principles, sGNNs facilitate robust multi-hop reasoning that preserves the integrity of long-range signaling logic [17].

Large language models (LLMs) provide a complementary source of biomedical knowledge through text, but they do not explicitly model sign-consistent propagation over graph topology. They retrieve pairwise associations from the literature, yet they are not designed to reliably propagate the product of signed directed edges across multi-hop regulatory paths, especially when such paths are not explicitly described in training corpora [18]. For this reason, graph-based learning with explicit sign semantics remains a distinct and necessary modeling framework for mechanistic prediction in biomedicine [19].

Recent studies applying sGNNs to biological graphs support the promise of sign-aware modeling, while also highlighting current limitations. SIGDR [20] uses sign-aware contrastive learning to model drug-disease associations, and CSGDN [21] applies signed graph diffusion to predict crop gene-phenotype associations. These studies demonstrate the utility of signed graph learning in biology, but they adopt supervised learning objectives, precluding self-supervised pre-training on unlabeled network topology, and operate on single-modality, small-scale graphs (< 20, 000 edges) that are not scalable to the multi-omics, million-edge setting required for comprehensive drug action modeling.

In parallel, graph foundation models have emerged as a promising direction for biomedicine by showing that pre-training on large KGs can yield transferable representations for drug repurposing, DDI prediction, and gene function annotation [8, 22]. TxGNN [5] is the most directly relevant predecessor: trained on a relational KG comprising 17,080 diseases and 7,957 drug candidates [7], it performs indication and contraindication ranking, substantially improving over earlier approaches. However, TxGNN relies on supervised metric-learning ob jectives defined over labeled drug-disease associations, which may limit generalization to diseases with scarce or no known therapeutic examples. Additionally, its underlying KG does not encode edge polarity and covers no early-stage small molecules, preventing sign-consistent multi-hop reasoning. BioPathNet [8] advances link prediction via path-based neural reasoning on an unsigned KG, but it is trained in a fully supervised manner. PT-KGNN [22] introduces self-supervised pre-training on unsigned biomedical KGs and demonstrates scaling with KG size, yet it applies homogeneous aggregation that discards relational and sign semantics. Collectively, these models expose a persistent gap: no existing framework combines self-supervised pre-training with explicit directed edge-sign semantics in a large-scale, multi-omics atlas.

We therefore hypothesize that explicitly encoding directed edge polarity in a large-scale multi-omics KG, and using this graph to pre-train a sign-aware GNN in a self-supervised manner, would produce more accurate and mechanistically interpretable predictions of drug actions than unsigned homogeneous or relational alternatives. To test this hypothesis, we constructed SIGMA-KG, a signed and directed multi-omics atlas of ~127,000 nodes, including 76,009 small molecules and 3.8 million signed edges integrated from 16 heterogeneous biomedical resources. We further developed FLASH, a computationally efficient signed heterogeneous GNN that enables the pre-training of large-scale signed KG in a self-supervised manner to learn transferable, balance-aware node representations without requiring task-specific supervision.

To evaluate the advantages of sign-aware representation learning, we benchmarked FLASH on three critical tasks spanning molecular to clinical scales: : target-specific mode-of-action (Ts-MoA), drug-induced clinical response (DCR), and DDI predictions. Across these benchmarks, FLASH matched or exceeded nine SOTA unsigned, relational, and signed baselines while maintaining high computational efficiency. Furthermore, we illustrated the model’s translational utility through explainable inductive drug repurposing, where high-confidence therapeutic candidates for complex diseases are recovered via biologically coherent, multi-hop signed paths. These results establish SIGMA-KG and FLASH as a scalable foundation for sign-aware, mechanistically grounded predictive modeling in drug discovery and clinical pharmacology.

2. Results

2.1. SIGMA-KG and FLASH establish a scalable sign-aware foundation-model framework for drug discovery

Here we introduce a scalable foundation-model framework for mechanistic drug discovery and drug repurposing, centered on SIGMA-KG and FLASH. Unlike conventional unsigned GNNs that overlook the polarity of biological interactions, our framework explicitly models signed biological effects and leverages structural balance principles to support sign-aware learning across molecular, clinical, and systems-level pharmacological settings.

We constructed SIGMA-KG, a large-scale, signed, multi-relational atlas that bridges molecular perturbations and clinical outcomes (Fig. 1a; Supplementary Table 1)). Specifically, we harmonized 16 large-scale resources spanning chemical genomics, transcriptomics, proteomics, and clinical phenotypes into a unified, signed, multi-layer KG. The resulting atlas comprises 127,422 nodes across four primary biomedical entity types (compounds, genes, proteins, and diseases) and 3.88 million signed edges (Table 1; Supplementary Table 2). These edges encode explicit biological polarities (a signed label per edge, +1 or −1), including activation versus inhibition, up-regulation versus down-regulation, and therapeutic indication versus adverse contraindication, thereby providing a mechanistically rich substrate for sign-aware modeling (Supplementary Figs. 1a and 1b).

Fig. 1. Sign-aware multi-omics graph foundation model for mechanistic drug action prediction across molecular-to-clinical scales.

Fig. 1

a, Construction of SIGMA-KG from harmonized identifiers and integrated multi-omics and clinical sources, forming a heterogeneous directed graph of compounds, genes, proteins, and diseases. Edge signs encode mechanistic polarity (positive or negative), including activation/inhibition, upregulation/downregulation, and indication/contraindication. Structural balance is enforced by pruning low-confidence edges involved in unbalanced signed cycles with sign product < 0. Pruning preferentially removes lower-confidence compound-gene or genedisease edges. The resulting graph contains ~127K nodes and ~3.8M signed edges. b, FLASH encodes each node using 1-hop neighborhood aggregation conditioned on edge-based signed motifs. Separate message-passing channels for positive and negative signed edges preserve transitive sign composition along multi-hop paths and mitigate sign mixing during aggregation. A multi-layer perceptron (MLP) produces the final embedding z for each node. c, Self-supervised pre-training of FLASH on the global topology of SIGMA-KG using sign-aware objectives to learn transferable graph representations. d, The frozen embeddings from the pre-trained foundation model is evaluated across downstream pharmacological tasks, including target-specific mode-of-action prediction (agonist/antagonist), drug-induced clinical response prediction (indication/contraindication), drug-drug interaction prediction, and a drug repurposing case study across four diseases. e, Mechanistic explainability via multi-hop signed metapath tracing in SIGMA-KG, aggregating intermediate disease-linked genes and proteins to support candidate drug-disease predictions.

Table 1. Statistical details of SIGMA-KG.

SIGMA-KG comprises 127,422 nodes and 3,887,829 signed edges connecting four biomedical entity types. Percentages (%) are relative to the total number of nodes or edges, respectively. Where multiple relation types exist between the same source and target entity types, they are listed together in a single row with their corresponding sign labels.

(a) Node statistics.

Entity type Description No. of nodes % of all nodes
Gene Human protein-coding genes (HGNC gene symbols) 19,245 15.1%
Protein Human proteins (UniProt identifiers) 18,609 14.6%
Compound Small molecules and drug-like chemicals (PubChem CID) 76,009 59.7%
Disease Human diseases and clinical phenotypes (Human Phenotype Ontology (HPO)) 13,559 10.6%

Total 127,422 100%
(b) Edge statistics by biomedical entity pair.
Source entity Target entity Relation type(s) with sign No. of edges % of all edges

Gene Disease Up-regulated in (+1); Down-regulated in (−1) 1,489,308 38.3%
Chemical Gene Up-regulates (+1); Down-regulates (−1) 1,322,830 34.0%
Compound Compound Similar to (+1) 581,872 15.0%
Protein Protein Interacts with (+1) 251,762 6.5%
Compound Disease Contraindication (+1); Indication (−1) 159,957 4.1%
Disease Disease Resembles (+1) 49,052 1.3%
Gene Protein Encodes (+1) 19,337 0.5%
Compound Protein Activates (+1); Inhibits (−1) 13,711 0.4%

Total 3,887,829 100%

SIGMA-KG exhibits high cross-modal density, with an average node degree of 30.5, enabling the tracing of continuous mechanistic chains from molecular perturbations to clinical phenotypes. In total, the atlas links 22,862 compounds directly to downstream transcriptomic signatures and proteomic targets, which are further connected to a unified disease landscape of 13,559 clinical phenotypes. This organization substantially expands the mechanistically reachable therapeutic space. Although only 3.8% of compounds (2,869/76,009) are directly linked to diseases through recorded clinical associations, the proportion of compounds connected to at least one disease phenotype increases to 28.5% (21,649/76,009) within two hops and 47.2% (35,905/76,009) within three hops. These results show that SIGMA-KG is not simply a collection of disjoint datasets, but a dense, searchable multi-omics map through which biological effects can be propagated across scales. We further examined the signed topology of SIGMA-KG through the lens of Heider’s structural balance theory [14] (Supplementary Fig. 1c). Across closed triads in the atlas, 99.6% were balanced, indicating that the signed graph structure is highly non-random and largely consistent with structural balance principles (Supplementary Fig. 1d). This high degree of triadic consistency supports the use of sign-aware graph learning as a biologically grounded inductive bias for modeling multi-hop effects in the network.

Leveraging SIGMA-KG, we developed FLASH, a scalable signed heterogeneous GNN designed for atlas-scale representation learning (Fig. 1b). Unlike standard relational GNNs, FLASH utilizes a motif-aware attention mechanism to encode seven biologically informed signed patterns, explicitly preserving signaling polarity and the ’double negative’ logic where inhibiting an inhibitor results in net activation. In self-supervised pre-training, FLASH uses observed edges and their signs in SIGMA-KG as supervision to optimize the link sign product loss, a sign-aware structural objective that encourages similar embeddings for positively connected nodes and dissimilar embeddings for negatively connected nodes (Methods). This process yields robust, transferable node embeddings that capture both topological context and sign-aware biomedical semantics (Fig. 1c).

FLASH is computationally efficient while retaining the expressive capacity required for multi-hop sign-aware reasoning. Compared with existing signed GNNs, the framework achieves a substantial reduction in training time, while producing a general-purpose graph foundation model that can be applied across diverse downstream pharmacological tasks. In the following sections, we evaluate this foundation model across three biological scales: Ts-MoA prediction at molecular level, DCR prediction at clinical level, and DDI prediction at systems level, followed by explainable inductive drug repurposing analyses (Fig. 1d and 1e).

To evaluate the predictive utility of the pretrained node embeddings, we implemented a standardized benchmarking protocol comparing FLASH against nine unsupervised GNN baselines. For molecular-level (Ts-MoA) and clinical-level (DCR) tasks, we employed an internal validation strategy using a randomized edge-split protocol. To prevent data leakage, SIGMA-KG was partitioned into three independent sets of held-out evaluation edges; the models were then pretrained in a self-supervised manner on the remaining graph topology to generate 256-dimensional node embeddings. These frozen embeddings were subsequently evaluated via a Multi-Layer Perceptron (MLP) head—trained on the internal training edges—to perform binary and three-class edge-sign prediction on the held-out sets.

In contrast, systems-level DDI prediction followed an external validation strategy. Models were pretrained on the complete SIGMA-KG atlas, and the resulting compound embeddings were directly assessed on two independent DDI benchmark datasets: DrugBank [23] and MUDI [24]. To ensure a controlled comparison, all models were initialized with identical 256-dimensional node features derived from truncated singular value decomposition (SVD). Evaluation across all tasks utilized standardized MLP architectures and hyperparameter tuning protocols (Methods). By standardizing architectural components and feature initialization, this experimental design isolates the specific predictive contribution of the signed multi-relational inductive bias from confounding factors. This enables a rigorous head-to-head comparison between FLASH and existing baselines while providing a clear basis for demonstrating the systematic performance advantage of sign-aware modeling over standard relational and unsigned graph architectures.

2.2. FLASH enables robust target-specific mode-of-action prediction

We first evaluated the capacity of the SIGMA-KG-pretrained foundation model to accurately predict Ts-MoA, a critical prerequisite for understanding the downstream functional consequences of drug-target engagement [3]. This task is central to pharmacology as it distinguishes whether a protein ligand acts as an agonist (target activator) or an antagonist (target inhibitor), a distinction that conventional GNNs often overlook by treating DTIs as unsigned proximities. We formulated Ts-MoA prediction as an edge-sign prediction task in both binary (activation, +1 vs. inhibition, −1) and three-class classification settings, with the latter incorporated a 5:1 ratio of non-existent associations (0) to labeled edges [25] (Methods). These evaluations were conducted on held-out datasets of Compound–Protein interactions that were entirely unseen during model pre-training, ensuring a rigorous assessment of the model’s ability to generalize to novel molecular mechanisms.

Our analysis showed that signed GNN architectures overall outperformed unsigned and relational baselines in Ts-MoA prediction (Supplementary Tables 3 and 4), with FLASH achieving SOTA or comparable performance across both binary and three-class settings (Fig. 2 and Supplementary Fig. 2). In the binary classification setting (Fig. 2a–c, Supplementary Table 5), the top signed baselines (FLASH and SNEA) significantly outperformed the top unsigned baseline (GraphSAGE) in both Accuracy (3.63% higher, p = 0.014, unpaired Welch’s t-test) and Macro-F1 (3.51% higher, p = 0.020). Within this setting, FLASH achieved the highest overall Accuracy (0.8990±0.0143) and showed performance statistically indistinguishable from the top-performing signed baseline (SGCN) in AUROC (p = 0.980) and Macro-F1 (p = 0.836; Supplementary Table 4.

Fig. 2. Performance comparison of FLASH with unsigned, relational, and signed graph neural network (GNN) baselines for predicting target-specific mode-of-action.

Fig. 2

a–c, Binary classification performance (activation or inhibition). Comparison across baselines using (a) AUROC, (b) accuracy, and (c) macro-F1 score. d–f, Three-class classification performance (activation, inhibition, or no interaction) evaluated using (d) AUROC, (e) accuracy, and (f) macro-F1 score. Bars represent the mean across three independent runs with random train–test splits, with error bars indicating the standard deviation. Statistical significance was assessed using two-tailed paired t-tests. FLASH was compared against the overall top-performing baseline or, in instances where FLASH established the state-of-the-art, the next-best performing model. Additionally, the leading signed baseline was compared against the top relational and unsigned alternatives to quantify the predictive gains afforded by explicit edge polarity in SIGMA-KG dataset. Significance annotations are shown above brackets (*p < 0.05, * * p < 0.01, n.s., not significant). Complete P-values are reported in Supplementary Tables 17 and 18, and full metric results for binary and three-class tasks are provided in Supplementary Tables 5 and 6. g, Accuracy–efficiency trade-off. Macro-F1 score for three-class classification plotted against training time per epoch for the top-performing signed and heterogeneous GNN baselines, illustrating the accuracy–efficiency trade-off. Marker size indicates approximate memory usage during training. FLASH achieves competitive or superior predictive performance while maintaining substantially lower training time compared with other top-performing signed GNN architectures.

Similarly, in the more challenging three-class setting (Fig. 2d–f), Supplementary Table 6, which requires joint discrimination of activation, inhibition, and the absence of an interaction, signed baselines demonstrated a substantial advantage in interaction detection and polarity classification, with FLASH significantly outperforming the best unsigned baseline (BGRL) in both Accuracy (1.32% higher, p = 0.0183) and Macro-F1 (4.67% higher, p = 0.0130). Furthermore, in the three-class setting, the top signed baseline (SNEA) significantly outperformed the top relational model (HGATE) by 4.48% in Macro AUROC (p ≤ 0.001), highlighting the advantage of signed architectures in distinguishing true functional interactions from biological noise. Although SNEA maintained a statistically significant lead in three-class Macro AUROC (p = 0.032), FLASH’s superior F1 scores for both inhibition and non-existent edges suggest a more balanced ability to capture functional polarity while limiting false positives (Supplementary Fig. 2.

Beyond its predictive accuracy, FLASH demonstrates superior computational scalability, a prerequisite for atlas-scale inference. Quantitatively, FLASH achieved SOTA Macro-F1 scores while requiring 4.5-fold less training time per epoch and substantially lower memory overhead compared to the next-best signed baseline, SGCN (Fig. 2g). These results establish FLASH as a signed graph foundation model capable of high-fidelity Ts-MoA inference at atlas scale while maintaining computational efficiency.

2.3. FLASH enables robust drug-induced clinical response prediction

Accurate DCR prediction o is essential for prioritizing therapeutic indications while mitigating adverse contraindications [5]. Distinguishing between these outcomes is a primary challenge in drug development, exacerbated by a severe class imbalance where documented contraindications (in SIGMA-KG) significantly outnumber therapeutic indications (~12.4:1). We evaluated the capacity of FLASH using its frozen embeddings to navigate this imbalanced landscape by predicting DCR, formulated as an edge-sign prediction task in both binary (indication, −1 vs. contraindication, +1) and three-class classification settings, with the latter incorporating a 5:1 ratio of non-existent associations (0) to labeled edges to serve as a rigorous proxy for real-world inductive drug repurposing [25] (Methods). These evaluations were performed on held-out datasets of Compound–Disease interactions entirely excluded from pre-training, testing the model’s ability to recover sparse, high-stakes therapeutic signals from the vast landscape of clinical associations.

Our comparative analysis revealed that signed baselines consistently outperformed top-performing unsigned and relational GNNs (Supplementary Tables 3 and 4). In the binary setting (Fig. 3a–c), the top-performing signed models, FLASH and SDGNN, exhibited a significant performance advantage in both Accuracy and Macro-F1, with FLASH outperforming the top unsigned baseline (BGRL; p = 0.002) and the leading relational baseline (HGATE; p ≤ 0.001) in Accuracy. This performance gap widened in the more demanding three-class setting (Fig. 3d–f), where the top signed model (SGCN) outperformed the leading unsigned model (DGI) by 11.46% in Macro-F1 (p = 0.003) and the top-performing relational model (HGATE) by 3.32% (p = 0.017) While unsigned models such as DGI achieved high AUPRC (0.9995) in the binary setting, their inability to model directionality led to a substantial failure in identifying minority-class therapeutic indications in the three-class setting (F1Indication ≤ 0.51), whereas signed models such as FLASH more effectively prioritized these relationships (F1Indication = 0.81; Supplementary Figs. 2b and 2d).

Fig. 3. Performance comparison of FLASH with unsigned, relational, and signed graph neural network (GNN) baselines for predicting DCR relationships.

Fig. 3

a–c, Binary classification performance. Comparison across baselines using (a) AUROC, (b) accuracy, and (c) macro-F1 score. d–f, Three-class classification performance (indication, contraindication, and no association) evaluated using (d) AUROC, (e) accuracy, and (f) macro-F1 score. Bars represent the mean across three independent runs with random train–test splits, with error bars indicating the standard deviation. Statistical significance was assessed using two-tailed paired t-tests. FLASH was compared against the overall top-performing baseline or, in instances where FLASH established the state-of-the-art, the next-best performing model. Additionally, the leading signed baseline was compared against the top relational and unsigned alternatives to quantify the predictive gains afforded by explicit edge polarity in SIGMA-KG dataset. Significance annotations are shown above brackets (*p < 0.05, * * p < 0.01, n.s., not significant). Complete P-values are reported in Supplementary Table A1, and full metric results for binary and three-class tasks are provided in Supplementary Tables 7 and 8. g, Accuracy–efficiency trade-off. Macro-F1 score for three-class classification plotted against training time per epoch for the top-performing signed and heterogeneous GNN baselines. Marker size reflects approximate GPU memory usage during training. FLASH achieves competitive or superior predictive performance while maintaining substantially lower training time compared with other high-performing signed GNN architectures.

Within this overall pattern, FLASH showed strong performance in the binary setting (Supplementary Table 7). In detail, FLASH attained the highest overall Accuracy (0.9893 ± 0.0003), significantly outperforming the runner-up baseline (SDGNN; p = 0.008) while establishing a 1.51% lead over the leading relational baseline (HGATE; p ≤ 0.001).

Beyond binary classification, FLASH also maintained strong performance in the more challenging three-class setting which includes non-existent associations (Supplementary Table 7). FLASH demonstrated strong discriminative power, achieving the highest Macro AUROC (0.9947 ± 0.0011). While SGCN showed a marginal statistical advantage in Accuracy (Δ = 0.22%, p = 0.002), FLASH maintained comparable Macro-F1 performance (p = 0.275), and achieved a 4.5-fold reduction in training time per epoch (Fig. 3g). This capacity to maintain high Macro AUROC and F1 scores under severe class imbalance, while remaining computationally efficient, provides a strong foundation for our subsequent therapeutic indication prediction task, ensuring that the predicted drug repurposing candidates are prioritized based on robust, sign-aware structural evidence.

2.4. FLASH significantly improves drug-drug interaction prediction across both warm- and cold-start scenarios

Accurate DDI prediction is a cornerstone of clinical safety and rational polypharmacy [26]. Yet this application remains a formidable challenge, particularly when models must generalize to novel therapeutic contexts where interaction data are sparse. To evaluate the pharmacological utility of the representations learned by our foundation model, we assessed its capacity to predict DDIs. Crucially, within SIGMA-KG, compound-compound edges are derived exclusively from molecular structural similarity, meaning that the foundation model receives no explicit pharmacological or clinical DDI knowledge during pre-training. The DDI benchmark thus evaluates whether FLASH can synthesize molecular structure with multi-omics signaling to predict complex interactions. This capability is critical for capturing DDIs that are not encoded within chemical structures alone, but instead emerge through the network-mediated propagation of biological changes across the atlas.

We benchmarked our framework using two external DDI benchmark datasets, DrugBank (version 5.1) [23] and MUDI [24]. DrugBank dataset contains clinically documented interactions curated from regulatory labels and literature, while MUDI dataset incorporates pharmacologically derived synergistic and antagonistic effects. Performance was evaluated via a binary classification task—distinguishing the presence or absence of interactions—under two distinct evaluation scenarios: transductive (warm-start) prediction based on random pair splitting, and inductive (cold-start) prediction using a drug-disjoint split to evaluate generalization to novel compounds (Fig. 4 and Supplementary Fig. 3). While the transductive scenario measures the model’s ability to interpolate within known drug spaces, the inductive scenario provides a more stringent, real-world assessment by requiring the model to predict interactions for compounds entirely absent from the training set (Methods).

Fig. 4. Performance comparison of FLASH with unsigned, relational, and signed graph neural network (GNN) baselines for predicting drug–drug interactions.

Fig. 4

Model performance was evaluated using AUROC, accuracy, and macro-F1. Efficiency was assessed by plotting average macro-F1 against training time per epoch, with marker size indicating GPU memory usage during training. a-c, Transductive prediction on DrugBank. d-f, Transductive prediction on MUDI. g, Efficiency analysis (transductive). h-j, Inductive prediction on DrugBank. k-m, Inductive prediction on MUDI. n, Efficiency analysis (inductive). Bars represent the mean across three independent runs with random splits, with error bars indicating the standard deviation. Statistical significance was assessed using two-tailed paired t-tests. FLASH was compared against the overall top-performing baseline or, in instances where FLASH established the state-of-the-art, the next-best performing model. The leading signed baseline was compared against the top relational and unsigned alternatives to quantify the predictive gains afforded by explicit edge polarity in SIGMA-KG. Significance annotations above brackets indicate (*p < 0.05, * * p < 0.01, n.s., not significant). Complete P-values are reported in, Supplementary Tables 17 and 18.

2.4.1. Transductive DDI prediction scenario

In the transductive setting, the results showed that signed models incorporating edge polarity consistently outperform both unsigned and relational baselines (Supplementary Tables 3 and 4). Across both DrugBank and MUDI benchmarks, the node representations learned by FLASH significantly outperformed the leading unsigned (DGI) and relational (HGATE) models in all evaluated metrics (Fig. 4a–f, Supplementary Tables 9 and 10). Specifically, on DrugBank, FLASH significantly improved AUROC and Macro-F1 by 1.72% and 2.94% (both p < 0.0001) respectively against the top unsigned model (DGI), and by 1.26% (p < 0.0001) and 2.07% (p = 0.0002) against leading relational model, HGATE. These margins were even more pronounced on the MUDI dataset, where FLASH exceeded the Macro-F1 of DGI and HGATE by 3.66% and 2.99%, respectively (both p = 0.0001).

In this systems-level evaluation, FLASH established new state-of-the-art (SOTA) performance across multiple benchmarks. On the DrugBank dataset (Fig. 4a–c), FLASH significantly outperformed the leading signed baseline (SDGNN) in AUROC (0.9742 ± 0.0004), AUPRC (0.9735 ± 0.0005), and Macro-F1 (0.9166 ± 0.0015; all three p < 0.001). This superior predictive capability was sustained on the MUDI dataset (Fig. 4d–f), where FLASH achieved a Macro-F1 of 0.9557 ± 0.0005, a significant 1.75% improvement over SDGNN (p < 0.001). Notably, these performance gains were achieved with high computational efficiency; FLASH required 3-fold less training time per epoch and significantly lower memory overhead compared to SDGNN, underscoring its practical utility for atlas-scale biomedical modeling (Fig. 4g).

2.4.2. Inductive DDI prediction scenario

We next assessed whether the inductive biases of signed graph pre-training enhance the model’s generalization to novel compounds. To ensure a rigorous evaluation, we implemented a drug-disjoint partitioning strategy (splitting the unique set of compounds into training, validation, and test sets) such that every test-set interaction involved at least one compound entirely unseen during the classifier’s training phase. Although predictive performance across all architectures decreased in this more challenging cold-start scenario, signed and relational models exhibited superior robustness compared to unsigned baselines, maintaining significantly higher generalizability (Fig. 4h–m, Supplementary Tables 3 and 4).

On the DrugBank dataset (Fig. 4h–j; Supplementary Table 11), FLASH outperformed all baselines, including the leading relational model (HGATE), achieving an AUROC of 0.7960 ± 0.0157 (p = 0.20) and an AUPRC of 0.7971 ± 0.0192 (p = 0.18). While the performance gap between the top models narrowed in this high-difficulty setting, FLASH remained among the leading methods across metrics. On MUDI dataset (Fig. 4k–m; Supplementary Table 12), FLASH achieved the highest Macro-F1 (0.6932 ± 0.0085; (Fig. 4n)) while showing comparable performance to the top baseline (SDGNN) in AUPRC (p = 0.344). Although SDGNN showed a marginal lead in Accuracy (0.6989 vs 0.6934, p = 0.445), the performance of FLASH remains highly competitive.

Taken together, these results show that FLASH offers a favorable balance between predictive performance and computational efficiency for large-scale biomedical applications. Across all scenarios, FLASH achieved comparable or superior Macro-F1 scores while benefiting from the reduction in training time and memory overhead compared to other signed and relational models (Fig. 4g and Fig. 4n). Systematic benchmarking across varying graph scales further confirms that FLASH maintains linear scalability and a significantly reduced memory footprint, as detailed in Supplementary Information and Supplementary Fig. 4. Collectively, these findings suggest that by leveraging the signed multi-omics structure of SIGMA-KG, FLASH learns latent representations that reflect functional and mechanistic similarity rather than molecular structure alone, enabling robust sign-aware prediction of drug actions.

2.5. FLASH-enabled therapeutic indication discovery across diverse disease pathologies

To evaluate the translational utility and cross-domain generalizability of our graph foundation model, we performed a systematic drug repurposing screen across four high-burden diseases with distinct etiologies: Alzheimer’s disease (AD), Parkinson’s disease (PD), Schizophrenia, and Chronic Obstructive Pulmonary Disease (COPD). By spanning neurodegenerative, psychiatric, and respiratory pathologies, these cases serve as a rigorous assessment of the ability of the model to navigate diverse therapeutic landscapes.

We screened 2,869 Food and Drug Administration (FDA)-approved drugs in an inductive setting, utilizing a frozen-encoder linear probing pipeline with a three-class MLP classifier trained on frozen FLASH embeddings(Methods). For each drug–target-disease pair, the model outputs probabilistic scores for each of three classes: therapeutic indication, adverse contraindication, and no association. Using a confidence threshold (Pindication > 0.5), we prioritized candidate drug–disease pairs as high-confidence therapeutic indications for each of the four disease pathologies (Figs. 5a–c).

Fig. 5. Framework and results of drug repurposing predictions generated by pre-trained FLASH on SIGMA-KG.

Fig. 5

a, Prediction framework. FLASH is first pre-trained self-supervisedly on SIGMA-KG to learn node embeddings. Drug–disease edge representations are then constructed by concatenating the corresponding node embeddings and passed to an MLP to predict the likelihood of indication, contraindication, and no association. b, Example predictions for Alzheimer’s disease. The model ranks candidate drugs by predicted probability of therapeutic indication, identifying both approved treatments and high-confidence candidates absent from the training data. c, Mechanistic interpretation. Multi-hop biological pathways linking drugs to diseases through proteins and genes are traced within SIGMA-KG, providing mechanistic hypotheses for predicted therapeutic effects. d, Repurposing screening workflow. A total of 2,869 drugs were screened across 4 diseases (Alzheimer’s disease (AD), Parkinson’s disease (PD), schizophrenia, and chronic obstructive pulmonary disease (COPD)), yielding 99 initial predicted drug–disease associations. After filtering documented indications present in SIGMA-KG, 23 prospective prioritized indications remained for further evaluation. e, Distribution of predicted therapeutics indications across diseases. Predictions include both recovered known indications within SIGMA-KG and prospective network-blind candidates not observed during training. f, External validation of prioritized candidates. Among the 23 prospective prioritized indications, 16 correspond to drugs already approved for the predicted disease, providing external clinical support (69.6%), while the remaining 7 represent novel repurposing candidates.

This screening recovered documented clinical knowledge while also identifying prospective therapeutic opportunities (Fig. 5d). In total, FLASH prioritized 99 high-confidence therapeutic indications, providing a map of candidate treatments across all four diseases. To assess the fidelity of our foundation model, we partitioned the 99 prioritized therapeutic indications into recovered and prospective sets based on their presence in the pre-training atlas. We found that FLASH recovered 76 known indications, that is, previously documented drug–disease edges in SIGMA-KG; this indicates that the model preserved and prioritized established pharmacological knowledge during embedding learning and downstream classification (Fig. 5e, Table 2 and Supplementary Table 13).

Table 2. Systematic screening and clinical validation of therapeutic indications across four high-burden diseases.

For each disease, 2,869 FDA-approved drugs in the SIGMA-KG universe were screened using pre-trained FLASH embeddings and a three-class MLP classifier. Total prioritized therapeutic indications denote drug–disease pairs with an indication probability P > 0.5; Recovered known indications represent predicted edges already documented within the SIGMA-KG training set; Prospective prioritized therapeutic indications correspond to network-blind candidates entirely absent from the knowledge graph during pre-training; Clinically validated indicates the subset of prospective candidates retrospectively confirmed via global regulatory records (FDA, European Medicines Agency, Medicines and Healthcare products Regulatory Agency); Novel repurposing drug candidates represent the high-confidence therapeutic hypotheses (including two candidates in Phase 4 trials) prioritized for experimental follow-up. COPD, chronic obstructive pulmonary disease.

Disease Total prioritized therapeutic indications Recovered known indications Prospective prioritized therapeutic indications Clinically validated (%) Novel repurposing drug candidates
Alzheimer’s disease 7 5 2 100 0
Parkinson’s disease 28 25 3 67 1
Schizophrenia 46 34 12 58 5
COPD 18 12 6 83 1

Total 99 76 23 69.6 (average) 7

To evaluate the predictive power of FLASH on associations absent from its training data, we then isolated 23 prospective prioritized therapeutic indications (2 for AD, 3 for PD, 12 for Schizophrenia, and 6 for COPD; Supplementary Table 14). These network-blind out-of-KG predictions represent drug–disease pairs that were not encoded as indication edges within SIGMA-KG.

Retrospective external clinical validation against global regulatory approval records revealed that the majority (69.6%) of these network-blind predictions correspond to established clinical approvals (Fig. 5f). By systematically searching regulatory databases and clinical registries (e.g., FDA, European Medicines Agency, Medicines and Healthcare products Regulatory Agency; Methods), we found that 16 of the 23 prospective prioritized indications (69.6%) have already received regulatory approval for their respective diseases (Supplementary Table 14). For instance, FLASH correctly identified the acetylcholinesterase inhibitor ipidacrine for AD and the bronchodilator formoterol for COPD, despite these associations being missing from the initial knowledge graph. The high concordance between FLASH’s network-blind predictions and independent regulatory approvals demonstrates the model’s capacity for high-fidelity mechanistic discovery, validating it as a robust framework for bridging systems-level computational insights with real-world therapeutic applications.

The remaining 7 novel repurposing drug candidates represent high-priority therapeutic hypotheses generated by the model. Notably, two of these novel candidates for schizophrenia are currently under active Phase 4 clinical investigation for the disease (NCT02593734 and NCT00485823; Supplementary Table 15). These findings underscore the utility of sign-aware foundation models in surfacing non-obvious therapeutic opportunities that are both mechanistically grounded and clinically plausible, providing a scalable pipeline for accelerating drug repurposing.

To verify that these candidates arise from biologically plausible mechanisms rather than opaque heuristics, we evaluated the mechanistic interpretability of FLASH by tracing five distinct multi-hop signed metapaths within SIGMA-KG. This approach identifies the gene and protein intermediaries that constitute the signaling cascades for each novel drug–disease indication, providing a transparent molecular basis for the model’s predictions (Fig. 5c; Methods). These signed metapaths are intrinsically designed to satisfy structural balance principles, ensuring that the inferred regulatory cascades are mechanistically consistent with biological logic (Supplementary Fig. 1d). For example, FLASH identified a plausible path for COPD (Drug → Protein → Gene → Disease) in which formoterol downregulates the beta-3 adrenergic receptor (encoded by ADRB3), a gene specifically upregulated in the COPD disease state. By aggregating these intermediate nodes, we identified disease-specific mechanistic target modules that exhibit high biological and functional coherence (a full list of the targets is provided in Supplementary Table 16). Comprehensive pathway and protein-protein interaction (PPI) enrichment analyses further confirmed that these target modules are significantly associated with established disease pathophysiology (Supplementary Fig. 5a–h; details provided in the Supplementary Information). Together, these results show that our graph foundation model provides not only high-confidence therapeutic prioritization but also a plausible and traceable biological rationale for each drug candidate.

3. Discussion

The transition from associative link prediction to sign-aware mechanistic modeling represents a critical shift in systems pharmacology. While traditional graph frameworks successfully capture topological proximities, they frequently fail to resolve the functional polarity (activation/indication versus inhibition/contraindication) that governs biological regulatory systems. Our results show that sign-awareness within FLASH foundation model is essential for accurately resolving the sign composition (transitive logic) that defines multi-step biological signaling. This theoretical advantage is empirically supported by 31 evaluation metrics across three key downstream tasks. Notably, statistical comparisons reveal that the top-performing signed architecture significantly outperformed unsigned models in 27 of the 31 metric evaluations and surpassed relational baselines in 19 instances, highlighting the robust and consistent superiority of sign-aware modeling. By embedding sign-awareness at the representation stage, we mitigate the polarity drift inherent in unsigned or purely relational pre-training, ensuring that directional semantics are preserved across diverse biological scales

A core strength of this framework lies in the alignment between Heider’s structural balance theory and biological regulatory motifs. This principle was empirically validated through a systematic drug repurposing screen across four complex diseases: AD, PD, schizophrenia, and COPD. Utilizing a frozen-encoder linear probing protocol, FLASH prioritized 23 network-blind therapeutic candidatesn (Drug–Disease pairs) that were entirely absent from the SIGMA-KG pre-training atlas. Retrospective cross-referencing against global regulatory registries (e.g., FDA and EMA) revealed that 69.6% of these prospective predictions correspond to established clinical approvals, such as ipidacrine for AD and formoterol for COPD (Methods). The predictive power of the model is further underscored by the identification of two candidates for schizophrenia, which are currently under Phase 4 clinical investigation (NCT02593734 and NCT00485823). These results highlights the capacity of sign-aware reasoning to uncover therapeutic overlaps that are mechanistically grounded yet invisible to unsigned methods.

The architectural efficiency of FLASH resolves the historical trade-off between model expressiveness and computational scalability. By utilizing a localized seven-motif aggregation strategy rather than exhaustive higher-order enumeration, FLASH approximates complex relational dependencies with linear scalability. This algorithmic innovation allowed for pre-training on a 3.8-million-edge atlas without the memory-exhaustion failures observed for the signed model SNEA in our largest configurations (Supplemental Information). This establishes FLASH as a scalable transferable foundation model, capable of performing inductive inference on out-of-distribution chemical entities (cold-start DDI prediction; Results) without the need for the task-specific fine-tuning required by predecessors of unsigned GNNs such as TxGNN [5] and signed GNNs such as SIGDR [20].

Beyond predictive accuracy, our framework enables a transparent-box approach to drug discovery. By tracing the algebraic product of signs along multi-hop regulatory paths within SIGMA-KG, FLASH provides testable mechanistic rationales that link molecular perturbations to systemic clinical outcomes. We identified disease-specific target modules that act as molecular intermediaries, linking 7 novel drug candidates to their respective clinical phenotypes. These modules exhibit significant functional inter-connectedness and pathway enrichment which confirms that our mechanistic rationales align with established pathophysiology rather than random topological associations. Such transparency is directly applicable to rational polypharmacy, where signed paths can identify drug combinations that act synergistically on a therapeutic target while neutralizing potential antagonistic effects.

Despite these advances, several limitations remain. The signed edges in SIGMA-KG represent consensus-based, context-averaged effects and may not capture the dynamic sign-reversals seen across different tissues, doses, or disease stages. Furthermore, the reliance on cell-line-derived transcriptomic data introduces a degree of ascertainment bias that may limit direct translation to in vivo human physiology. Additionally, the negative sampling strategy used in three-class benchmarking (Results), while used as standard practice [25], assumes that unobserved associations are true negatives, an assumption that potentially introduces label noise given the inherent incompleteness of current pharmacological databases. Finally, while SIGMA-KG significantly expands chemical coverage, the utility of the model remains constrained by the availability of high-quality transcriptomic signatures for novel chemical series.

Future efforts should prioritize the experimental validation of the novel candidates identified here, such as vigabatrin for PD (Supplementary Table 13), to bridge the gap between computational priority and clinical application. Expanding the SIGMA-KG atlas to include mutations (DNA variants), cell-types and tissues as biomedical entities, and also to incorporate single-cell annd bulk transcriptomics will be essential for moving toward cell-type-resolved and patient-stratified discovery. Ultimately, SIGMA-KG and FLASH provide a scalable, sign-aware foundation for mechanistically informed drug discovery, and we anticipate that the open release of these resources will facilitate more robust predictive modeling in systems pharmacology.

4. Methods

4.1. Signed multi-omics KG curation

SIGMA-KG was constructed by harmonizing and integrating data from 16 curated biomedical resources (Supplementary Table 1) into a unified signed directed heterogeneous KG. It was curated to provide a unified substrate for sign-aware representation learning across molecular and clinical layers. SIGMA-KG comprises 127,422 nodes spanning four biomedical entity types: compounds (76,009 nodes; 59.7%), genes (19,245 nodes; 15.1%), proteins (18,609 nodes; 14.6%), and Diseases/Clinical Phenotypes (13,559 nodes; 10.6%). The KG contains 3,887,829 edges representing twelve distinct relation types organized across inter-layer and intra-layer connections (Table 1, Supplementary Table 2). Critically, each edge carries a signed label (+1/−1) reflecting the polarity of the underlying biological effect. Fig. 1a and Supplementary Fig. 1 present the SIGMA-KG schema and summary statisics.

Identifier harmonization.

Entity identifiers were harmonized across resources using standard ontologies: HGNC symbols for genes, UniProt accession numbers for proteins, PubChem Compound Identifiers (CIDs) for compounds, and Human Phenotype Ontology (HPO) for diseases/clinical phenotypes. Disease, phenotype and side-effect terms were consolidated using SapBERT biomedical LLM embeddings [27]; term pairs with cosine similarity exceeding 0.8 were merged, reducing 69,207 initial terms to 13,559 unified disease nodes (Supplemental Information).

Data sources and integration.

Briefly, gene–disease transcriptional dysregulation edges were derived from AMP-AD [28], DiSignAtlas [29], GEO differential gene signatures [30], and Hetionet [6]; compound–gene perturbation edges were derived from LINCS L1000 Connectivity Map consensus signatures [31]; protein-protein interactions were obtained from STRING [32]; compound–disease indication/contraindication edges were integrated from PrimeKG [7], SIDER [33], OnSIDES [34], and repoDB [35]; disease–disease similarity edges were taken from PrimeKG [7]; gene–protein encoding edges were taken from UniProtKB [36]; and compound–protein activation/inhibition edges were taken from IUPHAR [37] and STITCH [38]. Compound–compound similarity edges were derived from the Chemical Checker [39] structural similarity level (A1), which harmonizes data from primary repositories including ChEMBL and PubChem. Complete preprocessing steps and filtering criteria are provided in Supplementary Information.

Edge construction and sign assignment.

Each directed edge was represented as a quadruple (source entity, relation type, sign, target entity), where the sign (+1 or −1) encoded the polarity of the underlying biological or clinical effect. Positive signs were assigned to: upregulated-in (gene→disease), upregulates (compound→gene), similar-to (compound↔compound), interacts-with (protein↔protein), resembles (disease↔disease), encodes (gene→protein), activates (compound→protein), and contraindication (compound→disease). Negative signs were assigned to: downregulated-in (gene→disease), downregulates (compound→gene), inhibits (compound→protein), and indication (compound→disease). This convention encodes contraindications (adverse associations) as positive and indications (therapeutic benefit) as negative, preserving semantics consistent with structural balance theory (Supplementary Fig. 1d). For instance, if a compound upregulates (+1) a gene that is downregulated (−1) in a disease, the inferred compound–disease relationship would be negative (−1), consistent with a therapeutic indication.

Signed network balance semantics.

To ensure the mechanistic and semantic consistency of the integrated multi-omics atlas, we enforced the principles of Heider’s structural balance theory throughout the SIGMA-KG framework. Following identifier harmonization and sign-aware edge construction, we performed a systematic topological refinement to resolve contradictory regulatory paths. Within this signed directed graph, a cycle is defined as structurally balanced if the product of its constituent edge signs is positive:

∏e∈Csgn(e)>0

where C represents the set of edges in a cycle (e.g., a ”double negative” resulting in net activation). Conversely, cycles yielding a negative product, representing ”frustrated” or biologically inconsistent relations, were resolved through a targeted pruning strategy. For every identified unbalanced cycle, we evaluated the constituent edges based on their source-specific confidence scores (e.g., z-scores from differential expression or curation evidence levels). We preferentially pruned the edge with the lowest confidence score. This pruning was primarily targeted at gene–disease or compound–gene associations, which are inherently noisier than high-confidence compound–disease or gene–protein encoding relations. This optimization step ensures that the resulting atlas provides a consistent substrate for sign-aware graph representation learning, mitigating the propagation of contradictory signals during model pre-training.

4.2. FLASH model architecture and pre-training strategy

To enable scalable representation learning over SIGMA-KG, we developed FLASH, a fast and lightweight signed heterogeneous GNN architecture designed to balance expressive power and computational efficiency.

Unlike conventional unsigned or relational GNNs, FLASH explicitly models edge polarity during message passing. The architecture is inspired by Signed Graph Attention Networks (SiGAT) [40], which extend attention mechanisms to signed directed graphs by distinguishing positive and negative neighborhoods. However, directly applying SiGAT to SIGMA-KG is computationally expensive due to the atlas-scale size of the network. We therefore redesign the aggregation strategy to suit the structural and biological properties of SIGMA-KG.

Signed neighborhood modeling for SIGMA-KG.

While SiGAT exhaustively enumerates dozens of signed motifs to capture higher-order neighborhood patterns of each node, such motif modeling scales poorly for large heterogeneous biomedical graphs like SIGMA-KG. To address this challenge, FLASH defines biologically meaningful motifs in the first-order neighborhood, chosen to preserve signed semantics while ensuring scalability: positive/negative neighbors (2 motifs), positive incoming/outgoing neighbors (2 motifs), negative incoming/outgoing neighbors (2 motifs), and positive bidirectional neighbors (1 motif).

The first two motifs capture polarity-specific aggregation without direction. The next four motifs encode polarity together with directionality, which is essential for mechanistic explainability in biological signaling chains (for example, Compound → Gene → Disease). The final motif captures edges such as similarity relations that are semantically positive and have bi-directionality (e.g. protein–protein and compound–compound edges). This seven-motif design preserves the essential signed and directional semantics required for structural balance reasoning while substantially reducing computational complexity compared to exhaustive motif enumeration. This curated seven-motif set serves as the minimal sufficient basis for the model to learn structural balance principles, enabling it to resolve global consistency across the signed multi-omics landscape.

Motif-aware attention-based aggregation.

Let G = (V, E) denote SIGMA-KG. For each node μ ∈ V, we initialize its feature representation as X(u).

Let ℳ denote the motif set, where |ℳ| = 7. For the i-th motif mi ∈ ℳ, we define a motif-aware neighborhood of node u:

𝒩mi(u)={v∣(u,v)∈Esatisfies motifmi}. (1)

We use the attention-based aggregator in GAT [41] to compute the motif-aware attention coefficient αuvm between node u and node v in 𝒩m(u):

αuvmi=exp(LeakyReLU(ami⊤[WmiX(u)‖WmiX(v)]))∑k∈𝒩mi(u)exp(LeakyReLU(ami⊤[WmiX(u)‖WmiX(k)])), (2)

where ami is a weighted vector, Wmi is a weighted matrix.

The motif-aware attention-based aggregation is then defined as:

Xmi(u)=∑v∈𝒩mi(u)αuvmiWmiX(v), (3)

Obtain final node embedding.

After computing motif-aware representations Xm(u) for all m ∈ ℳ, we concatenate them with the original node feature:

Xconcat (u)=CONCATX(u),Xm1(u),…,Xm7(u). (4)

The final node embedding is obtained through a two-layer feedforward network:

zu=W2σW1Xconcat(u)+b1+b2, (5)

where W1, W2 are learnable weight matrices, b1, b2 are bias terms, and σ is a sigmoid activation function.

Self-supervised pre-training ob jective.

FLASH is trained using a self-supervised objective, which models positive and negative relations via a signed proximity ob jective on the global topology of SIGMA-KG. Specifically, the model optimizes node embeddings by contrasting positive and negative neighborhoods.

For a node u, let 𝒩(u)+ and 𝒩(u)− denote its positive and negative neighbors, respectively. The objective encourages embeddings of positively connected nodes to be similar, while pushing apart those connected by negative edges. The training objective ℒ is defined as:

ℒ=−∑u[∑v∈𝒩(u)+log(σ(zu⊤zv))+λ∑v∈𝒩(u)−log(σ(−zu⊤zv))], (6)

where σ(·) is the sigmoid function, and λ is a balancing parameter to account for the imbalance between positive and negative edges.

This pre-training strategy explicitly enforces that positively connected nodes have similar embeddings while negatively connected nodes are dissimilar, thereby capturing signed structural patterns in the graph. As a result, the learned embeddings encode multi-hop sign propagation and structural balance properties through signed neighborhood aggregation. Importantly, no task-specific supervision (such as drug–disease labels) is used during pre-training, enabling FLASH to serve as a general-purpose signed graph foundation model transferable across molecular, clinical, and systems-level tasks.

4.3. Baselines and fair comparison protocol

To contextualize the performance of FLASH, we benchmarked it against nine GNN methods spanning three architectural categories: (1) unsigned (homogeneous) models (BGRL [42], DGI [43], InfoGraph [44], GraphSAGE [45]); (2) relational heterogeneous models (HGATE [46], DMGI [47]); and (3) signed graph representation learning models (SGCN [48], SNEA [49], SDGNN [50]). All baselines were implemented within a self-supervised pre-training framework to generate task-agnostic latent node representations for downstream evaluation. To ensure a fair and optimized comparison, each model was trained using the specific hyperparameter configurations recommended by their original authors (Supplementary Table 17).

To ensure a fair and rigorous comparison, all models were initialized with identical 256-dimensional node features derived from a truncated SVD of the adjacency matrix. We standardized the output at 256-dimensional node embeddings across all baselines; this dimensionality was selected as a necessary computational ceiling, as several complex architectures (specifically SNEA, SDGNN, and HGATE) consistently encountered memory-exhaustion failures on the 3.8M-edge SIGMA-KG atlas when attempting to generate 512-dimensional embeddings.

The information accessible to each model was strictly controlled to isolate the impact of signed multi-relational structures. Unsigned baselines were trained on an unsigned version of SIGMA-KG with all edge types and polarities discarded. Relational baselines utilized the full 12-relation multiplex structure, while signed baselines and FLASH received the signed network, with edges labeled as +1 or −1.

Model evaluation was conducted using a dual-stratified protocol across three independent replicates with matched random seeds to ensure statistical reproducibility. For molecular (Ts-MoA) and clinical (DCR) tasks, we employed an internal held-out validation strategy. Specifically, SIGMA-KG was partitioned into a training subgraph and an independent set of evaluation edges (compound–protein or compound–disease). Each foundation model was first pre-trained in a self-supervised manner on the training subgraph to generate node embeddings . Subsequently, an MLP classifier (binary and three-class) was trained on the training edges using these frozen embeddings and then evaluated on the held-out test set.

For systems-level DDI prediction, we implemented a more stringent external validation protocol. Baseline models were pre-trained on the complete SIGMA-KG atlas, and the resulting compound embeddings were directly transferred to two independent, unseen DDI benchmark datasets: DrugBank (version 5.1) [23] and MUDI [24]. To eliminate downstream classifier variance as a confounding factor, we employed a standardized MLP architecture and identical hyperparameters for FLASH and all baseline counterparts across every task (Supplementary Table 18). By standardizing the classifier and using frozen embeddings, we ensure that observed performance deltas are strictly attributable to the architectural inductive biases of the foundation models rather than variance in downstream supervised learning.

4.4. Downstream evaluation for edge-sign prediction via frozen-encoder linear probing

To rigorously assess the predictive utility of the FLASH foundation model, we implemented a standardized benchmarking framework across two critical pharmacological scales: molecular-level Ts-MoA and clinical-level DCR. Both tasks were formulated as edge-sign prediction problems using a frozen-encoder linear probing protocol. This design isolates the quality of the learned representations from the complexity of the downstream classifier, ensuring that performance gains are directly attributable to the sign-aware inductive biases of the foundation model.

4.4.1. Unified benchmarking protocol

For both the Ts-MoA and DCR tasks, we followed a consistent stratified evaluation pipeline to ensure that the foundation model remained entirely naive to evaluation data during pre-training:

  • Task-specific graph construction and data leakage prevention: Prior to self-supervised pre-training, all labeled edges for the target task (e.g., Compound–Protein or Compound–Disease) were randomly partitioned into training, validation, and test sets using a stratified split (0.8/0.1/0.1). To prevent information leakage, a task-specific pre-training KG was constructed by starting from the full SIGMA-KG and removing the validation and test edges corresponding to the specific task. The training edges and all remaining relation types in sigma-kg were retained. Initial node features (truncated SVD) were computed exclusively from this restricted pre-training topology.

  • Linear probing architecture: Following pre-training, node embeddings were extracted and frozen. For any candidate edge (u, v), a 512-dimensional feature vector was generated by concatenating the embeddings of the source and target nodes [zu ∥ zv]. This vector served as the input for a standardized MLP head. The classifier was optimized using Binary Cross-Entropy (BCE) loss for binary classification tasks and Categorical Cross-Entropy for three-class evaluations. For each task, the MLP was trained on the training edge split, with early stopping based on validation-set performance (Macro-F1). Final performance was measured exclusively on the 10% held-out test set.

  • Three-class open-world evaluation: To simulate realistic discovery settings, we introduced a ”non-existent” class (0) by uniformly sampling node pairs (compound–protein or compound–disease) absent from SIGMA-KG. Following established protocols [25], we maintained a 5:1 ratio of non-existent edges to signed edges. The MLP was thus trained to distinguish between three classes: activation/contraindication (+1), inhibition/indication (−1), and non-existent (0).

  • Replication and evaluation: To ensure statistical rigor, the entire pipeline, including graph partitioning, pre-training, and MLP evaluation, was repeated over three independent runs (matched random seeds for stratified edge split). Final metrics are reported as the mean ± standard deviation of AUROC, AUPRC, Accuracy, and Macro-F1 scores across these runs.

4.4.2. Ts-MoA prediction

The Ts-MoA task evaluates the model’s ability to distinguish the functional polarity of drug–target engagement. We labeled compound–protein edges based on their functional effect, where activation/agonist corresponds to a positive sign (+1) and inhibition/antagonist corresponds to a negative sign (−1). The dataset comprised 13,711 high-confidence interactions (6,389 activation and 7,322 inhibition), providing a balanced distribution for evaluating mechanistic reasoning. Binary and three-class classification results are detailed in Supplementary Table 5 and Supplementary Table 6, respectively.

4.4.3. DCR prediction

At the clinical scale, we evaluated the model’s capacity to bridge molecular perturbations with clinical phenotypes. Consistent with the signed semantics of SIGMA-KG, compound–disease edges were labeled as contraindication (+1; adverse association) or indication (−1; therapeutic association). The DCR dataset represents a severely imbalanced setting, containing 11,896 indication and 148,061 contraindication edges (approx. 1 : 12.4 ratio). To address this, the MLP loss function was class-weighted inversely to the frequencies observed in the training split. Results are reported in Supplementary Table 7 and Supplementary Table 8.

4.5. Drug-drug interaction prediction on external resources

We evaluated whether pre-trained embeddings generalize to DDIs absent from SIGMA-KG, using two external benchmark datasets whose interaction pairs were entirely unseen during pre-training.

DDI benchmark datasets.

Two benchmark datasets were used: DrugBank (version 5.1) [23] and MUDI [24]. The DrugBank dataset, manually curated from FDA and Health Canada drug labels and primary literature, annotates pairwise interactions across 113 relation types capturing metabolic, therapeutic, and adverse effect relationships. The MUDI dataset, constructed from a pharmacodynamic perspective using pharmacological text and multimodal chemical representations, annotates interactions as one of three clinically meaningful categories: Synergism, Antagonism, or New Effect. The unprocessed DrugBank dataset contained 222,696 interacting pairs across 1,876 compounds; the unprocessed MUDI dataset contained 310,532 interacting pairs across 1,198 compounds.

Pre-processing.

Compound identifiers were harmonized to PubChem CIDs via the PubChem Identifier Exchange Service (Supplementary Methods). Self-loop edges (a compound paired with itself ) were removed, and pairs with multiple interaction-type annotations were deduplicated to a single binary interaction label. To prevent inverse-pair leakage across data splits, edges were treated as unordered: each pair was canonicalized as (min{ci, cj}, max{ci, cj}) by sorting compound identifiers. Subsequently, only drug pairs for which both compounds were present in SIGMA-KG were retained. After preprocessing, DrugBank dataset comprised 195,967 interacting pairs across 1,645 compounds; MUDI dataset comprised 274,510 pairs across 1,129 compounds. For MUDI, the original train/test split was discarded in favor of the unified protocol below.

Task formulation.

Although both benchmark datasets provide multi-class interaction annotations, we formulated DDI prediction as a binary edge classification task (a pair of drugs interact vs. does not interact) for three reasons: (i) to maintain a consistent evaluation protocol across resources with different and non-aligned label taxonomies, and (ii) because the primary objective was to assess whether pre-trained embeddings capture general pharmacological relatedness rather than to discriminate among specific interaction mechanisms and (iii) both benchmarks exhibit severe class imbalance: in DrugBank, the three most prevalent interaction types comprise 66% of all pairs, whereas the majority of categories each constitute less than 1%; in MUDI, the Synergism type alone accounts for 83% of all annotations. This extreme distribution skew renders multi-class evaluation statistically unreliable for minority classes, justifying our focus on binary classification protocols.

Embedding extraction and negative sampling.

Each baseline was trained on the full SIGMA-KG in a self-supervised manner to obtain 256-dimensional node embeddings, which remained frozen during downstream evaluation. For each external dataset, only embeddings for drugs present in that preprocessed dataset were used. Positive edges corresponded to annotated interacting drug pairs from each dataset; negative edges were generated by randomly sampling unordered drug pairs– from the set of drugs present in each preprocessed dataset–that were not annotated as interacting, at a 1:1 ratio relative to positive pairs. This negative-sampling strategy implicitly assumes that unannotated pairs are non-interacting–a standard but imperfect assumption that may introduce label noise, particularly for incompletely curated resources.

Transductive (warm-start; random pair split) evaluation scenario.

For each dataset and each random seed, the combined positive and negative edge set was randomly partitioned into training/validation/test splits (0.8/0.1/0.1, stratified by class). As splits are edge-wise random, drugs may appear in both training and test sets (transductive setting). Each candidate pair (ci, cj) was represented by concatenating embeddings in canonical order, yielding a 512-dimensional feature vector [zmin(ci, cj) ∥ zmax(ci, cj)]. A binary MLP classifier (minimizing binary cross-entropy, with early stopping on validation Macro-F1) was trained to predict interaction status (interacting versus non-interacting); architecture and hyperparameters were held constant across baselines (Supplementary Table 18).

Inductive (cold-start) evaluation scenario.

To assess generalization to unseen drugs, we implemented a single-endpoint cold-start protocol. For each dataset and each random seed, the unique drug set D was partitioned into disjoint subsets Dtrain, Dval, and Dtest (0.8/0.1/0.1). Training positives comprised interactions between drugs both in Dtrain. Validation positives comprised interactions in which at least one endpoint belonged to Dval and the other to Dtrain ∪ Dval. Test positives comprised interactions in which at least one endpoint belonged to Dtest and the other to Dtrain ∪ Dtest. For each split, negatives pairs were sampled from the corresponding drug universe at a 1:1 ratio, sub ject to the same unannotated-equals-negative assumption. This design ensures that every validation and test edge involves at least one drug absent from classifier training, approximating the practical scenario of predicting interactions for a newly introduced compound. MLP architecture and hyperparameters were held constant across baselines (Supplementary Table 18).

Replication and evaluation.

For each baseline and dataset, data partitioning, negative sampling, and MLP training/evaluation were repeated across three independent random seeds; mean ± standard deviation are reported. The same seeds were used across baselines to ensure matched comparisons. For a given setting, the MLP architecture and hyperparameters were held constant across all baselines (Supplementary Table 4). Evaluation metrics included AUROC, AUPRC, accuracy, macro-F1, and class-specific F1. Full results for the transductive and inductive evaluation scenarios are reported in Supplementary Tables 9 and 10, and Supplementary Tables 11 and 12, respectively.

4.6. Computational efficiency and scalability benchmarking

To evaluate the computational efficiency and scalability of FLASH, we benchmarked its training time and peak GPU memory usage against nine GNN baselines spanning three architectural categories. All experiments report training time and peak GPU memory consumption measured over 100 epochs (Supplemental Information; Supplementary Fig. 4).

Subnetwork construction.

Scalability was assessed on subnetworks sampled from SIGMA-KG with systematically varied sizes. For node-scaling experiments, we varied the number of nodes from 20,000 to 120,000 while maintaining a constant average degree of 30.51 (matching the global average degree of SIGMA-KG). For edge-scaling experiments, we fixed the number of nodes at 100,000 and varied the number of edges from 500,000 to 2,500,000. This design isolates the impact of graph order (node count) and graph size (edge count) on computational cost.

Training protocol.

All models were trained in a self-supervised manner on each subnetwork using identical initial node features (128-dimensional truncated SVD of the adjacency matrix), the same number of epochs (100), and produced 128-dimensional node embeddings. Homogeneous baselines were trained on unsigned versions of the subnetworks in which edge signs and relation types were discarded.

Hardware environment.

All experiments were conducted on a Linux workstation equipped with dual AMD EPYC 7453 28-core processors (112 logical cores, up to 3.49 GHz), 2 TiB of RAM, and an NVIDIA L40S GPU (48 GB). Each training run was executed on a single dedicated GPU with no concurrent jobs to ensure consistent timing measurements.

4.7. Drug repurposing

4.7.1. Case studies

Disease selection.

To evaluate the translational utility of the foundation model, we applied FLASH to identify novel therapeutic indications for five diseases representing diverse pathophysiological mechanisms: AD (neurodegenerative), PD (neurodegenerative), schizophrenia (psychiatric), COPD (respiratory). This selection was designed to assess generalization across distinct therapeutic areas and molecular etiologies.

Drug candidate universe.

To focus the analysis on drugs with prior regulatory approval, the candidate drug universe was restricted to the 2,869 compounds in SIGMA-KG with at least one established clinical association. We systematically screened this universe against each target disease to prioritize candidate drugs for novel therapeutic indications.”

Edge representation and MLP classifier setup.

Drug repurposing was formulated as a three-class edge classification task ( 5a). We utilized 256-dimensional FLASH node embeddings, pre-trained on the full SIGMA-KG in a self-supervised manner, as frozen input features. For any given drug-disease pair, an edge representation was constructed by concatenating the drug embedding zc and the disease embedding zd into a 512-dimensional feature vector [zc ∥zd]. These vectors were fed into a MLP with a softmax output layer to predict three classes: (i) indication (therapeutic association; label −1), (ii) contraindication (adverse association; label +1), and (iii) non-existent (no known association; label 0).

Training strategy and class imbalance mitigation.

We implemented a stratified five-fold CV strategy. The training and validation sets were constructed from the 159,957 existing drug-disease edges in SIGMA-KG (labeled as indication or contraindication). To represent the non-existent class, we performed negative sampling by randomly pairing drug nodes (from the pool of 2,869) and disease nodes (from the pool of 13,559) that lacked any established association in SIGMA-KG. The number of sampled non-existent edges was set to five times the number of labeled edges; these were then uniformly distributed across the five folds. Notably, the ratio of indication to contraindication edges in SIGMA-KG is approximately 1:12.4. To account for this substantial class imbalance, the MLP was trained using a weighted cross-entropy loss function, with class weights set inversely proportional to their respective frequencies in the training data. MLP architecture and hyperparameter details are provided in Supplementary Methods.

Disease-specific screening and test set construction.

For each of the five target diseases, a separate screening pipeline was executed. While the model architecture and the global training/validation folds remained consistent across all five pipelines to ensure a generalized understanding of pharmacological patterns, the test set differed according to the target disease. Specifically, the test set for a given disease (e.g., AD) consisted of all 2,869 candidate drug–disease edges formed by pairing that disease node with every drug in the candidate universe—yielding edges such as AD–Drug1, AD–Drug2, … , AD–Drug2869 . This process was repeated independently for each target disease (AD, PD, Schizophrenia, and COPD).

Cross-validation and hyperparameter optimization.

During each fold of the CV, the MLP was trained on four folds (with the validation fold used for early stopping based on macro-F1) and then applied to the disease-specific test set, outputting softmax probabilities for each of the three classes for every test edge. Consequently, each testing edge (e.g., AD-Drug X) received five separate softmax probability distributions. MLP hyperparameters were tuned via nested five-fold CV on the training folds. Architecture and hyperparameter details are provided in Supplementary Table 18.

Prioritized therapeutic indications.

We calculated the final confidence score for each candidate indication by averaging the “indication” class probabilities across all five CV folds. Edges with a mean probability P (indication) > 0.5 were retained as prioritized therapeutic indications.

prospective prioritized therapeutic indications.

To ensure the discovery of novel repurposing candidates, we applied a rigorous filtering procedure: any predicted indication edge (drug–disease) already present in the SIGMA-KG (recovered known indications) was excluded from the list of prioritized therapeutic indications. Only network-blind predictions, drug–disease pairs absent from the pre-training KG, were retained as prospective prioritized therapeutic indications.

External clinical validation of prospective prioritized therapeutic indications.

Retrospective clinical validation was performed by cross-referencing predicted drug–disease pairs against regulatory approval records from FDA, European Medicines Agency (EMA), Medicines and Healthcare products Regulatory Agency (MHRA), and State Agency of Medicines of the Republic of Latvia (SAM), using standardized drug and disease identifiers. Validation was conducted in a systematic manner. A drug–disease pair was considered validated if the drug is officially approved for the indication in at least one of the referenced databases using standardized identifiers.

4.7.2. Mechanistic explainability and target inference

Metapath-based target identification.

To interpret the molecular basis of the predictions, we identified intermediate nodes connecting predicted drugs to their respective target diseases through SIGMA-KG. We investigated five specific multi-hop metapaths based on shortest-path connectivity:

  1. Drug → Gene → Disease

  2. Drug → Protein → Gene → Disease

  3. Drug → Drug → Gene → Disease

  4. Drug → Drug → Protein → Gene → Disease

  5. Drug → Drug → Protein → Protein → Gene → Disease

Pathway and PPI enrichment analysis.

Functional enrichment analysis was performed on the aggregated gene and protein targets for each disease using the STRING database [32]. Pathway enrichment was assessed with a FDR threshold of < 0.0001. PPI enrichment analysis was performed to evaluate the biological coherence of the target sets. Significant PPI enrichment p-value indicates that the identified targets form a functionally interconnected module rather than a random collection of proteins, supporting the mechanistic plausibility of the foundation model’s predictions.

4.8. Statistical Analysis

To rigorously assess the significance of the performance gains achieved by FLASH, all experiments were executed across three independent replicates (N = 3) using distinct random initializations to ensure reproducibility. To maintain a fair comparison, the same three random seeds were applied across all baselines for data partitioning and stochastic sampling. Statistical significance was evaluated using the unpaired Welch’s t-test, which is robust to potential differences in variance between model performances.

We implemented a two-tiered comparative analysis to evaluate the results:

  1. Inter-class comparison: We compared the top-performing signed baseline against the best-performing unsigned and relational baselines using a one-sided unpaired Welch’s t-test. This analysis was designed to specifically validate the inductive bias afforded by explicit edge polarity across different tasks (Supplementary Table 3).

  2. SOTA validation: For each metric, we conducted a direct comparison between FLASH and the runner-up baseline (in cases where FLASH established SOTA results) or the top-performing baseline (if FLASH matched existing performance). These comparisons utilized a two-sided unpaired Welch’s t-test to establish statistical superiority or parity (Supplementary Table 4).

Statistical significance was predefined as P < 0.05. Results are reported as mean ± standard deviation (s.d.). Where FLASH significantly outperformed the leading baselines, significance is denoted using a hierarchical notation: * for P < 0.05, ** for P < 0.01, and n.s. (not significant) for P ≥ 0.05. In instances where differences did not reach the significance threshold, FLASH is described as achieving performance comparable to existing SOTA methods. All statistical computations and distribution analyses were performed using SciPy (v1.12.0) [51]

Supplementary Material

Supplement 1
media-1.xlsx (27.5KB, xlsx)
Supplement 2
media-2.xlsx (44.6KB, xlsx)
Supplement 3
media-3.xlsx (17.2KB, xlsx)
4

Funding

This project has been funded with federal funds from the National Institute of General Medical Sciences of the National Institute of Health (grant no. R01GM122845, Lei Xie), the National Institute on Aging of the National Institute of Health (grant no. R01AG057555, Lei Xie; grant no. R33AG083302, Lei Xie) and the National Science Foundation (grant no. 2226183, Lei Xie).

Footnotes

Competing interests

The authors declare that they have no competing interests.

Data availability

All datasets utilized in this study are publicly available.

References

  • [1].Jiang W., Ye W., Tan X., Bao Y.-J.: Network-based multi-omics integrative analysis methods in drug discovery: a systematic review. BioData Mining 18 (2025) 10.1186/s13040-025-00442-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [2].Wang Z., Wei Z.: Pt-kgnn: A framework for pre-training biomedical knowledge graphs with graph neural networks. Computers in Biology and Medicine 178, 108768 (2024) 10.1016/j.compbiomed.2024.108768 [DOI] [PubMed] [Google Scholar]
  • [3].Molaee M., Charkari N.M., Ghaderi F.: Heterogeneous networks in drug-target interaction prediction (2025). https://arxiv.org/abs/2504.16152 [Google Scholar]
  • [4].Wang Y., Yang Z., Yao Q.: Accurate and interpretable drug-drug interaction prediction enabled by knowledge subgraph learning. Communications Medicine 4(1), 59 (2024) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [5].Huang K., Chandak P., Wang Q., Havaldar S., Vaid A., Leskovec J., Nadkarni G., Glicksberg B., Gehlenborg N., Zitnik M.: A foundation model for clinician centered drug repurposing. Nature Medicine (2023) 10.1101/2023.03.19.23287458 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [6].Himmelstein D.S., Lizee A., Hessler C., Brueggeman L., Chen S.L., Hadley D., Green A., Khankhanian P., Baranzini S.E.: Systematic integration of biomedical knowledge prioritizes drugs for repurposing. elife 6, 26726 (2017) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Chandak P., Huang K., Zitnik M.: Building a knowledge graph to enable precision medicine. scientific data 10, 1 (2023), 67. URL 10.1038/s41597-023-01960-3 (2023) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Hu E.Y., Oleshko S., Firmani S., Cheng H., Zhu Z., Ulmer M., Arnold M., Colomé-Tatché M., Tang J., Xhonneux S., et al. Enhancing link prediction in biomedical knowledge graphs with biopathnet. Nature Biomedical Engineering, 1–23 (2026) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Lo Surdo P., Iannuccelli M., Contino S., Castagnoli L., Licata L., Cesareni G., Perfetto L.: Signor 3.0, the signaling network open resource 3.0: 2022 update. Nucleic acids research 51(D1), 631–637 (2023) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Türei D., Schaul J., Palacio-Escat N., Bohár B., Bai Y., Ceccarelli F., Ç evrim E., Daley M., Darcan M., Dimitrov D., et al. Omnipath: integrated knowledgebase for multi-omics analysis. Nucleic Acids Research 54(D1), 652–660 (2026) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Han H., Cho J.-W., Lee S., Yun A., Kim H., Bae D., Yang S., Kim C.Y., Lee M., Kim E., et al. Trrust v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic acids research 46(D1), 380–386 (2018) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Yeang C.-H., Ideker T., Jaakkola T.: Physical network models. Journal of computational biology 11(2-3), 243–262 (2004) [DOI] [PubMed] [Google Scholar]
  • [13].Signorini L.F., Kupiec M., Sharan R.: Signing protein–protein interaction networks. Bioinformatics 42(1), 674 (2026) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Cartwright D., Harary F.: Structural balance: a generalization of heider’s theory. Psychological review 63(5), 277 (1956) [DOI] [PubMed] [Google Scholar]
  • [15].Allahyari N., Kargaran A., Hosseiny A., Jafari G.: The structure balance of gene-gene networks beyond pairwise interactions. Plos one 17(3), 0258596 (2022) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Cook F.A., Cook S.J.: Inhibition of raf dimers: it takes two to tango. Biochemical Society Transactions 49(1), 237–251 (2021) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Wu Y., Xie L.: Ai-driven multi-omics integration for multi-scale predictive modeling of genotype-environment-phenotype relationships. Computational and Structural Biotechnology Journal 27 (2025) 10.1016/j.csbj.2024.12.030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].You Y., Liu Z., Wen X., Zhang Y., Ai W.: Large language models meet graph neural networks: a perspective of graph mining. Mathematics 13(7), 1147 (2025) [Google Scholar]
  • [19].Zhang Z., Zhao P., Li X., Liu J., Zhang X., Huang J., Zhu X.: Signed graph representation learning: A survey. arXiv preprint arXiv:2402.15980 (2024) [Google Scholar]
  • [20].Cui H., Duan M., Yuan J., Ma Z., Zhang Y.: Sign-aware graph contrastive learning for drug repositioning. IEEE Journal of Biomedical and Health Informatics (2025) [DOI] [PubMed] [Google Scholar]
  • [21].Pan Y., Ji X., You J., Li L., Liu Z., Zhang X., Zhang Z., Wang M.: Csgdn: contrastive signed graph diffusion network for predicting crop gene–phenotype associations. Briefings in Bioinformatics 26(1), 062 (2025) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Wang Z., Wei Z.: Pt-kgnn: A framework for pre-training biomedical knowledge graphs with graph neural networks. Computers in Biology and Medicine 178, 108768 (2024) [DOI] [PubMed] [Google Scholar]
  • [23].Wishart D.S., Feunang Y.D., Guo A.C., Lo E.J., Marcu A., Grant J.R., Sajed T., Johnson D., Li C., Sayeeda Z., Assempour N., Iynkkaran I., Liu Y., Maciejewski A., Gale N., Wilson A., Chin L., Cummings R., Le D., Pon A., Knox C., Wilson M.: Drugbank 5.0: a major update to the drug-bank database for 2018. Nucleic Acids Research 46(D1), 1074–1082 (2017) 10.1093/nar/gkx1037 https://academic.oup.com/nar/article-pdf/46/D1/D1074/23162116/gkx1037.pdf [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Ngo T.-L., Tran B.-H., Can D.-C., Do T.-H., Chén O.Y., Le H.-Q.: Mudi: A multimodal biomedical dataset for understanding pharmacodynamic drug-drug interactions. In: Proceedings of the 33rd ACM International Conference on Multimedia. MM ’25, pp. 13243–13249. Association for Computing Machinery, New York, NY, USA (2025). 10.1145/3746027.3758281. https://doi.org/10.1145/3746027.3758281. [DOI] [Google Scholar]
  • [25].Lim H., Poleksic A., Yao Y., Tong H., He D., Zhuang L., Meng P., Xie L.: Large-scale off-target identification using fast and accurate dual regularized one-class collaborative filtering and its application to drug repurposing. PLoS computational biology 12(10), 1005135 (2016) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [26].Lu Y., Chen J., Fan N., Song W., Sheng H., Yang Y., Wang J.: Machine learning models for drug-drug interaction prediction from computational discovery to clinical application. npj Digital Medicine (2026) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Liu F., Shareghi E., Meng Z., Basaldella M., Collier N.: Self-Alignment Pre-training for Biomedical Entity Representations (2021). https://arxiv.org/abs/2010.11784 [Google Scholar]
  • [28].Hodes R.J., Buckholtz N.: Accelerating medicines partnership: Alzheimer’s disease (amp-ad) knowledge portal aids alzheimer’s drug discovery through open data sharing. Expert opinion on therapeutic targets 20(4), 389–391 (2016) [DOI] [PubMed] [Google Scholar]
  • [29].Zhai Z., Lin Z., Meng X., Zheng X., Du Y., Li Z., Zhang X., Liu C., Zhou L., Zhang X., et al. Disignatlas: an atlas of human and mouse disease signatures based on bulk and single-cell transcriptomics. Nucleic Acids Research 52(D1), 1236–1245 (2024) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [30].Barrett T., Wilhite S.E., Ledoux P., Evangelista C., Kim I.F., Tomashevsky M., Marshall K.A., Phillippy K.H., Sherman P.M., Holko M., et al. Ncbi geo: archive for functional genomics data sets—update. Nucleic acids research 41(D1), 991–995 (2012) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [31].Subramanian A., Narayan R., Corsello S.M., Peck D.D., Natoli T.E., Lu X., Gould J., Davis J.F., Tubelli A.A., Asiedu J.K., et al. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell 171(6), 1437–1452 (2017) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Szklarczyk D., Gable A.L., Lyon D., Junge A., Wyder S., Huerta-Cepas J., Simonovic M., Doncheva N.T., Morris J.H., Bork P., et al. String v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic acids research 47(D1), 607–613 (2019) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Kuhn M., Letunic I., Jensen L.J., Bork P.: The sider database of drugs and side effects. Nucleic acids research 44(D1), 1075–1079 (2016) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Tatonetti N.P., Ye P.P., Daneshjou R., Altman R.B.: Data-driven prediction of drug effects and interactions. Science translational medicine 4(125), 125–3112531 (2012) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Brown A.S., Patel C.J.: A standard database for drug repositioning. Scientific data 4(1), 1–7 (2017) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Uniprot: the universal protein knowledgebase in 2025. Nucleic Acids Research 53(D1), 609–617 (2025) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Harding S.D., Sharman J.L., Faccenda E., Southan C., Pawson A.J., Ireland S., Gray A.J., Bruce L., Alexander S.P., Anderton S., et al. The iuphar/bps guide to pharmacology in 2018: updates and expansion to encompass the new guide to immunopharmacology. Nucleic acids research 46(D1), 1091–1106 (2018) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Szklarczyk D., Santos A., Von Mering C., Jensen L.J., Bork P., Kuhn M.: Stitch 5: augmenting protein–chemical interaction networks with tissue and affinity data. Nucleic acids research 44(D1), 380–384 (2016) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39].Duran-Frigola M., Pauls E., Guitart-Pla O., Bertoni M., Alcalde V., Amat D., Juan-Blanco T., Aloy P.: Extending the small-molecule similarity principle to all levels of biology with the chemical checker. Nature Biotechnology 38(9), 1087–1096 (2020) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Huang J., Shen H., Hou L., Cheng X.: Signed graph attention networks. In: International Conference on Artificial Neural Networks, pp. 566–577 (2019). Springer [Google Scholar]
  • [41].Veličković P., Cucurull G., Casanova A., Romero A., Liò P., Bengio Y.: Graph Attention Networks (2018). https://arxiv.org/abs/1710.10903 [Google Scholar]
  • [42].Thakoor S., Tallec C., Azar M.G., Azabou M., Dyer E.L., Munos R., Veličković P., Valko M.: Large-Scale Representation Learning on Graphs via Bootstrapping (2023). https://arxiv.org/abs/2102.06514 [Google Scholar]
  • [43].Veličković P., Fedus W., Hamilton W.L., Liò P., Bengio Y., Hjelm R.D.: Deep graph infomax. In: International Conference on Learning Representations (2019). https://openreview.net/forum?id=rklz9iAcKQ [Google Scholar]
  • [44].Sun F.-Y., Hoffman J., Verma V., Tang J.: Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. In: International Conference on Learning Representations (2019) [Google Scholar]
  • [45].Hamilton W.L., Ying R., Leskovec J.: Inductive representation learning on large graphs. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS’17, pp. 1025–1035. Curran Associates Inc., Red Hook, NY, USA (2017) [Google Scholar]
  • [46].Wang W., Suo X., Wei X., Wang B., Wang H., Dai H.-N., Zhang X.: Hgate: Heterogeneous graph attention auto-encoders. IEEE Trans. on Knowl. and Data Eng. 35(4), 3938–3951 (2023) 10.1109/TKDE.2021.3138788 [DOI] [Google Scholar]
  • [47].Park C., Kim D., Han J., Yu H.: Unsupervised attributed multiplex network embedding. Proceedings of the AAAI Conference on Artificial Intelligence 34(04), 5371–5378 (2020) 10.1609/aaai.v34i04.5985 [DOI] [Google Scholar]
  • [48].Derr T., Ma Y., Tang J.: Signed graph convolutional networks. In: 2018 IEEE International Conference on Data Mining (ICDM), pp. 929–934 (2018) [Google Scholar]
  • [49].Li Y., Tian Y., Zhang J., Chang Y.: Learning signed network embedding via graph attention. Proceedings of the AAAI Conference on Artificial Intelligence 34(04), 4772–4779 (2020) 10.1609/aaai.v34i04.5911 [DOI] [Google Scholar]
  • [50].Huang J., Shen H., Hou L., Cheng X.: Sdgnn: Learning node representation for signed directed networks. Proceedings of the AAAI Conference on Artificial Intelligence 35(1), 196–203 (2021) 10.1609/aaai.v35i1.16093 [DOI] [Google Scholar]
  • [51].Virtanen P., Gommers R., Oliphant T.E., Haberland M., Reddy T., Cournapeau D., Burovski E., Peterson P., Weckesser W., Bright J., van der Walt S.J., Brett M., Wilson J., Millman K.J., Mayorov N., Nelson A.R.J., Jones I E., Kern R., Larson E., Carey C.J., Polat I., Feng Y., Moore E.W., Vander-Plas J., Laxalde D., Perktold J., Cimrman R., Henriksen I., Quintero E.A., Harris C.R., Archibald A.M., Ribeiro A.H., Pedregosa F., van Mulbregt P., SciPy 1.0 Contributors: SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nature Methods 17, 261–272 (2020) 10.1038/s41592-019-0686-2 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
media-1.xlsx (27.5KB, xlsx)
Supplement 2
media-2.xlsx (44.6KB, xlsx)
Supplement 3
media-3.xlsx (17.2KB, xlsx)
4

Data Availability Statement

All datasets utilized in this study are publicly available.


Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES