Abstract
Background:
The human microbiome plays a crucial role in regulating the efficacy and toxicity of drugs as well as in developing the drugs. Therefore, predicting the drug-related microbes is beneficial for analyzing the functional mechanisms of drugs. Recently, the graph learning based methods demonstrated their advantages in extracting the node features from the biological heterogeneous graphs. However, the previous methods failed to completely preserve the intrinsic structures of biological data and did not fully utilize the topological and positional information for predicting the drug-microbe associations.
Results:
We propose a new prediction model, structure-sensitive transformer and multi-view graph contrastive learning for microbe-drug association prediction (SMMDA), to encode and integrate the topological structures, semantics, and multiple-view embedding features of the drugs and microbes. Considering the sparsity of the original features of drugs and microbes, the learnable data augmentation strategy is designed to learn their global representations. Since similar drugs are more likely to associate with the similar microbes, a structure-sensitive transformer is proposed to integrate the topology structures composed of drugs (microbes) to form the multi-view embedding features. We design two contrastive learning strategies to exploit the complementary semantics across multiple views. As the embedding features from multiple views have various semantics, we design view-level attention to adaptively integrate these features.
Conclusions:
The extensive experimental results show that SMMDA outperforms several state-of-the-art methods for predicting the drug-related candidate microbes. The ablation studies show the effectiveness of the major innovations which include the learnable data augmentation, structure-sensitive transformer-based node feature learning, and multi-view contrastive learning. The case studies on three drugs also demonstrate SMMDA’s capability in retrieving the potential microbe candidates for the drugs.
Keywords: Learnable data augmentation, Structure-sensitive transformer, Multi-view contrastive learning, View-level attention mechanism
Introduction
The human body hosts trillions of microbes, including bacteria, archaea, fungi, protozoa, and viruses [1, 2], which form distinct microbiomes. These microbiomes are distributed across various parts of the body, such as the skin, oral cavity, nasal passages, gastrointestinal tract, and reproductive and urinary systems. They play a vital role in regulating human health by protecting against pathogens and promoting metabolism [3]. However, microorganisms can also negatively affect the body [4]. Research has shown that imbalances in the microbiome can lead to conditions such as diabetes, inflammatory bowel disease, and even cancer [5, 6].
Recent studies have underscored the significant role of microbes in influencing drug activity and toxicity [7, 8]. Conversely, drugs can alter the diversity and functionality of microbial communities. For example, methicillin-resistant Staphylococcus aureus (MRSA) produces the enzyme
-lactamase, which degrades the
-lactam ring in antibiotics [9], contributing to drug resistance. As drug-resistant pathogens become increasingly prevalent, understanding the interactions between microbes and drugs has become essential for advancing drug development [10].
Microbe–drug associations are traditionally identified through experimental techniques, but pinpointing target microbes remains a complex and challenging task, slowing the progress of new drug discoveries. Many studies have focused on reusing existing drugs or exploring drug combinations, but the rise of drug-resistant microbes complicates these efforts [11]. This underscores an urgent need for methods capable of predicting associations between drugs and novel target microbes. Traditional wet-lab approaches, while effective, are time-consuming, labor-intensive, and costly. Computational approaches provide a complementary solution, offering efficient and accurate predictions of associations [12, 13]. Rohit et al. [14] employ contrastive learning to enhance the prediction of drug–protein target interactions. Ma et al. [15] integrates six heterogeneous biological modalities by first constructing both protein–protein interaction and protein homology networks, then applying intra-view and inter-view contrastive learning to derive enriched, highly discriminative protein embeddings. Gu et al. [16] introduces a hierarchical graph Transformer that generates semantically informed super-node embeddings, adaptively aggregates them to form comprehensive graph representations, and uses contrastive learning as a regularization technique to maximize the consistency of the same graph’s embeddings across different views. Computational methods are also highly effective in identifying potential microbe–drug associations and reliable drug candidates for wet-lab validation. Predictive models such as graph attention networks (GAT) [17] and graph convolutional networks (GCN) [18] have been utilized to infer the probability of these associations. For instance, SCSMDA [19] employs graph contrastive learning and refined meta-paths to enhance drug and microbe representations. GNAEMDA [20] employs graph normalization to enhance drug and microbe representations. H Kuang et al. [21] developed an integrated framework, GARFMDA, which constructs multimodal heterogeneous networks of microbes and various drugs from multiple biomedical data sources. This framework incorporates a GAT, which assigns different weights to edges and nodes in a heterogeneous graph, and a bilayer random forest for filtering diversity information of drugs and microbes. DHDMP [22] constructs a dynamic hypergraph by treating specific relationships among multiple drugs and microbes as hyperedges and designs a spatial cross-attention mechanism to encode the attributes of node pairs. Despite their innovations, these methods have certain limitations. GARFMDA [23] uses bilayer random forest to filter information of drugs and microbes, without considering the potential impact of random forest on disrupting the intrinsic structure of the data. For example, the removal or addition of edges could significantly alter the identity of microbial data or even the effectiveness of drugs [23]. In addition, SCSMDA [19] applys a standard GCN to extract embedding features of drugs and microbes, while overlooks the structural and feature similarities between drugs (or microbes). The hyperedges in DHDMP [22] and the spatial cross-attention mechanism focus primarily on semantic information between nodes, while ignoring the neighboring topology information of the microbes (drugs). Moreover, these approaches struggle to address the noise introduced by numerous unobserved associations in the datasets, which can compromise prediction accuracy.
We propose structure-sensitive transformer and multi-view graph contrastive learning for microbe-drug association prediction (SMMDA), a novel model that predicts candidate microbes for drugs through multi-view contrastive learning. Our approach addresses limitations in previous methods by leveraging innovative strategies to improve predictive accuracy. The key contributions of our work are as follows:
We developed an innovative data augmentation method that generates new graph structures and representations for drugs and microbes. This approach preserves the intrinsic structure of drug–microbe graphs, facilitating the creation of multi-view information for contrastive learning. The learned graph structures and representations are utilized to derive rich, multi-view insights into drug–microbe interactions.
Drugs with similar structures are more likely to have comparable effects on related microbes. To simultaneously capture structural and attribute similarities between nodes, we designed a structure-sensitive transformer (SST) that generates multi-view embedding features. This transformer considers not only node attributes but also local structures by extracting subgraph representations around each node.
We introduced a novel framework to learn complementary information across multi-view drug–microbe interactions. This framework employs two types of losses: one that separates positive and negative embeddings within the same graph structure, and another that differentiates embeddings across different graph structures. By integrating learnable data augmentation and multi-view contrastive learning into a unified framework, our approach supports iterative optimization.
Since the embedding features generated for each drug–microbe view contribute differently to the prediction of microbe–drug associations, we implemented a view-level attention mechanism to effectively integrate multi-view embeddings. Our experimental results demonstrate that SMMDA significantly outperforms existing state-of-the-art methods in predictive performance.
Materials and methods
We propose a microbe–drug association prediction model, SMMDA (Fig. 1), which consists of four main components: a learnable data augmentation module, an SST module, a multi-view contrastive learning module, and a view-level attention module. A heterogeneous graph of drug–microbe interactions is constructed to integrate similarities and associations between drugs and microbes. Through learnable data augmentation, we generate new heterogeneous graphs and features to create multi-view representations of drug–microbe interactions. The SST is employed to capture embedding features for each view by considering both node attributes and local graph structures. Multi-view contrastive learning is then applied to enhance the complementary information across the views. Finally, the view-level attention mechanism assigns adaptive weights to embedding features from different views, producing the final embedding. This final representation is used to predict association scores between drugs and microbes.
Fig. 1.
The framework of the proposed SMMDA model consists of two main components: a Constructing a multi-view graph based on data augmentation and b Structure-sensitive feature representation learning with graph contrast
Dataset
The dataset used in our experiments is sourced from previous work, with raw data obtained from the MDAD database [24]. MDAD includes 5,505 clinical reports and experimentally validated associations between microbes and drugs. Many of these entries in the databases are redundant, referencing the same microbe–drug associations across different sources. To ensure data consistency, we merge entries that refer to the same drug–microbe pair. After this deduplication process, we obtain 2,470 unique associations involving 1,373 drugs and 173 microbes [19]. Specifically, drug–drug similarities were calculated as a weighted average of chemical substructure similarity and Gaussian similarity, while microbe–microbe similarities were derived using sequence-based methods [25].
Drug–drug similarity
We first build a structural similarity matrix
using the SIMCOMP2 method [25], which compares drugs based on their chemical structures. We then compute a Gaussian interaction profile kernel similarity from the drug–microbe interaction matrix
, capturing behavioral-level affinities. Finally, we integrate these two measures. If a drug pair has a nonzero structural similarity, we take the average of
and
; otherwise, we use
alone,
![]() |
1 |
where
denotes the Gaussian interaction profile kernel similarity between drug
and drug
.
Microbe–microbe similarity
We first compute a functional similarity matrix
based on microbial gene function information using the method proposed by Kamneva [26]. Next, we calculate a Gaussian interaction profile kernel similarity from the microbe–drug interaction data, denoted
. Finally, we integrate these two measures into a composite similarity,
![]() |
2 |
Construction of the drug-microbe heterogeneous graph
The matrix
captures the chemical substructure similarity between drugs, where the
-th row represents the similarity of drug
with all
drugs. Similarly, the matrix
encodes the similarity between microbe
and all
microbes, with the
-th row reflecting these relationships. The matrix
indicates known associations between drugs and microbes: if a connection exists between drug
and microbe
, then
; otherwise
. We construct a drug–microbe heterogeneous graph
, where node set
includes the drug nodes
and the microbe nodes
. The edge set
contains edges
that indicate whether a link exists between nodes
and
. The graph
comprises three types of edges: drug–drug similarity edges, microbe–microbe similarity edges, and drug–microbe association edges.
The adjacency matrix
of the heterogeneous graph
is given by,
![]() |
3 |
where
is the transpose of
, and
. The
-th row of
represents the similarity of drug
with all
drugs and its association with all
microbes. Similarly, the
-th row of
reflects the relationships of microbe
with the
drugs and its similarity to the
microbes. The diagonal elements of
correspond to the self-similarity of drugs or microbes and are all set to 1. Additionally, the
-th row of
represents the attribute vector of node
. Therefore, the adjacency matrix
doubles as an attribute matrix, encapsulating all the attributes of drugs and microbes, such as
.
Learnable data augmentation on drug and microbe features
Learnable node feature augmentation
The similarity between drugs is calculated based on their chemical substructures and Gaussian similarity [27]. The higher the similarity between drug
and drug
, the greater the likelihood that the microbes associated with drug
will also be associated with drug
[28–30]. To capture these key features, we propose a novel learnable node feature augmentation strategy [31, 32]. We introduce an indicator vector
, where the elements of
are randomly assigned as either 0 or 1 [33]. This vector
is used to mask elements of the original attribute matrix
, resulting in an augmented attribute matrix
,
![]() |
4 |
where
represents element-wise multiplication.
contains the similarity between drug
and
drugs, as well as its association with
microbes, while
represents the associations of microbe
with all drugs and its similarity with other microbes. To avoid excessive perturbations, we design the following loss function,
![]() |
5 |
where
is the squared Frobenius norm of the matrix, and
is a trainable parameter. This loss function ensures that
and
remain as consistent as possible while minimizing the
-norm. Under a specified masking ratio, we select columns in the weight matrix
with relatively small weights and set the corresponding elements of the vector
to 0, while the remaining elements are set to 1.
Learnable graph structure augmentation
In conventional data augmentation methods, the graph structure typically remains unchanged, i.e.,
. However, when the generated representation
is paired with the original adjacency matrix
, noise and redundancy may be introduced. By using Eq. (2), we obtain a better representation than
, which can subsequently be used to update
. This updated
better captures the intrinsic structure of the data. To achieve this [34, 35], we propose a learnable graph structure augmentation method that dynamically adjusts the graph structure to uncover the intrinsic relationships between drug and microbe data after augmentation. Specifically, we employ a novel graph learning method to emphasize the intrinsic structure of the data. The updated
is computed as follows,
![]() |
6 |
where
and
are trainable parameters,
is the exponential function, ReLU(.) is the activation function, and
is a row-normalized vector. The loss function for graph structure augmentation is defined as,
![]() |
7 |
where
denotes the augmented adjacency matrix and
is the weight of edge
,
is the squared Frobenius norm of the matrix, i.e., the sum of the squares of all its elements. The weighted structural difference term
penalizes edges in the augmented graph: if the two end-nodes are structurally similar, the difference is small, so retaining that edge incurs little loss, protecting critical connections from deletion; conversely, if an edge connects nodes with dissimilar structure, the penalty is large, driving the model to reduce
and thus avoid redundant or noisy edges. The Frobenius norm term
acts as an
penalty on all augmented edges, preventing excessively large weights and preserving the sparsity pattern of the augmented graph in line with the original topology. We integrate the losses from the data augmentation parts, Eq. (5) and Eq. (7), into a unified framework, as follows,
![]() |
8 |
Using the proposed learnable data augmentation method, multi-view information is generated based on the new representations
and feature graph
, as follows,
![]() |
9 |
In contrast to traditional data augmentation methods, our learnable approach updates both
and
, preventing excessive perturbation of insignificant nodes and edges, thereby preserving the structural integrity of the data [33].
Learning embedding features via structure-sensitive transformer
Building on the biological assumption that drugs with similar structures tend to have comparable effects on similar microbes, we derive the representation of a target node, whether drug or microbe, by aggregating the attributes of nodes exhibiting high structural similarity with it [27]. However, nodes with similar structures may be distant from the target node within the graph [36]. This global relationship is often overlooked by traditional GNN-based approaches, which primarily focus on local neighbors. Consequently, a global dependency exists across the attributes of all nodes in the heterogeneous network, highlighting the need to incorporate these relationships into the model.
Drawing inspiration from the Transformer model introduced in Chen et al. [37], we propose a structure-sensitive transformer to capture these global dependencies (as shown in the Fig. 2). Firstly, we use a
-subtree extractor to focus on the structural similarity between nodes through local subgraphs. This method applies to a graph neural network (GNN) with layers to the input graph and uses the output representation of node
as the subgraph representation of that node. Specifically, for each node, we collect all nodes within its
-hop neighborhood (including the node itself) as well as all edges among them. As shown in Fig. 2, the 1-hop subgraph of node
includes the set of nodes
and the corresponding edges
. The representation of node
in graph
is computed by,
![]() |
10 |
where
,
is the node feature matrix of
,
is the adjacency matrix of
,
is the degree matrix of
,
,
,
,
,
denotes the concatenation operation,
represents computing a weighted row vector by taking the column-wise mean of the resulting matrix. The similarity between node features in self-attention can be replaced with the similarity between vectors that incorporate structural information from a node’s local neighborhood,
![]() |
11 |
where
represents the inner product,
,
are the projection matrices for query and key, and
is the output dimension. Thus, we can compute the structure-sensitive attention vector of node
as follows,
![]() |
12 |
where
is a linear value function, where
is a trainable parameter matrix. The output of the structure-sensitive self-attention is then followed by a skip connection and a feedforward network (FFN), forming a SST layer, as follows,
![]() |
13 |
where
,
are parameter matrices, and
is the degree of node
. A multilayer stack of these structure-sensitive transformer layers constitutes the SST. Given the multi-view data (i.e.,
), we employ two SSTs (as shown in Fig. 2) to generate multi-view embeddings, as different graphs provide complementary information [33]. For instance, the original adjacency matrix
captures real-world similarities and associations between drugs and microbes, while the augmented adjacency matrix
collects similarities from the feature space. Specifically, since
and
share the same topology, SST1 is applied to
and
, facilitating the generation of high-quality
. Similarly,
and
are fed into SST2 to output embedding features
and
, while also updating
and
,
![]() |
14 |
Fig. 2.
Overview of an example SST layer that uses the k-subtree GNN extractor as its structure extractor. The structure extractor generates structure-sensitive node representations which are used to compute the query (Q) and key (K) matrices in the Transformer layer
Multi-view contrastive learning
and
(
and
) represent the embedding features of all drugs and microbes learned under different views. The enhanced features are derived from the original features, and the enhanced graph is derived from the enhanced features. Therefore,
and
(
and
) are closely related. To uncover more comprehensive contrastive relationships, we design a contrastive loss within the same graph to differentiate positive and negative embeddings within the same SST and graph structure [33]. Simultaneously, we design a contrastive loss across different views to separate positive and negative embeddings across different graph structures, focusing on embeddings derived from different graphs.
Fig. 3.
Illustration of the positive and negative pair selection strategy of SMMDA
Contrastive learning within the same view
The contrastive loss within the same view is designed to establish relationships within a single graph structure. Its objective is to bring the embedding features of the same node closer together across different views of the graph while pushing the features of one node farther from those of other nodes (whether in the same or different views) within the same structure. This is mathematically expressed as,
![]() |
15 |
![]() |
16 |
Here,
represents the similarity between two drug (or microbe) nodes, where
is the temperature parameter. The total contrastive loss within the same view for all drug and microbe nodes is defined as,
![]() |
17 |
Contrastive learning across different views
The goal of contrastive learning across different views is to compare embedding features between distinct graph structures, such as the original and augmented graphs. To achieve this, we design a contrastive loss that aligns the embedding features of the same drug (or microbe) node across different views, while maximizing the separation between the features of distinct nodes across these views [37]. This process is formulated as,
![]() |
18 |
![]() |
19 |
The total contrastive loss across different views, computed for all drug and microbe nodes, is defined as,
![]() |
20 |
Finally, by combining the contrastive loss within the same view and the contrastive loss across different views, the proposed multi-view contrastive learning loss is expressed as,
![]() |
21 |
The contrastive loss within the same view captures relationships within a single graph structure, while the contrastive loss across different views addresses relationships across distinct graph structures and views. By integrating these two components, Eq. (21) not only preserves the individual characteristics of each graph structure but also bridges the two graph structures. This approach enables the effective capture of complementary information between the original and augmented graph structures.
View-level attention mechanism
The embedding features of drugs (or microbes) from different perspectives may contribute differently to drug–microbe association prediction. To address this, we propose a view-level attention mechanism to integrate multi-view embeddings seamlessly [38]. Taking a drug
as an example (the attention scores for a microbe
,
, can be derived similarly), the relative importance scores
are learned using the view-level attention mechanism
as follows,
![]() |
22 |
where
represent the attention scores for
, respectively, and
denotes the attention operation. To compute the attention values, the embeddings are first transformed through a nonlinear transformation. A shared attention vector
is used to compute the attention score
,
![]() |
23 |
where
is the weight matrix, and
is the bias vector.Similarly,
can be computed using shared parameters. Next, these values are normalized using a Softmax function to obtain the weights
,
![]() |
24 |
where exp represents the exponential function. Similarly,
and
can be computed. Finally, the multi-view embeddings for drug
are aggregated to obtain the final embedding
,
![]() |
25 |
where
denotes the concatenation operation. Similarly, the final embedding for microbe
can be computed as,
![]() |
26 |
Final integration and optimization
The embeddings of drug
,
, and microbe
,
, are vertically stacked. After passing through a
convolution layer, the combined feature vector is fed into a fully connected layer to obtain the final prediction score
,
![]() |
27 |
where
represents the association score between drug
and microbe
, with a higher score indicating a greater likelihood of an association.
is the weight matrix of the multilayer perceptron (MLP), and
is the flattened feature vector obtained after convolution on the stacked features.
During training, we optimize the model using backpropagation and the gradient descent algorithm. The cross-entropy loss for the model’s predictions is defined as
,
![]() |
28 |
where
is the ground truth label,
is the size of the training batch, and
.
and
represent the probabilities of drug
and microbe
being associated or not being associated, respectively. Therefore, the overall loss of our model is defined as,
![]() |
29 |
Here,
is a tuning parameter that adjusts the impact of the learnable data augmentation and multi-view contrastive learning components on the model.
Experimental evaluations and discussions
Evaluation metrics
To assess the performance of SMMDA and other comparative methods, a five-fold cross-validation approach was utilized [39]. Known drug–microbe associations were classified as positive samples and randomly distributed into five folds, while unknown associations, which had not been validated through biological experiments, were regarded as negative samples. In each fold, four folds of positive samples, along with an equal number of randomly chosen negative samples, were employed for training. The remaining fold of positive samples, together with all unselected negative samples, constituted the test set.
If the association score between drug
and microbe
falls below a specified threshold
, the pair is classified as a negative sample. Conversely, if the score exceeds the threshold, the pair is labeled as a positive sample.To assess performance, we use the area under the receiver operating characteristic curve (AUC) [40], the area under the precision-recall curve (AUPR) [41], and the recall rate of the top-k candidate microbes associated with each drug as evaluation metrics.
Biologists frequently prioritize candidate samples at the top of their ranking lists for wet-lab validation. Therefore, it is crucial to maximize the number of positive samples that appear in these top ranks. To tackle this issue, we propose incorporating the recall rate of the top-k ranked samples as an additional evaluation metric. This metric is defined as the proportion of correctly identified positive samples within the top-k list relative to the total number of positive samples.
Parameter settings
The proposed model, SMMDA, was implemented using the PyTorch machine learning framework and trained on an NVIDIA GeForce 3070 GPU to accelerate computation. The Adam optimizer was used, and hyperparameters were fine-tuned to achieve optimal performance. To systematically evaluate how the masking rate influences the learnable data augmentation process, we conducted a sensitivity analysis across five masking rates
. As shown in Table T1, the model’s performance begins to degrade when the masking rate exceeds 0.2-excessively high masking rates result in the loss of critical information, while excessively low rates (e.g., 0.1) fail to generate sufficiently strong augmentations. The masking rate of 0.2 achieves the best performance in terms of both AUC and AUPR; thus, we set the masking rate to 0.2 in our experiments.For DHDMP, the learning rate was set to
. Both NHCN and GCNFP adopt two-layer encoders and construct 32 hyperedges per sample. The feature dimension is set to 50, and the number of attention heads is 8. GCNMDA is evaluated using five-fold cross-validation. The number of hidden units is set to 25, the negative sampling ratio
, the number of CRF iterations
, and the loss function weights are set to
,
, and
. GSAMDA sets the topological feature dimension to 128 and the attribute feature dimension to 32. The learning rates for both GAE and SAE modules are fixed at 0.01. MGAVAEMDA adopts a fixed learning rate of 0.01 and sets the hidden layer dimension to 128. SCSMDA achieves the best performance when the learning rate is
, the number of positive sample pairs is 10, the number of bins in self-paced negative sampling is 10, and both the MLP and GCN module consist of a single layer.
To systematically evaluate how the masking rate influences the learnable data augmentation process, we conducted a sensitivity analysis across five masking rates (0.1, 0.2, 0.3, 0.4, and 0.5). As shown in Table 1, the model’s performance begins to degrade when the masking rate exceeds 0.2-excessively high masking rates result in the loss of critical information, while excessively low rates (e.g., 0.1) fail to generate sufficiently strong augmentations. The masking rate of 0.2 achieves the best performance in terms of both AUC and AUPR. Thus, we set the masking rate to 0.2 in our experiments.
Table 1.
Ablation experimental results of masking rate
| Masking rate | AUC | AUPR |
|---|---|---|
| 0.10 | 0.971 | 0.842 |
| 0.20 | 0.973 | 0.845 |
| 0.30 | 0.970 | 0.841 |
| 0.40 | 0.964 | 0.837 |
| 0.50 | 0.959 | 0.834 |
The best result is highlighted in bold
The tuning parameter
, which controls the influence of matrix
on the masking operation, was selected from the range
. After testing,
was found to provide the best results. For the k-hop neighborhood,
was chosen from
, and
was selected for the final experiments. In the SST, the number of attention heads,
, was set to 8, and the output dimension,
, was set to 800. The hyperparameter
, which governs the balance between the contributions of multi-view contrastive learning and learnable data augmentation, was tested across the range
. The optimal value of
was found to yield the highest model performance, with an AUC of 0.973 and an AUPR of 0.845.
Ablation studies
To evaluate the effectiveness of our data augmentation method, multi-view contrastive loss, and SST, we introduced four ablation variants Proposed-R-A, Proposed-R-Inter, Proposed-R-Intra, and Proposed-R-SST. Specifically:
Proposed-R-A applies random perturbation for data augmentation and does not generate new augmented graphs.
Proposed-R-Inter removes the contrastive loss across different views (
).Proposed-R-Intra excludes the contrastive loss within the same view (
).Proposed-R-SST replaces SST with a GCN to generate multi-view embedding features.
Table 2.
Results of the ablation studies
| Method | AUC | AUPR |
|---|---|---|
| Proposed-R-A | 0.955 | 0.805 |
| Proposed-R-Inter | 0.967 | 0.819 |
| Proposed-R-Intra | 0.971 | 0.832 |
| Proposed-R-SST | 0.956 | 0.810 |
| Proposed | 0.973 | 0.845 |
The best result is highlighted in bold
The final model achieved the highest AUC (0.973) and AUPR (0.845). In comparison:
When random perturbation was used for data augmentation instead of learnable augmentation (Proposed-R-A), the AUC and AUPR decreased by 1.8% and 4.0%, respectively.
Removing the contrastive loss across different views (Proposed-R-Inter) resulted in a 0.6% decrease in AUC and a 2.6% decrease in AUPR.
Removing the contrastive loss within the same view (Proposed-R-Intra) caused the AUC and AUPR to drop by 0.2% and 1.3%, respectively.
Replacing SST with GCN (Proposed-R-SST) led to a 1.7% decline in AUC and a 3.5% decline in AUPR.
The ablation studies demonstrate that the learnable data-augmentation module yields the largest gain by dynamically masking regions based on node importance, preserving critical semantics and graph structure far more effectively than random augmentations. Structure-aware subgraph extraction (SST) is the second most impactful: without SST, the model relies solely on node attributes and ignores topology, whereas SST’s subgraph snapshots and attention mechanism capture vital local connectivity. Finally, both contrastive losses are essential: contrastive loss within the same view enforces consistency within each augmented view to counteract noise-induced drift, while contrastive loss across different views aligns representations across different views to harvest complementary information from multiple graph structures.
We projected an equal number of positive and negative drug–microbe pairs into two-dimensional space (Fig. 4). Without contrastive learning, the two classes appear loosely scattered; after contrastive learning, both positive and negative samples form tight clusters, markedly improving their separability.
Fig. 4.
t-SNE visualization of drug–microbe pair embeddings with and without contrastive learning
We conducted two cold-start experiments to assess SMMDA’s generalization: in the drug cold-start, test-set drugs are completely excluded from training so that the model must predict associations for entirely unseen drugs; in the microbe cold-start, test-set microbes are similarly withheld during training, forcing the model to generalize to novel microbial targets.
Comparison with other models
SMMDA was evaluated against six leading methods for microbe–drug association prediction: GCNMDA [18], EGATMDA [42], GSAMDA [43], SCSMDA [19], MGAVAEMDA [44], and DHDMP [22]. All methods, including SMMDA, were trained and tested using five-fold cross-validation on the same dataset splits, with hyperparameters set according to the respective publications. Below is a brief description of the comparison methods,
GCNMDA [18]: This method is based on a GCN and combines drug chemical information, microbe genetic information, and Gaussian interaction profiles to quantify similarity. It uses a random walk preprocessing scheme to capture features for predicting microbe–drug associations.
EGATMDA [42]: This method constructs a microbe–disease–drug network and employs a hierarchical attention mechanism to predict microbe–drug associations.
GSAMDA [43]: This model calculates drug (microbe) similarities using Gaussian and Hamming interaction profiles and employs a GAT along with a sparse autoencoder to learn node features.
SCSMDA [19]: This method builds a microbe–drug network using gene sequence data, Gaussian kernel interaction profiles, and drug chemical structures, applying graph contrastive learning to extract features for microbe and drug nodes.
MGAVAEMDA [44]: This computational framework integrates an improved graph attention variational autoencoder with biological information to predict potential microbe–drug associations.
DHDMP [22]: This method leverages dynamic hypergraph modeling and long-distance spatial correlation encoding to predict associations between drugs and microbes.
We first calculated the AUC and AUPR for all models and then computed the average AUC and AUPR across 1373 drugs. Table 3 shows that SMMDA achieved the highest average AUC of 0.973, surpassing the second-best model, DHDMP, by 1.4%, EGATMDA, by 3.3%, GCNMDA by 7.0%, MGAVAEMDA by 3.7%, GSAMDA by 7.1%, and SCSMDA by 5.7%. Additionally, SMMDA achieved the highest average AUPR of 84.5%, exceeding DHDMP by 2.2%, SCSMDA by 50.5%, GCNMDA by 53.0%, EGATMDA by 53.8%, MGAVAEMDA by 51.7%, and GSAMDA by 59.8%.
Table 3.
AUCs and AUPRs of different methods in comparison across all the 1373 drugs
| Networks | AUC (%) | AUPR (%) |
|---|---|---|
| GCNMDA | 90.3 | 31.5 |
| EGATMDA | 94.0 | 30.7 |
| GSAMDA | 90.2 | 24.7 |
| SCSMDA | 91.6 | 34.0 |
| MGAVAEMDA | 93.6 | 32.8 |
| DHDMP | 95.9 | 82.3 |
| SMMDA | 97.3 | 84.5 |
The best result is highlighted in bold
As shown Tables 4 and 5, we compare SMMDA against other methods under two cold-start scenarios. Although all approaches experience declines in AUC and AUPR, SMMDA remains highly robust and continues to deliver the best performance.
Table 4.
Drug cold-start evaluation
| Networks | AUC (%) | AUPR (%) |
|---|---|---|
| GCNMDA | 83.0 | 21.5 |
| EGATMDA | 88.5 | 26.4 |
| GSAMDA | 81.9 | 18.6 |
| SCSMDA | 84.2 | 22.1 |
| MGAVAEMDA | 87.1 | 25.5 |
| DHDMP | 91.8 | 67.2 |
| SMMDA | 93.3 | 70.7 |
The best result is highlighted in bold
Table 5.
Microbe cold-start evaluation
| Networks | AUC (%) | AUPR (%) |
|---|---|---|
| GCNMDA | 82.7 | 20.8 |
| EGATMDA | 88.9 | 27.0 |
| GSAMDA | 82.1 | 19.2 |
| SCSMDA | 84.5 | 23.0 |
| MGAVAEMDA | 86.8 | 26.0 |
| DHDMP | 92.0 | 68.0 |
| SMMDA | 93.5 | 71.5 |
The best result is highlighted in bold
Figure 5 illustrates the average recall rates of all drugs at different top-k candidate microbes. SMMDA, which leverages contrastive learning to enhance features with global information, outperformed all other methods across different top-k thresholds. For
, our model achieved the highest recall rate of 89. 3%, compared to 1. 2% for the second best model, DHDMP. MGAVAEMDA achieved the fifth-best recall rate of 45.2%, which was 1.9% lower than SCSMDA and 2.5% lower than EGATMDA.
Fig. 5.
The average recalls of drugs at different top
settings
When
,
, and
, SMMDA achieved the highest recall rates of 89.6%, 90.1%, and 90.3%, respectively. The second-best model, DHDMP, showed recall rates of 89.1%, 89.4%, and 89.5% at these thresholds. SCSMDA achieved recall rates of 65.4%, 72.7%, and 79%, outperforming MGAVAEMDA, which had recall rates of 61.5%, 66.8%, and 70.8%. EGATMDA performed better than both, with recall rates of 67.8%, 74.9%, and 80.1%. GSAMDA had lower recall rates of 53.6%, 67.6%, and 71.4%, but still exceeded GCNMDA, which achieved recall rates of 55.7%, 63.7%, and 68.4%.
Case studies on three drugs
To assess the effectiveness of SMMDA in identifying microbial candidates linked to specific medications, we performed case studies on ciprofloxacin, moxifloxacin, and ceftazidime. Ciprofloxacin is utilized for treating skin infections, typhoid fever, pneumonia, endocarditis, and various other bacterial infections. Moxifloxacin is commonly prescribed for pneumonia, tuberculosis, sinusitis, and chronic bronchitis. Ceftazidime, a third-generation cephalosporin antibiotic, is primarily employed in the treatment of pneumonia, sepsis, urinary tract infections, skin and soft tissue infections, and intra-abdominal infections.
In these case studies, the model utilized all known microbe–drug associations, supplemented by an equal number of randomly chosen unobserved associations to serve as negative samples. For each drug, the model produced a list of potential microbes, from which the top 20 ranked candidates were further examined.These case studies offer valuable insights into the predictive performance and potential utility of the SMMDA model in real-world applications. By comparing the model’s predicted microbe–drug associations with known established associations, the accuracy and reliability of the model in identifying drug-associated microbial candidates are further validated.
The MDAD [24] dataset contains information on known interactions between 1,373 drugs and 173 microbes. The aBiofilm [45] database focuses on the relationship between bacterial biofilm formation and drug resistance, providing data on bacterial responses to various drugs in the biofilm state, encompassing 1,720 drugs and 140 microbes. We validated the microbe–drug association predictions of SMMDA using the MDAD and aBiofilm datasets, as well as relevant literature. Among the top 20 candidate microbes associated with ciprofloxacin (Table 3), four were found in the MDAD and aBiofilm datasets, indicating their known relationship with ciprofloxacin, while 16 were further confirmed by literature evidence. For example, several microbes, including Salmonella Typhi [46], Streptococcus sanguinis [47], Streptococcus mutans [48], and Acinetobacter baumannii [49], were reported to be inhibited or killed by ciprofloxacin. Additionally, Enterococcus faecium was identified as resistant to ciprofloxacin [50].
Table 6.
Top 20 candidate microbes associated with Ciprofloxacin
| Rank | Microbe name | Evidence | Rank | Microbe name | Evidence |
|---|---|---|---|---|---|
| 1 | Streptococcus mutans | aBiofilm, MDAD | 11 | Vibrio cholerae | 26271050 |
| 2 | Pseudomonas aeruginosa | aBiofilm, MDAD | 12 | Burkholderia multivorans | 19633000 |
| 3 | Stenotrophomonas maltophilia | aBiofilm, MDAD | 13 | Candida tropicalis | 16849719 |
| 4 | Human immunodeficiency virus 1 | aBiofilm, MDAD | 14 | Salmonella Typhi | 10334265 |
| 5 | Acinetobacter baumannii | 20138741 | 15 | Enterococcus faecium | 20006472 |
| 6 | Propionibacterium acnes | 25541476 | 16 | Klebsiella pneumoniae | 10858336 |
| 7 | Streptococcus sanguinis | 11347679 | 17 | Human herpesvirus 1 | Unconfirmed |
| 8 | Proteus mirabilis | 22958285 | 18 | Serratia liquefaciens | 10965096 |
| 9 | Enterococcus faecalis | 27790716 | 19 | Bacillus anthracis | 12821500 |
| 10 | Streptococcus sanguis | 11347679 | 20 | Actinomyces oris | Unconfirmed |
Table 7.
Top 20 candidate microbes associated with Moxifloxacin
| Rank | Microbe name | Evidence | Rank | Microbe name | Evidence |
|---|---|---|---|---|---|
| 1 | Stenotrophomonas maltophilia | aBiofilm, MDAD | 11 | Pseudomonas aeruginosa | 31691651 |
| 2 | Staphylococcus aureus | aBiofilm, MDAD | 12 | Escherichia coli | 31542319 |
| 3 | Listeria monocytogenes | aBiofilm, MDAD | 13 | Vibrio harveyi | 27247095 |
| 4 | Haemophilus influenzae | aBiofilm, MDAD | 14 | Klebsiella pneumoniae | 27257956 |
| 5 | Bacillus subtilis | aBiofilm, MDAD | 15 | Acinetobacter baumannii | 20006472 |
| 6 | Listeria monocytogenes | aBiofilm, MDAD | 16 | Burkholderia cenocepacia | 27799222 |
| 7 | Burkholderia multivorans | 35754328 | 17 | Salmonella enterica | 15078598 |
| 8 | Actinomyces oris | 10858336 | 18 | Clostridium perfringens | 29486533 |
| 9 | Staphylococcus epidermidis | 28481197 | 19 | Burkholderia pseudomallei | 15731198 |
| 10 | Candida albicans | 31471074 | 20 | Serratia marcescens | Unconfirmed |
Table 8.
Top 20 candidate microbes associated with Ceftazidime
| Rank | Microbe name | Evidence | Rank | Microbe name | Evidence |
|---|---|---|---|---|---|
| 1 | Pseudomonas aeruginosa | aBiofilm,MDAD | 11 | Escherichia coli | 37574665 |
| 2 | Acinetobacter baumannii | aBiofilm,MDAD | 12 | Candida albicans | Unconfirmed |
| 3 | Haemophilus influenzae | 6376458 | 13 | Staphylococcus epidermis | 1730894 |
| 4 | Shigella flexneri | 31519769 | 14 | Klebsiella planticola | Unconfirmed |
| 5 | Pseudomonas aeruginosa | 34990760 | 15 | Candida spp | 6357068 |
| 6 | Bacillus subtilis | 31420587 | 16 | Eikenella corrodens | Unconfirmed |
| 7 | Mycobacterium tuberculosis | 20138741 | 17 | Stenotrophomonas maltophilia | 37615040 |
| 8 | Streptococcus pneumoniae serotype 4 | 8126192 | 18 | Candida tropicalis | Unconfirmed |
| 9 | Mycobacterium avium | 28922808 | 19 | Salmonella enterica | 19861080 |
| 10 | Proteus vulgaris | 19802966 | 20 | Pseudoalteromonas sp | 24031945 |
For the candidate microbes associated with moxifloxacin (Table 4), five were found in the MDAD and aBiofilm datasets, and 14 were supported by literature. For instance, moxifloxacin [51] exhibits antibacterial activity against Streptococcus pneumoniae, while Salmonella enterica [52] shows resistance to moxifloxacin. For the candidate microbes associated with ceftazidime (Table 5), two were present in the MDAD and aBiofilm datasets, and 14 candidates, including Shigella flexneri, were supported by literature. Among all 60 candidate microbes, six remain unconfirmed, indicating no supporting evidence for their associations with the target drugs. These findings demonstrate that SMMDA effectively identifies potential candidate microbes for the drugs under investigation.
Prediction of novel drug–microbe associations
SMMDA was comprehensively trained on all established drug–microbe associations to forecast potential candidate microbes for each drug. The top 20 predicted associations for every drug are provided in the supplementary Table ST1, which could assist biologists in identifying promising candidate microbes for further investigation.
Conclusion
We present a novel microbe–drug association prediction model that integrates both attribute and semantic information from drug and microbe nodes across multiple views using multi-view learning. This approach aims to predict microbe associations with drugs. Dynamic data augmentation is applied by selecting important features and perturbing nodes and edges. The SST model captures neighborhood structural features of drugs and microbes by extracting k-subtree information from nodes. Graph contrastive learning enhances the representation of drug and microbe nodes across multiple views by leveraging complementary information from two loss functions. A view-level attention mechanism assigns higher weights to more significant drug and microbe view features. Cross-validation experiments on public datasets demonstrate that SMMDA outperforms comparative methods in both AUC and AUPR. Additionally, the model’s average drug recall rate and case study analysis further confirm that SMMDA reliably identifies microbe candidates associated with drugs.
Limitations
In constructing drug–drug and microbe–microbe similarity matrices, we fuse Gaussian-kernel and structure/function measures to build drug–drug and microbe–microbe similarities, but only 2,470 known associations make these estimates noisy. To mitigate sparsity, we apply learnable masking augmentation and contrastive learning to align multi-view representations of each entity. However, by focusing only on within-entity consistency, we miss cross-entity functional relationships and finer similarity patterns. Therefore, in the next step, we will apply thresholding to retain only high-confidence similarities and explore richer positive/negative sampling strategies that incorporate drug-drug and microbe-microbe functional associations into the contrastive framework, thereby mining more nuanced similarity structures.
SMMDA is built based on a transformer backbone. Our small- subgraph extractor adds negligible cost versus self-attention, while the learnable augmentation and contrastive modules incur extra overheads, yielding a model of high computational complexity. A key future direction is to cut the high memory and time overhead of self-attention. Recent work on “linear Transformers”, such as Qin et al. [53] achieves true
time and space complexity, offering a promising approach to mitigate SMMDA’s computational demands.
Additional file
Acknowledgements
Not Applicable.
Author contributions
Ping Xuan: Designed the method and participated in manuscript writing. Rui Wang: Designed the experiments and participated in manuscript writing. Jing Gu: Participated in experiment design and manuscript writing. Hui Cui: Participated in experiment design. Tiangang Zhang: Participated in method design and manuscript writing.
Funding
This work was supported by Natural Science Foundation of Heilongjiang Province (LH2023F044). Natural Science Foundation of China (62172143, 62372282). Guangdong Basic and Applied Basic Research Foundation (2024A1515010176). STU Scientific Research Initiation Grant (NTF22032).
Data availability
The datasets are obtained from the previous work EGATMDA (https://github.com/uctoronto/EGATMDA) [42]. The datasets for training and testing our prediction model are freely available at https://github.com/pingxuan-hlju/SMMDA. The source code of our model is also contained by the GitHub link.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no Competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12859-025-06199-w.
References
- 1.THMP, Consortium. Structure, function and diversity of the healthy human microbiome. Nature. 2012;486:207–14. [DOI] [PMC free article] [PubMed]
- 2.Thiele I, Heinken A, Fleming RM. A systems biology approach to studying the role of microbes in human health. Curr Opin Biotechnol. 2013;24:4–12. [DOI] [PubMed] [Google Scholar]
- 3.ElRakaiby M, Dutilh BE, Rizkallah MR, Boleij A, Cole JN, Aziz RK. Pharmacomicrobiomics: the impact of human microbiome variations on systems pharmacology and personalized therapeutics. Omics J Integrat Biol. 2014;18:402–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Sprockett D, Fukami T, Relman DA. Role of priority effects in the early-life assembly of the gut microbiota. Nature Rev Gastroenterol Hepatol. 2018;15:197–205. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.de la Cuesta-Zuluaga J, Huus KE, Youngblut ND, Escobar JS, Ley RE. Obesity is the main driver of altered gut microbiome functions in the metabolically unhealthy. Gut Microbes. 2023;15:2246634. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Durack J, Lynch SV. The gut microbiome: relationships with disease and opportunities for therapy. J Experim Med. 2019;216:20–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Nejman D, Livyatan I, Fuks G, Gavert N, Zwang Y, Geller LT, Rotter-Maskowitz A, Weiser R, Mallel G, Gigi E. others The human tumor microbiome is composed of tumor type-specific intracellular bacteria. Science. 2020;368:973–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Algavi YM, Borenstein E. A data-driven approach for predicting the impact of drugs on the human microbiome. Nature Commun. 2023;14:3614. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Vestergaard M, Frees D, Ingmer H. Antibiotic resistance and the MRSA problem. Microbiol Spect. 2019;7:10–1128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hughes D, Andersson DI. Evolutionary trajectories to antibiotic resistance. Ann Rev Microbiol. 2017;71:579–96. [DOI] [PubMed] [Google Scholar]
- 11.Li F, Zhang Z, Guan J, Zhou S. Effective drug-target interaction prediction with mutual interaction neural network. Bioinformatics. 2022;38:3582–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.McCoubrey LE, Gaisford S, Orlu M, Basit AW. Predicting drug-microbiome interactions with machine learning. Biotechnol Adv. 2022;54:107797. [DOI] [PubMed] [Google Scholar]
- 13.Zimmermann M, Zimmermann-Kogadeeva M, Wegmann R, Goodman AL. Mapping human microbiome drug metabolism by gut bacteria and their genes. Nature. 2019;570:462–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Singh R. others Contrastive learning in protein language space predicts interactions between drugs and protein targets. Proc Natl Acad Sci. 2023;120:e2220778120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Ma W, Bi X, Jiang H, Wei Z, Zhang S. Annotating protein functions via fusing multiple biological modalities. Commun Biol. 2024;7:1705. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Gu Z, Luo X, Chen J, Deng M, Lai L. Hierarchical graph transformer with contrastive learning for protein function prediction. Bioinformatics. 2023;39:410. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Liu D, Liu J, Luo Y, He Q, Deng L. MGATMDA: Predicting microbe-disease associations via multi-component graph attention network. IEEE/ACM Trans Computat Biol Bioinform. 2021;19:3578–85. [DOI] [PubMed] [Google Scholar]
- 18.Long Y, Wu M, Kwoh CK, Luo J, Li X. Predicting human microbe-drug associations via graph convolutional network with conditional random field. Bioinformatics. 2020;36:4918–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Tian Z, Yu Y, Fang H, Xie W, Guo M. Predicting microbe-drug associations with structure-enhanced contrastive learning and self-paced negative sampling strategy. Brief Bioinform. 2023;24:634. [DOI] [PubMed] [Google Scholar]
- 20.Huang H, Sun Y, Lan M, Zhang H, Xie G. GNAEMDA: microbe-drug associations prediction on graph normalized convolutional network. IEEE J Biomed Health Inform. 2023;27:1635–43. [DOI] [PubMed] [Google Scholar]
- 21.Kuang H, Zhang Z, Zeng B, Liu X, Zuo H, Xu X, Wang L. A novel microbe-drug association prediction model based on graph attention networks and bilayer random forest. BMC Bioinform. 2024;25:124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Xuan P, Xu Z, Cui H. others Dynamic category-sensitive hypergraph inferring and homo-heterogeneous neighbor feature learning for drug-related microbe prediction. Bioinformatics. 2024;40:btae562. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.You Y, Chen T, Sui Y, Chen T, Wang Z, Shen Y. Graph Contrastive Learning with Augmentations. arXiv preprint 2020, arXiv:2010.13902
- 24.Sun Y-Z, Zhang D-H, Cai S-B, Ming Z, Li J-Q, Chen X. MDAD: a special resource for microbe-drug associations. Front Cellular Infect Microbiol. 2018;8:424. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Hattori M, Tanaka N, Kanehisa M. others SIMCOMP/SUBCOMP: chemical structure search servers for network analyses. Nucleic Acids Res. 2010;38:W652–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Kamneva OK. Genome composition and phylogeny of microbes predict their co-occurrence in the environment. PLOS Computat Biol. 2017;13:e1005366. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Peng W, Liu H, Dai W, Yu N, Wang J. Predicting cancer drug response using parallel heterogeneous graph convolutional networks with neighborhood interactions. Bioinformatics. 2022;38:4546–53. [DOI] [PubMed] [Google Scholar]
- 28.Wang T, Sun J, Zhao Q. Investigating cardiotoxicity related with hERG channel blockers using molecular fingerprints and graph attention mechanism. Comput Biol Med. 2023;153:106464. [DOI] [PubMed] [Google Scholar]
- 29.Meng R, Yin S, Sun J, Hu H, Zhao Q. scAAGA: single cell data analysis framework using asymmetric autoencoder with gene attention. Comput Biol Med. 2023;165:107414. [DOI] [PubMed] [Google Scholar]
- 30.Peng W, Chen T, Dai W. Predicting drug response based on multi-omics fusion and graph convolution. IEEE J Biomed Health Inform. 2021;26:1384–93. [DOI] [PubMed] [Google Scholar]
- 31.Wang Z, Nie F, Tian L, Wang R, Li X. Discriminative feature selection via a structured sparse subspace learning module. IJCAI. 2020;15:3009–15. [Google Scholar]
-
32.Nie F, Huang H, Cai X, Ding C. Efficient and robust feature selection via joint [CDATA[\ell ]]
2, 1-norms minimization. Adv Neural Inform Process Syst. 2010:23
- 33.Gan J, Hu R, Zhan M, Mo Y, Wan Y, Zhu X. Multi-view Unsupervised Graph Representation Learning. IJCAI. 2022;2987–2993.
- 34.Zhou P, Du L, Li X. Others unsupervised feature selection with adaptive multiple graph learning. Patt Recogn. 2020;105:107375. [Google Scholar]
- 35.Chen Y, Wu L, Zaki MJ. Deep Iterative and Adaptive Learning for Graph Neural Networks. arXiv preprintarXiv:1912.07832 2019.
- 36.Xuan P, Wang S, Cui H, Zhao Y, Zhang T, Wu P. Learning global dependencies and multi-semantics within heterogeneous graph for predicting disease-related lncRNAs. Brief Bioinform. 2022;23:bbac361. [DOI] [PubMed] [Google Scholar]
- 37.Chen D, O’Bray L, Borgwardt K. Structure-aware transformer for graph representation learning. International Conference on Machine Learning. 2022;3469–3489.
- 38.Wang X, Zhu M, Bo D, Cui P, Shi C, Pei J. Am-gcn: Adaptive multi-channel graph convolutional networks. Proceedings of the 26th ACM SIGKDD International conference on knowledge discovery & data mining. 2020;1243–1253.
- 39.Xuan P, Cao Y, Zhang T, Wang X, Pan S, Shen T. Drug repositioning through integration of prior knowledge and projections of drugs and diseases. Bioinformatics. 2019;35:4108–19. [DOI] [PubMed] [Google Scholar]
- 40.Huang J, Ling CX. Using AUC and accuracy in evaluating learning algorithms. IEEE Trans Knowl Data Eng. 2005;17:299–310. [Google Scholar]
- 41.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one. 2015;10:e0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Long Y, Wu M, Liu Y, Kwoh CK, Luo J, Li X. Ensembling graph attention networks for human microbe-drug association prediction. Bioinformatics. 2020;36:i779-86. [DOI] [PubMed] [Google Scholar]
- 43.Tan Y, Zou J, Kuang L, Wang X, Zeng B, Zhang Z, Wang L. GSAMDA: a computational model for predicting potential microbe–drug associations based on graph attention network and sparse autoencoder. BMC Bioinform. 2022;23:492. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Wang B, Ma F, Du X, Zhang G, Li J. Prediction of microbe-drug associations based on a modified graph attention variational autoencoder and random forest. Front Microbiol. 2024;15:1394302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Rajput A, Thakur A, Sharma S, Kumar M. aBiofilm: a resource of anti-biofilm agents and their potential implications in targeting antibiotic drug resistance. Nucleic Acids Res. 2018;46:D894–900. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Parry CM, Hien TT, Dougan G, White NJ, Farrar JJ. Ciprofloxacin-resistant Salmonella typhi and treatment failure. Lancet. 1999;353:1590–1. [DOI] [PubMed] [Google Scholar]
- 47.Suci PA, Tyler BJ. Selective killing of Aggregatibacter actinomycetemcomitans by ciprofloxacin during development of a dual species biofilm with Streptococcus sanguinis. Arch Microbiol. 2010;192:67–873. [DOI] [PubMed] [Google Scholar]
- 48.Arif W, Rana NF, Saleem I, Tanweer T, Khan MJ, Alshareef SA, Alaryani HM, Al-Kattan MO, Alatawi HA, Menaa F, Nadeem AY. Antibacterial activity of dental composite with ciprofloxacin loaded silver nanoparticles. Molecules. 2022;27:7182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Hamouda A, Amyes SGB. Novel gyrA and parC point mutations in two strains of Acinetobacter baumannii resistant to ciprofloxacin. J Antimicrob Chemoth. 2004;54:695–6. [DOI] [PubMed] [Google Scholar]
- 50.Sinel C, Cacaci M, Meignen P, Guérin F, Davies BW, Sanguinetti M, Giard J-C, Cattoir V. Subinhibitory concentrations of ciprofloxacin enhance antimicrobial resistance and pathogenicity of Enterococcus faecium. Antimicrob Agents Chemoth. 2017;61:e02763. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Lister PD, Sanders CC. Pharmacodynamics of moxifloxacin, levofloxacin and sparfloxacin against Streptococcus pneumoniae. J Antimicrob Chemoth. 2001;47:811–8. [DOI] [PubMed] [Google Scholar]
- 52.Kuzmenko AV, Gutyj BV, Sobolev OI, Gufriy DF, Darmohray LM, Gutyj OV, Guta ZA, Guta NM. Study of effectiveness of enrofloxacin and moxifloxacin in experimental salmonellosis of chickens. Sci Messeng LNU Veterin Med Biotechnol. 2015;17:61–6. [Google Scholar]
- 53.Qin Z, Sun W, Deng H, Li D, Wei Y, Lv B, Yan J, Kong L, Zhong Y. Cosformer: Rethinking Softmax in Attention. Proceedings of the International Conference on Learning Representations. 2022.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets are obtained from the previous work EGATMDA (https://github.com/uctoronto/EGATMDA) [42]. The datasets for training and testing our prediction model are freely available at https://github.com/pingxuan-hlju/SMMDA. The source code of our model is also contained by the GitHub link.


































