Skip to main content
BMC Bioinformatics logoLink to BMC Bioinformatics
. 2025 Sep 26;26:231. doi: 10.1186/s12859-025-06199-w

Structure-sensitive transformer and multi-view graph contrastive learning enhanced prediction of drug-related microbes

Ping Xuan 1,2, Rui Wang 2, Jing Gu 3, Hui Cui 4, Tiangang Zhang 1,3,✉
PMCID: PMC12465207  PMID: 41013188

Abstract

Background:

The human microbiome plays a crucial role in regulating the efficacy and toxicity of drugs as well as in developing the drugs. Therefore, predicting the drug-related microbes is beneficial for analyzing the functional mechanisms of drugs. Recently, the graph learning based methods demonstrated their advantages in extracting the node features from the biological heterogeneous graphs. However, the previous methods failed to completely preserve the intrinsic structures of biological data and did not fully utilize the topological and positional information for predicting the drug-microbe associations.

Results:

We propose a new prediction model, structure-sensitive transformer and multi-view graph contrastive learning for microbe-drug association prediction (SMMDA), to encode and integrate the topological structures, semantics, and multiple-view embedding features of the drugs and microbes. Considering the sparsity of the original features of drugs and microbes, the learnable data augmentation strategy is designed to learn their global representations. Since similar drugs are more likely to associate with the similar microbes, a structure-sensitive transformer is proposed to integrate the topology structures composed of drugs (microbes) to form the multi-view embedding features. We design two contrastive learning strategies to exploit the complementary semantics across multiple views. As the embedding features from multiple views have various semantics, we design view-level attention to adaptively integrate these features.

Conclusions:

The extensive experimental results show that SMMDA outperforms several state-of-the-art methods for predicting the drug-related candidate microbes. The ablation studies show the effectiveness of the major innovations which include the learnable data augmentation, structure-sensitive transformer-based node feature learning, and multi-view contrastive learning. The case studies on three drugs also demonstrate SMMDA’s capability in retrieving the potential microbe candidates for the drugs.

Keywords: Learnable data augmentation, Structure-sensitive transformer, Multi-view contrastive learning, View-level attention mechanism

Introduction

The human body hosts trillions of microbes, including bacteria, archaea, fungi, protozoa, and viruses [1, 2], which form distinct microbiomes. These microbiomes are distributed across various parts of the body, such as the skin, oral cavity, nasal passages, gastrointestinal tract, and reproductive and urinary systems. They play a vital role in regulating human health by protecting against pathogens and promoting metabolism [3]. However, microorganisms can also negatively affect the body [4]. Research has shown that imbalances in the microbiome can lead to conditions such as diabetes, inflammatory bowel disease, and even cancer [5, 6].

Recent studies have underscored the significant role of microbes in influencing drug activity and toxicity [7, 8]. Conversely, drugs can alter the diversity and functionality of microbial communities. For example, methicillin-resistant Staphylococcus aureus (MRSA) produces the enzyme Inline graphic-lactamase, which degrades the Inline graphic-lactam ring in antibiotics [9], contributing to drug resistance. As drug-resistant pathogens become increasingly prevalent, understanding the interactions between microbes and drugs has become essential for advancing drug development [10].

Microbe–drug associations are traditionally identified through experimental techniques, but pinpointing target microbes remains a complex and challenging task, slowing the progress of new drug discoveries. Many studies have focused on reusing existing drugs or exploring drug combinations, but the rise of drug-resistant microbes complicates these efforts [11]. This underscores an urgent need for methods capable of predicting associations between drugs and novel target microbes. Traditional wet-lab approaches, while effective, are time-consuming, labor-intensive, and costly. Computational approaches provide a complementary solution, offering efficient and accurate predictions of associations [12, 13]. Rohit et al. [14] employ contrastive learning to enhance the prediction of drug–protein target interactions. Ma et al. [15] integrates six heterogeneous biological modalities by first constructing both protein–protein interaction and protein homology networks, then applying intra-view and inter-view contrastive learning to derive enriched, highly discriminative protein embeddings. Gu et al. [16] introduces a hierarchical graph Transformer that generates semantically informed super-node embeddings, adaptively aggregates them to form comprehensive graph representations, and uses contrastive learning as a regularization technique to maximize the consistency of the same graph’s embeddings across different views. Computational methods are also highly effective in identifying potential microbe–drug associations and reliable drug candidates for wet-lab validation. Predictive models such as graph attention networks (GAT) [17] and graph convolutional networks (GCN) [18] have been utilized to infer the probability of these associations. For instance, SCSMDA [19] employs graph contrastive learning and refined meta-paths to enhance drug and microbe representations. GNAEMDA [20] employs graph normalization to enhance drug and microbe representations. H Kuang et al. [21] developed an integrated framework, GARFMDA, which constructs multimodal heterogeneous networks of microbes and various drugs from multiple biomedical data sources. This framework incorporates a GAT, which assigns different weights to edges and nodes in a heterogeneous graph, and a bilayer random forest for filtering diversity information of drugs and microbes. DHDMP [22] constructs a dynamic hypergraph by treating specific relationships among multiple drugs and microbes as hyperedges and designs a spatial cross-attention mechanism to encode the attributes of node pairs. Despite their innovations, these methods have certain limitations. GARFMDA [23] uses bilayer random forest to filter information of drugs and microbes, without considering the potential impact of random forest on disrupting the intrinsic structure of the data. For example, the removal or addition of edges could significantly alter the identity of microbial data or even the effectiveness of drugs [23]. In addition, SCSMDA [19] applys a standard GCN to extract embedding features of drugs and microbes, while overlooks the structural and feature similarities between drugs (or microbes). The hyperedges in DHDMP [22] and the spatial cross-attention mechanism focus primarily on semantic information between nodes, while ignoring the neighboring topology information of the microbes (drugs). Moreover, these approaches struggle to address the noise introduced by numerous unobserved associations in the datasets, which can compromise prediction accuracy.

We propose structure-sensitive transformer and multi-view graph contrastive learning for microbe-drug association prediction (SMMDA), a novel model that predicts candidate microbes for drugs through multi-view contrastive learning. Our approach addresses limitations in previous methods by leveraging innovative strategies to improve predictive accuracy. The key contributions of our work are as follows:

  • We developed an innovative data augmentation method that generates new graph structures and representations for drugs and microbes. This approach preserves the intrinsic structure of drug–microbe graphs, facilitating the creation of multi-view information for contrastive learning. The learned graph structures and representations are utilized to derive rich, multi-view insights into drug–microbe interactions.

  • Drugs with similar structures are more likely to have comparable effects on related microbes. To simultaneously capture structural and attribute similarities between nodes, we designed a structure-sensitive transformer (SST) that generates multi-view embedding features. This transformer considers not only node attributes but also local structures by extracting subgraph representations around each node.

  • We introduced a novel framework to learn complementary information across multi-view drug–microbe interactions. This framework employs two types of losses: one that separates positive and negative embeddings within the same graph structure, and another that differentiates embeddings across different graph structures. By integrating learnable data augmentation and multi-view contrastive learning into a unified framework, our approach supports iterative optimization.

  • Since the embedding features generated for each drug–microbe view contribute differently to the prediction of microbe–drug associations, we implemented a view-level attention mechanism to effectively integrate multi-view embeddings. Our experimental results demonstrate that SMMDA significantly outperforms existing state-of-the-art methods in predictive performance.

Materials and methods

We propose a microbe–drug association prediction model, SMMDA (Fig. 1), which consists of four main components: a learnable data augmentation module, an SST module, a multi-view contrastive learning module, and a view-level attention module. A heterogeneous graph of drug–microbe interactions is constructed to integrate similarities and associations between drugs and microbes. Through learnable data augmentation, we generate new heterogeneous graphs and features to create multi-view representations of drug–microbe interactions. The SST is employed to capture embedding features for each view by considering both node attributes and local graph structures. Multi-view contrastive learning is then applied to enhance the complementary information across the views. Finally, the view-level attention mechanism assigns adaptive weights to embedding features from different views, producing the final embedding. This final representation is used to predict association scores between drugs and microbes.

Fig. 1.

Fig. 1

The framework of the proposed SMMDA model consists of two main components: a Constructing a multi-view graph based on data augmentation and b Structure-sensitive feature representation learning with graph contrast

Dataset

The dataset used in our experiments is sourced from previous work, with raw data obtained from the MDAD database [24]. MDAD includes 5,505 clinical reports and experimentally validated associations between microbes and drugs. Many of these entries in the databases are redundant, referencing the same microbe–drug associations across different sources. To ensure data consistency, we merge entries that refer to the same drug–microbe pair. After this deduplication process, we obtain 2,470 unique associations involving 1,373 drugs and 173 microbes [19]. Specifically, drug–drug similarities were calculated as a weighted average of chemical substructure similarity and Gaussian similarity, while microbe–microbe similarities were derived using sequence-based methods [25].

Drug–drug similarity

We first build a structural similarity matrix Inline graphic using the SIMCOMP2 method [25], which compares drugs based on their chemical structures. We then compute a Gaussian interaction profile kernel similarity from the drug–microbe interaction matrix Inline graphic, capturing behavioral-level affinities. Finally, we integrate these two measures. If a drug pair has a nonzero structural similarity, we take the average of Inline graphic and Inline graphic; otherwise, we use Inline graphic alone,

graphic file with name d33e417.gif 1

where Inline graphic denotes the Gaussian interaction profile kernel similarity between drug Inline graphic and drug Inline graphic.

Microbe–microbe similarity

We first compute a functional similarity matrix Inline graphic based on microbial gene function information using the method proposed by Kamneva [26]. Next, we calculate a Gaussian interaction profile kernel similarity from the microbe–drug interaction data, denoted Inline graphic. Finally, we integrate these two measures into a composite similarity,

graphic file with name d33e461.gif 2

Construction of the drug-microbe heterogeneous graph

The matrix Inline graphic captures the chemical substructure similarity between drugs, where the Inline graphic-th row represents the similarity of drug Inline graphic with all Inline graphic drugs. Similarly, the matrix Inline graphic encodes the similarity between microbe Inline graphic and all Inline graphic microbes, with the Inline graphic-th row reflecting these relationships. The matrix Inline graphic indicates known associations between drugs and microbes: if a connection exists between drug Inline graphic and microbe Inline graphic, then Inline graphic; otherwise Inline graphic. We construct a drug–microbe heterogeneous graph Inline graphic, where node set Inline graphic includes the drug nodes Inline graphic and the microbe nodes Inline graphic. The edge set Inline graphic contains edges Inline graphic that indicate whether a link exists between nodes Inline graphic and Inline graphic. The graph Inline graphic comprises three types of edges: drug–drug similarity edges, microbe–microbe similarity edges, and drug–microbe association edges.

The adjacency matrix Inline graphic of the heterogeneous graph Inline graphic is given by,

graphic file with name d33e620.gif 3

where Inline graphic is the transpose of Inline graphic, and Inline graphic. The Inline graphic-th row of Inline graphic represents the similarity of drug Inline graphic with all Inline graphic drugs and its association with all Inline graphic microbes. Similarly, the Inline graphic-th row of Inline graphic reflects the relationships of microbe Inline graphic with the Inline graphic drugs and its similarity to the Inline graphic microbes. The diagonal elements of Inline graphic correspond to the self-similarity of drugs or microbes and are all set to 1. Additionally, the Inline graphic-th row of Inline graphic represents the attribute vector of node Inline graphic. Therefore, the adjacency matrix Inline graphic doubles as an attribute matrix, encapsulating all the attributes of drugs and microbes, such as Inline graphic.

Learnable data augmentation on drug and microbe features

Learnable node feature augmentation

The similarity between drugs is calculated based on their chemical substructures and Gaussian similarity [27]. The higher the similarity between drug Inline graphic and drug Inline graphic, the greater the likelihood that the microbes associated with drug Inline graphic will also be associated with drug Inline graphic [28–30]. To capture these key features, we propose a novel learnable node feature augmentation strategy [31, 32]. We introduce an indicator vector Inline graphic, where the elements of Inline graphic are randomly assigned as either 0 or 1 [33]. This vector Inline graphic is used to mask elements of the original attribute matrix Inline graphic, resulting in an augmented attribute matrix Inline graphic,

graphic file with name d33e824.gif 4

where Inline graphic represents element-wise multiplication. Inline graphic contains the similarity between drug Inline graphic and Inline graphic drugs, as well as its association with Inline graphic microbes, while Inline graphic represents the associations of microbe Inline graphic with all drugs and its similarity with other microbes. To avoid excessive perturbations, we design the following loss function,

graphic file with name d33e875.gif 5

where Inline graphic is the squared Frobenius norm of the matrix, and Inline graphic is a trainable parameter. This loss function ensures that Inline graphic and Inline graphic remain as consistent as possible while minimizing the Inline graphic-norm. Under a specified masking ratio, we select columns in the weight matrix Inline graphic with relatively small weights and set the corresponding elements of the vector Inline graphic to 0, while the remaining elements are set to 1.

Learnable graph structure augmentation

In conventional data augmentation methods, the graph structure typically remains unchanged, i.e., Inline graphic. However, when the generated representation Inline graphic is paired with the original adjacency matrix Inline graphic, noise and redundancy may be introduced. By using Eq. (2), we obtain a better representation than Inline graphic, which can subsequently be used to update Inline graphic. This updated Inline graphic better captures the intrinsic structure of the data. To achieve this [34, 35], we propose a learnable graph structure augmentation method that dynamically adjusts the graph structure to uncover the intrinsic relationships between drug and microbe data after augmentation. Specifically, we employ a novel graph learning method to emphasize the intrinsic structure of the data. The updated Inline graphic is computed as follows,

graphic file with name d33e978.gif 6

where Inline graphic and Inline graphic are trainable parameters, Inline graphic is the exponential function, ReLU(.) is the activation function, and Inline graphic is a row-normalized vector. The loss function for graph structure augmentation is defined as,

graphic file with name d33e1010.gif 7

where Inline graphic denotes the augmented adjacency matrix and Inline graphic is the weight of edge Inline graphic, Inline graphic is the squared Frobenius norm of the matrix, i.e., the sum of the squares of all its elements. The weighted structural difference term Inline graphic penalizes edges in the augmented graph: if the two end-nodes are structurally similar, the difference is small, so retaining that edge incurs little loss, protecting critical connections from deletion; conversely, if an edge connects nodes with dissimilar structure, the penalty is large, driving the model to reduce Inline graphic and thus avoid redundant or noisy edges. The Frobenius norm term Inline graphic acts as an Inline graphic penalty on all augmented edges, preventing excessively large weights and preserving the sparsity pattern of the augmented graph in line with the original topology. We integrate the losses from the data augmentation parts, Eq. (5) and Eq. (7), into a unified framework, as follows,

graphic file with name d33e1067.gif 8

Using the proposed learnable data augmentation method, multi-view information is generated based on the new representations Inline graphic and feature graph Inline graphic, as follows,

graphic file with name d33e1086.gif 9

In contrast to traditional data augmentation methods, our learnable approach updates both Inline graphic and Inline graphic, preventing excessive perturbation of insignificant nodes and edges, thereby preserving the structural integrity of the data [33].

Learning embedding features via structure-sensitive transformer

Building on the biological assumption that drugs with similar structures tend to have comparable effects on similar microbes, we derive the representation of a target node, whether drug or microbe, by aggregating the attributes of nodes exhibiting high structural similarity with it [27]. However, nodes with similar structures may be distant from the target node within the graph [36]. This global relationship is often overlooked by traditional GNN-based approaches, which primarily focus on local neighbors. Consequently, a global dependency exists across the attributes of all nodes in the heterogeneous network, highlighting the need to incorporate these relationships into the model.

Drawing inspiration from the Transformer model introduced in Chen et al. [37], we propose a structure-sensitive transformer to capture these global dependencies (as shown in the Fig. 2). Firstly, we use a Inline graphic-subtree extractor to focus on the structural similarity between nodes through local subgraphs. This method applies to a graph neural network (GNN) with layers to the input graph and uses the output representation of node Inline graphic as the subgraph representation of that node. Specifically, for each node, we collect all nodes within its Inline graphic-hop neighborhood (including the node itself) as well as all edges among them. As shown in Fig. 2, the 1-hop subgraph of node Inline graphic includes the set of nodes Inline graphic and the corresponding edges Inline graphic. The representation of node Inline graphic in graph Inline graphic is computed by,

graphic file with name d33e1191.gif 10

where Inline graphic, Inline graphicis the node feature matrix of Inline graphic, Inline graphic is the adjacency matrix of Inline graphic, Inline graphic is the degree matrix of Inline graphic, Inline graphic, Inline graphic, Inline graphic, Inline graphic, Inline graphic denotes the concatenation operation, Inline graphicrepresents computing a weighted row vector by taking the column-wise mean of the resulting matrix. The similarity between node features in self-attention can be replaced with the similarity between vectors that incorporate structural information from a node’s local neighborhood,

graphic file with name d33e1278.gif 11

where Inline graphic represents the inner product, Inline graphic, Inline graphic are the projection matrices for query and key, and Inline graphic is the output dimension. Thus, we can compute the structure-sensitive attention vector of node Inline graphic as follows,

graphic file with name d33e1316.gif 12

where Inline graphic is a linear value function, where Inline graphic is a trainable parameter matrix. The output of the structure-sensitive self-attention is then followed by a skip connection and a feedforward network (FFN), forming a SST layer, as follows,

graphic file with name d33e1335.gif 13

where Inline graphic, Inline graphic are parameter matrices, and Inline graphic is the degree of node Inline graphic. A multilayer stack of these structure-sensitive transformer layers constitutes the SST. Given the multi-view data (i.e., Inline graphic), we employ two SSTs (as shown in Fig. 2) to generate multi-view embeddings, as different graphs provide complementary information [33]. For instance, the original adjacency matrix Inline graphic captures real-world similarities and associations between drugs and microbes, while the augmented adjacency matrix Inline graphic collects similarities from the feature space. Specifically, since Inline graphic and Inline graphic share the same topology, SST1 is applied to Inline graphic and Inline graphic, facilitating the generation of high-quality Inline graphic. Similarly, Inline graphic and Inline graphic are fed into SST2 to output embedding features Inline graphic and Inline graphic, while also updating Inline graphic and Inline graphic,

graphic file with name d33e1460.gif 14

Fig. 2.

Fig. 2

Overview of an example SST layer that uses the k-subtree GNN extractor as its structure extractor. The structure extractor generates structure-sensitive node representations which are used to compute the query (Q) and key (K) matrices in the Transformer layer

Multi-view contrastive learning

Inline graphic and Inline graphic (Inline graphic and Inline graphic) represent the embedding features of all drugs and microbes learned under different views. The enhanced features are derived from the original features, and the enhanced graph is derived from the enhanced features. Therefore, Inline graphic and Inline graphic (Inline graphic and Inline graphic) are closely related. To uncover more comprehensive contrastive relationships, we design a contrastive loss within the same graph to differentiate positive and negative embeddings within the same SST and graph structure [33]. Simultaneously, we design a contrastive loss across different views to separate positive and negative embeddings across different graph structures, focusing on embeddings derived from different graphs.

Fig. 3.

Fig. 3

Illustration of the positive and negative pair selection strategy of SMMDA

Contrastive learning within the same view

The contrastive loss within the same view is designed to establish relationships within a single graph structure. Its objective is to bring the embedding features of the same node closer together across different views of the graph while pushing the features of one node farther from those of other nodes (whether in the same or different views) within the same structure. This is mathematically expressed as,

graphic file with name d33e1532.gif 15
graphic file with name d33e1538.gif 16

Here, Inline graphic represents the similarity between two drug (or microbe) nodes, where Inline graphic is the temperature parameter. The total contrastive loss within the same view for all drug and microbe nodes is defined as,

graphic file with name d33e1557.gif 17

Contrastive learning across different views

The goal of contrastive learning across different views is to compare embedding features between distinct graph structures, such as the original and augmented graphs. To achieve this, we design a contrastive loss that aligns the embedding features of the same drug (or microbe) node across different views, while maximizing the separation between the features of distinct nodes across these views [37]. This process is formulated as,

graphic file with name d33e1570.gif 18
graphic file with name d33e1576.gif 19

The total contrastive loss across different views, computed for all drug and microbe nodes, is defined as,

graphic file with name d33e1583.gif 20

Finally, by combining the contrastive loss within the same view and the contrastive loss across different views, the proposed multi-view contrastive learning loss is expressed as,

graphic file with name d33e1590.gif 21

The contrastive loss within the same view captures relationships within a single graph structure, while the contrastive loss across different views addresses relationships across distinct graph structures and views. By integrating these two components, Eq. (21) not only preserves the individual characteristics of each graph structure but also bridges the two graph structures. This approach enables the effective capture of complementary information between the original and augmented graph structures.

View-level attention mechanism

The embedding features of drugs (or microbes) from different perspectives may contribute differently to drug–microbe association prediction. To address this, we propose a view-level attention mechanism to integrate multi-view embeddings seamlessly [38]. Taking a drug Inline graphic as an example (the attention scores for a microbe Inline graphic, Inline graphic, can be derived similarly), the relative importance scores Inline graphic are learned using the view-level attention mechanism Inline graphic as follows,

graphic file with name d33e1635.gif 22

where Inline graphic represent the attention scores for Inline graphic, respectively, and Inline graphic denotes the attention operation. To compute the attention values, the embeddings are first transformed through a nonlinear transformation. A shared attention vector Inline graphic is used to compute the attention score Inline graphic,

graphic file with name d33e1673.gif 23

where Inline graphic is the weight matrix, and Inline graphic is the bias vector.Similarly, Inline graphic can be computed using shared parameters. Next, these values are normalized using a Softmax function to obtain the weights Inline graphic,

graphic file with name d33e1705.gif 24

where exp represents the exponential function. Similarly, Inline graphic and Inline graphic can be computed. Finally, the multi-view embeddings for drug Inline graphic are aggregated to obtain the final embedding Inline graphic,

graphic file with name d33e1736.gif 25

where Inline graphic denotes the concatenation operation. Similarly, the final embedding for microbe Inline graphic can be computed as,

graphic file with name d33e1756.gif 26

Final integration and optimization

The embeddings of drug Inline graphic, Inline graphic, and microbe Inline graphic, Inline graphic, are vertically stacked. After passing through a Inline graphic convolution layer, the combined feature vector is fed into a fully connected layer to obtain the final prediction score Inline graphic,

graphic file with name d33e1803.gif 27

where Inline graphic represents the association score between drug Inline graphic and microbe Inline graphic, with a higher score indicating a greater likelihood of an association. Inline graphic is the weight matrix of the multilayer perceptron (MLP), and Inline graphic is the flattened feature vector obtained after convolution on the stacked features.

During training, we optimize the model using backpropagation and the gradient descent algorithm. The cross-entropy loss for the model’s predictions is defined as Inline graphic,

graphic file with name d33e1849.gif 28

where Inline graphic is the ground truth label, Inline graphic is the size of the training batch, and Inline graphic. Inline graphic and Inline graphic represent the probabilities of drug Inline graphic and microbe Inline graphic being associated or not being associated, respectively. Therefore, the overall loss of our model is defined as,

graphic file with name d33e1899.gif 29

Here, Inline graphic is a tuning parameter that adjusts the impact of the learnable data augmentation and multi-view contrastive learning components on the model.

Experimental evaluations and discussions

Evaluation metrics

To assess the performance of SMMDA and other comparative methods, a five-fold cross-validation approach was utilized [39]. Known drug–microbe associations were classified as positive samples and randomly distributed into five folds, while unknown associations, which had not been validated through biological experiments, were regarded as negative samples. In each fold, four folds of positive samples, along with an equal number of randomly chosen negative samples, were employed for training. The remaining fold of positive samples, together with all unselected negative samples, constituted the test set.

If the association score between drug Inline graphic and microbe Inline graphic falls below a specified threshold Inline graphic, the pair is classified as a negative sample. Conversely, if the score exceeds the threshold, the pair is labeled as a positive sample.To assess performance, we use the area under the receiver operating characteristic curve (AUC) [40], the area under the precision-recall curve (AUPR) [41], and the recall rate of the top-k candidate microbes associated with each drug as evaluation metrics.

Biologists frequently prioritize candidate samples at the top of their ranking lists for wet-lab validation. Therefore, it is crucial to maximize the number of positive samples that appear in these top ranks. To tackle this issue, we propose incorporating the recall rate of the top-k ranked samples as an additional evaluation metric. This metric is defined as the proportion of correctly identified positive samples within the top-k list relative to the total number of positive samples.

Parameter settings

The proposed model, SMMDA, was implemented using the PyTorch machine learning framework and trained on an NVIDIA GeForce 3070 GPU to accelerate computation. The Adam optimizer was used, and hyperparameters were fine-tuned to achieve optimal performance. To systematically evaluate how the masking rate influences the learnable data augmentation process, we conducted a sensitivity analysis across five masking rates Inline graphic. As shown in Table T1, the model’s performance begins to degrade when the masking rate exceeds 0.2-excessively high masking rates result in the loss of critical information, while excessively low rates (e.g., 0.1) fail to generate sufficiently strong augmentations. The masking rate of 0.2 achieves the best performance in terms of both AUC and AUPR; thus, we set the masking rate to 0.2 in our experiments.For DHDMP, the learning rate was set to Inline graphic. Both NHCN and GCNFP adopt two-layer encoders and construct 32 hyperedges per sample. The feature dimension is set to 50, and the number of attention heads is 8. GCNMDA is evaluated using five-fold cross-validation. The number of hidden units is set to 25, the negative sampling ratio Inline graphic, the number of CRF iterations Inline graphic, and the loss function weights are set to Inline graphic, Inline graphic, and Inline graphic. GSAMDA sets the topological feature dimension to 128 and the attribute feature dimension to 32. The learning rates for both GAE and SAE modules are fixed at 0.01. MGAVAEMDA adopts a fixed learning rate of 0.01 and sets the hidden layer dimension to 128. SCSMDA achieves the best performance when the learning rate is Inline graphic, the number of positive sample pairs is 10, the number of bins in self-paced negative sampling is 10, and both the MLP and GCN module consist of a single layer.

To systematically evaluate how the masking rate influences the learnable data augmentation process, we conducted a sensitivity analysis across five masking rates (0.1, 0.2, 0.3, 0.4, and 0.5). As shown in Table 1, the model’s performance begins to degrade when the masking rate exceeds 0.2-excessively high masking rates result in the loss of critical information, while excessively low rates (e.g., 0.1) fail to generate sufficiently strong augmentations. The masking rate of 0.2 achieves the best performance in terms of both AUC and AUPR. Thus, we set the masking rate to 0.2 in our experiments.

Table 1.

Ablation experimental results of masking rate

Masking rate AUC AUPR
0.10 0.971 0.842
0.20 0.973 0.845
0.30 0.970 0.841
0.40 0.964 0.837
0.50 0.959 0.834

The best result is highlighted in bold

The tuning parameter Inline graphic, which controls the influence of matrix Inline graphic on the masking operation, was selected from the range Inline graphic. After testing, Inline graphic was found to provide the best results. For the k-hop neighborhood, Inline graphic was chosen from Inline graphic, and Inline graphic was selected for the final experiments. In the SST, the number of attention heads, Inline graphic, was set to 8, and the output dimension, Inline graphic, was set to 800. The hyperparameter Inline graphic, which governs the balance between the contributions of multi-view contrastive learning and learnable data augmentation, was tested across the range Inline graphic. The optimal value of Inline graphic was found to yield the highest model performance, with an AUC of 0.973 and an AUPR of 0.845.

Ablation studies

To evaluate the effectiveness of our data augmentation method, multi-view contrastive loss, and SST, we introduced four ablation variants Proposed-R-A, Proposed-R-Inter, Proposed-R-Intra, and Proposed-R-SST. Specifically:

  • Proposed-R-A applies random perturbation for data augmentation and does not generate new augmented graphs.

  • Proposed-R-Inter removes the contrastive loss across different views (Inline graphic).

  • Proposed-R-Intra excludes the contrastive loss within the same view (Inline graphic).

  • Proposed-R-SST replaces SST with a GCN to generate multi-view embedding features.

Table 2.

Results of the ablation studies

Method AUC AUPR
Proposed-R-A 0.955 0.805
Proposed-R-Inter 0.967 0.819
Proposed-R-Intra 0.971 0.832
Proposed-R-SST 0.956 0.810
Proposed 0.973 0.845

The best result is highlighted in bold

The final model achieved the highest AUC (0.973) and AUPR (0.845). In comparison:

  • When random perturbation was used for data augmentation instead of learnable augmentation (Proposed-R-A), the AUC and AUPR decreased by 1.8% and 4.0%, respectively.

  • Removing the contrastive loss across different views (Proposed-R-Inter) resulted in a 0.6% decrease in AUC and a 2.6% decrease in AUPR.

  • Removing the contrastive loss within the same view (Proposed-R-Intra) caused the AUC and AUPR to drop by 0.2% and 1.3%, respectively.

  • Replacing SST with GCN (Proposed-R-SST) led to a 1.7% decline in AUC and a 3.5% decline in AUPR.

The ablation studies demonstrate that the learnable data-augmentation module yields the largest gain by dynamically masking regions based on node importance, preserving critical semantics and graph structure far more effectively than random augmentations. Structure-aware subgraph extraction (SST) is the second most impactful: without SST, the model relies solely on node attributes and ignores topology, whereas SST’s subgraph snapshots and attention mechanism capture vital local connectivity. Finally, both contrastive losses are essential: contrastive loss within the same view enforces consistency within each augmented view to counteract noise-induced drift, while contrastive loss across different views aligns representations across different views to harvest complementary information from multiple graph structures.

We projected an equal number of positive and negative drug–microbe pairs into two-dimensional space (Fig. 4). Without contrastive learning, the two classes appear loosely scattered; after contrastive learning, both positive and negative samples form tight clusters, markedly improving their separability.

Fig. 4.

Fig. 4

t-SNE visualization of drug–microbe pair embeddings with and without contrastive learning

We conducted two cold-start experiments to assess SMMDA’s generalization: in the drug cold-start, test-set drugs are completely excluded from training so that the model must predict associations for entirely unseen drugs; in the microbe cold-start, test-set microbes are similarly withheld during training, forcing the model to generalize to novel microbial targets.

Comparison with other models

SMMDA was evaluated against six leading methods for microbe–drug association prediction: GCNMDA [18], EGATMDA [42], GSAMDA [43], SCSMDA [19], MGAVAEMDA [44], and DHDMP [22]. All methods, including SMMDA, were trained and tested using five-fold cross-validation on the same dataset splits, with hyperparameters set according to the respective publications. Below is a brief description of the comparison methods,

  • GCNMDA [18]: This method is based on a GCN and combines drug chemical information, microbe genetic information, and Gaussian interaction profiles to quantify similarity. It uses a random walk preprocessing scheme to capture features for predicting microbe–drug associations.

  • EGATMDA [42]: This method constructs a microbe–disease–drug network and employs a hierarchical attention mechanism to predict microbe–drug associations.

  • GSAMDA [43]: This model calculates drug (microbe) similarities using Gaussian and Hamming interaction profiles and employs a GAT along with a sparse autoencoder to learn node features.

  • SCSMDA [19]: This method builds a microbe–drug network using gene sequence data, Gaussian kernel interaction profiles, and drug chemical structures, applying graph contrastive learning to extract features for microbe and drug nodes.

  • MGAVAEMDA [44]: This computational framework integrates an improved graph attention variational autoencoder with biological information to predict potential microbe–drug associations.

  • DHDMP [22]: This method leverages dynamic hypergraph modeling and long-distance spatial correlation encoding to predict associations between drugs and microbes.

We first calculated the AUC and AUPR for all models and then computed the average AUC and AUPR across 1373 drugs. Table 3 shows that SMMDA achieved the highest average AUC of 0.973, surpassing the second-best model, DHDMP, by 1.4%, EGATMDA, by 3.3%, GCNMDA by 7.0%, MGAVAEMDA by 3.7%, GSAMDA by 7.1%, and SCSMDA by 5.7%. Additionally, SMMDA achieved the highest average AUPR of 84.5%, exceeding DHDMP by 2.2%, SCSMDA by 50.5%, GCNMDA by 53.0%, EGATMDA by 53.8%, MGAVAEMDA by 51.7%, and GSAMDA by 59.8%.

Table 3.

AUCs and AUPRs of different methods in comparison across all the 1373 drugs

Networks AUC (%) AUPR (%)
GCNMDA 90.3 31.5
EGATMDA 94.0 30.7
GSAMDA 90.2 24.7
SCSMDA 91.6 34.0
MGAVAEMDA 93.6 32.8
DHDMP 95.9 82.3
SMMDA 97.3 84.5

The best result is highlighted in bold

As shown Tables 4 and 5, we compare SMMDA against other methods under two cold-start scenarios. Although all approaches experience declines in AUC and AUPR, SMMDA remains highly robust and continues to deliver the best performance.

Table 4.

Drug cold-start evaluation

Networks AUC (%) AUPR (%)
GCNMDA 83.0 21.5
EGATMDA 88.5 26.4
GSAMDA 81.9 18.6
SCSMDA 84.2 22.1
MGAVAEMDA 87.1 25.5
DHDMP 91.8 67.2
SMMDA 93.3 70.7

The best result is highlighted in bold

Table 5.

Microbe cold-start evaluation

Networks AUC (%) AUPR (%)
GCNMDA 82.7 20.8
EGATMDA 88.9 27.0
GSAMDA 82.1 19.2
SCSMDA 84.5 23.0
MGAVAEMDA 86.8 26.0
DHDMP 92.0 68.0
SMMDA 93.5 71.5

The best result is highlighted in bold

Figure 5 illustrates the average recall rates of all drugs at different top-k candidate microbes. SMMDA, which leverages contrastive learning to enhance features with global information, outperformed all other methods across different top-k thresholds. For Inline graphic, our model achieved the highest recall rate of 89. 3%, compared to 1. 2% for the second best model, DHDMP. MGAVAEMDA achieved the fifth-best recall rate of 45.2%, which was 1.9% lower than SCSMDA and 2.5% lower than EGATMDA.

Fig. 5.

Fig. 5

The average recalls of drugs at different top Inline graphic settings

When Inline graphic, Inline graphic, and Inline graphic, SMMDA achieved the highest recall rates of 89.6%, 90.1%, and 90.3%, respectively. The second-best model, DHDMP, showed recall rates of 89.1%, 89.4%, and 89.5% at these thresholds. SCSMDA achieved recall rates of 65.4%, 72.7%, and 79%, outperforming MGAVAEMDA, which had recall rates of 61.5%, 66.8%, and 70.8%. EGATMDA performed better than both, with recall rates of 67.8%, 74.9%, and 80.1%. GSAMDA had lower recall rates of 53.6%, 67.6%, and 71.4%, but still exceeded GCNMDA, which achieved recall rates of 55.7%, 63.7%, and 68.4%.

Case studies on three drugs

To assess the effectiveness of SMMDA in identifying microbial candidates linked to specific medications, we performed case studies on ciprofloxacin, moxifloxacin, and ceftazidime. Ciprofloxacin is utilized for treating skin infections, typhoid fever, pneumonia, endocarditis, and various other bacterial infections. Moxifloxacin is commonly prescribed for pneumonia, tuberculosis, sinusitis, and chronic bronchitis. Ceftazidime, a third-generation cephalosporin antibiotic, is primarily employed in the treatment of pneumonia, sepsis, urinary tract infections, skin and soft tissue infections, and intra-abdominal infections.

In these case studies, the model utilized all known microbe–drug associations, supplemented by an equal number of randomly chosen unobserved associations to serve as negative samples. For each drug, the model produced a list of potential microbes, from which the top 20 ranked candidates were further examined.These case studies offer valuable insights into the predictive performance and potential utility of the SMMDA model in real-world applications. By comparing the model’s predicted microbe–drug associations with known established associations, the accuracy and reliability of the model in identifying drug-associated microbial candidates are further validated.

The MDAD [24] dataset contains information on known interactions between 1,373 drugs and 173 microbes. The aBiofilm [45] database focuses on the relationship between bacterial biofilm formation and drug resistance, providing data on bacterial responses to various drugs in the biofilm state, encompassing 1,720 drugs and 140 microbes. We validated the microbe–drug association predictions of SMMDA using the MDAD and aBiofilm datasets, as well as relevant literature. Among the top 20 candidate microbes associated with ciprofloxacin (Table 3), four were found in the MDAD and aBiofilm datasets, indicating their known relationship with ciprofloxacin, while 16 were further confirmed by literature evidence. For example, several microbes, including Salmonella Typhi [46], Streptococcus sanguinis [47], Streptococcus mutans [48], and Acinetobacter baumannii [49], were reported to be inhibited or killed by ciprofloxacin. Additionally, Enterococcus faecium was identified as resistant to ciprofloxacin [50].

Table 6.

Top 20 candidate microbes associated with Ciprofloxacin

Rank Microbe name Evidence Rank Microbe name Evidence
1 Streptococcus mutans aBiofilm, MDAD 11 Vibrio cholerae 26271050
2 Pseudomonas aeruginosa aBiofilm, MDAD 12 Burkholderia multivorans 19633000
3 Stenotrophomonas maltophilia aBiofilm, MDAD 13 Candida tropicalis 16849719
4 Human immunodeficiency virus 1 aBiofilm, MDAD 14 Salmonella Typhi 10334265
5 Acinetobacter baumannii 20138741 15 Enterococcus faecium 20006472
6 Propionibacterium acnes 25541476 16 Klebsiella pneumoniae 10858336
7 Streptococcus sanguinis 11347679 17 Human herpesvirus 1 Unconfirmed
8 Proteus mirabilis 22958285 18 Serratia liquefaciens 10965096
9 Enterococcus faecalis 27790716 19 Bacillus anthracis 12821500
10 Streptococcus sanguis 11347679 20 Actinomyces oris Unconfirmed

Table 7.

Top 20 candidate microbes associated with Moxifloxacin

Rank Microbe name Evidence Rank Microbe name Evidence
1 Stenotrophomonas maltophilia aBiofilm, MDAD 11 Pseudomonas aeruginosa 31691651
2 Staphylococcus aureus aBiofilm, MDAD 12 Escherichia coli 31542319
3 Listeria monocytogenes aBiofilm, MDAD 13 Vibrio harveyi 27247095
4 Haemophilus influenzae aBiofilm, MDAD 14 Klebsiella pneumoniae 27257956
5 Bacillus subtilis aBiofilm, MDAD 15 Acinetobacter baumannii 20006472
6 Listeria monocytogenes aBiofilm, MDAD 16 Burkholderia cenocepacia 27799222
7 Burkholderia multivorans 35754328 17 Salmonella enterica 15078598
8 Actinomyces oris 10858336 18 Clostridium perfringens 29486533
9 Staphylococcus epidermidis 28481197 19 Burkholderia pseudomallei 15731198
10 Candida albicans 31471074 20 Serratia marcescens Unconfirmed

Table 8.

Top 20 candidate microbes associated with Ceftazidime

Rank Microbe name Evidence Rank Microbe name Evidence
1 Pseudomonas aeruginosa aBiofilm,MDAD 11 Escherichia coli 37574665
2 Acinetobacter baumannii aBiofilm,MDAD 12 Candida albicans Unconfirmed
3 Haemophilus influenzae 6376458 13 Staphylococcus epidermis 1730894
4 Shigella flexneri 31519769 14 Klebsiella planticola Unconfirmed
5 Pseudomonas aeruginosa 34990760 15 Candida spp 6357068
6 Bacillus subtilis 31420587 16 Eikenella corrodens Unconfirmed
7 Mycobacterium tuberculosis 20138741 17 Stenotrophomonas maltophilia 37615040
8 Streptococcus pneumoniae serotype 4 8126192 18 Candida tropicalis Unconfirmed
9 Mycobacterium avium 28922808 19 Salmonella enterica 19861080
10 Proteus vulgaris 19802966 20 Pseudoalteromonas sp 24031945

For the candidate microbes associated with moxifloxacin (Table 4), five were found in the MDAD and aBiofilm datasets, and 14 were supported by literature. For instance, moxifloxacin [51] exhibits antibacterial activity against Streptococcus pneumoniae, while Salmonella enterica [52] shows resistance to moxifloxacin. For the candidate microbes associated with ceftazidime (Table 5), two were present in the MDAD and aBiofilm datasets, and 14 candidates, including Shigella flexneri, were supported by literature. Among all 60 candidate microbes, six remain unconfirmed, indicating no supporting evidence for their associations with the target drugs. These findings demonstrate that SMMDA effectively identifies potential candidate microbes for the drugs under investigation.

Prediction of novel drug–microbe associations

SMMDA was comprehensively trained on all established drug–microbe associations to forecast potential candidate microbes for each drug. The top 20 predicted associations for every drug are provided in the supplementary Table ST1, which could assist biologists in identifying promising candidate microbes for further investigation.

Conclusion

We present a novel microbe–drug association prediction model that integrates both attribute and semantic information from drug and microbe nodes across multiple views using multi-view learning. This approach aims to predict microbe associations with drugs. Dynamic data augmentation is applied by selecting important features and perturbing nodes and edges. The SST model captures neighborhood structural features of drugs and microbes by extracting k-subtree information from nodes. Graph contrastive learning enhances the representation of drug and microbe nodes across multiple views by leveraging complementary information from two loss functions. A view-level attention mechanism assigns higher weights to more significant drug and microbe view features. Cross-validation experiments on public datasets demonstrate that SMMDA outperforms comparative methods in both AUC and AUPR. Additionally, the model’s average drug recall rate and case study analysis further confirm that SMMDA reliably identifies microbe candidates associated with drugs.

Limitations

In constructing drug–drug and microbe–microbe similarity matrices, we fuse Gaussian-kernel and structure/function measures to build drug–drug and microbe–microbe similarities, but only 2,470 known associations make these estimates noisy. To mitigate sparsity, we apply learnable masking augmentation and contrastive learning to align multi-view representations of each entity. However, by focusing only on within-entity consistency, we miss cross-entity functional relationships and finer similarity patterns. Therefore, in the next step, we will apply thresholding to retain only high-confidence similarities and explore richer positive/negative sampling strategies that incorporate drug-drug and microbe-microbe functional associations into the contrastive framework, thereby mining more nuanced similarity structures.

SMMDA is built based on a transformer backbone. Our small- subgraph extractor adds negligible cost versus self-attention, while the learnable augmentation and contrastive modules incur extra overheads, yielding a model of high computational complexity. A key future direction is to cut the high memory and time overhead of self-attention. Recent work on “linear Transformers”, such as Qin et al. [53] achieves true Inline graphic time and space complexity, offering a promising approach to mitigate SMMDA’s computational demands.

Additional file

Supplementary file 1. (805.5KB, xlsx)

Acknowledgements

Not Applicable.

Author contributions

Ping Xuan: Designed the method and participated in manuscript writing. Rui Wang: Designed the experiments and participated in manuscript writing. Jing Gu: Participated in experiment design and manuscript writing. Hui Cui: Participated in experiment design. Tiangang Zhang: Participated in method design and manuscript writing.

Funding

This work was supported by Natural Science Foundation of Heilongjiang Province (LH2023F044). Natural Science Foundation of China (62172143, 62372282). Guangdong Basic and Applied Basic Research Foundation (2024A1515010176). STU Scientific Research Initiation Grant (NTF22032).

Data availability

The datasets are obtained from the previous work EGATMDA (https://github.com/uctoronto/EGATMDA)  [42]. The datasets for training and testing our prediction model are freely available at https://github.com/pingxuan-hlju/SMMDA. The source code of our model is also contained by the GitHub link.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no Competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12859-025-06199-w.

References

  • 1.THMP, Consortium. Structure, function and diversity of the healthy human microbiome. Nature. 2012;486:207–14. [DOI] [PMC free article] [PubMed]
  • 2.Thiele I, Heinken A, Fleming RM. A systems biology approach to studying the role of microbes in human health. Curr Opin Biotechnol. 2013;24:4–12. [DOI] [PubMed] [Google Scholar]
  • 3.ElRakaiby M, Dutilh BE, Rizkallah MR, Boleij A, Cole JN, Aziz RK. Pharmacomicrobiomics: the impact of human microbiome variations on systems pharmacology and personalized therapeutics. Omics J Integrat Biol. 2014;18:402–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Sprockett D, Fukami T, Relman DA. Role of priority effects in the early-life assembly of the gut microbiota. Nature Rev Gastroenterol Hepatol. 2018;15:197–205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.de la Cuesta-Zuluaga J, Huus KE, Youngblut ND, Escobar JS, Ley RE. Obesity is the main driver of altered gut microbiome functions in the metabolically unhealthy. Gut Microbes. 2023;15:2246634. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Durack J, Lynch SV. The gut microbiome: relationships with disease and opportunities for therapy. J Experim Med. 2019;216:20–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Nejman D, Livyatan I, Fuks G, Gavert N, Zwang Y, Geller LT, Rotter-Maskowitz A, Weiser R, Mallel G, Gigi E. others The human tumor microbiome is composed of tumor type-specific intracellular bacteria. Science. 2020;368:973–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Algavi YM, Borenstein E. A data-driven approach for predicting the impact of drugs on the human microbiome. Nature Commun. 2023;14:3614. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Vestergaard M, Frees D, Ingmer H. Antibiotic resistance and the MRSA problem. Microbiol Spect. 2019;7:10–1128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Hughes D, Andersson DI. Evolutionary trajectories to antibiotic resistance. Ann Rev Microbiol. 2017;71:579–96. [DOI] [PubMed] [Google Scholar]
  • 11.Li F, Zhang Z, Guan J, Zhou S. Effective drug-target interaction prediction with mutual interaction neural network. Bioinformatics. 2022;38:3582–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.McCoubrey LE, Gaisford S, Orlu M, Basit AW. Predicting drug-microbiome interactions with machine learning. Biotechnol Adv. 2022;54:107797. [DOI] [PubMed] [Google Scholar]
  • 13.Zimmermann M, Zimmermann-Kogadeeva M, Wegmann R, Goodman AL. Mapping human microbiome drug metabolism by gut bacteria and their genes. Nature. 2019;570:462–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Singh R. others Contrastive learning in protein language space predicts interactions between drugs and protein targets. Proc Natl Acad Sci. 2023;120:e2220778120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Ma W, Bi X, Jiang H, Wei Z, Zhang S. Annotating protein functions via fusing multiple biological modalities. Commun Biol. 2024;7:1705. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Gu Z, Luo X, Chen J, Deng M, Lai L. Hierarchical graph transformer with contrastive learning for protein function prediction. Bioinformatics. 2023;39:410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Liu D, Liu J, Luo Y, He Q, Deng L. MGATMDA: Predicting microbe-disease associations via multi-component graph attention network. IEEE/ACM Trans Computat Biol Bioinform. 2021;19:3578–85. [DOI] [PubMed] [Google Scholar]
  • 18.Long Y, Wu M, Kwoh CK, Luo J, Li X. Predicting human microbe-drug associations via graph convolutional network with conditional random field. Bioinformatics. 2020;36:4918–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tian Z, Yu Y, Fang H, Xie W, Guo M. Predicting microbe-drug associations with structure-enhanced contrastive learning and self-paced negative sampling strategy. Brief Bioinform. 2023;24:634. [DOI] [PubMed] [Google Scholar]
  • 20.Huang H, Sun Y, Lan M, Zhang H, Xie G. GNAEMDA: microbe-drug associations prediction on graph normalized convolutional network. IEEE J Biomed Health Inform. 2023;27:1635–43. [DOI] [PubMed] [Google Scholar]
  • 21.Kuang H, Zhang Z, Zeng B, Liu X, Zuo H, Xu X, Wang L. A novel microbe-drug association prediction model based on graph attention networks and bilayer random forest. BMC Bioinform. 2024;25:124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Xuan P, Xu Z, Cui H. others Dynamic category-sensitive hypergraph inferring and homo-heterogeneous neighbor feature learning for drug-related microbe prediction. Bioinformatics. 2024;40:btae562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.You Y, Chen T, Sui Y, Chen T, Wang Z, Shen Y. Graph Contrastive Learning with Augmentations. arXiv preprint 2020, arXiv:2010.13902
  • 24.Sun Y-Z, Zhang D-H, Cai S-B, Ming Z, Li J-Q, Chen X. MDAD: a special resource for microbe-drug associations. Front Cellular Infect Microbiol. 2018;8:424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Hattori M, Tanaka N, Kanehisa M. others SIMCOMP/SUBCOMP: chemical structure search servers for network analyses. Nucleic Acids Res. 2010;38:W652–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Kamneva OK. Genome composition and phylogeny of microbes predict their co-occurrence in the environment. PLOS Computat Biol. 2017;13:e1005366. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Peng W, Liu H, Dai W, Yu N, Wang J. Predicting cancer drug response using parallel heterogeneous graph convolutional networks with neighborhood interactions. Bioinformatics. 2022;38:4546–53. [DOI] [PubMed] [Google Scholar]
  • 28.Wang T, Sun J, Zhao Q. Investigating cardiotoxicity related with hERG channel blockers using molecular fingerprints and graph attention mechanism. Comput Biol Med. 2023;153:106464. [DOI] [PubMed] [Google Scholar]
  • 29.Meng R, Yin S, Sun J, Hu H, Zhao Q. scAAGA: single cell data analysis framework using asymmetric autoencoder with gene attention. Comput Biol Med. 2023;165:107414. [DOI] [PubMed] [Google Scholar]
  • 30.Peng W, Chen T, Dai W. Predicting drug response based on multi-omics fusion and graph convolution. IEEE J Biomed Health Inform. 2021;26:1384–93. [DOI] [PubMed] [Google Scholar]
  • 31.Wang Z, Nie F, Tian L, Wang R, Li X. Discriminative feature selection via a structured sparse subspace learning module. IJCAI. 2020;15:3009–15. [Google Scholar]
  • 32.Nie F, Huang H, Cai X, Ding C. Efficient and robust feature selection via joint [CDATA[\ell ]]Inline graphic2, 1-norms minimization. Adv Neural Inform Process Syst. 2010:23
  • 33.Gan J, Hu R, Zhan M, Mo Y, Wan Y, Zhu X. Multi-view Unsupervised Graph Representation Learning. IJCAI. 2022;2987–2993.
  • 34.Zhou P, Du L, Li X. Others unsupervised feature selection with adaptive multiple graph learning. Patt Recogn. 2020;105:107375. [Google Scholar]
  • 35.Chen Y, Wu L, Zaki MJ. Deep Iterative and Adaptive Learning for Graph Neural Networks. arXiv preprintarXiv:1912.07832 2019.
  • 36.Xuan P, Wang S, Cui H, Zhao Y, Zhang T, Wu P. Learning global dependencies and multi-semantics within heterogeneous graph for predicting disease-related lncRNAs. Brief Bioinform. 2022;23:bbac361. [DOI] [PubMed] [Google Scholar]
  • 37.Chen D, O’Bray L, Borgwardt K. Structure-aware transformer for graph representation learning. International Conference on Machine Learning. 2022;3469–3489.
  • 38.Wang X, Zhu M, Bo D, Cui P, Shi C, Pei J. Am-gcn: Adaptive multi-channel graph convolutional networks. Proceedings of the 26th ACM SIGKDD International conference on knowledge discovery & data mining. 2020;1243–1253.
  • 39.Xuan P, Cao Y, Zhang T, Wang X, Pan S, Shen T. Drug repositioning through integration of prior knowledge and projections of drugs and diseases. Bioinformatics. 2019;35:4108–19. [DOI] [PubMed] [Google Scholar]
  • 40.Huang J, Ling CX. Using AUC and accuracy in evaluating learning algorithms. IEEE Trans Knowl Data Eng. 2005;17:299–310. [Google Scholar]
  • 41.Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one. 2015;10:e0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Long Y, Wu M, Liu Y, Kwoh CK, Luo J, Li X. Ensembling graph attention networks for human microbe-drug association prediction. Bioinformatics. 2020;36:i779-86. [DOI] [PubMed] [Google Scholar]
  • 43.Tan Y, Zou J, Kuang L, Wang X, Zeng B, Zhang Z, Wang L. GSAMDA: a computational model for predicting potential microbe–drug associations based on graph attention network and sparse autoencoder. BMC Bioinform. 2022;23:492. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wang B, Ma F, Du X, Zhang G, Li J. Prediction of microbe-drug associations based on a modified graph attention variational autoencoder and random forest. Front Microbiol. 2024;15:1394302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Rajput A, Thakur A, Sharma S, Kumar M. aBiofilm: a resource of anti-biofilm agents and their potential implications in targeting antibiotic drug resistance. Nucleic Acids Res. 2018;46:D894–900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Parry CM, Hien TT, Dougan G, White NJ, Farrar JJ. Ciprofloxacin-resistant Salmonella typhi and treatment failure. Lancet. 1999;353:1590–1. [DOI] [PubMed] [Google Scholar]
  • 47.Suci PA, Tyler BJ. Selective killing of Aggregatibacter actinomycetemcomitans by ciprofloxacin during development of a dual species biofilm with Streptococcus sanguinis. Arch Microbiol. 2010;192:67–873. [DOI] [PubMed] [Google Scholar]
  • 48.Arif W, Rana NF, Saleem I, Tanweer T, Khan MJ, Alshareef SA, Alaryani HM, Al-Kattan MO, Alatawi HA, Menaa F, Nadeem AY. Antibacterial activity of dental composite with ciprofloxacin loaded silver nanoparticles. Molecules. 2022;27:7182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Hamouda A, Amyes SGB. Novel gyrA and parC point mutations in two strains of Acinetobacter baumannii resistant to ciprofloxacin. J Antimicrob Chemoth. 2004;54:695–6. [DOI] [PubMed] [Google Scholar]
  • 50.Sinel C, Cacaci M, Meignen P, Guérin F, Davies BW, Sanguinetti M, Giard J-C, Cattoir V. Subinhibitory concentrations of ciprofloxacin enhance antimicrobial resistance and pathogenicity of Enterococcus faecium. Antimicrob Agents Chemoth. 2017;61:e02763. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Lister PD, Sanders CC. Pharmacodynamics of moxifloxacin, levofloxacin and sparfloxacin against Streptococcus pneumoniae. J Antimicrob Chemoth. 2001;47:811–8. [DOI] [PubMed] [Google Scholar]
  • 52.Kuzmenko AV, Gutyj BV, Sobolev OI, Gufriy DF, Darmohray LM, Gutyj OV, Guta ZA, Guta NM. Study of effectiveness of enrofloxacin and moxifloxacin in experimental salmonellosis of chickens. Sci Messeng LNU Veterin Med Biotechnol. 2015;17:61–6. [Google Scholar]
  • 53.Qin Z, Sun W, Deng H, Li D, Wei Y, Lv B, Yan J, Kong L, Zhong Y. Cosformer: Rethinking Softmax in Attention. Proceedings of the International Conference on Learning Representations. 2022.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary file 1. (805.5KB, xlsx)

Data Availability Statement

The datasets are obtained from the previous work EGATMDA (https://github.com/uctoronto/EGATMDA)  [42]. The datasets for training and testing our prediction model are freely available at https://github.com/pingxuan-hlju/SMMDA. The source code of our model is also contained by the GitHub link.


Articles from BMC Bioinformatics are provided here courtesy of BMC

RESOURCES