Abstract
Artificial intelligence holds great promise for the design of antimicrobial peptides (AMPs); however, current models face limitations in generating AMPs with sufficient novelty and diversity, and they are rarely applied to the generation of antifungal peptides. Here, we develop an alternative pipeline grounded in a diffusion model and molecular dynamics for the de novo design of AMPs. The peptides generated by our pipeline have lower similarity and identity than those of other reported methodologies. Among the 40 peptides synthesized for an experimental validation, 25 exhibit either antibacterial or antifungal activity. AMP-29 shows selective antifungal activity against Candida glabrata and in vivo antifungal efficacy in a murine skin infection model. AMP-24 exhibits potent in vitro activity against Gram-negative bacteria and in vivo efficacy against both skin and lung Acinetobacter baumannii infection models. The proposed approach offers a pipeline for designing diverse AMPs to counteract the threat of antibiotic resistance.
A pipeline using latent diffusion model generates antimicrobial peptides to combat antibiotic resistance.
INTRODUCTION
The advent of antibiotics revolutionized the battle against microbial infections (1). However, the rampant misuse of these drugs has led to widespread antibiotic resistance (2). This resistance poses a substantial challenge to global health care systems, and antibiotic-resistant infections claim the lives of 700,000 individuals annually (3). By 2050, antibiotic-resistant infections may eclipse cancer as the leading cause of death (4). To combat the problem of microbial resistance to antibiotics, alternative antimicrobial agents are needed.
Antimicrobial peptides (AMPs) primarily interact with cell membranes via electrostatic interactions, disrupting membrane structures and leading to cell deaths (5). Compared with traditional single-target antibiotics, AMPs have a lower tendency to induce drug resistance in their target microorganisms due to their rapid, potent membrane activity and diverse inhibitory mechanisms (6). Consequently, AMPs are excellent prototypes for the development of alternative antimicrobial agents, such as the approved drugs daptomycin and polymyxin.
A formidable challenge encountered when designing AMPs is the immense chemical space that must be explored, which is time-consuming and costly (7). Recent advances in artificial intelligence (AI) have led to promising results in terms of AMP designs (8). AI-based methodologies for the design of AMPs can be broadly classified into two distinct strategies (9). The first strategy entails the use of one or more predictive models designed to ascertain the key properties of a given peptide sequence (10–12). By harnessing the power of these models, researchers can effectively filter existing peptide libraries, randomly generated sequences, proteomes, genomes, and microbiome data to identify promising candidate peptide sequences for further investigation (4, 13–15). While the computational time required for inferring individual peptides is minimal, as the peptide length increases, the corresponding chemical space expands exponentially, rendering global searches computationally impractical (16). The second strategy involves the development of a generative model that is adept at producing peptide sequences with particular attributes (17). Compared with the first strategy, generative model–based methods can explore larger spaces with lower computational costs.
In recent years, diffusion models, a type of generative model, have exhibited remarkable performance across a diverse range of applications (18), particularly in image synthesis (19) and the design of protein structures and sequences. These models are delineated into two principal approaches, direct diffusion and diffusion through latent spaces, based on their modeling strategies. Gruver et al. (20) and Alamdari et al. (21) directly designed proteins within the sequence space. Simultaneously, Zhang et al. (22) introduced PRO-LDM, a method that enables the conditional generation of protein sequences through a latent framework. The fundamental principle guiding these methodologies is the diffusion process, whereby base data are gradually transformed into noise adhering to a Gaussian distribution; this is described as the forward process (23). In tandem with this approach, such models learn the inverse process of the diffusion process, aiding in the reinterpretation of Gaussian noise into required samples (24). This innovative integration of bidirectional processes enables diffusion models to yield high-quality, diverse samples from complex data distributions (25).
In this research, we developed a pipeline that is capable of acquiring AMP sequences by integrating diffusion models with molecular dynamics. Our pipeline encompasses two key phases: generation and filtration. In the generative phase, we used a variational autoencoder (VAE) (26) that can map peptide sequences with variable lengths to uniform-dimensional latent variables (Fig. 1). Then, we conditionally generated latent variables for the candidate AMPs with the help of a diffusion model, and the VAE decoded the latent variables to generate peptide sequences. Compared with the previously reported VAE and generative adversarial network methodologies for AMP generation, our pipeline achieved markedly better performance, as evidenced by the greater novelty, greater diversity, and larger chemical space of the generated peptide sequences. The filtering phase included a three-step strategy comprising classifier predictions, sequence clustering, and molecular dynamics simulations.
Fig. 1. The process of generating AMPs using a latent diffusion model.
(A) Training set acquisition. (B) Encoding of peptide sequences as latent variables. (C) Model training. (D) Generation of candidate peptides using the latent diffusion model. (E) Pipeline for further filtering the candidate AMPs.
To verify the performance of this pipeline, we obtained 40 candidate peptide sequences with the help of the pipeline. An experimental validation revealed that 25 peptides had antibacterial or antifungal activity. Nine were particularly active [minimum inhibitory concentration (MIC) ≤ 12.5 μM] and were further investigated as potential therapeutic agents. This potent, innovative strategy for discovering AMPs has the potential to make valuable contributions for addressing the pressing challenge of antimicrobial resistance in contemporary health care.
RESULTS
Pipeline for generating and filtering AMP candidates
Our pipeline encompasses two primary stages: generating candidate sequences and subsequently performing filtration (27). Diffusion models represent a novel and groundbreaking category within the family of deep generative models, demonstrating exceptional efficacy across a broad spectrum of domains (28). In this study, we opted for latent diffusion models to generate candidate AMP sequences (fig. S1). A latent diffusion model comprises two principal components: an encoder-decoder structure and a diffusion model. The two principal components were trained sequentially. After training the VAE, the latent variables generated by the VAE were used to train the diffusion model. The encoder-decoder structure is tasked with acquiring the latent variables of peptides and decoding these latent variables into peptide sequences. Moreover, the diffusion model is used to generate the latent variables of peptides.
To identify a suitable encoder-decoder for generating peptide sequences, we analyzed various types of encoder-decoder structures in terms of their ability to produce latent variables (26, 29). The training dataset for these encoder-decoder structures consisted of all peptide sequences from the UniProt database (30) that were fewer than 50 amino acids in length (see Materials and Methods). The reconstruction error reflects, at least in part, the representational learning capacity of the diffusion model. The reconstruction error is, therefore, used as a criterion to select the most suitable encoder and decoder for the diffusion model. Because the SD of the latent variables was scaled to one, matching the Gaussian noise distribution used in the training process of diffusion models, a smaller SD of the latent variables requires a larger scaling factor. According to the reconstruction error formula (fig. S2), a larger scaling factor leads to a smaller reconstruction error in the diffusion model. When diffusion models are trained with latent variables (its SD has been scaled to 1) generated by different encoder-decoder structures, the resulting losses of diffusion model eventually converge to similar values. Among these encoder-decoder structures (fig. S2), transformer-based VAE (31) using a larger scaling factor achieved smallest reconstruction error for the diffusion model, thereby guaranteeing the diffusion model’s reconstruction accuracy. In addition, we assessed the changes in the Euclidean distance between the latent variables of the original peptide sequences and those generated after randomly mutating one to five amino acids within a peptide sequence. The latent variables corresponding to the peptide sequences with similar mutations generated by transformer-based VAEs had the smallest Euclidean distances (fig. S3). On the basis of these findings, we selected a transformer-based VAE as the encoder and decoder for the latent diffusion model.
Given the feature extraction efficiency and the strong adaptability of bidirectional encoder representations from transformers (BERT), we opted for a BERT encoder as the backbone of our diffusion model. We implemented a dual-phase training regimen, incorporating both pretraining and fine-tuning techniques to augment the performance and extend the versatility of the diffusion models. This strategy has proven effective in many applications, such as natural language processing and computer vision (25). For the pretraining process, we extracted 12,000 peptide sequences of each length from the VAE training set, resorting to the full dataset for categories with fewer than 12,000 sequences. During the fine-tuning phase, our training data comprised antimicrobial (both antibacterial and antifungal) peptides collected from multiple databases, as well as non-AMPs that were filtered on the basis of certain criteria to match the length distribution of the AMP data (refer to Materials and Methods for more details). Both the pretraining and fine-tuning stages used the same training methodology, differing primarily in the application of conditional constraints. Specifically, during the pretraining phase, the model was trained using unlabeled peptide sequences to develop its capacity for generating latent peptide variables in an unconstrained manner. In the fine-tuning phase, conditional embeddings were introduced, and the model was trained with both antimicrobial and non-AMP data to enhance its ability to conditionally generate AMPs. In the generation phase, noise XT, which followed a Gaussian distribution, was randomly produced. Then, the diffusion model sequentially produced Xt−1 from Xt using a step-by-step sampling approach until X0 was achieved. Subsequently, the VAE decoder was used to convert the generated latent variables back into their peptide sequences.
The filtering process is divided into three distinct phases: classification, clustering, and coarse-grained molecular dynamics simulations. For the classification model, we drew upon the methodology of Ma et al. (4), using the same training set used for fine-tuning the diffusion model but with an augmented dataset wherein non-AMPs outnumbered AMPs by an order of magnitude. We trained the classification model using various foundational structures, including recurrent neural networks (RNNs), RNNs with attention (RNN-Attention) networks, convolutional neural networks (CNNs), recurrent CNNs (RCNNs), and transformers. The ensemble model exhibiting the highest performance on the test set was selected as the classifier (table S1). After performing filtering through the classifier, we clustered the filtered sequences. Clustering was performed using CD-HIT software with a threshold of 0.6. After clustering, the diversity and novelty of the sequences showed a slight increase compared to the original sequences, which were not subjected to any filtering (tables S2 and S3). Subsequently, we used a random forest model and coarse-grained molecular dynamics simulations (for further details, see Materials and Methods), classifying the peptides obtained after clustering based on their interactions with membranes during the simulations. The classifier model attained a precision of 91% for AMPs at a threshold of 0.65 (table S4). The candidate sequences were ranked from highest to lowest on the basis of their predicted values.
Analyzing the peptide sequences generated and filtered through the pipeline
We leveraged ESM2 650M to extract features from the training set and the peptides generated by the diffusion model without filtration; this was followed by dimensionality reduction using t-distributed stochastic neighbor embedding (t-SNE; Fig. 2A). It was observed that some AMPs within the training set could be distinguished from the non-AMP data in the protein language model, and the distribution of AMPs in the protein language model largely overlaps with the distribution generated by the diffusion model without filtration. This similarity proved that the peptide sequences generated by our model shared a comparable distribution with the AMP sequences contained in the training set within the feature space of the protein language model. Because both the protein language model and t-SNE are unsupervised learning algorithms, they do not incorporate specific information about AMPs during training. This lack of specific information prevents them from distinctly differentiating between AMPs and non-AMPs.
Fig. 2. Comparison of the physicochemical properties of peptides generated by different methods without the assistance of a classifier and the novelty and diversity of the peptides after filtering them with different methods.
(A) t-Distributed stochastic neighbor embedding (t-SNE) visualization of AMPs, non-AMPs, and generated AMPs (sample size, n = 1000). (B to F) Evaluation of unfiltered peptides generated by the models in unconstrained mode: (B) Amino acid distribution. (C) Aromaticity. (D) Helix fraction. (E) Isoelectric point. (F) Net charge [(B) to (F): All sequences for positive and negative; n = 22,000 for UniProt; n = 5120 for others]. (G to J) Evaluation of peptides filtered by each published method for experimental validation: (G) Sequence identities compared to the training set. (H) Sequence similarities compared to the training set. (I) Sequence identities compared to other generated sequences. (J) Sequence similarities compared to other generated sequences. The significance levels of corresponding two-sided Mann-Whitney tests are denoted above the bars as follows: n.s., P ≥ 0.05; *P ≤ 0.05; ***P ≤ 0.001.
Subsequently, we analyzed the physicochemical properties of the candidate AMPs generated by the diffusion model without filtration and compared them with the sequences generated by recently developed methodologies and model-free approaches to assess their differences. All model-based methods (diffusion model, HydrAMP in unconstrained mode, PepCVAE, and VAE) were retrained using the same training set before conducting the comparison. Because of their generative rules, the model-free approaches produced amino acid distributions that were identical to that of the training set. All models were generated under unconstrained conditions, using AMP as the sole control condition. Compared to the negative dataset, our generated sequences exhibited higher occurrences of Gly, Lys, Arg, and Cys and lower occurrences of Met, Glu, Ser, Asp, and Thr. The generated AMP sequences had higher Leu, Gln, and His contents and lower Ala, Trp, and Lys contents than did the training AMP sequences (Fig. 2B). Furthermore, we assessed differences in the physicochemical properties and secondary structures of the candidate AMPs generated by different methods. Our method and the amino acid distribution sampling method produced candidate sequences with aromaticity similar to that of the AMPs in the training set, with no significant differences (Fig. 2C). Only our approach showed no significant differences between the helix fractions of its peptide secondary structures and those of the AMPs in the training set (Fig. 2D). Although the distributions of the isoelectric points of the peptide sequences generated by all methods differed from that of the training set, the quartiles of the isoelectric point distribution of the peptides generated by our method were close to those of the training set (Fig. 2, E and F). These findings suggest that diffusion models are capable of distinguishing between specific AMPs and non-AMPs, thereby generating candidate cationic peptides with mechanisms that are potentially akin to those of known AMPs. This highlights the superior ability of latent diffusion models to produce peptides that exhibit the desired physicochemical properties. Subsequently, we evaluated the capacities of diffusion models versus those of other models to explore the peptide chemical space. Upon generating more than 600,000 sequences, the diffusion model maintained a sequence repetition rate below 1%, while the repetition rates for HydrAMP and PepCVAE were 9.55 and 98.33%, respectively. These results demonstrate the ability of diffusion models to explore extensive chemical spaces. Subsequently, we used the proposed filtration pipeline to sift through the 600,000 generated peptide sequences, selecting 40 for a further experimental investigation.
Ultimately, we compared the candidate AMPs generated for wet-lab validation using the latent diffusion model and other state-of-the-art methods, assessing their novelty and diversity (13, 17, 32, 33). Four metrics were used for the comparison, including the similarities and identities of the candidate AMPs relative to those in the training set and the similarities and identities of some generated candidate AMPs relative to those of other generated candidate AMPs (intrasequence similarity). Because we are evaluating HydrAMP in the analog generation mode, it was not suitable for intrasequence similarity comparisons, and the similarities are expected to be lower than for HydrAMP in the unconstrained mode. The similarity of generated peptides to those in the training set is as low as 0.5686 ± 0.0720, outperforming the state-of-the-art methods including the CLaSS method (0.7499) forming t by Das et al. (32), HydrAMP (0.7655 ± 0.1082) (17), MLPep method (0.7763 ± 0.0574) by Capecchi et al. (33), and ML framework (0.7662 ± 0.0571) (Fig. 2, G and H, and table S2) by Huang et al. (13). The similarity and identity levels of some generated candidate AMPs relative to those of the other generated candidate AMPs were 0.4162 ± 0.0626 and 0.2623 ± 0.0486, respectively, which were markedly lower than those of the recently developed CLaSS, HydrAMP, MLPep, and ML pipeline approaches (Fig. 2, I and J, and table S3). The results demonstrate that our pipeline is capable of generating more novel and diverse candidate AMPs.
Validation of the activity of the generated AMPs
The antimicrobial activity of the top 40 candidate sequences was assessed against six strains of pathogenic fungi (Candida albicans SC5314, Candida glabrata CG13, Candida tropicalis CT-Q-2, Candida parapsilosis CP001, Candida auris CBS15108, and Cryptococcus neoformans H99) and five strains of pathogenic bacteria (the Gram-positive Staphylococcus aureus ATCC29213 strain and the more difficult-to-treat Gram-negative pathogens isolated from clinical specimens: Pseudomonas aeruginosa GD001, Klebsiella pneumoniae GD002, Acinetobacter baumannii GD003, and Escherichia coli GD004). We defined peptides with their MICs less than or equal to 200 μM as AMPs. Last, we obtained 25 peptides with antimicrobial activity from the 40 candidate sequences, with an accuracy of 62.5%.
Peptides 1, 2, 3, 8, 9, 11, 12, 14, 16, 20, 24, 25, 29, 30, 31, 32, 37, 38, and 39 had MICs ≤ 200 μM against pathogenic fungi. In particular, peptide 29 had potent antifungal activity against C. glabrata CG13, and peptides 3, 11, 12, and 20 were active against C. neoformans H99, with MICs ranging from 3.125 to 6.25 μM (Table 1 and tables S5 and S6). Peptide 29 was further evaluated against nine other C. glabrata strains, including the azole-resistant CG2 and CG14 strains and the azole-susceptible CG3, CG4, CG8, CG11, CG12, CG15, and CG17 strains (table S7). Notably, peptide 29 had more potent antifungal activity than (pre)clinical-phase peptides LL-37, against C. glabrata (table S8) (34–36). The antifungal activity of peptides 3, 11, 12, and 20 was further assessed against C. neoformans strains 108, 117, 129, 134, and 138 (table S9). These five peptides had potent activity against all tested strains, with MICs of 3.125 to 12.5 μM, and were labeled AMP-3, AMP-11, AMP-12, AMP-20, and AMP-29, respectively.
Table 1. AMPs with excellent activity (micromolar).
The bold entries indicate exceptional antimicrobial activities for the specified AMPs.
| C. glabrata | C. neoformans | P. aeruginosa | K. pneumoniae | A. baumannii | E. coli | |
|---|---|---|---|---|---|---|
| CG13 | H99 | GD001 | GD002 | GD003 | GD004 | |
| AMP-2 | >400 | 50 | 6.25 | 50 | 6.25 | 25 |
| AMP-3 | >400 | 6.25 | 50 | >400 | >400 | >400 |
| AMP-11 | >400 | 6.25 | 3.125 | 12.5 | 3.125 | 6.25 |
| AMP-12 | >400 | 6.25 | >400 | >400 | >400 | >400 |
| AMP-17 | >400 | >400 | 3.125 | 3.125 | 3.125 | 3.125 |
| AMP-20 | >400 | 6.25 | >400 | >400 | >400 | >400 |
| AMP-24 | >400 | 200 | 100 | 50 | 6.25 | 6.25 |
| AMP-25 | >400 | >400 | 400 | >400 | 12.5 | 50 |
| AMP-29 | 6.25 | >400 | >400 | >400 | >400 | >400 |
Fifteen peptides exhibited antibacterial activity. Three peptides (1, 24, and 31) had antibacterial activity against the Gram-positive bacterium S. aureus, and thirteen peptides (2, 3, 4, 5, 9, 11, 13, 14, 17, 24, 25, 36, and 40) were active against the more difficult-to-treat Gram-negative bacteria (Table 1 and tables S5 and S6). Peptides 2, 11, 17, and 24 showed comparable antibacterial activity to that of several (pre)clinical-phase peptides, such as LL-37 and iseganan, against P. aeruginosa, A. baumannii, and E. coli (table S10) (37–39). Notably, peptide 24 was active against both Gram-negative and Gram-positive bacteria (table S5). Ampicillin, piperacillin, amoxicillin, ceftaroline, fosamil, and cefotaxime sodium were used as positive controls to determine the drug resistance levels of these strains. All P. aeruginosa, A. baumannii, and E. coli strains were resistant to multiple drugs, with MICs > 200 μM; K. pneumoniae GD002 was highly resistant to ampicillin and amoxicillin and moderately resistant to the other three drugs. Peptide 2 had potent activity against P. aeruginosa and A. baumannii, with an MIC of 6.25 μM. Peptides 11 and 17 had MICs of 3.125 to 12.5 μM against all Gram-negative bacteria. Peptide 24 had an MIC of 6.25 μM against A. baumannii and E. coli, and peptide 25 had an MIC of 12.5 μM against A. baumannii. It is worth noting that peptides 2, 24, and 25 exhibited selectivity against specific bacterial species (Table 1 and table S6). Peptides 2, 11, 17, 24, and 25 were selected for further study and labeled as AMP-2, AMP-11, AMP-17, AMP-24, and AMP-25, respectively.
The broad-spectrum activity of these five peptides was examined by assessing their antibacterial activity against four sensitive A. baumannii strains, three sensitive E. coli strains, and two drug-resistant E. coli strains. All of these compounds exhibited potent antibacterial activity, with MICs between 1.5625 and 12.5 μM (table S11).
Hemolytic activity and cytotoxicity assays
The safety profiles of the nine AMPs with MICs < 25 μM were evaluated by testing their hemolytic activity against rat erythrocytes and their cytotoxicity against mammalian cells. All of the AMPs with antifungal activity resulted in hemolysis, except AMP-29, which had hemolytic activity at a concentration of only 50 μM (Fig. 3A). Among the tested antibacterial peptides, AMP-17 and AMP-24 resulted in <5% hemolysis at 200 μM, similar to the positive control LL-37, whereas AMP-2, AMP-11, AMP-25, and another positive control, iseganan, had potent hemolytic activity (Fig. 3A). 3-[4,5-dimethylthiazol-2-yl]-2,5 diphenyl tetrazolium bromide (MTT) assays using human immortalized keratinocyte cells (HaCaTs) and human umbilical vein endothelial cells (HUVECs) showed that AMP-24 and AMP-29 had low cytotoxicity and that AMP-2, AMP-11, AMP-17, and AMP-25 had moderate cytotoxicity; the remaining peptides had relatively potent cytotoxicity (table S12). The concentrations at which AMP-24 induced hemolysis and cytotoxicity were at least 32-fold greater than its MICs, and the concentrations at which AMP-29 elicited hemolytic activity and cytotoxicity were 4- and >32-fold greater than its MICs, respectively. These results suggest that AMP-24 and AMP-29 have potential for in vivo use.
Fig. 3. Characterizing peptides.
(A) Hemolytic activity assays of peptides with potent antimicrobial activity (MIC ≤ 12.5 μM). LL-37, iseganan, and Amphotericin B (AMB) were used as positive controls for antibacterial or antifungal. (B) All-atom molecular dynamics simulations of peptide-membrane systems. (C) Transmission electron microscopy (TEM) images of bacterial or fungal pathogens treated with AMP-24 or AMP-29, respectively. A. baumannii GD003 and E. coli GD004 were treated with 50 μM AMP-24, whereas C. glabrata CG14 was treated with 50 μM AMP-29. After 3 hours of treatment, the cells were collected for TEM observation. Scale bars, 500 nm. (D) Resistance analyses of AMP-24 against A. baumannii GD003 and AMP-29 against C. glabrata CG14.
Mechanistic analyses
Because of their potent antimicrobial activity and favorable safety profiles, antibacterial peptide AMP-24 and antifungal peptide AMP-29 were selected for mechanistic analyses. The AlphaFold2 algorithm predicted an alpha-helical structure for AMP-24 and a random linear structure for AMP-29. All-atom molecular dynamics simulations were performed using the predicted structures of the peptides in the presence of lipid membranes to observe their binding mechanisms (Fig. 3B). The simulation results demonstrated that peptides having antimicrobial properties can align parallel to and infiltrate deeply within the membrane. In contrast, peptide 26, which lacks antimicrobial activity, binds exclusively to the membrane surface (figs. S4 and S5). The simulation results demonstrated that both peptides bind the lipid membrane in parallel, which is consistent with the toroidal pore model proposed by Davidson and co-workers (40). Parallel peptide binding disrupts the alignment of the polar head groups of the lipids, perturbing the acyl chain interactions of the lipids. The resulting membrane curvature changes and destabilization of the membrane surface integrity culminate in membrane destabilization.
To determine whether AMP-24 and AMP-29 affected the integrity of the plasma membrane, propidium iodide (PI), a membrane-impermeable fluorescent nucleic acid stain that penetrates damaged or permeabilized cell membranes and emits red fluorescence, was used. A. baumannii GD003 and E. coli GD004 were treated with 50 μM AMP-24, and C. glabrata CG13 and CG14 were treated with 50 μM AMP-29. Confocal microscopy showed that PI accumulated in cells treated with AMPs, indicating cell membrane damage (fig. S6). Transmission electron microscopy (TEM) revealed that the cell membranes were intact in untreated cells but destroyed in AMP-treated cells. In bacterial cells treated with AMP-24, the bacterial cell membranes were notably curled shrunk and peeled from the cell walls. In the fungal cells treated with AMP-29, the thicknesses of their cell walls increased, and their plasma membranes were damaged (Fig. 3C). Collectively, these observations suggested that the disruption of cell membranes by AMPs contributes to the death of fungal and bacterial cells.
Resistance analyses
The repeated use of antibiotics often results in the development of drug resistance (41). To evaluate the ability of AMP-24 and AMP-29 to induce resistance, A. baumannii GD003 cells and C. glabrata CG14 cells were exposed to sub-MIC levels of AMP-24 and AMP-29, respectively, for 30 days using a serial passage method in 96-well plates. After 30 passages of induced resistance, the MIC of AMP-29 against C. glabrata CG14 was unchanged, whereas the MIC of AMP-24 against A. baumannii GD003 doubled from 6.25 to 12.5 μM (Fig. 3D). These results demonstrated that AMP-24 and AMP-29 have very low potential for inducing drug resistance.
In vivo evaluation of AMP-24 and AMP-29
Because of the potential hemolytic activity of AMP-29, we used a murine skin infection model to evaluate its in vivo efficacy (Fig. 4A). After 24 hours of treatment with AMP-29, the fungal load in the skin of the mice was significantly reduced (Fig. 4B), and inflammation was alleviated (Fig. 4C). Similarly, treating skin infections caused by A. baumannii with AMP-24 significantly reduced the bacterial load (Fig. 4D) and reduced the infiltration of inflammatory cells into the skin after 4 hours of treatment (Fig. 4E).
Fig. 4. Evaluation of the in vivo efficacy of AMP-24 and AMP-29.
(A) Schematic diagram of using AMP-29 and AMP-24 for treatment in a murine skin infection model. h, hours. (B) The fungal burden in murine skin infected with C. glabrata CG14 after treatment with 2% (w/w) AMP-29 (n = 3). A statistical analysis was conducted using Student’s t test. (C) Histological images of H&E and periodic acid–Schiff (PAS)–stained of skins infected with C. glabrata CG14 and treated with AMP-29 (scale bars, 400 μm). The insets in the top right part are enlarged images of the areas enclosed in rectangles. The arrow in the enlarged image of the vehicle-treated group indicates an example of PAS-stained fungi. (D) The bacterial burden in murine skin infected with A. baumannii GD003 after treatment with 2% (w/w) AMP-24 (n = 3). A statistical analysis was conducted using Student’s t test. (E) Histological images of H&E-stained skin inoculated with A. baumannii and treated with AMP-24 (scale bars, 400 μm). (F) Schematic diagram of using APM-24 for treatment in a murine lung infection model. d, days. (G) Bacterial burden in the lungs after treatment with AMP-29. The mice were intranasally inoculated with A. baumannii (1 × 108 CFU per mouse) and then treated with PBS or AMP-24 (40 mg/kg). A statistical analysis was conducted using Student’s t test (n = 5). (H) Histological images of H&E- and Masson-stained of lungs inoculated with A. baumannii and treated with AMP-24 (scale bars, 200 μm).
A. baumannii pneumonia causes mortality in healthy individuals or in hospital settings, with a high mortality rate varying from 40 to 70% (42). Because of the low hemolytic activity and cytotoxicity of AMP-24, we treated mice with AMP-24 via intravenous injection to assess the efficacy of this peptide against A. baumannii pneumonia (Fig. 4F). The bacterial load was reduced in the mice treated with AMP-24 (Fig. 4G). Hematoxylin and eosin (H&E) and Masson’s staining of fixed lung sections revealed that AMP-24 markedly alleviated lung inflammation and fibrosis (Fig. 4H). These results suggested the in vivo efficacy of AMP-24 for the treatment of A. baumannii pneumonia. In addition, other off-target tissues, including the heart, liver and kidney, were subjected to H&E staining, and no obvious immune cell infiltrations were observed (fig. S7), suggesting the low toxicity of AMP-24 in vivo.
DISCUSSION
AMPs have emerged as powerful alternatives to traditional antibiotics, with exceptional efficacy against multidrug-resistant bacterial pathogens. AI is a dynamic tool for navigating the complex chemical landscape of AMPs and markedly expediting the development of innovative AMPs to aid in the global fight against antibiotic resistance. Continued research is crucial for addressing the current limitations and refining methodologies for generating highly diverse and effective AMP candidates.
Data-driven approaches have demonstrated considerable potential over several years. However, the current AI approaches are limited by their insufficient sequence diversity and constrained chemical spaces and are seldom used to generate antifungal peptides; thus, overcoming these limitations is essential for identifying novel AMPs with substantial application potential.
In this study, our primary objective was to address the challenge of enhancing the diversity of effective AMPs generated by AI models while maintaining a certain level of accuracy. To achieve this goal, we developed a pipeline combining latent diffusion and filtering methods to generate potential AMP candidates. Our pipeline has two advantages. First, our pipeline designed sequences with unparalleled diversity (compared with those produced by contemporary methodologies). Second, its tailored training protocol can be easily adapted to other peptide generation tasks, such as those involving antitumor or antidiabetes peptides.
The efficacy of the pipeline was ascertained by synthesizing and examining 40 peptide sequences, among which 25 exhibited antibacterial or antifungal activities. To our knowledge, this is the first study to develop antifungal peptides by using AI methods. Among the 25 AMPs, five had selective activity against specific fungal species, and three showed some degree of selectivity against specific bacterial species, suggesting that these AMPs do not have universal activity. The bindings of different AMPs to components of the cell wall or cell membrane may vary among fungal or bacterial species. Elucidating these critical interactions will aid in the development of AMPs that target specific pathogens with reduced side effects on beneficial bacteria or fungi. The membrane permeability barrier in Gram-negative bacteria restricts the discovery of antibiotics (43). Among the AMPs generated by the latent diffusion model, AMP-24 exhibited potent activity against Gram-negative bacteria, including P. aeruginosa, K. pneumoniae, A. baumannii, and E. coli, which are on the priority list of antibiotic-resistant bacteria published by the World Health Organization (44). AMP-24 also exhibited therapeutic efficacy in vivo in a mouse model of A. baumannii pneumonia, supporting its potential as a lead molecule for the development of novel antibacterial peptides. AMP-29 showed potent but selective antifungal activity against C. glabrata and in vivo antifungal efficacy in a murine skin infection model. The activity of AMP-29 against C. glabrata was much better than that of (pre)clinical-phase peptide LL-37. These results demonstrate that our pipeline developed herein provides a route for developing innovative peptide drugs to combat drug resistance.
Despite these results, room for improvement remains. Integrating parameters such as physicochemical attributes, secondary structures, and activity metrics as conditions within the generation process would allow precision-driven AMP generation. In addition to AMPs, this innovative approach holds immense potential for the de novo creation of bioactive molecules aimed at distinct biological targets, advancing the production of groundbreaking therapeutic agents.
MATERIALS AND METHODS
Data preparation
We collected four datasets to train a VAE, a latent diffusion model, a conditional generation latent diffusion model, and various classifiers. To collect a more comprehensive collection of AMP sequences, we did not differentiate the specific strains targeted by these peptides. Instead, we uniformly classified them as AMPs.
Variational autoencoder
From the entire UniProt database (30), we gathered all peptide sequences with lengths of fewer than 50 amino acids. After eliminating duplicates and nonnatural amino acids (X, B, Z, U, and O), we obtained 2,880,719 sequences.
Latent diffusion model
We generated the training set for the latent diffusion model by sampling 12,000 sequences per length from the VAE training set. If the number of sequences was less than 12,000, then we retained all of them. Ultimately, we collected 480,358 sequences.
Conditional generation latent diffusion model
We created a training set for the conditional generation latent diffusion model comprising AMPs and non-AMPs. The AMPs were obtained by combining experimentally validated antibacterial or antifungal sequences from dbAMP (45), Dramp (46), GRAMPA (47), and starPep (48), after removing duplicates. The non-AMPs were assumed to be biologically inactive and were manually selected using UniProt filters, excluding properties such as anticancer, antimicrobial, antibacterial, antifungal, antiviral, antibiotic, antiproliferative, fungicide, cytotoxic, hemolytic, defensin, defense, secreted, toxin, toxic, transit peptide, inhibitor, and inhibits. BLAST searches of all AMPs and non-AMPs were performed using blastp v.2.13.0 (49), and sequences with greater than 80% coverage and 60% identity were removed. Non-AMPs were further filtered using CD-HIT (50) with a threshold of 0.9. Last, we sampled the non-AMPs to obtain a length distribution similar to that of the AMPs (fig. S8).
Classifiers
We created two training sets for classifiers from the filtered UniProt database using CD-HIT. In the first set, the number and length distribution of the AMPs and non-AMPs were comparable. In the second set, the length distribution of the AMPs and non-AMPs remained similar, but the number of non-AMPs was 10 times greater than the number of AMPs.
Variational autoencoder
Model structure
In this work, we used a VAE consisting of a three-layer transformer as both the encoder and the decoder. Initially, peptide sequences are transformed from their one-hot input representations into vectors of 128 dimensions via an embedding layer. Subsequently, these vectors undergo further processing through three transformer encoding layers. Dollar et al. demonstrated that compression via convolutional layers enhances reconstruction performance when contrasted with the use of linear layers. Drawing inspiration from their findings, we implemented convolutional layers for the compression of latent variables in our methodology (31).
Before reparameterization, we used a fully connected layer to predict the sequence length. Following the reparameterization process, the decompression of latent variables is initially conducted through a deconvolution layer. Before the decoding phase, a mask is created, tailored to the predicted length of the sequence. This mask, in conjunction with the latent variables, is then used to decode and reconstruct the peptide sequences.
Loss function
The loss function used to train the VAE consists of three components: reconstruction loss, KL divergence, and length loss. Because the distribution of amino acids in peptides is imbalanced, it is crucial to balance the losses based on frequency to prevent the model from merely learning the most frequently used amino acids. The predicted amino acids (Ypred) and the true amino acids (Ytrue) at sequence reconstruction loss were used in the calculation. The amino acids were weighted by their proportional log frequency and then scaled to values ranging from 0.5 to 1.0. In the formulas below, σ is the predicted mean, and μ2 is the predicted variance. The predicted length (Lpred) and the true length (Ltrue) were also included in the calculation of length loss
Training
The VAE was implemented using PyTorch (51) and trained on an NVIDIA Tesla V100 32G * 8. The model was optimized using the Adam optimizer, and each amino acid was represented by a 128-dimensional vector. The batch size was set to 512. We tested the VAE’s loss at different learning rates and ultimately selected 0.0007 as the final learning rate to be used (fig. S9). The λ value was linearly annealed from 0 to 0.5 over the course of 300 epochs. The VAE that we used was compared with other methods as shown in table S13.
Latent diffusion model
Model structure
The latent diffusion model comprised two embedding layers (conditional and position), three fully connected neural networks (time embedding, input feature processing, and output feature processing), a BERT encoding layer, a layer normalization layer, and a dropout layer. First, the time embedding, control embedding, and position embeddings are computed and then added to the latent variables processed by a fully connected layer. Subsequently, after normalization with normalization layer, they are fed into a BERT encoding layer built by transformers python package. The diffusion and generation processes were inspired by previous studies of image and sequence diffusion models (52, 53).
Loss function
When the diffusion step was not zero, the loss function for training the latent diffusion model was the mean squared error of Xpredict and Xstart. When the diffusion step was zero, the model predicted X0. In the formulas below, βt∈(0, 1) is the variance of Gaussian noise added at step t, and T is the number of diffusion steps
Training
The latent diffusion model was implemented using PyTorch and trained on an NVIDIA Tesla A100 40G * 8 using the AdamW optimizer. The data before VAE convolutional layer compression were used as the diffusion data. The batch size was 512. We tested the diffusion loss during the train process at different learning rates and ultimately selected 0.0001 as the final learning rate to be used (fig. S10). After consideration of both the training loss and the running time, we determined that 500 diffusion steps represent the optimal balance for our model (fig. S11). In this task, the VAE was trained for 200 epochs. The model was first trained for unconditional generation using 480,358 unlabeled data and then trained for conditional generation using AMP and non-AMP data (fig. S12).
Biological filtering criteria
Biological criteria were used to approximate expert evaluations of peptide synthesizability. In this work, we adopted the method proposed by Szymczak et al. (17). Sequences that contained more than three positively charged amino acids (Lys and Arg) within a window of five amino acids were excluded. Sequences with three consecutive hydrophobic amino acids (Ala, Val, Leu, Ile, Phe, Trp, and Met) or three consecutive repeated amino acids were also removed. Last, sequences containing Cys were excluded.
Classifiers
Model structure
To filter the generated AMPs, we trained several different models. García-Jacas et al. (54) and Sidorczuk et al. (55) demonstrated that shallow models have comparable performance to deep models for AMP prediction. Therefore, we selected the RNN, RNN-Attention, CNN, RCNN, and transformer as the base structures to train the classifier model using two separate training sets.
Loss function
We trained all models using the Cross-Entropy loss function.
Training
All models were implemented in PyTorch and trained on an NVIDIA GeForce RTX3090 24G using the Adam optimizer. The batch size was 512. For each model, we conducted training at learning rates of 0.01, 0.001, 0.0001, 0.00001, and 0.000001. The optimal learning rate was then determined on the basis of the highest level of performance achieved on the validation set. The ExponentialLR function was used with a gamma of 0.99 to adjust the learning rate.
Sampling, generation, and properties calculation
Sampling
As noted by Ho et al. (52), a step-by-step sampling from X0 is not required to generate Xt. Xt can be directly sampled from X0. , ,
Generation
Although our model was able to directly predict X0, it needed to resample Xt−1 to generate step by step due to the inaccurate results of one prediction. According to Bayes’ theorem, we can resample from to generate Xt−1. The generation was run on an NVIDIA Tesla V100 32G. During the generation, the model’s default batch size was set to 512, and it was run 10 times to generate a total of 5120 sequences
Filter
Initially, we used an ensemble classifier for preliminary filtering of generated sequences. Subsequently, biological filters were applied, followed by clustering using CD-HIT at a threshold of 0.6. The clustered results were then further screened through coarse-grained molecular dynamics simulations and random forest model.
Property calculation
To evaluate the physicochemical properties of sequences generated by the diffusion model, we initially selected two model-free methods based on amino acid distribution and amino acid position distribution. The amino acid distribution method involves statistically calculating the frequency of each amino acid and the peptide length, then randomly generating candidate peptide sequences based on these frequencies. The amino acid position distribution method differs in that the amino acid distribution is independently calculated for each position. The HydrAMP, PepCVAE, and VAE models were retrained using the code provided by Szymczak et al., using the same dataset as our diffusion model for retraining. The HydrAMP, PepCVAE, and VAE models all use an unconstrained generation approach, using AMP as the sole control condition. The terms “positive” and “negative” refer, respectively, to the antimicrobial and non-AMP data used to train the models. The UniProt dataset was derived from sequences randomly selected from peptides under 50 amino acids in length collected from UniProt. The random dataset consists of sequences generated by assigning each amino acid with an equal probability. We used the peptide Python package to calculate the physicochemical properties of the peptides.
The similarity between two peptides was defined as the length of the longest contiguous common sequence (56). The similarity between a peptide and the peptide library was defined as the similarity between the peptide and the most similar peptide in the library. Intrasequence similarity involves calculate the extent of similarity and identity that each peptide in a given library shares with the rest of the sequences within that library, excluding the comparison with itself. All training sets used for similarity comparisons were the ones used by the authors of each method to train their respective models. For HydrAMP, MLPep, and ML pipelines, we compare using the AMPs from the training sets provided by the authors. For the CLaSS model, as the authors did not provide a complete training set, we collected AMPs from the corresponding databases based on the authors’ descriptions to create the training set. The sequences compared were those ultimately used for validation in wet-lab experiments. The sequences compared have been subjected to distinct filtration processes by their authors. PepCVAE and VAE did not undergo wet-lab validation; therefore, they are not compared in terms of novelty and diversity. Now, all sequence similarities are calculated using the Needleman-Wunsch global alignment algorithm provided by the needle module in the EMBOSS software to compute the identity and similarity between two sequences.
Molecular dynamics
Coarse-grained molecular dynamics
Tinker (57) was used to generate a Protein Data Bank (PDB) file of the all-atom representation of the peptide. All peptides were modeled as structures with alpha helices and dihedral angles of ϕ = −60 and ψ = −50. The three-dimensional structure of each peptide was optimized using the Amber99SB molecular force field.
We used martinize.py (58) to convert the all-atom representation of the peptide into a coarse-grained system and then built the peptide-membrane system using insane.py (59). The solvent was 9:1 water:antifreeze particles, and the membrane was 3:1 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC):1-Palmitoyl-2-oleoyl-sn-glycero-3-phosphatidylglycerol (POPG) (32). The addition of antifreeze particles is intended to prevent potential impacts of ice formation on the simulation. The system had dimensions of 8 nm by 8 nm by 12 nm, with the membrane perpendicular to the longest axis. To neutralize the system, 0.15 M ions were added. For all coarse-grained molecular dynamics simulations, the Martini force field was used (60).
The coarse-grained system was minimized for 50,000 steps using Gromacs 2021.4 (61) and the 2.0 version of the Martini force field. After minimization, the production run was performed for 500 ns with a time step of 20 fs. The temperature was maintained at 310 K by applying velocity rescaling (62) independently to the peptide, lipid, and solvent groups. The pressure was maintained at 1 atm using the Parrinello-Rahman method. After sampling for 500 ns, we used the MDAnalysis Python package (63) to count the average distance of all coarse-grained peptide atoms from the membrane. We use the average distances at different times within 500 ns as features for training. Last, we used the random forest algorithm in the sklearn Python package (64) to classify the AMPs and non-AMPs. We used the RandomizedSearchCV from sklearn to determine the optimal hyperparameters for the random forest model. We simulated 428 AMPs and 409 non-AMPs to generate the training set.
All-atom molecular dynamics
We used the CHARMM-GUI (65) webservice to prepare inputs for MD simulations. The peptide structures predicted by AlphaFold2 (66) are used as inputs, and the system is constructed using the Membrane Builder in CHARMM-GUI. For the simulations of antibacterial peptides, the membrane composition was 3:1 POPC:POPG (32). For the antifungal peptide simulations, we used a membrane composition of 32:11:10:10:35 POPC:1-Palmitoyl-2-oleoyl-sn-glycero-3-phosphate (POPA):1-Palmitoyl-2-oleoyl-sn-glycero-3-phospho-L-serine sodium (POPS):1-palmitoyl-2-oleoyl-sn-glycero-3-phosphoinositol (POPI):Erogosterol (67). The systems were solvated in a water box of 10.5 nm by 10.5 nm by 10.5nm, with NaCl at the concentration of 0.15 M. We used CHARMM36m force field for peptides and ions, and transferable intermolecular potential with 3 points (TIP3P) model for water. The models processed by CHARMM-GUI were then used as inputs to GROMACS (68) for molecular dynamics simulations. Negative distances indicate that the peptide is in the solution, away from the membrane, while positive distances suggest that it is in contact with, or inserted into, the membrane.
Evaluation
To evaluate the VAE model, we selected the bilingual evaluation understudy (BLEU) score, reconstruction error rate and token accuracy as metrics. The sentence_bleu function in the nltk Python package was used to calculate the BLEU score. We used the accuracy_score, roc_auc_score, and classification_report functions in the sklearn Python package to evaluate the performance of the classifier. We also calculated true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN)
Strains, culture conditions, and chemicals
Six species of fungi were used in antimicrobial activity assays: C. albicans (SC5314), C. glabrata (CG2, CG3, CG4, CG8, CG11, CG12, CG13, CG14, CG15, and CG17), C. tropicalis (CT-Q-2), C. parapsilosis (CP001), C. auris (CBS15108), and C. neoformans (H99, 108, 117, 129, 134, and 138). In addition, five species of bacteria were used: S. aureus (ATCC29213), P. aeruginosa (GD001), K. pneumoniae (GD002), A. baumannii (GD003, GD4948, GD4906, GD4955, and GD144013), and E. coli (GD004, GD7188, GD7135, GD7465, GD6738, and GD7354). The strains of C. neoformans were obtained from G. Liao of Southwest University. All strains of C. glabrata and C. tropicalis were obtained from Shandong Provincial Qianfoshan Hospital. C. parapsilosis CP001 was obtained from the Central Hospital of Jinan, and C. auris CBS15108 was donated by W.Q. Liao of Second Military Medical University. All strains of S. aureus, P. aeruginosa, K. pneumoniae, A. baumannii, and E. coli were obtained from Guangzhou Medical University. The strains were stored in preservation solution supplemented with 20% (v/v) glycerol at −80°C. Fungal strains were propagated from frozen stocks on yeast peptone dextrose (YPD) agar plates (1% yeast extract, 2% peptone, 2% dextrose, and 2% agar) using a sterile inoculation loop and then incubated overnight at 30°C. Bacterial strains were propagated on Luria-Bertani (LB) agar plates (0.5% yeast extract, 1% peptone, 1% NaCl, and 1.2% agar) and incubated at 37°C overnight. For further experiments, selected colonies were inoculated in YPD liquid medium at 30°C for fungal strains and LB liquid medium at 37°C for bacterial strains and agitated (200 rpm) overnight to midlogarithmic and stationary growth phase, respectively. The cells were harvested by centrifugation, adjusted to 1 × 107 cells/ml, and diluted to the desired inoculum concentration based on the optical density at 600 nm. In subsequent assays, Mops-buffered RPMI 1640 (pH 7.4) or Mueller-Hinton broth (MHB) medium was used for fungal strains or bacterial strains, respectively.
Ampicillin (Solarbio, China), oxacillin (Solarbio, China), piperacillin (Yanye, China), amoxicillin (Bomei, China), ceftaroline fosamil (Macklin, China), cefotaxime sodium (TCI, Japan), and fluconazole (Solarbio, China) were prepared in dimethyl sulfoxide (DMSO). All peptides were dissolved in sterile water for further experiments.
MIC determination
The MICs of the peptides against fungal strains were determined by the microbroth dilution method according to the guidelines of Clinical and Laboratory Standards Institute CLSI M27-A3 (69). Briefly, overnight cultures of fungi were diluted to a cell density of 1 × 103 cells/ml in RPMI 1640 medium, and 100 μl was added to each well of a 96-well flat-bottomed microtiter plate containing a peptide concentration gradient of 1.5625 to 200 μM. Fluconazole was used as a positive control. Wells in which the cell suspensions were not exposed to peptides were used as untreated controls. The plates were incubated at 35°C for 24 hours. The MIC value was determined as the lowest drug concentration that led to no visible fungal growth.
The MICs of the peptides against bacterial strains were determined by the microbroth dilution method according to the guidelines of Clinical and Laboratory Standards Institute CLSI M100-S31 (70). Briefly, overnight cultures were diluted to 5 × 105 cells/ml in MHB medium, and 100 μl was added to each well of a 96-well plate containing a peptide concentration gradient of 1.5625 to 200 μM. Wells without peptide were used as negative controls. The positive controls were oxazoline for S. aureus and ampicillin, piperacillin, amoxicillin, ceftaroline fosamil, and cefotaxime sodium for Gram-negative bacteria. The positive AMPs, LL-37 and iseganan, were also used in the assays. The bacterial plates were incubated at 37°C for 24 hours. The MIC value was determined as the lowest drug concentration that led to no visible bacterial growth.
Membrane permeabilization
Cell membrane permeability was assessed using the fluorescent probe PI. Cells in logarithmic phase were collected and washed with phosphate-buffered saline (PBS), and the cell density was adjusted to 5 × 106 cells/ml. Cultures of GD003 and GD004 bacteria were exposed to 50 μM AMP-24, and cultures of CG13 and CG14 fungi were treated with 50 μM AMP-29. Untreated cultures served as negative controls. After incubation for 3 hours, the cells were centrifuged, resuspended in PBS, and stained with 5 μM PI for 30 min. After staining, the cells were washed thrice with PBS, and the fluorescence intensity was measured using a confocal microscope (LSM 900 with AiryScan 2) with a 63× oil-immersion objective lens with excitation at 561 nm.
Transmission electron microscopy
Changes in cell membrane morphology and structure induced by AMPs were observed by examining the ultrastructure of fungal and bacterial cells by TEM. Logarithmic-phase fungal or bacterial cells were collected and washed with PBS. The cell density was assessed by spectrophotometry and diluted to 5 × 106 cells/ml. Cultures of GD003 and GD004 bacteria were exposed to 50 μM AMP-24, and cultures of CG14 fungi were exposed to 50 μM AMP-29. Untreated cultures served as negative controls. After incubation for 3 hours, the cultures were washed, fixed, dehydrated, embedded, and sectioned into 70- to 90-nm slices. These slices were then stained, air dried, and observed under a transmission electron microscope (Hitachi HT7700).
Drug resistance development
A. baumannii GD003 and C. glabrata CG14 were used in this assay. After activation in liquid culture, these two strains were cultured for 12 hours to generate primary cells. The primary bacterial cells were collected and diluted to 5 × 105 cells/ml in MHB liquid medium. The primary fungal cells were collected and diluted to 1 × 103 cells/ml in RPMI 1640 liquid medium. The diluted cells were then added to a 96-well sterile microplate for drug testing using the microbroth dilution method. Secondary cells at half of the MIC value were collected for further culture, and MICs were assessed as described above. This process was repeated over a period of 30 days to generate a curve of MIC variations. Because of highly resistance to ordinary antibiotics for these two strains, we failed to choose corresponding drugs as the control to develop drug resistance.
Hemolysis assay
Blood was collected from the abdominal aorta of Sprague-Dawley rats using heparin sodium anticoagulant. Red blood cells (RBCs) were isolated by centrifugation, washed, and diluted with saline to obtain a 10% suspension. AMPs were diluted with saline solution to obtain different concentrations using the twofold dilution method. Negative control samples containing only physiological saline and positive control samples containing 1% Triton X-100 were included with each set of samples. The samples were incubated with 10% RBCs at 37°C for 1 hour and then centrifuged for 10 min. The supernatant was collected, and the absorbance at 540 nm was measured using a multimode reader to evaluate hemolysis.
The hemolysis rate was calculated using the formula (Am − An)/(Ap − An) × 100%, where Am, An, and Ap are the OD values of the experimental, negative control, and positive control samples, respectively.
Cytotoxicity assay
The HUVECs and HaCaTs were purchased from Shanghai Institutes for Biological Sciences. Logarithmic-phase HUVECs and HaCaTs were digested with trypsin and adjusted to 5 × 104 cells/ml. The cells were then seeded in 96-well plates and incubated overnight for adhesion. Different AMP concentrations were added, and the plates were incubated for 24 hours at 37°C with 5% CO2. Wells without AMPs were used as untreated negative controls. Following incubation, MTT solution (5 mg/ml) was added to each well (10 μl) and incubated for 4 hours. The supernatant was discarded, and 100 μl of DMSO was added and mixed thoroughly. Last, the absorbance at 490 nm was measured using a multimode reader to evaluate cytotoxicity.
In vivo studies: Murine skin and lung infection models
Male BALB/c mice aged 6 to 8 weeks and weighing 18 to 20 g were purchased from Beijing Vital River Laboratory Animal Technology Company. Mice were randomly assigned to three groups: the uninfected group, the infection group, and the AMP-treated group. For the murine skin infection model, the mice were anesthetized with 1% sodium pentobarbital solution (50 mg/kg), and dorsal dehairing was performed by scraping the backs of the mice with sterile Velcro. Bacterial or fungal suspension (20 μl) containing 1 × 108 CFUs was applied on the wounded skin or 20 μl of saline in the control group and bandaged. The mice were treated with 2% AMP-24 gel 4 hours after A. baumannii skin infection or with 2% AMP-29 gel 24 hours after C. glabrata skin infection. Twenty-four hours after drug administration, the bacterial or fungal load in the injured skin was determined. H&E staining was performed for both skin infection models; in addition, periodic acid–Schiff staining was performed for the C. glabrata infection model.
To construct the murine bacteremic pneumonia model, mice were immunosuppressed by subcutaneous injection of cortisone (100 mg/kg) for three consecutive days before infection, and A. baumannii was inoculated intranasally (1 × 108 CFU per model). AMP-24 (40 mg/kg) was administered intravenously for three consecutive days; the control group received saline. The bacterial load in the lungs was evaluated by spot plate calculation, and H&E and Masson’s trichrome staining were performed for histological examinations.
Statistical analysis
Data were statistically analyzed using Student’s t test or Mann–Whitney U test to compare values between two specific groups. Statistical significance was determined according to the P value. *P < 0.05, **P < 0.01, and ***P < 0.001; and n.s. means not significant. Statistical details are found in the figures or figure legends.
Ethics statement
The animal experiments in this study were conducted under a protocol authorized by the Animal Care and Use Committee at Shandong University under approval number 21-107. Animal experiments were minimized, and the murine research methods were designed to minimize mouse suffering.
Acknowledgments
We thank G. Liao of Southwest University for donating the C. neoformans strains used in this study.
Funding: This work was supported by the National Natural Science Foundation of China, nos. 82273975 (W.C.), 82293682 (H.L.), and 82173703 (H.L.).
Author contributions: Conceptualization: W.C., H.Lo., and Y.W. Methodology: W.C., Y.W., M.S., W.F., and G.L. Formal analysis: Y.W., R.H., and W.Y. Investigation: W.C., Y.W., M.S., F.L., Z.L., and X.F. Validation: H.Lo., Y.W., M.S., F.L., Z.L., R.H., and Y.D. Visualization: W.C., H.Lo., and Y.W., Resources: W.C., H.Lo., H.Lu., and W.Y. Funding acquisition: W.C. and H.Lo. Data curation: W.C., Y.W., and H.Lu. Supervision: W.C., H.Lo., and Y.D. Software: Y.W. Project administration: W.C. and H.Lo. Writing—original draft: W.C., Y.W., and M.S., Writing—review and editing: W.C., H.Lo., and Y.W.
Competing interests: W.C. is the inventor of a Chinese patent with applying no. 202311205302.8, related to this manuscript. The authors declare that they have no other competing interests.
Data and materials availability: All data needed to evaluate the conclusions in the paper are present in the paper and/or the Supplementary Materials. The main data and codes supporting the findings of this study are available at https://zenodo.org/records/13762213 or https://github.com/Wangyj2023/PepDiffusion.
Supplementary Materials
This PDF file includes:
Figs. S1 to S12
Tables S1 to S19
REFERENCES AND NOTES
- 1.Aminov R. I., A brief history of the antibiotic era: Lessons learned and challenges for the future. Front. Microbiol. 1, 134 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Larsson D. G. J., Flach C. F., Antibiotic resistance in the environment. Nat. Rev. Microbiol. 20, 257–269 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ramazi S., Mohammadi N., Allahverdi A., Khalili E., Abdolmaleki P., A review on antimicrobial peptides databases and the computational tools. Database 2022, baac011 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ma Y., Guo Z., Xia B., Zhang Y., Liu X., Yu Y., Tang N., Tong X., Wang M., Ye X., Feng J., Chen Y., Wang J., Identification of antimicrobial peptides from the human gut microbiome using deep learning. Nat. Biotechnol. 40, 921–931 (2022). [DOI] [PubMed] [Google Scholar]
- 5.Pushpanathan M., Gunasekaran P., Rajendhran J., Antimicrobial peptides: Versatile biological properties. Int. J. Pept. 2013, 675391 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Magana M., Pushpanathan M., Santos A. L., Leanse L., Fernandez M., Ioannidis A., Giulianotti M. A., Apidianakis Y., Bradfute S., Ferguson A. L., Cherkasov A., Seleem M. N., Pinilla C., de la Fuente-Nunez C., Lazaridis T., Dai T., Houghten R. A., Hancock R. E. W., Tegos G. P., The value of antimicrobial peptides in the age of resistance. Lancet Infect. Dis. 20, e216–e230 (2020). [DOI] [PubMed] [Google Scholar]
- 7.Li J., Koh J. J., Liu S., Lakshminarayanan R., Verma C. S., Beuerman R. W., Membrane active antimicrobial peptides: Translating mechanistic insights to design. Front. Neurosci. 11, 73 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Chen C. H., Bepler T., Pepper K., Fu D., Lu T. K., Synthetic molecular evolution of antimicrobial peptides. Curr. Opin. Biotechnol. 75, 102718 (2022). [DOI] [PubMed] [Google Scholar]
- 9.Cardoso M. H., Orozco R. Q., Rezende S. B., Rodrigues G., Oshiro K. G. N., Cândido E. S., Franco O. L., Computer-aided design of antimicrobial peptides: Are we generating effective drug candidates? Front. Microbiol. 10, 3097 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Yan J., Bhadra P., Li A., Sethiya P., Qin L., Tai H. K., Wong K. H., Siu S. W. I., Deep-AmPEP30: Improve short antimicrobial peptides prediction with deep learning. Mol. Ther. Nucleic Acids. 20, 882–894 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kavousi K., Bagheri M., Behrouzi S., Vafadar S., Atanaki F. F., Lotfabadi B. T., Ariaeenejad S., Shockravi A., Moosavi-Movahedi A. A., IAMPE: NMR-assisted computational prediction of antimicrobial peptides. J. Chem. Inf. Model. 60, 4691–4701 (2020). [DOI] [PubMed] [Google Scholar]
- 12.Li C., Sutherland D., Hammond S. A., Yang C., Taho F., Bergman L., Houston S., Warren R. L., Wong T., Hoang L. M. N., Cameron C. E., Helbing C. C., Birol I., AMPlify: Attentive deep learning model for discovery of novel antimicrobial peptides effective against WHO priority pathogens. BMC Genomics 23, 1–15 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Huang J., Xu Y., Xue Y., Huang Y., Li X., Chen X., Xu Y., Zhang D., Zhang P., Zhao J., Ji J., Identification of potent antimicrobial peptides via a machine-learning pipeline that mines the entire space of peptide sequences. Nat. Biomed. Eng. 7, 797–810 (2023). [DOI] [PubMed] [Google Scholar]
- 14.Wan F., Torres M. D. T., Peng J., de la Fuente-Nunez C., Deep-learning-enabled antibiotic discovery through molecular de-extinction. Nat. Biomed. Eng. 8, 854–871 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Santos-Júnior C. D., Torres M. D. T., Duan Y., Rodríguez Del Río Á., Schmidt T. S. B., Chong H., Fullam A., Kuhn M., Zhu C., Houseman A., Somborski J., Vines A., Zhao X.-M., Bork P., Huerta-Cepas J., de la Fuente-Nunez C., Coelho L. P., Discovery of antimicrobial peptides in the global microbiome with machine learning. Cell 187, 3761–3778.e16 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Aronica P. G. A., Reid L. M., Desai N., Li J., Fox S. J., Yadahalli S., Essex J. W., Verma C. S., Computational methods and tools in antimicrobial peptide research. J. Chem. Inf. Model. 61, 3172–3196 (2021). [DOI] [PubMed] [Google Scholar]
- 17.Szymczak P., Możejko M., Grzegorzek T., Jurczak R., Bauer M., Neubauer D., Sikora K., Michalski M., Sroka J., Setny P., Kamysz W., Szczurek E., Discovering highly potent antimicrobial peptides with deep generative model HydrAMP. Nat. Commun. 14, 1453 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.A. Ulhaq, N. Akhtar, G. Pogrebna, Efficient diffusion models for vision: A survey. arXiv:2210.09292 [cs.CV] (2022).
- 19.Dhariwal P., Nichol A., Diffusion models beat GANs on image synthesis. Adv. Neural Inf. Process. Syst. 11, 8780–8794 (2021). [Google Scholar]
- 20.N. Gruver, S. Stanton, N. Frey, T. G. J. Rudner, I. Hotzel, J. Lafrance-Vanasse, A. Rajpal, K. Cho, A. G. Wilson, Protein design with guided discrete diffusion. arXiv:2305.20009 [cs.LG] (2023).
- 21.Sarah Alamdari, Nitya Thakkar, Rianne van den Berg, Alex X. Lu, Nicolo Fusi, Ava P. Amini, K. K. Yang, Protein generation with evolutionary diffusion: sequence is all you need. bioRxiv 556673 [Preprint] (2023). 10.1101/2023.09.11.556673. [DOI]
- 22.S. Zhang, Z. Jiang, R. Huang, S. Mo, L. Zhu, P. Li, Z. Zhang, E. Pan, X. Chen, Y. Long, PRO-LDM: Protein sequence generation with a conditional latent diffusion model. bioRxiv 554145 [Preprint] (2023). 10.1101/2023.08.22.554145. [DOI]
- 23.J. Song, C. Meng, S. Ermon, Denoising diffusion implicit models, paper presented at the 9th International Conference on Learning Representations (ICLR 2021), Vienna, Austria, 4 May 2021. [Google Scholar]
- 24.Z. Guo, J. Liu, Y. Wang, M. Chen, D. Wang, D. Xu, J. Cheng, Diffusion models in bioinformatics: A new wave of deep learning revolution in action. arXiv:2302.10907 [cs.LG] (2023).
- 25.Croitoru F. A., Hondru V., Ionescu R. T., Shah M., Diffusion models in vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 45, 10850–10869 (2023). [DOI] [PubMed] [Google Scholar]
- 26.D. P. Kingma, M. Welling, Auto-encoding variational Bayes, paper presented at the 2nd International Conference on Learning Representations (ICLR 2014 Conference Track), Banff, Canada, 14 to 16 April 2014. [Google Scholar]
- 27.Y. Wang, M. Song, F. Liu, Z. Liang, R. Hong, Y. Dong, H. Luan, X. Fu, W. Yuan, W. Fang, G. Li, H. Lou, W. Chang, Artificial intelligence using a latent diffusion model enables the generation of diverse and potent antimicrobial peptides, Zenodo (2024); 10.5281/zenodo.13762213. [DOI]
- 28.R. Rombach, A. Blattmann, D. Lorenz, P. Esser, B. Ommer, “High-resolution image synthesis with latent diffusion models” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (IEEE, 2022), pp. 10674–10685. [Google Scholar]
- 29.Girin L., Leglaive S., Bie X., Diard J., Hueber T., Alameda-Pineda X., Dynamical variational autoencoders: A comprehensive review. Found. Trends Mach. Learn. 15, 1–175 (2021). [Google Scholar]
- 30.Bateman A., Martin M. J., Orchard S., Magrane M., Agivetova R., Ahmad S., Alpi E., Bowler-Barnett E. H., Britto R., Bursteinas B., UniProt: The universal protein knowledgebase in 2021. Nucleic Acids Res. 49, D480–D489 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Dollar O., Joshi N., Beck D. A. C., Pfaendtner J., Attention-based generative models for: De novo molecular design. Chem. Sci. 12, 8362–8372 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Das P., Sercu T., Wadhawan K., Padhi I., Gehrmann S., Cipcigan F., Chenthamarakshan V., Strobelt H., dos Santos C., Chen P. Y., Yang Y. Y., Tan J. P. K., Hedrick J., Crain J., Mojsilovic A., Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nat. Biomed. Eng. 5, 613–623 (2021). [DOI] [PubMed] [Google Scholar]
- 33.Capecchi A., Cai X., Personne H., Köhler T., van Delden C., Reymond J. L., Machine learning designs non-hemolytic antimicrobial peptides. Chem. Sci. 12, 9221–9232 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Scarsini M., Tomasinsig L., Arzese A., D’Este F., Oro D., Skerlavaj B., Antifungal activity of cathelicidin peptides against planktonic and biofilm cultures of Candida species isolated from vaginal infections. Peptides 71, 211–221 (2015). [DOI] [PubMed] [Google Scholar]
- 35.Fais R., Rizzato C., Franconi I., Tavanti A., Lupetti A., Synergistic activity of the human lactoferricin-derived peptide hLF1-11 in combination with caspofungin against candida species. Microbiol. Spectr. 10, e01240-22 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.de Breij A., Riool M., Cordfunke R. A., Malanovic N., de Boer L., Koning R. I., Ravensbergen E., Franken M., van der Heijde T., Boekema B. K., The antimicrobial peptide SAAP-148 combats drug-resistant bacteria and biofilms. Sci. Transl. Med. 10, eaan4044 (2018). [DOI] [PubMed] [Google Scholar]
- 37.Boge L., Umerska A., Matougui N., Bysell H., Ringstad L., Davoudi M., Eriksson J., Edwards K., Andersson M., Cubosomes post-loaded with antimicrobial peptides: Characterization, bactericidal effect and proteolytic stability. Int. J. Pharm. 526, 400–412 (2017). [DOI] [PubMed] [Google Scholar]
- 38.Mosca D. A., Hurst M. A., So W., Viajar B. S. C., Fujii C. A., Falla T. J., IB-367, a protegrin peptide with in vitro and in vivo activities against the microflora associated with oral mucositis. Antimicrob. Agents Chemother. 44, 1803–1808 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Feng Q., Huang Y., Chen M., Li G., Chen Y., Functional synergy of α-helical antimicrobial peptides and traditional antibiotics against Gram-negative and Gram-positive bacteria in vitro and in vivo. Eur. J. Clin. Microbiol. Infect. Dis. 34, 197–204 (2015). [DOI] [PubMed] [Google Scholar]
- 40.Mookherjee N., Anderson M. A., Haagsman H. P., Davidson D. J., Antimicrobial host defence peptides: Functions and clinical potential. Nat. Rev. Drug Discov. 19, 311–332 (2020). [DOI] [PubMed] [Google Scholar]
- 41.De Oliveira D. M. P., Forde B. M., Kidd T. J., Harris P. N. A., Schembri M. A., Beatson S. A., Paterson D. L., Walker M. J., Antimicrobial resistance in ESKAPE pathogens. Clin. Microbiol. Rev. 33, e00181-19 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Ibrahim S., Al-Saryi N., Al-Kadmy I. M. S., Aziz S. N., Multidrug-resistant Acinetobacter baumannii as an emerging concern in hospitals. Mol. Biol. Rep. 48, 6987–6998 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Imai Y., Meyer K. J., Iinishi A., Favre-Godal Q., Green R., Manuse S., Caboni M., Mori M., Niles S., Ghiglieri M., Honrao C., Ma X., Guo J. J., Makriyannis A., Linares-Otoya L., Böhringer N., Wuisan Z. G., Kaur H., Wu R., Mateus A., Typas A., Savitski M. M., Espinoza J. L., O’Rourke A., Nelson K. E., Hiller S., Noinaj N., Schäberle T. F., D’Onofrio A., Lewis K., A new antibiotic selectively kills Gram-negative pathogens. Nature 576, 459–464 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Piddock L. J. V., Alimi Y., Anderson J., de Felice D., Moore C. E., Røttingen J.-A., Skinner H., Beyer P., Advancing global antibiotic research, development and access. Nat. Med. 30, 2432–2443 (2024). [DOI] [PubMed] [Google Scholar]
- 45.Jhong J. H., Yao L., Pang Y., Li Z., Chung C. R., Wang R., Li S., Li W., Luo M., Ma R., Huang Y., Zhu X., Zhang J., Feng H., Cheng Q., Wang C., Xi K., Wu L. C., Chang T. H., Horng J. T., Zhu L., Chiang Y. C., Wang Z., T. Y., Lee, dbAMP 2.0: Updated resource for antimicrobial peptides with an enhanced scanning method for genomic and proteomic data. Nucleic Acids Res. 50, D460–D470 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Shi G., Kang X., Dong F., Liu Y., Zhu N., Hu Y., Xu H., Lao X., Zheng H., DRAMP 3.0: An enhanced comprehensive data repository of antimicrobial peptides. Nucleic Acids Res. 50, D488–D496 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Wang C., Garlick S., Zloh M., Deep learning for novel antimicrobial peptide design. Biomolecules 11, 1–17 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Aguilera-Mendoza L., Marrero-Ponce Y., Beltran J. A., Ibarra R. T., Guillen-Ramirez H. A., Brizuela C. A., Graph-based data integration from bioactive peptide databases of pharmaceutical interest: Toward an organized collection enabling visual network analysis. Bioinformatics 35, 4739–4747 (2019). [DOI] [PubMed] [Google Scholar]
- 49.Camacho C., Coulouris G., Avagyan V., Ma N., Papadopoulos J., Bealer K., Madden T. L., BLAST+: Architecture and applications. BMC Bioinformatics 10, 1–9 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Li W., Godzik A., Cd-hit: A fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics 22, 1658–1659 (2006). [DOI] [PubMed] [Google Scholar]
- 51.Paszke A., Gross S., Massa F., Lerer A., Bradbury J., Chanan G., Killeen T., Lin Z., Gimelshein N., Antiga L., Desmaison A., Köpf A., Yang E., DeVito Z., Raison M., Tejani A., Chilamkurthy S., Steiner B., Fang L., Bai J., Chintala S., PyTorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 721, 8026–8037 (2019). [Google Scholar]
- 52.Ho J., Jain A., Abbeel P., Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 6840–6851 (2020). [Google Scholar]
- 53.Nichol A., Dhariwal P., Improved denoising diffusion probabilistic models. Proc. Mach. Learn. Res. 139, 8162–8171 (2021). [Google Scholar]
- 54.García-Jacas C. R., Pinacho-Castellanos S. A., García-González L. A., Brizuela C. A., Do deep learning models make a difference in the identification of antimicrobial peptides? Brief. Bioinform. 23, bbac094 (2022). [DOI] [PubMed] [Google Scholar]
- 55.Sidorczuk K., Gagat P., Pietluch F., Kała J., Rafacz D., Bąkała L., Słowik J., Kolenda R., Rödiger S., Fingerhut L. C. H. W., Cooke I. R., MacKiewicz P., Burdukiewicz M., Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Brief. Bioinform. 23, bbac343 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Li G., Iyer B., Prasath V. B. S., Ni Y., Salomonis N., DeepImmuno: Deep learning-empowered prediction and generation of immunogenic peptides for T-cell immunity. Brief. Bioinform. 22, bbab160 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Rackers J. A., Wang Z., Lu C., Laury M. L., Lagardère L., Schnieders M. J., Piquemal J. P., Ren P., Ponder J. W., Tinker 8: Software tools for molecular design. J. Chem. Theory Comput. 14, 5273–5289 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.De Jong D. H., Singh G., Bennett W. F. D., Arnarez C., Wassenaar T. A., Schäfer L. V., Periole X., Tieleman D. P., Marrink S. J., Improved parameters for the martini coarse-grained protein force field. J. Chem. Theory Comput. 9, 687–697 (2013). [DOI] [PubMed] [Google Scholar]
- 59.Wassenaar T. A., Ingólfsson H. I., Böckmann R. A., Tieleman D. P., Marrink S. J., Computational lipidomics with insane: A versatile tool for generating custom membranes for molecular simulations. J. Chem. Theory Comput. 11, 2144–2155 (2015). [DOI] [PubMed] [Google Scholar]
- 60.Marrink S. J., Risselada H. J., Yefimov S., Tieleman D. P., De Vries A. H., The MARTINI force field: Coarse grained model for biomolecular simulations. J. Phys. Chem. B. 111, 7812–7824 (2007). [DOI] [PubMed] [Google Scholar]
- 61.Páll S., Zhmurov A., Bauer P., Abraham M., Lundborg M., Gray A., Hess B., Lindahl E., Heterogeneous parallelization and acceleration of molecular dynamics simulations in GROMACS. J. Chem. Phys. 153, 134110 (2020). [DOI] [PubMed] [Google Scholar]
- 62.Bussi G., Donadio D., Parrinello M., Canonical sampling through velocity rescaling. J. Chem. Phys. 126, 014101 (2007). [DOI] [PubMed] [Google Scholar]
- 63.R. Gowers, M. Linke, J. Barnoud, T. Reddy, M. Melo, S. Seyler, J. Domański, D. Dotson, S. Buchoux, I. Kenney, O. Beckstein, “MDAnalysis: A python package for the rapid analysis of molecular dynamics simulations” in Proceedings of the 15th Python in Science Conference (SciPy, 2016), vol. 98, pp. 98–105. [Google Scholar]
- 64.Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., Vanderplas J., Passos A., Cournapeau D., Brucher M., Perrot M., Duchesnay É., Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 12, 2825–2830 (2011). [Google Scholar]
- 65.Jo S., Kim T., Iyer V. G., Im W., CHARMM-GUI: A web-based graphical user interface for CHARMM. J. Comput. Chem. 29, 1859–1865 (2008). [DOI] [PubMed] [Google Scholar]
- 66.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A., Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Ahmed T., Nisler C. R., Fluck E. C. III, Walujkar S., Sotomayor M., Moiseenkova-Bell V. Y., Structure of the ancient TRPY1 channel from Saccharomyces cerevisiae reveals mechanisms of modulation by lipids and calcium. Structure 30, 139–155.e5 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Abraham M. J., Murtola T., Schulz R., Páll S., Smith J. C., Hess B., Lindahl E., GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 1-2, 19–25 (2015). [Google Scholar]
- 69.Clinical and Laboratory Standards Institute (CLSI), Reference Method for Broth Dilution Antifungal Susceptibility Testing of Yeasts; Approved Standard, CLSI document M27-A (CLSI, ed. 33, 2008). [Google Scholar]
- 70.Clinical and Laboratory Standards Institute (CLSI), M100: Performance Standards for Antimicrobial Susceptibility Testing (CLSI, ed. 31, 2021). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figs. S1 to S12
Tables S1 to S19




