Skip to main content
Synthetic and Systems Biotechnology logoLink to Synthetic and Systems Biotechnology
. 2026 Jul 28;16:155–175. doi: 10.1016/j.synbio.2026.05.019

Artificial intelligence catalyzes antimicrobial peptide design

Yongqiang Liu a,b, Jie Hu a, Ning Zhang a, Yinqi Bai c,d,⁎, Yuan Yao e,f,g,⁎, Gaoxiang Chen a,⁎
PMCID: PMC13449796  PMID: 42569460

Abstract

With broad-spectrum, low resistance, and multifunctional properties, antimicrobial peptides (AMPs) are promising therapeutic agents against drug-resistant pathogens, yet their discovery and optimization still remain challenging due to the complexity of sequence-function associations. Artificial intelligence (AI), through the construction of comprehensive data-driven models that assisted with miscellaneous learning strategies, enables de novo peptide design by learning latent representations inherent in peptide sequences as well as their biological properties to ensure physically plausible and biologically relevant predictions. Consequently, this paradigm enhances the likelihood of designing peptide candidates with significantly improved therapeutic potential, reducing resource-intensive trial-and-error processes and revealing the transformative impact of computational innovation in advancing next-generation therapeutics. Here, we provide a snapshot of this field and survey two modes of AI-driven technologies for AMP design, one concentrated on identifying whether current data possess antimicrobial activity (identification-oriented) and the other on generating AMP candidates with potential therapeutic properties (generation-oriented). We also highlight the challenges and limitations that still hinder AMP development even accelerated by AI, as well as the foreseeable prospects, from finer-grained explorations to model-driven data enrichment and model enhancement.

Keywords: Antimicrobial peptides, Antimicrobial resistance, Microorganism, Computational biology, Artificial intelligence

1. Introduction

Antibiotics [1], [2], [3], [4], [5], [6], [7] exerted a significant influence on mitigating global mortality incurred by microbial infections. However, the emergence of antimicrobial resistance and other accompanying long-lasting side-effects severely limited their advancement in the pharmaceutical field [8], [9], [10], [11], [12], [13], [14], [15], [16], [17]. Antimicrobial peptides (AMPs) [18], [19], [20], [21], [22], [23], [24], [25], a part of the immune defense family that always consists of short amino acid sequences, possess the characteristics that enable the termination of various bacteria via mechanisms such as disrupting cell membranes, immunomodulation, specific target binding, and interference with metabolic processes [26], [27], [28], [29]. Consequently, AMPs have emerged as effective candidates for the development of novel alternative therapies. Their broad-spectrum activity and lower risk of antimicrobial resistance position them as super-potential agents in drug development, thereby mitigating the threat posed to the lives of patients and the economy of our society [30], [31]. The extensive exploration of AMPs from diverse sources, encompassing both natural and synthetic origins, has demonstrated their capacity to combat microbial pathogens. In addition to their antimicrobial properties, AMPs have also been identified as conferring additional benefits, such as promoting wound healing, exhibiting anti-tumor activities, and modulating immune responses [32], [33], [34], [35]. Recombinant or synthetic AMPs offer a viable and safe alternative for therapeutic applications in aquaculture [36], [37], [38], plant protection [39], food preservatives [40], [41], [42], [43], and even in biomaterials [44], [45], [46], [47], [48]. The heightened imperative to identify and design more effective AMPs has thus been underscored.

However, the systematic exploration and stepwise refinement of AMPs for clinical translation is a complex, multidisciplinary endeavor. It demands integration across diverse fields such as biology, materials science, chemistry, statistics, computer science, bioinformatics, and molecular informatics, posing significant challenges. Moreover, recent evidence reveals that AMPs exhibit an unexpected level of specificity and significant potential for synergistic interactions [49], [50], [51], [52], [53], [54]. This indicates that studying a single AMP in isolation may be insufficient to fully comprehend its mechanism of action. Instead, combining it with other antimicrobial peptides or antibiotics could potentially achieve a more potent therapeutic effect, which adds to the complexity of exploring and understanding the mechanisms. To this end, close cooperation among multiple parties is regarded as the key to success. Meanwhile, the employment of an intelligent strategy or a versatile tool [55] can certainly better facilitate and accelerate the process of exploring AMPs.

As a transformative technology accelerating progress across multiple domains, artificial intelligence (AI) [56], [57], [58], [59], [60], [61], [62], [63], [64] has emerged as the brightest rising star, capturing widespread attention. Individual intelligence systems, along with associated methods and models, have been widely and successfully applied in multiple fields such as natural language processing, image recognition, robotics, and games. By training complex weight-driven models (such as deep neural networks, DNNs) to learn patterns and representations from large amounts of data, these methods are able to generate new required data or make decisions or predictions for various tasks. In the past few years, to get rid of the limitations that time-consuming and low-throughput brings in biological exploration, these AI-based learning models have made a lot of contributions, e.g., genomic analysis [65], [66], [67], [68], structural modeling and prediction [69], [70], [71], [72], [73], [74], medical imaging and disease diagnosis [75], [76], and novel drug discovery [77], [78], [79], [80], [81].

To date, numerous computational methodologies have been devised for expediting the design of novel AMPs. In this review, we survey how these AI-based models have assisted various aspects of AMPs’ design. Our analytical synthesis traverses the developmental arc of AI-assisted exploration of AMPs. We present a snapshot of the operational intricacies inherent to these methodologies, delineating the sequential stages of data aggregation, dataset feature extraction, model training, and inference. Particularly, we investigate two crucial modes for AMP design, the identification-oriented mode and the generation-oriented mode, as well as an overview of the methodologies employed therein. Finally, we analyze the prevailing challenges impeding AMP advancement and proffer prospective developments in this field.

2. Preliminaries

With lower risk of antimicrobial resistance, broad-spectrum antibacterial activity, low propensity to develop toxicity, multiple anti-activity (antibacterial, antiviral, antifungal, anticancer, antiprotozoal, etc.) [82], [83], [84], [85], [86], AMP has emerged as a treasure aiding clinical therapeutics and a warrior fighting against kinds of diseases, such as respiratory diseases [87], [88], [89], [90], eyes diseases [91], [92], gastrointestinal diseases [93], [94], [95], [96], bone and joint infection [97], [98], wound healing and skin infections [99], [100], and oral diseases [101], [102]. Typically, they rely on disrupting cell membranes to destroy targets [103], [104], interfere with metabolic processes, or promote immune cell activation and immune response by lowering the barrier to cell entry [105], [106], [107]. During the detailed exploration of AMPs’ mechanisms of action, some external manifestation rules have been summarized, as follows.

  • •

    Element-level. The composition of amino acids and the physicochemical properties, such as charge and amphiphilicity. Generally, the positive charge of AMPs1 enables selective interaction with anionic bacterial membranes, while their hydrophobic regions facilitate engagement with the hydrophobic core of these membranes [105].

  • •

    Sequence-level. Sequence length, distribution (position), and motifs. For instance, the length of the AMPs is usually small (6–100 amino acids) [111].

  • •

    Structure-level. Specific structural characteristics are more likely to be associated with AMPs, such as α-helical, β-sheet, lasso, and cyclic [107], [112], [113], [114], [115].

Over the past period of time, researchers have identified and developed novel AMPs based on these enlightenments. However, the resultant methodologies have notably exhibited limitations in throughput and efficiency. AI-based methods, characterized by their data-driven trait, facilitate the capture of a broader range of data features, significantly accelerating the design of AMPs. Fig. 1 provides an overview of AI-driven AMP design. AI encompasses three fundamental components: input data, models or algorithms, and output results, denoted as X, f(.), and Y, respectively. At its core, it operates on the principle of transforming input data through algorithmic processing to generate output results, succinctly represented by the formula Y=f(X). For instance, the antimicrobial activity (Y) of an input sequence (X) can be determined by passing it through an AI-guided model (f(.)). Generally, AI consists of two parts, training and inference. The training phase entails the systematic optimization of a model’s parameters on a training dataset by utilizing algorithms such as stochastic gradient descent to minimize a predefined loss function, which quantifies the discrepancy between the model’s predictions and the ground truth. The inference phase capitalizes on the model’s previously learned parameters to infer outputs from new inputs, achieving the desired learning purpose, such as identification and generation. To enhance generalization and address overfitting and bias–variance tradeoffs, AI models incorporate techniques such as minimal empirical risk optimization, regularization, data augmentation using domain-invariant transformations, cross-validation with stratified sampling, and ensemble methods. A well-designed AI model requires exceptional performance, robust generalization, interpretable reasoning, ethical compliance, and scalable adaptability across diverse contexts.

Fig. 1.

Fig. 1

Overview of designing AMPs accelerated by AI.

AI’s efficient framework makes it highly applicable across various fields, including AMP exploration. As illustrated in Fig. 1, constructing an AI model for AMP design begins with collecting relevant biological data and extracting representations to serve as inputs. Here, we summarize the main public AMPs’ databases [109], [116] in Table C.1 Appendix C. The five representational types are summarized by Wan et al. [117], including global descriptors (include sequence composition, structural features, and physicochemical properties) [118], [119], [120], [121], [122], [123], [124], sequence-based representation [125], [126], graph-based representation [127], [128], [129], 3D representation [130], [131], [132], and data-driven representation (AI-based automatically learning features) [133], [134], [135], [136]. After training on these represented data, the model can generate core predictions to expedite novel AMP discovery. Table C.2 Appendix C presents major AI-based methods dedicated to advancing AMP development.

In recent years, AI-based models have been applied to AMP design, which can be summarized into two levels: identification and generation. Identification emphasizes determining whether a given peptide sequence exhibits antimicrobial activity, encompassing tasks such as classification and regression. This foundational step provides critical insights into sequence-function relationships and serves as a prerequisite for the subsequent generative level. At the generative level, AI-guided models are employed to design novel AMP candidates de novo, creating sequences with potential therapeutic properties from scratch. This is a process where sequences go from nothing at all, and hence often requires iterative validation with multiple identification models to ensure the generated candidates meet desired activity, specificity, and safety criteria. In the following sections, we will delve into each of these levels, exploring their methodologies, advancements, and limitations.

3. Identification-oriented

Identification refers to distinguishing whether a given sequence possesses antimicrobial activity. Over the past few decades, AI-based identification methods for AMP design have undergone significant development. The basic tasks of identification are classification and regression, as shown in Fig. 2. The former directly maps the inputs into distinct categories, e.g., non-AMPs and AMPs, while the latter calculates the regression values of the inputs, e.g., the prediction of minimum inhibitory concentration (MIC) values, which not only reflect whether it is an AMP by setting a threshold but also indicate its antimicrobial effectiveness. Based on our investigation of AI-based methods for AMP identification, we summarize three model-building and design styles: retraining, ensemble, and hybrid.

Fig. 2.

Fig. 2

Overview of identification.

3.1. Retraining models

Retraining existing models on new datasets is essential for maintaining model accuracy in adapting to new tasks with specified conditions or scientific insights. First, retraining ensures models remain accurate and effective over time, particularly in dynamic environments where data evolves and new discoveries emerge. For instance, as peptide sequences continue to be identified as AMPs, enriching the database, retraining allows models to incorporate the latest findings, thereby enhancing predictive capabilities and ensuring alignment with current scientific knowledge. Second, retraining serves as one of the foundations of transfer learning, often excelling in new tasks, such as identifying AMPs. To be well-compatible with new tasks, designers typically need to select appropriate representations according to task characteristics and modify the model to fit these requirements precisely, like a glove. There are two specific modes of this model modification: full parameter training and fine-tuning.

Full parameter training refers to the process of training a model by updating all of its parameters based on the training data. It allows models to be thoroughly optimized for specific scientific tasks and datasets. It is particularly necessary when the target task significantly differs from the original pre-training task or when the available data provides novel insights that the model needs to capture comprehensively. For instance, Sun et al. [137] transfer the sequence-based data into the graph-based data to retrain a two-layer GCN model to identify AMPs. They tested that a two-layer GCN can achieve the optimal performance, while too many GCN layers caused the model to be over-smoothing. During training, all model weights were updated, and the activation function, window size, convolution size, learning rate, loss rate, and training epoch were set to ReLU, 15, 200, 0.01, 0.5, and 150, respectively. Lee et al. [138] retrain SVM-based classifiers to investigate α-helical AMPs and the interrelated nature of their functional commonality and sequence homology. Specifically, they employed the L2-norm and trained three SVM classifiers with three distinct kernels, linear, polynomial, and radial basis function (RBF). Similarly, all weights were updated during training, and the hyperparameters were optimized via a grid search strategy. Through testing and evaluation, they recommended adopting the linear kernel. This recommendation is attributed to its superior performance and interpretability, coupled with the fact that the performance improvement offered by non-linear kernels proved insufficient. Fine-tuning refers to the process of adapting a pre-trained model to a specific task or domain by making minor adjustments to its parameters. This process enables the model to learn specific patterns and features relevant to the target task, thereby improving its accuracy and effectiveness. It allows models to leverage knowledge from pre-training while being customized for specific applications. Fine-tuning bridges the gap between general-purpose pre-trained models and the complex requirements of scientific tasks without requiring extensive computational resources or large datasets. For instance, large language models (LLMs) like BERT, with numerous parameters, are well-suited for fine-tuning in downstream tasks such as AMP identification [139]. Wang et al. [140] first trained an AMP-centric language model as the foundation model, called AMP-GPT, on a dataset of peptides extracted from UniProt [141]. Subsequently, they fine-tuned AMP-GPT to develop two submodels, AMP-Prompt and AMP-MIC. AMP-Prompt adopted a contrastive prompt tuning strategy by initializing the prompt embedding layer based on the word embedding of AMP-GPT. Only the parameters of this layer were updated, while the parameters of AMP-GPT were held fixed.AMP-MIC fine-tuned the AMP-GPT on the MIC datasets of three bacterial species, namely S. aureus, E. coli, and P. aeruginosa.

Although retraining offers simplicity and efficiency in operational implementation, it still has limitations. A single model that performs exceptionally well in only a few fields makes it difficult to fit perfectly into newly migrated problems. Ignoring the particular features of the new situation will often lead to overfitting and poor generalization ability.

3.2. Ensemble frameworks

An ensemble framework integrates multiple distinct models with an ensemble strategy, such as boosting and voting, to enhance prediction accuracy by coordinating multiple inferred results. These models can be of the same type (homogeneous) or different types (heterogeneous). The benefit of ensemble frameworks lies in harnessing the collective knowledge of multiple models to achieve superior performance compared to individual models, which are often constrained by their inherent biases and variances. In particular, ensemble models exhibit greater stability than individual models, as they are less susceptible to minor fluctuations in training data, which is particularly crucial for advancing scientific discovery, including AMP discovery. For instance, Ma et al. [111] implemented three models, ATT, LSTM, and BERT, for mining human gut microbiomes. By setting the unanimous voting strategy, they recognized 2349 sequences as candidate AMPs. Among these candidates, 216 were chemically synthesized, and 181 showed antimicrobial activity. Chen et al. [142] followed the same pipeline to mine the global marine microbial database and identified 121 unique candidate AMPs, of which 10 novel AMPs are successfully validated. Similarly, based on the same ensemble method, Xu et al. [143] mined 27 candidate AMPs from sludge and successfully synthesized 25 peptides among them. As a result, 21 peptides with antibacterial activity are experimentally verified against 4 strains.

The ensemble paradigm demonstrates notable efficacy in augmenting model generalization capability and operational robustness through the systematic integration of diverse constituent learners. This methodological framework capitalizes on the complementary strengths of heterogeneous base models, effectively mitigating individual model biases while enhancing collective decision-making through sophisticated aggregation mechanisms. By strategically amalgamating predictions from multiple learners, the ensemble architecture substantially reduces variance-related errors and improves resistance to data distribution shifts, thereby achieving statistically significant performance elevation compared with singular modeling approaches. Although this approach collects opinions is of wide benefit, it still has limitations. First, multiple models working together means more computational overhead. Second, the ensemble methods are very dependent on the strategy, which has limited effectiveness. For instance, the unanimous voting strategy greatly reduces the output space, thus missing some potential AMPs. In summary, the choice of ensemble technique should align with the specific task, data representations, and available computational resources to maximize benefits.

3.3. Hybrid building

Beyond directly retraining existing classical models or constructing ensemble frameworks, some designers propose novel architectures tailored to the unique challenges and data characteristics of AMP identification. A common approach is to combine modules2 from different existing models to rebuild a new architecture. The ensemble framework focuses on how to select or combine outputs from different models, whereas hybrid building focuses on integrating modules with diverse functionalities into a novel architecture. Each module fulfills its specific objective and functionality. Through mutual splicing and sequential cooperation, these modules collectively constitute a robust model. Not as compartmentalized and independent as the ensemble methods, the hybrid approach aims to embody all the strengths in one entity. For instance, iAMP-CA2L [144] is a deep neural network architecture combining CNN, BiLSTM, and SVM. The CNN part extracts amino acid sequence features effectively via feature mapping, the BiLSTM part captures contextual sequence features, and the SVM part is responsible for classification. Different modules assume distinct responsibilities. Li et al. proposed AMPlify [145], an attentive deep learning model for AMP identification, which applies two types of attention mechanisms layered on a BiLSTM layer. The Bi-LSTM layer encodes positional information from the input sequence, while one attention layer refines sequence representation using multiple weight vectors, and the other generates a summary vector by learning contextual information from the previous layer.

Hybrid approaches allow for the integration of diverse module capabilities, thereby creating a more versatile and comprehensive model that leverages the advantages of each constituent component. This amalgamation offers the potential for performance improvements and innovative solutions tailored to specific tasks. However, due to the necessity of integrating the modules as a cohesive whole, it is typically feasible to integrate only a limited number of individual elements rather than combining an extensive array of models, such as the dozen or more utilized in ensemble methods. Furthermore, the design of novel architectures or hybrid models demands substantial expertise and involves intricate design decisions.

To summarize, AI-based approaches enable the systematic exploration of a more profound feature space, thus capturing intricate relationships between sequences, distant source correlations, and salient features pertinent to identification. Notwithstanding the aforementioned distinctions among the three model types, several common limitations persist.

3.4. Challenges and limitations

Although there has been a proliferation of AI models designed for identifying AMPs, numerous challenges and dilemmas remain. We present a systematic summary of these issues across three dimensions, including task, data, and model.

Task. In general, data-driven and end-to-end models tend to understate the problem’s complexity, thereby neglecting the exploration of the intrinsic mechanistic nature of the task. Here, we summarize two points: the similarity trap and the absence of granular focalization. The similarity trap refers to the barriers to solution transferability, manifested through over-reliance on prior experiences and established frameworks while neglecting the embeddedness of problems within unique contextual ecosystems. Superficially analogous tasks often conceal divergent core mechanisms, necessitating context-driven solution reformulation, even in small aspects. For instance, although highly accurate structural prediction models emerge one after another, unfortunately, a single predicted structure (or conformation) may not sufficiently represent a given peptide because the short peptide’s conformational flexibility [117], [146], [147], [148], [149]. Therefore, the methods that rely on predicting structures to discover macromolecular proteins, such as enzymes, might not be quite suitable for identifying AMPs. The absence of granular focalization refers to the tasks constrained by broad generalizations and superficial assessments, while lacking systematic exploration into the essential and underlying mechanisms. As a result, it remains a niche area, with significant challenges to be overcome for broader applications. Synergistic mechanisms and chemical modifications are two classical archetypes. For example, many works have revealed that naturally co-occurring AMPs with distinct functions can synergize together [150], [151], [152], [153], but few AI-guided models directly explore or predict synergistic mechanisms among AMPs. The complexity of this problem and the difficulty of modeling constitute one aspect, while the other is the lack of necessary synergistic data. It is beneficial to enrich relevant data by collecting more measurement feedback about the synergisms among AMPs in vivo, rather than merely testing MIC values of individual components in vitro [154]. Moreover, although chemical modification is one of the important strategies to enhance antibacterial activity or improve stability, such as lipidation and N-terminal acetylation, current AI-guided approaches are generally confined to the 20 basic amino acids and rarely directly contribute to chemical modification.3 For instance, in the context of natural language models (NLMs) applied to AMP design, the vocabulary typically only consists of the 20 basic natural amino acids [111].

Data. The quality and quantity of training data are paramount to the model’s ability to learn effectively and generalize well, directly impacting its predictive accuracy and robustness. While publicly available datasets (Table C.1) exist for designing AMPs, their size is limited compared to extensive datasets available in computer science, due to the cost associated with wet labs. For instance, the classical computer vision dataset ImageNet contains over 14 million labeled samples [156], whereas APD3 [157] has only 5099 peptides for AMP designing. Furthermore, three limitations we have concluded in terms of AMPs’ data. First, specific data is scarce and noisy. While wet experimental validation is the sole criterion for discriminating whether a peptide is an AMP or not, not all AMPs have MIC data. Meanwhile, due to the varying environmental standards of wet labs, MIC data has noise, posing a challenge in the regression tasks. Second, coarse-grained. Antimicrobial activities, such as antibacterial, antifungal, and antivirus, are highly prevalent, yet there is a dearth of individual-level anti-labels, such as anti-E. coli or anti-S. aureus. Although datasets specifically designed to provide MIC values, such as GRAMPA [158], record more targeted information, they are outdated due to stagnant updates. Unfortunately, even though AMPs are broad-spectrum, they are not effective against all bacteria and still exhibit specificity. For instance, Ma et al. demonstrated their predicted AMP c_AMP67 could effectively destroy three strains of K. pneumoniae (NK04047, NK06129, and NK08334) but had little effect on NK01067 [111]. Therefore, coarse-grained categorization prevents researchers from knowing exactly which species a known AMP is effective against, which is not conducive to experimental replication and validation. There remains a need to develop more fine-grained works, such as the microbial strain-specific AMP Identification [159]. Third, negative data. While it is relatively easy to access several AMPs’ databases as summarized in Table C.1, there is a notable absence of specialized databases for non-AMPs, which is one of the major dilemmas in this field. Researchers often collected unlabeled samples as negative data, for example, a peptide without certain keywords related to ‘antimicrobial’ from UniProt as a negative one. However, this approach potentially leads to false-negative data that interfere with, or even mislead, AI models trained on them. Several strategies have been employed to enhance the selection of negative data, such as filtering similarity, label smoothing, and positive-unlabeled learning. These strategies are only temporary fixes rather than permanent solutions. All these methods are based on the common assumption that negative samples exist and are detectable. Unfortunately, it is not feasible to determine a negative sample since only wet-lab experimental validation can definitively determine whether a peptide is active or not, and it is impossible to test against all bacteria. Proving the positive is straightforward, but proving the negative is extremely challenging, which is referred to as unsurveyability [160]. Therefore, we cannot expect precise results from a vague categorization task. To this end, a change of perspective is needed to make the negative provable. Fine-grained classification could be a solution, focusing only on individual-level anti-labels as previously mentioned. For instance, experimental verifications can ascertain whether a peptide has anti-E. coli activity, simultaneously identifying positive and negative samples.

Model. High-quality data needs to be supported by a suitable model to enhance the effectiveness of AMP identification. However, AI models are not a panacea and still have limitations. First, they lack interpretability, especially for models within the deep neural network (DNN) framework. Their black-box attribute obscures the relationship between data and predictions, hindering the elucidation of intrinsic mechanisms. While various efforts have been made to enhance their explainability [161], [162], [163], [164], [165], [166], such as the attributional paradigm Integrated Gradients (IG) [161], these methods are post-hoc and primarily input-feature oriented. For instance, using physicochemical features inferred by explainable approaches and validated through wet experiments is a conventional strategy. However, providing specific attributional interpretations of the hidden space remains challenging, especially for deep learning models. Second, the absence of uniform criteria among models may result in inequitable comparisons. AI-based approaches exhibit significant variability in data quality, feature extraction, core algorithms, evaluation strategies, and metrics, with each study following its own protocols. Although several works [29], [167] have endeavored to conduct fair comparisons of various models for identifying AMPs, their comparative outcomes are inherently time-sensitive and diminish in relevance with the continuous emergence of novel methodologies. Additionally, the dynamic nature of training datasets presents another layer of complexity. Unlike static datasets in fields such as computer vision that are fixed and molded, antimicrobial peptide databases [157] are constantly evolving and expanding. Consequently, whenever a new model is designed, it is invariably trained on the latest version of the dataset. Simultaneously, authors often neglect to retrain baseline models on the newly collected datasets when benchmarking their new models. As a result, it is increasingly difficult to ascertain whether superior predictive performance is primarily due to the innovative design of the new method or the enhanced quality of the training data. Third, generalization. Generalization refers to a model’s ability to perform effectively on independent, unseen data rather than merely on the training data to which it has been fitted. This capability is crucial because the ultimate objective of AI-based approaches is to extrapolate from specific training instances to broader, general scenarios. Such models mitigate overfitting or underfitting to the peculiarities of the training set, thereby ensuring robust performance across a diverse range of inputs. Although several techniques, such as regularization and cross-validation, can enhance generalization, some critical aspects are often overlooked. For instance, the similarity between the training and the test dataset is frequently neglected in discussions. High similarity between these datasets may lead to seemingly high generalization performance. However, this setting is ultimately meaningless since it is vulnerable to subsequent unseen data that may exhibit significant variability. Fourth, reproducibility. Heil et al. [168] proposed three standards for computational reproducibility in 2021, including bronze, silver, and gold. Recently, Sidorczuk et al. [167] investigated the reproducibility of 26 models for AMP prediction and found that approximately 70% of the models represented non-reproducible work. Only 8 models met the minimal bronze standard. This study advocates for authors to provide exhaustive documentation regarding their models, with particular emphasis on training details. This comprehensive transparency would substantially enhance the reproducibility, thereby precluding futile endeavors, conserving valuable resources, and maintaining the momentum of scientific advancement in this domain.

3.5. Discussion

We advocate that rigorous architectural comparison demands identical training and test dataset deployments across all competitors, since only under such controlled conditions can genuine algorithmic superiority be isolated from data variance. Expanding the comparisons to fine-grained metrics is equally desirable, yet this inevitably obliges a complete retraining cycle for every model. However, it is an undertaking that is both time-intensive and computationally expensive. For this reason, few reviews satisfy the standard in its strictest form, and it is not the primary objective of our work. Therefore, no quantitative inter-model comparison is undertaken in this paper. Fortunately, some benchmark-oriented studies [139], [167], [169] prioritize uniform evaluation conditions and deliver precisely this level-playing-field benchmark. For instance, Gao et al. retrained the models on a unified AMP dataset and reported comprehensive metrics [139]. Rather than duplicating their effort, we therefore recommend directly referencing this type of benchmark-oriented literature.

4. Generation-oriented

Building upon the capability to discern antimicrobial sequences, the subsequent challenge lies in the development of methodologies capable of generating a de novo repertoire of sequences endowed with antimicrobial potency. Beyond the determinations that are predicted from the identifying models, generative models facilitate the creation of novel data points that are consistent with the training data, such as a set of candidates with antimicrobial activity, reflecting a deeper understanding and insight into the probabilistic nature and structure of the data. This advancement would critically enhance the antimicrobial arsenal, offering novel therapeutic options and bolstering the ability to combat resistant pathogens. The generation process, involving the creation of new candidate sequences from scratch, is fundamentally centered on sampling. Specifically, it can be categorized into two aspects, sampling from the observable space and sampling from the latent space, as illustrated in Fig. 3.

Fig. 3.

Fig. 3

Overview of generation.

4.1. Generation from observable space

Equipped with pre-trained identification models, exploring the observable space to check if a target sequence has antimicrobial activity is straightforward. Within this paradigm, peptide sequences in the observable space authenticated by the identifier(s) can be regarded as newly generated AMPs. For instance, the simplest approach is to screen the entire peptide sequence space. Huang et al. [170] mined the entire virtual library of peptides composed of 6–9 amino acids to identify potent antimicrobial peptides with an identifier. They synthesized the top-10 predicted peptides via solid-phase synthesis and tested their MIC values against S. aureus. All 10 peptides exhibited antimicrobial activities. This special antimicrobial peptide generation pattern is attributed to the limited discrete entire space, which benefits from the relatively short sequence lengths. In contrast, for continuous domains, employing exhaustive exploration through traversal is impractical. However, this approach, although effective, is not optimistic in terms of resource consumption for exploring the entire space, especially as the sequence length increases. Therefore, mining from the partial space is more desirable. Some candidate datasets with a greater likelihood of the presence of potential AMPs are favored [111], [142], [143], [171]. For instance, Ma et al. [111] mined AMPs from the human gut microbiome, while Chen et al. [142] mined the global marine microbial database. Both of them discovered novel AMPs successfully. The quality of this mining-driven mode hinges jointly on the candidate datasets and the identifiers employed. Selecting an inappropriate candidate dataset may prove futile. Furthermore, given that different identification models can produce varying results and no single model is universally superior, it is crucial to devise a suitable identification strategy prior to large-scale mining. Additionally, with the continuous emergence of new models, the results of mining-driven generation remain time-sensitive.

Another prevalent pattern draws inspiration from biological mutations and evolutionary processes. These approaches first sample sequences from AMP databases and then simulate the process of evolution or mutation by systematically altering the target sequences to generate novel candidates. Among these methods, genetic algorithms (GAs) [172], [173], [174] and neural language models (NLMs) [175], [176], [177], [178], [179] are particularly favored. GAs maintain a population of candidate solutions and iteratively evolve them through operations such as selection, crossover, and mutation, thereby efficiently exploring the solution space. For example, Porto et al. leveraged a genetic algorithm to explore the potential space of the guava peptide, Pg-AMP1, and generated the guavanin peptides, one of which displayed potent activity against Gram-negative bacteria [173]. NLMs employ a sophisticated mechanism involving masking and generation to comprehend and reconstruct sequences. They predict concealed tokens based on contextual information derived from unmasked tokens, thereby learning the nuanced relationships between amino acids and their surrounding environment. However, despite exhibiting clear goal-oriented behavior and facilitating efficient exploration of local spaces, these methods are inefficient in exploring the global space. Consequently, the diversity of sequences generated by such approaches is limited.

4.2. Generation from latent space

Different from the generative modes that rely on sampling from the observable space, some DNN-based models [180], such as variational autoencoders (VAEs) [181], [182], [183], [184], generative adversarial networks (GANs) [185], [186], [187], [188], and diffusion models [189], [190], [191], [192], [193], [194], generate novel candidates by sampling from the latent space. The latent space serves as a compact, continuous representation that encapsulates the intrinsic structure and variability of the training data, enabling the model to learn a more efficient and meaningful encoding of the data’s underlying features. Take VAEs as an example, by first mapping the input data to a lower-dimensional latent space and then decoding it back to the original data space, these models produce a wide array of outputs while maintaining coherence and relevance to the learned data distribution. Sampling from the pre-learned latent space, VAE-based models, such as PepVAE [181], LSSAMP [184], and PepCVAE [183], facilitate the generation of new samples analogous to the training data but with novel characteristics. In the case of GAN-based models, the generator learns to produce high-quality candidate sequences from the latent space through continuous correction and guided feedback from the discriminator. Lin et al. [185] synthesized 8 AMPs generated from their proposed GAN-based models, two of which exhibited broad-spectrum antibacterial effects and were effective against antibiotic-resistant bacterial strains, such as S. aureus. Diffusion-based models rely on adding noise to build the latent space, such as Gaussian noise, and then refining random noise into data samples by reversing a diffusion process. Finally, by sampling the latent space, the model can generate new elements that are similar to the training data. Based on this principle, Chen et al. proposed AMP-Diffusion [195], Qi et al. proposed CDiffusion-AMP [196], and Wang et al. proposed Diff-AMP [197] for AMPs generation.

4.3. Complete generative pipeline

We here summarize the complete generative pipeline, spanning from in silico generation to wet-lab validation, including six core steps, as follows.

  • •

    Task anchoring. In light of objective conditions, such as available computational resources and the accessibility of in-house datasets, researchers need to crystallize the task objective and to chart the entire implementation pathway. For example, generation from the latent space or generation from the observable space.

  • •

    Data collection and representation. AMP datasets are customarily assembled through the systematic integration of multiple repositories, thereby augmenting sample size. Concomitantly, a diverse repertoire of feature-extraction protocols has been elaborated to capture the signatures of these sequences. The corresponding databases and strategic feature-encoding schemes are catalogued in Appendix B for reference.

  • •

    Model designing, training, and inference. The model architecture should be tailored to the precise objectives of the task and the characteristics of the curated data. Some hyperparameters might inevitably rely on empirically informed choices to secure peak performance. Computational resources must be meticulously scheduled in advance to obviate superfluous expenditure.

  • •

    Computational evaluation and filtering. Generated candidates undergo multi-layered filtering: predictions for antimicrobial activity, novelty and diversity, and biological criteria. We summarize these evaluation and filtration strategies in Appendix A.

  • •

    Wet-lab validation. The ultimately shortlisted candidate sequences will undergo wet-lab validation via a curated suite of assays, including organisms and antimicrobial assay [111], [140], [173], [182], [185], [198], [199], hemolysis assay [111], [140], [170], [173], [198], plasma stability test [140], time-kill assay [140], [170], [173], membrane permeabilization [170], [173], [182], [200], cytotoxicity assay [111], [170], [182], [200], NMR experiment [173], [174], [201], CD spectroscopy [173], resistance assay [140], [170], [182], and in vivo activity test [140], [173], [174], [182], [199].

  • •

    Feedback. Two feedback types need to be accommodated. At the knowledge-base aspect, wet-lab outcomes are scrutinized to determine whether they corroborate established tenets, overturn prevailing dogma, or unveil previously unrecognized biological principles. At the modeling aspect, the wet-lab outcomes are channeled back into the pipeline, enabling reinforcement-driven retraining (reinforcement learning) that iteratively refines and evolutionarily advances the model.

Fig. 4 shows an overview of the pipeline. The initial four steps constitute the foundational and uniform core of the pipeline; nevertheless, wet-lab validation is not universally furnished across the studies, especially in computation-oriented literature. Analogously, the reinforcement-cycle feedback invoked in Step 6 is also far from being a consistently implemented component of every modeling framework. A fully elaborated exemplar is therefore indispensable, as it affords a step-wise point from which to apprehend the complete pipeline in operation. To break the “bacteria-only” bottleneck of current AMP generators, Wang et al. [200] collected four datasets and extracted one-hot features to train their designed modeling framework PepDiffusion, which can generate antibacterial and antifungal peptides from the latent space. Then, they used three filtrations (classification, clustering, and coarse-grained MD simulations)4 to sift through the 600,000 generated peptide sequences, and thus selected 40 for wet-lab validations, including MIC determination assay, membrane permeabilization assay, drug resistance assay, hemolysis assay, cytotoxicity assay, and in vivo studies (Murine). The result shows that 25 filtered peptides exhibited antibacterial or antifungal activities. Among them, 5 had selective activity against specific fungal species, suggesting that these AMPs have limited universal activity.

Fig. 4.

Fig. 4

Overview of pipeline.

4.4. Challenges and limitations

The generative models not only facilitate the generation of a substantial number of candidate sequences but also endeavor to minimize the inclusion of negative samples during the training phase, thereby effectively circumventing the influence of the ineradicable noise inherent in negative samples, as elaborated in Section 3. Nevertheless, in addition to sharing some common drawbacks with identification models, such as non-interpretability, they are also subject to several additional limitations.

Duplicate candidates. In contrast to the mining-driven generation mode, which does not produce duplicate candidates, generative models often yield multiple duplicate sequences. For example, Lin et al. [185] had to remove 1225 duplicate GAN-designed peptides before proceeding to experimental validation. Unfortunately, such repeated sequences are difficult to avoid, as generative models essentially compress the entire sequence space into the functional sequence space. Duplicates arise both from sampling the same elements in the latent space and from mapping different sample points from the latent space to the same sequence.

Delayed and indirect feedback. Wet experimental validation stands as the sole definitive method for identifying antimicrobial activity. However, this procedure is time-consuming and costly. Consequently, an intermediate step is frequently indispensable to refine the selection of candidates prior to experimental validation. This step typically involves supplementary identification models, such as those predicting the MIC values, to rank and filter the generated candidates. As a consequence, the selection is narrowed down to those sequences with the highest likelihood of antimicrobial activity, thereby enhancing the efficiency and accuracy of subsequent experimental validation. Nevertheless, this approach introduces an additional layer of influence from the identification models employed in the double-check, which interferes with the direct reflection of the intrinsic generative quality of the original generative models.

Ambiguous metrics. The pre-segmentation of the test dataset prior to the training phase enables the quantitative assessment of identification models through metrics such as precision, recall, and F1-score. In contrast, evaluating generative models is more intricate, as it aims to generate both high-quality and diverse samples. This complexity necessitates a combination of quantitative metrics to ensure the models’ effectiveness and utility. Meanwhile, due to the absence of a unified standard, the subjective nature of quality assessment in generative models introduces variability, as different evaluators have varying perceptions of the generated samples’ quality. Take diversity as an example, Wang et al. utilized Levenshtein distance [202] to evaluate the diversity of candidates generated by their VAE-based model LSSAMP [184], while Van Oort et al. employed the Gotoh global alignment algorithm [203] for diversity assessment.

4.5. Discussion

We here turn to a discussion of feature extraction, a common procedural component that is indispensable to both discriminative and generative paradigms. In Section 2, five feature extraction schemes are briefly introduced, following a comprehensive review [204]. We survey the specific implementation details of these five categories in Appendix B. Then, we investigate whether there exist some features that are unique to AMPs. To date, no feature has yet been devised specifically for this purpose. Researchers prioritize generalized protein representations, refraining from subjective weighting of individual features and instead adopting rigorous statistical strategies (e.g., covariance analysis [205]) for subsequent feature selection (dimensionality reduction). Conversely, several feature strategies that are commonplace in other protein prediction domains remain largely absent from AMP tasks. For instance, three-dimensional features. Although some studies extract three-dimensional features as AMP representations [206], such approaches are adopted far less frequently than in PDB (protein data bank)-centric prediction tasks such as protein–protein interaction site (PPIS) identification [131], [207]. Moreover, even when three-dimensional features are extracted, this representation does not serve as the sole input to models like sequence representations, but rather combines with numerous other distinct representations as inputs [206]. It is attributable to the fact that AMPs are short sequences, rendering their three-dimensional conformations inherently unstable and thus introducing substantial uncertainty. In summary, researchers have exercised considerable caution in incorporating three-dimensional features into AMP design tasks.

5. Flexible design

Enhanced by artificial intelligence, researchers can adaptively design AMPs from diverse perspectives, offering a rich tapestry of solutions and opportunities for innovation.

Diverse perspectives. Multifaceted analysis, which examines issues from diverse perspectives, is instrumental in enhancing our understanding and resolution of complex problems, thereby providing a significant impetus for the advancement of the AMP design. For instance, existing works have taken the perspective on the interaction with membranes [208], [209], molecular dynamics (MD) simulations [182], [210], [211], site prediction [212], and functional domains [174], [213]. Each view contributes uniquely to the advancement of AMP-based characteristics, highlighting the necessity of embracing diverse perspectives in the pursuit of novel and effective antimicrobial solutions. For example, MD simulation provides atomic-level details of AMPs interacting with lipid bilayers and other biological membranes. This elucidates the intricate mechanisms of how AMPs disrupt bacterial membranes while sparing cells, aiding in the design of more stable and effective peptides, and thus serving as a bridge between computational design and virtual experimental validation [182]. Various perspectives have jointly contributed to the flourishing of the process of AMP design, holding the promise of exploring the design-function space of AMPs at an increased rate and depth.

Multipath. Three crucial principal paths govern AMP development, mining, redesign, and de novo design. Each of them demonstrates distinct operational paradigms within the AMP discovery pipeline. Mining exploits existing databases or potential in-house datasets, such as marine microbial databases [142] and human gut microbiomes [111], to rapidly discover novel AMP candidates, accelerating the identification of peptides with specific antimicrobial properties. This primary data-driven mode accelerates candidate identification while minimizing experimental validation requirements, though constrained by existing sequence bias, enabling rapid preliminary screening of established antimicrobial scaffolds. Rational redesign enhances existing AMP performance through strategic modifications. This iterative process improves pharmacological parameters (such as target affinity, proteolytic stability, and cytotoxicity) while maintaining core antimicrobial motifs. Compared with mining, redesign is optimization-focused, typically building on known peptides, which makes the motives and demands more specific. De novo design utilizes generative architectures to explore uncharted chemical space, thus creating novel sequences with emergent antimicrobial mechanisms. This paradigm-shifting mode explores a vast space without explicit templating on existing sequences, yet requires rigorous validation of functionality and biosafety profiles. Although the distinction in protein design between redesign and de novo design is more nuanced than the dichotomy, since both of them incorporate functional motifs from natural sequence and structural elements to varying degrees in the training phase [204], de novo design demonstrates superior sequence diversity through enhanced exploration of potential sequence space. Together, these three paths accelerate discovery, enhance performance, and expand the therapeutic potential of AMPs.

Multimodality. The performance of AI models is typically enhanced by the quantity and quality of training data, as well as by the integration of multiple data modalities to deepen the model’s comprehension and interaction. Notin et al. concluded three core modalities for functional protein design, including sequences, structures, and functional labels [204]. Wan et al. further summarized five representational modalities based on them, including global descriptors, sequence-based representation, graph-based representation, 3D representation, and data-driven representation [117]. Each modality has its own strengths and limitations, contingent upon the specific tasks and data availability. For instance, global descriptors succinctly summarize peptide properties, which is particularly advantageous as it captures specific information relevant to the modeled properties. Compared with it, graph-based representations are better inputs for geometry-related tasks because they capture connectivity information. By integrating multiple data modalities, AI models enable a more comprehensive understanding and generation of information, leading to improving data complementarity, enriching semantic representation, reducing data dependency, and enhancing contextual understanding [214], [215], [216].

Multistage. The advent of diverse technological tools [120] and computational platforms [217] for peptide engineering has catalyzed a paradigm shift from singular-method approaches to sophisticated multistage design pipelines that synergistically integrate heterogeneous methodologies, thereby enabling the implementation of diversified design strategies with distinct objectives at sequential phases. For instance, target-constrained tertiary architectures generated via RFdiffusion [193] undergo inverse folding through ProteinMPNN [218] to probabilistically decode sequence ensembles, which are subsequently double-checked via ESMFold [74] to ensure sequence-structure congruence. Despite the potential of protein structure-function-guided design methods [219], [220], the inherent instability of AMPs poses a significant challenge due to their small size and flexible conformation [211], [221]. This instability impedes their practical application and necessitates the development of strategies to enhance their robustness. Another paradigmatic illustration lies in the synergistic collaboration between identification models and generative models. When generative models produce candidate sequences, there exists a critical need for subsequent multi-criteria discriminative screening through identification models. This iterative multistage cross-validation mechanism establishes a self-reinforcing feedback loop that significantly enhances model robustness by mitigating confirmation bias inherent in single-model approaches. Moreover, cutting-edge implementations (such as AMP-designer [140]) are increasingly integrating wet-lab experimental data with reinforcement learning frameworks to create closed-loop optimization systems. This feedback enables continuous model retraining, thereby progressively enhancing the physicochemical validity and biological relevance of generated sequences. Such hybrid computational–experimental pipelines have demonstrated particular efficacy in de novo protein design challenges where traditional purely in silico methods struggle with the combinatorial complexity of sequence-structure-function relationships. In summary, this multistage closed-loop paradigm enables continuous model refinement via retraining and multi-objective optimization, systematically improving robustness in identifying sequences that satisfy specific functional constraints while mitigating confirmation biases inherent in single-stage approaches.

6. Discussions

Numerous strategies have been developed to investigate novel AMPs. Bioassay-guided approaches, such as chromatography and fluorescence screening [222], demonstrate high precision yet are constrained by their time-intensive and costly characteristic. Empirical design, for example, enriching cationic or hydrophobic residues that are strongly associated with antimicrobial activity, holds merit but heavily relies on established knowledge, which results in bottlenecks in the scope of exploration. Due to these limitations, large-scale implementation remains impractical and ineffective. Traditional bioinformatics methods, such as BLAST (Basic Local Alignment Search Tool) [223], offer significant advantages in sequence comparison, potentially identifying AMPs. Despite its effectiveness, the flat filter style of this method restricts the diversity of its conclusions. In contrast, AI-based methods, capable of analyzing vast datasets with high accuracy and identifying complex patterns and trends, have effectively accelerated AMP design. However, this field is still in its early stages, and numerous challenges remain to be addressed and resolved. We summarize the challenges and prospects throughout our survey as follows.

6.1. Challenges

Vast peptide space, undesirable physicochemical properties, and unspecific mechanisms of action stall the discovery of novel AMPs. Meanwhile, the high cost of peptide synthesis as well as the generation of more industrial wastes hinder the implementation of translating AMPs as therapeutic drugs to treat infectious diseases [224]. Furthermore, although AMPs exhibit more favorable pharmacodynamics than conventional antibiotics in terms of preventing resistance evolution, bacteria can still evolve resistance to AMPs. For instance, Sinorhizobium meliloti has been investigated to evolve a peptidase that protects it from the harmful effects of AMPs [225]. These challenges underscore the intricate nature of AMP development and emphasize the need for more in-depth and sustained exploration. AI-guided methods have emerged as a promising pipeline to accelerate this process by efficiently identifying potential AMP candidates from large libraries. Nevertheless, artificial intelligence is not a panacea. The limitations in terms of identified and generative models for designing AMPs have been discussed in Sections 3.4, 4.4, respectively.

In addition to the aforementioned limitations, AI-guided technologies are constrained by inherent restrictions, termed impossibility theorems. These theorems are classified into five mechanism-based categories: deduction, indistinguishability, induction, tradeoffs, and intractability [226]. For instance, the renowned Gödel’s Incompleteness theorems [227] fall under deduction, while the well-known No Free Lunch (NFL) theorems [228], [229] are categorized into induction. These impossibility theorems provide crucial guidance and caution on building AI systems for AMP design. Take the induction impossibility theorems for an example, many AI-guided approaches, such as statistical learning and deep learning, rely on inductive reasoning, which leads to Goodman’s problem [230], underscoring that predictions of AMPs are not logically guaranteed. Specifically, one of the prediction biases stems from the implicit assumptions embedded in AI algorithms, such as the assumption of independent and identically distributed (i.i.d.) data. This assumption posits that data samples are drawn independently from the same probability distribution. However, in the scenario of AMP design, candidates deviate from this ideal condition, thereby potentially reducing the accuracy of model predictions. Furthermore, AI-guided methodologies tend to converge to sequence clusters that resemble those in the observed datasets and generally lack the capacity to explore entirely novel sequence clusters. This phenomenon accentuates the inherent limitations of AI in terms of de novo exploration and making groundbreaking discoveries, as AI systems predominantly rely on pre-existing data and established patterns. Therefore, it is essential to integrate multi-scale approaches and specialized knowledge to collectively drive its development. Subsequently, as novel data continuously accumulates, the quality driven by AI can be continuously improved, thereby deducing more valuable AMP candidates.

6.2. Prospects

The various challenges inherent in antimicrobial peptide design impel persistent exploration and resolution. While presenting formidable hurdles, they simultaneously act as catalysts for innovation and discovery, continually giving rise to numerous opportunities for advancement. Concurrently, as datasets continue to be enriched and expanded, models are iteratively refined, and knowledge bases are augmented, the design of novel AMPs will be significantly promoted. We summarize three aspects for future research in AI-accelerated AMP design, including finer-grained exploration, model-driven data enrichment, and model enhancement.

Finer-grained explorations. AMPs exhibit broad-spectrum antimicrobial activity. Nevertheless, emerging evidence has revealed previously unrecognized target-specific activity patterns [231], [232], [233]. Studies on these specificity mechanisms constitute a critical strategy for enhancing drug targeting capabilities [234], [235]. Systematically delineating the mechanistic interplay between AMPs and diverse biological targets may guide the development of more potent systemic applications. Similarly, investigating synergistic interactions among AMP combinations represents a potential approach to augment antimicrobial efficacy and mitigate resistance risk. Additional promising research avenues include strategic chemical modification for improving stability, systematic elucidation of fundamental biophysical properties, and conditional generation with finer control for precision design. Without granular mechanistic studies, unresolved barriers will consistently impede clinical advancement. For instance, the current restriction of AMPs to topical formulations rather than systemic administration represents a substantial obstacle to their implementation as viable pharmaceutical agents.

Model-driven data enrichment The paucity of high-resolution training datasets constrains AI-assisted AMP design. For instance, the scarcity of data regarding AMP synergies precludes the development of models capable of discerning novel cooperative relationships among AMPs. This limitation underscores the imperative for specialized data acquisition aligned with model-driven requirements, a trajectory that reflects the escalating importance of interdisciplinary collaboration. Biologists, in partnership with modeling specialists, can prioritize the collection of targeted data to enhance model functionality and robustness. For example, rather than relying exclusively on in vitro MIC testing of individual components, they can measure AMP synergisms both in vivo and in vitro. Although this collection demands significant resources and temporal investment, it represents a crucial foundation for sustaining high-quality research initiatives over the long term. Another example concerns chemical modification. Enriching training datasets with non-standard peptides instead of only 20 standard amino acids will significantly expand the potential design space and enhance model generalizability. Incorporating these tailored libraries into training datasets represents a promising avenue for future progress. In parallel, the integration of multi-scale data into unified modeling architectures is gaining prominence, exemplified by multimodal models. As a result, dataset enrichment now extends beyond increasing data volume (depth) to encompass a broader range of data types (breadth).

Model enhancement AI models function as intermediaries between observational data and predictive insights alongside continuous refinement and advancement. Two trends in model enhancement may catalyze the accelerated design of AMPs: scalable generalization and refined specialization. The scalable generalization benefits from the scaling laws [236], which delineate the quantitative relationships between model performance and factors such as model size, dataset size, and computational resources, typically observed as power-law relationships. Additionally, during the accumulation of these resources, a phenomenon termed “emergence” [237], [238] occurs, wherein new capabilities or characteristics suddenly manifest in a system following gradual and seemingly minor alterations. For instance, the integration of multi-scale data exemplifies how multimodal models excel in capturing heterogeneous data across scales, enabling comprehensive representation learning and cross-modal synergy. Similarly, the Mixtral of Experts (MoE) models [239], [240], [241], [242] enhance model robustness through dynamic expert routing for adaptive computation. These large-scale models offer a macroscopic and comprehensive perspective to guide the design of AMPs. In parallel, functional AI agent system [243], [244] conveniently serve end-to-end applications [245]. This complete pipeline will predict the results of novel AMPs from de novo design through all pre-clinical validations for reference, for instance, toxicity concerns, delivery systems, and immune response. The refined specialization refers to optimizing the model for specific tasks to enhance its learning and inference capabilities. For instance, fine-tuning [246], [247], [248], [249] can adapt a pre-trained model to a specific task or dataset by further training, thereby enabling efficient transfer learning and improved performance. Furthermore, purely data-driven models exhibit suboptimal performance when dealing with scarce or noisy data. By incorporating physical constraints, the robustness and interpretability of the model are enhanced. A prime example is IC-Light [250], which incorporates the physical principle of consistent light transmission. With this consistency, it enables uniformly handling various data sources and facilitating a physically grounded model behavior, hence guaranteeing accurate light modification while preserving intrinsic properties. These physics-informed models [251], [252], [253] improve the validity and reliability of predictions and exhibit proficiency in conducting finer-grained data analysis and learning for specific tasks. Similarly, in the foreseeable future, a proliferation of models incorporating biological principles is anticipated to emerge, for example, equipped with biophysics principles to provide the necessary inductive knowledge to extrapolate beyond the observable samples.

Artificial intelligence catalyzes antimicrobial peptide design. By combining computational efficiency with the biological environment and bridging computational innovation with practical therapeutic solutions, this transformative paradigm not only accelerates the design cycle but also expands the chemical and functional diversity of AMPs, representing a cornerstone for next-generation antimicrobial agents tailored to combat evolving pathogens with escalating antimicrobial resistance. However, despite the immense development potential and promising applications of AMPs, it is imperative to proceed with caution and avoid the pitfalls of overeager advancement, as Lazzaro et al. suggested in 2020 [154], “If AMPs are to be used clinically, it is crucial to understand their natural biology in order to lessen the risk of collateral harm and avoid the crisis of resistance now facing conventional antibiotics”.

CRediT authorship contribution statement

Yongqiang Liu: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Resources, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Jie Hu: Validation, Investigation, Data curation. Ning Zhang: Validation, Investigation, Data curation. Yinqi Bai: Writing – original draft, Validation, Formal analysis, Data curation. Yuan Yao: Writing – original draft, Validation, Funding acquisition. Gaoxiang Chen: Writing – review & editing, Writing – original draft, Validation, Project administration, Funding acquisition, Conceptualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

This research is supported by the ‘Pioneer’ and ‘Leading Goose’ R&D Program of Zhejiang (No. 2025C01114, No. 2025C01100, No. 2024SSYS0007, No. 2024C03004, and No. 2025C01102), Zhejiang Provincial Natural Science Foundation of China under Grant (No. LR25C100001), the National Natural Science Foundation of China General Program (No. 32471372). Also, we would like to thank the iBiofoundary and Core Facility of ZJU-Hangzhou Global Scientific and Technological Innovation Center.

Peer review under the responsibility of Editorial Board of Synthetic and Systems Biotechnology.

1

Negative charge in minority cases. For instance, dermcidin is negatively charged, but its activities can be enhanced by combination with zinc or highly cationic residues [108], [109], [110].

5

The other type, k-mer: every k amino acid as a group is regarded as a word. Sequences are divided from beginning to end. When the end of the sequence is less than k AAs, the remaining one forms a word.

2

A model is composed of multiple modules. For example, a convolutional neural network comprises convolutional modules, pooling modules, and fully connected modules, etc.

3

Liu et al. [155] proposed TransSAFP, a deep learning-guided model for the de novo design of antimicrobial materials, which integrates non-natural amino acids for enhanced peptide self-assembly and effectively predicts the functional activity of the self-assembling peptide materials.

4

Prior to the wet experiments, the biological filtering criteria were also conducted, which is used to approximate expert evaluations of peptide synthesizability [198].

Contributor Information

Yinqi Bai, Email: baiyinqi@genomics.cn.

Yuan Yao, Email: yyao1@zju.edu.cn.

Gaoxiang Chen, Email: cgx@zhejianglab.edu.cn.

Appendix A. Computational evaluation and filtration types of AI-driven AMP design

Between the candidate sequences generated from generative models and the subsequent execution of wet-lab validation lies a critical intermediate phase, computational evaluations and filtrations. Based on a systematic survey of the literature, beyond basic operations like deduplication, we consolidate three evaluation and filtration types, predictions for antimicrobial activity, novelty and diversity, and biological criteria, as follows.

Predictions for antimicrobial activity. Antimicrobial activity is the most important property of AMPs. Therefore, it is essential to evaluate whether candidate sequences possess this property, as follows.

  • •

    Identification models (classifiers) with statistical metrics [140], [182], [183], [185], [186], [197], [198], [200]. As the most frequently employed computational validation strategy, an identification model reassesses each generated candidate peptide. Sequences must satisfy a comprehensive suite of statistical criteria derived from the four cardinal indices of the confusion matrix, true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), and their derivatives, such as F1 score, Recall value, MCC (Matthews Correlation Coefficient); only those classified as antimicrobially active are permitted to proceed the next filtration.

  • •

    Molecular dynamics (MD) simulations [182], [198], [199], [200]. MD simulations are implemented using the predicted structures of the peptides in the presence of lipid membranes to test their binding mechanisms. Candidates with antimicrobial properties can align parallel to and infiltrate deeply within the membrane [28], [254]. For instance, GROMACS [255] can be employed to execute MD simulations: negative distances denote that the peptide remains solvated and distant from the membrane surface, consistent with a non-antimicrobial phenotype, whereas positive distances signify membrane contact or insertion, indicative of antimicrobial activity.

Novelty and diversity. The principal merit of generative models lies in their capacity to yield candidate peptides that depart from the extant antimicrobial repertoire. Consequently, the novelty and diversity of the generated sequences assume equal importance. Moreover, sequences warranting costly wet-lab scrutiny are typically those exhibiting pronounced divergence, since revalidating the activity of already-documented peptides is scarcely justified. Although all downstream metrics ultimately quantify the degree of divergence, the evaluation criteria adopted by state-of-the-art methodologies diverge in detail and remain heterogeneous. Below, we investigate the prevailing assessment strategies.

  • •

    Clustering [200]. Clustering candidate sequences against their parental set is a widely adopted strategy to quantitatively demarcate novelty and to safeguard sequence diversity. For instance, PepDiffusion implements clustering by using CD-HIT software [256] with a threshold of 0.6.

  • •

    Percentage of peptides that are different from the real AMPs [140].

  • •

    Difference between the generated peptides and the training AMP set. For instance, for each generated sequence, Wang et al. [184] searched the training dataset for a peptide that has the smallest Levenshtein distance [202] from it and normalized the distance according to its length. Then they calculated the average length as the novelty.

  • •

    Similarity among the generated peptides, measured by calculating the Levenshtein distance between every two sequences and normalizing it by the sequence length [184], [198].

  • •

    Relative similarity [176], [186] of two sets of sequences through the global alignment algorithm [203].

  • •

    Language model perplexity. Das et al. proposed PepCVAE [183] and used a character-level LSTM language model (LM) to estimate the perplexity (PPL) of the generated sequences.

  • •

    Relative n-gram entropy [183] estimated by mixing generated samples with original samples at a 1:1 ratio

  • •

    Number of shared n-grams [257] for different values of n between the generated sequences and the original ones [183].

  • •

    Uniqueness [140], [184]. Uniqueness denotes the percentage of unique peptides generated.

  • •

    Physicochemical characteristics [140], [172], [198]. Physicochemical attributes of the generated candidates, such as charge and amphiphilicity, can be predicted by bioinformatic tools, e.g., modlAMP [122]. Discrepancies between the physicochemical property distributions of the generated sequences and those of extant AMPs likewise serve as indicators of novelty and diversity.

Biological criteria.

  • •

    Biosynthesis criteria [198], [200]. To approximate expert assessments of peptide synthesizability, it is better to implement the biosynthesis criteria. For instance, sequences are excluded if (i) any pentapeptide window contained more than three cationic residues, (ii) three or more consecutive hydrophobic residues (Ala, Val, Leu, Ile, Phe, Trp, Met) or three identical residues in succession were present, or (iii) Cys is detected.

  • •

    Virtual screening for biological phenotype [140], [182], [199]. Some wet-lab results can be predicted through in-silico filters, affording an additional layer of pre-experimental filtration. For instance, the hemolysis score is predicted by HemoPI [258], and the toxicity score is predicted by ToxinPred [259], [260].

Appendix B. AMP feature extraction strategies

We surveyed the feature-extraction strategies that underpin current AI-driven AMP design and adopted the five core paradigms previously systematized by Notin et al. [204], as follows.

Global descriptors

  • •

    Amino acid composition (AAC) [261], [262], [263], [264], [265]. AAC represents the proportion of each amino acid in a peptide sequence.

  • •

    Pseudo-amino acid composition (PAAC) [261], [262], [263], [265], [266]. The feature set of PAAC comprises over 20 attributes, incorporating various physicochemical properties such as hydrophobicity values, hydrophilicity, side chain masses, and sequence-order information [267].

  • •

    Composition, transition and distribution descriptor (CTDD) [261], [262], [263], [264], [265], [268]. These features consist of 13 different physicochemical properties, including hydrophobicity, normalized van der Waals volume, polarity, polarizability, charge, secondary structures, and solvent accessibility, which can be extracted by a Python package iFeature [269].

  • •

    Grouped amino acid composition (GAAC) [261], [262]. The GAAC divides the 20 AAs into five groups, which consist of the aliphatic group, aromatic group, positive charge group, negative charge group, and uncharged group.

  • •

    Dipeptide composition (DPC) [262], [264], [265], [266]. DPC aims to retrieve information about AAs using regional amino acid ordering by counting how frequently any two adjacent residues appear. Thus, the feature dimension is 400 [270].

  • •

    Tripeptide composition (TPC) [262]. Similar to DPC but for three adjacent residues [271].

  • •

    K-mer [268]. K-mer represents the occurrence frequency of all peptide sub-sequence with a length k.

  • •

    Distance-Pairs (DP) [268]. It is effective to represent the composition and position information by introducing a concept called “distance-pair”.

  • •

    Distance-based Top-n-gram approach (DT) [268]. It integrates the evolutionary information and the sequence information into a single vector.

  • •

    Composition of k-spaced amino acid pairs (CKSAAP) [262], [265]. The CKSAAP proposed to extract protein structure-related features [272].

  • •

    Conjoint triad (CT) [262]. Conjoint triad considered the properties of one AA and its vicinal AAs and regarded any three continuous AAs as a unit [273].

  • •

    ProtDCal [274], [275]. ProtDCal is a software that can provide 45,494 features, based on features from individual residues such as hydrophobicity and electronic charge [276].

Sequence-based representations

  • •

    Sequence (one-hot) [200], [277], [278], [279]: Each of the 20 AAs is initially encoded as a unique one-hot vector. A zero-padding approach is employed to ensure that all peptide sequences are aligned to a uniform length. As a result, a sequence S of length L turns into an L×20 matrix.

  • •

    Sequence (tokenization) [111], [170], [280], [281], [282], [283], [284]: The sequences comprise AAs represented by official IUPAC single-letter symbols.5 Additionally, two special class tokens are inserted at the beginning and end of the sequence to denote it as a sentence, respectively. In cases where the sequence length is less than the fixed maximum length, padding ([PAD]) tokens are used to fill the empty spaces, ensuring that all sequences have a consistent feature dimension. These standardized sequences are then embedded into X-dimensional vectors to serve as input features.

  • •

    BLOSUM62 [278], [279]. The BLOSUM62 matrix [285] aims to assess the similarity of protein sequences, which is a widely employed method for indicating protein evolutionary information. To encode a given peptide sequence, each residue is represented as a vector possessing 20 dimensions. As a result, a sequence S of length L turns into an L×20 matrix.

  • •

    Z-Scale [278]. The Z-Scale matrix [286] aims to represent amino acid physicochemical properties by characterizing each amino acid with five numerical values, including hydrophobicity and hydrophilicity, steric bulk properties and polarizability, polarity, and electronic effects. Therefore, a peptide S of length L can be encoded as an L×5 matrix.

  • •

    AAIndex [261], [264], [279]. AAindex is a comprehensive database that encompasses 531 physicochemical properties. Each of the 20 AAs is assigned a specific numerical value corresponding to each property within the AAindex database [287].

  • •

    Position-specific scoring matrix (PSSM) [266]. PSSM [288] is employed in biological sequences to express patterns, represented in a two-dimensional matrix, where the number of rows is determined by the length of the peptide sequence, while the number of columns is determined by the size of the alphabet (20).

Graph-based representations

Table C.1.

AMPs’ Databasesa.

Database Description Last updatedb URL Reference(s)
CyBase Cyclotides. Database of proteins that possess a cyclic
backbone in which the N and C termini have been
joined.
2025.04 https://www.cybase.org.au/ [114], [115]

APD3 Including natural AMPs from the six life kingdoms,
synthetic, and predicted AMPs with 25 activities.
2025.01 https://aps.unmc.edu/ [157], [289], [290]

dbAMP Including predicted AMPs and natural AMPs with 23
activities.
2024.11 https://awi.cuhk.edu.cn/dbAMP/index.php [291], [292]

DRAMP Including sequences, structures, activities,
physicochemical, patent, clinical, and reference.
2024.11 http://dramp.cpu-bioinfor.org/ [293], [294], [295], [296]

AMPSphere A collection of AMPs from the global microbiome,
containing 863,498 distinct sequences, clustered into
10,715 AMP families with at least 7 sequences each.
N/A
(2024.07)
https://ampsphere.big-data-biology.org/about [171]

MBPDB Includes all known bioactive peptides derived from
milk proteins from any species.
2024.06 https://mbpdb.nws.oregonstate.edu/about_us/ [297], [298]

AMPDB A feature-rich knowledge base on AMPs that assimilate
information from 30 databases with 88 functional
activities.
2023.04 https://bblserver.org.in/ampdb/ [299]

CAMPR4 Including predicted AMPs and natural AMPs on
sequence, structure, activity, source organism, target
organisms, protein family descriptions, modifications,
and links.
N/A
(2023.01)
https://camp.bicnirrh.res.in/index.php [300], [301], [302], [303], [304]

MEGARes Hand-curated resistance genes for antimicrobial
drugs, biocides and metals, incorporating 4 data
sources.
2022.09 https://www.meglab.org/megares/ [305], [306], [307]

B-AMP A structural and functional repository consisting of
5766 AMPs from natural and synthetic sources.
2022.08 https://b-amp.karishmakaushiklab.com/index.html [308], [309]

DBAASP Including ribosomal, non-ribosomal, and
synthetic peptides that show antimicrobial activity as
Monomers, Multimers, and Multi-Peptides.
N/A
(2021.01)
https://dbaasp.org/home [310], [311], [312]

AHTPDB Database of experimentally validated Antihypertensive
peptides, containing 5978 entries of about 1694
unique peptides.
N/A
(2015.01)
https://webs.iiitd.edu.in/raghava/ahtpdb/ [313]

CancerPPD Database of experimentally verified anticancer
peptides and proteins, containing 3491 entries.
N/A
(2015.01)
https://webs.iiitd.edu.in/raghava/cancerppd/ [314]

starPepDB An integrated graph database holding bioactive
peptides and metadata from a large variety of biological
data sources.
2014.12 https://mobiosd-hub.com/starpep/ [315], [316]

ParaPep Database of experimentally validated anti-parasitic
peptides, containing 863 anti-parasite peptide entries,
which have been tested against 12 different types of
parasites.
N/A
(2014.06)
https://webs.iiitd.edu.in/raghava/parapep/ [317]

YADAMP Particular attention on short alpha helix peptides
interacting with cell membranes.
2013.03 http://yadamp.unisa.it/ [318]
a

The date of the latest update and verification is April 17, 2025.

b

The content within the parentheses indicates the publication date of the latest literature.

  • •

    The graph-based representations are better for capturing geometry-related features, consisting of nodes and edges. Nodes can be sequences, atoms, or residues, and edges can be chemical bonds, ingredient attribution, or the geometric distance between the nodes. For instance, LABAMPsGCN [137] is a graph convolutional neural network. Its nodes consist of k-mer residues and sequences, and its edges are determined by affiliation relationships.

Three-dimensional representations

  • •

    Three-dimensional coordinates [206]. The three-dimensional coordinates provide information about the distances between amino acid residues, reflecting their spatial interactions. The positions of the Cα atoms in all standard AAs are selected to represent the residue-level coordinates.

Data-driven representations

  • •

    Embeddings (LLM or pLM) [209], [282]: By utilizing a Large Language Model (LLM) or pre-trained protein Language Model (pLM), e.g., BERT or ESM2, the embedding features for each peptide sequence extracted from the language model can be transferred to any task requiring a numerical representation.

  • •

    Predicted protein secondary structure (PPSS) [274]. Protein secondary structure refers to the local conformation of the polypeptide backbone of proteins. For instance, the secondary structure feature of all peptides predicted from DeepCNF [319] has eight types, including 3 types for helix, 2 types for strand, and 3 types for coil.

  • •

    Predicted surface-level features [206]. The geometric and chemical features for each vertex on the protein surface can be extracted from the three-dimensional structures of peptide sequences predicted by AlphaFold2 [72]. Each vertex is characterized by five geometric or physicochemical properties [320].

  • •

    Cellular automata image (CAI) [144]. By generating image representations for biological sequences, the cellular automata image approach could find some implicit sequence features [321], [322].

Appendix C. Databases and models

Public AMP databases are summarized in Table C.1, whereas the snapshot of AI methods for designing AMPs is presented in Table C.2.

Table C.2.

AI-guided methods for designing AMPsa.

Reference Year AI-module(s) Task Wet-lab Web Server Open-source codes
EvoGradient [277] 2025 CNN, LSTM, ATT, and TF Classification Y N Y
Shen et al. [280] 2025 CNN, LSTM, and BERT Classification Y N N
TranSAFP [155] 2025 CNN, LSTM, ATT, and GRU Classification Y N N
AMP-Designer [140] 2025 TF, ATT, and RNN Generation Y N Y
Zhang et al. [174] 2025 GA and KNN Generation Y N N
Gao et al. [201] 2025 ESM and GA Generation Y N Y
PepDiffusion [200] 2025 TF, VAE, and Diffusion Generation Y N Y
Chung et al. [261] 2024 CNN, BiLSTM, RF, and TF Classification N N Y
dmSLAY [209] 2024 RF, SVM, and DT Classification Y N Y
iAMP-Attenpred [282] 2024 CNN, BERT, BiLSTM, and ATT Classification N N Y
AGRAMP [323] 2024 RF Classification Y Y N
Diff-AMP [197] 2024 Diffusion and CNN Generation N Y Y
AMP-Diffusion [195] 2024 Diffusion Generation N N N
HydrAMP [198] 2023 VAE Generation Y Y Y
Teimouri et al. [324] 2023 SVM and LR Classification N N Y
AMP-BERT [284] 2023 BERT Classification N N Y
Cao et al. [187] 2023 GAN and BERT Generation Y N Y
Huang et al. [170] 2023 XGBoost, CNN, RF, and LSTM Classification Y N Y
DeepAFP [278] 2023 BERT, CNN, and BiLSTM Classification N N Y
iAMPCN [279] 2023 CNN Classification N N Y
MBC-Attention [262] 2023 CNN and ATT Classification N N Y
Ma et al. [111] 2022 ATT, LSTM, and BERT Classification Y N Y
sAMP-PFPDeep [325] 2022 CNN Classification N N Y
LSSAMP [184] 2022 VAE Generation Y N Y
AMPpred-EL [268] 2022 LightGBM and LR Classification N N N
Sun et al. [137] 2022 GCN Classification N Y (error: 404) N
AMPlify [145] 2022 BiLSTM and ATT Classification Y N Y
MLBP [281] 2022 CNN and RNN Classification N N Y
Das et al. [182] 2021 LSTM, WAE, and CLaSS Generation Y N Y
Capecchi et al. [179] 2021 RNN Generation Y N Y
PepVAE [181] 2021 VAE Generation Y N N
Wang et al. [178] 2021 LSTM and BiLSTM Generation N N N
AMPGAN v2 [186] 2021 BiCGAN Generation N N Y
iAMP-CA2L [144] 2021 CNN, BiLSTM, and SVM Classification N N Y
Zhang et al. [283] 2021 BERT Classification N N Y
a

Abbreviations: Convolutional Neural Network (CNN) [326], Bi-directional Long Short-Term Memory (BiLSTM) [327], Random Forest (RF) [328], Transformer (TF) [329], Attension (ATT) [330], Long Short-Term Memory (LSTM) [331], Gated Recurrent Unit (GRU) [332], ESMFold (ESM) [74], Bidirectional Encoder Representations from Transformers (BERT) [333], Support vector machine (SVM), Decision Tree (DT) [334], Genetic Algorithm (GA) [335], Fuzzy k-nearest Neighbor (fuzzy k-NN) [336], Variational Autoencoder (VAE) [337], Logistic Regression (LR) [338], Generative Adversarial Network (GAN) [339], [340], EXtreme Gradient Boosting (XGBoost) [341], Light Gradient-Boosting Machine (LightGBM) [342], Graph Convolutional Network (GCN) [343], Wasserstein Autoencoder (WAE) [344], Recurrent Neural Network (RNN) [345], Bidirectional Conditional Generative adversarial Network (BiCGAN) [346].

References

  • 1.Luther A., Urfer M., Zahn M., Müller M., Wang S.-Y., Mondal M., Vitale A., Hartmann J.-B., Sharpe T., Monte F.L., et al. Chimeric peptidomimetic antibiotics against gram-negative bacteria. Nature. 2019;576(7787):452–458. doi: 10.1038/s41586-019-1665-6. [DOI] [PubMed] [Google Scholar]
  • 2.Wong F., Zheng E.J., Valeri J.A., Donghia N.M., Anahtar M.N., Omori S., Li A., Cubillos-Ruiz A., Krishnan A., Jin W., et al. Discovery of a structural class of antibiotics with explainable deep learning. Nature. 2024;626(7997):177–185. doi: 10.1038/s41586-023-06887-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Czaplewski L., Bax R., Clokie M., Dawson M., Fairhead H., Fischetti V.A., Foster S., Gilmore B.F., Hancock R.E., Harper D., et al. Alternatives to antibiotics—a pipeline portfolio review. Lancet Infect Dis. 2016;16(2):239–251. doi: 10.1016/S1473-3099(15)00466-1. [DOI] [PubMed] [Google Scholar]
  • 4.Forsberg K.J., Reyes A., Wang B., Selleck E.M., Sommer M.O., Dantas G. The shared antibiotic resistome of soil bacteria and human pathogens. Science. 2012;337(6098):1107–1111. doi: 10.1126/science.1220761. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Zheng E.J., Valeri J.A., Andrews I.W., Krishnan A., Bandyopadhyay P., Herneisen A., Schulte F., Linnehan B., Wong F., Stokes J.M., et al. Discovery of antibiotics that selectively kill metabolically dormant bacteria. Cell Chem Biol. 2024;31(4):712–728. doi: 10.1016/j.chembiol.2023.10.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Liu G., Catacutan D.B., Rathod K., Swanson K., Jin W., Mohammed J.C., Chiappino-Pepe A., Syed S.A., Fragis M., Rachwalski K., et al. Deep learning-guided discovery of an antibiotic targeting acinetobacter baumannii. Nat Chem Biol. 2023;19(11):1342–1350. doi: 10.1038/s41589-023-01349-8. [DOI] [PubMed] [Google Scholar]
  • 7.Melo M.C., Maasch J.R., de la Fuente-Nunez C. Accelerating antibiotic discovery through artificial intelligence. Commun Biol. 2021;4(1):1050. doi: 10.1038/s42003-021-02586-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Holmes A.H., Moore L.S., Sundsfjord A., Steinbakk M., Regmi S., Karkey A., Guerin P.J., Piddock L.J. Understanding the mechanisms and drivers of antimicrobial resistance. Lancet. 2016;387(10014):176–187. doi: 10.1016/S0140-6736(15)00473-0. [DOI] [PubMed] [Google Scholar]
  • 9.Magana M., Sereti C., Ioannidis A., Mitchell C.A., Ball A.R., Magiorkinis E., Chatzipanagiotou S., Hamblin M.R., Hadjifrangiskou M., Tegos G.P. Options and limitations in clinical investigation of bacterial biofilms. Clin Microbiol Rev. 2018;31(3):10–1128. doi: 10.1128/CMR.00084-16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Shan Y., Brown Gandt A., Rowe S.E., Deisinger J.P., Conlon B.P., Lewis K. ATP-dependent persister formation in Escherichia coli. MBio. 2017;8(1):10–1128. doi: 10.1128/mBio.02267-16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lewis K., Shan Y. Persister awakening. Mol Cell. 2016;63(1):3–4. doi: 10.1016/j.molcel.2016.06.025. [DOI] [PubMed] [Google Scholar]
  • 12.Conlon B.P., Rowe S.E., Gandt A.B., Nuxoll A.S., Donegan N.P., Zalis E.A., Clair G., Adkins J.N., Cheung A.L., Lewis K. Persister formation in staphylococcus aureus is associated with ATP depletion. Nat Microbiol. 2016;1(5):1–7. doi: 10.1038/nmicrobiol.2016.51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Lewis K. Persister cells. Annu Rev Microbiol. 2010;64(1):357–372. doi: 10.1146/annurev.micro.112408.134306. [DOI] [PubMed] [Google Scholar]
  • 14.Mulukutla A., Shreshtha R., Deb V.K., Chatterjee P., Jain U., Chauhan N. Recent advances in antimicrobial peptide-based therapy. Bioorg Chem. 2024 doi: 10.1016/j.bioorg.2024.107151. [DOI] [PubMed] [Google Scholar]
  • 15.Lachat J., Lextrait G., Jouan R., Boukherissa A., Yokota A., Jang S., Ishigami K., Futahashi R., Cossard R., Naquin D., et al. Hundreds of antimicrobial peptides create a selective barrier for insect gut symbionts. Proc Natl Acad Sci. 2024;121(25) doi: 10.1073/pnas.2401802121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Laxminarayan R., Duse A., Wattal C., Zaidi A.K., Wertheim H.F., Sumpradit N., Vlieghe E., Hara G.L., Gould I.M., Goossens H., et al. Antibiotic resistance—the need for global solutions. Lancet Infect Dis. 2013;13(12):1057–1098. doi: 10.1016/S1473-3099(13)70318-9. [DOI] [PubMed] [Google Scholar]
  • 17.de Kraker M.E., Stewardson A.J., Harbarth S. Will 10 million people die a year due to antimicrobial resistance by 2050? PLoS Med. 2016;13(11) doi: 10.1371/journal.pmed.1002184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Goldberg K., Lobov A., Antonello P., Shmueli M.D., Yakir I., Weizman T., Ulman A., Sheban D., Laser E., Kramer M.P., et al. Cell-autonomous icell-autonomousnnate immunity by proteasome-derived defence peptides. Nature. 2025:1–10. doi: 10.1038/s41586-025-08615-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Torres M.D., Melo M.C., Flowers L., Crescenzi O., Notomista E., de la Fuente-Nunez C. Mining for encrypted peptide antibiotics in the human proteome. Nat Biomed Eng. 2022;6(1):67–75. doi: 10.1038/s41551-021-00801-1. [DOI] [PubMed] [Google Scholar]
  • 20.Ferrazzano L., Catani M., Cavazzini A., Martelli G., Corbisiero D., Cantelmi P., Fantoni T., Mattellone A., De Luca C., Felletti S., et al. Sustainability in peptide chemistry: current synthesis and purification technologies and future challenges. Green Chem. 2022;24(3):975–1020. [Google Scholar]
  • 21.Chaudhary S., Mahfouz M.M. Molecular farming of antimicrobial peptides. Nat Rev Bioeng. 2024;2(1):3–5. [Google Scholar]
  • 22.Silva O.N., Torres M.D., Cao J., Alves E.S., Rodrigues L.V., Resende J.M., Lião L.M., Porto W.F., Fensterseifer I.C., Lu T.K., et al. Repurposing a peptide toxin from wasp venom into antiinfectives with dual antimicrobial and immunomodulatory properties. Proc Natl Acad Sci U S A. 2020;117(43):26936–26945. doi: 10.1073/pnas.2012379117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Zasloff M. Antimicrobial peptides of multicellular organisms. Nature. 2002;415(6870):389–395. doi: 10.1038/415389a. [DOI] [PubMed] [Google Scholar]
  • 24.Ji S., An F., Zhang T., Lou M., Guo J., Liu K., Zhu Y., Wu J., Wu R. Antimicrobial peptides: An alternative to traditional antibiotics. Eur J Med Chem. 2023 doi: 10.1016/j.ejmech.2023.116072. [DOI] [PubMed] [Google Scholar]
  • 25.Cullen T., Schofield W., Barry N., Putnam E., Rundell E., Trent M., Degnan P., Booth C., Yu H., Goodman A. Antimicrobial peptide resistance mediates resilience of prominent gut commensals during inflammation. Science. 2015;347(6218):170–175. doi: 10.1126/science.1260580. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Mookherjee N., Hancock R. Cationic host defence peptides: innate immune regulatory peptides as a novel approach for treating infections. Cell Mol Life Sci. 2007;64:922–933. doi: 10.1007/s00018-007-6475-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Magana M., Pushpanathan M., Santos A.L., Leanse L., Fernandez M., Ioannidis A., Giulianotti M.A., Apidianakis Y., Bradfute S., Ferguson A.L., et al. The value of antimicrobial peptides in the age of resistance. Lancet Infect Dis. 2020;20(9):e216–e230. doi: 10.1016/S1473-3099(20)30327-3. [DOI] [PubMed] [Google Scholar]
  • 28.Mookherjee N., Anderson M.A., Haagsman H.P., Davidson D.J. Antimicrobial host defence peptides: functions and clinical potential. Nat Rev Drug Discov. 2020;19(5):311–332. doi: 10.1038/s41573-019-0058-8. [DOI] [PubMed] [Google Scholar]
  • 29.Xu J., Li F., Leier A., Xiang D., Shen H.-H., Marquez Lago T.T., Li J., Yu D.-J., Song J. Comprehensive assessment of machine learning-based methods for predicting antimicrobial peptides. Brief Bioinform. 2021;22(5):bbab083. doi: 10.1093/bib/bbab083. [DOI] [PubMed] [Google Scholar]
  • 30.Torres Salazar B.O., Dema T., Schilling N.A., Janek D., Bornikoel J., Berscheid A., Elsherbini A.M., Krauss S., Jaag S.J., Lämmerhofer M., et al. Commensal production of a broad-spectrum and short-lived antimicrobial peptide polyene eliminates nasal staphylococcus aureus. Nat Microbiol. 2024;9(1):200–213. doi: 10.1038/s41564-023-01544-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Murray C.J., Ikuta K.S., Sharara F., Swetschinski L., Aguilar G.R., Gray A., Han C., Bisignano C., Rao P., Wool E., et al. Global burden of bacterial antimicrobial resistance in 2019: a systematic analysis. Lancet. 2022;399(10325):629–655. doi: 10.1016/S0140-6736(21)02724-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Alencar-Silva T., Díaz-Martín R.D., Sousa dos Santos M., Saraiva R.V.P., Leite M.L., de Oliveira Rodrigues M.T., Pogue R., Andrade R., Falconi Costa F., Brito N., et al. Screening of the skin-regenerative potential of antimicrobial peptides: Clavanin a, clavanin-MO, and mastoparan-MO. Int J Mol Sci. 2024;25(13):6851. doi: 10.3390/ijms25136851. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.da Silva A.M.B., Silva-Goncalves L.C., Oliveira F.A., Arcisio-Miranda M. Pro-necrotic activity of cationic mastoparan peptides in human glioblastoma multiforme cells via membranolytic action. Mol Neurobiol. 2018;55:5490–5504. doi: 10.1007/s12035-017-0782-1. [DOI] [PubMed] [Google Scholar]
  • 34.Alencar-Silva T., Braga M.C., Santana G.O.S., Saldanha-Araujo F., Pogue R., Dias S.C., Franco O.L., Carvalho J.L. Breaking the frontiers of cosmetology with antimicrobial peptides. Biotech Adv. 2018;36(8):2019–2031. doi: 10.1016/j.biotechadv.2018.08.005. [DOI] [PubMed] [Google Scholar]
  • 35.Wang G., Wang W., Chen Z., Hu T., Tu L., Wang X., Hu W., Li S., Wang Z. Photothermal microneedle patch loaded with antimicrobial peptide/MnO2 hybrid nanoparticles for chronic wound healing. Chem Eng J. 2024;482 [Google Scholar]
  • 36.Wang J., Cheng Y. Enhancing aquaculture disease resistance: Antimicrobial peptides and gene editing. Rev Aquac. 2024;16(1):433–451. [Google Scholar]
  • 37.Xia J., Ge C., Yao H. Antimicrobial peptides: An alternative to antibiotic for mitigating the risks of antibiotic resistance in aquaculture. Environ Res. 2024;251 doi: 10.1016/j.envres.2024.118619. [DOI] [PubMed] [Google Scholar]
  • 38.Huo X., Wang P., Zhao F., Liu Q., Yang C., Zhang Y., Su J. High-efficiency expression of a novel antimicrobial peptide I20 with superior bactericidal ability and biocompatibility in pichia pastoris and its efficiency enhancement to aquaculture. Aquaculture. 2024;579 [Google Scholar]
  • 39.Satchanska G., Davidova S., Gergova A. Diversity and mechanisms of action of plant, animal, and human antimicrobial peptides. Antibiotics. 2024;13(3):202. doi: 10.3390/antibiotics13030202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Si S., Huang X., Wang Q., Manickam S., Zhao D., Liu Y. Enhancing refrigerated chicken breasts preservation: Novel composite hydrogels incorporated with antimicrobial peptides, bacterial cellulose, and polyvinyl alcohol. Int J Biiol Macromol. 2024;281 doi: 10.1016/j.ijbiomac.2024.136505. [DOI] [PubMed] [Google Scholar]
  • 41.Choi S., Ingale S., Kim J., Park Y., Kwon I., Chae B. An antimicrobial peptide-A3: effects on growth performance, nutrient retention, intestinal and faecal microflora and intestinal morphology of broilers. Br Poult Sci. 2013;54(6):738–746. doi: 10.1080/00071668.2013.838746. [DOI] [PubMed] [Google Scholar]
  • 42.Liu Q., Yao S., Chen Y., Gao S., Yang Y., Deng J., Ren Z., Shen L., Cui H., Hu Y., et al. Use of antimicrobial peptides as a feed additive for juvenile goats. Sci Rep. 2017;7(1):12254. doi: 10.1038/s41598-017-12394-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Montesinos E., Bardaji E. Synthetic antimicrobial peptides as agricultural pesticides for plant-disease control. Chem Biodivers. 2008;5(7):1225–1237. doi: 10.1002/cbdv.200890111. [DOI] [PubMed] [Google Scholar]
  • 44.Dong J., Chen F., Yao Y., Wu C., Ye S., Ma Z., Yuan H., Shao D., Wang L., Wang Y. Bioactive mesoporous silica nanoparticle-functionalized titanium implants with controllable antimicrobial peptide release potentiate the regulation of inflammation and osseointegration. Biomaterials. 2024;305 doi: 10.1016/j.biomaterials.2023.122465. [DOI] [PubMed] [Google Scholar]
  • 45.Uygun Z.O., Tasoglu S. Impedimetric antimicrobial peptide biosensor for the detection of human immunodeficiency virus envelope protein gp120. iScience. 2024;27(3) doi: 10.1016/j.isci.2024.109190. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Alves P.M., Barrias C.C., Gomes P., Martins M.C.L. How can biomaterial-conjugated antimicrobial peptides fight bacteria and be protected from degradation? Acta Biomater. 2024;181:98–116. doi: 10.1016/j.actbio.2024.04.043. [DOI] [PubMed] [Google Scholar]
  • 47.Caselli L., Traini T., Micciulla S., Sebastiani F., Köhler S., Nielsen E.M., Diedrichsen R.G., Skoda M.W., Malmsten M. Antimicrobial peptide coating of TiO2 nanoparticles for boosted antimicrobial effects. Adv Funct Mater. 2024 [Google Scholar]
  • 48.Fonseca D.R., Alves P.M., Neto E., Custódio B., Guimarães S., Moura D., Annis F., Martins M., Gomes A., Teixeira C., et al. One-pot microfluidics to engineer chitosan nanoparticles conjugated with antimicrobial peptides using “photoclick” chemistry: Validation using the gastric bacterium helicobacter pylori. ACS Appl Mater Interfaces. 2024;16(12):14533–14547. doi: 10.1021/acsami.3c18772. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Cao Z., Shi Z., Tong M., Yang D., Liu L. Synergistic antimicrobial mechanism of the ultrashort antimicrobial peptide R3W4V with a tadpole-like conformation. J Chem Inf Model. 2024;64(17):6838–6849. doi: 10.1021/acs.jcim.4c01100. [DOI] [PubMed] [Google Scholar]
  • 50.Yu G., Baeder D.Y., Regoes R.R., Rolff J. Combination effects of antimicrobial peptides. Antimicrob Agents Chemother. 2016;60(3):1717–1724. doi: 10.1128/AAC.02434-15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Rahnamaeian M., Cytryńska M., Zdybicka-Barabas A., Dobslaff K., Wiesner J., Twyman R.M., Zuchner T., Sadd B.M., Regoes R.R., Schmid-Hempel P., et al. Insect antimicrobial peptides show potentiating functional interactions against gram-negative bacteria. Proc R Soc B Biol Sci. 2015;282(1806) doi: 10.1098/rspb.2015.0293. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Westerhoff H.V., Zasloff M., Rosner J.L., Hendler R.W., DE Waal A., Gomes A.V., Jongsma A.P., Riethorst A., Juretić D. Functional synergism of the magainins PGLa and magainin-2 in escherichia coli, tumor cells and liposomes. Eur J Biochem. 1995;228(2):257–264. [PubMed] [Google Scholar]
  • 53.Yan H., Hancock R.E. Synergistic interactions between mammalian antimicrobial defense peptides. Antimicrob Agents Chemother. 2001;45(5):1558–1560. doi: 10.1128/AAC.45.5.1558-1560.2001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Taheri-Araghi S. Synergistic action of antimicrobial peptides and antibiotics: current understanding and future directions. Front Microbiol. 2024;15 doi: 10.3389/fmicb.2024.1390765. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Porto W., Pires A., Franco O. Computational tools for exploring sequence databases as a resource for antimicrobial peptides. Biotech Adv. 2017;35(3):337–349. doi: 10.1016/j.biotechadv.2017.02.001. [DOI] [PubMed] [Google Scholar]
  • 56.Krizhevsky A, Sutskever I, Hinton GE. ImageNet Classification with Deep Convolutional Neural Networks. In: Conference on neural information processing systems 2012. neurIPS. 2012, p. 1106–14.
  • 57.Krizhevsky A., Sutskever I., Hinton G.E. ImageNet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84–90. [Google Scholar]
  • 58.Silver D., Huang A., Maddison C.J., Guez A., Sifre L., Van Den Driessche G., Schrittwieser J., Antonoglou I., Panneershelvam V., Lanctot M., et al. Mastering the game of go with deep neural networks and tree search. Nature. 2016;529(7587):484–489. doi: 10.1038/nature16961. [DOI] [PubMed] [Google Scholar]
  • 59.He K., Zhang X., Ren S., Sun J. Conference on computer vision and pattern recognition, CVPR. IEEE Computer Society; 2016. Deep residual learning for image recognition. [Google Scholar]
  • 60.Zoph B., Le Q.V. Conference on learning representations, ICLR. OpenReview.net; 2017. Neural architecture search with reinforcement learning. [Google Scholar]
  • 61.Finn C., Abbeel P., Levine S. vol. 70. PMLR; 2017. Model-agnostic meta-learning for fast adaptation of deep networks; pp. 1126–1135. (International conference on machine learning, ICML). [Google Scholar]
  • 62.Lake B.M., Salakhutdinov R., Tenenbaum J.B. Human-level concept learning through probabilistic program induction. Science. 2015;350(6266):1332–1338. doi: 10.1126/science.aab3050. [DOI] [PubMed] [Google Scholar]
  • 63.Radford A., Kim J.W., Hallacy C., Ramesh A., Goh G., Agarwal S., Sastry G., Askell A., Mishkin P., Clark J., Krueger G., Sutskever I. International conference on machine learning, ICML. vol. 139. PMLR; 2021. Learning transferable visual models from natural language supervision; pp. 8748–8763. (Proceedings of machine learning research). [Google Scholar]
  • 64.Radford A, Metz L, Chintala S. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In: Conference on learning representations, ICLR. 2016.
  • 65.Kim H.K., Min S., Song M., Jung S., Choi J.W., Kim Y., Lee S., Yoon S., Kim H. Deep learning improves prediction of CRISPR–Cpf1 guide RNA activity. Nature Biotechnol. 2018;36(3):239–241. doi: 10.1038/nbt.4061. [DOI] [PubMed] [Google Scholar]
  • 66.Zhou J., Troyanskaya O.G. Predicting effects of noncoding variants with deep learning–based sequence model. Nature Methods. 2015;12(10):931–934. doi: 10.1038/nmeth.3547. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Lopez R., Regier J., Cole M.B., Jordan M.I., Yosef N. Deep generative modeling for single-cell transcriptomics. Nature Methods. 2018;15(12):1053–1058. doi: 10.1038/s41592-018-0229-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Alipanahi B., Delong A., Weirauch M.T., Frey B.J. Predicting the sequence specificities of DNA-and RNA-binding proteins by deep learning. Nature Biotechnol. 2015;33(8):831–838. doi: 10.1038/nbt.3300. [DOI] [PubMed] [Google Scholar]
  • 69.Wu R., Ding F., Wang R., Shen R., Zhang X., Luo S., Su C., Wu Z., Xie Q., Berger B., et al. 2022. High-resolution de novo structure prediction from primary sequence. 2022–07. [Google Scholar]
  • 70.Baek M., DiMaio F., Anishchenko I., Dauparas J., Ovchinnikov S., Lee G.R., Wang J., Cong Q., Kinch L.N., Schaeffer R.D., et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373(6557):871–876. doi: 10.1126/science.abj8754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Senior A.W., Evans R., Jumper J., Kirkpatrick J., Sifre L., Green T., Qin C., Žídek A., Nelson A.W., Bridgland A., et al. Improved protein structure prediction using potentials from deep learning. Nature. 2020;577(7792):706–710. doi: 10.1038/s41586-019-1923-7. [DOI] [PubMed] [Google Scholar]
  • 72.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A., et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A.J., Bambrick J., et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024:1–3. doi: 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Lin Z., Akin H., Rao R., Hie B., Zhu Z., Lu W., Smetanin N., Verkuil R., Kabeli O., Shmueli Y., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–1130. doi: 10.1126/science.ade2574. [DOI] [PubMed] [Google Scholar]
  • 75.Shen D., Wu G., Suk H.-I. Deep learning in medical image analysis. Annu Rev Biomed Eng. 2017;19(1):221–248. doi: 10.1146/annurev-bioeng-071516-044442. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Rajkomar A., Oren E., Chen K., Dai A.M., Hajaj N., Hardt M., Liu P.J., Liu X., Marcus J., Sun M., et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1(1):1–10. doi: 10.1038/s41746-018-0029-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Stokes J.M., Yang K., Swanson K., Jin W., Cubillos-Ruiz A., Donghia N.M., MacNair C.R., French S., Carfrae L.A., Bloom-Ackermann Z., et al. A deep learning approach to antibiotic discovery. Cell. 2020;180(4):688–702. doi: 10.1016/j.cell.2020.01.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Vamathevan J., Clark D., Czodrowski P., Dunham I., Ferran E., Lee G., Li B., Madabhushi A., Shah P., Spitzer M., et al. Applications of machine learning in drug discovery and development. Nat Rev Drug Discov. 2019;18(6):463–477. doi: 10.1038/s41573-019-0024-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Li S., Wan F., Shu H., Jiang T., Zhao D., Zeng J. MONN: a multi-objective neural network for predicting compound-protein interactions and affinities. Cell Syst. 2020;10(4):308–322. [Google Scholar]
  • 80.Ge Y., Tian T., Huang S., Wan F., Li J., Li S., Wang X., Yang H., Hong L., Wu N., et al. An integrative drug repositioning framework discovered a potential therapeutic agent targeting COVID-19. Signal Transduct Target Ther. 2021;6(1):165. doi: 10.1038/s41392-021-00568-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Zhavoronkov A., Ivanenkov Y.A., Aliper A., Veselov M.S., Aladinskiy V.A., Aladinskaya A.V., Terentiev V.A., Polykovskiy D.A., Kuznetsov M.D., Asadulaev A., et al. Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nature Biotechnol. 2019;37(9):1038–1040. doi: 10.1038/s41587-019-0224-x. [DOI] [PubMed] [Google Scholar]
  • 82.Brakel A., Grochow T., Fritsche S., Knappe D., Krizsan A., Fietz S.A., Alber G., Hoffmann R., Müller U. Evaluation of proline-rich antimicrobial peptides as potential lead structures for novel antimycotics against cryptococcus neoformans. Front Microbiol. 2024;14 doi: 10.3389/fmicb.2023.1328890. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Buda De Cesare G., Cristy S.A., Garsin D.A., Lorenz M.C. Antimicrobial peptides: a new frontier in antifungal therapy. MBio. 2020;11(6):10–1128. doi: 10.1128/mBio.02123-20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Hu Y., Ling Y., Qin Z., Huang J., Jian L., et al. Isolation, identification, and synergistic mechanism of a novel antimicrobial peptide and phenolic compound from fermented walnut meal and their application in rosa roxbughii tratt spoilage fungus. Food Chem. 2024;433 doi: 10.1016/j.foodchem.2023.137333. [DOI] [PubMed] [Google Scholar]
  • 85.Qu B., Yuan J., Liu X., Zhang S., Ma X., Lu L. Anticancer activities of natural antimicrobial peptides from animals. Front Microbiol. 2024;14 doi: 10.3389/fmicb.2023.1321386. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Gong H., Wang X., Hu X., Liao M., Yuan C., Lu J.R., Gao L., Yan X. Effective treatment of helicobacter pylori infection using supramolecular antimicrobial peptide hydrogels. Biomacromolecules. 2024;25(3):1602–1611. doi: 10.1021/acs.biomac.3c01141. [DOI] [PubMed] [Google Scholar]
  • 87.Mardirossian M., Pompilio A., Degasperi M., Runti G., Pacor S., Di Bonaventura G., Scocchi M. D-BMAP18 antimicrobial peptide is active in vitro, resists to pulmonary proteases but loses its activity in a murine model of pseudomonas aeruginosa lung infection. Front Chem. 2017;5:40. doi: 10.3389/fchem.2017.00040. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Zhang H., Sun J., Xin X., Xu W., Shen J., Song Z., Yuan S. Modulating self-assembly behavior of a salt-free peptide amphiphile (PA) and zwitterionic surfactant mixed system. J Colloid Interface Sci. 2016;467:43–50. doi: 10.1016/j.jcis.2015.12.005. [DOI] [PubMed] [Google Scholar]
  • 89.Yang L., Gao Y., Zhang J., Tian C., Lin F., Song D., Zhou L., Peng J., Guo G. Antimicrobial peptide DvAMP combats carbapenem-resistant acinetobacter baumannii infection. Int J Antimicro Ag. 2024;63(4) doi: 10.1016/j.ijantimicag.2024.107106. [DOI] [PubMed] [Google Scholar]
  • 90.Primo L.M.D.G., Roque-Borda C.A., Canales C.S.C., Caruso I.P., de Lourenço I.O., Colturato V.M.M., Sábio R.M., de Melo F.A., Vicente E.F., Chorilli M., et al. Antimicrobial peptides grafted onto the surface of N-acetylcysteine-chitosan nanoparticles can revitalize drugs against clinical isolates of mycobacterium tuberculosis. Carbohydr Polymers. 2024;323 doi: 10.1016/j.carbpol.2023.121449. [DOI] [PubMed] [Google Scholar]
  • 91.Mohammed I., Said D.G., Dua H.S. Human antimicrobial peptides in ocular surface defense. Prog Retin Eye Res. 2017;61:1–22. doi: 10.1016/j.preteyeres.2017.03.004. [DOI] [PubMed] [Google Scholar]
  • 92.Silva C.N.d., Silva F.R.d., Dourado L.F.N., Reis P.V.M.d., Silva R.O., Costa B.L.d., Nunes P.S., Amaral F.A., Santos V.L.d., de Lima M.E., et al. A new topical eye drop containing lyetxi-b, a synthetic peptide designed from a lycosa erithrognata venom toxin, was effective to treat resistant bacterial keratitis. Toxins. 2019;11(4):203. doi: 10.3390/toxins11040203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Xiao Y., Jin L., Zhang C. From a hunger-regulating hormone to an antimicrobial peptide: gastrointestinal derived circulating endocrine hormone-peptide YY exerts exocrine antimicrobial effects against selective gut microbiota. Gut Microbes. 2024;16(1) doi: 10.1080/19490976.2024.2316927. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Herrera C.V., O’Connor P.M., Ratrey P., Ross R.P., Hill C., Hudson S.P. Anionic liposome formulation for oral delivery of thuricin CD, a potential antimicrobial peptide therapeutic. Int J Pharm. 2024;654 doi: 10.1016/j.ijpharm.2024.123918. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Neshani A., Zare H., Akbari Eidgahi M.R., Hooshyar Chichaklu A., Movaqar A., Ghazvini K. Review of antimicrobial peptides with anti-helicobacter pylori activity. Helicobacter. 2019;24(1) doi: 10.1111/hel.12555. [DOI] [PubMed] [Google Scholar]
  • 96.Narayana J.L., Huang H.-N., Wu C.-J., Chen J.-Y. Efficacy of the antimicrobial peptide TP4 against helicobacter pylori infection: in vitro membrane perturbation via micellization and in vivo suppression of host immune responses in a mouse model. Oncotarget. 2015;6(15):12936. doi: 10.18632/oncotarget.4101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Melicherčík P., Nešuta O., Čeřovskỳ V. Antimicrobial peptides for topical treatment of osteomyelitis and implant-related infections: study in the spongy bone. Pharmaceuticals. 2018;11(1):20. doi: 10.3390/ph11010020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Murdoch A.D., Grady L.M., Ablett M.P., Katopodi T., Meadows R.S., Hardingham T.E. Chondrogenic differentiation of human bone marrow stem cells in transwell cultures: generation of scaffold-free cartilage. Stem Cells. 2007;25(11):2786–2796. doi: 10.1634/stemcells.2007-0374. [DOI] [PubMed] [Google Scholar]
  • 99.Pfalzgraff A., Brandenburg K., Weindl G. Antimicrobial peptides and their therapeutic potential for bacterial skin infections and wounds. Front Pharmacol. 2018;9:281. doi: 10.3389/fphar.2018.00281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Guo L., McLean J.S., Yang Y., Eckert R., Kaplan C.W., Kyme P., Sheikh O., Varnum B., Lux R., Shi W., et al. Precision-guided antimicrobial peptide as a targeted modulator of human microbial ecology. Proc Natl Acad Sci. 2015;112(24):7569–7574. doi: 10.1073/pnas.1506207112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Donnelly J.P., Bellm L.A., Epstein J.B., Sonis S.T., Symonds R.P. Antimicrobial therapy to prevent or treat oral mucositis. Lancet Infect Dis. 2003;3(7):405–412. doi: 10.1016/s1473-3099(03)00668-6. [DOI] [PubMed] [Google Scholar]
  • 102.Fjell C.D., Hiss J.A., Hancock R.E., Schneider G. Designing antimicrobial peptides: form follows function. Nat Rev Drug Discov. 2012;11(1):37–51. doi: 10.1038/nrd3591. [DOI] [PubMed] [Google Scholar]
  • 103.Stephani J.C., Gerhards L., Khairalla B., Solov’yov I.A., Brand I. How do antimicrobial peptides interact with the outer membrane of gram-negative bacteria? Role of lipopolysaccharides in peptide binding, anchoring, and penetration. ACS Infect Dis. 2024;10(2):763–778. doi: 10.1021/acsinfecdis.3c00673. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Zhang M., Li S., Zhao J., Shuang Q., Xia Y., Zhang F. A novel endogenous antimicrobial peptide MP-4 derived from koumiss of inner mongolia by peptidomics, and effects on staphylococcus aureus. LWT. 2024;191 [Google Scholar]
  • 105.Luong H.X., Thanh T.T., Tran T.H. Antimicrobial peptides–advances in development of therapeutic applications. Life Sci. 2020;260 doi: 10.1016/j.lfs.2020.118407. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Galambos N., Vincent-Monegat C., Vallier A., Parisot N., Heddi A., Zaidman-Remy A. Cereal weevils’ antimicrobial peptides: at the crosstalk between development, endosymbiosis and immune response. Philos Trans R Soc B. 2024;379(1901) doi: 10.1098/rstb.2023.0062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Jangra M., Travin D.Y., Aleksandrova E.V., Kaur M., Darwish L., Koteva K., Klepacki D., Wang W., Tiffany M., Sokaribo A., et al. A broad-spectrum lasso peptide antibiotic targeting the bacterial ribosome. Nature. 2025:1–9. doi: 10.1038/s41586-025-08723-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Bahar A.A., Ren D. Antimicrobial peptides. Pharmaceuticals. 2013;6(12):1543–1575. doi: 10.3390/ph6121543. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Ramazi S., Mohammadi N., Allahverdi A., Khalili E., Abdolmaleki P. A review on antimicrobial peptides databases and the computational tools. Database. 2022;2022:baac011. doi: 10.1093/database/baac011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Clark S., Jowitt T.A., Harris L.K., Knight C.G., Dobson C.B. The lexicon of antimicrobial peptides: a complete set of arginine and tryptophan sequences. Commun Biol. 2021;4(1):605. doi: 10.1038/s42003-021-02137-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Ma Y., Guo Z., Xia B., Zhang Y., Liu X., Yu Y., Tang N., Tong X., Wang M., Ye X., et al. Identification of antimicrobial peptides from the human gut microbiome using deep learning. Nature Biotechnol. 2022;40(6):921–931. doi: 10.1038/s41587-022-01226-0. [DOI] [PubMed] [Google Scholar]
  • 112.Huang Y., Huang J., Chen Y. Alpha-helical cationic antimicrobial peptides: relationships of structure and function. Protein Cell. 2010;1:143–152. doi: 10.1007/s13238-010-0004-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Beevers A.J., Dixon A.M. Helical membrane peptides to modulate cell function. Chem Soc Rev. 2010;39(6):2146–2157. doi: 10.1039/b912944h. [DOI] [PubMed] [Google Scholar]
  • 114.Mulvenna J.P., Wang C., Craik D.J. Cybase: a database of cyclic protein sequence and structure. Nucleic Acids Res. 2006;34(suppl_1):D192–D194. doi: 10.1093/nar/gkj005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Wang C.K., Kaas Q., Chiche L., Craik D.J. Cybase: a database of cyclic protein sequences and structures, with applications in protein discovery and engineering. Nucleic Acids Res. 2007;36(suppl_1):D206–D210. doi: 10.1093/nar/gkm953. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Bin Hafeez A., Jiang X., Bergen P.J., Zhu Y. Antimicrobial peptides: an update on classifications and databases. Int J Mol Sci. 2021;22(21):11691. doi: 10.3390/ijms222111691. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Wan F., Wong F., Collins J.J., de la Fuente-Nunez C. Machine learning for antimicrobial peptide identification and design. Nat Rev Bioeng. 2024;2(5):392–407. doi: 10.1038/s44222-024-00152-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118.Kawashima S., Kanehisa M. AAindex: amino acid index database. Nucleic Acids Res. 2000;28(1):374. doi: 10.1093/nar/28.1.374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Haney E.F., Straus S.K., Hancock R.E. Reassessing the host defense peptide landscape. Front Chem. 2019;7 doi: 10.3389/fchem.2019.00043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Osorio D., Rondón-Villarreal P., Torres R. Peptides: a package for data mining of antimicrobial peptides. Small. 2015;12:44–444. [Google Scholar]
  • 121.van Westen G.J., Swier R.F., Wegner J.K., IJzerman A.P., van Vlijmen H.W., Bender A. Benchmarking of protein descriptor sets in proteochemometric modeling (part 1): comparative study of 13 amino acid descriptor sets. J Cheminform. 2013;5:1–11. doi: 10.1186/1758-2946-5-41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Müller A.T., Gabernet G., Hiss J.A., Schneider G. ModlAMP: Python for antimicrobial peptides. Bioinformatics. 2017;33(17):2753–2755. doi: 10.1093/bioinformatics/btx285. [DOI] [PubMed] [Google Scholar]
  • 123.Saeys Y., Inza I., Larranaga P. A review of feature selection techniques in bioinformatics. Bioinformatics. 2007;23(19):2507–2517. doi: 10.1093/bioinformatics/btm344. [DOI] [PubMed] [Google Scholar]
  • 124.Chen Z., Zhao P., Li F., Leier A., Marquez-Lago T.T., Wang Y., Webb G.I., Smith A.I., Daly R.J., Chou K.-C., et al. Ifeature: a python package and web server for features extraction and selection from protein and peptide sequences. Bioinformatics. 2018;34(14):2499–2502. doi: 10.1093/bioinformatics/bty140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.ElAbd H., Bromberg Y., Hoarfrost A., Lenz T., Franke A., Wendorff M. Amino acid encoding for deep learning applications. BMC Bioinformatics. 2020;21:1–14. doi: 10.1186/s12859-020-03546-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Chung J., Gülçehre Ç., Cho K., Bengio Y. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555. [Google Scholar]
  • 127.Yan K., Lv H., Guo Y., Peng W., Liu B. sAMPpred-GAT: prediction of antimicrobial peptide by graph attention network and predicted peptide structure. Bioinformatics. 2023;39(1):btac715. doi: 10.1093/bioinformatics/btac715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Ganea O, Pattanaik L, Coley CW, Barzilay R, Jensen KF, Jr. WHG, Jaakkola TS. GeoMol: Torsional Geometric Generation of Molecular 3D Conformer Ensembles. In: Conference on neural information processing systems 2021, neurIPS. 2021, p. 13757–69.
  • 129.Jin W, Wohlwend J, Barzilay R, Jaakkola TS. Iterative Refinement Graph Neural Network for Antibody Sequence-Structure Co-design. In: The tenth international conference on learning representations, ICLR. 2022.
  • 130.Maturana D., Scherer S.A. International conference on intelligent robots and systems, IROS. IEEE; 2015. VoxNet: A 3D convolutional neural network for real-time object recognition; pp. 922–928. [Google Scholar]
  • 131.Jiménez J., Doerr S., Martínez-Rosell G., Rose A.S., De Fabritiis G. DeepSite: protein-binding site predictor using 3D-convolutional neural networks. Bioinformatics. 2017;33(19):3036–3042. doi: 10.1093/bioinformatics/btx350. [DOI] [PubMed] [Google Scholar]
  • 132.Jones D., Kim H., Zhang X., Zemla A., Stevenson G., Bennett W.D., Kirshner D., Wong S.E., Lightstone F.C., Allen J.E. Improved protein–ligand binding affinity prediction with structure-based deep fusion inference. J Chem Inf Model. 2021;61(4):1583–1592. doi: 10.1021/acs.jcim.0c01306. [DOI] [PubMed] [Google Scholar]
  • 133.Bengio Y., Courville A.C., Vincent P. Representation learning: A review and new perspectives. IEEE Trans Pattern Anal Mach Intell. 2013;35(8):1798–1828. doi: 10.1109/TPAMI.2013.50. [DOI] [PubMed] [Google Scholar]
  • 134.Grill J, Strub F, Altché F, Tallec C, Richemond PH, Buchatskaya E, Doersch C, Pires BÁ, Guo Z, Azar MG, Piot B, Kavukcuoglu K, Munos R, Valko M. Bootstrap Your Own Latent - A New Approach to Self-Supervised Learning. In: Conference on neural information processing systems 2020, neurIPS. 2020.
  • 135.Mikolov T, Chen K, Corrado G, Dean J. Efficient Estimation of Word Representations in Vector Space. In: 1st international conference on learning representations, ICLR. 2013.
  • 136.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler DM, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D. Language Models are Few-Shot Learners. In: Conference on neural information processing systems 2020, neurIPS. 2020.
  • 137.Sun T.-J., Bu H.-L., Yan X., Sun Z.-H., Zha M.-S., Dong G.-F. LABAMPsGCN: A framework for identifying lactic acid bacteria antimicrobial peptides based on graph convolutional neural network. Front Genet. 2022;13 doi: 10.3389/fgene.2022.1062576. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Lee E.Y., Fulan B.M., Wong G.C., Ferguson A.L. Mapping membrane activity in undiscovered peptide sequence space using machine learning. Proc Natl Acad Sci. 2016;113(48):13588–13593. doi: 10.1073/pnas.1609893113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Gao W., Zhao J., Gui J., Wang Z., Chen J., Yue Z. Comprehensive assessment of BERT-based methods for predicting antimicrobial peptides. J Chem Inf Model. 2024 doi: 10.1021/acs.jcim.4c00507. [DOI] [PubMed] [Google Scholar]
  • 140.Wang J., Feng J., Kang Y., Pan P., Ge J., Wang Y., Wang M., Wu Z., Zhang X., Yu J., et al. Discovery of antimicrobial peptides with notable antibacterial potency by an LLM-based foundation model. Sci Adv. 2025;11(10):eads8932. doi: 10.1126/sciadv.ads8932. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Consortium U. UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res. 2019;47(D1):D506–D515. doi: 10.1093/nar/gky1049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Chen J., Jia Y., Sun Y., Liu K., Zhou C., Liu C., Li D., Liu G., Zhang C., Yang T., et al. Global marine microbial diversity and its potential in bioprospecting. Nature. 2024:1–9. doi: 10.1038/s41586-024-07891-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Xu J., Xu X., Jiang Y., Fu Y., Shen C. Waste to resource: Mining antimicrobial peptides in sludge from metagenomes using machine learning. Environ Int. 2024;186 doi: 10.1016/j.envint.2024.108574. [DOI] [PubMed] [Google Scholar]
  • 144.Xiao X., Shao Y.-T., Cheng X., Stamatovic B. iAMP-CA2L: a new CNN-BiLSTM-SVM classifier based on cellular automata image for identifying antimicrobial peptides and their functional types. Brief Bioinform. 2021;22(6):bbab209. doi: 10.1093/bib/bbab209. [DOI] [PubMed] [Google Scholar]
  • 145.Li C., Sutherland D., Hammond S.A., Yang C., Taho F., Bergman L., Houston S., Warren R.L., Wong T., Hoang L.M., et al. AMPlify: attentive deep learning model for discovery of novel antimicrobial peptides effective against WHO priority pathogens. BMC Genomics. 2022;23(1):77. doi: 10.1186/s12864-022-08310-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Lai Z., Yuan X., Chen H., Zhu Y., Dong N., Shan A. Strategies employed in the design of antimicrobial peptides with enhanced proteolytic stability. Biotech Adv. 2022;59 doi: 10.1016/j.biotechadv.2022.107962. [DOI] [PubMed] [Google Scholar]
  • 147.Persikov A.V., Ramshaw J.A., Brodsky B. Prediction of collagen stability from amino acid sequence. J Biol Chem. 2005;280(19):19343–19349. doi: 10.1074/jbc.M501657200. [DOI] [PubMed] [Google Scholar]
  • 148.Wang F., Sangfuang N., McCoubrey L.E., Yadav V., Elbadawi M., Orlu M., Gaisford S., Basit A.W. Advancing oral delivery of biologics: Machine learning predicts peptide stability in the gastrointestinal tract. Int J Pharm. 2023;634 doi: 10.1016/j.ijpharm.2023.122643. [DOI] [PubMed] [Google Scholar]
  • 149.Guo X., Miao X., An Y., Yan T., Jia Y., Deng B., Cai J., Yang W., Sun W., Wang R., et al. Novel antimicrobial peptides modified with fluorinated sulfono-γ-AA having high stability and targeting multidrug-resistant bacteria infections. Eur J Med Chem. 2024;264 doi: 10.1016/j.ejmech.2023.116001. [DOI] [PubMed] [Google Scholar]
  • 150.Antunes B., Zanchi C., Johnston P.R., Maron B., Witzany C., Regoes R.R., Hayouka Z., Rolff J. The evolution of antimicrobial peptide resistance in pseudomonas aeruginosa is severely constrained by random peptide mixtures. PLoS Biol. 2024;22(7) doi: 10.1371/journal.pbio.3002692. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Baeder D.Y., Yu G., Hozé N., Rolff J., Regoes R.R. Antimicrobial combinations: Bliss independence and loewe additivity derived from mechanistic multi-hit models. Phil Trans R Soc B. 2016;371(1695) doi: 10.1098/rstb.2015.0294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Russ D., Kishony R. Additivity of inhibitory effects in multidrug combinations. Nat Microbiol. 2018;3(12):1339–1345. doi: 10.1038/s41564-018-0252-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153.Rabel D., Charlet M., Ehret-Sabatier L., Cavicchioli L., Cudic M., Otvos L., Bulet P. Primary structure and in vitro antibacterial properties of the drosophila melanogaster attacin c pro-domain. J Biol Chem. 2004;279(15):14853–14859. doi: 10.1074/jbc.M313608200. [DOI] [PubMed] [Google Scholar]
  • 154.Lazzaro B.P., Zasloff M., Rolff J. Antimicrobial peptides: Application informed by evolution. Science. 2020;368(6490):eaau5480. doi: 10.1126/science.aau5480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Liu H., Song Z., Zhang Y., Wu B., Chen D., Zhou Z., Zhang H., Li S., Feng X., Huang J., et al. De novo design of self-assembling peptides with antimicrobial activity guided by deep learning. Nat Mater. 2025:1–12. doi: 10.1038/s41563-025-02164-3. [DOI] [PubMed] [Google Scholar]
  • 156.Deng J., Dong W., Socher R., Li L., Li K., Fei-Fei L. Conference on computer vision and pattern recognition (CVPR. IEEE Computer Society; 2009. ImageNet: A large-scale hierarchical image database; pp. 248–255. [Google Scholar]
  • 157.Wang G., Li X., Wang Z. APD3: the antimicrobial peptide database as a tool for research and education. Nucleic Acids Res. 2016;44(D1):D1087–D1093. doi: 10.1093/nar/gkv1278. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 158.Witten J., Witten Z. 2019. Deep learning regression model for antimicrobial peptide design. BioRxiv. [Google Scholar]
  • 159.Vishnepolsky B., Grigolava M., Managadze G., Gabrielian A., Rosenthal A., Hurt D.E., Tartakovsky M., Pirtskhalava M. Comparative analysis of machine learning algorithms on the microbial strain-specific AMP prediction. Brief Bioinform. 2022;23(4):bbac233. doi: 10.1093/bib/bbac233. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160.Bassler O.B. The surveyability of mathematical proof: A historical perspective. Synthese. 2006;148:99–133. [Google Scholar]
  • 161.Sundararajan M., Taly A., Yan Q. Proceedings of the 34th international conference on machine learning, ICML. vol. 70. PMLR; 2017. Axiomatic attribution for deep networks; pp. 3319–3328. (Proceedings of machine learning research). [Google Scholar]
  • 162.Shrikumar A., Greenside P., Kundaje A. Proceedings of the 34th international conference on machine learning, ICML. vol. 70. PMLR; 2017. Learning important features through propagating activation differences; pp. 3145–3153. (Proceedings of machine learning research). [Google Scholar]
  • 163.Ribeiro M.T., Singh S., Guestrin C. Proceedings of the 22nd ACM international conference on knowledge discovery and data mining, SIGKDD. ACM; 2016. ”Why should I trust you?”: Explaining the predictions of any classifier; pp. 1135–1144. [Google Scholar]
  • 164.Yuan H., Yu H., Gui S., Ji S. Explainability in graph neural networks: A taxonomic survey. IEEE Trans Pattern Anal Mach Intell. 2022;45(5):5782–5799. doi: 10.1109/TPAMI.2022.3204236. [DOI] [PubMed] [Google Scholar]
  • 165.Lundberg SM, Lee S. A Unified Approach to Interpreting Model Predictions. In: Conference on neural information processing systems, neurIPS. 2017, p. 4765–74.
  • 166.Jiménez-Luna J., Grisoni F., Schneider G. Drug discovery with explainable artificial intelligence. Nat Mach Intell. 2020;2(10):573–584. [Google Scholar]
  • 167.Sidorczuk K., Gagat P., Pietluch F., Kala J., Rafacz D., Bakala L., Slowik J., Kolenda R., Rödiger S., Fingerhut L.C.H.W., Cooke I., Mackiewicz P., Burdukiewicz M. Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Brief Bioinform. 2022;23(5):bbac343. doi: 10.1093/bib/bbac343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 168.Heil B.J., Hoffman M.M., Markowetz F., Lee S.-I., Greene C.S., Hicks S.C. Reproducibility standards for machine learning in the life sciences. Nature Methods. 2021;18(10):1132–1135. doi: 10.1038/s41592-021-01256-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 169.García-Jacas C.R., Pinacho-Castellanos S.A., García-González L.A., Brizuela C.A. Do deep learning models make a difference in the identification of antimicrobial peptides? Brief Bioinform. 2022;23(3):bbac094. doi: 10.1093/bib/bbac094. [DOI] [PubMed] [Google Scholar]
  • 170.Huang J., Xu Y., Xue Y., Huang Y., Li X., Chen X., Xu Y., Zhang D., Zhang P., Zhao J., et al. Identification of potent antimicrobial peptides via a machine-learning pipeline that mines the entire space of peptide sequences. Nat Biomed Eng. 2023;7(6):797–810. doi: 10.1038/s41551-022-00991-2. [DOI] [PubMed] [Google Scholar]
  • 171.Santos-Júnior C.D., Torres M.D., Duan Y., Del Río Á.R., Schmidt T.S., Chong H., Fullam A., Kuhn M., Zhu C., Houseman A., et al. Discovery of antimicrobial peptides in the global microbiome with machine learning. Cell. 2024;187(14):3761–3778. doi: 10.1016/j.cell.2024.05.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 172.Yoshida M., Hinkley T., Tsuda S., Abul-Haija Y.M., McBurney R.T., Kulikov V., Mathieson J.S., Reyes S.G., Castro M.D., Cronin L. Using evolutionary algorithms and machine learning to explore sequence space for the discovery of antimicrobial peptides. Chem. 2018;4(3):533–543. [Google Scholar]
  • 173.Porto W.F., Irazazabal L., Alves E.S., Ribeiro S.M., Matos C.O., Pires Á.S., Fensterseifer I.C., Miranda V.J., Haney E.F., Humblot V., et al. In silico optimization of a guava antimicrobial peptide enables combinatorial exploration for peptide design. Nat Commun. 2018;9(1):1490. doi: 10.1038/s41467-018-03746-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 174.Zhang H., Wang Y., Zhu Y., Huang P., Gao Q., Li X., Chen Z., Liu Y., Jiang J., Gao Y., et al. Machine learning and genetic algorithm-guided directed evolution for the development of antimicrobial peptides. J Adv Res. 2025;68:415–428. doi: 10.1016/j.jare.2024.02.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 175.Muller A.T., Hiss J.A., Schneider G. Recurrent neural network model for constructive peptide design. J Chem Inf Model. 2018;58(2):472–479. doi: 10.1021/acs.jcim.7b00414. [DOI] [PubMed] [Google Scholar]
  • 176.Mao J., Guan S., Chen Y., Zeb A., Sun Q., Lu R., Dong J., Wang J., Cao D. Application of a deep generative model produces novel and diverse functional peptides against microbial resistance. Comput Struct Biotechnol J. 2023;21:463–471. doi: 10.1016/j.csbj.2022.12.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 177.Nagarajan D., Nagarajan T., Roy N., Kulkarni O., Ravichandran S., Mishra M., Chakravortty D., Chandra N. Computational antimicrobial peptide design and evaluation against multidrug-resistant clinical isolates of bacteria. J Biol Chem. 2018;293(10):3492–3509. doi: 10.1074/jbc.M117.805499. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 178.Wang C., Garlick S., Zloh M. Deep learning for novel antimicrobial peptide design. Biomolecules. 2021;11(3):471. doi: 10.3390/biom11030471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 179.Capecchi A., Cai X., Personne H., Köhler T., van Delden C., Reymond J.-L. Machine learning designs non-hemolytic antimicrobial peptides. Chem Sci. 2021;12(26):9221–9232. doi: 10.1039/d1sc01713f. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 180.Wan F., Kontogiorgos-Heintz D., de la Fuente-Nunez C. Deep generative models for peptide design. Digit Discov. 2022;1(3):195–208. doi: 10.1039/d1dd00024a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 181.Dean S.N., Alvarez J.A.E., Zabetakis D., Walper S.A., Malanoski A.P. PepVAE: variational autoencoder framework for antimicrobial peptide generation and activity prediction. Front Microbiol. 2021;12 doi: 10.3389/fmicb.2021.725727. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 182.Das P., Sercu T., Wadhawan K., Padhi I., Gehrmann S., Cipcigan F., Chenthamarakshan V., Strobelt H., Dos Santos C., Chen P.-Y., et al. Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nat Biomed Eng. 2021;5(6):613–623. doi: 10.1038/s41551-021-00689-x. [DOI] [PubMed] [Google Scholar]
  • 183.Das P., Wadhawan K., Chang O., Sercu T., Santos C.D., Riemer M., Chenthamarakshan V., Padhi I., Mojsilovic A. 2018. PepCVAE: Semi-supervised targeted design of antimicrobial peptide sequences. arXiv preprint arXiv:1810.07743. [Google Scholar]
  • 184.Wang D, Wen Z, Fei Y, Li L, Zhou H. Accelerating antimicrobial peptide discovery with latent sequence-structure model. In: ICLR 2023-machine learning for drug discovery workshop. 2022.
  • 185.Lin T.-T., Yang L.-Y., Lin C.-Y., Wang C.-T., Lai C.-W., Ko C.-F., Shih Y.-H., Chen S.-H. Intelligent de novo design of novel antimicrobial peptides against antibiotic-resistant bacteria strains. Int J Mol Sci. 2023;24(7):6788. doi: 10.3390/ijms24076788. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 186.Van Oort C.M., Ferrell J.B., Remington J.M., Wshah S., Li J. AMPGAN v2: machine learning-guided design of antimicrobial peptides. J Chem Inf Model. 2021;61(5):2198–2207. doi: 10.1021/acs.jcim.0c01441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 187.Cao Q., Ge C., Wang X., Harvey P.J., Zhang Z., Ma Y., Wang X., Jia X., Mobli M., Craik D.J., et al. Designing antimicrobial peptides using deep learning and molecular dynamic simulations. Brief Bioinform. 2023;24(2):bbad058. doi: 10.1093/bib/bbad058. [DOI] [PubMed] [Google Scholar]
  • 188.Yu L., Zhang W., Wang J., Yu Y. Proceedings of the thirty-first AAAI conference on artificial intelligence. AAAI Press; 2017. SeqGAN: Sequence generative adversarial nets with policy gradient; pp. 2852–2858. [Google Scholar]
  • 189.Anand N., Achim T. 2022. Protein structure and sequence generation with equivariant denoising diffusion probabilistic models. arXiv preprint arXiv:2205.15019. [Google Scholar]
  • 190.Guo Z., Liu J., Wang Y., Chen M., Wang D., Xu D., Cheng J. Diffusion models in bioinformatics and computational biology. Nat Rev Bioeng. 2024;2(2):136–154. doi: 10.1038/s44222-023-00114-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 191.Wu K.E., Yang K.K., van den Berg R., Alamdari S., Zou J.Y., Lu A.X., Amini A.P. Protein structure generation via folding diffusion. Nat Commun. 2024;15(1):1059. doi: 10.1038/s41467-024-45051-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 192.Yi K, Zhou B, Shen Y, Lió P, Wang Y. Graph Denoising Diffusion for Inverse Protein Folding. In: Conference on neural information processing systems, neurIPS. 2023.
  • 193.Watson J.L., Juergens D., Bennett N.R., Trippe B.L., Yim J., Eisenach H.E., Ahern W., Borst A.J., Ragotte R.J., Milles L.F., et al. De novo design of protein structure and function with RFdiffusion. Nature. 2023;620(7976):1089–1100. doi: 10.1038/s41586-023-06415-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 194.Cao H., Tan C., Gao Z., Xu Y., Chen G., Heng P., Li S.Z. A survey on generative diffusion models. IEEE Trans Knowl Data Eng. 2024;36(7):2814–2830. [Google Scholar]
  • 195.Chen T., Vure P., Pulugurta R., Chatterjee P. AMP-diffusion: Integrating latent diffusion with protein language models for antimicrobial peptide generation. BioRxiv. 2024 2024–03. [Google Scholar]
  • 196.Qi Y, Jiang X, Jiang Y, Yang Y, Zhang Q, Tian Y. Antimicrobial Peptide Sequence Generation Based on Conditional Diffusion Model. In: Proceedings of the 2024 16th international conference on bioinformatics and biomedical technology. 2024, p. 102–7.
  • 197.Wang R., Wang T., Zhuo L., Wei J., Fu X., Zou Q., Yao X. Diff-AMP: tailored designed antimicrobial peptide framework with all-in-one generation, identification, prediction and optimization. Brief Bioinform. 2024;25(2):bbae078. doi: 10.1093/bib/bbae078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 198.Szymczak P., Możejko M., Grzegorzek T., Jurczak R., Bauer M., Neubauer D., Sikora K., Michalski M., Sroka J., Setny P., et al. Discovering highly potent antimicrobial peptides with deep generative model HydrAMP. Nat Commun. 2023;14(1):1453. doi: 10.1038/s41467-023-36994-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 199.Bolatchiev A., Baturin V., Shchetinin E., Bolatchieva E. Novel antimicrobial peptides designed using a recurrent neural network reduce mortality in experimental sepsis. Antibiotics. 2022;11(3):411. doi: 10.3390/antibiotics11030411. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 200.Wang Y., Song M., Liu F., Liang Z., Hong R., Dong Y., Luan H., Fu X., Yuan W., Fang W., et al. Artificial intelligence using a latent diffusion model enables the generation of diverse and potent antimicrobial peptides. Sci Adv. 2025;11(6):eadp7171. doi: 10.1126/sciadv.adp7171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 201.Gao Q., Ge L., Wang Y., Zhu Y., Liu Y., Zhang H., Huang J., Qin Z. An explainable few-shot learning model for the directed evolution of antimicrobial peptides. Int J Biiol Macromol. 2025;285 doi: 10.1016/j.ijbiomac.2024.138272. [DOI] [PubMed] [Google Scholar]
  • 202.Levenshtein V. Binary codes capable of correcting deletions, insertions, and reversals. Proc Sov Phys Dokl. 1966 [Google Scholar]
  • 203.Gotoh O. An improved algorithm for matching biological sequences. J Mol Biol. 1982;162(3):705–708. doi: 10.1016/0022-2836(82)90398-9. [DOI] [PubMed] [Google Scholar]
  • 204.Notin P., Rollins N., Gal Y., Sander C., Marks D. Machine learning for functional protein design. Nature Biotechnol. 2024;42(2):216–228. doi: 10.1038/s41587-024-02127-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 205.Olcay B., Ozdemir G.D., Ozdemir M.A., Ercan U.K., Guren O., Karaman O. Prediction of the synergistic effect of antimicrobial peptides and antimicrobial agents via supervised machine learning. BMC Biomed Eng. 2024;6(1):1. doi: 10.1186/s42490-024-00075-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 206.Sun Z., Xu J., Zhang Y., Zhang Y., Wang Z., Wang X., Li S., Guo Y., Shen H.H., Song J. Multimodal geometric learning for antimicrobial peptide identification by leveraging alphafold2-predicted structures and surface features. Brief Bioinform. 2025;26(3) doi: 10.1093/bib/bbaf261. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 207.Gao X., Cao H., Li J., Qiu J., Chen G., Heng P.-A. Gated-GPS: enhancing protein–protein interaction site prediction with scalable learning and imbalance-aware optimization. Brief Bioinform. 2025;26(3):bbaf248. doi: 10.1093/bib/bbaf248. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 208.Vishnepolsky B., Pirtskhalava M. Prediction of linear cationic antimicrobial peptides based on characteristics responsible for their interaction with the membranes. J Chem Inf Model. 2014;54(5):1512–1523. doi: 10.1021/ci4007003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 209.Randall J.R., Vieira L.C., Wilke C.O., Davies B.W. Deep mutational scanning and machine learning for the analysis of antimicrobial-peptide features driving membrane selectivity. Nat Biomed Eng. 2024;8(7):842–853. doi: 10.1038/s41551-024-01243-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 210.Shi C., Xu M., Zhu Z., Zhang W., Zhang M., Tang J. 8th international conference on learning representations, ICLR. OpenReview.net; 2020. GraphAF: a flow-based autoregressive model for molecular graph generation. [Google Scholar]
  • 211.Janson G., Valdes-Garcia G., Heo L., Feig M. Direct generation of protein conformational ensembles via machine learning. Nat Commun. 2023;14(1):774. doi: 10.1038/s41467-023-36443-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 212.Maasch J.R., Torres M.D., Melo M.C., de la Fuente-Nunez C. Molecular de-extinction of ancient antimicrobial peptides enabled by machine learning. Cell Host Microbe. 2023;31(8):1260–1274. doi: 10.1016/j.chom.2023.07.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 213.Schissel C.K., Mohapatra S., Wolfe J.M., Fadzen C.M., Bellovoda K., Wu C.-L., Wood J.A., Malmberg A.B., Loas A., Gómez-Bombarelli R., et al. Deep learning to design nuclear-targeting abiotic miniproteins. Nat Chem. 2021;13(10):992–1000. doi: 10.1038/s41557-021-00766-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 214.Song A.H., Chen R.J., Jaume G., Vaidya A.J., Baras A.S., Mahmood F. International conference on machine learning, ICML. OpenReview.net; 2024. Multimodal prototyping for cancer survival prediction. [Google Scholar]
  • 215.Lei D., Xu M., Wang S. A deep multimodal network for multi-task trajectory prediction. Inf Fusion. 2025;113 [Google Scholar]
  • 216.Lian J., Luo X., Shan C., Han D., Zhang C., Vardhanabhuti V., Li D., Qiu L. Personalized progression modelling and prediction in parkinson’s disease with a novel multi-modal graph approach. Npj Parkinson’s Dis. 2024;10(1):229. doi: 10.1038/s41531-024-00832-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 217.Chen Z., Zhao P., Li C., Li F., Xiang D., Chen Y.-Z., Akutsu T., Daly R.J., Webb G.I., Zhao Q., et al. ILearnPlus: a comprehensive and automated machine-learning platform for nucleic acid and protein sequence analysis, prediction and visualization. Nucleic Acids Res. 2021;49(10):e60. doi: 10.1093/nar/gkab122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 218.Dauparas J., Anishchenko I., Bennett N., Bai H., Ragotte R.J., Milles L.F., Wicky B.I., Courbet A., de Haas R.J., Bethel N., et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science. 2022;378(6615):49–56. doi: 10.1126/science.add2187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 219.Torres M.D., Pedron C.N., Higashikuni Y., Kramer R.M., Cardoso M.H., Oshiro K.G., Franco O.L., Silva Junior P.I., Silva F.D., Oliveira Junior V.X., et al. Structure-function-guided exploration of the antimicrobial peptide polybia-CP identifies activity determinants and generates synthetic therapeutic candidates. Commun Biol. 2018;1(1):221. doi: 10.1038/s42003-018-0224-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 220.Boaro A., Ageitos L., Torres M.D.T., Blasco E.B., Oztekin S., de la Fuente-Nunez C. Structure-function-guided design of synthetic peptides with anti-infective activity derived from wasp venom. Cell Rep Phys Sci. 2023;4(7) doi: 10.1016/j.xcrp.2023.101459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 221.Bello-Madruga R., Burgas M.T. The limits of prediction: Why intrinsically disordered regions challenge our understanding of antimicrobial peptides. Comput Struct Biotechnol J. 2024;23:972–981. doi: 10.1016/j.csbj.2024.02.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 222.Yang B., Yang H., Liang J., Chen J., Wang C., Wang Y., Wang J., Luo W., Deng T., Guo J. A review on the screening methods for the discovery of natural antimicrobial peptides. J Pharm Anal. 2024 doi: 10.1016/j.jpha.2024.101046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 223.Altschul S.F., Gish W., Miller W., Myers E.W., Lipman D.J. Basic local alignment search tool. J Mol Biol. 1990;215(3):403–410. doi: 10.1016/S0022-2836(05)80360-2. [DOI] [PubMed] [Google Scholar]
  • 224.Wong F., de la Fuente-Nunez C., Collins J.J. Leveraging artificial intelligence in the fight against infectious diseases. Science. 2023;381(6654):164–170. doi: 10.1126/science.adh1114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 225.Arnold M.F., Shabab M., Penterman J., Boehme K.L., Griffitts J.S., Walker G.C. Genome-wide sensitivity analysis of the microsymbiont sinorhizobium meliloti to symbiotically important, defensin-like host peptides. MBio. 2017;8(4):10–1128. doi: 10.1128/mBio.01060-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 226.Brcic M., Yampolskiy R.V. Impossibility results in AI: a survey. Acm Comput Surv. 2023;56(1):1–24. [Google Scholar]
  • 227.Gödel K. Über formal unentscheidbare sätze der principia mathematica und verwandter systeme i. Monatshefte Math Phys. 1931;38:173–198. [Google Scholar]
  • 228.Wolpert D.H. The existence of a priori distinctions between learning algorithms. Neural Comput. 1996;8(7):1391–1420. [Google Scholar]
  • 229.Wolpert D.H., Macready W.G. No free lunch theorems for optimization. IEEE Trans Evol Comput. 1997;1(1):67–82. [Google Scholar]
  • 230.Goodman N. A query on confirmation. J Philos. 1946;43(14):383–385. [Google Scholar]
  • 231.Zanchi C., Johnston P.R., Rolff J. Evolution of defence cocktails: Antimicrobial peptide combinations reduce mortality and persistent infection. Mol Ecol. 2017;26(19):5334–5343. doi: 10.1111/mec.14267. [DOI] [PubMed] [Google Scholar]
  • 232.Unckless R.L., Howick V.M., Lazzaro B.P. Convergent balancing selection on an antimicrobial peptide in drosophila. Curr Biol. 2016;26(2):257–262. doi: 10.1016/j.cub.2015.11.063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 233.Unckless R.L., Lazzaro B.P. The potential for adaptive maintenance of diversity in insect antimicrobial peptides. Phil Trans R Soc B. 2016;371(1695) doi: 10.1098/rstb.2015.0291. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 234.Tang K., Wang X., Niu M., Wang X., Zhou G., Shi J., Yu Y., Chen Z., Li C. Augmenting the precise targeting of antimicrobial peptides (AMPs) and AMP-based drug delivery via affinity-filtering strategy. Adv Funct Mater. 2022;32(17) [Google Scholar]
  • 235.Jyakhwo S., Serov N., Dmitrenko A., Vinogradov V.V. Machine learning reinforced genetic algorithm for massive targeted discovery of selectively cytotoxic inorganic nanoparticles. Small. 2024;20(6) doi: 10.1002/smll.202305375. [DOI] [PubMed] [Google Scholar]
  • 236.Kaplan J., McCandlish S., Henighan T., Brown T.B., Chess B., Child R., Gray S., Radford A., Wu J., Amodei D. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. [Google Scholar]
  • 237.Wei J., Tay Y., Bommasani R., Raffel C., Zoph B., Borgeaud S., Yogatama D., Bosma M., Zhou D., Metzler D., et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682. [Google Scholar]
  • 238.Schaeffer R, Miranda B, Koyejo S. Are Emergent Abilities of Large Language Models a Mirage?. In: Conference on neural information processing systems, neurIPS. 2023.
  • 239.Csordás R, Piekos P, Irie K, Schmidhuber J. SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention. In: Conference on neural information processing systems, neurIPS. 2024.
  • 240.Zhu X, Guan Y, Liang D, Chen Y, Liu Y, Bai X. MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks. In: Conference on neural information processing systems, neurIPS. 2024.
  • 241.Shazeer N., Mirhoseini A., Maziarz K., Davis A., Le Q., Hinton G., Dean J. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538. [Google Scholar]
  • 242.Jacobs R.A., Jordan M.I., Nowlan S.J., Hinton G.E. Adaptive mixtures of local experts. Neural Comput. 1991;3(1):79–87. doi: 10.1162/neco.1991.3.1.79. [DOI] [PubMed] [Google Scholar]
  • 243.Durante Z., Huang Q., Wake N., Gong R., Park J.S., Sarkar B., Taori R., Noda Y., Terzopoulos D., Choi Y., Ikeuchi K., Vo H., Fei-Fei L., Gao J. 2024. Agent ai: Surveying the horizons of multimodal interaction. arXiv preprint arXiv:2401.03568. [Google Scholar]
  • 244.Rapp J.T., Bremer B.J., Romero P.A. Self-driving laboratories to autonomously navigate the protein fitness landscape. Nat Chem Eng. 2024;1(1):97–107. doi: 10.1038/s44286-023-00002-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 245.Yu T., Boob A.G., Singh N., Su Y., Zhao H. In vitro continuous protein evolution empowered by machine learning and automation. Cell Syst. 2023;14(8):633–644. doi: 10.1016/j.cels.2023.04.006. [DOI] [PubMed] [Google Scholar]
  • 246.Wortsman M., Ilharco G., Gadre S.Y., Roelofs R., Lopes R.G., Morcos A.S., Namkoong H., Farhadi A., Carmon Y., Kornblith S., Schmidt L. International conference on machine learning, ICML. vol. 162. PMLR; 2022. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time; pp. 23965–23998. (Proceedings of machine learning research). [Google Scholar]
  • 247.Zhu B, Cui J, Zhang H. Robust Fine-tuning of Zero-shot Models via Variance Reduction. In: Conference on neural information processing systems, neurIPS. 2024.
  • 248.Wortsman M., Ilharco G., Kim J.W., Li M., Kornblith S., Roelofs R., Lopes R.G., Hajishirzi H., Farhadi A., Namkoong H., Schmidt L. IEEE/CVF conference on computer vision and pattern recognition, CVPR. IEEE; 2022. Robust fine-tuning of zero-shot models. [Google Scholar]
  • 249.Malladi S, Gao T, Nichani E, Damian A, Lee JD, Chen D, Arora S. Fine-Tuning Language Models with Just Forward Passes. In: Conference on neural information processing systems, neurIPS. 2023.
  • 250.Zhang L., Rao A., Agrawala M. International conference on learning representations, ICLR. 2025. Scaling in-the-wild training for diffusion-based illumination harmonization and editing by imposing consistent light transport. [Google Scholar]
  • 251.Karniadakis G.E., Kevrekidis I.G., Lu L., Perdikaris P., Wang S., Yang L. Physics-informed machine learning. Nat Rev Phys. 2021;3(6):422–440. [Google Scholar]
  • 252.Malbranke C., Bikard D., Cocco S., Monasson R., Tubiana J. Machine learning for evolutionary-based and physics-inspired protein design: Current and future synergies. Curr Opin Struct Biol. 2023;80 doi: 10.1016/j.sbi.2023.102571. [DOI] [PubMed] [Google Scholar]
  • 253.Doerr S., Majewski M., Pérez A., Kramer A., Clementi C., Noe F., Giorgino T., De Fabritiis G. TorchMD: A deep learning framework for molecular simulations. J Chem Theory Comput. 2021;17(4):2355–2363. doi: 10.1021/acs.jctc.0c01343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 254.Bond P.J., Holyoake J., Ivetac A., Khalid S., Sansom M.S. Coarse-grained molecular dynamics simulations of membrane proteins and peptides. J Struct Biol. 2007;157(3):593–605. doi: 10.1016/j.jsb.2006.10.004. [DOI] [PubMed] [Google Scholar]
  • 255.Abraham M.J., Murtola T., Schulz R., Páll S., Smith J.C., Hess B., Lindahl E. GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX. 2015;1:19–25. [Google Scholar]
  • 256.Li W., Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 2006;22(13):1658–1659. doi: 10.1093/bioinformatics/btl158. [DOI] [PubMed] [Google Scholar]
  • 257.Osmanbeyoglu H.U., Ganapathiraju M.K. N-gram analysis of 970 microbial organisms reveals presence of biological language models. BMC Bioinformatics. 2011;12(1):12. doi: 10.1186/1471-2105-12-12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 258.Chaudhary K., Kumar R., Singh S., Tuknait A., Gautam A., Mathur D., Anand P., Varshney G.C., Raghava G.P. A web server and mobile app for computing hemolytic potency of peptides. Sci Rep. 2016;6(1):22843. doi: 10.1038/srep22843. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 259.Gupta S., Kapoor P., Chaudhary K., Gautam A., Kumar R., Raghava G.P. Computational peptidology. Springer; 2014. Peptide toxicity prediction; pp. 143–157. [DOI] [PubMed] [Google Scholar]
  • 260.Gupta S., Kapoor P., Chaudhary K., Gautam A., Kumar R., Consortium O.S.D.D., Raghava G.P. In silico approach for predicting toxicity of peptides and proteins. PloS One. 2013;8(9) doi: 10.1371/journal.pone.0073957. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 261.Chung C.-R., Chien C.-Y., Tang Y., Wu L.-C., Hsu J.B.-K., Lu J.-J., Lee T.-Y., Bai C., Horng J.-T. An ensemble deep learning model for predicting minimum inhibitory concentrations of antimicrobial peptides against pathogenic bacteria. Iscience. 2024;27(9) doi: 10.1016/j.isci.2024.110718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 262.Yan J., Zhang B., Zhou M., Campbell-Valois F.-X., Siu S.W. A deep learning method for predicting the minimum inhibitory concentration of antimicrobial peptides against escherichia coli using multi-branch-CNN and attention. Msystems. 2023;8(4):e00345–23. doi: 10.1128/msystems.00345-23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 263.Yan J., Bhadra P., Li A., Sethiya P., Qin L., Tai H.K., Wong K.H., Siu S.W. Deep-AmPEP30: improve short antimicrobial peptides prediction with deep learning. Mol Therapy-Nucleic Acids. 2020;20:882–894. doi: 10.1016/j.omtn.2020.05.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 264.Manavalan B., Shin T.H., Kim M.O., Lee G. AIPpred: sequence-based prediction of anti-inflammatory peptides using random forest. Front Pharmacol. 2018;9:276. doi: 10.3389/fphar.2018.00276. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 265.Nasiri F., Atanaki F.F., Behrouzi S., Kavousi K., Bagheri M. CpACpP: in silico cell-penetrating anticancer peptide prediction using a novel bioinformatics framework. ACS Omega. 2021;6(30):19846–19859. doi: 10.1021/acsomega.1c02569. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 266.Jan A., Hayat M., Wedyan M., Alturki R., Gazzawe F., Ali H., Alarfaj F.K. Target-AMP: Computational prediction of antimicrobial peptides by coupling sequential information with evolutionary profile. Comput Biol Med. 2022;151 doi: 10.1016/j.compbiomed.2022.106311. [DOI] [PubMed] [Google Scholar]
  • 267.Chou K.-C. Prediction of signal peptides using scaled window. Peptides. 2001;22(12):1973–1979. doi: 10.1016/s0196-9781(01)00540-x. [DOI] [PubMed] [Google Scholar]
  • 268.Lv H., Yan K., Guo Y., Zou Q., Hesham A.E.-L., Liu B. AMPpred-EL: An effective antimicrobial peptide prediction model based on ensemble learning. Comput Biol Med. 2022;146 doi: 10.1016/j.compbiomed.2022.105577. [DOI] [PubMed] [Google Scholar]
  • 269.Chen Z., Zhao P., Li F., Leier A., Marquez-Lago T.T., Wang Y., Webb G.I., Smith A.I., Daly R.J., Chou K.-C., et al. Ifeature: a python package and web server for features extraction and selection from protein and peptide sequences. Bioinformatics. 2018;34(14):2499–2502. doi: 10.1093/bioinformatics/bty140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 270.Ali F., Hayat M. Machine learning approaches for discrimination of extracellular matrix proteins using hybrid feature space. J Theoret Biol. 2016;403:30–37. doi: 10.1016/j.jtbi.2016.05.011. [DOI] [PubMed] [Google Scholar]
  • 271.Bhasin M., Raghava G.P. Classification of nuclear receptors based on amino acid composition and dipeptide composition. J Biol Chem. 2004;279(22):23262–23266. doi: 10.1074/jbc.M401932200. [DOI] [PubMed] [Google Scholar]
  • 272.Chen Z., Zhou Y., Song J., Zhang Z. hCKSAAP_UbSite: improved prediction of human ubiquitination sites by exploiting amino acid pattern and properties. Biochim Biophys Acta (BBA)-Proteins Proteom. 2013;1834(8):1461–1467. doi: 10.1016/j.bbapap.2013.04.006. [DOI] [PubMed] [Google Scholar]
  • 273.Shen J., Zhang J., Luo X., Zhu W., Yu K., Chen K., Li Y., Jiang H. Predicting protein–protein interactions based only on sequences information. Proc Natl Acad Sci. 2007;104(11):4337–4341. doi: 10.1073/pnas.0607879104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 274.Youmans M., Spainhour J.C., Qiu P. Classification of antibacterial peptides using long short-term memory recurrent neural networks. IEEE/ACM Trans Comput Biol Bioinform. 2019;17(4):1134–1140. doi: 10.1109/TCBB.2019.2903800. [DOI] [PubMed] [Google Scholar]
  • 275.Pinacho-Castellanos S.A., García-Jacas C.R., Gilson M.K., Brizuela C.A. Alignment-free antimicrobial peptide predictors: improving performance by a thorough analysis of the largest available data set. J Chem Inf Model. 2021;61(6):3141–3157. doi: 10.1021/acs.jcim.1c00251. [DOI] [PubMed] [Google Scholar]
  • 276.Ruiz-Blanco Y.B., Paz W., Green J., Marrero-Ponce Y. ProtDCal: A program to compute general-purpose-numerical descriptors for sequences and 3D-structures of proteins. BMC Bioinformatics. 2015;16(1):162. doi: 10.1186/s12859-015-0586-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 277.Wang B., Lin P., Zhong Y., Tan X., Shen Y., Huang Y., Jin K., Zhang Y., Zhan Y., Shen D., et al. Explainable deep learning and virtual evolution identifies antimicrobial peptides with activity against multidrug-resistant human pathogens. Nat Microbiol. 2025:1–16. doi: 10.1038/s41564-024-01907-3. [DOI] [PubMed] [Google Scholar]
  • 278.Yao L., Zhang Y., Li W., Chung C.-R., Guan J., Zhang W., Chiang Y.-C., Lee T.-Y. DeepAFP: an effective computational framework for identifying antifungal peptides based on deep learning. Prot Sci. 2023;32(10) doi: 10.1002/pro.4758. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 279.Xu J., Li F., Li C., Guo X., Landersdorfer C., Shen H.-H., Peleg A.Y., Li J., Imoto S., Yao J., et al. iAMPCN: a deep-learning approach for identifying antimicrobial peptides and their functional activities. Brief Bioinform. 2023;24(4):bbad240. doi: 10.1093/bib/bbad240. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 280.Shen H., Li Y., Pi Q., Tian J., Xu X., Huang Z., Huang J., Pian C., Mao S. Unveiling novel antimicrobial peptides from the ruminant gastrointestinal microbiomes: A deep learning-driven approach yields an anti-MRSA candidate. J Adv Res. 2025 doi: 10.1016/j.jare.2025.01.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 281.Tang W., Dai R., Yan W., Zhang W., Bin Y., Xia E., Xia J. Identifying multi-functional bioactive peptide functions using multi-label deep learning. Brief Bioinform. 2022;23(1):bbab414. doi: 10.1093/bib/bbab414. [DOI] [PubMed] [Google Scholar]
  • 282.Xing W., Zhang J., Li C., Huo Y., Dong G. iAMP-attenpred: a novel antimicrobial peptide predictor based on bert feature extraction method and CNN-BiLSTM-attention combination model. Brief Bioinform. 2024;25(1):bbad443. doi: 10.1093/bib/bbad443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 283.Zhang Y., Lin J., Zhao L., Zeng X., Liu X. A novel antibacterial peptide recognition algorithm based on BERT. Brief Bioinform. 2021;22(6):bbab200. doi: 10.1093/bib/bbab200. [DOI] [PubMed] [Google Scholar]
  • 284.Lee H., Lee S., Lee I., Nam H. AMP-BERT: Prediction of antimicrobial peptide function based on a s model. Prot Sci. 2023;32(1) doi: 10.1002/pro.4529. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 285.Henikoff S., Henikoff J.G. Amino acid substitution matrices from protein blocks. Proc Natl Acad Sci. 1992;89(22):10915–10919. doi: 10.1073/pnas.89.22.10915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 286.Sandberg M., Eriksson L., Jonsson J., Sjöström M., Wold S. New chemical descriptors relevant for the design of biologically active peptides. a multivariate characterization of 87 amino acids. J Med Chem. 1998;41(14):2481–2491. doi: 10.1021/jm9700575. [DOI] [PubMed] [Google Scholar]
  • 287.Kawashima S., Kanehisa M. AAindex: amino acid index database. Nucleic Acids Res. 2000;28(1):374. doi: 10.1093/nar/28.1.374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 288.Jones D.T. Protein secondary structure prediction based on position-specific scoring matrices. J Mol Biol. 1999;292(2):195–202. doi: 10.1006/jmbi.1999.3091. [DOI] [PubMed] [Google Scholar]
  • 289.Wang Z., Wang G. APD: the antimicrobial peptide database. Nucleic Acids Res. 2004;32(suppl_1):D590–D592. doi: 10.1093/nar/gkh025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 290.Wang G., Li X., Wang Z. APD2: the updated antimicrobial peptide database and its application in peptide design. Nucleic Acids Res. 2009;37(suppl_1):D933–D937. doi: 10.1093/nar/gkn823. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 291.Jhong J.-H., Yao L., Pang Y., Li Z., Chung C.-R., Wang R., Li S., Li W., Luo M., Ma R., et al. dbAMP 2.0: updated resource for antimicrobial peptides with an enhanced scanning method for genomic and proteomic data. Nucleic Acids Res. 2022;50(D1):D460–D470. doi: 10.1093/nar/gkab1080. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 292.Yao L., Guan J., Xie P., Chung C.-R., Zhao Z., Dong D., Guo Y., Zhang W., Deng J., Pang Y., et al. dbAMP 3.0: updated resource of antimicrobial activity and structural annotation of peptides in the post-pandemic era. Nucleic Acids Res. 2025;53(D1):D364–D376. doi: 10.1093/nar/gkae1019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 293.Fan L., Sun J., Zhou M., Zhou J., Lao X., Zheng H., Xu H. DRAMP: a comprehensive data repository of antimicrobial peptides. Sci Rep. 2016;6(1):24482. doi: 10.1038/srep24482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 294.Kang X., Dong F., Shi C., Liu S., Sun J., Chen J., Li H., Xu H., Lao X., Zheng H. DRAMP 2.0, an updated data repository of antimicrobial peptides. Sci Data. 2019;6(1):148. doi: 10.1038/s41597-019-0154-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 295.Shi G., Kang X., Dong F., Liu Y., Zhu N., Hu Y., Xu H., Lao X., Zheng H. DRAMP 3.0: an enhanced comprehensive data repository of antimicrobial peptides. Nucleic Acids Res. 2022;50(D1):D488–D496. doi: 10.1093/nar/gkab651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 296.Ma T., Liu Y., Yu B., Sun X., Yao H., Hao C., Li J., Nawaz M., Jiang X., Lao X., et al. DRAMP 4.0: an open-access data repository dedicated to the clinical translation of antimicrobial peptides. Nucleic Acids Res. 2025;53(D1):D403–D410. doi: 10.1093/nar/gkae1046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 297.Nielsen S.D., Beverly R.L., Qu Y., Dallas D.C. Milk bioactive peptide database: A comprehensive database of milk protein-derived bioactive peptides and novel visualization. Food Chem. 2017;232:673–682. doi: 10.1016/j.foodchem.2017.04.056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 298.Nielsen S.D.-H., Liang N., Rathish H., Kim B.J., Lueangsakulthai J., Koh J., Qu Y., Schulz H.-J., Dallas D.C. Bioactive milk peptides: An updated comprehensive overview and database. Crit Rev Food Sci Nutr. 2024;64(31):11510–11529. doi: 10.1080/10408398.2023.2240396. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 299.Mondal R.K., Sen D., Arya A., Samanta S.K. Developing anti-microbial peptide database version 1 to provide comprehensive and exhaustive resource of manually curated AMPs. Sci Rep. 2023;13(1):17843. doi: 10.1038/s41598-023-45016-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 300.Waghu F.H., Gopi L., Barai R.S., Ramteke P., Nizami B., Idicula-Thomas S. CAMP: Collection of sequences and structures of antimicrobial peptides. Nucleic Acids Res. 2014;42(D1):D1154–D1158. doi: 10.1093/nar/gkt1157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 301.Thomas S., Karnik S., Barai R.S., Jayaraman V.K., Idicula-Thomas S. CAMP: a useful resource for research on antimicrobial peptides. Nucleic Acids Res. 2010;38(suppl_1):D774–D780. doi: 10.1093/nar/gkp1021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 302.Waghu F.H., Barai R.S., Gurung P., Idicula-Thomas S. CAMPR3: a database on sequences, structures and signatures of antimicrobial peptides. Nucleic Acids Res. 2016;44(D1):D1094–D1097. doi: 10.1093/nar/gkv1051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 303.Waghu F.H., Idicula-Thomas S. Collection of antimicrobial peptides database and its derivatives: Applications and beyond. Prot Sci. 2020;29(1):36–42. doi: 10.1002/pro.3714. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 304.Gawde U., Chakraborty S., Waghu F.H., Barai R.S., Khanderkar A., Indraguru R., Shirsat T., Idicula-Thomas S. CAMPR4: a database of natural and synthetic antimicrobial peptides. Nucleic Acids Res. 2023;51(D1):D377–D383. doi: 10.1093/nar/gkac933. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 305.Lakin S.M., Dean C., Noyes N.R., Dettenwanger A., Ross A.S., Doster E., Rovira P., Abdo Z., Jones K.L., Ruiz J., et al. MEGARes: an antimicrobial resistance database for high throughput sequencing. Nucleic Acids Res. 2017;45(D1):D574–D580. doi: 10.1093/nar/gkw1009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 306.Doster E., Lakin S.M., Dean C.J., Wolfe C., Young J.G., Boucher C., Belk K.E., Noyes N.R., Morley P.S. MEGARes 2.0: a database for classification of antimicrobial drug, biocide and metal resistance determinants in metagenomic sequence data. Nucleic Acids Res. 2020;48(D1):D561–D569. doi: 10.1093/nar/gkz1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 307.Bonin N., Doster E., Worley H., Pinnell L.J., Bravo J.E., Ferm P., Marini S., Prosperi M., Noyes N., Morley P.S., et al. MEGARes and AMR++, v3. 0: an updated comprehensive database of antimicrobial resistance determinants and an improved software pipeline for classification using high-throughput sequencing. Nucleic Acids Res. 2023;51(D1):D744–D752. doi: 10.1093/nar/gkac1047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 308.Mhade S., Panse S., Tendulkar G., Awate R., Narasimhan Y., Kadam S., Yennamalli R.M., Kaushik K.S. AMPing up the search: A structural and functional repository of antimicrobial peptides for biofilm studies, and a case study of its application to corynebacterium striatum, an emerging pathogen. Front Cell Infect Microbiol. 2021;11 doi: 10.3389/fcimb.2021.803774. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 309.Ravichandran S., Avatapalli S., Narasimhan Y., Kaushik K.S., Yennamalli R.M. ’Targeting’the search: An upgraded structural and functional repository of antimicrobial peptides for biofilm studies (b-AMP v2. 0) with a focus on biofilm protein targets. Front Cell Infect Microbiol. 2022;12 doi: 10.3389/fcimb.2022.1020391. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 310.Gogoladze G., Grigolava M., Vishnepolsky B., Chubinidze M., Duroux P., Lefranc M.-P., Pirtskhalava M. DBAASP: database of antimicrobial activity and structure of peptides. FEMS Microbiol Lett. 2014;357(1):63–68. doi: 10.1111/1574-6968.12489. [DOI] [PubMed] [Google Scholar]
  • 311.Pirtskhalava M., Gabrielian A., Cruz P., Griggs H.L., Squires R.B., Hurt D.E., Grigolava M., Chubinidze M., Gogoladze G., Vishnepolsky B., et al. DBAASP v. 2: an enhanced database of structure and antimicrobial/cytotoxic activity of natural and synthetic peptides. Nucleic Acids Res. 2016;44(D1):D1104–D1112. doi: 10.1093/nar/gkv1174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 312.Pirtskhalava M., Amstrong A.A., Grigolava M., Chubinidze M., Alimbarashvili E., Vishnepolsky B., Gabrielian A., Rosenthal A., Hurt D.E., Tartakovsky M. DBAASP v3: database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Res. 2021;49(D1):D288–D297. doi: 10.1093/nar/gkaa991. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 313.Kumar R., Chaudhary K., Sharma M., Nagpal G., Chauhan J.S., Singh S., Gautam A., Raghava G.P. AHTPDB: a comprehensive platform for analysis and presentation of antihypertensive peptides. Nucleic Acids Res. 2015;43(D1):D956–D962. doi: 10.1093/nar/gku1141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 314.Tyagi A., Tuknait A., Anand P., Gupta S., Sharma M., Mathur D., Joshi A., Singh S., Gautam A., Raghava G.P. CancerPPD: a database of anticancer peptides and proteins. Nucleic Acids Res. 2015;43(D1):D837–D843. doi: 10.1093/nar/gku892. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 315.Aguilera-Mendoza L., Marrero-Ponce Y., Beltran J.A., Tellez Ibarra R., Guillen-Ramirez H.A., Brizuela C.A. Graph-based data integration from bioactive peptide databases of pharmaceutical interest: toward an organized collection enabling visual network analysis. Bioinformatics. 2019;35(22):4739–4747. doi: 10.1093/bioinformatics/btz260. [DOI] [PubMed] [Google Scholar]
  • 316.Ayala-Ruano S., Marrero-Ponce Y., Aguilera-Mendoza L., Perez N., Aguero-Chapin G., Antunes A., Aguilar A.C. Network science and group fusion similarity-based searching to explore the chemical space of antiparasitic peptides. ACS Omega. 2022;7(50):46012–46036. doi: 10.1021/acsomega.2c03398. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 317.Mehta D., Anand P., Kumar V., Joshi A., Mathur D., Singh S., Tuknait A., Chaudhary K., Gautam S.K., Gautam A., et al. ParaPep: a web resource for experimentally validated antiparasitic peptide sequences and their structures. Database. 2014;2014:bau051. doi: 10.1093/database/bau051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 318.Piotto S.P., Sessa L., Concilio S., Iannelli P. YADAMP: yet another database of antimicrobial peptides. Int J Antimicro Ag. 2012;39(4):346–351. doi: 10.1016/j.ijantimicag.2011.12.003. [DOI] [PubMed] [Google Scholar]
  • 319.Wang S., Peng J., Ma J., Xu J. Protein secondary structure prediction using deep convolutional neural fields. Sci Rep. 2016;6(1):18962. doi: 10.1038/srep18962. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 320.Gainza P., Sverrisson F., Monti F., Rodola E., Boscaini D., Bronstein M.M., Correia B.E. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods. 2020;17(2):184–192. doi: 10.1038/s41592-019-0666-6. [DOI] [PubMed] [Google Scholar]
  • 321.Xiao X., Shao S., Ding Y., Huang Z., Chen X., Chou K.-C. Using cellular automata to generate image representation for biological sequences. Amino Acids. 2005;28(1):29–35. doi: 10.1007/s00726-004-0154-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 322.Xiao X., Lin W.-Z., Chou K.-C. Recent advances in predicting protein classification and their applications to drug development. Curr Top Med Chem. 2013;13(14):1622–1635. doi: 10.2174/15680266113139990113. [DOI] [PubMed] [Google Scholar]
  • 323.Shao J., Zhao Y., Wei W., Vaisman I.I. AGRAMP: machine learning models for predicting antimicrobial peptides against phytopathogenic bacteria. Front Microbiol. 2024;15 doi: 10.3389/fmicb.2024.1304044. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 324.Teimouri H., Medvedeva A., Kolomeisky A.B. Bacteria-specific feature selection for enhanced antimicrobial peptide activity predictions using machine-learning methods. J Chem Inf Model. 2023;63(6):1723–1733. doi: 10.1021/acs.jcim.2c01551. [DOI] [PubMed] [Google Scholar]
  • 325.Hussain W. sAMP-PFPDeep: Improving accuracy of short antimicrobial peptides prediction using three different sequence encodings and deep neural networks. Brief Bioinform. 2022;23(1):bbab487. doi: 10.1093/bib/bbab487. [DOI] [PubMed] [Google Scholar]
  • 326.LeCun Y., Bottou L., Bengio Y., Haffner P. Gradient-based learning applied to document recognition. Proc IEEE. 1998;86(11):2278–2324. [Google Scholar]
  • 327.Graves A., Schmidhuber J. Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Netw. 2005;18(5–6):602–610. doi: 10.1016/j.neunet.2005.06.042. [DOI] [PubMed] [Google Scholar]
  • 328.Breiman L. Random forests. Mach Learn. 2001;45:5–32. [Google Scholar]
  • 329.Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser L., Polosukhin I. Advances in neural information processing systems (neurIPS) 2017. Attention is all you need; pp. 5998–6008. [Google Scholar]
  • 330.Mnih V, Heess N, Graves A, Kavukcuoglu K. Recurrent Models of Visual Attention. In: Conference on neural information processing systems, neurIPS. 2014, p. 2204–12.
  • 331.Hochreiter S., Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735–1780. doi: 10.1162/neco.1997.9.8.1735. [DOI] [PubMed] [Google Scholar]
  • 332.Cho K., van Merrienboer B., Gülçehre Ç., Bahdanau D., Bougares F., Schwenk H., Bengio Y. EMNLP. ACL; 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation; pp. 1724–1734. [Google Scholar]
  • 333.Devlin J., Chang M., Lee K., Toutanova K. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, NAACL-HLT. Association for Computational Linguistics; 2019. BERT: pre-training of deep bidirectional transformers for language understanding; pp. 4171–4186. [Google Scholar]
  • 334.Quinlan J.R. Induction of decision trees. Mach Learn. 1986;1:81–106. [Google Scholar]
  • 335.Holland J.H. MIT Press; 1992. Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence. [Google Scholar]
  • 336.Cabello D., Barro S., Salceda J., Ruiz R., Mira J. Fuzzy K-nearest neighbor classifiers for ventricular arrhythmia detection. Int J Bio-Med Comput. 1991;27(2):77–93. doi: 10.1016/0020-7101(91)90089-w. [DOI] [PubMed] [Google Scholar]
  • 337.Kingma DP, Welling M. Auto-Encoding Variational Bayes. In: 2nd international conference on learning representations, ICLR. 2014.
  • 338.Dreiseitl S., Ohno-Machado L. Logistic regression and artificial neural network classification models: a methodology review. J Biomed Inform. 2002;35(5–6):352–359. doi: 10.1016/s1532-0464(03)00034-0. [DOI] [PubMed] [Google Scholar]
  • 339.Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville AC, Bengio Y. Generative Adversarial Nets. In: Conference on neural information processing systems, neurIPS. 2014, p. 2672–80.
  • 340.Goodfellow I., Pouget-Abadie J., Mirza M., Xu B., Warde-Farley D., Ozair S., Courville A., Bengio Y. Generative adversarial networks. Commun ACM. 2020;63(11):139–144. [Google Scholar]
  • 341.Chen T., Guestrin C. International conference on knowledge discovery and data mining, SIGKDD. ACM; 2016. XGBoost: A scalable tree boosting system; pp. 785–794. [Google Scholar]
  • 342.Ke G., Meng Q., Finley T., Wang T., Chen W., Ma W., Ye Q., Liu T.-Y. Lightgbm: A highly efficient gradient boosting decision tree. Conf Neural Inf Process Syst, NeurIPS. 2017;30 [Google Scholar]
  • 343.Kipf T.N., Welling M. International conference on learning representations, ICLR. OpenReview.net; 2017. Semi-supervised classification with graph convolutional networks. [Google Scholar]
  • 344.Tolstikhin I.O., Bousquet O., Gelly S., Schölkopf B. 2017. Wasserstein auto-encoders. CoRR, abs/1711.01558. [Google Scholar]
  • 345.Jordan M.I. vol. 8. 1986. Attractor dynamics and parallelism in a connectionist sequential machine. (Proceedings of the annual meeting of the cognitive science society). [Google Scholar]
  • 346.Donahue J., Krähenbühl P., Darrell T. 2016. Adversarial feature learning. arXiv preprint arXiv:1605.09782. [Google Scholar]

Articles from Synthetic and Systems Biotechnology are provided here courtesy of KeAi Publishing

RESOURCES