Abstract
Understanding how molecular representations encode structure–property relationships is a central challenge in chemoinformatics, particularly for complex biomolecular systems such as antimicrobial peptides (AMPs). Although numerous computational models have been developed to predict peptide hemolysis, less attention has been given to how different descriptor representations influence both predictive robustness and mechanistic interpretability. Here, we present a comparative computational analysis of sequence-derived and structure-based molecular descriptors to identify the physicochemical properties governing AMP-induced hemolysis. Our analysis identifies a reduced set of key descriptors that preserve the predictive performance of the process. It shows that toxicity is primarily associated with hydrophobic clustering, amphipathic polarity patterning, solvent accessibility, and specific dipeptide motifs, whereas reduced toxicity correlates with higher aggregation propensity and earlier accumulation of polarizable residues. Complementary molecular descriptors suggest that periodic organization of electronic and aromatic properties and localized charge distributions contribute to membrane-disruptive behavior. These findings demonstrate how the representation choice might provide mechanistic insights and guiding principles for descriptor-based analysis and rational design of selective antimicrobial peptides.
Introduction
Understanding how molecular structure determines biological activity remains a central goal in studies of biological processes and molecular design. Machine learning approaches have achieved substantial success in predicting important contributions of multiple molecular properties, including bioactivity, toxicity, and physicochemical behavior. − Despite these advances, a persistent challenge is the interpretability of such models, particularly in complex biomolecular systems where predictive accuracy alone provides limited insight into underlying structure–property relationships. In many situations, the machine-learning predictions are not able to connect these properties with molecular processes.
Antimicrobial peptides (AMPs) provide a particularly valuable model system for addressing this challenge. These predominantly cationic and amphipathic biomolecules exhibit broad-spectrum activity against diverse pathogens, − yet their clinical application is limited by hemolytic toxicity toward host cells. , Peptide hemolysis describes the rupture of membranes of red blood cells, causing them to release hemoglobin, which contributes to the toxicity of these biomolecules and limits their applications. Understanding the physicochemical features that distinguish toxic from nontoxic peptides is therefore essential for designing selective AMPs with improved therapeutic potential.
Most computational studies of AMP toxicity rely primarily on sequence-derived descriptors, such as amino acid composition, dipeptide motifs, hydrophobicity indices, and charge distributions. While these descriptors have enabled more interpretable classification models and provided some insights into toxicity-associated sequence motifs − as well as hydrophobicity and charge patterns linked to hemolytic activity, − sequence-only representations do not directly capture molecular characteristics such as steric organization, electronic properties, or spatial charge distributions and other types of intramolecular correlations. Structure-based molecular descriptors, widely used in other areas of chemoinformatics, can encode such information, but their systematic integration into AMP toxicity studies remains limited.
Recent studies integrating sequence and structural representations demonstrate that different descriptor types provide complementary information and can improve the prediction of peptide properties. − For example, diverse representations including sequence-derived descriptors, structural descriptors, and graph-based encodings have been successfully applied to model peptide permeability. Furthermore, simple domain-specific molecular representations such as fingerprints have been shown to capture key biochemical features of peptides and to perform as well as or better than complex deep learning models. These findings suggest that short-range structural patterns and local physicochemical characteristics together may be sufficient to encode essential functional information about biomolecules. However, it remains unclear how different levels of molecular descriptions systematically influence both predictive performance and mechanistic interpretability.
In this work, we present a systematic comparative analysis of sequence-derived and structure-based molecular descriptors for predicting and interpreting AMP-induced hemolysis. Consistent feature-selection procedures are applied across representations to identify the most informative descriptors. Through this analysis, we aim to clarify how different molecular representations illuminate complementary aspects of peptide behavior, to identify the physicochemical properties most strongly associated with high toxicity or reduced toxicity, and to provide general insights into the role of descriptor representations in chemoinformatics modeling.
Methods
Overview
We analyzed antimicrobial peptide (AMP) hemolytic toxicity using a comparative descriptor-based framework, shown in Figure . A previous data set of toxic and nontoxic AMPs was used to construct three complementary molecular representations: sequence-derived descriptors, structure-based molecular descriptors, and their combined feature set. Each descriptor set was processed using identical preprocessing steps, including normalization to a [0,1] range and stratified partitioning into training and held-out testing data sets. After preprocessing, two feature-selection methods were applied to each representation to identify subsets of descriptors most strongly associated with toxicity. The predictive performance of these reduced feature sets was then evaluated and compared with that of the full descriptor sets to confirm their relevance for distinguishing toxic and nontoxic AMPs. Finally, the most influential descriptors from each representation were examined and interpreted in terms of their physicochemical meaning and potential roles in the mechanisms underlying AMP-induced hemolysis.
2.
Schematic overview of the computational workflow. AMP hemolysis data were represented using sequence-derived descriptors, molecular descriptors, and their combined feature set. After preprocessing and stratified data partitioning, feature-selection methods were applied to identify descriptors most strongly associated with toxicity. Classification models were then trained on both full and reduced feature sets and evaluated using multiple performance metrics. The most influential descriptors were subsequently interpreted to provide physicochemical and mechanistic insight into AMP-induced hemolysis.
Data Set
The data set, obtained from Khabbaz et al. and originally sourced from a database of AMPs (DBAASP , ), consists of AMP sequences with associated hemolysis measurements and binary toxicity labels (toxic = 1, nontoxic = 0) assigned according to previously defined hemolysis thresholds (see Supporting Information).
Figure presents examples of hemolytic and nonhemolytic peptides derived from wasps. Figure (a) shows the toxic 12-residue linear peptide Orancis-protonectin with AlphaFold , -predicted peptide structure (residues 57–68 in AF–P86147-F1-v6; UniprotKB: P86147), peptide sequence, and chemical structure showing the amidated C-terminus. Figure (b) shows the nontoxic 12-residue linear peptide Protonectin peptide structure (PDB6N68 ), peptide sequence, and chemical structure showing the amidated C-terminus. Similarities between the sequences are shown in green and underlined, and similarities between structures are shown in green boxes. Chemical structures were generated by first converting sequence to SMILES in the PepSMI tool by NovoProLabs (https://www.novoprolabs.com/tools/convert-peptide-to-smiles-string), then converting SMILES to chemical structures in ChemDraw (Revvity Signals, version 22.2.0). The two peptides differ in sequence by only four amino acids at different positions within the sequence, and their peptide structures are both α-helical, but despite these similarities, one peptide is hemolytic and the other is nonhemolytic, further illustrating the need for identifying physicochemical features that can distinguish between them.
1.

Examples of (a) hemolytic and (b) nonhemolytic peptides derived from wasps showing peptide structure, peptide sequence, and chemical formula.
The curated data set includes only peptides composed of canonical amino acids with acetyl or amide terminal groups, reported hemolysis values in consistent units (μM), and sequence lengths between 6 and 50 amino acids.
The original data set was used without modification. However, 411 peptides were excluded because one or more descriptors could not be computed from their sequences. After descriptor generation, 2418 peptides remained for analysis, including 1320 toxic and 1098 nontoxic sequences Figure .
Descriptor Generation
Sequence-Derived Descriptors (ProPy and Prior-Study Variables)
A total of 1547 sequence-derived descriptors were computed using the ProPy computational tool, including amino acid composition patterns, physicochemical indices, sequence-order features, and autocorrelation descriptors. In addition, 15 sequence-related variables from the data set that are not generated by ProPy were also included: sequence length, normalized hydrophobic moment, normalized hydrophobicity, N- and C-terminal indicators, net charge, isoelectric point, charge density, disorder, steric hindrance, solvation, hydropathy, amphiphilicity, and aggregation propensity in vitro and in vivo. The resulting sequence-derived descriptor matrix contained 1562 properties.
Molecular (Mordred) Descriptors
To capture molecular composition, topology, and electronic properties, each peptide sequence was converted to a SMILES representation using the cheminformatics package RDKit. Mordred descriptors were then computed from these SMILES structures. Because three-dimensional coordinates were not generated, many Mordred descriptors that relied on these coordinates were inapplicable and produced missing values. These descriptors, along with columns containing non-numerical values, were removed. After column removal, the molecular descriptor matrix contained 1437 descriptors.
Combined Descriptors
For the combined representation, the sequence-derived and molecular descriptor matrices were aligned by peptide sequence and concatenated column-wise, yielding 2999 descriptors.
Preprocessing
Each descriptor set was partitioned into training, test, and validation sets using a stratified 60/20/20 split to preserve proportions of toxic versus nontoxic peptides: 60% for training, 20% for test, and 20% for validation. Each descriptor column in the training set was scaled to [0,1] by min–max normalization. For a given descriptor value x, the normalized value x′ was computed as
| 1 |
where x min and x max are the minimum and maximum values of that descriptor across the data set. This normalization ensures that all descriptors are directly comparable in the analysis. The scaler was fit on the training set only and subsequently applied to the validation and test sets.
Feature-Selection Methods and Prediction Models
To ensure that the results were not specific to a single analytical approach, two independent feature-selection strategies were applied to identify reduced subsets of informative physicochemical features: a linear support vector machine (SVM) , and a Random Forest (RF) selector. In addition, we evaluated four commonly used prediction models: logistic regression, linear SVM, SVM with an RBF kernel, and SVM with a polynomial kernel. , Performance of the prediction models was evaluated using various statistical metrics such as accuracy, Matthews correlation coefficient (MCC), F1 score, precision, and recall. Comparing multiple models and selection strategies allows us to determine whether reducing the number of descriptors consistently preserves predictive performance across diverse analytical conditions.
Hyperparameter Tuning
For SVM, we tuned the hyperparameter C that influences the number of features selected. Smaller values of C enforce stronger sparsity, while for larger values of C, more features are allowed to remain. We tuned C by 5-fold cross-validation on the training set over C ∈ {0.01, 0.1, 1, 10}, choosing the value that maximized cross-validated accuracy. The best values were obtained for C = 10 for sequence-derived features, for C = 1 for molecular features, and for C = 10 for combined features.
Five-Fold Feature Overlap (Stability Selection)
To identify physicochemical descriptors that were robust to variations in the training data, feature selection was evaluated across multiple resampled subsets of the data set. Specifically, each selection procedure was repeated using a 5-fold stratified cross-validation applied to the training set, ensuring that each fold preserved the same proportion of toxic and nontoxic peptides. Within each fold, feature importance was computed independently using the respective selection method.
To obtain a stable set of predictors, only descriptors that were consistently selected across all five folds were retained. This criterion ensured that the final feature set reflected physicochemical properties that repeatedly contributed to toxicity discrimination, rather than those selected due to random variations in the data. For descriptors retained under this stability criterion, importance values were summarized across folds using the mean and the standard error, providing an estimate of both the average contribution and its variability. This procedure yielded a reduced set of robust, interpretable features that were used for subsequent modeling and analysis.
Normalization of Feature Coefficients for Visualization
For interpretability plots, SVM coefficients mean values were min–max normalized to [−1,1] to preserve directionality (positive values associated with the toxic class; negative with the nontoxic class). RF mean importance parameters were normalized to [0,1] because they do not contain directionality information. For comparability, SE values were normalized in proportion to the range of the mean importance values. Smaller SE values correspond to more stable features across folds, while larger SE values indicate greater variability in importance depending on data partitioning.
Additional Data Set
We included a data set containing secondary structure information for an independent analysis of the relationship between peptide secondary structure and hemolysis, and additionally as an independent test set for external validation. This data set, originally obtained from the Antimicrobial Peptide Database (APD), contained 3081 natural peptides with diverse structures and sources that were labeled hemolytic or nonhemolytic. It is important to note that in the description of the data set the criteria used to define peptides as hemolytic versus nonhemolytic did not appear to be specified. Therefore, it is unclear whether these criteria are consistent with those used for the primary data set analyzed in the current study.
Analysis of Secondary Structure
For the analysis of the relationship between secondary structure and hemolysis, peptides known secondary structure were retained. APD classifies peptide secondary structure into four categories: α helix, β-sheet, combined α-β, and nonalpha-β, the latter including other conformations such as random coils. Peptides with unknown conformations were excluded. 1840 peptides with unknown secondary structures were excluded from the analysis, leaving 1106 nonhemolytic peptides and 135 hemolytic peptides.
The relative frequency of each structural category was calculated separately for hemolytic and nonhemolytic peptide data sets as the percentage of peptides assigned to each category. Comparisons between hemolytic and nonhemolytic groups were performed independently for each structural category using chi-square tests of independence on 2 × 2 contingency tables constructed from observed category counts.
External Validation
The top-performing models and feature subsets identified from the main data set (Figure ) were selected for external validation. Of the 3081 peptides in the external APD-derived data set, most descriptors could be calculated for 2950, including 2646 nontoxic peptides and 304 toxic peptides. However, even for these 2950 peptides, several descriptors that were successfully calculated for the main data set could not be generated. Consequently, only the subset of descriptors that could be calculated consistently across the external data set was used for prediction (see feature set in Table ).
4.

Effect of reducing the number of physicochemical descriptors on prediction of antimicrobial peptide toxicity. The heatmap shows changes in accuracy (ΔAcc) and recall (ΔRec) relative to using the full descriptor set (B) for four prediction models (M1–M4: LR, linear SVM, SVM-RBF, SVM-polynomial). Reduced descriptor subsets were obtained using two independent selection methods (F1: SVM and F2: RF) or their intersection (F1 ∩ F2). Boxes highlight cases where both accuracy and recall are maintained or improved relative to the full descriptor set.
1. Performance (Accuracy, MCC, F1, Precision, Recall) of the Training and External-Validation Independent Test Set across the Top-Performing Models from the Main Data Set (see Figure ).
| feature set | model | accuracy | MCC | F1 | precision | recall |
|---|---|---|---|---|---|---|
| Sequencex SVM-RBF | ||||||
| Baseline (1547) | train | 0.80 | 0.60 | 0.82 | 0.79 | 0.86 |
| val | 0.43 | 0.15 | 0.24 | 0.14 | 0.86 | |
| SVM (122) | train | 0.82 | 0.63 | 0.84 | 0.80 | 0.88 |
| val | 0.43 | 0.14 | 0.23 | 0.13 | 0.83 | |
| Sequencex SVM-poly | ||||||
| Baseline (1547) | train | 0.90 | 0.81 | 0.91 | 0.91 | 0.92 |
| val | 0.45 | 0.17 | 0.24 | 0.14 | 0.86 | |
| RF (418) | train | 0.96 | 0.92 | 0.96 | 0.95 | 0.98 |
| val | 0.50 | 0.19 | 0.26 | 0.15 | 0.84 | |
| Molecularx LR | ||||||
| Baseline (1434) | train | 0.77 | 0.53 | 0.80 | 0.76 | 0.83 |
| val | 0.46 | 0.15 | 0.24 | 0.14 | 0.82 | |
| RF (383) | train | 0.76 | 0.52 | 0.79 | 0.76 | 0.83 |
| val | 0.45 | 0.15 | 0.24 | 0.14 | 0.84 | |
Sequence-level features selected from the sequence descriptor set.
Support Vector Machine with radial basis function (RBF) kernel.
Support Vector Machine with polynomial kernel.
Molecular-level features selected from the sequence descriptor set.
Logistic Regression.
The same training set used for model development, testing, and validation on the main data set was retained for external application. However, because not all descriptors computed for the main data set could be generated for the external data set, the feature space was restricted to the subset of descriptors available in both data sets. The scaler was therefore fitted using only this reduced descriptor set from the original training data, and the top-performing models were retrained accordingly using the same subset of features. Apart from this restriction to shared descriptors, no information from the external data set was used during model training or parameter tuning. The external data set was used exclusively for model evaluation after all preprocessing and retraining steps were completed.
Results and Discussion
Analysis of Secondary Structure
Figure shows the distribution of peptide secondary structure classes for hemolytic and nonhemolytic peptides, with significant differences between peptide sets in α-helical and nonalpha-β structures. For hemolytic peptides, there was a significantly higher frequency of α-helical secondary structures compared to nonhemolytic peptides. For nonhemolytic peptides, there was a significantly higher frequency of nonalpha-β secondary structures compared to hemolytic peptides. There was no significant difference between the two sets of peptides in β-sheet or combined α-β secondary structures.
3.
Distribution of peptide secondary structure classes in hemolytic and nonhemolytic peptide data sets. Bars represent the percentage of peptides assigned to each structural category. Red hatched bars indicate hemolytic peptides, while solid green bars indicate nonhemolytic peptides. Significance annotations are shown above bar pairs as asterisks (*=p < 0.05, **=p < 0.01, ***=p < 0.001) or n.s. (not significant).
Model Evaluation
Evaluation of prediction performance revealed that several combinations of models and reduced descriptor sets maintained or improved both accuracy and recall relative to the full feature set (Figure , boxed regions). These cases demonstrate that substantial reduction in number of descriptors can be achieved without sacrificing predictive performance. Notably, this occurred for molecular descriptors selected by RF (method F2) and prediction by LR (Model M1), sequence descriptors selected by SVM (method F1) and prediction by SVM-RBF (Model M3), and sequence descriptors selected by RF (method F2) and prediction by SVM-polynomial kernel (Model M4).
External Validation
External validation performance on the independent APD-derived data set is summarized in Table . The models selected from the main data set (see Figure ) maintained sensitivity to toxic peptides, with recall values ranging from 0.82 to 0.86 across sequence- and molecular-level feature sets. The sequence-level SVM-poly model using RF-selected features achieved the highest external validation accuracy (0.50) and MCC (0.19), demonstrating that informative patterns learned from the main data set were at least partially transferable to the external data set. Despite lower precision and F1 values compared with training performance, the models’ ability to detect toxic peptides in an independent set highlights their potential utility for general toxicity screening. Differences in validation metrics may reflect variations in how peptides were defined as toxic versus nontoxic in the external data set, as specific criteria for toxicity were not clearly reported. It is possible that thresholds for defining toxicity differed between the main and APD-derived data sets, meaning peptides considered toxic in one data set may not meet the same criteria in the other, and vice versa.
The lower performance metrics in the external validation findings underscore the importance of clearly reporting toxicity definitions in future data sets to facilitate model generalization. Even with these challenges, the models successfully identified a substantial proportion of toxic peptides, indicating that the top-performing feature sets and models retain predictive value beyond the original training set.
Selected Features
Our computational approach allowed us to identify the specific sequential and molecular physicochemical properties that are associated with the toxicity of AMPs. Table S5 in the Supporting Information shows the detailed explanations of all features discussed here.
Sequence-Derived Descriptors (ProPy and Prior-Study Variables)
Figure shows the five most important features that distinguish toxic from nontoxic peptides for the sequence-based set of descriptors. As shown in Figure (a), the strong positive weight of the CY dipeptide motif suggests that cysteine–tyrosine pairings may promote toxic activity, possibly by favoring structural motifs that stabilize disulfide bridges (via cysteine) or enhance aromatic stacking and membrane insertion (via tyrosine). Such combinations could increase peptide rigidity and membrane anchoring, thereby enhancing lytic potential. , Conversely, descriptors linked to van der Waals volume (NormalizedVDWVD1075, NormalizedVDWVD3050) and polarizability (PolarizabilityD3050, PolarizabilityD2025) were strongly negative, associating them with nontoxicity. Specifically, NormalizedVDWVD1075 records the relative sequence position at which 75% of low–van der Waals volume residues have been encountered, while NormalizedVDWVD3050 marks the position at which 50% of high–volume residues have appeared. Both were negatively weighted, indicating that peptides in which low-volume residues accumulate later, or high-volume residues cluster earlier, tend to be less toxic. Mechanistically, this pattern suggests that the spatial distribution of sterically bulky groups within the sequence influences peptide–membrane interactions: early clustering of bulky side chains or delayed presence of compact ones may hinder efficient amphipathic folding and insertion, thereby biasing peptides toward nontoxicity. It is important to note that variance was small for volume and polarizability but larger for CY, indicating that some toxic peptides had larger proportions of CY in their composition, while others had smaller proportions.
5.
Relative importance of sequence-derived physicochemical features distinguishing nontoxic (negative coefficients) from toxic (positive coefficients) determined from (a) SVM and (b) RF feature-selection methods. For each analysis, the top 5 features are displayed after sorting the features by their contribution magnitude.
As shown in Figure (b), the dominant role of the aggregationPropensityInVivo property is consistent with the idea that toxic peptides often self-associate, either preassembling in solution or clustering upon membrane binding. Aggregation can facilitate cooperative membrane disruption through pore formation or micellization. , Similarly, increased hydrophobicity (HydrophobicityC3, normalizedHydrophobicity) and solvent accessibility (SolventAccessibilityC1) support deeper partitioning into lipid bilayers, a prerequisite for destabilizing red blood cell membranes. Elevated polarity (PolarityC1), while seemingly counterintuitive, may promote amphipathic segregationthe alignment of polar and nonpolar residues into distinct faceswhich is a hallmark of hemolytic α-helices. It is important to note that there was moderate variance in all of these features, indicating that the exact characteristic values varied between folds.
Molecular (Mordred) Descriptors
Figure shows the five most important features that distinguish toxic from nontoxic peptides for the molecular set of descriptors. Among the top-ranked SVM features, shown in Figure (a), JGI1 showed the strongest positive contribution to toxicity. This descriptor reflects the mean topological charge index of order 1, capturing localized charge distribution across the molecular graph. A strong positive weight suggests that peptides with greater charge asymmetry are more likely to interact with negatively charged membranes, thereby enhancing toxicity. In contrast, features associated with fused ring structures contributed negatively. Both n9FRing and nFRing descriptors, which count nine-membered and general fused rings, respectively, were linked to nontoxicity, consistent with the idea that rigid cyclic structures reduce conformational adaptability and membrane insertion ability. − Additional nontoxic associations included VE2A, a valence connectivity descriptor reflecting branching and substitution patterns, and JGI10, the mean topological charge index of order 10, which captures long-range charge delocalization. The negative weights of these descriptors suggest that extensive branching and charge delocalization reduce electrostatic selectivity and hinder effective bilayer disruption, an observation which is supported by previous studies. −
6.
Relative importance of molecular physicochemical features distinguishing nontoxic (negative coefficients) from toxic (positive coefficients) determined from (a) SVM and (b) RF feature-selection methods. For each analysis, the top 5 features are displayed after sorting the features by their contribution magnitude.
The RF analysis emphasized a complementary set of descriptors, shown in Figure (b). Notably, several descriptors corresponded to autocorrelation functions at lag 6. ATSC6pe, MATS6pe, and ATSC6se descriptors reflect periodic correlations in Sanderson electronegativity or electronic state indices, while AATSC6are and MATS6are capture periodicity in aromaticity-related properties. The consistent identification of lag-6 descriptors suggests that periodic spatial patterns in electronic and aromatic properties are critical for distinguishing between toxic and nontoxic peptides. , Aromatic residues arranged with this periodicity may enhance membrane anchoring via π–π stacking or cation−π interactions, , while electronegativity periodicity may favor amphipathic segregation and insertion into lipid bilayers. ,
Combined Descriptors
When the sequence-derived and molecular descriptor sets were combined and subjected to feature selection, both SVM and RF methods highlighted descriptors consistent with our earlier findings. As shown in Figure , the top five combined features were all sequence-derived descriptors, reflecting that the dominant predictive signal was from sequence-derived features. This suggests that it may be valuable to consider sequence-derived and molecular descriptors separately to retain insights from both levels of feature representation.
7.
Relative importance of combined sequence-derived and molecular physicochemical features distinguishing nontoxic (negative coefficients) from toxic (positive coefficients) determined from (a) SVM and (b) RF feature-selection methods. For each analysis, the top 5 features are displayed after sorting the features by their contribution magnitude.
Among the top SVM-selected features, shown in Figure (a), CY was also present in the previous sequence-derived-only analysis. In addition, a polarizability distribution descriptor reappeared, but with a shift: in the previous analysis, PolarizabilityD2050 was selected, corresponding to the sequence position where 50% of medium-polarizability residues have appeared. In the combined analysis, PolarizabilityD3025 emerged instead, which marks the position of the first 25% of high-polarizability residues. Despite this change in the group and percentile captured, both descriptors carried the same interpretation: earlier accumulation of polarizable residues along the sequence was negatively associated with toxicity. This suggests that clustering of polarizable residues reduces the stability of amphipathic packing and dissipates membrane interaction, which could lead to reduced toxicity. Finally, the aggregationPropensityInVivo descriptor was previously in the RF subset but reappeared here in both the SVM and RF feature subsets.
Among new features, the QH property showed the strongest positive association with toxicity. This descriptor quantifies the fractional composition of dipeptides containing glutamine followed by histidine (Q–H) within a peptide sequence. Mechanistically, glutamine provides hydrogen-bonding capacity while histidine contributes pH-dependent charge, meaning that Q–H motifs may enhance peptide polarity and amphipathic organization, favoring stronger peptide–membrane interactions.
The MoranAutoHydrophobicity6 descriptor was strongly negatively weighted, suggesting that highly regular hydrophobic correlations bias toward nontoxicity. This reflects a broader pattern across analyses: periodic physicochemical patterning shapes amphipathic folding and can either promote or hinder effective bilayer disruption. In this case, greater regularity may reduce the flexibility required for membrane destabilization. The results of RF feature selection, shown in Figure (b), emphasized the same descriptors as in the sequence-derived-only analysis, but with shifts in rank and magnitude.
Summary and Conclusions
In this work, antimicrobial peptides were explored as a testing ground to examine how different descriptor families relate to both predictive performance and interpretability for specific functions. We compared sequence-derived (ProPy) and molecular (Mordred) descriptors, individually and in combination, after feature selection to investigate this question in detail. This analysis identified several recurring structure–activity relationships (SARs) that appear to govern the hemolytic toxicity of antimicrobial peptides (AMPs). Across multiple descriptor representations and feature-selection approaches, hemolytic activity was consistently associated with hydrophobic clustering, increased solvent accessibility, amphipathic polarity patterning, localized charge concentration, and specific sequence motifs, particularly CY (Cys–Tyr) and QH (Gln–His). These features collectively support a mechanistic model in which toxicity arises from enhanced membrane association, insertion, and disruption, driven by favorable hydrophobic and electrostatic interactions with erythrocyte membranes. The identification of topological charge descriptors and periodic electronic and aromatic patterns further suggests that not only overall composition but also the spatial organization of physicochemical properties contributes to membrane-disruptive behavior.
In contrast, reduced hemolytic toxicity was associated with earlier accumulation of highly polarizable residues, greater aggregation propensity, more regular hydrophobic periodicity, sterically unfavorable distributions of bulky residues, and molecular descriptors indicative of charge delocalization or increased structural rigidity. These features may interfere with the formation of optimal amphipathic structures, reduce selective membrane interactions, or limit the conformational flexibility required for efficient membrane insertion and lysis. Together, these observations suggest that peptide toxicity is influenced not only by the presence of specific physicochemical characteristics but also by their positional distribution and higher-order organization along the sequence.
Importantly, the findings highlight that sequence-derived descriptors capture much of the predictive signal underlying AMP hemolysis, whereas molecular descriptors provide complementary mechanistic insights into electronic, topological, and charge-related factors that may not be readily apparent from sequence information alone. The convergence of multiple descriptor families on common physicochemical themes strengthens confidence that hydrophobic organization, amphipathicity, charge distribution, and residue patterning are key determinants of hemolytic activity.
Nevertheless, several limitations should be considered when interpreting these results. The identified relationships were derived from a specific DBAASP-derived hemolysis data set and therefore represent statistical associations rather than experimentally validated causal mechanisms. In particular, the CY and QH motifs should be viewed as hypothesis-generating observations that warrant prospective experimental validation before being considered as general design rules. Furthermore, hemolysis represents only one aspect of peptide toxicity, the data set aggregates measurements obtained under heterogeneous experimental conditions, and the descriptor space explored here does not encompass all potentially relevant structural or biophysical features. Accordingly, the proposed SAR trends should be interpreted as guiding principles rather than universal predictors of AMP behavior.
Overall, the results support a unified SAR framework in which hemolytic toxicity is promoted by localized hydrophobic and electrostatic organization that facilitates membrane disruption, whereas features that disperse these interactions, increase structural constraints, or alter residue distribution patterns tend to reduce toxicity. These findings provide testable hypotheses for the rational design of more selective antimicrobial peptides and demonstrate the value of interpretable descriptor-based approaches for uncovering mechanistic insights in complex biomolecular systems.
Supplementary Material
Acknowledgments
The work was supported by the Welch Foundation (C-1559), the NIH (R01GM148537), and the Center for Theoretical Biological Physics sponsored by the NSF (PHY-2019745).
The data set used in this work, the analysis scripts, and the complete analysis outputs and details are available on Figshare at the following URL: https://figshare.com/s/e13e7a7fcfe32ab3fc5a
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.6c01031.
Additional explanation of methods, detailed evaluation metrics for each prediction model, detailed discussion and possible mechanistic interpretation of features reproduced from the prior study, and description of all selected features mentioned in the main text from each analysis (PDF)
A.M. designed the research. A.M., K.K., C.V., A.R., and A.B. performed the research. A.M. and A.B.K. wrote the article.
The authors declare no competing financial interest.
References
- Szymczak P., Zarzecki W., Wang J., Duan Y., Wang J., Coelho L. P., de la Fuente-Nunez C., Szczurek E.. AI-Driven Antimicrobial Peptide Discovery: Mining and Generation. Acc. Chem. Res. 2025;58:1453–1871. doi: 10.1021/acs.accounts.0c00594. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ramos-Llorens M., Bello-Madruga R., Valle J., Andreu D., Torrent M.. PyAMPA: a high-throughput prediction and optimization tool for antimicrobial peptides. MSystems. 2024;9:e01358-23. doi: 10.1128/msystems.01358-23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Júnior N. G. O., Souza C. M., Buccini D. F., Cardoso M. H., Franco O. L.. Antimicrobial peptides: structure, functions and translational applications. Nat. Rev. Microbiol. 2025;23:687–700. doi: 10.1038/s41579-025-01200-y. [DOI] [PubMed] [Google Scholar]
- Roque-Borda C. A., Primo L. M. D. G., Medina-Alarcón K. P., Campos I. C., de Fátima Nascimento C., Saraiva M. M., Junior A. B., Fusco-Almeida A. M., Mendes-Giannini M. J. S., Perdigão J.. et al. Antimicrobial peptides: a promising alternative to conventional antimicrobials for combating polymicrobial biofilms. Adv. Sci. 2025;12:2410893. doi: 10.1002/advs.202410893. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang S., Zeng X., Yang Q., Qiao S.. Antimicrobial peptides as potential alternatives to antibiotics in food animal industry. Int. J. Mol. Sci. 2016;17(5):603. doi: 10.3390/ijms17050603. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ebbensgaard A., Mordhorst H., Overgaard M. T., Nielsen C. G., Aarestrup F. M., Hansen E. B.. Comparative evaluation of the antimicrobial activity of different antimicrobial peptides against a range of pathogenic bacteria. PLoS One. 2015;10:e0144611. doi: 10.1371/journal.pone.0144611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benfield A. H., Henriques S. T.. Mode-of-Action of Antimicrobial Peptides: Membrane Disruption vs. Intracellular Mechanisms. Front. Med. Technol. 2020;2:610997. doi: 10.3389/fmedt.2020.610997. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bezu L., Kepp O., Cerrato G., Pol J., Fucikova J., Spisek R., Zitvogel L., Kroemer G., Galluzzi L.. Trial watch: peptide-based vaccines in anticancer therapy. Oncoimmunology. 2018;7:e1511506. doi: 10.1080/2162402X.2018.1511506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang L.-j., Gallo R. L.. Antimicrobial peptides. Curr. Biol. 2016;26:R14–R19. doi: 10.1016/j.cub.2015.11.017. [DOI] [PubMed] [Google Scholar]
- Zhu Y., Hao W., Wang X., Ouyang J., Deng X., Yu H., Wang Y.. Antimicrobial peptides, conventional antibiotics, and their synergistic utility for the treatment of drug-resistant infections. Med. Res. Rev. 2022;42:1377–1422. doi: 10.1002/med.21879. [DOI] [PubMed] [Google Scholar]
- Lei J., Sun L., Huang S., Zhu C., Li P., He J., Mackey V., Coy D. H., He Q.. The antimicrobial peptides and their potential clinical applications. Am. J. Transl. Res. 2019;11(7):3919–3931. [PMC free article] [PubMed] [Google Scholar]
- Gautam A., Chaudhary K., Singh S., Joshi A., Anand P., Tuknait A., Mathur D., Varshney G. C., Raghava G. P.. Hemolytik: a database of experimentally determined hemolytic and non-hemolytic peptides. Nucleic Acids Res. 2014;42:D444–D449. doi: 10.1093/nar/gkt1008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chaudhary K., Kumar R., Singh S., Tuknait A., Gautam A., Mathur D., Anand P., Varshney G. C., Raghava G. P.. A web server and mobile app for computing hemolytic potency of peptides. Sci. Rep. 2016;6:22843. doi: 10.1038/srep22843. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rathore A. S., Kumar N., Choudhury S., Mehta N. K., Raghava G. P.. Prediction of hemolytic peptides and their hemolytic concentration. Commun. Biol. 2025;8:176. doi: 10.1038/s42003-025-07615-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kumar A., Chadha S., Sharma M., Kumar M.. Deciphering optimal molecular determinants of non-hemolytic, cell-penetrating antimicrobial peptides through bioinformatics and Random Forest. Briefings Bioinf. 2025;26:bbaf049. doi: 10.1093/bib/bbaf049. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bhatnagar P., Khandelwal Y., Mishra S., G S. K., Dutta A., Dutta A., Mitra D., Mitra D., Biswas S.. Predicting antibacterial activity, efficacy, and hemotoxicity of peptides using an explainable machine learning framework. Process Biochem. 2024;145:163–174. doi: 10.1016/j.procbio.2024.06.027. [DOI] [Google Scholar]
- Tsai C.-T., Lin C.-W., Ye G.-L., Wu S.-C., Yao P., Lin C.-T., Wan L., Tsai H.-H. G.. Accelerating antimicrobial peptide discovery for who priority pathogens through predictive and interpretable machine learning models. ACS Omega. 2024;9:9357–9374. doi: 10.1021/acsomega.3c08676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Khabbaz H., Karimi-Jafari M. H., Saboury A. A., BabaAli B.. Prediction of antimicrobial peptides toxicity based on their physico-chemical properties using machine learning techniques. BMC Bioinf. 2021;22:549. doi: 10.1186/s12859-021-04468-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tan X., Liu Q., Fang Y., Zhu Y., Chen F., Zeng W., Ouyang D., Dong J.. Predicting peptide permeability across diverse barriers: a systematic investigation. Mol. Pharm. 2024;21:4116–4127. doi: 10.1021/acs.molpharmaceut.4c00478. [DOI] [PubMed] [Google Scholar]
- Kang Y., Zhang H., Wang X., Yang Y., Jia Q.. MMDB: Multimodal dual-branch model for multi-functional bioactive peptide prediction. Anal. Biochem. 2024;690:115491. doi: 10.1016/j.ab.2024.115491. [DOI] [PubMed] [Google Scholar]
- Adamczyk, J. ; Ludynia, P. ; Czech, W. . Molecular Fingerprints Are Strong Models for Peptide Function Prediction. 2025, arXiv:2501.17901. arXiv:2501.17901. https://arxiv.org/abs/2501.17901. [DOI] [PMC free article] [PubMed]
- Pirtskhalava M., Amstrong A. A., Grigolava M., Chubinidze M., Alimbarashvili E., Vishnepolsky B., Gabrielian A., Rosenthal A., Hurt D. E., Tartakovsky M.. DBAASP v3: database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Res. 2021;49:D288–D297. doi: 10.1093/nar/gkaa991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pirtskhalava M., Gabrielian A., Cruz P., Griggs H. L., Squires R. B., Hurt D. E., Grigolava M., Chubinidze M., Gogoladze G., Vishnepolsky B.. et al. DBAASP v. 2: an enhanced database of structure and antimicrobial/cytotoxic activity of natural and synthetic peptides. Nucleic Acids Res. 2016;44:D1104–D1112. doi: 10.1093/nar/gkv1174. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Abd El-Wahed A., Yosri N., Sakr H. H., Du M., Algethami A. F., Zhao C., Abdelazeem A. H., Tahir H. E., Masry S. H., Abdel-Daim M. M.. et al. Wasp venom biochemical components and their potential in biological applications and nanotechnological interventions. Toxins. 2021;13:206. doi: 10.3390/toxins13030206. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A.. et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–589. doi: 10.1038/s41586-021-03819-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fleming J., Magana P., Nair S., Tsenkov M., Bertoni D., Pidruchna I., Afonso M. Q. L., Midlik A., Paramval U., Žídek A.. et al. AlphaFold protein structure database and 3D-Beacons: new data and capabilities. J. Mol. Biol. 2025;437:168967. doi: 10.1016/j.jmb.2025.168967. [DOI] [PubMed] [Google Scholar]
- Martins D. B., Fadel V., Oliveira F. D., Gaspar D., Alvares D. S., Castanho M. A., dos Santos Cabrera M. P.. Protonectin peptides target lipids, act at the interface and selectively kill metastatic breast cancer cells while preserving morphological integrity. J. Colloid Interface Sci. 2021;601:517–530. doi: 10.1016/j.jcis.2021.05.115. [DOI] [PubMed] [Google Scholar]
- Cao D. S., Xu Q. S., Liang Y.. Z+ propy: a tool to generate various modes of Chou’s PseAAC. Bioinformatics. 2013;29:960–962. doi: 10.1093/bioinformatics/btt072. [DOI] [PubMed] [Google Scholar]
- Bento A. P., Hersey A., Félix E., Landrum G., Gaulton A., Atkinson F., Bellis L. J., De Veij M., Leach A. R.. An open source chemical structure curation pipeline using RDKit. J. Cheminf. 2020;12:51. doi: 10.1186/s13321-020-00456-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Moriwaki H., Tian Y.-S., Kawashita N., Takagi T.. Mordred: a molecular descriptor calculator. J. Cheminf. 2018;10:4. doi: 10.1186/s13321-018-0258-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Balfer J., Bajorath J.. Visualization and interpretation of support vector machine activity predictions. J. Chem. Inf. Model. 2015;55:1136–1147. doi: 10.1021/acs.jcim.5b00175. [DOI] [PubMed] [Google Scholar]
- Lee E. Y., Fulan B. M., Wong G. C., Ferguson A. L.. Mapping membrane activity in undiscovered peptide sequence space using machine learning. Proc. Natl. Acad. Sci. U.S.A. 2016;113:13588–13593. doi: 10.1073/pnas.1609893113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Menze B. H., Kelm B. M., Masuch R., Himmelreich U., Bachert P., Petrich W., Hamprecht F. A.. A comparison of random forest and its Gini importance with standard chemometric methods for the feature selection and classification of spectral data. BMC Bioinf. 2009;10:213. doi: 10.1186/1471-2105-10-213. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pandey D., Niwaria K., Chourasia B.. Machine learning algorithms: a review. Mach. Learn. 2019;6:916–922. [Google Scholar]
- Teimouri H., Medvedeva A., Kolomeisky A. B.. Bacteria-specific feature selection for enhanced antimicrobial peptide activity predictions using machine-learning methods. J. Chem. Inf. Model. 2023;63:1723–1733. doi: 10.1021/acs.jcim.2c01551. [DOI] [PubMed] [Google Scholar]
- Wang G., Li X., Wang Z.. APD3: the antimicrobial peptide database as a tool for research and education. Nucleic Acids Res. 2016;44:D1087–D1093. doi: 10.1093/nar/gkv1278. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Plisson F., Ramírez-Sánchez O., Martínez-Hernández C.. Machine learning-guided discovery and design of non-hemolytic peptides. Sci. Rep. 2020;10:16581. doi: 10.1038/s41598-020-73644-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rana R., Singhal R.. Chi-square test and its application in hypothesis testing. J. Pract. Cardiovasc. Sci. 2015;1:69–71. doi: 10.4103/2395-5414.157577. [DOI] [Google Scholar]
- Tam J. P., Wang S., Wong K. H., Tan W. L.. Antimicrobial peptides from plants. Pharmaceuticals. 2015;8:711–757. doi: 10.3390/ph8040711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Q.-Y., Yan Z.-B., Meng Y.-M., Hong X.-Y., Shao G., Ma J.-J., Cheng X.-R., Liu J., Kang J., Fu C.-Y.. Antimicrobial peptides: mechanism of action, activity and clinical potential. Mil. Med. Res. 2021;8:48. doi: 10.1186/s40779-021-00343-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Serian M., Mason A. J., Lorenz C. D.. Emergent conformational and aggregation properties of synergistic antimicrobial peptide combinations. Nanoscale. 2024;16:20657–20669. doi: 10.1039/D4NR03043E. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chen P., Zhang T., Li C., Praveen P., Parisi K., Beh C., Ding S., Wade J. D., Hong Y., Li S.. et al. Aggregation-prone antimicrobial peptides target gram-negative bacterial nucleic acids and protein synthesis. Acta Biomater. 2025;192:446–460. doi: 10.1016/j.actbio.2024.12.002. [DOI] [PubMed] [Google Scholar]
- Huan Y., Kong Q., Mou H., Yi H.. Antimicrobial peptides: classification, design, application and research progress in multiple fields. Front. Microbiol. 2020;11:582779. doi: 10.3389/fmicb.2020.582779. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pirtskhalava, M. ; Vishnepolsky, B. ; Grigolava, M. . Transmembrane and antimicrobial peptides. Hydrophobicity, amphiphilicity and propensity to aggregation. 2013, arXiv:1307.6160. arXiv.org e-Printarchive. https://arxiv.org/abs/1307.6160.
- Todeschini, R. ; Consonni, V. . Handbook of Molecular Descriptors; John Wiley & Sons, 2008. [Google Scholar]
- Linker S. M., Schellhaas C., Ries B., Roth H.-J., Fouché M., Rodde S., Riniker S.. Polar/apolar interfaces modulate the conformational behavior of cyclic peptides with impact on their passive membrane permeability. RSC Adv. 2022;12:5782–5796. doi: 10.1039/D1RA09025A. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Coppins, R. L. Cyclic Antibiotic Peptide Design: Structure and Membrane Interaction, Technical report, University of Illinois at Urbana-Champaign; 2001. https://chemistry.illinois.edu/system/files/inline-files/s02_Coppins.pdf. February 14, 2001. [Google Scholar]
- Beck K., Nandy J., Hoernke M.. Strong Membrane Permeabilization Activity Can Reduce Selectivity of Cyclic Antimicrobial Peptides. J. Phys. Chem. B. 2025;129:2446–2460. doi: 10.1021/acs.jpcb.4c05019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vejzovic D., Piller P., Cordfunke R. A., Drijfhout J. W., Eisenberg T., Lohner K., Malanovic N.. Where electrostatics matter: bacterial surface neutralization and membrane disruption by antimicrobial peptides SAAP-148 and OP-145. Biomolecules. 2022;12:1252. doi: 10.3390/biom12091252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boudjemaa R., Cabriel C., Dubois-Brissonnet F., Bourg N., Dupuis G., Gruss A., Lévêque-Fort S., Briandet R., Fontaine-Aupart M.-P., Steenkeste K.. Impact of bacterial membrane fatty acid composition on the failure of daptomycin to kill Staphylococcus aureus. Antimicrob. Agents Chemother. 2018;62:10–1128. doi: 10.1128/AAC.00023-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Felsztyna I., Galassi V. V., Wilke N.. Selectivity of membrane-active peptides: the role of electrostatics and other membrane biophysical properties. Biophys. Rev. 2025;17:591–604. doi: 10.1007/s12551-025-01309-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Edison J. R., Spencer R. K., Butterfoss G. L., Hudson B. C., Hochbaum A. I., Paravastu A. K., Zuckermann R. N., Whitelam S.. Conformations of peptoids in nanosheets result from the interplay of backbone energetics and intermolecular interactions. Proc. Natl. Acad. Sci. U.S.A. 2018;115:5647–5651. doi: 10.1073/pnas.1800397115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson B. C., Battigelli A., Connolly M. D., Edison J., Spencer R. K., Whitelam S., Zuckermann R. N., Paravastu A. K.. Evidence for cis amide bonds in peptoid nanosheets. J. Phys. Chem. Lett. 2018;9:2574–2578. doi: 10.1021/acs.jpclett.8b01040. [DOI] [PubMed] [Google Scholar]
- Wexselblatt E., Esko J. D., Tor Y.. On guanidinium and cellular uptake. J. Org. Chem. 2014;79:6766–6774. doi: 10.1021/jo501101s. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chi Y., Peng Y., Zhang S., Tang S., Zhang W., Dai C., Ji S.. A rapid in vivo toxicity assessment method for antimicrobial peptides. Toxics. 2024;12:387. doi: 10.3390/toxics12060387. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Daniel O., Thiebault T., Guigon E.. Medicinal applications and environmental fate of antimicrobial peptides: a review. Environ. Chem. Lett. 2025;23:1745–1776. doi: 10.1007/s10311-025-01854-3. [DOI] [Google Scholar]
- Kumar K., Woo S. M., Siu T., Cortopassi W. A., Duarte F., Paton R. S.. Cation-π interactions in protein-ligand binding: theory and data-mining reveal different roles for lysine and arginine. Chem. Sci. 2018;9:2655–2665. doi: 10.1039/C7SC04905F. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shigedomi K., Osada S., Jelokhani-Niaraki M., Kodama H.. Systematic design and validation of ion channel stabilization of amphipathic α-helical peptides incorporating tryptophan residues. ACS Omega. 2021;6:723–732. doi: 10.1021/acsomega.0c05254. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Giménez-Andrés M., Čopič A., Antonny B.. The many faces of amphipathic helices. Biomolecules. 2018;8:45. doi: 10.3390/biom8030045. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liang Y., Zhang Y., Huang Y., Xu C., Chen J., Zhang X., Huang B., Gan Z., Dong X., Huang S.. et al. Helicity-directed recognition of bacterial phospholipid via radially amphiphilic antimicrobial peptides. Sci. Adv. 2024;10:eadn9435. doi: 10.1126/sciadv.adn9435. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data set used in this work, the analysis scripts, and the complete analysis outputs and details are available on Figshare at the following URL: https://figshare.com/s/e13e7a7fcfe32ab3fc5a







