Abstract
Despite substantial progress in nanopore sensing, residue-by-residue peptide sequencing remains a major challenge. Herein, we present EANPSeq, an exopeptidase-assisted nanopore peptide identification strategy based on peptide libraries to decode the peptide sequence. By continuously recognizing the resulting fragments from digesting peptides stepwise through a nanopore, this approach could achieve the identification of peptide sequence based on the comparison of fragment data with libraries of shortened and mutated peptides, with the assistance of machine learning. Notably, compared with previously reported nanopore peptide sensing strategies, EANPSeq shows sufficient resolution to recognize the continuous sequence of peptides containing adjacent identical residues and to precisely localize post-translational modification (PTM) sites within consecutive residues. These proof-of-concept results highlight our nanopore-based strategy as a new avenue for single-molecule protein sequencing.
Subject terms: Sequencing, Nanobiotechnology, Nanopores, Nanopores, Peptides
There is significant interest in developing nanopore-based methods for peptide sensing and sequencing. Here the authors report, EANPSeq, a method which enables residue-by-residue peptide sequencing using an exopeptidase-assisted approach, decoding stepwise digestion fragments with machine-learning to achieve single-molecule sequence identification and PTM localization.
Introduction
Decoding the peptide sequences is important for understanding peptide structure, function, and biological processes1–3. Currently, the most widely used sequencing methods include Edman degradation4,5 and mass spectrometry (MS)6,7. Despite efforts toward automation to reduce detection time, Edman degradation cannot easily sequence peptides from complex peptide mixtures and struggles with blocked N-termini. Meanwhile, MS faces several challenges, including limits of detection8, restricted abundance dynamic range9, and difficulty in unlabeled identifying post-translational modifications (PTMs)10, which hinder the sequencing of low-abundance peptides. These limitations have driven the development of new single-molecule technology for label-free peptide sequencing, especially for PTMs and low-abundance peptides.
Nanopore technology is a powerful analytical technique for single-molecule sensing. In nanopore sensing, when an individual analyte molecule passes through the nanopore under an applied voltage, the ionic flow changes, causing a detectable blockage in the ionic current11–13. Analyzing this current blockage can provide insight into the molecular properties of the analyte, such as size, charge, composition, and conformation14–19. So far, nanopore has previously been applied to nucleic acid sequencing20,21, single amino acid sensing and peptide sensing, including discrimination of all 20 proteinogenic amino acids and their modifications22–25, protein identification26, peptide PTMs detection27–29, peptide enantiomers discrimination30–33, and peptide structural variants analysis34. In the past few years, two main nanopore-based peptide sequence identification strategies have been proposed. One is to use motor enzymes (e.g., helicases and unfoldases) to control the stepwise translocation of peptides terminally bonded with a capture tag through the protein nanopore35–38. This labeled strategy can effectively discriminate single negatively charged amino acid substitutions. But peptides enriched in positive or neutral residues would tangle up inside the nanopore, making it difficult to reliably extract sequence information39. The other is the detection of free amino acids released from exopeptidase digestion of the peptide23–25, where the likely peptide sequence is deduced from abundance variations of amino acids. This approach faces challenges in resolving peptides with repeated residues. In addition, due to the low capture efficiency of amino acids entering the nanopore, relatively high peptide concentrations are required.
To achieve precise decoding of continuous peptide sequences, we develop EANPSeq, an exopeptidase-assisted nanopore peptide sequence identification strategy (Fig. 1). We analyze the unlabeled peptide fragments generated from exopeptidase digestion using a biological nanopore. The resulting single-molecule events are matched to peptide libraries with the assistance of artificial intelligence (AI), enabling stepwise sequence readout. To evaluate our strategy, we analyze peptide fragments generated by carboxypeptidase Y (CPY) using the T274L/N226Q/S228K Aerolysin (LQK AeL) nanopore at defined digestion times. Because CPY sequentially removes residues from the peptide C terminus, the resulting fragment series provides temporal information for peptide sequence identification. After comparing these data with a shortened and mutated peptide library, we can precisely resolve fragment lengths and continuous sequences through the developed repeated verification strategy. Especially, this method offers sufficient resolution to sequence peptides with adjacent identical residues. More importantly, post-translational modification sites can be localized with high precision, which is demonstrated by the unambiguous identification of aspartic acid (Asp, D) isomerization within consecutive D residues. Moreover, our results show that the trained AI model potentially enables reliable identification of peptides differing in abundance differences of up to 106-fold, thereby supporting its potential applicability to peptide identification within a broad abundance dynamic range. This advance opens new avenues for label-free peptide sequencing, bringing the field closer to single-molecule proteomics.
Fig. 1. Schematic representation of recognizing peptide sequence using the method EANPSeq.

a The peptide is digested into peptide fragments by carboxypeptidase Y (CPY) for the nanopore direct sensing. For proof-of-concept, the produced peptide fragments are directly detected by T274L/N226Q/S228K Aerolysin (LQK AeL). b The ionic current data acquired at various digestion time points. c Features from recording data of peptide fragments are analyzed by using the AI model for the shortened library (peptides with varying lengths) to determine the corresponding length and employing a mutated peptide library (single amino acid differences at the same position of the peptide) to obtain relative sequence, respectively. Here, peptide sequences were randomly selected as representative models to demonstrate the general applicability of EANPSeq, with all sequences listed in Supplementary Table 1. The peptide sequence can be decoded from the cross-validation results obtained at various digestion time points.
Results
Preparing a shortened peptide library for sequence length calling
Aerolysin (AeL) nanopore possesses a β-barrel structure with a long lumen (~10 nm) and narrow diameter (~1 nm), enabling extensive interactions with peptides and generating rich ionic current fingerprints for peptide recognition40. The previous studies revealed that the mutant sites N226Q and S228K of AeL enhanced electroosmotic flow (EOF) to promote peptide capture, while forming an electrostatic trap that decelerated peptide translocation41. To further enhance the sensitivity, the threonine at site 274 was mutated to leucine with a bulky volume (denoted as T274L), producing the mutant T274L/N226Q/S228K Aerolysin (LQK AeL). Molecular dynamics (MD) simulations showed that electrostatic trap around N226Q and S228K sites was strengthened, thereby promoting prolonged intensive interactions between the pore and peptides (Supplementary Fig. 1). To evaluate the sensing performance of this mutant nanopore, we used a series of AeL mutants to identify two peptides (P8-5H and P8-5G) differing by a single-amino-acid differences at the fifth position in the sequence N’-RHFSXGED-C’ (X is amino acid H or G) (Fig. 2a, b, Supplementary Figs. 2–7, and Tables 1–4). As shown in the scatter plot of duration time (τ) versus ionic current blockage (I/I0) in Fig. 2a, the LQK AeL exhibits a high resolution for peptide discrimination, with a resolution value (R) of 2, demonstrating that the two I/I0 distribution peaks corresponding to the two peptides are well separated (Supplementary Table 2). Furthermore, the statistical duration time generated by the mutant LQK AeL was more than twice that of the others, enabling the acquisition of events with richer fingerprints for improving peptide identification (Fig. 2b and Supplementary Table 3). The maximum theoretical number of resolvable peptides (I/I0 peak capacity) for the LQK AeL reaches 20.0, nearly three times higher than that of other nanopores (Fig. 2b and Supplementary Table 4, calculation details shown in “Methods”). Under the synergistic effect of electrophoretic force (EPF) and EOF (Supplementary Fig. 5, details of EOF calculation are shown in the “Methods” and “Supplementary Discussion” Section), the peptides with positive, negative, or neutral charges can all be captured and detected by the LQK AeL, with event frequencies (fpeptide) values exceeding 100 min−1 (Supplementary Fig. 8 and Supplementary Table 5).
Fig. 2. Identification of peptides with a single amino acid length difference.

a The scatter plot of duration time (τ) versus current blockage (I/I0) and corresponding I/I0 histogram of P8-5H and P8-5G peptide mixture through T274L/N226Q/S228K Aerolysin (LQK AeL) nanopore. Events with an I/I0 value greater than approximately 0.55 are bumping events (Supplementary Fig. 7). b The I/I0 peak capacity and statistical duration time of P8-5H and P8-5G mixture detected by WT, T232K, T274V, T274L, and LQK AeL. Related calculation details are shown in the Methods. The concentrations of both peptides were 4.0 μM. c Representative peptide events measured with LQK AeL. The final concentrations of peptides P12, P11, P10, P9, P8, and P7 were 2.0, 2.0, 2.0, 2.0, 8.0, and 2.0 μM, respectively. d Workflow of the machine learning algorithm. Features were extracted from datasets and input into a Bagging classifier to train the AI model. Details of the training model are shown in the Methods “Machine Learning”. e The confusion matrix of classifying peptides using the AI model. TPR represents the proportion of correct classification of each true class (Data shown in Supplementary Table 8). f Representative ionic current traces of hydrolyzed residual peptide fragments generated at various digestion times. t0, t1, t2, t3, t4, t5, and t6 represent digestion time at 0, 2, 5, 8, 12, and 16 min, respectively. g Digestion pathway of P12 obtained by the AI model. The pathway was determined by identifying the fragments with the highest abundance at each time point and the sequential appearance of their peak abundances (Data shown in Supplementary Table 9). All the experiments were performed in 1 M KCl and 10 mM Tris-HCl buffer (pH 8.0) at 23 ± 3 °C under a voltage of +80 mV. Source data are provided as a Source Data file.
To establish a shortened peptide library for nanopore sequence identification, the LQK AeL was further employed to detect a series of peptides with a single amino acid length difference at the C-terminus (P12, P11, P10, P9, P8, and P7). The ionic current traces of these peptides showed that all peptides generated identifiable blockage events at a sufficient frequency to allow reliable statistical analysis, with mean I/I0 values exceeding 0.30, and relative signal-to-noise ratios (SNR) larger than 10 (Fig. 2c, Supplementary Figs. 9–11, and Supplementary Table 6). The results showed that I/I0 increases as peptide length decreases (Supplementary Figs. 12 and 13). Then, a machine-learning algorithm was used to identify the unique fingerprint of blockage events corresponding to each peptide (details shown in “Methods”). According to previous studies42,43, the features extracted from the time, frequency, and time-frequency domains are correlated with the coupled interactions among nanopore inner-wall residues, ions, and the analyte, thereby enabling a comprehensive representation of each nanopore event. Therefore, a total of 72 features from these three domains were extracted from each event to enable discrimination of peptide sequence differences. After evaluating and selecting features by employing permutation importance estimation44, 27 features as a vector were input into the Bagging ensemble framework consisting of 30 decision-tree base learners to train the artificial intelligence (AI) model for the shortened peptide library (Fig. 2d and Supplementary Table 7). Using the 5-fold cross-validation, this model achieved 96.2% accuracy, indicating that these six peptides (P12 ~ P7) could be directly distinguished (Fig. 2e and Supplementary Table 8). This AI model could be applied to determine the length of peptide fragments resulting from carboxypeptidase Y (CPY) digestion (Fig. 2f, g).
Considering the amino acid-dependent cleavage efficiency of CPY45, we systematically optimized the enzyme digestion conditions, including pH, enzymatic addition protocol, and enzyme dosage, to rapidly generate all possible peptide fragments resulting from cleavage (Supplementary Fig. 14). The results demonstrated that under optimized stepwise enzyme addition conditions (pH 5.4, initial peptide-to-CPY volume ratio VS, t = 0 min:VE, t = 0 min = 60:0.6), with sequential additions of 5 and 10 μL CPY at the 5th and 14th minutes of digestion, respectively (Condition E in Supplementary Fig. 14), all possible peptide fragments (P11, P10, P9, P8, and P7) were generated and detected within only 16 min (Fig. 2f). Since CPY sequentially removes amino acid residues from the peptide C-terminus45, the longest peptide fragments (P11) initially increase in concentration. The abundance of P11 in the mixture peaked within the first minute of digestion and then decreased as it became a substrate for further cleavage, while those of P10, P9, P8, and P7 sequentially increased over the digestion time (Fig. 2g, Supplementary Fig. 15, and Supplementary Table 9). These time-dependent changes were further confirmed by mass spectrometry (Supplementary Fig. 16). The resulting abundance-time profile reflects the enzymatic cleavage order and allows cross-validation of peptide identification between digestion time points, assisting in determining the peptide lengths. As for peptide P12, the adjacent amino acid residues at positions 8 and 9 are both aspartic acid residues. After CPY digestion, the peptide fragments P9, P8, and P7 could be sequentially identified, allowing the amino acid residues at adjacent positions 8 and 9 to be independently recognized and verified (Fig. 2f, g). In addition to peptide fragments, CPY cleavage also produces free C-terminal amino acids. These free amino acids, together with the CPY enzyme itself, generate no detectable single-molecule events and therefore do not interfere with peptide fragment detection (Supplementary Figs. 17 and 18). Based on the determination of the fragment length, this method could achieve reading the peptide sequence even if consecutive amino acid residues in a peptide are identical.
Preparing the mutated peptide library for sequence calling
After determining the peptide lengths, the recorded nanopore data were compared with a mutated peptide library containing single amino acid substitutions for further resolving the fragment sequences (Fig. 1c). To conduct a proof-of-concept, we established the 8-mer mutated peptide library (the sequence is N’-RHFSXGED-C’, and X represents 20 natural amino acids, shown in Supplementary Table 1). Each peptide was measured by the LQK AeL. It was observed that all the peptides showed distinct current blockade (Fig. 3a, b, and Supplementary Fig. 19). Peptides P8-5X with mutated charged amino acid residues of D, E, R, and K exhibited lower I/I₀ values, which may be attributed to stronger electrostatic interactions within the nanopore. In contrast, nonpolar amino acid residues of M, F, W, I, A, P, V, L, and G showed relatively higher I/I₀ values, potentially arising from weaker peptide-nanopore interactions and reduced perturbation of ionic transport. Compared with previous studies using Aerolysin nanopores for peptide detection15,46–48, this mutant nanopore demonstrates the strongest capability for peptide sequence identification without labeling and introducing extra charged carrier. To improve the accuracy and efficiency of peptide identification, 48 features were extracted and applied in training the AI model using the Bagging classifier (Supplementary Fig. 20 and Supplementary Tables 10–12). Furthermore, the correlations between features and peptide properties were analyzed (Fig. 3c). The results demonstrated that the top-ranked features, including I, I/I₀, coefficient of variation (cov), and range of current amplitude for each blockage (range), exhibit strong correlations with dipole moment, suggesting that peptide-nanopore interaction and orientation-dependent electrostatic effects play an important role in event differentiation. In contrast, features correlated with molecular weight show lower importance scores in permutation-based feature ranking, indicating that these properties contribute less to the peptide discrimination. These results are consistent with a previous study49, where peptide-pore interactions were found to be more critical than molecular weight in determining event characteristics. Among these features, sample entropy (SampleEn), quantifying the complexity of time-series data, is one of the most useful features for peptide recognition in addition to I/I0 (Supplementary Fig. 21), according to the evaluation of feature stability and separability. Our results showed that all 20 peptides with single amino acid differences could be identified with a high accuracy of 90.1% (Fig. 3d and Supplementary Table 13).
Fig. 3. Identification of 20 peptides with a single amino acid difference.

a Typical events generated by 20 peptides, P8-5X, measured using LQK AeL. Sequences of peptides P8-5X are provided in Supplementary Table 1. b Experimentally determined I/I0 values of peptides P8-5X. Data are presented as mean ± s.d. c Relationship between features and peptide or amino acid physicochemical properties (mutated amino acid dipole moment (μAA), peptide grand average of hydropathicity (GRAVY), peptide molecular weight (M.W.), and peptide isoelectric point (pI)). For each property, the top five features with the highest absolute Spearman correlation coefficients are shown, with circle size indicating ranking (larger circles indicate stronger correlation). Feature importance is provided in Supplementary Fig. 20. Calculation details of the coefficients are described in the Data Analysis “Feature analysis”. d The confusion matrix of classifying peptides P8-5X by the AI model (Data shown in Supplementary Table 13). e Quantification of each peptide from the constructed mixture consisted of randomly sampling events from independently acquired P8-5I and P8-5W datasets. The NP8-5I/NP8-5W ratios were predicted using AI model, with NP8-5W normalized to 1 (Details shown in the “Data Analysis” section). The black diagonal (y = x) indicates perfect agreement between predicted and actual values. Data are presented as mean ± s.d. from 3 independent sampling and prediction experiments. All data were collected in the 1 M KCl and 10 mM Tris-HCl buffer (pH 8.0) at 23 ± 3 °C under a voltage of +80 mV. The concentrations of peptides P8-5C, P8-5I, P8-5Q, P8-5S, P8-5T, and P8-5V were 10.0 μM. The concentrations of peptides P8-5F and P8-5W were 1.0 μM. The concentration of peptide P8-5D was 2.0 μM. The concentrations of remaining peptides were 4.0 μM. Source data are provided as a Source Data file.
Since the peptide abundance in the single cell spans several orders of magnitude50, the sequence identification method should be able to identify peptides across a wide range of abundance differences. Therefore, our method was subsequently applied to quantify constructed peptide datasets generated by randomly sampling events from the labeled P8-5I and P8-5W datasets at varying ratios (Fig. 3e and details shown in the Methods section “Quantification of Peptide Abundance”). As shown in Fig. 3e, the predicted abundances of two peptides closely matched the true event ratios, with an average precision higher than 99.96% (Supplementary Table 14). Notably, the model can quantify peptide samples with an approximately 106-fold difference in abundance, which has not been realized by other Aerolysin nanopore methods. This performance is largely attributed to the mutant LQK AeL, which prolongs events of peptides compared to other pores, thereby generating richer features (Fig. 2b). These improved event characteristics enable the AI model to more effectively extract discriminative information, leading to improved quantitative accuracy. Therefore, this reliable method can be applied for the precise sequence identification and relative quantification of peptides with single amino acid differences.
Recognition of peptide sequence
To demonstrate the proof-of-concept study of recognizing peptide sequence, the peptide P12-5I was used as a model peptide and hydrolyzed by CPY for evaluation. Peptide fragments generated during CPY digestion were collected at various time points and subsequently detected using LQK AeL (Fig. 4a, b). The representative raw current trace collected at digestion time points demonstrated that events with progressively larger current blockage (I/I0) gradually appeared as hydrolysis proceeded, indicating the continuous generation of shorter peptide fragments during CPY digestion (Fig. 4c). This approach can expand a single peptide dataset into multiple fragment datasets with different peptide lengths, thereby increasing the diversity of signal features for peptide identification. To quantitatively determine peptide length during hydrolysis, the events extracted from each hydrolysis time point were analyzed using the AI model trained with the shortened peptide library (Fig. 4d). The prediction results revealed a clear time-dependent evolution of peptide products from the original 12-mer peptide toward progressively shorter peptide fragments, consistent with the expected CPY digestion pathway (Fig. 4e). To evaluate the reproducibility and quantitative reliability of the approach, three independent hydrolysis experiments were performed (Supplementary Table 15). The results obtained from all repeats were consistent. As shown in Fig. 4e and Supplementary Table 15, after 16 min of digestion, approximately 62.9% of the detected events were assigned as P9′ (9-mer peptides, PX′ notes a fragment with a confirmed length X), while approximately 10.0% corresponded to P8′.
Fig. 4. Proof-of-concept demonstration of peptide sequence identification method EANPSeq.

a Peptide fragments generated by carboxypeptidase Y (CPY) digestion for defined time periods. b Schematic illustration of peptide fragments detected by the engineered LQK AeL nanopore. c Typical current trace of fragments at each digestion time point. t0, t1, t2, t3, t4, t5, and t6 represent digestion time at 0, 2, 5, 8, 12, and 16 min, respectively. d Fragment events were predicted by the AI model trained on the shortened peptide library. TPR represents the proportion of correct classification of each true class. e Peptide hydrolysis pathway map of P12-5I for obtaining peptide length. f Fragment sequences were identified by calculating similarity probability scores between the known-length fragment data and sequences in the mutated peptide library. g Similarity ranking of P8′ obtained by digesting P12-5I for 16 min to determine the most probable peptide sequence. The rank denotes the probability score, with a smaller rank corresponding to a higher similarity probability. Here, the P8′ showed the highest similarity to the mutated peptide P8-5I. This approach allows for multiple rounds of independent sequence validation, while cross-validation across time points helps reduce false positives. Details of identifying peptide length and sequence are shown in “Recognition of the Peptide Sequence” of Data Analysis in the Methods section. All data were collected in the 1 M KCl and 10 mM Tris-HCl buffer (pH 8.0) at 23 ± 3 °C under a voltage of +80 mV. Source data are provided as a Source Data file.
After peptide length determination, the most probable sequence of each fragment was assigned by evaluating the similarity between the data of the known-length peptide fragment and the mutated peptide library (Fig. 4f, details in Methods “Recognition of the Peptide Sequence”). As shown in Fig. 4g, the dataset of P8′ generated after 16 min hydrolysis showed the highest similarity to peptide P8-5I in the mutated peptide library, suggesting that the most likely sequence is N’-RHFSIGED-C’, in agreement with the actual sequence (Supplementary Fig. 22). Because peptide fragments of identical length can be generated and analyzed at multiple hydrolysis time points, this strategy enables intrinsic cross-validation of sequence assignments, thereby effectively reducing false-positive identifications. Notably, the P8′ fragment generated after 12 min hydrolysis of P12-5I was also most confidently assigned to P8-5I, further confirming the consistency of sequence calling across different hydrolysis time points (Supplementary Fig. 23). Furthermore, peptide fragments of other lengths generated during digestion may also be matched to corresponding sequences in the mutated peptide library, potentially extending the cross-validation strategy and further reducing sequence identification errors.
Identification of post-translational modification (PTM) sites
To evaluate the effectiveness of this method in identifying the sequence of peptide containing PTM, we hydrolyzed peptide P12-9iD (N’-RHFSQGEDiDFLV-C’, where iD represents isoaspartic acid residue, shown in Fig. 5a) using CPY. Asp isomerization is a modification with unchanged molecular mass, making its identification challenging. Traditionally, it is detected by chemical labeling for mass spectrometric analysis51. The digestion pathway and the corresponding peptide fragment lengths of P12-9iD were determined using AI model for shortened peptide library (Fig. 5a and Supplementary Fig. 24a). The results indicate CPY progressively cleaves P12-9iD to generate shorter fragments, including the peptides P11′ (11-mer) and P10′ (10-mer). Notably, both the event distribution and the inferred digestion pathway reveal a distinct cleavage pattern compared with that of P12 and P12-5I. Specifically, after 12 and 16 min of incubation with CPY, P12-9iD was hydrolyzed only down to the 10-mer peptide P10′ (Fig. 5b and Supplementary Fig. 25a). In contrast, P12-8iD yielded a shorter fragment, the 9-mer peptide (P9′), although a residual population of P10′ persisted even after 12 and 16 min of digestion (Fig. 5c and Supplementary Figs. 24b, 25b). Unlike P12 and P12-5I, these two peptides containing iD neither produced P8′ (an 8-residue fragment). These results are attributed to the inability of CPY to cleave the bond between an isoaspartic acid (iD) residue and its C-terminal neighboring residue52,53. Therefore, the site of Asp isomerization is precisely recognized at either the ninth or eighth position on the N-terminal in peptide chains without any labels.
Fig. 5. Identification of PTM-containing peptide sequences.

a P12-9iD digestion pathway (left), the scatter plot of τ versus I/I0 and corresponding I/I0 histograms at 16 min (right). The scatter plot shows a broader I/I0 distribution than that of P12-8iD, possibly due to increased conformational heterogeneity or impurities. P10′ is a 10-mer peptide (PX′ notes a fragment with a confirmed length X). b The corresponding ionic current traces at various digestion times and similarity-based analysis of residual peptide fragments P10′. Sequences of peptide P10-9iD and P10-8iD are provided in Supplementary Table 1. c P12-8iD digestion pathway (left), the scatter plot of τ versus I/I0 and corresponding I/I0 histograms at 16 min (right). d The corresponding ionic current traces at various digestion times and similarity analysis of residual peptide fragments P9′. Sequences of peptide P9-8iD and P9-9iD are provided in Supplementary Table 1. t0, t1, t2, t3, t4, t5, and t6 represent digestion time at 0, 2, 5, 8, 12, and 16 min, respectively. All experiments were performed in 1 M KCl and 10 mM Tris-HCl buffer (pH 8.0) with an applied voltage of +80 mV at 23 ± 3 °C. Source data are provided as a Source Data file.
To further validate that the proposed framework can resolve specific peptide sequences, the similarity between the P10′ generated by digesting P12-9iD and each sequence in the PTM-position library was calculated. This library includes P10-9iD, P10-8iD, and P10, which have the identical length and molecular weight but differ by a single amino acid structural variation. The results show that P10′ most likely corresponds to P10-9iD, ranking highest in similarity score (Fig. 5b). Likewise, the 9-mer fragment (P9′) obtained from hydrolyzing P12-8iD showed the highest similarity to P9-8iD (Fig. 5d). Both predicted sequences are consistent with the actual peptide sequences. Importantly, these assignments were fully consistent across independent experiments, confirming the robustness of the prediction framework (Supplementary Fig. 26). This analysis demonstrates the method’s capability to distinguish peptide fragments containing single-residue modifications and to precisely localize PTMs within consecutive identical residues. With the development of a more comprehensive PTM library, this approach could further enable the identification of diverse PTMs by evaluating the similarity between experimental data and the library.
Discussion
In summary, we established EANPSeq, a continuous peptide sequence identification method based on the nanopore. Rather than focusing on freely diffused amino acids released from peptides, this approach directly recognizes intact peptide fragments, improving resolution in cases with adjacent identical residues. The combined use of shortened and mutated peptide libraries enables accurate length prediction of fragments and sequence assignment, as demonstrated by the successful decoding of the model peptide P12-5I.
One notable strength of EANPSeq is its ability to detect and localize PTMs without chemical labeling, such as aspartic acid isomerization. Furthermore, the result demonstrates the trained AI model’s capacity to potentially quantify low-abundance peptides within mixtures exhibiting up to 106-fold differences in abundance. The strategy potentially extends the dynamic range of peptide sensing and enables access to peptides that would otherwise remain undetected.
Looking forward, building more comprehensive PTM and sequence libraries, supported by AI-driven feature extraction, could significantly expand the method’s applicability to unknown peptides. Furthermore, learning patterns from diverse and information-rich datasets, such as event key features closely associated with peptide properties, could significantly reduce the effort required to construct peptide libraries. Integrating aminopeptidases could complement C-terminal digestion, ultimately achieving near-complete sequence coverage. This method could realize rapid sequencing of peptides with known sequences, such as the sequence identification of Ang peptides, thereby facilitating fast disease diagnosis. Before enzyme digestion, peptides/proteins with secondary structures or their aggregates require typically unfolded by thermal or chemical denaturation (e.g., guanidinium chloride54), along with reduction using dithiothreitol (DTT)26. Chemically treated samples are then desalted and digested with sequence-specific enzymes, such as trypsin, into shorter peptides (~14 amino acids on average55), enabling the following sequence identification by EANPSeq. Importantly, for peptides containing multiple PTM sites, enzymatic digestion may separate PTM sites into different peptide fragments depending on the cleavage positions. This fragmentation pattern, combined with subsequent CPY-assisted stepwise degradation, can facilitate PTM site localization. Together with the previous report demonstrating the nanopore detection of peptides in biofluids56, our method can be extended to peptide analysis in complex biological samples. However, for practical applications in real biofluids, appropriate sample preparation steps, such as desalination and enrichment, would likely be necessary to reduce matrix effects and ensure reliable enzymatic digestion and event detection. Meanwhile, this method could serve as a valuable platform for peptide-based drug discovery by enabling the rapid and accurate analysis of synthetic candidates. Its ability to resolve subtle sequence variations and modifications makes it particularly promising for high-throughput screening workflows.
Methods
Materials
Potassium chloride (KCl), tris(hydroxymethyl)aminomethane (Tris), sodium phosphate (Na3PO4), sodium chloride (NaCl), imidazole, 2-morpholinoethanesulfonic acid hydrate (MES hydrate), carboxypeptidase Y (CPY, 50 U/mg protein), decane, hexadecane, hexane, and 0.25% trypsin-EDTA solution were purchased from Sigma-Aldrich. 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC) was purchased from Avanti Polar Lipids. N-carbobenzoxy-L-phenylalanine-L-alanine (Z-FA) was purchased from QYAOBIO ChinaPeptides. All peptides with high purity (≥98%) were purchased from GL Biochem. E. coli BL21(DE3) pLysS cells were purchased from TransGen Biotech. Isopropyl-β-d-thiogalactopyranoside (IPTG) was purchased from Thermo Fisher Scientific.
Instruments
All nanopore experiments were performed using a portable electrochemical instrument named “Cube-D2” and a home-designed software “SmartNano.” Peptide products were analyzed using an Orbitrap Q Exactive high-resolution mass spectrometer (Thermo Fisher, Germany) equipped with a heated electrospray ionization source. The absorbance of Z-FA was measured using a UV–VIS-NIR spectrometer (UV-3600, Shimadzu Co., Ltd., Japan).
Nanopore preparation
WT, T232K, T274V, T274L, and T274L/N226Q/S228K (LQK) proaerolysin were expressed and purified in our laboratory according to the previous studies57. Briefly, each plasmid encoding WT, T232K, T274V, T274L, and LQK proaerolysin was separately transformed into E. coli BL21(DE3) pLysS cells, and each proaerolysin was expressed individually. Then, cells were spread onto Luria-Bertani (LB) agar plates containing 100 μg mL−1 ampicillin and incubated overnight at 37 °C. A single colony was inoculated into 15 mL LB broth containing ampicillin (100 μg mL−1) and cultured overnight at 37 °C with shaking at 220 r.p.m. The mixture was transferred into 1 L fresh LB broth containing ampicillin and grown at 37 °C until the optical density at 600 nm (OD600) reached 0.8. Protein expression was induced by adding 0.5 mM IPTG, followed by overnight incubation at 20 °C with shaking at 220 r.p.m. Cells were harvested by centrifugation and resuspended in 100 mL lysis buffer containing 500 mM NaCl and 20 mM sodium phosphate (pH 7.4). Then, cell disruption was performed using a high-pressure cell crusher at 70 MPa. The lysate was centrifuged for 30 min at 4 °C to remove cell debris. The supernatant containing soluble His-tagged proaerolysin was loaded onto a HisTrap HP His tag protein purification column (Cytiva) at 2 min mL−1. After washing to remove nonspecifically bound proteins, proaerolysin was eluted using buffer (20 mM Na3PO4, 500 mM NaCl, and 300 mM imidazole, pH 7.4). Purified proteins were analyzed by 12% SDS–PAGE to confirm protein purity and molecular weight.
Nanopore experiments
Nanopore measurements were performed in two chambers separated by a Teflon film containing an aperture 30–50 μm in diameter at the center. The electrically grounded chamber is defined as cis, and the other is defined as the trans chamber. After pretreating the aperture with 1% (v/v) hexadecane in hexane solution, both chambers were filled with 500 μL 1 M KCl buffer (1 M KCl, 10 mM Tris-HCl, pH 8.0), unless otherwise stated. The peptides were added to the solution of the cis chamber. A drop of 1,2-diphytanoyl-sn-glycero-3-phosphocholine (30 mg/mL in decane) was added into the two chambers to form a lipid bilayer. After proaerolysin activation with trypsin-EDTA, the activated protein was introduced into the buffer solution in the cis chamber to initiate pore insertion. Voltages were applied by a pair of Ag/AgCl electrodes immersed in the buffer solution of each chamber. All recordings were acquired at a sampling rate of 100 kHz and low-pass filtered at 5 kHz.
Molecular dynamics (MD) simulation
All MD simulations were performed in the software package NAMD. The Aerolysin-lipid system was constructed according to the previous work49. The LQK AeL was constructed using the Mutator Plugin in VMD58. The data analysis details are included in the Supplementary Information.
The measurement of the reversal potential (Vr) for AeLs
To measure the reversal voltage (Vr) of nanopores, asymmetrical KCl concentrations (1.0 M KCl buffer in cis and 0.2 M KCl in trans chambers) were introduced. To ensure the stability of the electrodes in the two chambers, agarose salt bridges containing saturated KCl were used between the electrodes and the buffer solution. The Aerolysin protein was then added to the cis chamber to form the nanopore. Finally, the reversal voltages (Vr) of nanopores were obtained (Supplementary Fig. 5).
Peptide digestion and single-molecule sensing
The digestion reaction was initiated by adding 0.60 μL of CPY solution (50 U/mL in digestion buffer, pH 6.75) to 60.0 μL of the target peptide solution (500 μM) prepared in digestion buffer (50 mM MES, pH 5.40). CPY sequentially hydrolyzes peptides from the C-terminus in a stepwise manner, releasing one amino acid at a time59–61. The catalytic constant (kcat) of CPY is determined as 23.2 s−1 (Supplementary Fig. 27 and details shown in the Methods section “Enzymatic Assay of Carboxypeptidase Y”). Under stepwise enzyme addition conditions, an additional 5.0 and 10.0 μL aliquots of CPY solution were added at the 5th and 14th minutes, respectively. The reaction mixture was incubated at 25 °C, and aliquots were collected at 2, 5, 8, 12, and 16 min. The reaction was terminated by heating at 80 °C for 4 min. Following digestion, samples were analyzed using nanopore measurements to acquire single-molecule event data.
Enzymatic assay of carboxypeptidase Y
The catalytic constant (kcat) of carboxypeptidase Y (CPY) was determined using a continuous spectrophotometric assay based on the hydrolysis of N-carbobenzoxy-L-phenylalanine-L-alanine (Z-FA)62. The reaction was carried out at 25 °C in 50 mM MES buffer. The substrate Z-FA was prepared in 50 mM MES buffer (pH 5.40) containing 2% (v/v) methanol. CPY was freshly prepared in 50 mM MES buffer (pH 6.75). For the enzyme activity assay, the 10 μL enzyme CPY solution (10 μg CPY) was added to the 3.00 mL substrate solution. A blank control was prepared identically by adding 10 μL 50 mM MES buffer (pH 6.75) without adding CPY into 3.00 mL substrate solution. The absorbances at 230 nm were recorded continuously for 5 min using a UV–VIS-NIR spectrometer with a 1 cm path length.
Data analysis
The initial time, duration time , and ionic current blockage (I/I0) values of each event were extracted using PyNanoLab software63. Scatter plots and I/I0 histograms were generated using Origin 2023b software. The I/I0 of each peptide was the peak position value obtained by fitting a Gaussian function to the corresponding I/I0 histogram.
Calculation of resolution values (R)
The R values of LQK AeL for peptides P8-5H and P8-5G were calculated using equation 164:
| 1 |
where and are the peak values obtained from the Gaussian distributions of I/I0 histogram of peptides P8-5H and P8-5G through the AeL, while and represent the corresponding full width at half maximum (FWHM) for I/I0 histogram from P8-5H and P8-5G, respectively.
Calculation of statistical duration times ()
The values of the mixture of peptides P8-5H and P8-5G through the WT, T232K, T274V, and T274L AeL were obtained by using Origin 2023b software as follows:
The I/I0 histograms were fitted using Gaussian functions to determine the FWHM. Events with I/I0 values in this range of FWHM were selected.
The logarithmic values of duration times for the selected events were calculated, and a Gaussian function was applied to fit the histogram of these logarithmic values () to obtain the peak value ().
The statistical duration time was acquired by reversing the logarithmic transformation for .
The calculation of value from LQK AeL follows the same steps as (2 and 3) described above, but in step (1), a bimodal Gaussian fitting is performed to select data within the corresponding FWHM ranges.
Calculation of I/I0 peak capacity
The I/I0 peak capacity values of nanopores were acquired according to the following Eq. 2:
| 2 |
where represents the corresponding full width at half maximum (FWHM) of I/I0 histogram generated by peptides. For WT, T232K, T274V, and T274L AeL nanopores, only a single peak was observed in the I/I₀ histograms for sensing P8-5H and P8-5G mixture. The single Gaussian fitting was applied to extract the FWHM values. FWHM values from sensing P8-5H and P8-5G mixture with LQK AeL were obtained by fitting the I/I0 histogram using a two-peak Gaussian fitting. This value represents the capability of distinguishing peptides in complex samples using I/I0 histogram. A higher value indicates that a greater number of peptides can be distinguished in a single measurement.
Calculation of peptide event frequency ()
As described in our previous study47, the frequency of peptide events () is obtained as follows:
| 3 |
where is the number of the corresponding peptide events within effective time (). was calculated by subtracting the cumulative duration of the corresponding peptide events from the total recording time. Herein, for each peptide was calculated using 180 s recording data.
Statistical signal-to-noise ratio (SNR)
SNR values of peptides were calculated according to the following equation65:
| 4 |
where is the standard deviation of the open-pore current, obtained by fitting the histogram of the recording current values using GaussAmp functions in Origin 2023b software.
Calculation of catalytic constant ()
The catalytic constant () of CPY was calculated by the equation:
| 5 |
where is the amount of the enzyme CPY. represents the maximum rate of Z-FA hydrolysis, expressed as the amount of substrate hydrolyzed per second. was obtained by fitting the initial rate versus substrate Z-FA concentration using the Michaelis-Menten equation. Initial rates were determined from the linear portion of the absorbance change of Z-FA at 230 nm within the first 1 min of the reaction.
Machine learning
The events (0.5 ms < < 100 ms and I/I0 within FWHM of its I/I0 distribution) for each peptide class were collected and labeled. To form a labeled dataset, the rising and falling edges of each event were not included. Seventy-two features, including duration, skewness, and range, were extracted from each event in the peptide library. Subsequently, the extracted features are evaluated to reduce information redundancy and improve classification accuracy. For the shortened library, 27 features through evaluation and selection from each event were assembled into feature vectors and used to train a Bagging ensemble framework consisting of 30 decision-tree base learners to obtain the AI model. While 48 selected features from the mutated library were input into the same classifier to train the model for identifying mutated peptides. For each individual learner, a bootstrap-sampled subset of the training dataset was generated through random subsampling with replacement. The predictions from all base learners were then aggregated using a majority-voting strategy to produce the final output. By increasing the diversity among individual decision trees through independent data resampling, the ensemble model effectively reduced overfitting and improved the overall generalization performance and robustness of the classifier. Finally, the model was evaluated based on the performance in 5-fold cross-validation66, and expert experience, with a comprehensive comparison to reduce the model’s overfitting. Specifically, the dataset was randomly divided into five subsets. All features extracted from the same event were kept together during the splitting process. In each fold, four subsets were used for training the model, while the remaining subset was used for testing. This procedure was repeated five times so that each subset served once as the testing set, and the final accuracy was calculated by averaging the results across all folds. The AI model achieving the highest accuracy for the shortened peptide library was then applied to predict unlabeled data obtained from peptide digestion experiments. All machine learning procedures were conducted using MATLAB R2022b software.
Feature analysis
In the mutated library, Spearman correlation coefficients between each feature value and physicochemical properties of peptides or amino acids were calculated. The absolute value of the coefficient was used to evaluate the relationship between features and properties, with larger values indicating stronger correlations67.
Quantification of peptide abundance
To evaluate the capability of the trained model for detecting low-abundance peptides, combined datasets with different relative peptide abundance ratios were constructed and analyzed. The datasets were generated by randomly sampling events from independently acquired P8-5I and P8-5W events, rather than from experimentally prepared peptide mixtures. Specifically, a certain number of events were randomly selected from the independently collected P8-5W dataset and combined with the P8-5I dataset containing 32,000 events to generate datasets with P8-5I:P8-5W ratios ranging from 10:1 to 104:1.
For higher abundance ratios (105:1–106:1), the P8-5I dataset was further expanded from 3.2 × 104 to 3.2 × 106 events using the bootstrap resampling method68. The constructed datasets were subsequently analyzed using the trained classifier, and the predicted abundances of P8-5I and P8-5W events were compared with the expected ratios to evaluate the quantitative detection performance of the model.
Recognition of the peptide sequence
All the events of peptide fragments from enzyme digestion were extracted using PyNanoLab software. Then, the event features were put into the AI model trained by the shortened peptide library (consisting of peptides P12, P11, P10, P9, P8, and P7) to predict the lengths of these fragments. At 0 min, the predicted fragment with the highest event number corresponds to the full length of the detected peptide. For each product at other digestion times, the number of all length fragments identified by the AI model was normalized. The abundance of each length fragment is the ratio of the fragment count to the total number of fragments at every digestion time. This normalization allowed for direct comparison of fragment abundance changes across different digestion times.
After determining the peptide fragment length at different digestion times, the similarity between the experimental data and the mutated peptide library was calculated to identify the most probable sequence. Four features (I, cov, Peakwave, and Pulsewave) were selected to construct 4-dimensional feature vectors. Euclidean distances between the data vector and each class in the mutated library were computed and normalized into similarity probabilities using the Softmax function69. The peptide sequence associated with the highest probability is regarded as the most probable sequence for the peptide fragment.
Statistics and reproducibility
All statistical results were based on three independent measurements (unless indicated differently), and data are presented as the mean ± standard deviation values. No statistical methods were used to predetermine sample size. No randomization was performed in this study. The investigators were not blinded to allocation during experiments and outcome assessment.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Source data
Acknowledgements
We thank K. Liu for assistance with solution preparation and helpful comments during manuscript revision, J.-G. Li for nanopore preparation, W. Liu for discussion on enzyme digestion, M.-Y. Li for the valuable discussions on manuscript preparation.
Author contributions
Y.-L.Y. and Y.-T.L. conceived the concept and supervised the project. Y.-H.F. and X.L. performed the experiments. N.-N.W. designed the machine learning algorithms. Y.-H.F. designed other data analysis algorithms for analyzing feature correlation and recognizing peptide sequences. Y.-H.F., N.-N.W., and Y.-L.Y. performed data analysis. K.-L.X. conducted the molecular dynamics simulations. Y.-H.F., N.-N.W., L.-M.Z., C.Y., F.Y., Y.-L.Y., and Y.-T.L. prepared the figures and the manuscript. All authors discussed the findings and revised the manuscript.
Peer review
Peer review information
Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.
Funding
This study was supported by the National Key Research and Development Program of China (2023YFF1205802 to L.-M.Z.), National Natural Science Foundation of China (22525403 to Y.-L.Y., 22334006 to Y.-T.L., and 22090054 to C.Y.), Fundamental Research Funds for the Central Universities (020514380356 to Y.-L.Y.), and Jiangsu Province “Double First-Class” Initiative Grant (0205-1480601101 to Y.-L.Y.).
Data availability
Data supporting the findings of this study are given in the main text and the Supplementary Information, which are also available within the Source data provided with this paper. Source data are available via Zenodo at https://doi.org/10.5281/zenodo.21203081 (ref. 70). Source data are provided with this paper.
Code availability
The source code is available at this link https://doi.org/10.5281/zenodo.21203081.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Ying-Huan Fu, Nan-Nan Wei.
Contributor Information
Cheng Yang, Email: cyang@nju.edu.cn.
Yi-Tao Long, Email: yitaolong@nju.edu.cn.
Yi-Lun Ying, Email: yilunying@nju.edu.cn.
Supplementary information
The online version contains Supplementary material available at https://doi.org/10.1038/s41467-026-75942-5.
References
- 1.Fowler, D. M. et al. High-resolution mapping of protein sequence-function relationships. Nat. Methods7, 741–746 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Liu, S. et al. Structures of wild-type and H451N mutant human lymphocyte potassium channel KV1.3. Cell Discov.7, 39 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Edman, P. Method for determination of the amino acid sequence in peptides. Acta Chem. Scand.4, 283–293 (1950). [Google Scholar]
- 5.Deol, H. et al. After 75 years, an alternative to Edman degradation: a mechanistic and efficiency study of a base-induced method for N-terminal peptide sequencing. J. Am. Chem. Soc.147, 13973–13982 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Aebersold, R. & Mann, M. Mass spectrometry-based proteomics. Nature422, 198–207 (2003). [DOI] [PubMed] [Google Scholar]
- 7.Chait, B. T. Mass spectrometry: bottom-up or top-down? Science314, 65–66 (2006). [DOI] [PubMed] [Google Scholar]
- 8.Easterling, L. F., Yerabolu, R., Kumar, R., Alzarieni, K. Z. & Kenttämaa, H. I. Factors affecting the limit of detection for HPLC/tandem mass spectrometry experiments based on gas-phase ion–molecule reactions. Anal. Chem.92, 7471–7477 (2020). [DOI] [PubMed] [Google Scholar]
- 9.Domon, B. & Aebersold, R. Options and considerations when selecting a quantitative proteomics strategy. Nat. Biotechnol.28, 710–721 (2010). [DOI] [PubMed] [Google Scholar]
- 10.Timp, W. & Timp, G. Beyond mass spectrometry, the next step in proteomics. Sci. Adv.6, eaax8978 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Zhao, C. et al. Direct and continuous monitoring of multicomponent antibiotic gentamicin in blood at single-molecule resolution. ACS Nano18, 9137–9149 (2024). [DOI] [PubMed] [Google Scholar]
- 12.Prajapati, J. D. & Kleinekathöfer, U. Voltage-dependent transport of neutral solutes through nanopores: a molecular view. J. Phys. Chem. B124, 10718–10731 (2020). [DOI] [PubMed] [Google Scholar]
- 13.Pastoriza-Gallego, M. et al. Dynamics of unfolded protein transport through an aerolysin pore. J. Am. Chem. Soc.133, 2923–2931 (2011). [DOI] [PubMed] [Google Scholar]
- 14.Iesu, L. et al. Single-molecule nanopore sensing of proline cis/trans amide isomers. Chem. Sci.16, 9730–9738 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Piguet, F. et al. Identification of single amino acid differences in uniformly charged homopolymeric peptides with aerolysin nanopore. Nat. Commun.9, 966 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Sauciuc, A., Morozzo della Rocca, B., Tadema, M. J., Chinappi, M. & Maglia, G. Translocation of linearized full-length proteins through an engineered nanopore under opposing electrophoretic force. Nat. Biotechnol.42, 1275–1281 (2024). [DOI] [PubMed] [Google Scholar]
- 17.Yu, L. et al. Unidirectional single-file transport of full-length proteins through a nanopore. Nat. Biotechnol.41, 1130–1139 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Zhang, S. et al. Bottom-up fabrication of a proteasome–nanopore that unravels and processes single proteins. Nat. Chem.13, 1192–1199 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Liu, P., Honda, M. & Kawano, R. Determination of the position of DNA methylation and base modifications using a biological nanopore. Small Methods9, 2401760 (2025). [DOI] [PubMed] [Google Scholar]
- 20.Cherf, G. M. et al. Automated forward and reverse ratcheting of DNA in a nanopore at 5-Å precision. Nat. Biotechnol.30, 344–348 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Manrao, E. A. et al. Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase. Nat. Biotechnol.30, 349–353 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wang, F. et al. MoS2 nanopore identifies single amino acids with sub-1 Dalton resolution. Nat. Commun.14, 2895 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Wang, K. et al. Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nat. Methods21, 92–101 (2024). [DOI] [PubMed] [Google Scholar]
- 24.Zhang, M. et al. Real-time detection of 20 amino acids and discrimination of pathologically relevant peptides with functionalized nanopore. Nat. Methods21, 609–618 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Zhang, Y. et al. Peptide sequencing based on host–guest interaction-assisted nanopore sensing. Nat. Methods21, 102–109 (2024). [DOI] [PubMed] [Google Scholar]
- 26.Lucas, F. L. R., Versloot, R. C. A., Yakovlieva, L., Walvoort, M. T. C. & Maglia, G. Protein identification by nanopore peptide profiling. Nat. Commun.12, 5795 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Stierlen, A. et al. Nanopore discrimination of coagulation biomarker derivatives and characterization of a post-translational modification. ACS Cent. Sci.9, 228–238 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Zhao, X., Qin, H., Tang, M., Zhang, X. & Qing, G. Nanopore: emerging for detecting protein post-translational modifications. Trends Anal. Chem.173, 117658 (2024). [Google Scholar]
- 29.Li, Z. et al. Nanopore-based high-resolution detection of multiple post-translational modifications in protein. Angew. Chem. Int. Ed.64, e202423801 (2025). [DOI] [PubMed] [Google Scholar]
- 30.Wang, J. et al. Identification of single amino acid chiral and positional isomers using an electrostatically asymmetric nanopore. J. Am. Chem. Soc.144, 15072–15078 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Abraham Versloot, R. C. et al. Seeing the invisibles: detection of peptide enantiomers, diastereomers, and isobaric ring formation in lanthipeptides using nanopores. J. Am. Chem. Soc.145, 18355–18365 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Ratinho, L., Bacri, L., Thiebot, B., Cressiot, B. & Pelta, J. Identification and detection of a peptide biomarker and its enantiomer by nanopore. ACS Cent. Sci.10, 1167–1178 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Zhang, H. et al. Chiral discrimination of all proteinogenic amino acid enantiomers by nanopore sensing. Angew. Chem. Int. Ed.64, e202515531 (2025). [DOI] [PubMed] [Google Scholar]
- 34.Meyer, N. et al. Discrimination of oxytocin, a behavioral neuropeptide hormone, and its structural variants by nanopore. ACS Nano19, 28690–28701 (2025). [DOI] [PubMed] [Google Scholar]
- 35.Brinkerhoff, H., Kang, A. S. W., Liu, J., Aksimentiev, A. & Dekker, C. Multiple rereads of single proteins at single–amino acid resolution using nanopores. Science374, 1509–1513 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Chen, Z. et al. Controlled movement of ssDNA conjugated peptide through Mycobacterium smegmatis porin A (MspA) nanopore by a helicase motor for peptide sequencing application. Chem. Sci.12, 15750–15756 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Motone, K. et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature633, 662–669 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Yan, S. et al. Single molecule ratcheting motion of peptides in a Mycobacterium smegmatis porin A (MspA) nanopore. Nano Lett.21, 6703–6710 (2021). [DOI] [PubMed] [Google Scholar]
- 39.Lu, C., Bonini, A., Viel, J. H. & Maglia, G. Toward single-molecule protein sequencing using nanopores. Nat. Biotechnol.43, 312–322 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Li, J.-G., Ying, Y.-L. & Long, Y.-T. Aerolysin nanopore electrochemistry. Acc. Chem. Res.58, 517–528 (2025). [DOI] [PubMed] [Google Scholar]
- 41.Niu, H., Li, M.-Y., Ying, Y.-L. & Long, Y.-T. An engineered third electrostatic constriction of aerolysin to manipulate heterogeneously charged peptide transport. Chem. Sci.13, 2456–2461 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Li, X., Ying, Y.-L., Fu, X.-X., Wan, Y.-J. & Long, Y.-T. Single-molecule frequency fingerprint for ion interaction networks in a confined nanopore. Angew. Chem. Int. Ed.60, 24582–24587 (2021). [DOI] [PubMed] [Google Scholar]
- 43.Fu, Y.-H. et al. Exploring the single-molecule transient interactions with nanopore frequency spectrum. J. Phys. Chem. C128, 1110–1115 (2024). [Google Scholar]
- 44.Loh, W.-Y. Regression tress with unbiased variable selection and interaction detection. Stat. Sin.12, 361–386 (2002). [Google Scholar]
- 45.Helbig, A. O. & Tholey, A. Exopeptidase assisted N-and C-terminal proteome sequencing. Anal. Chem.92, 5023–5032 (2020). [DOI] [PubMed] [Google Scholar]
- 46.Ouldali, H. et al. Electrical recognition of the twenty proteinogenic amino acids using an aerolysin nanopore. Nat. Biotechnol.38, 176–181 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Jiang, J. et al. Protein nanopore reveals the renin–angiotensin system crosstalk with single-amino-acid resolution. Nat. Chem.15, 578–586 (2023). [DOI] [PubMed] [Google Scholar]
- 48.Niu, H. et al. Direct mapping of tyrosine sulfation states in native peptides by nanopore. Nat. Chem. Biol.21, 716–726 (2025). [DOI] [PubMed] [Google Scholar]
- 49.Li, M.-Y. et al. Revisiting the origin of nanopore current blockage for volume difference sensing at the atomic level. JACS Au1, 967–976 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Redit, C., Cha, S. & Ai, N. Single-cell proteomics: challenges and prospects. Nat. Methods20, 317–318 (2023). [DOI] [PubMed] [Google Scholar]
- 51.Liu, M. et al. Protein isoaspartate methyltransferase-mediated 18O-labeling of isoaspartic acid for mass spectrometry analysis. Anal. Chem.84, 1056–1062 (2012). [DOI] [PubMed] [Google Scholar]
- 52.Romero-Severson, J. et al. A seed-endophytic Bacillus safensis strain with antimicrobial activity has genes for novel bacteriocin-like antimicrobial peptides. Front. Microbiol.12, 734216 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Sung, Y.-S., Khvalbota, L., Dhaubhadel, U., Špánik, I. & Armstrong, D. W. Teicoplanin aglycone media and carboxypeptidase Y: tools for finding low-abundance D-amino acids and epimeric peptides. Chirality35, 461–468 (2023). [DOI] [PubMed] [Google Scholar]
- 54.Pastoriza-Gallego, M. et al. Evidence of unfolded protein translocation through a protein nanopore. ACS Nano8, 11350–11360 (2014). [DOI] [PubMed] [Google Scholar]
- 55.Burkhart, J. M., Schumbrutzki, C., Wortelkamp, S., Sickmann, A. & Zahedi, R. P. Systematic and quantitative comparison of digest efficiency and specificity reveals the impact of trypsin quality on MS-based proteomics. J. Proteom.75, 1454–1462 (2012). [DOI] [PubMed] [Google Scholar]
- 56.Greive, S. J., Bacri, L., Cressiot, B. & Pelta, J. Identification of conformational variants for bradykinin biomarker peptides from a biofluid using a nanopore and machine learning. ACS Nano18, 539–550 (2024). [DOI] [PubMed] [Google Scholar]
- 57.Wu, X.-Y. et al. Precise construction and tuning of an aerolysin single-biomolecule interface for single-molecule sensing. CCS Chem.1, 304–312 (2019). [Google Scholar]
- 58.Humphrey, W., Dalke, A. & Schulten, K. VMD: visual molecular dynamics. J. Mol. Graph.14, 33–38 (1996). [DOI] [PubMed] [Google Scholar]
- 59.Hayashi, R., Bai, Y. & Hata, T. Kinetic studies of carboxypeptidase Y: I. Kinetic parameters for the hydrolysis of synthetic substrates1. J. Biochem.77, 69–79 (1975). [PubMed] [Google Scholar]
- 60.Hayashi, R., Moore, S. & Stein, W. H. Serine at the active center of yeast carboxypeptidase. J. Biol. Chem.248, 8366–8369 (1973). [PubMed] [Google Scholar]
- 61.Qiang, J., Xu, Z., Li, Y., Wang, H. & Zhang, Y. Carboxypeptidase Y assisted disulfide-bond identification with linearized database search. Anal. Chem.93, 14940–14945 (2021). [DOI] [PubMed] [Google Scholar]
- 62.Sigma-Aldrich. Enzymatic assay of carboxypeptidase Y, https://www.sigmaaldrich.com/deepweb/assets/sigmaaldrich/product/documents/189/273/c3888enz.pdf (1994).
- 63.Liu, S.-C. et al. An advanced optical–electrochemical nanopore measurement system for single-molecule analysis. Rev. Sci. Instrum.92, 121301 (2021). [DOI] [PubMed] [Google Scholar]
- 64.Hu, Z.-L. et al. Real-time and accurate identification of single oligonucleotide photoisomers via an aerolysin nanopore. Anal. Chem.90, 4268–4272 (2018). [DOI] [PubMed] [Google Scholar]
- 65.Smeets, R. M. M., Keyser, U. F., Dekker, N. H. & Dekker, C. Noise in solid-state nanopores. Proc. Natl. Acad. Sci. USA105, 417–421 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Fushiki, T. Estimation of prediction error by using K-fold cross-validation. Stat. Comput.21, 137–146 (2011). [Google Scholar]
- 67.Schober, P., Boer, C. & Schwarte, L. A. Correlation coefficients: appropriate use and interpretation. Anesth. Analg.126, 1763–1768 (2018). [DOI] [PubMed] [Google Scholar]
- 68.Austin, P. C. & Tu, J. V. Bootstrap methods for developing predictive models. Am. Stat.58, 131–137 (2004). [Google Scholar]
- 69.Bridle, J. Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. Adv. Neural Inf. Process. Syst. 2, 211–217 (1989).
- 70.Fu, Y.-H. et al. Exopeptidase-assisted nanopore peptide sequence identification. Zenodo 10.5281/zenodo.21203081 (2026). [DOI] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data supporting the findings of this study are given in the main text and the Supplementary Information, which are also available within the Source data provided with this paper. Source data are available via Zenodo at https://doi.org/10.5281/zenodo.21203081 (ref. 70). Source data are provided with this paper.
The source code is available at this link https://doi.org/10.5281/zenodo.21203081.
