Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Apr 18.
Published in final edited form as: Nat Biotechnol. 2025 Mar 17;43(3):312–322. doi: 10.1038/s41587-025-02587-y

Toward single-molecule protein sequencing using nanopores

Chunzhe Lu 1, Andrea Bonini 1, Jakob Viel 1, Giovanni Maglia 1,*
PMCID: PMC12006967  NIHMSID: NIHMS2069637  PMID: 40097683

Abstract

Over the past three decades, biological nanopore sequencing has grown from a research curiosity to a mature technology to sequence nucleic acids at the single molecule level. Now, recent achievements suggest nanopores might be able to sequence proteins soon. In this perspective, we analyze the different approaches that have been proposed to measure proteins and peptides using nanopores. We predict that, more likely than not, nanopores will be capable of identifying full length proteins at the single-molecule level and with single amino acid resolution, paving the way to single-molecule protein sequencing. This would allow several applications in proteomics that are at present challenging, including measuring the heterogeneity of post translational modifications, quantifying low abundance proteins, and characterizing protein splicing.


The sequencing of the human genome has been a scientific landmark achievement, and it has had major implications for our understanding of human health.1 However, excluding contributing factors from the epigenetic state of the genome, all 200 cell types in our body have the same DNA. What finally distinguishes them are the proteins they produce. Proteins determine the correct functioning of our cells, playing a pivotal role in health, disease, diagnoses, and potential cures. For this reason, it is essential to not only understand our genome, but also all the proteins that are expressed in our body, our proteome. However, accessing our proteome like we did for our genome poses a considerable challenge.

Compared to DNA, proteins are more complex in many ways. Instead of having four bases and a uniform charge as DNA, proteins are made of 20 amino acids with diverse charge contents and chemical compositions. Moreover, multiple proteins may be synthesized from a single gene through alternative mRNA splicing and post translational modifications (PTMs). These so-called proteoforms and their distribution are crucial for functionality, and they substantially increase the challenge of protein sequencing. Effectively, identifying just part of a protein, or even its whole amino acid sequence, doesn’t accurately represent the set of proteoforms produced from a specific gene. Finally, some proteins and proteoforms in our cells are several orders of magnitude more abundant than others. This dynamic range complicates the detection of proteins that are present in small amounts. A protein sequencing method that can overcome all these challenges would lay the groundwork for many new fields of study.

Currently, most proteomic analysis is done with bottom-up mass spectrometry (MS). In this method, protein samples from cell extracts are enzymatically digested, the resulting peptides separated by chromatography, and the masses in the resulting peptide-mixture are determined.2 With the genetic or proteomic information of the protein sample available, the predicted in silico digestion products are compared with the detected fragments, and the proteins from which the fragments originated are identified.

To sequence proteins without reference data, de novo protein sequencing, a combination of Edman degradation and tandem MS (MS/MS) is used.2,3 In Edman degradation, a peptide sequence is determined one residue at the time, by the chemical removal and subsequent identification of amino acids from the N-terminus.4 It requires purified protein samples with concentrations in the low picomolar range and can sequence up to ~30 amino acids in length.5 With MS/MS de novo sequencing, the amino acid sequence of a protein can be identified through fragmentation and subsequent fragment-mass determination.6 Here, a bottom-up approach can also be used.7 A purified protein is pre-digested, after which the resulting peptides are separated by chromatography. The peptides are measured by a first analyzer (MS1), then fragmented, and finally sequenced by a second analyzer (MS2).

The methods described above give an incredible insight into our proteome, but they generally require purification and fragmentation methods that cause proteoform information to be convoluted or lost.8 To tap into this extra layer of information, methods to address an entire protein are being developed, like top-down MS/MS, where intact rather than pre-digested protein samples are measured.9 Although this method is a major improvement for proteoform identification, locating PTMs to specific amino acids for each proteoform remains for now challenging.9 Top-down MS/MS, already employed for single cell proteomics, still faces challenges with dynamic range, resolution, and coverage.9,10

Even though developments in the field of top-down proteomics are promising, considering the clinical relevance of de novo protein sequencing, an alternative approach is warranted. A major challenge in proteomic analysis is that proteins cannot be amplified. This is often a drawback, as many relevant proteins are present in heterogeneous mixtures and in low abundance. These proteins are typically missed by the ensemble methods used in discovery-based proteomics. Here, single molecule nanopore sequencing, especially when coupled to high throughput analysis, can play an important role.

Nanopores used for sequencing

A nanopore is a small water-filled nanoscale aperture on an insulating membrane. Nanopores may be drilled in solid-state membranes,11 or can be made of proteins,12,13 DNA14 or peptides15. To date, nanopores that are successfully used in DNA sequencing applications are made of proteins. These ‘biological nanopores’ are proteins that play a crucial role in a variety of biological processes, facilitating the regulated translocation of ions, water, small molecules, and other substances. Advantages of biological nanopores are that they can be made cheaply and reproducibly, while not changing over the time of the measurement, and they can be engineered with atomic precision. Therefore, here we only consider biological nanopores (hereafter nanopores) for sequencing.

The size of nanopores can vary considerably, ranging from a few angstroms to tens of nanometres. Typically, nanopores with a narrow constriction such as Mycobacterium smegmatis porin A (MspA) and Curlin sigma S-dependent growth subunit G pore (CsgG) are used for polymer sequencing, whereas nanopores with an extended barrel such as alpha hemolysin (α-HL), aerolysin, and cytotoxin K (CytK), and alpha helical nanopores such as fragaceatoxin C (FraC) are used for peptide analysis (Figure 1ad). Although at present only a limited number of nanopores can be found in nature, computer designed nanopores might also be generated,1618 suggesting that nanopores with bespoke size and shapes will soon become available.

Figure 1. Nanopore analysis.

Figure 1.

Structure of MspA (a), CsgG (b), FraC (c), α-HL (d), aerolysin (e) and CytK (f) nanopores, which are widely used in DNA and protein sequencing applications. g, Schematic representation of a polypeptide translocating across a nanopore (grey) embedded into an amphipathic membrane (yellow). When a potential is applied, ions (e.g., K+ and Cl-) are transported across the nanopore. h, An idealized example of an ionic current signal generated by a single nanopore inserted into the amphipathic membrane. When a polymer translocates across the nanopore, fluctuation of the signal can provide information about the polymer.

Over the past decades, nanopores have been developed to identify molecules, study chemical and enzymatic reactions and, most notably, to sequence nucleic acids at the single-molecule level (Box 1). In nanopore analysis, a membrane separates two compartments, named cis and trans. A potential is applied across the membrane, which generates an electric field and a resulting current of hydrated ions across the nanopore (Figure 1g,h). Molecules lodged inside the nanopore generate ‘blockades’ providing information about the molecule’s structure, translocation dynamics, and chemical identity. This current change highly depends on the type of nanopore used and the analyte investigated.

Box 1: Nanopore DNA sequencing.

Arguably the biggest success of nanopore analysis has been the sequencing of single DNA molecules. From its initial conceptualization in 198957 to the commercialization of the first portable nanopore DNA sequencing device - Oxford Nanopore Technologies’ MinION - in 2014,58 nanopore DNA sequencing has taken 25 years to be realized. Crucial discoveries were the ability of the applied potential to induce the translocation of single strand DNA across a nanopore,59 the recognition of two60 and then four61 DNA bases in immobilized DNA strands (and four individual deoxyribonucleoside 5’-monophosphates62), and the stepwise motion of DNA by enzymes.63,64 Beside the above milestone achievements, however, the success of nanopore DNA sequencing relies on several key elements. First, a high-resolution nanopore is required. For example, nanopores such as MspA and CsgG have a relatively narrow constriction (~1nm), which allow large differences between nucleobases to be detected. Second, a driving force must be used to facilitate the threading of the DNA strand and to keep the strand linearized throughout the nanopore. Owing to the uniform and negative charge properties of DNA, this force can be given by an externally applied potential. Third, a double stranded DNA strand must first be unwound, and a resulting single strand passed through a nanopore in a base-by-base manner. This task is provided by molecular motor proteins such as DNA polymerase phi2964, helicase 308,65 and helicase Dda.57 A complication is that multiple bases are read simultaneously at the constriction of nanopores (i.e., >4 by MspA).66 Hence, a final challenge has been the deconvolution of the signal into a sequence. Key to this process is the unidirectional stepwise motion of DNA across the nanopore allowing the same base to be addressed multiple times. Using this method, single DNA molecules can now be sequenced with an over 99% precision during single nanopore passes.67

Figure Box 1.

Figure Box 1.

CsgG nanopore used for DNA sequencing. a, Schematic diagram of DNA sequencing using the CsgG nanopore (PDB 4UV3) and ATP dependent helicase Dda (PDB 3UPU). The helicase unwinds a double stranded DNA molecule and delivers the resulting single DNA strand into the CsgG nanopore. The sensing region of this nanopore covers up to five consecutive nucleotides. b, Example current readout of nanopore DNA sequencing signal.

Challenges of using nanopores to address single proteins

Taking DNA sequencing as a benchmark (Box 1), the sequencing of proteins with nanopores requires overcoming two fundamental challenges: identifying the 20 natural amino acids and the plethora of PTMs, and transporting polypeptides unidirectionally in a linear manner across nanopores at a speed that is compatible with reading individual amino acids.

The issue of identifying twenty amino acids

Over the past three decades, many studies revealed that nanopores can differentiate among molecules, including single amino acids in peptides,1923 unfolded proteins24 or folded proteins,25 and PTMs.24,2632 Most strikingly, differences between two enantiomers, the smallest chemical difference in two molecules, have been reported using several nanopores.26,3335 This suggests that nanopore currents should be capable of identifying any difference between two amino acids in protein.

Several studies have investigated whether all 20 amino acids can be distinguished simultaneously by the same nanopore. Initial work used a wild-type aerolysin nanopore sampled the 20 proteinogenic amino acids attached at the N-terminus of a poly-R7 sequence.36 It was found that, assuming a residence time to 200 ms, 16 peptide species could be identified with a probability of 90% or higher. Some amino acids (such as methionine and tyrosine) could be further separated through selective chemical modifications.36 Besides, two recent studies demonstrated that both a Ni2+-modified MspA37 and a copper (II)-functionalized MspA38 can differentiate among the 20 individual amino acids, albeit many amino acids showed overlapping Gaussian distributions. Immobilized strands might be used to mimic the pausing of enzymes between steps. A recent work with α-HL nanopores used a host-guest interaction to immobilize a set of 20 peptides containing a poly-anionic tail differing by the 20 proteinogenic amino acids at a specific position. Despite some amino acids showing overlap in their Gaussian distributions, the authors found that 14 amino acids could be differentiated by the wild-type pore, with the remaining 6 distinguished using an engineered pore.39

Another research focus is whether all individual amino acids of a polypeptide can be sequenced as it moves through the pore. The first report of amino acid resolution in moving strands used idealized peptides attached to single stranded DNA (ssDNA) stand ratcheted by a helicase and revealed that at least some amino acids (e.g., aspartic acid, glycine, and tryptophan)40 can be addressed by MspA-M2 nanopores. More recent work using idealized polypeptides moved by an unfoldase across CsgG nanopore showed that differences in individual amino acids are based on charge or steric exclusion.41 Using a rereading strategy, the authors achieved a 61% accuracy in the 20 amino acids classification.

These works showed that, overall, there might be enough bandwidth in the nanopore signal to uniquely identify most amino acids by a single nanopore. However, it is worth noting that differences between similar amino acids are not always significant during single reads and signal overlapping occurs, which may affect the measurement accuracy. In addition, in several studies amino acids could be distinguished using the mean values from multiple measurements, with overlap in the Gaussian distributions. Hence, if re-reading approaches are not used - as discussed below - the precise identification of amino acids during a single pass might be challenging. Furthermore, a yet bigger challenge might be the presence of multiple amino acids within the constriction of the nanopore (about 17 in CsgG42). This complicates the signal notably. In DNA sequencing (Box 1), the deconvolution of the signal deals with four canonical nucleobases. In proteins, the signal must account for twenty amino acids (plus modifications). Thus, de novo sequencing of proteins provides a more difficult combinatorial problem than sequencing DNA.

However, it is also worth noting that with a step size of 2 amino acids as measured for ClpX,41 the same amino acid is read multiple times, which should help the deconvolution of the signal. Furthermore, the separation and deconvolution of signals from individual amino acids might not be required. This is because a crucial advantage of proteins compared to DNA is that a proteome has a limited number of genes (~20,000 in humans), which limits the number of possible amino acid combinations. Therefore, protein sequencing might be achieved by solving three problems: 1) finding a nanopore that provides unique signals for all proteins (or peptides) in a proteome, 2) establishing a correlation between the signal and the protein sequence and 3) identifying the variations from the reference signal with single amino acid resolution.

The problem of delivery to the nanopore

A second main challenge in nanopore analysis is the delivery of a molecule to the nanopore sensing region, which is often near or around the narrower part of the nanopore - the constriction (Figure 1af). A charged molecule experiences an electrophoretic force (FEP) that dominates the diffusion of charged molecules across the nanopore. In the case of DNA, electrophoretic forces are crucial for the initial threading of the DNA into the nanopore and to linearize the DNA strand. Polypeptides have a weak and variable charge. Hence, the FEP on the polypeptides cannot be used to thread proteins across nanopores. This has long been thought to pose a fundamental limitation to nanopore protein sequencing. However, it is also known that in ion selective nanopores the external potential creates a directional fluid motion, which results in an electroosmotic flow (FEOF) on the fluid and nearby molecules. This FEOF affects the diffusion and capture of neutral molecules,43 folded proteins,44,45 and (poly)peptides,19,29,46 inside a nanopore. Therefore, the FEOF might be used to transport polypeptides across nanopores, providing that for charged polypeptides the FEOF is stronger than the FEP.

During our initial work with FraC nanopores, we recognized that peptides are captured by an FEOF.19 However, subsequent work with Aerolysin4648 showed that peptides might not translocate the nanopore against an applied potential. Meanwhile, using FraC nanopores,49 we demonstrated that the translocation of charged peptides against an applied potential was only achievable by reducing the peptide’s charge through lowering the pH to 3.8.49 This raised questions about whether an FEOF could overcome the FEP during polypeptide translocation in sequencing applications.

Full-length proteins unfolded by guanidine hydrochloride (GdmCl) have been initially studied using weakly selective α-HL or aerolysin nanopores.50,51 Recently, it was discovered that adding a charged polymer (such as DNA or polyaspartate) at the protein terminus is required for most proteins to translocate across the nanopore.52 Interestingly, molecular dynamics (MD) simulations suggested that the binding of GdmCl cations to the lumen of the nanopore induce an FEOF, facilitating protein transport.52 Another recent study showed that an FEOF generated by introducing a ring of positive charges into the barrel of a α-HL nanopore allows the capture of a folded protein elongated with a neutral polypeptide.29 However, the unfolded translocation of the captured protein, which was also observed with the weakly selective α-HL nanopore,24,5355 was promoted by both the FEP and the FEOF. In a study published at around the same time, we set out to investigate whether unraveled proteins without additional tags could be transported across a nanopore against an FEP. We found that if a CytK nanopore’s inner surface is lined with at least three rings of charges, an FEOF can be engineered that is strong enough to transport unstructured polypeptides and urea-unfolded proteins against relatively strong FEP.56 This work proved that the FEOF can be used on proteins in a similar way as the FEP is used on DNA, indicating that no fundamental limitation exists for the sequencing of proteins using nanopores.

Approaches to sequence proteins using nanopores

Several approaches have been formulated for the identification of proteins using nanopores. First, a distinction might be made whether information from single proteins is retrieved. This is an important distinction, because despite nanopores being single-molecule sensors, the information extracted is often aggregated over many reads. Extracting sequence information from single reads is relevant, for example, to establish the heterogeneity of a protein sample or to measure complex mixtures. Nanopore devices for ensemble protein analysis could provide portable, low-maintenance, low-cost, and rapid proteomics measurements when MS is not available. Here we will focus on approaches that allow the identification of single proteins or peptides, an important focus of proteomic analysis.

A second distinction might be made whether full-length proteins or peptides are identified. Peptides are important in many biological processes (e.g., in immunology), and sequencing them through MS analysis is often challenging. A single-molecule peptide sequencing approach would also provide a complementary method to address the limitation of MS-based bottom-up proteomics, the workhorse of proteomics analysis. The sequencing of full-length proteins, however, is likely to deliver the most complete information about a protein. Full-length sequencing methods will preserve information about the heterogeneity in the polypeptide chain, allowing to address complex information such as the exact distribution of PTMs or information about spliced proteins. This information is often obscured by the sample fragmentation or inherently convoluted in bottom-up MS analysis.9

Chop-and-drop analysis of full-length proteins

A starting point for protein identification is through protein fingerprinting. We68 and others69 showed that proteins pre-digested with specific peptidases (e.g., trypsin) are identified according to their nanopore peptide spectrum or fingerprint. This approach - named nanopore peptide spectrometry21 - may require a robust database if broadly applied. However, it was also shown by others21 and our group49 that nanopore signals correlate with specific properties of peptides such as their volume. Therefore, when genetic or proteomic information of the protein sample is available, predicted in silico digestion products might be used to identify proteins de novo.

An alternative method called exopeptidase sequencing, which is the equivalent of a method proposed for DNA sequencing,62 has been proposed where amino acids are sequentially released from a peptide C- or N-terminus using an exopeptidase, and subsequently identified. This method, which was initially conceptualized to be used with MS,70 might also be used with nanopore technology.71 Exopeptidase nanopore sequencing is an enzymatic-nanopore equivalent of Edman degradation, albeit with a potential for higher speed and throughput, and lower detection limits.

Both identification methods might be turned into single-molecule approaches by attaching a peptidase above the nanopore by a chop-and-drop mechanism. Early steps have been made by our laboratory in nanopore spectrometry, in which an archaeal proteasome has been attached above a designed REG nanopore (Figure 2a).72 A VAT unfoldase has been shown to deliver proteins to the nanopore, which are gradually ‘chopped up’ into peptides. This work is a stepping stone, showing the capture of peptides from individual proteins. Bioinformatic analysis revealed that 97% of proteins in the human proteome can be identified using this approach.73

Figure 2. Approaches of nanopore peptide and protein sequencing.

Figure 2.

a, Single-molecule nanopore peptide spectrometry using a VAT unfoldase-proteasome-nanopore-based by chop-and-drop. b, Helicase-assisted peptide sequencing. c, Agent-assisted unfolding and translocation for protein identification. A capture tag is required for threading the nanopore. d, EOF-induced transport of unfolded proteins. No tag is required to induce initial transport. Unstructured polypeptides or protein unfolded by using denaturants can be used e, Protein identification by a cis-to-trans thread-and- pull approach. Proteins are electrophoretically captured and remain folded on the cis side. A ClpX unfoldase pull the substrate from trans across a 〈-HL nanopores. f, Protein sequencing by a trans-to-cis thread-and-pull approach. Substrates are first electrophoretically translocated through CsgG to the trans side and then pulled out by ClpX to the cis side. g, Protein sequencing by a thread-and-read approach using VAT unfoldase and proteasome-nanopore. Unfolded substrates are immediately captured by the nanopore.

Strand sequencing: peptides

One approach might identify polypeptides as they are transported in single file across a nanopore. In 2021, three independent groups attached a ssDNA molecule to a peptide terminus. The DNA was then used by enzymes - Hel308 helicase,40 Phi29 DNA polymerase,74 or MTA helicase75 - to move the polypeptide stepwise through a MspA-M2 nanopore. These methods have shown differences in individual amino acid substitutions on negatively charged peptides (Figure 2b). For example, using the DNA-peptide conjugate substrates, MspA-M2 nanopores can discriminate a single amino acid (aspartic acid, glycine, and tryptophan) difference in the peptide sequence using either Hel30840 or Phi29 DNA polymerase.74 Hel308 is a particularly suitable motor protein because it moves DNA by half nucleotide during ATP binding and hydrolysis, which is roughly the step size required to address individual amino acid. PTMs such as phosphorylation and sulfation in weakly charged peptides have been detected.30,76 One group75 used a ‘sandwich’ strategy, whereas a 23-amino-acid neutral peptide sequence was embedded between two DNA strands, allowing to overcome the issue of using a negative peptide strand. In this configuration, using MTA helicase, single amino acid substitutions, such as aspartic acid, glutamic acid, lysine, glutamine, and phenylalanine, as well as PTMs like phosphorylation, within a neutral glycine homopolymer environment were observed. One challenge of this approach is that positive or neutral peptides could not be stretched by the electric field, and they tangled up inside the nanopore.75 Another limitation is the reading length for peptides, at present ~25 amino acids, which depends on the length of the nanopore (MspA and CsgG are ~10 nm, Figure 1a). Using a nanopore with longer lengths should allow reading longer peptides.

Strand sequencing: full-length proteins

The ultimate goal of proteomic analysis is the ability to sequence full-length proteins. For this, proteins must be unfolded and linearly transported across a nanopore at a speed that is compatible with single amino acid analysis. Two approaches have been proposed, one of which uses an enzyme to unfold the target protein, and the other does not.

Enzyme-free characterization of a protein by consensus identification

In one approach proteins are unfolded and threaded across nanopores by electrophoretic and/or electroosmotic forces. Early work showed that in the presence of guanidine hydrochloride proteins maybe be translocated across an α-HL nanopore.50,51 Later, it was shown that an electrophoretic tag attached at a protein terminus might be required to obtain the translocation of proteins (Figure 2c).52 The threading orientation and identity of two proteins extended with poly aspartate extensions could be distinguished with accuracies larger than 90%,52 suggesting that single-translocation events may contain sufficient information to fingerprint proteins. We recently showed that nanopore with a strong FEOF allowed characterizing urea-destabilized proteins without the need of introducing electrophoretic appendices at the N- or C-termini,56 providing a simplified path to protein analysis (Figure 2d).

Native proteins may be elongated with a long unstructured tag to allow capture by the nanopore, may be unfolded spontaneously or by using a small amount of denaturant, and transported across nanopores.24,29,53,54 Differences in phosphorylation, glutathionylation, and glycosylation could be observed during the pauses of domain unfolding.24,29 However, in freely translocating strands, phosphorylated residues required chemical modifications to enhance the signal.77 Free translocation could provide a fast and efficient way to measure PTMs in proteins. However, these findings suggest that methods to slow down molecular transport across the nanopore (e.g., by engineering the nanopore’s inner surface78) and effectively stretching polypeptides may be required to identify differences in individual amino acids.

It is possible, nonetheless, that free transport across the nanopore will be too erratic to allow single pass resolution, and possibly, many of these methods will require to compile a consensus signal from the transport of multiple copies of the same protein (Figure 3a). Although this might limit their use for the characterization of simple proteomic mixtures, re-reading approaches - as described below - might increase accuracy on single molecule reads. Furthermore, this approach has the advantage of being rapid, not requiring sample preparation, and being potentially low-cost.

Figure 3. Approaches to full length nanopore protein identification.

Figure 3.

a, In consensus identification, protein signals are grouped, averaged, and then matched to known signals. Differences in the matched squiggles (e.g., because of PTMs or mutations) might be observed by variation of the signal. b, In reference-based identifications, individual protein signals are of good enough quality to be matched after a single pass. Deviations from the predicted signal are used to identify, for example, PTMs. c, In reference-based sequencing, segments of the signal are matched to known protein signals. Although multiple possible sequences are possible, the correct sequence is selected by analyzing genomic data. Stepwise motion would help recognition as groups of amino acids are read multiple times. In the example the step-size is three amino acids, and the constriction reads 5 amino acids at a time. d, In de novo sequencing, high quality signals allow the sequencing of individual amino acids. Amino-acid-by-amino-acid motion would help recognition as each amino acid is read multiple times.

Enzyme-based full-length analysis during a single pass

The identification of proteins during a single pass will allow the single-molecule characterization and sequencing of proteins in mixtures. As observed for DNA, enzymes might provide the favorable solution for controlling polypeptide translocation speed to allow single-amino-acid resolution.

In a first example, a protein containing a C-terminal ssrA tag (for recognition by an unfoldase) and a 65-amino-acid long charged polypeptide was electrophoretically threaded across an α-HL nanopore from the cis side (Figure 2e).79 Full translocation was prevented by a folded protein introduced after the charged polypeptide. Then, the unfoldase ClpX placed on the trans side was used to pull and unfold the protein across the nanopore. A follow-up work demonstrated the ability to discriminate between different proteins and their variant domains.80 This cis to trans thread-and-pull approach used the chemical energy of ATP to pull the polypeptide through the nanopore. The narrow entry of the nanopore was then used to mechanically unfold the protein against the narrow nanopore entry. However, if the unfolding of the proteins rather than the polypeptide passing the constriction of the nanopore dominates the signal, this approach might not allow the sequencing of proteins.

Another issue is that the unfoldase and the substrate protein are introduced on opposite sides of the membrane, which limits their use in commercial nanopore platforms where the trans side is sealed. This issue was tackled by electrophoretically threading an unfolded polypeptide from the cis side extended by a blocking protein containing a ssrA tag (Figure 2f).41 Then, a ClpX unfoldase subsequently added from the cis side was used to pull the unfolded polypeptide back out of the nanopore. Single point mutations were tested using 59-amino-acid repeating sequence blocks (named PASTOR) each containing a single amino acid substitution. Five consecutive PASTOR regions allowed the analysis of up to five different amino acids in a single read. These idealized polypeptide sequences allowed characterizing many important features. Ionic currents from all twenty amino acids revealed that both the size and charge of amino acids contribute to the nanopore signal. The PASTOR design allowed measuring the step size of the unfoldase (ClpX), which was found to be of two amino acids in accordance with previous studies.81 Furthermore, the authors showed the detection of mono and poly-phosphorylation, an important proof-of-concept application of nanopore analysis.

Finally, the authors tested two PASTOR substrates containing non-idealized proteins: a titin I27V15P and I27V15P,C47E,C63E domains, which was destabilized by the V15P mutation,82 and an amyloid beta peptide 1–42 (Aβ peptide) and its shorter derivative amyloid beta peptide 1–15. The protein domains could be transported to the trans side of the nanopore, where they refolded, and then pulled back out by the ClpX to the cis side. For titin, the ClpX-mediated unfolding was observed by deep current blockades most likely reflecting the occupancy of titin within the nanopore constriction, followed by a sequence-specific signal for the transport of unfolded titin domains. The nanopore signals of the different peptides and protein domains were similar but distinct features in the unfolding states could be observed.

This work strongly suggests that the transport of different proteins and their modifications across a nanopore is likely to produce unique signals. However, the present form of this trans to cis thread-and-pull method presents limitations. Notably, only unstructured or weakly folded proteins can be addressed, and the unfolding of the target protein against the nanopore might mask the sequence-specific current signal. Furthermore, since the action of the unfoldase cannot be started at the nanopore, subsequent steps of loading, washing and unfolding/reading are required to address each single protein, reducing throughput compared to free-flowing systems.

The nanopore-enzymatic machine described earlier (chop-and-drop),72 in which the proteasome is inactivated or removed, might provide a solution for many of these issues. In the thread-and-read mode (Figure 2g), the VAT unfoldase was shown to deliver proteins to the nanopore-proteasome and allow multiple unraveled translocations of folded proteins from cis-to-trans (Figure 2g) in a free-flowing manner. However, several challenges must still be overcome. The signal from the translocating protein was too short for full protein characterization. Most likely, the length of the proteasome-nanopore was too long to characterize the entire protein, the FEOF too weak to allow smooth translocation, and the nanopore too wide to show features in the translocating protein. If these issues can be solved, this approach should be compatible with commercial nanopore sequencing platforms, and it should be able to address full length proteins.

The power of resequencing

The ability of sequencing/recognizing proteins or peptides during a single pass might be challenging, as thermal fluctuations during protein translocation could make the signal noisy. Work with peptides showed that increasing the concentration of the helicase (i.e., to > 1 μM) the same peptide can be re-read multiple times.40 Similarly, the incorporation a ‘slippery’ amino acid sequence near the N terminus of a PASTOR induced the unfoldase to ‘lose grip’ on the strand allowing the protein to rethread into the pore by electrophoresis. By combining data from multiple re-reads, identification accuracy can be substantially increased, allowing for low single-molecule error rates, even when single-pass accuracy is limited, thus greatly improving the accuracy of amino acid classification. Furthermore, if single proteins can be recaptured multiple times after their first translocation, for example by reversing the applied potential as already done for DNA,83 then proteins might be identified even without an enzyme controlling the transport speed across the nanopore.

En route to protein sequencing

Although important first steps have been taken, several key technical problems must be solved to allow nanopore protein sequencing. Most enzymes that can transport proteins across a nanopore recognize a specific sequence of amino acids or require a DNA strand at the C- or N-terminus. Hence, these moieties must be attached to the protein of interest. Chemical or enzymatic methods that can modify the N- or C- terminus of proteins have been described in the literature.84,85 However, it is unknown whether they act on all proteins and /or peptides or whether they are efficient or selective enough for nanopore characterization.

Chop-and-drop methods should prove whether all generated fragments can be captured in sequence by the nanopore, and whether the signal from a single blockade can identify the molecule. Theoretical work assessing a similar method proposed for DNA highlighted the challenge of capturing individual molecules.86 Most likely, an enclosing peptidase-nanopore will be required to aid the capture of protein fragments. In this respect, exo-sequencing has the additional challenges that no exopeptidase is known with a toroidal shape, and that the FEOF is weaker on smaller molecules. Finally, no exopeptidase is known to unfold proteins, suggesting only peptides might be sequenced.

Strand sequencing approaches might be at a higher readiness level. They have shown the ability to move (poly)peptides across nanopores and to recognize at least a few amino acids in moving strands. In peptide sequencing, it will be important to show that the peptides can be stretched by the FEP. A more challenging task will be decoding the signal of the peptide translocating the nanopore. Since peptides are short (~25 amino acids can be read) and multiple amino acids are lodged within the nanopore constriction simultaneously, most likely large training sets will be required to learn the peptide signals.

Full-length strand sequencing methods are the ultimate goal of nanopore sequencing. Initial work using the thread-and-pull approach with idealized polypeptides showed that the different amino acids induce significant and specific changes to the nanopore signal. Furthermore, sequence specific signals - squiggles - from natural domains (amyloid beta and titin-I27V15P) have been shown. Although in the present state the thread-and-pull approach might only address partially structured domains and the signal from the unfolding of the protein against the nanopore might interfere with the squiggle, methods using enzymes to unfold proteins and feed the unstructured polypeptide cis to trans might overcome these issues. Nanopores with a strong FEOF have been shown to capture and thread unraveled proteins and unstructured polypeptides. Work with DNA indicated that the enzymes above the nanopore usually do not significantly affect the current signal from the translocating polymer, suggesting that the signal should only reflect polypeptides passing the nanopore constriction. The challenge is to ensure the enzyme starts the unfolding process at the nanopore rather than in solution. An important question is whether a motor protein can provide consistent and stepwise movement of the polymer to enable high accuracy protein sequencing.

De novo sequencing versus identification with single amino acid resolution

Compared to a genome, which has an astronomical number of possible nucleobase arrangements, typically, a proteome only has a few thousands of unique protein sequences. Therefore, many proteomic approaches simply require identifying proteins, ideally with single amino acid resolution (i.e., all amino acids should contribute to the signal). Hence, we expect that the simplest approaches to nanopore proteomics will be a reference-based identification approach, in which a database of squiggles from known purified proteins will be first collected and then matched to the protein analyzed (Figure 3b).

A range of applications can be envisaged. Single molecule resolution combined with high throughput analysis will allow characterization of low abundance proteins, a challenge in MS analysis. Full-length protein analysis will reveal information not easily accessible to bottom-up proteomics, such as profiling the heterogeneity of PTMs in proteins. Portable, low-cost, and high throughput devices will allow, for example, assessing the quality of protein-based formulations, such as antibodies, in real-time at the site of protein production, at the clinic or at home.

If the nanopore signal can be learned, then de novo recognition might become possible, expanding the application to all proteins that cannot be easily or cheaply pre-analyzed. In the first instance, a reference-based (de novo) sequencing is envisaged where a genetic reference will simplify the recognition of protein sequences by limiting the combinatorial space of the number of possible amino acids (Figure 3c). Eventually, if all 20 amino acids provide distinguishable signals in a moving strand and an unfoldase is found to move processively ideally one amino acid at the time (a possible requirement to reduce the signal complexity), de novo protein sequencing should also become possible even without a genetic blueprint (Figure 3d).

Conclusion

The development of DNA sequencing technologies that allowed the low-cost sequencing of the human genome has brought a revolution in science and medicine. The ability to understand the proteome at the single amino acid level will likely bring a similar outcome. Industry efforts are now underway to make single-molecule proteomics a reality (Box 2). Nanopores have advantages compared to other techniques, especially those based on fluorescence, as they process native and intact molecules, they do not require costly antibody recognition, and they might eventually be capable of de novo sequencing of proteins. Furthermore, the nanopore signal can be easily interfaced with low-cost silicon-based devices, and arrays of hundreds of thousands nanopores allow high-throughput analysis.

Box 2. Industry efforts focusing on single-molecule protein sequencing.

Commercial efforts are underway to offer single-molecule sequencing of proteins and peptides using nanopores. Portal Biotech is developing full-length protein sequencing methods87 based on the published technologies from our laboratory.56,72 Oxford Nanopore Technology is exploring peptide and protein sequencing methods88 like the one recently published for peptide40,74,75 and full-length protein analysis.41

A few other single-molecule approaches, based on different reading technologies, are already offering partial peptide sequencing. Quantum SI technology offers a solution that relies on the identification of the N-terminus of immobilized peptides using three different fluorescently labelled recognizers (typically binding one or more amino acids at time). The recognition is obtained by decoding parameters such as their temporal order and on/off binding kinetics, allowing the identification of several amino acids and some PTMs. Exopeptidases are then used to remove the last amino acid in the peptide.89 In a similar fashion, Encodia uses affinity reagent binders to recognize the N-terminus of an immobilized peptide. The binders, however, introduce a DNA barcode tag to create a peptide-DNA chimera molecule that is eventually sequenced after the last amino acid to be identified is removed. Erisyon technology chemically labels a subset of amino acids with fluorophores. The entire peptide’s fluorescence is read before a step of Edman degradation is performed to remove one amino acid at the N-terminus. These approaches take advantage of the massively parallel platforms developed for the second-generation DNA sequencing technologies but have a main limitation that only peptides can be sequenced. The Nautilus platform offers the possibility to identify full-length proteins using a short epitope mapping (PrISM) labelling technique,90 where arrays of billions of single proteins are immobilized on DNA pads91 and different cycles of the multi-affinity PrISM library probe are then probed. The platform translates single interaction events into protein counts and can recognize billions of different single-molecule proteins in one measuring run. For sequencing proteins, the development of binders that identify single amino acids within an intact and unfolded polypeptide strand might be required.

Single-molecule protein identification appears to be on the cusp of realization, already promising a wealth of applications in proteomics, including full-length PTM profiling. De novo protein sequencing may require overcoming additional challenges. First, the rich nanopore signal should be translated into sequence information - a task for which artificial intelligence tools might prove especially suited. Second, it might be necessary to bioengineer the molecular components required for sequencing proteins. Existing nanopores appear to already be capable of detecting the subtle chemical differences among the 20 amino acids, but bespoke engineering might be required to measure polypeptides as they move through nanopores. Under electroosmotic forces, unfoldases are expected to move substrates unidirectionally through nanopores. However, it is yet unknown whether they can move polypeptides single amino-acid-by-single-amino-acid, which might be necessary to simplify the deconvolution of the nanopore signal.

Acknowledgements

The authors A.B.C. discloses support for the research of this work from Funder NWO-VICI [grant number 192.068] and Funder NIH [grant 1R01HG012554].

Footnotes

Competing interest statement.

Giovanni Maglia is a founder, director, and shareholder of Portal Biotech Limited, a company engaged in the development of nanopore technologies.

References

  • 1.Hood L & Rowen L The human genome project: big science transforms biology and medicine. Genome Med 5, 79 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Dupree EJ et al. A Critical Review of Bottom-Up Proteomics: The Good, the Bad, and the Future of This Field. Proteomes vol. 8 Preprint at 10.3390/proteomes8030014 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.König S, Obermann WMJ & Eble JA The Current State-of-the-Art Identification of Unknown Proteins Using Mass Spectrometry Exemplified on De Novo Sequencing of a Venom Protease from Bothrops moojeni. Molecules vol. 27 Preprint at 10.3390/molecules27154976 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Edman P & Begg G A Protein Sequenator. Eur J Biochem 1, 80–91 (1967). [DOI] [PubMed] [Google Scholar]
  • 5.Chang E et al. N-Terminal Amino Acid Sequence Determination of Proteins by N-Terminal Dimethyl Labeling: Pitfalls and Advantages When Compared with Edman Degradation Sequence Analysis. J Biomol Tech 27, 61–74 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Vyatkina K De Novo Sequencing of Top-Down Tandem Mass Spectra: A Next Step towards Retrieving a Complete Protein Sequence. Proteomes vol. 5 Preprint at 10.3390/proteomes5010006 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Miller RM & Smith LM Overview and considerations in bottom-up proteomics. Analyst 148, 475–486 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Su T, Hollas MAR, Fellers RT & Kelleher NL Identification of Splice Variants and Isoforms in Transcriptomics and Proteomics. Annu Rev Biomed Data Sci 6, 357–376 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Po A & Eyers CE Top-Down Proteomics and the Challenges of True Proteoform Characterization. J Proteome Res 22, 3663–3675 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Brown KA, Melby JA, Roberts DS & Ge Y Top-down proteomics: challenges, innovations, and applications in basic and clinical research. Expert Rev Proteomics 17, 719–733 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Xue L et al. Solid-state nanopore sensors. Nat Rev Mater 5, 931–951 (2020). [Google Scholar]
  • 12.Mayer SF, Cao C & Dal Peraro M Biological nanopores for single-molecule sensing. iScience 25, 104145 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ayub M & Bayley H Engineered transmembrane pores. Curr Opin Chem Biol 34, 117–126 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Howorka S Building membrane nanopores. Nat Nanotechnol 12, 619–630 (2017). [DOI] [PubMed] [Google Scholar]
  • 15.Scott AJ et al. Constructing ion channels from water-soluble α-helical barrels. Nat Chem (2021) doi: 10.1038/s41557-021-00688-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Xu C et al. Computational design of transmembrane pores. Nature 585, 129–134 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Berhanu S et al. Sculpting conducting nanopore size and shape through de novo protein design. Science (1979) 385, 282–288 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Goverde CA et al. Computational design of soluble and functional membrane protein analogues. Nature 631, 449–458 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Huang G, Willems K, Soskine M, Wloka C & Maglia G Electro-osmotic capture and ionic discrimination of peptide and protein biomarkers with FraC nanopores. Nat Commun 8, 935 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Piguet F et al. Identification of single amino acid differences in uniformly charged homopolymeric peptides with aerolysin nanopore. Nat Commun 9, 966 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chavis AE et al. Single Molecule Nanopore Spectrometry for Peptide Detection. ACS Sens 2, 1319–1328 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Afshar Bakshloo M et al. Discrimination between Alpha-Synuclein Protein Variants with a Single Nanometer-Scale Pore. ACS Chem Neurosci 14, 2517–2526 (2023). [DOI] [PubMed] [Google Scholar]
  • 23.Asandei A, Rossini AE, Chinappi M, Park Y & Luchian T Protein Nanopore-Based Discrimination between Selected Neutral Amino Acids from Polypeptides. Langmuir 33, 14451–14459 (2017). [DOI] [PubMed] [Google Scholar]
  • 24.Rosen CB, Rodriguez-Larrea D & Bayley H Single-molecule site-specific detection of protein phosphorylation with a nanopore. Nat Biotechnol 32, 179–181 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Huang G et al. PlyAB Nanopores Detect Single Amino Acid Differences in Folded Haemoglobin from Blood**. Angewandte Chemie 134, (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Ensslen T, Sarthak K, Aksimentiev A & Behrends JC Resolving Isomeric Posttranslational Modifications Using a Biological Nanopore as a Sensor of Molecular Shape. J Am Chem Soc 144, 16060–16068 (2022). [DOI] [PubMed] [Google Scholar]
  • 27.Versloot RCA et al. Quantification of Protein Glycosylation Using Nanopores. Nano Lett 22, 5357–5364 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Restrepo-Pérez L, Wong CH, Maglia G, Dekker C & Joo C Label-Free Detection of Post-translational Modifications with a Nanopore. Nano Lett 19, 7957–7964 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Martin-Baniandres P et al. Enzyme-less nanopore detection of post-translational modifications within long polypeptides. Nat Nanotechnol (2023) doi: 10.1038/s41565-023-01462-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Nova IC et al. Detection of phosphorylation post-translational modifications along single peptides with nanopores. Nat Biotechnol (2023) doi: 10.1038/s41587-023-01839-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Huo M, Hu Z, Ying Y & Long Y Enhanced identification of Tau acetylation and phosphorylation with an engineered aerolysin nanopore. Proteomics 2100041 (2021) doi: 10.1002/pmic.202100041. [DOI] [PubMed] [Google Scholar]
  • 32.Lan W-H, He H, Bayley H & Qing Y Location of Phosphorylation Sites within Long Polypeptide Chains by Binder-Assisted Nanopore Detection. J Am Chem Soc 146, 24265–24270 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kang X-F, Cheley S, Guan X & Bayley H Stochastic detection of enantiomers. J Am Chem Soc 128, 10684–5 (2006). [DOI] [PubMed] [Google Scholar]
  • 34.Abraham Versloot RC et al. Seeing the Invisibles: Detection of Peptide Enantiomers, Diastereomers, and Isobaric Ring Formation in Lanthipeptides Using Nanopores. J Am Chem Soc 145, 18355–18365 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Wang J et al. Identification of Single Amino Acid Chiral and Positional Isomers Using an Electrostatically Asymmetric Nanopore. J Am Chem Soc 144, 15072–15078 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Ouldali H et al. Electrical recognition of the twenty proteinogenic amino acids using an aerolysin nanopore. Nat Biotechnol 38, 176–181 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Wang K et al. Unambiguous discrimination of all 20 proteinogenic amino acids and their modifications by nanopore. Nat Methods 21, 92–101 (2024). [DOI] [PubMed] [Google Scholar]
  • 38.Zhang M et al. Real-time detection of 20 amino acids and discrimination of pathologically relevant peptides with functionalized nanopore. Nat Methods 21, 609–618 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Zhang Y et al. Peptide sequencing based on host–guest interaction-assisted nanopore sensing. Nat Methods 21, 102–109 (2024). [DOI] [PubMed] [Google Scholar]
  • 40.Brinkerhoff H, Kang ASW, Liu J, Aksimentiev A & Dekker C Multiple rereads of single proteins at single–amino acid resolution using nanopores. Science (1979) (2021) doi: 10.1126/science.abl4381. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Motone K et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature 633, 662–669 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Cardozo N et al. Multiplexed direct detection of barcoded protein reporters on a nanopore array. Nat Biotechnol 40, 42–46 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Gu L-Q, Cheley S & Bayley H Electroosmotic enhancement of the binding of a neutral molecule to a transmembrane pore. Proceedings of the National Academy of Sciences 100, 15498–15503 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Soskine M et al. An engineered ClyA nanopore detects folded target proteins by selective external association and pore entry. Nano Lett 12, 4895–4900 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Firnkes M et al. Electrically facilitated translocations of proteins through silicon nitride nanopores: conjoint and competitive action of diffusion, electrophoresis, and electroosmosis. Nano Lett 10, 2162–2167 (2010). [DOI] [PubMed] [Google Scholar]
  • 46.Li S, Cao C, Yang J & Long Y Detection of Peptides with Different Charges and Lengths by Using the Aerolysin Nanopore. ChemElectroChem 6, 126–129 (2019). [Google Scholar]
  • 47.Bakshloo MA et al. Polypeptide analysis for nanopore-based protein identification. Nano Res 15, 9831–9842 (2022). [Google Scholar]
  • 48.Versloot RCA, Straathof SAP, Stouwie G, Tadema MJ & Maglia G β-Barrel Nanopores with an Acidic–Aromatic Sensing Region Identify Proteinogenic Peptides at Low pH. ACS Nano 16, 7258–7268 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Huang G, Voet A & Maglia G FraC nanopores with adjustable diameter identify the mass of opposite-charge peptides with 44 dalton resolution. Nat Commun 10, 835 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Pastoriza-Gallego M et al. Dynamics of unfolded protein transport through an aerolysin pore. J Am Chem Soc 133, 2923–2931 (2011). [DOI] [PubMed] [Google Scholar]
  • 51.Oukhaled G et al. Unfolding of Proteins and Long Transient Conformations Detected by Single Nanopore Recording. Phys Rev Lett 98, 158101 (2007). [DOI] [PubMed] [Google Scholar]
  • 52.Yu L et al. Unidirectional single-file transport of full-length proteins through a nanopore. Nat Biotechnol (2023) doi: 10.1038/s41587-022-01598-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Rodriguez-Larrea D & Bayley H Multistep protein unfolding during nanopore translocation. Nat Nanotechnol 8, 288–295 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Rodriguez-Larrea D & Bayley H Protein co-translocational unfolding depends on the direction of pulling. Nat Commun (2014) doi: 10.1038/ncomms5841. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Pastoriza-Gallego M et al. Evidence of unfolded protein translocation through a protein nanopore. ACS Nano 8, 11350–11360 (2014). [DOI] [PubMed] [Google Scholar]
  • 56.Sauciuc A, Morozzo della Rocca B, Tadema MJ, Chinappi M & Maglia G Translocation of linearized full-length proteins through an engineered nanopore under opposing electrophoretic force. Nat Biotechnol (2023) doi: 10.1038/s41587-023-01954-x. [DOI] [PubMed] [Google Scholar]
  • 57.Branton D & Deamer D Nanopore Sequencing. (WORLD SCIENTIFIC, 2018). doi:doi: 10.1142/10995. [DOI] [Google Scholar]
  • 58.Oxford Nanopore Technologies. Company history. https://nanoporetech.com/about/history.
  • 59.Kasianowicz JJ, Brandin E, Branton D, Deamer DW & Deamer David W.. Characterization of individual polynucleotide molecules using a membrane channel. Proceedings of the National Academy of Sciences 93, 13770–13773 (1996). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Ashkenasy N, Sanchez-Quesada J, Bayley H & Ghadiri MR Recognizing a single base in an individual DNA strand: a step toward DNA sequencing in nanopores. Angewandte Chemie-International Edition 44, 1401–1404 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Stoddart D, Heron AJ, Mikhailova E, Maglia G & Bayley H Single-nucleotide discrimination in immobilized DNA oligonucleotides with a biological nanopore. Proc Natl Acad Sci U S A 106, 7702–7707 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Astier Y, Braha O & Bayley H Toward single molecule DNA sequencing: direct identification of ribonucleoside and deoxyribonucleoside 5’-monophosphates by using an engineered protein nanopore equipped with a molecular adapter. J Am Chem Soc 128, 1705–1710 (2006). [DOI] [PubMed] [Google Scholar]
  • 63.Cockroft SL, Chu J, Amorin M & Ghadiri MR A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution. J Am Chem Soc 130, 818–820 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Lieberman KR et al. Processive Replication of Single DNA Molecules in a Nanopore Catalyzed by phi29 DNA Polymerase. J Am Chem Soc 132, null-null (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Craig JM et al. Revealing dynamics of helicase translocation on single-stranded DNA using high-resolution nanopore tweezers. Proc Natl Acad Sci U S A 114, 11932–11937 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Derrington IM et al. Nanopore DNA sequencing with MspA. Proc Natl Acad Sci U S A 107, 16060–16065 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Oxford Nanopore Technologies. Nanopore Sequencing Accuracy. https://nanoporetech.com/platform/accuracy.
  • 68.Lucas FLR, Versloot RCA, Yakovlieva L, Walvoort MTC & Maglia G Protein identification by nanopore peptide profiling. Nat Commun 12, 5795 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Afshar Bakshloo M et al. Nanopore-Based Protein Identification. J Am Chem Soc 144, 2716–2725 (2022). [DOI] [PubMed] [Google Scholar]
  • 70.Helbig AO & Tholey A Exopeptidase Assisted N- and C-Terminal Proteome Sequencing. Anal Chem 92, 5023–5032 (2020). [DOI] [PubMed] [Google Scholar]
  • 71.Bonini A, Sauciuc A & Maglia G Engineered nanopores for exopeptidase protein sequencing. Nat Methods 21, 16–17 (2024). [DOI] [PubMed] [Google Scholar]
  • 72.Zhang S et al. Bottom-up fabrication of a proteasome–nanopore that unravels and processes single proteins. Nat Chem 13, 1192–1199 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.de Lannoy C, Lucas FLR, Maglia G & de Ridder D In silico assessment of a novel single-molecule protein fingerprinting method employing fragmentation and nanopore detection. iScience 24, 103202 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Yan S et al. Single Molecule Ratcheting Motion of Peptides in a Mycobacterium smegmatis Porin A (MspA) Nanopore. Nano Lett 21, 6703–6710 (2021). [DOI] [PubMed] [Google Scholar]
  • 75.Chen Z et al. Controlled movement of ssDNA conjugated peptide through Mycobacterium smegmatis porin A (MspA) nanopore by a helicase motor for peptide sequencing application. Chem Sci (2021) doi: 10.1039/D1SC04342K. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Chen X et al. Resolving sulfation PTMs on a plant peptide hormone using nanopore sequencing. bioRxiv 2024.05.08.593138 (2024) doi: 10.1101/2024.05.08.593138. [DOI] [Google Scholar]
  • 77.Lan W-H, He H, Bayley H & Qing Y Location of phosphorylation sites within long polypeptide chains by binder-assisted nanopore detection. bioRxiv 2024.04.29.590540 (2024) doi: 10.1101/2024.04.29.590540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Sauciuc A & Maglia G Controlled Translocation of Proteins through a Biological Nanopore for Single-Protein Fingerprint Identification. Nano Lett 24, 14118–14124 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Nivala J, Marks DB & Akeson M Unfoldase-mediated protein translocation through an alpha-hemolysin nanopore. Nat Biotechnol 31, 247–250 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Nivala J, Mulroney L, Li G, Schreiber J & Akeson M Discrimination among Protein Variants Using an Unfoldase-Coupled Nanopore. ACS Nano 8, 12365–12375 (2014). [DOI] [PubMed] [Google Scholar]
  • 81.Fei X et al. Structures of the ATP-fueled ClpXP proteolytic machine bound to protein substrate. Elife 9, (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Oguro T et al. Structural Stabilities of Different Regions of the Titin I27 Domain Contribute Differently to Unfolding upon Mitochondrial Protein Import. J Mol Biol 385, 811–819 (2009). [DOI] [PubMed] [Google Scholar]
  • 83.Gershow M & Golovchenko JA Recapturing and trapping single molecules with a solid-state nanopore. Nat Nanotechnol 2, 775–779 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Shadish JA & DeForest CA Site-Selective Protein Modification: From Functionalized Proteins to Functional Biomaterials. Matter 2, 50–77 (2020). [Google Scholar]
  • 85.Toplak A et al. From thiol-subtilisin to omniligase: Design and structure of a broadly applicable peptide ligase. Comput Struct Biotechnol J 19, 1277–1287 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Reiner JE et al. The effects of diffusion on an exonucleasenanopore-based DNA sequencing engine. Journal of Chemical Physics (2012) doi: 10.1063/1.4766363. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Portal Biotech. https://www.portalbiotech.com/.
  • 88.Proteomic methods at Oxford Nanopore Technology. https://nanoporetech.com/oxford-nanopore-proteomics.
  • 89.Reed BD et al. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device. Science (1979) 378, 186–192 (2022). [DOI] [PubMed] [Google Scholar]
  • 90.Egertson JD et al. A theoretical framework for proteome-scale single-molecule protein identification using multi-affinity protein binding reagents. bioRxiv 2021.10.11.463967 (2021) doi: 10.1101/2021.10.11.463967. [DOI] [Google Scholar]
  • 91.Aksel T et al. Highly dense and scalable protein arrays for single-molecule studies. bioRxiv 2022.05.02.490328 (2022) doi: 10.1101/2022.05.02.490328. [DOI] [Google Scholar]

RESOURCES