Abstract
The cellular proteome is complex and dynamic, with proteins playing a critical role in cell-level biological processes that contribute to homeostasis, stimuli response, and disease pathology, among others. As such, protein analysis and characterization are of extreme importance in both research and clinical settings. In the last few decades, most proteomics analysis has relied on mass spectrometry, affinity reagents, or some combination thereof. However, these techniques are limited by their requirements for large sample amounts, low resolution, and insufficient dynamic range, making them largely insufficient for the characterization of proteins in low-abundance or single-cell proteomic analysis. Despite unique technical challenges, several single-molecule protein sequencing (SMPS) technologies have been proposed in recent years to address these issues. In this review, we outline several approaches to SMPS technologies and discuss their advantages, limitations, and potential contributions toward an accurate, sensitive, and high-throughput platform.
INTRODUCTION
The role of proteins in regulating the behavior of the cell is complex and dynamic. Proteins control, signal, and catalyze cellular processes, driving both cellular phenotypes and disease pathology. As a major functional product of gene expression, proteins constitute a major component of the regulome: the proteins, mRNA, and other components in the cell that play a role in homeostasis or in responding to internal (genetic) and external (environmental) stimuli.1 As such, the proteome contains a wealth of information regarding cell-level biological processes, including protein–protein interactions, protein trafficking, protein localization, and more, that cannot be discerned by analyzing the genome or transcriptome alone.2,3
Recent single-cell analyses have revealed a surprising degree of variability between individual cells, shedding light on inherently dynamic and stochastic cell processes contributing to pathological phenotypes such as diabetes and oncogenesis.4,5 It is therefore of critical importance to the fields of synthetic biology, biomedical engineering, and personalized medicine to develop technologies capable of interrogating the proteomic content of not only an organism or tissue, but individual proteins within a single cell itself. Single-molecule protein sequencing (SMPS) technologies should therefore be capable of detecting and quantitatively sequencing all proteomic constituents of a single cell to unambiguously determine protein identity despite small starting volumes or protein abundances. Such technologies would provide unprecedented power to researchers for engineering predictive biological systems, discovering novel drug targets, and identifying candidate disease biomarkers. However, while there have been significant advances in next-generation DNA sequencing techniques, sequencing the proteomic content of a single cell has proven more difficult. Therefore, developing accurate, sensitive, and high-throughput methodologies for single-molecule protein identification requires additional engineering considerations.
Challenges for SMPS
While the genomics revolution was made possible in part due to the rapid progress in next-generation DNA sequencing technologies, high-throughput protein sequencing technologies that are both accurate and sensitive lag behind their DNA counterparts.6 This delay can be attributed to several causes, including the extreme diversity of cellular proteomes compared to the corresponding protein-coding regions of genomes, in addition to the dynamic range of cellular protein quantity.3,7
First, compared to the genome or transcriptome, the complexity and heterogeneity of the proteome pose a problem for protein sequencing technologies. While DNA is comprised of four nucleotides (A, G, C, and T), proteins are made up of at least 20 different amino acids. For sequencing technologies that rely on the use of fluorophore attachment to specific peptide constituents, this fivefold increase in potential readout complexity makes discriminating between individual signals difficult due to the subnanometer imaging resolution required and the possibility of spectral overlap. The relatively large number of amino acids also complicates any chemistry involved in residue labeling, while potentially disrupting the native protein sequence.
Second, a comprehensive proteomic analysis would consider not only gene-level variations such as residue order, but also the wide array of proteoforms that arise from a combination of alternative splicing events, single amino acid polymorphisms, and post-translational modifications (PTMs) (Fig. 1). While the human genome contains ∼20 000 non-modified (canonical) protein-coding genes,8 between 40% and 60% of human genes give rise to isoforms resulting from the alternative splicing of mRNA,9,10 increasing to 80% for genes on specific chromosomes.11 These transcriptional variants can create over 100 protein isoforms per gene on average,8 for a total number easily exceeding over one million possible proteoforms. This astounding level of heterogeneity is further complicated by minor allele frequencies of >20% resulting from single-nucleotide polymorphisms12 and the existence of more than 200 post-translational processing mechanisms, including the addition of functional groups, proteolytic processing events, and protein splicing events.13,14 Therefore, single molecule protein sequencing technologies used to detect dynamic disease states or engineer predictive biological systems must be capable of identifying proteins based not only on their canonical sequence but also their many possible isoforms.
FIG. 1.
Proteome complexity builds off of the genome and transcriptome. Relative to DNA sequencing, the dynamic complexity of the human proteome poses unique challenges for SMPS technologies. While the human genome contains ∼20 000 genes and remains relatively stable over time, subsequent transcription and mRNA processing events, such as alternative splicing of exons, alternative transcription start sites, and alternative poly-A sites, result in a dynamic transcriptome composed of ∼100 000 transcripts. These transcript variants play roles in protein synthesis, transcriptional regulation, gene silencing, and more. Following the translation of mRNA into proteins, proteolytic cleavage, addition of PTMs, and other protein processing events results in a protein population capable of exhibiting well over 1 000 000 proteoforms. Due to dynamic complexity of the human proteome, SMPS technologies must be able to accurately detect not only the result of genomic-level variations, but variation at the transcriptome and proteome levels, including more than 200 possible PTMs.
Finally, the massive parallelization employed by next-generation sequencing (NGS) techniques for DNA is not an option for protein sequencing. The first NGS/second-generation technologies such as Roche 454 in 200515 and Illumina Solexa in 200616 rely on clonal amplification of DNA. The DNA polymerase (DNAP) and short oligonucleotides employed in these methods create billions of amplified DNA molecules, allowing millions of sequencing reactions to occur in parallel.17 Amplification also allows for the detection of analytes that are low-abundance due to small sample sizes or total protein quantity by increasing analyte concentrations to reach the limit of detection. However, unlike DNA, proteins lack biochemical amplification methods such as polymerase chain reaction (PCR), creating a major engineering hurdle to achieving sensitive, high-throughput protein sequencing methods.
Compared to the relative “simplicity” of DNA, SMPS technologies require several additional engineering considerations to create a massively parallel protein sequencing technology capable of identifying and quantifying proteins for use in single-cell proteomics and medical diagnostics. In this paper, we outline the limitations of current protein identification methods and present a review of the SMPS methods currently in development that focus on sequencing or fingerprinting linearized proteins and peptides.
MASS SPECTROMETRY
A variety of protein identification and analysis methods are in wide use, including chromatography, x-ray crystallography, nuclear magnetic resonance (NMR) spectroscopy, antibody recognition, and electron microscopy, among others. Among these, mass spectrometry (MS) is the current method of choice for protein identification, quantification, and characterization without the use of affinity reagents.18 Two of the most commonly used MS approaches in proteomics research are bottom-up/peptide-centric proteomics, which utilizes proteolytic digestion to perform large-scale analyses of peptides,19,20 and top-down/native MS proteomics, which analyzes intact proteins.21 While the choice of MS technique employed depends on the question being addressed, top-down MS is better suited toward analyzing high numbers of intact proteins. However, current sequence coverage and depth of top-down methods attempting to characterize proteins are relatively low, with mean sequence coverage of ∼33% for a single sample.22 Therefore, for the purposes of this review, we focus on bottom-up/peptide-centric MS techniques, briefly outlined below.
Peptide-centric MS operates under the assumption that proteins in a sample can be identified by the mass of their constituent peptides. Conventional sample preparation for peptide-centric MS first requires enzymatically digesting a protein population into short, <10-residue peptides via sequence-specific proteases, usually trypsin (Fig. 2). The resulting peptide mixture is de-salted and further separated via liquid chromatography (LC) techniques, such as high-performance liquid chromatography (HPLC) or ion exchange.20,23,24 Both sample preparatory steps are crucial prior to the introduction of peptides to the mass spectrometer, since peptides shorter than 20 residues are more soluble and stable than native proteins, while reliable homology searches require lengths of five residues or longer.23,25 Additionally, analysis of smaller peptides reduces the chance of detecting multiple PTMs per peptide. Directly following separation, peptides are converted to gas-phase ions by electrospray ionization (ESI) or matrix-assisted laser desorption/ionization (MALDI) in a vacuum chamber. In the first phase of MS acquisition, the resulting ions are separated by the magnetic field of a mass analyzer. Here, ionized peptides are distributed by their mass-to-charge ratio (m/z) to generate a mass spectrum, where signal intensity increases with peptide quantity.20,23
FIG. 2.
Peptide-centric mass spectrometry. Peptide-centric MS operates under the assumption that proteins in a sample can be identified by the mass of their constituent peptides. Sample peptides are prepared for the mass spectrometer by first extracting proteins from cells or tissues and then digesting the protein extract using proteases with predictable cleavage patterns. Peptides are then further separated, usually on the basis of hydrophobicity, using liquid chromatography (LC) techniques to enrich for a specific population, reduce spectral complexity, and remove contaminants. Following sample preparation, the resulting peptides (<10 amino acids) are ionized for the mass spectrometer via ESI or matrix-assisted laser desorption/ionization (MALDI) and sent to the mass analyzer. In the first mass analysis (MS), precursor ions are selected and fragmented via kinetic energy by an inert gas [collision-induced dissociation (CID)] to form product ions or single/short peptides, including isomers and isobars. Product ions are then sent again to a mass analyzer for the second mass analysis, or tandem mass spectrometry (MS/MS), where resulting peptide spectra can then be compared to proteomics databases for identification.
To obtain peptide sequence information, peptide ions obtained in the MS survey scan, also known as precursor ions, are sequentially isolated for tandem MS (MS/MS). Isolated ions are then fragmented at the peptide backbone via collision-induced dissociation (CID) to produce peptide fragments, or product ions. Product ions are then introduced to the mass analyzer once more to generate a mass spectrum of individual peptides, which, as in the survey scan, can be distinguished from one another by their mass-to-charge ratio. The resulting spectral data are then either compared to simulated spectra generated in silico, or the full protein sequence is assembled from parallel digestions for identification via bioinformatics tools.20,23,26–28 While several single-cell MS technologies are currently in development (reviewed in Refs. 3, 29, and 30), here we highlight some of the challenges and shortcomings posed by MS-based protein sequencing.
Challenges faced by mass spectrometry for SMPS
Single-molecule protein identification technologies require high dynamic range to detect less-abundant proteins of clinical significance in complex biological samples. For example, protein dynamic range spans four orders of magnitude in yeast31,32 and five orders of magnitude in mouse fibroblasts, with low abundance proteins (<100 copies per cell) escaping detection.33 In humans, an analysis of protein dynamic range in plasma showed that protein analytes spanned 10 orders of magnitude, detecting picograms of the inflammation marker interleukin six in samples that also contained milligrams of serum albumin. However, MS runs into challenges when attempting to identify low-abundance proteins in complex samples due to both limited total protein quantity and technical limitations regarding dynamic range. For modern mass spectrometers, the dynamic range has low proteomic coverage and typically spans only four to five orders of magnitude, exhibiting sensitivities in the femtomole to attomole range.34,35 As low-copy-number proteins present in a cell are present at less than 1000 molecules,36 the MS technologies currently available lack the sensitivity and dynamic range to sequence clinically relevant biomarkers present in a sample at low abundance. Using MS to identify low-abundance species in complex samples requires a priori enrichment of these species via protein fractionation or affinity reagents in order to outcompete spectra generated from high-abundance species that may be irrelevant to clinical studies.35,37
In MS, chromatic co-elution of similar peptides or high-abundance peptides with similar masses can mask lower-abundance species that may be of clinical significance.26,38 While enrichment of known targets can lower their detection limit, these methods also require corresponding increases in the amount of starting sample due to the absence of protein amplification methods.35 Deep proteomic coverage of complex samples also requires large sample sizes (∼1 mg sample to detect ∼9000 proteins), with some MS-based proteomic studies using bulk samples comprising multiple cell lines.30 Therefore, these methods cannot easily account for sample heterogeneity or perform comprehensive proteomic analysis from small amounts of tissue, such as patient biopsies, much less at the single-cell level.
Complete proteomic analysis of complex samples by MS remains challenging, as conventional MS workflows yield data that are incomplete, requiring repeated analyses that generate overlapping datasets and contribute to inconsistencies in both dataset size and quality.37 However, inter-sample reproducibility can be difficult as a result of random sampling or undersampling high-abundance proteins since only a subset of precursor ions is selected in each analysis.26,38,39 Unanticipated cleavage products from initial proteolytic digestions, low signal-to-noise events, PTMs, and the lack of complete reference databases also mean that complex samples can create populations of unmatched spectra that further complicate peptide analysis and identification, with 75% of collected spectra remaining unidentified.20,40
Nanopore mass spectrometers
Despite these challenges, new efforts to engineer a mass spectrometer capable of sequencing biopolymers at single-molecule resolution by combining both MS and solid-state nanopore technologies are under way. Conventional ESI MS emits relatively large droplets containing charged molecules directly into a gas chamber, where evaporation causes droplet shrinkage. This shrinkage forces charged molecules closer together until they overcome the surface tension of the solvent droplet due to the repulsive force between molecules carrying like charges. This force of repulsion leads to a chain of explosions (Coulomb explosions) until single ions with high mass-to-charge ratios emerge.41 While these explosions are necessary to create single ions that are analyzable by the machine, they also result in large sample losses due to ion dispersal and render biopolymer monomer order impossible to determine.42,43
To overcome these issues, the Stein laboratory engineered a nanopore ion source capable of transferring single ions directly into a vacuum chamber via a sub-micrometer-diameter nozzle (<500 nm). This design builds off of previous nanospray techniques, where smaller droplets emitted by micrometer-sized tips (1–5 μm) into a vacuum result in fewer Coulomb explosions.44 By using a nanometer-sized tip, Bush et al. attempted to bypass Coulomb explosions entirely in an attempt to retain the order of the peptide. In addition to detecting larger sample amounts, the small size of the nozzle could force the biopolymer to adapt the linear configuration necessary to retain monomer sequence order for identification. While this technique has yet to translocate or sequence amino acids or peptides, it was able to detect simple salts and the DNA base cytosine in low volatility solvents such as formamide and water.42 Using the nanopore mass spectrometer, spectra generated from sodium iodide (NaI) in formamide matched previous data,45,46 with clustered sodium and iodine ion peaks separated by the molecular mass of formamide, while sodium chloride (NaCl) in water generated less complex spectra than traditional methods.47–49 Ions from one DNA base, cytosine, were also resolved in the spectrum after only 10 min, although contamination from formamide molecules was also present.42 Although the nanopore mass spectrometer has yet to sequence peptides, with improvement this technique could eventually allow for sequential detection of ionized peptide residues as they pass through the nozzle, allowing for reconstruction of the original peptide sequence. Additionally, unlike current MS workflows that rely on sequencing short (<10 residues) peptides, this technique has the potential for long-read lengths, which are less computationally demanding to reassemble in silico.
FLUORESCENCE-BASED FINGERPRINTING TECHNIQUES
Fluorosequencing using Edman degradation
Edman degradation was first published by Edman in 1949 as a technique for identifying the amino acid sequence of peptide fragments.50 Using stepwise chemical degradation, peptide sequencing occurs in an N- to C-terminus direction and proceeds through multiple experimental cycles of labeling and cleaving the N-terminal amino acid while leaving the remainder of the peptide chain intact. Specifically, the Edman reagent (phenylisothiocyanate, PITC) reacts with the N-terminal amino group of a protein under mildly alkaline conditions. Once labeled, the modified amino group is cleaved from the peptide chain under acidic anhydrous conditions, to later be extracted and identified as a PTH (phenylthiocarbamyl) derivative using chromatography or electrophoresis. The newly exposed N-terminus then serves as the site of the next degradation reaction, and the process is serially repeated to determine the full peptide sequence.51
While MS replaced Edman degradation for most polypeptide sequencing experiments by the mid-1990s,7 recent work from the Anslyn and Marcotte laboratories optimized Edman degradation for use in highly parallel single-molecule peptide sequencing. Their approach labels short (<30 residues) digested peptide fragments with a small number of covalently bound, residue-specific fluorophores and utilizes total internal reflection fluorescence (TIRF) microscopy to achieve single-molecule imaging resolution across millions of molecules per run52 [Fig. 3(a)]. As stepwise PITC degradation occurs across experimental cycles and N-terminal amino acids are cleaved from the peptide fragment, the partial sequence readout resulting from decreases in fluorescence, or fluorosequence, can be compared to a reference proteome to determine protein identity [Fig. 3(b)].
FIG. 3.
Fluorosequencing of sample peptides using Edman degradation. (a) Sample preparation first requires digestion of a protein sample, followed by peptide labeling with covalently bound fluorescent dyes (green and red stars) specific to a single or a small subset of amino acids. Labeled peptides are then immobilized on a glass surface in a perfusion chamber for imaging using TIRF microscopy. (b) A schematic of a labeled peptide undergoing consecutive rounds of fluorosequencing. The fluorescence of peptides labeled at cysteine (green star) or lysine (red star) residues is monitored and recorded prior to each Edman degradation cycle. The N-terminal residue is then cleaved from the peptide fragment via PITC degradation, exposing the next N-terminal residue and leaving the peptide one amino acid shorter. The fluorescence of the shortened peptide is again recorded using TIRF microscopy, with drops in fluorescence intensity attributed to the removal of residues bonded to the corresponding fluorophore. The cycle repeats until the anchoring residue is reached. Readouts for cysteine and lysine dyes have been consolidated into one graph.
Initial Monte Carlo simulations from Swaminathan et al. found that under ideal conditions, sequencing analyses of residues labeled with four distinct amino-acid-specific fluorophores were capable of uniquely identifying 98% of the human proteome. After accounting for the most common error rates derived from dye labeling inefficiencies, photobleaching, and failed or delayed residue removal due to inefficient degradation chemistry, proteome coverage approached 60% when peptides were labeled with only two distinct fluorophores.52
Follow-up work from Marcotte's group demonstrated experimental proof of principle of single-molecule protein identification using peptides containing fluorescently tagged cysteine, lysine, and phosphoserine residues immobilized on a glass coverslip at high picomolar concentrations. Both the instrumentation setup and fluorescent dyes used in peptide labeling were chosen for robustness based on their ability to withstand the harsh reagents used in Edman degradation for extended periods of time without otherwise negatively affecting the reaction. The correct position of labeled residues was identified in 40% of tagged peptides, with the largest source of error stemming from defective dyes or failed dye attachment (7%).34 This method was also capable of distinguishing between zeptomolar concentrations of two peptide pair mixtures carrying fluorescently tagged cysteines at different positions.
Advantages and disadvantages to fluorosequencing
Fluorosequencing utilizes existing proteome databases to identify proteins based on partial reads, therefore bypassing the need to fluorescently label every amino acid and subsequently resolve every fluorophore. This method also shares similar protein isolation and digestion methods with peptide-centric MS techniques; however, the single-molecule nature of the Edman degradation-based peptide fluorosequencer suggests that it would be amenable to smaller sample sizes in the sub-femtomolar range. Additionally, upstream proteolysis for peptide preparation can result in millions of peptide fragments, allowing for highly parallel reads from mixed samples, since millions of labeled peptides can be immobilized on the platform for imaging.
While fluorosequencing presents an all-chemical approach to SMPS, the TIRF microscopy setup employed must be chemically resistant to the harsh organic acids, bases, and heat required by the Edman degradation reaction. Errors resulting from fluorophore loss (5% per cycle) and defective dyes or failed dye attachment (7% per cycle) can also be falsely attributed to N-terminal peptide cleavage, causing erroneous residue additions during fluorosequencing. The frequency of dye-destruction events currently prevents fluorosequencing of peptides longer than 30 residues or full-length proteins, resulting in short reads that require more demanding computational analysis.
All fluorescent-based approaches to SMPS rely on known labeling chemistries between fluorophores and chemically reactive functional groups of specific amino acids or PTMs. However, the residues capable of being labeled are primarily limited to those containing amino, carboxyl, or sulfhydryl groups in their side chains; moreover, only a small number of PTMs have known labeling chemistries.53 Although the ability to identify 60% of proteins based on a subset of labeled residues somewhat ameliorates this issue, previous computational modeling demonstrates that Edman degradation-based fluorosequencing would attain higher proteome coverage with a greater number of labeled residues.52 To this end, there has been recent work toward the discovery of additional efficient and selective labeling chemistries with little cross-reactivity, specifically targeting side chains containing aryl and carboxyl groups using both solution-phase (KDYWEC) and solid-phase (KDYWE) synthesis that could potentially prove useful in improving the discriminatory power of this technique.54
Additionally, fluorosequencing is only possible when analyzing proteins with free N-terminal residues that are not buried, modified, or otherwise sterically inaccessible. Finally, the reaction process itself is orders of magnitude slower than alternate forms of SMPS, with one experimental cycle taking 90 min to cleave a single residue. While the current setup limits the length of peptides to <30 residues, efficiency improvements will be required to adapt this method for longer polypeptide chain sequencing.
ClpXP FRET fingerprinting
While Edman degradation relies solely on chemical reactions, work from the Meyer and Joo laboratories recently demonstrated a fluorescence-based approach to SMPS utilizing the bacterial protease ClpXP in conjunction with Förster resonance energy transfer (FRET) at the single-molecule level. FRET is a distance-dependent phenomenon in which a donor fluorophore transfers its excitation energy to a nearby acceptor fluorophore only if the two are within the Förster radius (3–6 nm).55,56 Using their method, it is possible to sequentially detect two residues, cysteine (C) and lysine (K), labeled with distinguishable acceptor-fluorophores as the substrate translocates through a donor-labeled ClpXP enzyme.
ClpXP is a AAA+ (ATPase associated with diverse cellular activities) barrel-shaped protein complex composed of a hexameric ATPase (ClpX) on top of two co-axially stacked heptameric proteases (ClpP).57–59 Substrate degradation initiates when the axial pore of ClpX interacts with a disordered peptide tag on a protein, usually located at the N- or C-terminus60 [Fig. 4(a)]. Subsequent mechanical forces generated by adenosine triphosphate (ATP) hydrolysis allow the ClpX hexamer to unfold and translocate the polypeptide through its central pore in a unidirectional, sequential manner61 until it reaches the ClpP peptidase chamber, where it is degraded.59,62,63 Previous work has shown ClpXP to be highly processive, capable of translocating over eight concatenated monomers spanning upward of 1000 amino acids.61
FIG. 4.
Peptide fingerprinting using ClpXP-based smFRET. (a) ClpXP is a barrel-shaped protease that utilizes ATP hydrolysis to power substrate proteolysis. A substrate bearing the disordered, 11-residue ssrA degradation tag that engages with the narrow (∼1 nm) ClpX pore (left) is captured and subsequently unfolded through ATP-driven conformational changes of ClpX pore loops. Unfolding of the attached native protein occurs by repeatedly pulling on the structure through successive power strokes, after which the unfolded polypeptide is translocated and degraded in the ClpP peptidase chamber. (b) A schematic of the ClpXP fingerprinting platform. ClpXP is labeled with a donor fluorophore and then immobilized on a passivated glass slide via biotin-streptavidin conjugation in a flow chamber. The acceptor-labeled substrate (10–20 nM) carries one internal cysteine labeled with Cy3 (red star), an N-terminal site labeled with Cy5 (green star), and a C-terminal ssrA-degradation tag. This substrate is then introduced into the chamber and allowed to localize to the immobilized ClpXP, where a combination of TIRF microscopy and ALEX imaging monitors FRET events as the substrate translocates through the pore. (c) FRET between an acceptor-labeled cysteine residue and the donor-labeled ClpP (left). Representation of a time trace from three-color ALEX (top). Initial binding of the acceptor-labeled substrate to ClpXP is shown by the simultaneous appearance of signals from Cy3 (middle) and Cy5 (bottom) upon excitation at 532 and 637 nm, respectively. As substrates carrying an ssrA tag move through the protease in a C-terminal to N-terminal direction, upon excitation at 437 nm the first FRET event is observed between Cy3 and Alexa488 (○) followed by FRET between Cy5 and Alexa488 (x). The order and distance of the recorded FRET events result in a fingerprint that can be compared to proteomics databases for identification (right).
Initial computational modeling by Yao et al. demonstrated that a FRET-based approach using a substrate carrying only two separate, distinguishable acceptor labels at their cysteine (C) and lysine (K) residues could bypass the issue of labeling and resolving 20 different amino acids. This concept strikes a balance between the daunting task of having to label and resolve each amino acid vs labeling so few residues as to render any resulting sequence information useless. Using their method, an acceptor-labeled substrate is localized to the ClpXP protease via an appended ssrA-degradation tag, where it is then translocated, or “scanned,” in a linear manner through the axial pore of a ClpXP. When acceptor-labeled residues are within the appropriate distance of the donor-labeled ClpP proteolytic chamber, a FRET readout, or “fingerprint,” is generated. This linear translocation of substrates allows FRET signals to appear in the same order as the labeled residues on the substrate (CK read) and can be compared to proteomics databases to determine identity. Cysteine and lysine were specifically chosen due to their frequency in canonical and isoform protein databases and because they can be labeled efficiently and orthogonally with minimal cross-labeling.64
Under ideal conditions with no experimental error, the probability of identifying the correct sequence (P) of a human protein was 90% when only the order of labeled residues was considered, and approached 100% when both order and distance were incorporated into the FRET readout analysis. After iteratively introducing a combination of random experimental errors, such as deletions, insertions, and transpositions, their method could identify 70%–80% of proteins at higher error rates (10%–30%) when both order and distance were considered.64
Subsequent work from the Meyer and Joo groups used peptides, mono-, and dimeric versions of titin to illustrate proof of concept for the ClpXP FRET fingerprinter with nanometer-scale accuracy. The ClpXP protease was modified to include a Cy3 or Alexa488 donor fluorophore in the ClpP proteolytic chamber located ∼12 nm away to prevent FRET between substrates carrying acceptor-fluorophores not yet bound to the ClpX complex. ClpXP was then immobilized onto a passivated quartz surface via biotin-streptavidin conjugation, followed by addition of labeled substrate into the microfluidic flow chamber [Fig. 4(b)]. An internal cysteine residue and the N-terminal site of the substrate were labeled with FRET acceptor dyes Cy3 and Cy5 using maleimide or NHS ester chemistry, respectively. Since the radius of the ClpP chamber is ∼5 nm, only substrates threaded through the chamber carrying fluorophore-labeled residues will be in close enough proximity (4–6 nm) to the donor fluorophore for FRET to occur [Fig. 4(c)]. FRET events between peptide substrates and ClpXP were monitored using a combination of TIRF microscopy and alternating laser excitation (ALEX) imaging to prevent early photobleaching and spectral overlap between donor and acceptor fluorophores.65
Labeled short polypeptides and small proteins (mono- or dimeric substrates of the titin I27 domain) were localized to the protease via a C-terminal ssrA-degradation tag and translocated the pore in a unidirectional, C-to-N-terminus direction and at constant speed, with an average processing time of 24 amino acids/s. The FRET scanner shows the potential to detect low-abundance substrates, demonstrating equal sensitivity for each donor–acceptor FRET pair when acceptor fluorophores were added at differing ratios.65 Finally, although the ClpX unfoldase is known to stall on substrates with complex secondary structures,66,67 the process of labeling cysteine residues appeared to sufficiently destabilize the small protein substrates used to allow for efficient processing, indicating that this technique may be able to process proteins regardless of their structural stability.65
Advantages and disadvantages to ClpXP FRET fingerprinting
Current protein sequencing techniques, such as mass spectroscopy, require large sample amounts for accurate readout and fail to recognize low-abundance species.37 In contrast, the accuracy and sensitivity exhibited by the ClpXP fingerprinter for low-abundance proteins in substrate mixtures show potential for next-generation protein sequencing platforms. This technique also shows promise in applications where protein separation is either not possible or not desirable, such as single-cell proteomics and health diagnostics involving small samples such as patient biopsies.
The choice of attaching a fluorophore to lysine, one of the most abundant amino acids, makes this technique applicable to almost all proteins. Even after computationally modeling indicated error rates of 30%, the fingerprinter accurately identified substrates containing ten or more C/K residues more than 50% of the time, with an increase in detection precision corresponding with substrate length. This amount of ten or more C/K residues is far below the average number of C/K residues per protein of ∼45 in humans, at which level the detection precision was predicted to approach 80%.64 Incorporating consideration of the distance between each FRET read (CK-distance reads) was also computationally shown to increase protein identification accuracy compared to CK order alone.64,65 Relative to other sequencing methods that rely on spectral analysis, CK-distance reads can be compared to currently available proteomics databases for identification.
Importantly, since this method relied on FRET between the donor and acceptor molecules, photobleaching of the donor molecule occurred after ∼40 s. As the substrate processing speed for donor-labeled ClpXP averaged 23.9 amino acids per second and was proportional to the length of protein substrates, a total of ∼950 amino acids could be processed prior to observed photobleaching. This is well above the median protein length of 350 amino acids in eukaryotes,64 making this method acceptable for fingerprinting most proteins. However, the substrate processing time to photobleaching could be further improved by increasing fluorophore photostability. While van Ginkel et al. used glucose oxidase and catalase in combination with Trolox as oxygen scavengers, the use of Ni2+ in place of Trolox could be beneficial as it has been shown to increase Cy3 photostability up to 6.5-fold at higher laser excitation powers (31 mW), out to almost 190 s without significantly reducing fluorescence intensity.68
One of the main limitations to this technique is that substrates require specific degradation tags located at the N- or C-terminus in order to localize to ClpXP. Substrates used by van Ginkel et al. all carried the ssrA degradation tag, a disordered 11-residue sequence added to nascent proteins stalled at the ribosome in Escherichia coli. Additionally, as with other fluorosequencing techniques, a specific subset of residues must be labeled. While the labeling shown by van Ginkel et al. sufficiently denatures their model protein analytes to allow efficient translocation, additional residue labeling is limited by the availability of known labeling chemistries. Relatedly, this method does not currently allow for the identification of minor sequence variation, especially among unlabeled residues, or PTMs.
RECOGNITION TUNNELING
Based on the principles of quantum tunneling, recognition tunneling is a variation of electron tunneling spectroscopy.69 In recognition tunneling, electrons are passed through a nanoscopic gap (∼2 nm) sandwiched between a conductive probe and substrate, both functionalized with strongly bonded recognition molecules. When a small voltage (<1 V) is applied across the junction, an analyte transiently interacts with the recognition molecules via relatively weak non-covalent bonds.70 Thermal vibrations caused by the analyte orienting itself within the junction create a unique distribution of current spikes that constitute a binding motif between the analyte and recognition molecules. Machine learning algorithms use signal features obtained from known analytes, such as temporal, spectral, and amplitudinal information, to generate an electronic fingerprint of each binding motif used to identify unknown molecules in the junction.70,71
To adapt recognition tunneling for SMPS, the Lindsay group used a scanning tunneling microscope (STM) equipped with a 2-nm junction between a palladium probe and palladium substrate. Both the probe and substrate were functionalized with 4(5)-substituted-1-H-imidazole-2-carboxamide (ICA) recognition molecules via thiol chemistry and partially insulated with polyethylene to achieve single-molecule resolution.72,73 After the application of a sub-1-V voltage across the junction, the resulting clusters of ion current spikes carrying analyte signal features were analyzed by a machine-learning algorithm, the support vector machine (SVM), for identification [Fig. 5(a)]. Initial experiments distinguished between the amino acid glycine and its methylated counterpart sarcosine, although accuracy was limited to 70% when only individual signal features, such as those associated with amplitude distributions or pulse shapes, were used. However, when these features were used in combination with each other, call accuracy increased to 95%. The use of multiple data parameters by the algorithm allowed the group to reproducibly distinguish between pairs of amino acids containing only minor structural differences, such as enantiomers and isobaric amino acids, to an accuracy of 80% or greater. Discrimination between short, simple peptides (3–4 amino acids) was also over 90% accurate, although the signal features of peptide residues differed relative to the same residue free in solution.73
FIG. 5.
Peptide sequencing using recognition tunneling and nanopores. (a) Recognition tunneling using two STM-coupled palladium electrodes functionalized with ICA recognition molecules. To create an ionic current, voltage is applied across the 2 nm junction. Analytes that pass through the gap and interact with the recognition molecules create thermal vibrations that generate a series of current spikes containing analyte-specific spectral features. Spectral data can then be used to identify the analyte via machine-learning algorithms trained on empirically determined signal features. (b) Synthetic nanopores consist of a small hole (<30 nm in diameter) in an insulating membrane separating two electrolyte-filled compartments. Voltage is applied between the two chambers to create an ionic current within the pore (left). In a mock current trace (right), a substrate approaches the pore (i) and traverses through it (ii), blocking ion flow and generating a measurable current blockade based on the size and shape of the substrate before leaving the pore (iii). (c) Biological nanopores operate in a similar manner to synthetic pores, but utilize a lipid bilayer membrane containing a trans-membrane protein containing a nanoscale-sized pore lumen (<10 nm) through which substrates may pass. (d) Pulling and unfolding of an ssrA-tagged substrate through an α-hemolysin pore used in conjunction with the molecular motor ClpX in the trans compartment.
In an effort to improve the signal-to-noise ratio, alternate setups used gold electrodes created from mechanically controllable break junctions. Using this method, the Kawai group obtained a gap width of a single amino acid (0.55–0.70 nm) capable of recognizing 12 out of 20 amino acids in total at the single-molecule level. Seven of these amino acids could be detected in mixed solution, in addition to the PTM phosphotyrosine.74 Work from the Taniguchi group furthered efforts to reduce ion-derived noise and increase speed by using gold electrodes created from mechanically controllable break junctions insulated with silicon dioxide.75,76
Advantages and disadvantages to recognition tunneling
Recognition tunneling techniques are both sensitive and highly accurate, capable of discriminating between individual amino acids, pairs of amino acids sharing similar chemical and physical properties, short peptides, and residues carrying PTMs. The instruments employed in recognition tunneling are smaller, less expensive, and less complex than other forms of peptide identification such as mass spectrometry. Additionally, compared to fluorescence-based techniques, recognition tunneling precludes the need for substrate labeling, allowing this method to distinguish PTMs lacking known labeling chemistry.73
As with other technologies employing current-trace readouts, the data obtained are complex and unintuitive, and require rigorous testing by machine algorithms to construct a unique identity for each molecule. Since there are many possible ways for the residue to position itself between the electrodes, current SVM algorithms can create reproducible signals for certain amino acids in solution but cannot identify the same residues individually when incorporated into a peptide. Instead of making calls based on the signal features of individual residues, accuracy is therefore dependent upon SVM algorithms learning to distinguish between short peptides or their constituent residues in context with the amino acids or modifications surrounding them.73 While this may present an initial hurdle, the complexity of current spike data may eventually advantage the SVM-based approach, which becomes more accurate as the number of signal features incorporated into analysis increases.
NANOPORE TECHNOLOGIES
The fundamental principle behind nanopore technologies is that modulations to an ionic current caused by an analyte traversing a nanoscale pore under an applied potential can be measured and analyzed with single-molecule sensitivity. In nanopore experiments, an ultra-thin, impermeable membrane containing a nanometer-sized hole is placed between two electrolyte-filled chambers. An electrical potential is then applied across the pore to create an electrical circuit. As an analyte translocates through the pore, the resulting current blockade causes fluctuations in the trans-pore ionic current thought mainly to be proportional to the volume of the analyte, although noncovalent interactions between pore and analyte also affect conductance.77 From this, data can be obtained regarding macromolecule size, charge, and conformation.78,79 Information regarding native protein state, DNA–protein interactions, individual amino acids, and PTMs are also possible to obtain using nanopore technologies.80,81
For peptide sequencing purposes, as long as the flow of a peptide analyte through the nanopore is unidirectional, it is theoretically possible to determine its sequence as each individual unit passes through the pore. Current research into nanopore-based single-molecule peptide sequencing focuses on two types of technologies: synthetic and biological nanopores. Synthetic or solid-state setups use nanofabrication to create small-diameter solid-state pores, through which peptides traverse after being treated with denaturing agents, such as sodium dodecyl sulfate (SDS) and β-mercaptoethanol (BME) [Fig. 5(b)].81,82 Alternatively, biological nanopores utilize a transmembrane pore protein in conjunction with a means of disrupting protein tertiary structure, such as an unfoldase enzyme or low concentrations of denaturing chemical agent [Fig. 5(c)].66,83–86
Initial nanopore projects utilized polymers and DNA/RNA substrates as opposed to proteins. In 1994, Bezrukov et al. demonstrated the traversal of polyethylene glycol through a molecular pore,87 and in 1996, Kasianowicz et al. showed that single-stranded DNA and RNA molecules could be translocated through the central ion channel of α-hemolysin. After applying an electrical field, the ion current fluctuations produced by biopolymer transport through the pore allowed them to determine its length.88 Decades later, modern-day biological nanopores are used in conjunction with a helicase or polymerase to control analyte translocation speed for label-free, single-molecule sequencing of linear nucleic acids.89 Currently available platforms, such as those from Oxford Nanopore,90 are fast and high-throughput, capable of achieving maximum read lengths in the hundreds of thousands. For example, the ProMethION platform released in 2017 utilizes 48 independently running flow cells each containing 3000 nanopores and can produce >130 GB data/flow cell for a total of 3–6 Tb data per run.17 For more in-depth analysis on the use of nanopores for DNA sequencing, and we refer to other reviews.91,92
While nanopore-based DNA/RNA sequencing technologies are impressive, their use in SMPS lags behind. Previous studies have used nanopore technologies to identify and analyze native proteins and their primary structures, while experiments studying peptide translocation probed the dynamics of model polypeptides and their interaction with the pore. Biophysical characteristics of short polypeptides, such as length, charge, stability, and conformational state, could be resolved as unique ionic fluctuations during translocation.93–96 Despite these achievements, the idiosyncratic characteristics of proteins present additional challenges for single-molecule nanopore sequencing, such as variable charge and complex tertiary structures. A nucleic acid biopolymer such as DNA adopts a strong and homogeneous negatively charged coil conformation, relying on electrophoresis to drive the biopolymer in a single-file, unidirectional manner through a nanopore following the application of voltage to create an ion current. However, since the net charge of a protein is variable and uneven owing both to the charge of its constituent residue side chains and the solution pH, the driving force through the pore is often a combination of electrophoresis, electroosmosis, and diffusion.97 Proteins are dynamic, folded polymers that exist in their native states or transition between several conformational states, and their complex tertiary structures must be unfolded and linearized before translocation through the nanopore is even possible. Biological pores, such as α-HL and aerolysin, have a small internal pore constriction (∼1.4 nm for α-HL98 and ∼1.0 nm for aerolysin99) that forbids the passage of folded proteins while allowing access to linearized ones. To optimize nanopores for use in sequencing and to ensure an accurate, replicable blockade current analysis, protein tertiary structures must be disrupted to permit sequential, linear translocation through these pores at a steady velocity, despite variable residue charges.
Protein denaturation
Both physical and chemical methods have been employed for protein denaturation, with the method of choice dependent on the nanopore setup employed. Any chemical denaturant used in conjunction with a biological nanopore must simultaneously be strong enough to disrupt analyte tertiary structure without denaturing the motor protein itself.100 This disruption can be accomplished by selecting for pore robustness and using less concentrated denaturants. For example, biological pores such as α-HL are resistant to tertiary disruption84,101 and are capable of withstanding denaturation at concentrations ranging from ∼4.5–7 M urea84 and 0.8–1.5 M guanidinium chloride102,103 before pore collapse occurs. By comparison, the mechanical robustness of solid-state nanopores made from silicon nitride (SiN) allows for peptide translocation at higher concentrations of denaturants such as 8 M urea,104 6 M guanidinium chloride,105 1% sodium dodecyl sulfate (SDS),81 or some combination thereof.82 In addition to chemical methods of denaturation, both thermal (50 °C and 70 °C)106,107 and electromotive (250–800 mV)106,108–110 denaturation techniques have been employed. However, while chemical denaturants such as those listed above prevent a protein from refolding once disrupted, temperature denaturation provides no such guarantee, and proteins may refold once the temperature is dropped. Additionally, thermal and electromotive denaturing methods also increase the rate of analyte translocation through the pore and therefore make separating resulting spectra more difficult.
Controlling peptide speed and translocation
Once a protein is denatured, its direction and speed as it traverses the nanopore must be optimized. Since nanopore technologies require an electric circuit between two chambers of electrolytes, the direction of electrophoretic transport is susceptible to both electroosmotic flow and protein charge.97 Therefore, the variable charge along a protein chain can alter the electric field of the pore lumen during substrate translocation and contribute to fluctuations in translocation velocity.86 To address this issue, a variety of methods has been employed to control peptide translocation direction through the nanopore. The use of SDS as a protein denaturant imparts a homogenous negative charge to the protein, ensuring unidirectional translocation through the pore.82 In 2017, Restrepo-Pérez and John et al. used both single-molecule experiments with solid-state SiN nanopores and molecular dynamics simulations to probe the interaction between SDS detergent and three model proteins. Substrate proteins were first denatured via boiling, after which the addition of SDS imparted a negative charge while preventing the substrate from refolding once temperatures were lowered. This newly imparted negative charge facilitated unidirectional protein translocation in the direction of the electrophoretic force.81
Other groups have modulated pore translocation velocity by appending negatively charged oligonucleotides to protein termini.85,111–113 Similarly, the Oukhaled group recently created a single-molecule trap by linking each substrate residue to arginine heptapeptide, a polycationic carrier that produces well-defined current blockades.114,115 Using an aerolysin nanopore, their method detected and differentiated all 20 amino acids from the polycationic carrier alone, while 15 out of 20 amino acids could be distinguished from each other. Although not capable of sequencing, their method was able to discriminate between amino acids with the same molecular mass as well as those with similar chemical structures.115
While substrate translocation can be directed by altering the charge of the protein, the Akeson group showed in 2013 that the speed and direction of a substrate can also be controlled using not only electrical force but also a biological pore in conjunction with the AAA+ unfoldase ClpX. ClpX was chosen since it generates sufficient force (∼20 pN) to denature secondary and tertiary structures and its translocation rate of up to 80 amino acids per second is appropriate for sequence analysis.61,116 Their setup consisted of an α-HL nanopore within a lipid bilayer that bisected the chamber into a cis compartment containing their substrate and a trans compartment containing ClpX and ATP [Fig. 5(d)]. To promote capture by the ClpX unfoldase in the trans chamber opposite the pore, their substrate contained a folded Smt3 domain (98 residues) coupled to a 65-residue-long polyanionic tail, immediately followed by a C-terminal ssrA degradation tag.
Following the application of a voltage bias, the electric field provided by the pore alone proved sufficient to capture the polyanionic tail, leaving the ssrA tag of the folded protein exposed on the trans side of the chamber. Here, ClpX present in solution interacted with the ssrA tag of the substrate, using energy derived from ATP hydrolysis to generate the mechanical force necessary to unfold the Smt3 domain that remained on the cis side of the chamber. The resulting ionic current fluctuations indicated that the combination of the pore electric field and the mechanical forces exerted by the ClpX unfoldase was capable of unfolding and translocating the remainder of the protein. By modifying the substrate with additional Smt3 domains, Nivala et al. were also able to demonstrate that sequence-dependent features of the substrate such as additional Smt3 domains and extended amino acid linkers could be identified as it was pulled through the α-HL pore.86
To demonstrate the sensitivity of this technique, follow-up experiments from this group were able to distinguish between different protein domains as well as substrate point mutations, unfolding intermediates, proteolytic cleavage events, and circular permutations with 86%–99% accuracy using machine learning tools trained on dwell time, mean current, or current root mean square (RMS) noise.66 Similar to SDS-assisted peptide translocation, the addition of an unfoldase both denatured the tertiary structure of protein substrates and maintained unidirectional translocation. While this method has the added benefit of slowing translocation enough to allow for sequencing, it also requires appending degradation tags to analytes for proper localization to the motor protein. This drawback is analogous to other nanopore-based methods that append negatively charged oligonucleotides to analytes to control polypeptide direction.
Several groups have also explored alternative pores capable of retaining peptides in their lumen long enough for accurate characterization and identification. Past work from the Maglia group used the pore-forming hemolytic protein cytolysin A (ClyA) due to the smaller trans opening of the pore (3 nm) relative to its cis opening (6 nm). This difference in size allowed analytes to remain trapped in the pore for extended periods of time, long enough to discriminate between isomeric, mono-, and poly-ubiquitinated proteins.117 Recent work using a nanopore electro-osmotic trap (NEOtrap) used electro-osmosis to allow a 35-nm-diameter DNA-origami sphere to trap label-free, single proteins for up to 11 h. This setup was capable of sub-millisecond discrimination between the size, shape, and nucleotide-dependent conformations of full-length proteins.118
Very recently, two groups demonstrated controlled translocation of a peptide chain through a nanopore while also obtaining sequence-dependent results based on current blockages. Work in the Huang laboratory combined peptide-oligonucleotide conjugates (POCs) with nanopore-induced phase-shift sequencing (NIPSS). The “phase shift” comes from the fixed distance of 14–15 nucleotides between their biological nanopore, MspA, and the highly processive motor protein, phi29 DNA polymerase (DNAP), which allows coupling of the electrochemical DNA reading from the pore with the enzymatic DNA ratcheting created by phi29 DNAP.119
In their setup, the oligonucleotide portion of the POC serves as the drive strand, initiating the enzyme-driven primer extension through the phi29 DNAP in a precise, stepwise manner against the electrophoretic force and preventing the POC from translocating the pore too quickly. Initial experiments used N-terminally conjugated POC strands and were capable of generating sequence-dependent current variances, mostly in the oligonucleotide segment of the POC, although single amino acid substitutions were also detectable. When C-terminally conjugated POC strands were used, the peptide section of the trace was only comparable to other events using C-terminally conjugated POC strands but showed clear discrimination between different peptide sequences. However, phi29 DNAP cannot translocate peptide segments, limiting NIPSS measurements to ∼14 nucleotides, corresponding to 5–6 amino acids and in principle necessitating enzymatic digestion of protein samples. The necessary protein sample preparation and purification methods are low-yield and costly, and POC synthesis is also restricted by the availability of chemical modifications for terminal residues.119
The Dekker laboratory used a similar DNA-peptide polymer methodology to determine site-specific information for a peptide, replacing phi29 DNAP with a DNA helicase, Hel308. Hel308 is stable in high salt concentrations and has enough pulling force (>50 pN) to likely denature protein secondary structures, Hel308 is also capable of translocating the polymer through the MspA nanopore in small, half-nucleotide steps of ∼0.33 nm, closer to that of the distance between adjacent peptide residues (∼0.36).120,121 In their design, the ssDNA-peptide polymer also includes an extender for helicase loading, as well as its complementary oligo which is capable of associating with the lipid bilayer via a 3′ cholesterol modification. This oligo blocks translocation by the helicase until the polymer is pulled into the nanopore, after which it is sheared off by MspA, allowing the helicase to translocate the DNA and subsequent peptide.120
Sequence traces were capable of detecting sequence differences consisting of single amino acid substitutions (D, W, or G) sandwiched between negatively charged amino acids (D and E), with hidden Markov modeling identifying the correct variant with ∼87% accuracy. Interestingly, re-reads of the same peptide were possible at high concentrations of the helicase (1 μM) due to successive helicase queuing, and re-reads showed a significant improvement in read accuracy. Even ∼30 rereads of an individual peptide decreased the error rate to <1 in 106. Although the read length is dependent upon the DNA-peptide linker (∼25 amino acids), this length is still larger than the peptide fragments used for analysis in mass spectrometry (<10 amino acids).120
Advantages and disadvantages to nanopore technologies
As with protein fluorosequencing using smFRET,64,65 nanopores could allow for sequencing of longer substrates since sample preparation does not require digestion by proteases. Additionally, biological nanopore technologies allow for the engineering of consistently reproducible protein pores to better improve resolution and detection capabilities. For example, using the biological pore actinoporin fragaceatoxin C, the Maglia group recently demonstrated that the introduction of aromatic residues inside the lumen improved peptide capture and recognition.122 Likewise, modifications of different chemical and physical properties of solid-state nanopores are possible. Several groups have explored a variety of 2D materials, including boron nitride,123,124 graphene,125 and molybdenum disulfide126–129 to decrease the thickness of the pore, thereby increasing spatial resolution and detection sensitivity.
Similar to tunneling currents, nanopores present a possibility for de novo sequencing without the use of chemical labeling. However, as with tunneling currents, these technologies first require establishing, either empirically or through computational modeling, a reproducible ionic current pattern for each analyte in order to generate a library of spectral data. While increasing resolution may help ameliorate this issue in the future, cataloguing each amino acid or PTM represents an immense undertaking that is made more complex if specific spectral patterns change based on surrounding residues, etc.
CONCLUDING REMARKS
SMPS is a complex and dynamic field. In the last decade, the explosion of SMPS technologies has moved us closer to the realm of personalized medicine and single-cell proteomics. When used in conjunction with the wealth of data obtained by genomics and transcriptomics, we can begin to develop a truly comprehensive understanding of cell-level biological processes and homeostasis. While no technique is yet capable of high-throughput, de novo protein sequencing at the single-molecule level, the five methods highlighted in this review are exciting developments toward an accurate and sensitive platform (Table I). A major hurdle to overcome is achieving the dynamic range required to detect low-abundance proteins that serve as important disease biomarkers. Biomarkers in complex samples, such as human blood, serve as early predictors of cancer or infectious diseases, or correlate with therapeutic outcomes, yet they often fall below the detection limit of current technologies such as MS.19,130,131 Additionally, current MS peptide-sequencing technologies are limited to short (<10 amino acids) peptide fragments and require expensive equipment and high operational costs relative to other technologies discussed here. While the nanopore mass spectrometer developed by the Stein lab has yet to test peptides or full-length proteins, a working model capable of linearizing polymers and sequentially feeding monomers to the mass analyzer would allow a mass spectrometer to achieve longer peptide reads with less sample, possibly reducing the need for running modern transcriptomic technologies alongside MS to achieve better accuracy and sensitivity. However, since this method currently relies on a single nanopore ion source, it lacks the throughput of alternate SMPS technologies capable of parallelization described here.
TABLE I.
Summary of single-molecule protein fingerprinting and sequencing methods.
| Method | Goal | Labeling | Demonstrated read length | Residues identified | Order determination | Identity determination | Refs. |
|---|---|---|---|---|---|---|---|
| Mass spectrometry | |||||||
| Nanopore mass spectrometer | De novo sequencing | No | Salts | N/A | Polymer linearization due to nozzle size | Mass-to-charge ratio | 42 |
| DNA bases | |||||||
| Fluorescence-based | |||||||
| Edman degradation | Identification | Yes | Peptides (<30 AAs) | AAs (C, K) | Cleavage of fluorescent N-terminal residue | Labeled residues | 52 |
| (Fluorosequencing) | (>93% efficient) | PTMs (phosphoserine) | (60% accurate)a | 34 | |||
| (98% accurate)b | |||||||
| ClpXP FRET scanner | Identification | Yes | Peptides (29–51 AAs) | AAs (C, K) | Estimated from substrate dwell time | Labeled residues | 64 |
| (Fingerprinting) | (95% efficient) | Native protein | (70%–80% accurate)a | 65 | |||
| Recognition tunneling | |||||||
| Tunneling currents | De novo sequencing | No | Free AAs | AAs (12) | N/A | Machine learning algorithms trained on multiple signal features | 73 |
| Peptides (4 AAs) | PTMs | 74 | |||||
| (99% accurate) | |||||||
| Nanopores | |||||||
| Helicase-/polymerase-coupled | De novo sequencing | No | Peptide-oligonucleotide conjugates | AAs | Residues estimated from current trace | Signal features resulting from current blockage | 120 |
| (87% accurate) | 119 | ||||||
| Unfoldase-coupled | De novo sequencing | No | Unfolded and native protein | Structural motifs | Structural motifs estimated from current trace | Signal features resulting from current blockage | 86 |
| AAs inferred from structural motifs | (86%–99% accurate) | 66 | |||||
| Solid-state SiN | De novo sequencing | No | Unfolded protein | Quadromers | Quadromers estimated from current trace | Signal features resulting from current blockage | 82 |
| (77% accurate) | |||||||
Note: AA = amino acids.
Computationally determined, two labeled peptides.
Computationally determined, four labeled peptides.
Fluorosequencing technologies such as Edman degradation and the ClpXP fingerprinter are promising in the immediate future as they utilize relatively inexpensive, well-established methodologies and platform fabrication methods capable of parallelization. These technologies have demonstrated the ability to fingerprint peptides and, in the case of the ClpXP FRET-scanner, individual proteins. Unlike tunneling current or nanopore methods reliant upon reference current trace readouts, fluorosequencing technologies do not require a corresponding trace library of individual residues, PTMs, or their spatial relation to each other on a linearized polypeptide. However, researchers are therefore limited in selecting residues or PTMs with known labeling chemistries, which will ultimately be a hindrance as nanopore technologies and their respective current trace libraries develop further. Additionally, there remain complicating issues regarding fluorescence-based imaging analysis, such as photobleaching, photoblinking, and dye-destruction, all of which contribute to experimental error rates and require further optimization. In the case of Edman degradation, recent work focused on decreasing error rates by improving purification and labeling shows promise for increasing yields and potentially allowing detection of low-abundance peptides.132
Both Edman degradation and the ClpXP FRET scanner technologies only determine partial peptide sequence based on the order and distance of labeled residues. As such, these methods rely on existing proteome databases to infer the missing sequences, which can introduce bias since proteins not yet available in these databases may be overlooked or mis-assigned. This problem is potentially exacerbated by the fact that only a small subset of residues is labeled in these techniques. Further, provided that the labeling chemistry is known, appending several different fluorescent labels to individual peptides may result in bulky molecules and subsequent steric interference. Ultimately, the inherent design of these technologies is inadequate for de novo protein sequencing, but their uses in a priori, long-read protein fingerprinting experiments could still be optimized toward less abundant proteins and disease biomarkers.
Unlike fluorescent-based fingerprinting methods that rely on proteomic databases, the potential exists for de novo sequencing using tunneling currents and nanopore technologies. Tunneling currents have demonstrated the ability to sequence short peptides (4 amino acids) and PTMs in conjunction with machine learning algorithms at high accuracy. Additional optimization using nano-fluidic sample injection via nanopores could reduce sample concentrations to the picomolar scale,73 allowing de novo sequencing of short peptides derived from protein digests. Similarly, recent biological nanopore technologies using MspA in conjunction with DNA–peptide conjugates can read between 14 and 25 amino acids and are sensitive to single amino acid changes.119,120 While this represents an improvement in read length over short peptides (<10 amino acids) used in mass spectrometry, protein fragmentation and high sequence coverage are still required to obtain sufficient read accuracy. In the case of Brinkerhoff et al., their setup obtained higher coverage by conducting multiple rereads of an individual peptide due to successive helicase bindings onto the pore, reducing stochastic error caused by enzyme backtracking or unresolved steps to <1 in 106.120
The issue of deconvoluting the ion current trace to determine the identity of any single amino acid in a peptide remains a challenge for nanopore technologies. Nanopore strategies that fall back on protein fingerprinting or other a priori sequence knowledge, rather than true de novo sequencing, have been suggested to compensate for this.120 Similar to sample preparation techniques used in mass spectrometry, pre-digestion of sample proteins could serve as a viable fingerprinting strategy, as demonstrated by recent work using a FraC pore to recognize protein substrates following digestion of Gallus-gallus lysozyme via spectral matching. This method would allow for low-cost, high-throughput protein analysis due to the relative ease with which electrical output signals can be incorporated into other devices.133 Related, work from the Meller laboratory using relatively low-cost TiO2 solid-state nanopores suggests that current trace deconvolution can be enhanced using multicolor fluorescence, something that is otherwise challenging in SiN pores due to the signal-to-background ratio caused by high photoluminescence of the membrane.134
Proteomics is a fast-evolving, multi-disciplinary field of research, of which SMPS technologies will no doubt play a major role. As explored in this review, the dynamic complexity of the proteome creates unique technical challenges not faced by similar next-generation DNA sequencing technologies. However, these challenges are daunting but not insurmountable, and research toward an SMPS platform capable of achieving the sensitivity and dynamic range required for protein analysis at the single-cell level is ongoing. Much in the same way that next-generation DNA sequencing revolutionized the field of genomics, a high-throughput SMPS platform would revolutionize proteomics for both basic research and medical standards.
ACKNOWLEDGMENTS
M.M.B. was supported by Grant No. 16SMPS01 from the Netherlands Foundation of Scientific Research Institute.
AUTHOR DECLARATIONS
Conflict of Interest
A. S. Meyer declares she has no conflict of interest. M. M. Brady declares he has no conflict of interest.
Author Contributions
A. S. Meyer and M. M. Brady conceptualized the data; M. M. Brady performed literature search and analysis; M. M. Brady wrote original draft preparation; A. S. Meyer and M. M. Brady reviewed and edited the document.
DATA AVAILABILITY
Data sharing is not applicable to this article as no new data were created or analyzed in this study.
References
- 1. Haider S. and Pal R., “ Integrated analysis of transcriptomic and proteomic data,” Curr. Genomics 14, 91–110 (2013). 10.2174/1389202911314020003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Chandramouli K. and Qian P. Y., “ Proteomics: Challenges, techniques and possibilities to overcome biological sample complexity,” Hum Genomics Proteomics 2009, 239204. 10.4061/2009/239204 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Vistain L. F. and Tay S., “ Single-cell proteomics,” Trends Biochem. Sci. 46, 661–672 (2021). 10.1016/j.tibs.2021.01.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Tritschler S., Theis F. J., Lickert H., and Bottcher A., “ Systematic single-cell analysis provides new insights into heterogeneity and plasticity of the pancreas,” Mol. Metab. 6, 974–990 (2017). 10.1016/j.molmet.2017.06.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Zheng H., Pomyen Y., Hernandez M. O., Li C., Livak F., Tang W., Dang H., Greten T. F., Davis J. L., Zhao Y., Mehta M., Levin Y., Shetty J., Tran B., Budhu A., and Wang X. W., “ Single-cell analysis reveals cancer stem cell heterogeneity in hepatocellular carcinoma,” Hepatology 68, 127–140 (2018). 10.1002/hep.29778 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Omenn G. S., “ The strategy, organization, and progress of the HUPO human proteome project,” J. Proteomics 100, 3–7 (2014). 10.1016/j.jprot.2013.10.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Domon B. and Aebersold R., “ Mass spectrometry and protein analysis,” Science 312, 212–217 (2006). 10.1126/science.1124619 [DOI] [PubMed] [Google Scholar]
- 8. Ponomarenko E. A., Poverennaya E. V., Ilgisonis E. V., Pyatnitskiy M. A., Kopylov A. T., Zgoda V. G., Lisitsa A. V., and Archakov A. I., “ The size of the human proteome: The width and depth,” Int. J. Anal. Chem. 2016, 7436849. 10.1155/2016/7436849 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Brett D., Pospisil H., Valcarcel J., Reich J., and Bork P., “ Alternative splicing and genome complexity,” Nat. Genet. 30, 29–30 (2002). 10.1038/ng803 [DOI] [PubMed] [Google Scholar]
- 10. Modrek B. and Lee C., “ A genomic view of alternative splicing,” Nat. Genet. 30, 13–19 (2002). 10.1038/ng0102-13 [DOI] [PubMed] [Google Scholar]
- 11. Kampa D., Cheng J., Kapranov P., Yamanaka M., Brubaker S., Cawley S., Drenkow J., Piccolboni A., Bekiranov S., Helt G., Tammana H., and Gingeras T. R., “ Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22,” Genome Res. 14, 331–342 (2004). 10.1101/gr.2094104 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Marth G., Yeh R., Minton M., Donaldson R., Li Q., Duan S., Davenport R., Miller R. D., and Kwok P. Y., “ Single-nucleotide polymorphisms in the public domain: How useful are they?,” Nat. Genet. 27, 371–372 (2001). 10.1038/86864 [DOI] [PubMed] [Google Scholar]
- 13. Mann M. and Jensen O. N., “ Proteomic analysis of post-translational modifications,” Nat. Biotechnol. 21, 255–261 (2003). 10.1038/nbt0303-255 [DOI] [PubMed] [Google Scholar]
- 14. Olsen J. V. and Mann M., “ Status of large-scale analysis of post-translational modifications by mass spectrometry,” Mol. Cell. Proteomics 12, 3444–3452 (2013). 10.1074/mcp.O113.034181 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Margulies M., Egholm M., Altman W. E., Attiya S., Bader J. S., Bemben L. A., Berka J., Braverman M. S., Chen Y. J., Chen Z., Dewell S. B., Du L., Fierro J. M., Gomes X. V., Godwin B. C., He W., Helgesen S., Ho C. H., Irzyk G. P., Jando S. C., Alenquer M. L., Jarvie T. P., Jirage K. B., Kim J. B., Knight J. R., Lanza J. R., Leamon J. H., Lefkowitz S. M., Lei M., Li J., Lohman K. L., Lu H., Makhijani V. B., McDade K. E., McKenna M. P., Myers E. W., Nickerson E., Nobile J. R., Plant R., Puc B. P., Ronan M. T., Roth G. T., Sarkis G. J., Simons J. F., Simpson J. W., Srinivasan M., Tartaro K. R., Tomasz A., Vogt K. A., Volkmer G. A., Wang S. H., Wang Y., Weiner M. P., Yu P., Begley R. F., and Rothberg J. M., “ Genome sequencing in microfabricated high-density picolitre reactors,” Nature 437, 376–380 (2005). 10.1038/nature03959 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Bentley D. R., Balasubramanian S., Swerdlow H. P., Smith G. P., Milton J., Brown C. G., Hall K. P., Evers D. J., Barnes C. L., Bignell H. R., Boutell J. M., Bryant J., Carter R. J., Keira Cheetham R., Cox A. J., Ellis D. J., Flatbush M. R., Gormley N. A., Humphray S. J., Irving L. J., Karbelashvili M. S., Kirk S. M., Li H., Liu X., Maisinger K. S., Murray L. J., Obradovic B., Ost T., Parkinson M. L., Pratt M. R., Rasolonjatovo I. M., Reed M. T., Rigatti R., Rodighiero C., Ross M. T., Sabot A., Sankar S. V., Scally A., Schroth G. P., Smith M. E., Smith V. P., Spiridou A., Torrance P. E., Tzonev S. S., Vermaas E. H., Walter K., Wu X., Zhang L., Alam M. D., Anastasi C., Aniebo I. C., Bailey D. M., Bancarz I. R., Banerjee S., Barbour S. G., Baybayan P. A., Benoit V. A., Benson K. F., Bevis C., Black P. J., Boodhun A., Brennan J. S., Bridgham J. A., Brown R. C., Brown A. A., Buermann D. H., Bundu A. A., Burrows J. C., Carter N. P., Castillo N., Chiara E. C. M., Chang S., Neil Cooley R., Crake N. R., Dada O. O., Diakoumakos K. D., Dominguez-Fernandez B., Earnshaw D. J., Egbujor U. C., Elmore D. W., Etchin S. S., Ewan M. R., Fedurco M., Fraser L. J., Fuentes Fajardo K. V., Scott Furey W., George D., Gietzen K. J., Goddard C. P., Golda G. S., Granieri P. A., Green D. E., Gustafson D. L., Hansen N. F., Harnish K., Haudenschild C. D., Heyer N. I., Hims M. M., Ho J. T., Horgan A. M., Hoschler K., Hurwitz S., Ivanov D. V., Johnson M. Q., James T., Huw Jones T. A., Kang G. D., Kerelska T. H., Kersey A. D., Khrebtukova I., Kindwall A. P., Kingsbury Z., Kokko-Gonzales P. I., Kumar A., Laurent M. A., Lawley C. T., Lee S. E., Lee X., Liao A. K., Loch J. A., Lok M., Luo S., Mammen R. M., Martin J. W., McCauley P. G., McNitt P., Mehta P., Moon K. W., Mullens J. W., Newington T., Ning Z., Ling Ng B., Novo S. M., O'Neill M. J., Osborne M. A., Osnowski A., Ostadan O., Paraschos L. L., Pickering L., Pike A. C., Pike A. C., Chris Pinkard D., Pliskin D. P., Podhasky J., Quijano V. J., Raczy C., Rae V. H., Rawlings S. R., Chiva Rodriguez A., Roe P. M., Rogers J., Rogert Bacigalupo M. C., Romanov N., Romieu A., Roth R. K., Rourke N. J., Ruediger S. T., Rusman E., Sanches-Kuiper R. M., Schenker M. R., Seoane J. M., Shaw R. J., Shiver M. K., Short S. W., Sizto N. L., Sluis J. P., Smith M. A., Ernest Sohna Sohna J., Spence E. J., Stevens K., Sutton N., Szajkowski L., Tregidgo C. L., Turcatti G., Vandevondele S., Verhovsky Y., Virk S. M., Wakelin S., Walcott G. C., Wang J., Worsley G. J., Yan J., Yau L., Zuerlein M., Rogers J., Mullikin J. C., Hurles M. E., McCooke N. J., West J. S., Oaks F. L., Lundberg P. L., Klenerman D., Durbin R., and Smith A. J., “ Accurate whole human genome sequencing using reversible terminator chemistry,” Nature 456, 53–59 (2008). 10.1038/nature07517 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Gupta N. and Verma V. K., “ Next-generation sequencing and its application: Empowering in public health beyond reality,” in Microbial Technology for the Welfare of Society, edited by Arora P., 1st ed. ( Springer, Singapore, 2019), pp. 313–341. [Google Scholar]
- 18. Liu Y., Beyer A., and Aebersold R., “ On the dependency of cellular protein levels on mRNA abundance,” Cell 165, 535–550 (2016). 10.1016/j.cell.2016.03.014 [DOI] [PubMed] [Google Scholar]
- 19. Aebersold R. and Mann M., “ Mass spectrometry-based proteomics,” Nature 422, 198–207 (2003). 10.1038/nature01511 [DOI] [PubMed] [Google Scholar]
- 20. Duncan M. W., Aebersold R., and Caprioli R. M., “ The pros and cons of peptide-centric proteomics,” Nat. Biotechnol. 28, 659–664 (2010). 10.1038/nbt0710-659 [DOI] [PubMed] [Google Scholar]
- 21. Loo J. A., Edmonds C. G., and Smith R. D., “ Primary sequence information from intact proteins by electrospray ionization tandem mass spectrometry,” Science 248, 201–204 (1990). 10.1126/science.2326633 [DOI] [PubMed] [Google Scholar]
- 22. Samaras P., Schmidt T., Frejno M., Gessulat S., Reinecke M., Jarzab A., Zecha J., Mergner J., Giansanti P., Ehrlich H. C., Aiche S., Rank J., Kienegger H., Krcmar H., Kuster B., and Wilhelm M., “ ProteomicsDB: A multi-omics and multi-organism resource for life science research,” Nucl. Acids Res. 48(D1), D1153–D1163 (2020). 10.1093/nar/gkz974 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Steen H. and Mann M., “ The ABC's (and XYZ's) of peptide sequencing,” Nat. Rev. Mol. Cell. Biol. 5, 699–711 (2004). 10.1038/nrm1468 [DOI] [PubMed] [Google Scholar]
- 24. Cox J., Hubner N. C., and Mann M., “ How much peptide sequence information is contained in ion trap tandem mass spectra?,” J. Am. Soc. Mass Spectrom. 19, 1813–1820 (2008). 10.1016/j.jasms.2008.07.024 [DOI] [PubMed] [Google Scholar]
- 25. Medzihradszky K. F. and Chalkley R. J., “ Lessons in de novo peptide sequencing by tandem mass spectrometry,” Mass Spectrom. Rev. 34, 43–63 (2015). 10.1002/mas.21406 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Lesur A. and Domon B., “ Advances in high-resolution accurate mass spectrometry application to targeted proteomics,” Proteomics 15, 880–890 (2015). 10.1002/pmic.201400450 [DOI] [PubMed] [Google Scholar]
- 27. Muth T. and Renard B. Y., “ Evaluating de novo sequencing in proteomics: Already an accurate alternative to database-driven peptide identification?,” Briefings Bioinf. 19, 954–970 (2018). 10.1093/bib/bbx033 [DOI] [PubMed] [Google Scholar]
- 28. Sinitcyn P., Rudolph J. D., and Cox J., “ Computational methods for understanding mass spectrometry-based shotgun proteomics data,” Annu. Rev. Biomed. Data Sci. 1, 207–234 (2018). 10.1146/annurev-biodatasci-080917-013516 [DOI] [Google Scholar]
- 29. Kelly R. T., “ Single-cell proteomics: Progress and prospects,” Mol. Cell. Proteomics 19, 1739–1748 (2020). 10.1074/mcp.R120.002234 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Richards A. L., Merrill A. E., and Coon J. J., “ Proteome sequencing goes deep,” Curr. Opin. Chem. Biol. 24, 11–17 (2015). 10.1016/j.cbpa.2014.10.017 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Ghaemmaghami S., Huh W. K., Bower K., Howson R. W., Belle A., Dephoure N., O'Shea E. K., and Weissman J. S., “ Global analysis of protein expression in yeast,” Nature 425, 737–741 (2003). 10.1038/nature02046 [DOI] [PubMed] [Google Scholar]
- 32. Huh W. K., Falvo J. V., Gerke L. C., Carroll A. S., Howson R. W., Weissman J. S., and O'Shea E. K., “ Global analysis of protein localization in budding yeast,” Nature 425, 686–691 (2003). 10.1038/nature02026 [DOI] [PubMed] [Google Scholar]
- 33. Schwanhausser B., Busse D., Li N., Dittmar G., Schuchhardt J., Wolf J., Chen W., and Selbach M., “ Global quantification of mammalian gene expression control,” Nature 473, 337–342 (2011). 10.1038/nature10098 [DOI] [PubMed] [Google Scholar]
- 34. Swaminathan J., Boulgakov A. A., Hernandez E. T., Bardo A. M., Bachman J. L., Marotta J., Johnson A. M., Anslyn E. V., and Marcotte E. M., “ Highly parallel single-molecule identification of proteins in zeptomole-scale mixtures,” Nat. Biotechnol. 36, 1076–1082 (2018). 10.1038/nbt.4278 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Thakur S. S., Geiger T., Chatterjee B., Bandilla P., Frohlich F., Cox J., and Mann M., “ Deep and highly sensitive proteome coverage by LC-MS/MS without prefractionation,” Mol. Cell. Proteomics 10, M110.003699 (2011). 10.1074/mcp.M110.003699 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Huang B., Wu H., Bhaya D., Grossman A., Granier S., Kobilka B. K., and Zare R. N., “ Counting low-copy number proteins in a single cell,” Science 315, 81–84 (2007). 10.1126/science.1133992 [DOI] [PubMed] [Google Scholar]
- 37. Domon B. and Aebersold R., “ Options and considerations when selecting a quantitative proteomics strategy,” Nat. Biotechnol. 28, 710–721 (2010). 10.1038/nbt.1661 [DOI] [PubMed] [Google Scholar]
- 38. Bilbao A., Varesio E., Luban J., Strambio-De-Castillia C., Hopfgartner G., Muller M., and Lisacek F., “ Processing strategies and software solutions for data-independent acquisition in mass spectrometry,” Proteomics 15, 964–980 (2015). 10.1002/pmic.201400323 [DOI] [PubMed] [Google Scholar]
- 39. Macklin A., Khan S., and Kislinger T., “ Recent advances in mass spectrometry based clinical proteomics: Applications to cancer research,” Clin. Proteomics 17, 17 (2020). 10.1186/s12014-020-09283-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Griss J., Perez-Riverol Y., Lewis S., Tabb D. L., Dianes J. A., Del-Toro N., Rurik M., Walzer M. W., Kohlbacher O., Hermjakob H., Wang R., and Vizcaino J. A., “ Recognizing millions of consistently unidentified spectra across hundreds of shotgun proteomics datasets,” Nat. Methods 13, 651–656 (2016). 10.1038/nmeth.3902 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Banerjee S. and Mazumdar S., “ Electrospray ionization mass spectrometry: A technique to access the information beyond the molecular weight of the analyte,” Int. J. Anal. Chem. 2012, 282574. 10.1155/2012/282574 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Bush J., Maulbetsch W., Lepoitevin M., Wiener B., Mihovilovic Skanata M., Moon W., Pruitt C., and Stein D., “ The nanopore mass spectrometer,” Rev. Sci. Instrum. 88, 113307 (2017). 10.1063/1.4986043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Maulbetsch W., Wiener B., Poole W., Bush J., and Stein D., “ Preserving the sequence of a biopolymer's monomers as they enter an electrospray mass spectrometer,” Phys. Rev. Appl. 6, 054006 (2016). 10.1103/PhysRevApplied.6.054006 [DOI] [Google Scholar]
- 44. Wilm M. and Mann M., “ Analytical properties of the nanoelectrospray ion source,” Anal. Chem. 68, 1–8 (1996). 10.1021/ac9509519 [DOI] [PubMed] [Google Scholar]
- 45. Chiu Y.-H., Austin B. L., Dressler R. A., Levandier D., Murray P. T., Lozano P., and Martinez-Sanchez M., “ Mass spectrometric analysis of colloid thruster ion emission from selected propellants,” J. Propul. Power 21, 416–423 (2012). 10.2514/1.9690 [DOI] [Google Scholar]
- 46. Luedtke W. D., Landman U., Chiu Y. H., Levandier D. J., Dressler R. A., Sok S., and Gordon M. S., “ Nanojets, electrospray, and ion field evaporation: Molecular dynamics simulations and laboratory experiments,” J. Phys. Chem. A 112, 9628–9649 (2008). 10.1021/jp804585y [DOI] [PubMed] [Google Scholar]
- 47. Cook K. D., “ Electrohydrodynamic mass spectrometry,” Mass Spectrom. Rev. 5, 467–519 (1986). 10.1002/mas.1280050404 [DOI] [Google Scholar]
- 48. Zolotoi N. B., Karpov G. V., and Tal'roze V. L., “ Field-evaporation mass spectrography for ions from water and aqueous solutions: Aqueous solutions of sodium iodide and sucrose,” Zh. Anal. Khim. 35, 1781–1791 (1980). [Google Scholar]
- 49. Zolotoi N. B., Karpov G. V., and Skurat V. E., “ Field-evaporation mass spectroscopy applied to crown ether complexes with alkali and alkaline-earth metals,” Theor. Exp. Chem. 24, 235–238 (1988). 10.1007/BF00531205 [DOI] [Google Scholar]
- 50. Edman P., “ A method for the determination of amino acid sequence in peptides,” Arch. Biochem. 22, 475 (1949). [PubMed] [Google Scholar]
- 51. Han K. K., Belaiche D., Moreau O., and Briand G., “ Current developments in stepwise edman degradation of peptides and proteins,” Int. J. Biochem. 17, 429–445 (1985). 10.1016/0020-711X(85)90138-7 [DOI] [Google Scholar]
- 52. Swaminathan J., Boulgakov A. A., and Marcotte E. M., “ A theoretical justification for single molecule peptide sequencing,” PLoS Comput. Biol. 11, e1004080 (2015). 10.1371/journal.pcbi.1004080 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. Collins B. C. and Aebersold R., “ Proteomics goes parallel,” Nat. Biotechnol. 36, 1051–1053 (2018). 10.1038/nbt.4288 [DOI] [PubMed] [Google Scholar]
- 54. Hernandez E. T., Swaminathan J., Marcotte E. M., and Anslyn E. V., “ Solution-phase and solid-phase sequential, selective modification of side chains in KDYWEC and KDYWE as models for usage in single-molecule protein sequencing,” New J. Chem. 41, 462–469 (2017). 10.1039/C6NJ02932A [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55. Förster T., “ Energiewanderung und fluoreszenz,” Naturwissenschaften 33, 166–175 (1946). 10.1007/BF00585226 [DOI] [Google Scholar]
- 56. Sekar R. B. and Periasamy A., “ Fluorescence resonance energy transfer (FRET) microscopy imaging of live cell protein localizations,” J. Cell. Biol. 160, 629–633 (2003). 10.1083/jcb.200210140 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Grimaud R., Kessel M., Beuron F., Steven A. C., and Maurizi M. R., “ Enzymatic and structural similarities between the Escherichia coli ATP-dependent proteases, ClpXP and ClpAP,” J. Biol. Chem. 273, 12476–12481 (1998). 10.1074/jbc.273.20.12476 [DOI] [PubMed] [Google Scholar]
- 58. Baker T. A. and Sauer R. T., “ ClpXP, an ATP-powered unfolding and protein-degradation machine,” Biochim. Biophys. Acta 1823, 15–28 (2012). 10.1016/j.bbamcr.2011.06.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Olivares A. O., Baker T. A., and Sauer R. T., “ Mechanical protein unfolding and degradation,” Annu. Rev. Physiol. 80, 413–429 (2018). 10.1146/annurev-physiol-021317-121303 [DOI] [PubMed] [Google Scholar]
- 60. Gottesman S., Roche E., Zhou Y., and Sauer R. T., “ The ClpXP and ClpAP proteases degrade proteins with carboxy-terminal peptide tails added by the ssrA-tagging system,” Genes Dev. 12, 1338–1347 (1998). 10.1101/gad.12.9.1338 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Aubin-Tam M. E., Olivares A. O., Sauer R. T., Baker T. A., and Lang M. J., “ Single-molecule protein unfolding and translocation by an ATP-fueled proteolytic machine,” Cell 145, 257–267 (2011). 10.1016/j.cell.2011.03.036 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Martin A., Baker T. A., and Sauer R. T., “ Diverse pore loops of the AAA + ClpX machine mediate unassisted and adaptor-dependent recognition of ssrA-tagged substrates,” Mol. Cell 29, 441–450 (2008). 10.1016/j.molcel.2008.02.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Martin A., Baker T. A., and Sauer R. T., “ Pore loops of the AAA + ClpX machine grip substrates to drive translocation and unfolding,” Nat. Struct. Mol. Biol. 15, 1147–1151 (2008). 10.1038/nsmb.1503 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64. Yao Y., Docter M., van Ginkel J., de Ridder D., and Joo C., “ Single-molecule protein sequencing through fingerprinting: Computational assessment,” Phys. Biol. 12, 055003 (2015). 10.1088/1478-3975/12/5/055003 [DOI] [PubMed] [Google Scholar]
- 65. van Ginkel J., Filius M., Szczepaniak M., Tulinski P., Meyer A. S., and Joo C., “ Single-molecule peptide fingerprinting,” Proc. Natl. Acad. Sci. U.S.A. 115, 3338–3343 (2018). 10.1073/pnas.1707207115 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Nivala J., Mulroney L., Li G., Schreiber J., and Akeson M., “ Discrimination among protein variants using an unfoldase-coupled nanopore,” ACS Nano 8, 12365–12375 (2014). 10.1021/nn5049987 [DOI] [PubMed] [Google Scholar]
- 67. Cordova J. C., Olivares A. O., Shin Y., Stinson B. M., Calmat S., Schmitz K. R., Aubin-Tam M. E., Baker T. A., Lang M. J., and Sauer R. T., “ Stochastic but highly coordinated protein unfolding and translocation by the ClpXP proteolytic machine,” Cell 158, 647–658 (2014). 10.1016/j.cell.2014.05.043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Glembockyte V., Lincoln R., and Cosa G., “ Cy3 photoprotection mediated by Ni2+ for extended single-molecule imaging: Old tricks for new techniques,” J. Am. Chem. Soc. 137, 1116–1122 (2015). 10.1021/ja509923e [DOI] [PubMed] [Google Scholar]
- 69. Lindsay S., “ The promises and challenges of solid-state sequencing,” Nat. Nanotechnol. 11, 109–111 (2016). 10.1038/nnano.2016.9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Lindsay S., He J., Sankey O., Hapala P., Jelinek P., Zhang P., Chang S., and Huang S., “ Recognition tunneling,” Nanotechnology 21, 262001 (2010). 10.1088/0957-4484/21/26/262001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Chang S., Huang S., He J., Liang F., Zhang P., Li S., Chen X., Sankey O., and Lindsay S., “ Electronic signatures of all four DNA nucleosides in a tunneling gap,” Nano Lett. 10, 1070–1075 (2010). 10.1021/nl1001185 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72. Liang F., Li S., Lindsay S., and Zhang P., “ Synthesis, physicochemical properties, and hydrogen bonding of 4(5)-substituted 1-H-imidazole-2-carboxamide, a potential universal reader for DNA sequencing by recognition tunneling,” Chemistry 18, 5998–6007 (2012). 10.1002/chem.201103306 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Zhao Y., Ashcroft B., Zhang P., Liu H., Sen S., Song W., Im J., Gyarfas B., Manna S., Biswas S., Borges C., and Lindsay S., “ Single-molecule spectroscopy of amino acids and peptides by recognition tunnelling,” Nat. Nanotechnol. 9, 466–473 (2014). 10.1038/nnano.2014.54 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Ohshiro T., Tsutsui M., Yokota K., Furuhashi M., Taniguchi M., and Kawai T., “ Detection of post-translational modifications in single peptides using electron tunnelling currents,” Nat. Nanotechnol. 9, 835–840 (2014). 10.1038/nnano.2014.193 [DOI] [PubMed] [Google Scholar]
- 75. Morikawa T., Yokota K., Tanimoto S., Tsutsui M., and Taniguchi M., “ Detecting single-nucleotides by tunneling current measurements at sub-MHz temporal resolution,” Sensors (Basel) 17, 885 (2017). 10.3390/s17040885 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Morikawa T., Yokota K., Tsutsui M., and Taniguchi M., “ Fast and low-noise tunnelling current measurements for single-molecule detection in an electrolyte solution using insulator-protected nanoelectrodes,” Nanoscale 9, 4076–4081 (2017). 10.1039/C6NR09278K [DOI] [PubMed] [Google Scholar]
- 77. Li M. Y., Ying Y. L., Yu J., Liu S. C., Wang Y. Q., Li S., and Long Y. T., “ Revisiting the origin of nanopore current blockage for volume difference sensing at the atomic level,” JACS Au 1, 967–976 (2021). 10.1021/jacsau.1c00109 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78. Bell N. A. and Keyser U. F., “ Digitally encoded DNA nanostructures for multiplexed, single-molecule protein sensing with nanopores,” Nat. Nanotechnol. 11, 645–651 (2016). 10.1038/nnano.2016.50 [DOI] [PubMed] [Google Scholar]
- 79. Waduge P., Hu R., Bandarkar P., Yamazaki H., Cressiot B., Zhao Q., Whitford P. C., and Wanunu M., “ Nanopore-based measurements of protein size, fluctuations, and conformational changes,” ACS Nano 11, 5706–5716 (2017). 10.1021/acsnano.7b01212 [DOI] [PubMed] [Google Scholar]
- 80. Carson S. and Wanunu M., “ Challenges in DNA motion control and sequence readout using nanopore devices,” Nanotechnology 26, 074004 (2015). 10.1088/0957-4484/26/7/074004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Restrepo-Perez L., John S., Aksimentiev A., Joo C., and Dekker C., “ SDS-assisted protein transport through solid-state nanopores,” Nanoscale 9, 11685–11693 (2017). 10.1039/C7NR02450A [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82. Kennedy E., Dong Z., Tennant C., and Timp G., “ Reading the primary structure of a protein with 0.07 nm(3) resolution using a subnanometre-diameter pore,” Nat. Nanotechnol. 11, 968–976 (2016). 10.1038/nnano.2016.120 [DOI] [PubMed] [Google Scholar]
- 83. Christensen C., Baran C., Krasniqi B., Stefureac R. I., Nokhrin S., and Lee J. S., “ Effect of charge, topology and orientation of the electric field on the interaction of peptides with the alpha-hemolysin pore,” J. Pept. Sci. 17, 726–734 (2011). 10.1002/psc.1393 [DOI] [PubMed] [Google Scholar]
- 84. Pastoriza-Gallego M., Oukhaled G., Mathe J., Thiebot B., Betton J. M., Auvray L., and Pelta J., “ Urea denaturation of alpha-hemolysin pore inserted in planar lipid bilayer detected by single nanopore recording: Loss of structural asymmetry,” FEBS Lett. 581, 3371–3376 (2007). 10.1016/j.febslet.2007.06.036 [DOI] [PubMed] [Google Scholar]
- 85. Pastoriza-Gallego M., Rabah L., Gibrat G., Thiebot B., van der Goot F. G., Auvray L., Betton J. M., and Pelta J., “ Dynamics of unfolded protein transport through an aerolysin pore,” J. Am. Chem. Soc. 133, 2923–2931 (2011). 10.1021/ja1073245 [DOI] [PubMed] [Google Scholar]
- 86. Nivala J., Marks D. B., and Akeson M., “ Unfoldase-mediated protein translocation through an alpha-hemolysin nanopore,” Nat. Biotechnol. 31, 247–250 (2013). 10.1038/nbt.2503 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87. Bezrukov S. M., Vodyanoy I., and Parsegian V. A., “ Counting polymers moving through a single ion channel,” Nature 370, 279–281 (1994). 10.1038/370279a0 [DOI] [PubMed] [Google Scholar]
- 88. Kasianowicz J. J., Brandin E., Branton D., and Deamer D. W., “ Characterization of individual polynucleotide molecules using a membrane channel,” Proc. Natl. Acad. Sci. U.S.A. 93, 13770–13773 (1996). 10.1073/pnas.93.24.13770 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89. Caldwell C. C. and Spies M., “ Helicase SPRNTing through the nanopore,” Proc. Natl. Acad. Sci. U.S.A. 114, 11809–11811 (2017). 10.1073/pnas.1716866114 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90. Lu H., Giordano F., and Ning Z., “ Oxford nanopore MinION sequencing and genome assembly,” Genomics Proteomics Bioinf. 14, 265–279 (2016). 10.1016/j.gpb.2016.05.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91. Mardis E. R., “ DNA sequencing technologies: 2006–2016,” Nat. Protoc. 12, 213–218 (2017). 10.1038/nprot.2016.182 [DOI] [PubMed] [Google Scholar]
- 92. Kraft F. and Kurth I., “ Long-read sequencing to understand genome biology and cell function,” Int. J. Biochem. Cell Biol. 126, 105799 (2020). 10.1016/j.biocel.2020.105799 [DOI] [PubMed] [Google Scholar]
- 93. Sutherland T. C., Long Y. T., Stefureac R., Bediako-Amoa I., Heinz-Bernhard K., and Lee J. S., “ Structure of peptides investigated by nanopore analysis,” Nano Lett. 4, 1273–1277 (2004). 10.1021/nl049413e [DOI] [Google Scholar]
- 94. Stefureac R., Long Y. T., Kraatz H. B., Howard P., and Lee J. S., “ Transport of alpha-helical peptides through alpha-hemolysin and aerolysin pores,” Biochemistry 45, 9172–9179 (2006). 10.1021/bi0604835 [DOI] [PubMed] [Google Scholar]
- 95. Goodrich C. P., Kirmizialtin S., Huyghues-Despointes B. M., Zhu A., Scholtz J. M., Makarov D. E., and Movileanu L., “ Single-molecule electrophoresis of beta-hairpin peptides by electrical recordings and Langevin dynamics simulations,” J. Phys. Chem. B 111, 3332–3335 (2007). 10.1021/jp071364h [DOI] [PubMed] [Google Scholar]
- 96. Movileanu L., “ Interrogating single proteins through nanopores: Challenges and opportunities,” Trends Biotechnol. 27, 333–341 (2009). 10.1016/j.tibtech.2009.02.008 [DOI] [PubMed] [Google Scholar]
- 97. Firnkes M., Pedone D., Knezevic J., Doblinger M., and Rant U., “ Electrically facilitated translocations of proteins through silicon nitride nanopores: Conjoint and competitive action of diffusion, electrophoresis, and electroosmosis,” Nano Lett. 10, 2162–2167 (2010). 10.1021/nl100861c [DOI] [PubMed] [Google Scholar]
- 98. Song L., Hobaugh M., Shustak C., Cheley S., Bayley H., and Gouaux J., “ Structure of staphylococcal alpha-hemolysin, a heptameric transmembrane pore,” Science 274, 1859–1866 (1996). 10.1126/science.274.5294.1859 [DOI] [PubMed] [Google Scholar]
- 99. Iacovache I., De Carlo S., Cirauqui N., Dal Peraro M., van der Goot F., and Zuber B., “ Cryo-EM structure of aerolysin variants reveals a novel protein fold and the pore-formation process,” Nat. Commun. 7, 12062 (2016). 10.1038/ncomms12062 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100. Kurz V., Nelson E. M., Shim J., and Timp G., “ Direct visualization of single-molecule translocations through synthetic nanopores comparable in size to a molecule,” ACS Nano 7, 4057–4069 (2013). 10.1021/nn400182s [DOI] [PubMed] [Google Scholar]
- 101. Bortoleto R. K. and Ward R. J., “ A stability transition at mildly acidic pH in the alpha-hemolysin (alpha-toxin) from staphylococcus aureus,” FEBS Lett. 459, 438–442 (1999). 10.1016/S0014-5793(99)01246-6 [DOI] [PubMed] [Google Scholar]
- 102. Oukhaled G., Mathe J., Biance A. L., Bacri L., Betton J. M., Lairez D., Pelta J., and Auvray L., “ Unfolding of proteins and long transient conformations detected by single nanopore recording,” Phys. Rev. Lett. 98, 158101 (2007). 10.1103/PhysRevLett.98.158101 [DOI] [PubMed] [Google Scholar]
- 103. Merstorf C., Cressiot B., Pastoriza-Gallego M., Oukhaled A., Betton J. M., Auvray L., and Pelta J., “ Wild type, mutant protein unfolding and phase transition detected by single-nanopore recording,” ACS Chem. Biol. 7, 652–658 (2012). 10.1021/cb2004737 [DOI] [PubMed] [Google Scholar]
- 104. Talaga D. S. and Li J., “ Single-molecule protein unfolding in solid state nanopores,” J. Am. Chem. Soc. 131, 9287–9297 (2009). 10.1021/ja901088b [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105. Li J., Fologea D., Rollings R., and Ledden B., “ Characterization of protein unfolding with solid-state nanopores,” Protein Pept. Lett. 21, 256–265 (2014). 10.2174/09298665113209990077 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106. Freedman K. J., Jurgens M., Prabhu A., Ahn C. W., Jemth P., Edel J. B., and Kim M. J., “ Chemical, thermal, and electric field induced unfolding of single protein molecules studied using nanopores,” Anal. Chem. 83, 5137–5144 (2011). 10.1021/ac2001725 [DOI] [PubMed] [Google Scholar]
- 107. Payet L., Martinho M., Pastoriza-Gallego M., Betton J. M., Auvray L., Pelta J., and Mathe J., “ Thermal unfolding of proteins probed at the single molecule level using nanopores,” Anal. Chem. 84, 4071–4076 (2012). 10.1021/ac300129e [DOI] [PubMed] [Google Scholar]
- 108. Oukhaled A., Cressiot B., Bacri L., Pastoriza-Gallego M., Betton J. M., Bourhis E., Jede R., Gierak J., Auvray L., and Pelta J., “ Dynamics of completely unfolded and native proteins through solid-state nanopores as a function of electric driving force,” ACS Nano 5, 3628–3638 (2011). 10.1021/nn1034795 [DOI] [PubMed] [Google Scholar]
- 109. Cressiot B., Oukhaled A., Patriarche G., Pastoriza-Gallego M., Betton J. M., Auvray L., Muthukumar M., Bacri L., and Pelta J., “ Protein transport through a narrow solid-state nanopore at high voltage: Experiments and theory,” ACS Nano 6, 6236–6243 (2012). 10.1021/nn301672g [DOI] [PubMed] [Google Scholar]
- 110. Freedman K. J., Haq S. R., Edel J. B., Jemth P., and Kim M. J., “ Single molecule unfolding and stretching of protein domains inside a solid-state nanopore by electric field,” Sci. Rep. 3, 1638 (2013). 10.1038/srep01638 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111. Rodriguez-Larrea D. and Bayley H., “ Multistep protein unfolding during nanopore translocation,” Nat. Nanotechnol. 8, 288–295 (2013). 10.1038/nnano.2013.22 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112. Rosen C. B., Rodriguez-Larrea D., and Bayley H., “ Single-molecule site-specific detection of protein phosphorylation with a nanopore,” Nat. Biotechnol. 32, 179–181 (2014). 10.1038/nbt.2799 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113. Biswas S., Song W., Borges C., Lindsay S., and Zhang P., “ Click addition of a DNA thread to the N-termini of peptides for their translocation through solid-state nanopores,” ACS Nano 9, 9652–9664 (2015). 10.1021/acsnano.5b04984 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114. Piguet F., Ouldali H., Pastoriza-Gallego M., Manivet P., Pelta J., and Oukhaled A., “ Identification of single amino acid differences in uniformly charged homopolymeric peptides with aerolysin nanopore,” Nat. Commun. 9, 966 (2018). 10.1038/s41467-018-03418-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115. Ouldali H., Sarthak K., Ensslen T., Piguet F., Manivet P., Pelta J., Behrends J. C., Aksimentiev A., and Oukhaled A., “ Electrical recognition of the twenty proteinogenic amino acids using an aerolysin nanopore,” Nat. Biotechnol. 38, 176–181 (2020). 10.1038/s41587-019-0345-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116. Maillard R. A., Chistol G., Sen M., Righini M., Tan J., Kaiser C. M., Hodges C., Martin A., and Bustamante C., “ ClpX(P) generates mechanical force to unfold and translocate its protein substrates,” Cell 145, 459–469 (2011). 10.1016/j.cell.2011.04.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117. Wloka C., Van Meervelt V., van Gelder D., Danda N., Jager N., Williams C. P., and Maglia G., “ Label-free and real-time detection of protein ubiquitination with a biological nanopore,” ACS Nano 11, 4387–4394 (2017). 10.1021/acsnano.6b07760 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118. Schmid S., Stömmer P., Dietz H., and Dekker C., “ Nanopore electro-osmotic trap for the label-free study of single proteins and their conformations,” Nat. Nanotechnol. 16, 1244–1250 (2021). 10.1038/s41565-021-00958-5 [DOI] [PubMed] [Google Scholar]
- 119. Yan S., Zhang J., Wang Y., Guo W., Zhang S., Liu Y., Cao J., Wang Y., Wang L., Ma F., Zhang P., Chen H. Y., and Huang S., “ Single molecule ratcheting motion of peptides in a mycobacterium smegmatis porin A (MspA) nanopore,” Nano Lett. 21, 6703–3710 (2021). 10.1021/acs.nanolett.1c02371 [DOI] [PubMed] [Google Scholar]
- 120. Brinkerhoff H., Kang A., Liu J., Aksimentiev A., and Dekker C., “ Multiple rereads of single proteins at single-amino acid resolution using nanopores,” Science 374, eabl4381 (2021). 10.1126/science.abl4381 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121. Eisenberg D., “ The discovery of the alpha-helix and beta-sheet, the principal structural features of proteins,” Proc. Natl. Acad. Sci. U.S.A. 100, 11207–11210 (2003). 10.1073/pnas.2034522100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122. Lucas F. L. R., Sarthak K., Lenting E. M., Coltan D., van der Heide N. J., Versloot R. C. A., Aksimentiev A., and Maglia G., “ The manipulation of the internal hydrophobicity of FraC nanopores augments peptide capture and recognition,” ACS Nano 15, 9600–9613 (2021). 10.1021/acsnano.0c09958 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123. Liu S., Lu B., Zhao Q., Li J., Gao T., Chen Y., Zhang Y., Liu Z., Fan Z., Yang F., You L., and Yu D., “ Boron nitride nanopores: Highly sensitive DNA single-molecule detectors,” Adv. Mater. 25, 4549–4554 (2013). 10.1002/adma.201301336 [DOI] [PubMed] [Google Scholar]
- 124. Gu Z., Zhang Y., Luan B., and Zhou R., “ DNA translocation through single-layer boron nitride nanopores,” Soft Matter 12, 817–823 (2016). 10.1039/C5SM02197A [DOI] [PubMed] [Google Scholar]
- 125. Al-Dirini F., Mohammed M. A., Hossain M. S., Hossain F. M., Nirmalathas A., and Skafidas E., “ Tuneable graphene nanopores for single biomolecule detection,” Nanoscale 8, 10066–10077 (2016). 10.1039/C5NR05274B [DOI] [PubMed] [Google Scholar]
- 126. Liu K., Feng J., Kis A., and Radenovic A., “ Atomically thin molybdenum disulfide nanopores with high sensitivity for DNA translocation,” ACS Nano 8, 2504–2511 (2014). 10.1021/nn406102h [DOI] [PubMed] [Google Scholar]
- 127. Farimani A. B., Min K., and Aluru N. R., “ DNA base detection using a single-layer MoS2,” ACS Nano 8, 7914–7922 (2014). 10.1021/nn5029295 [DOI] [PubMed] [Google Scholar]
- 128. Chen H., Li L., Zhang T., Qiao Z., Tang J., and Zhou J., “ Protein translocation through a MoS2 nanopore: A molecular dynamics study,” J. Phys. Chem. C 122, 2070–2080 (2018). 10.1021/acs.jpcc.7b07842 [DOI] [Google Scholar]
- 129. Nicolai A., Barrios Perez M. D., Delarue P., Meunier V., Drndic M., and Senet P., “ Molecular dynamics investigation of polylysine peptide translocation through MoS2 nanopores,” J. Phys. Chem. B 123, 2342–2353 (2019). 10.1021/acs.jpcb.8b10634 [DOI] [PubMed] [Google Scholar]
- 130. Anderson N. L. and Anderson N. G., “ The human plasma proteome: History, character, and diagnostic prospects,” Mol. Cell. Proteomics 1, 845–867 (2002). 10.1074/mcp.R200007-MCP200 [DOI] [PubMed] [Google Scholar]
- 131. Pan S., Aebersold R., Chen R., Rush J., Goodlett D. R., McIntosh M. W., Zhang J., and Brentnall T. A., “ Mass spectrometry based targeted protein quantification: Methods and applications,” J. Proteome Res. 8, 787–797 (2009). 10.1021/pr800538n [DOI] [PMC free article] [PubMed] [Google Scholar]
- 132. Howard C. J., Floyd B. M., Bardo A. M., Swaminathan J., Marcotte E. M., and Anslyn E. V., “ Solid-phase peptide capture and release for bulk and single-molecule proteomics,” ACS Chem. Biol. 15, 1401–1407 (2020). 10.1021/acschembio.0c00040 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133. Lucas F., Versloot R., Yakovlieva L., Walvoort M., and Maglia G., “ Protein identification by nanopore peptide profiling,” Nat. Commun. 12, 5795 (2021). 10.1038/s41467-021-26046-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134. Wang R., Gilboa T., Song J., Huttner D., Grinstaff M. W., and Meller A., “ Single-molecule discrimination of labeled DNAs and polypeptides using photoluminescent-free TiO2 nanopores,” ACS Nano 12, 11648–11656 (2018). 10.1021/acsnano.8b07055 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data sharing is not applicable to this article as no new data were created or analyzed in this study.





