Abstract
Proteins are a class of macromolecules with essential roles in processes and structures associated with life. Protein sequencing technologies are, therefore, fundamental for understanding cell metabolic pathways, disease mechanisms, and how pathogenic agents and toxins function. Emerging next-generation protein sequencing (NGPS) technologies promise a dramatic improvement in proteomics, enabling the identification of pathogens and toxins with unparalleled sensitivity and precision. The Quantum-Si (QSi) Platinum Sequencer is an emerging single-molecule protein sequencing technology capable of single amino acid resolution. In this work, we conducted significant optimization of the QSi protein library preparation protocol, reducing sample preparation time from 32 to 10 h without sacrificing substantial sequencing quality, allowing for a sample-to-answer timeline in less than 24 h. The modified protocol was applied for analyzing a set of proteins including 16 single-domain antibodies with diverse sequences and a nontoxic derivative of staphylococcal enterotoxin B. We were further able to determine the library dilution threshold: losing the ability to sequence beyond a 100× dilution. Finally, we were able to successfully obtain protein sequences within a crude bacterial cell lysate background, demonstrating the effectiveness of sequencing in complex protein mixtures. Improvements in sequencing chemistry and data processing may soon lessen or eliminate the dependence on reference sequences, a current obstacle for efficiently characterizing unknown proteins. By further condensing and optimizing library preparation, this technique presents a potential application for proteomics that requires rapid characterization of highly complex biological systems, significantly improving protein-based diagnostic technologies.
Introduction
Protein and peptide sequencing techniques are fundamental tools in molecular biology and biotechnology, providing insights into the primary structures of proteins, which allow us to explore their functional roles in living systems. This information may be critical for understanding protein function and disease mechanisms and provides indispensable data for the development of a broad range of drugs and therapies. Over the decades, sequencing approaches have evolved from foundational methods, such as Edman degradation, to advanced techniques utilizing mass spectrometry and fluorosequencing, offering unprecedented speed and resolution. Today, innovations such as single-molecule sequencing, top-down proteomics, and machine learning , are redefining the field by expanding its applications in research, personalized medicine, and diagnostics.
Single-molecule protein sequencing has emerged as a transformative tool for identifying proteins and peptides with unparalleled sensitivity and precision. Unlike traditional methods that require bulk, homogeneous protein samples, single-molecule techniques analyze individual protein molecules, potentially enabling the detection of rare or low-abundance peptides within complex protein mixtures. This capability may prove particularly useful for pathogen and toxin identification, where rapid detection from minimally processed samples, either clinical or environmental, is critical for early diagnosis of infectious disease or detection of toxins in the environment. The ability to characterize proteins directly without reliance on traditional proteomic technologies (such as culture-, immunoassay-, or genomic-based approaches) plausibly accelerates diagnostic timelines significantly while possibly uncovering new biomarkers and virulence factors in complex biological samples. Several protein sequencing technologies have been developed, each with its own trade-offs (for a comprehensive review on contemporary protein sequencing techniques see Alfaro et al.). Edman degradation, the first developed method to sequence a pure peptide, while transformative in the 1950s when it was first introduced, is arduous and in its recent evolution known as fluorosequencing requires chemically labeled peptides for multiplexed sequencing peptides of high purity. Mass spectrometry (MS) is another method currently used for proteome characterization and is capable of sensitive identification of proteins in complex mixtures. De novo sequencing by MS, however, relies on existing signature databases, where sequencing unknown proteins remains an unsolved challenge. Rapid advancements in applying biological and solid-state nanopores show promise as a peptide sequencing methodology, − this technology is yet to mature, thus limiting these approaches to nucleotide sequencing. Similarly, quantum field-effect nanogap sensing, which uses the quantum transport method coupled with machine learning to allow for high precision nucleotide sequencing, has been validated for DNA sequencing and speculated to work for protein sequencing, yet no commercialized device has been developed for either nanopore or nanogap protein sequencing.
The Quantum-Si (QSi) Platinum protein sequencer currently presents the only benchtop, commercialized single-molecule sequencing platform. Protein library preparation for this system is significantly simpler than the procedures typically associated with traditional sequencing methods. It achieves single-molecule protein sequencing in which individual peptides are probed, in real time, using a mixture of dye-labeled N-terminal amino acid recognizers. Proteins are digested into peptides, which are subsequently immobilized onto wells within a semiconductor chip. The sequencing process relies on iterative cycles of N-terminal amino acid recognition and cleavage by an aminopeptidase. Fluorescent-signal-based signatures provide data used to infer amino acid sequences of the immobilized peptides. These are subsequently used for alignment with the reference sequences of a specified target(s). The advantage of a benchtop sequencing system lies in the ability to decentralize protein sequencing with an accessible alternative to the complex and costly infrastructure of mass spectrometry. Similarly to next-generation nucleic acid sequencing, the ease-of-use and portability of this system permit a broader range of use-cases without requiring extensive training compared to lower-throughput methods like Edman degradation.
The platform, however, faces multiple limitations. Not every amino acid currently has a corresponding recognizer, and some amino acids share a single recognizer. This does introduce ambiguity in amino acid calling. The output of the sequencing run is thus a sequence of recognizers rather than a definitive amino acid sequence. This series of recognizers is then aligned to a reference sequence or library. Theoretical analyses have reviewed fundamental concerns on whether single-molecule protein sequencing can aptly address the wide dynamic range and complexity inherent within proteomes. To date, however, the QSi platform has primarily been evaluated with purified samples with no more than ten different proteins. Recent work has demonstrated the platform’s ability to distinguish proteoform-informative peptides, including paralogous variants that differ by single isobaric residues and post-translational modifications. Further work has demonstrated the complementarity of single-molecule sequencing with mass spectrometry, given that neither technology alone achieves full sequence coverage. It remains apparent, however, that no single technology sufficiently captures entire proteome diversity. Systematic optimization of both sample preparation and downstream analytical workflows, therefore, is needed to enable single-molecule sequencing for applications in complex biological matrices.
In this work, we address these challenges by evaluating and optimizing the QSi Platinum single-molecule workflow using version 3 (V3) sequencing chemistry. Our aims here were, therefore, to (1) modify the existing QSi library preparation protocol to demonstrate a complete sample-to-answer workflow within 24 h without substantial loss of sequencing performance, (2) evaluate the system’s performance across proteins with diverse sequences and sizes: from single-domain antibodies (sdAbs −Table ) to toxoids (SEBv) and (3) determine the platform’s capability for sequencing target proteins in mixtures of varying complexity, including crude bacterial lysates. Altogether, data here aim to advance the QSi platform toward applications that require rapid characterization of proteins in complex biological matrices, such as pathogen and toxin identification.
1. Proteins Used in the Study.
| name | type | length (aa) | MW (kD) | reference |
|---|---|---|---|---|
| V2B3 | sdAb | 135 | 14.3 | |
| V2C3 | sdAb | 126 | 13.7 | |
| V3A8f | sdAb | 138 | 15.0 | |
| WD11f | sdAb | 136 | 14.6 | |
| WE11f | sdAb | 139 | 15.1 | |
| WB9 | sdAb | 127 | 13.9 | |
| WF4 | sdAb | 126 | 13.7 | |
| WC10 | sdAb | 127 | 13.9 | |
| WH11 | sdAb | 129 | 14.2 | |
| WE10 | sdAb | 129 | 14.2 | |
| a16 | sdAb | 136 | 15.0 | |
| a18 | sdAb | 141 | 15.3 | |
| a19 | sdAb | 133 | 14.5 | |
| a86 | sdAb | 134 | 14.6 | |
| a155 | sdAb | 134 | 14.5 | |
| ACVE | sdAb | 153 | 15.9 | |
| SEBv | toxoid | 248 | 29.2 | , |
Results and Discussion
Determination of Run Performance Criteria
We used the QSi Platinum protein sequencing system in offline configuration for this study (see Methods), which does not explicitly provide an assessment of the quality of each sequencing run. Several values, however, including chip loading (CL) and number of high-quality reads (HQR) are available after performing primary analysis of sequencing data. CL, expressed as a percentage, reflects the proportion of sequencing chip wells (apertures) occupied by peptides after a library loading step. The CL value was calculated based on the count of wells occupied by one or more peptides. However, only wells containing a single peptide provide reliable sequencing results. For the 79 runs analyzed here, the observed CL values ranged between 0.2% and 80.4%, and HQR ranged from 24 to 125151. Chip loading, and to a certain degree HQR, ranges appear dependent on the protein analyzed (Figure S1G,H). Another metric considered for run quality assessment was the number of alignments for sequencing reads of an analyzed protein with its own sequence. For the purposes of this work, we considered two alignment values: total alignments (TA) and high-quality alignments (HQA), which only included alignments with false detection rates (FDRs) less than 0.05 (see Methods).
To determine the impact of chip loading on run quality and the relationship between high-quality read and alignment numbers, we calculated coefficients of determination (R 2): (1) between loading (CL) versus HQR, TA, and HQA, (2) between reads (HQR) versus alignment values (TA and HQA), and (3) between total and high-quality alignments (TA and HQA) shown in Figure S1. We observed a weak linear correlation (R 2 = 0.058) between CL and HQR (Figure S1A). We observed even weaker linear correlation (R 2 = 0.008) between CL and both TA and HQA (Figure S1B,C) in both cases. Both TA and HQA correlate with HQR (R 2 = 0.58, Figure S1E,F) and with each other (R 2 = 0.99, Figure S1C). HQR was used as the primary metric for run quality throughout, as it is a reference-independent metric for the Platinum’s sequencing output. Alignment metrics (i.e., TA and HQA) are critical for target detection but are reliant on the sequencability of the protein being sequenced along with the availability of an accurate reference sequence. An instrumentally successful run, therefore, may output high HQR but low HQA. For these reasons, TA and HQA are primarily used to assess target detection within a successful run.
The lack of strong correlation between CL and HQR is especially noticeable for loading values higher than 30%. Analysis of runs with CL at or lower than 30% was performed using only SEBv sequencing runs, as the only sample type for which a sufficient number of runs were performed in the range of CL values between 0 and 30% (Figure S1A inset). Although CL can be modulated by changing the library dilution prior to loading, resource constraints limited testing this effect with other protein libraries, therefore, limiting analysis to SEBv. On one hand, the analysis shows that for CL values <30%, HQRs are strongly correlated with CL (R 2 = 0.8) where decreasing loading is associated with a significant decrease in high-quality reads. On the other hand, changes of CL above ∼10% appear to have minimal impact on HQRs, thus consequently on sequencing results. We, therefore, deemed runs with CL above ∼10% successful while also considering HQR and HQA numbers. Exceptions here are experiments with deliberate use of lower peptide library concentrations to test the sequencer’s limit of detection.
Sample-to-Answer Protocol Optimization
Results from all sequencing rounds discussed below can be found in Table S1. We tested various method optimization steps (Figures and S2) in an attempt to either increase the quality and/or quantity of the protein sequence output or shorten the sample-to-answer protocol originally developed by QSi , without significantly compromising the sequencing results. The following modifications were tested: (a) use of an alternative protein digestion enzyme, (b) shortening the library preparation time by reduction of protein digestion and K-linker conjugation incubations, and (c) shortening the sequencing run times.
1.
Development of library preparation and sequencing protocol. (A) High-quality read (HQR) counts for V2B3 (pink triangles), V2C3 (purple squares), and V3A8f (lime circles); sdAbs decrease when digested with Tryp instead of Lys-C during library preparation. Library preparation conditions: 16-h protease digestion time (DT), 16-h K-linker incubation time (KT). Sequencing conditions: 0.4 nM loading concentration (LC), 10-h sequencing run time (RT). (B) HQR counts for V2B3 (pink triangles), V2C3 (purple squares), and V3A8f (lime circles) sdAbs remain relatively constant when decreasing Lys-C digestion time from 16 to two hours (KT:16 h, LC: 0.4 nM, RT: 10 h). (C) HQR counts for SEBv (blue diamonds) decrease for K-linker incubation shorter than 3 h (DT: 2-h, LC: 0.2 nM, RT: 10 h). (D). HQR counts for SEBv (blue diamonds) and V2C3 (purple triangles) decrease with lower sequencing run times (for both proteins - DT: 2-h, LC: 0.2 nM; V2C3 - KT: 16 h, SEBv - KT: 3 h). Linear correlation coefficient values (R 2) equal 0.89 and 0.88 for SEBv and V2C3, respectively.
Use of Trypsin for Protein Digestion
To expand the possible search space of lysed peptides, we tested the use of Trypsin (Tryp) instead of Lys-C used in the standard QSi protocol. Lys-C cleaves peptide bonds at the carboxyl side of lysine residues, while Tryp cleaves at the carboxyl side of both lysine and arginine residues. As a result, digestion with Tryp produces a higher number of shorter peptides. The main benefit of using Tryp instead of Lys-C is the introduction of additional sequencing initiation points in peptides generated by cutting at arginine (see Figure S3). While Tryp increases the number of sequencing initiation points, due to the chemistry used for peptide attachment to the chip surface by the original QSi library preparation protocol, the molecules needed to link peptides to the chip (K-linkers) cannot be conjugated with peptides ending with arginines. As a result, these peptides are unable to attach to the chip during the loading step and therefore, are not sequenced. Another complication when using Tryp for peptide generation is that QSi sequencing analysis software is designed around Lys-C digestion, thus it does not recognize arginine cutting sites when aligning reads. To bypass this limitation, arginines were replaced with lysines in the reference sequences used to generate alignments for libraries prepared using Tryp. To test the impact of Tryp substitution, we prepared libraries of three single-domain antibodies (sdAbs; ∼15 kDa binding domains): V2B3, V2C3, and V3A8f. For each of these sdAbs, two libraries were prepared: (i) the QSi standard 16-h Lys-C digestion protocol and (ii) 16-h Tryp digestion. Since Lys-C libraries had lower than expected concentration after the K-linker attachment step, they were loaded at 0.4 nM for sequencing (while the standard 0.2 nM concentrations of Tryp libraries were used). All six libraries were sequenced using the standard QSi protocol. The CL values and comparison of HQR, TA, and HQA between these two library types are shown in Figure (Figure A), Figure S2A and Table . All the CL values were found to be above 20%. We observed CL for Lys-C libraries between 63.1 and 75.9% and between 23.9 and 38.3% for Tryp libraries. To approximate an apt comparison between loading concentrations (0.4 and 0.2 nM, respectively) of Lys-C and Tryp libraries, we applied a scaling factor of 2 to the Tryp library values. The scaling factor of 2 selected is based on the observation that HQR generally scales with library loading concentration (Supporting Information, Table S1). This normalization; however, is approximate, thus, subsequent quantitative comparisons between Lys-C and Tryp libraries may contain unaccounted confounding differences and be interpreted with caution. For all libraries, HQR (both before and after the concentration-related adjustment) values are lower for Tryp than for Lys-C regardless of normalization. Specifically, the alignment numbers (both TA and HQA) are higher for V2B3 and V2C3 Tryp libraries even when not adjusted for library concentration, suggesting an increased proportion of sequencable peptides in these Tryp libraries compared with Lys-C. This was not observed in the case of V3A8f, which produced negligible TA and zero HQA for both conditions, suggesting high dependency on the composition of the protein. Given that these comparisons are confounded by the differences in loading concentration, truly definitive, quantitative conclusions are difficult to draw from this data. The results nonetheless do warrant future investigation of trypsin-based library preparation with quantitatively standardized concentrations to determine the benefit, or detriment, of trypsin digestion.
2. Use of Trypsin for Protein Digestion.
| HQR |
TA |
HQA |
HQR |
TA |
HQA |
||||
|---|---|---|---|---|---|---|---|---|---|
| sample | enzyme | library concentration (nM) | CL (%) | actual | normalized | ||||
| V2B3 | Lys-C | 0.4 | 67.7 | 15041 | 91 | 64 | 15041 | 91 | 64 |
| V2B3 | Trp | 0.2 | 31.1 | 2583 | 221 | 198 | 5166 | 442 | 396 |
| V2C3 | Lys-C | 0.4 | 63.1 | 120680 | 13017 | 12920 | 120680 | 13017 | 12920 |
| V2C3 | Trp | 0.2 | 38.3 | 47936 | 15884 | 15827 | 95872 | 31768 | 31654 |
| V3A8f | Lys-C | 0.4 | 75.9 | 6924 | 16 | 0 | 6924 | 16 | 0 |
| V3A8f | Trp | 0.2 | 23.9 | 1242 | 18 | 0 | 2484 | 36 | 0 |
Reduction of Sample-to-Answer Time
To assess the viability of shortening the library preparation time, we tested variations in (i) protease digestion time and (ii) k-linker incubation time.
-
i.
Protease Digestion Time. To test the impact of Lys-C digestion time, we prepared libraries of three sdAbs: V2B3, V2C3, and V3A8f. Two libraries were prepared for each protein with either the standard QSi 16-h Lys-C digestion protocol or 2-h Lys-C digestion. All six libraries were sequenced using the standard QSi protocol, with reported CL values between 34.8% and 74.7%. The CL values and comparison of HQR, TA, and HQA numbers between these two library types are shown in Figures B, S2B, and Table . When comparing libraries incubated for 16 and 2 h, we observed relatively minor variation in HQR for V2B3, and V2C3. We, however, observed a 2-fold decrease in HQRs for 2-h incubation in the case of V3A8f. The alignment numbers (both TA and HQA), regardless were significantly higher for all libraries using a 2-h incubation time. For V2B3 and V3A8f, the HQA numbers increased from 0 to 428 and 470 respectively, while the increase was approximately 3-fold (11024 vs 32307) for V2C3. The data show that shortening the Lys-C protease digestion time for the proteins tested did not negatively impact the quality of the sequencing run, we observe the opposite effect, indicating improved quality of obtained sequences. This result is consistent with the recommended Lys-C digestion time of 2–4 h. , The original digestion time of 16 h outlined in the QSi protocol was most likely designed around the practicality of an overnight incubation. This extended digestion time may lead to nonspecific overdigestion, causing a decrease in the proportion of high-quality sequencable peptides.
3. Protein Digestion Time.
| sample | digestion time | CL (%) | HQR | TA | HQA |
|---|---|---|---|---|---|
| V2B3 | 16 | 74.7 | 21048 | 123 | 0 |
| V2B3 | 2 | 62.1 | 21962 | 976 | 961 |
| V2C3 | 16 | 61.8 | 125151 | 11087 | 11024 |
| V2C3 | 2 | 68.7 | 110034 | 32359 | 32307 |
| V3A8f | 16 | 62.6 | 19779 | 46 | 0 |
| V3A8f | 2 | 34.8 | 8521 | 486 | 470 |
-
ii.
K-Linker Incubation Time. The K-linker incubation was another time-consuming step of the standard QSi library preparation protocol targeted for optimization. To explore the impact of the K-linker incubation time on sequencing quality, we prepared four libraries of an inactivated staphylococcal enterotoxin B derivative (SEBv, see Methods). The libraries were prepared using a modified QSi protocol with short, 2 h. Lys-C incubation, with four K-linker incubation times at 16, 3, 2, or 1 h. All four libraries were sequenced using the standard QSi protocol with reported CL values between 26.5% (3-h incubation) and 0.2% (1-h incubation), shown in Figures D, S2D, and Table .
4. K-Linker Conjugation Time.
| sample | conjugation time | CL (%) | HQR | TA | HQA |
|---|---|---|---|---|---|
| SEBv | 16 | 19.6 | 20178 | 2355 | 2054 |
| SEBv | 3 | 26.5 | 25159 | 2962 | 2617 |
| SEBv | 2 | 13.9 | 21388 | 2079 | 1901 |
| SEBv | 1 | 0.3 | 35 | 0 | 0 |
The CL values for the libraries with K-linker incubation of 2 h or more fluctuated around 20%, with the highest value (26.5%) for the 3-h incubation and the lowest for the 2-h incubation (13.9%). Despite the low loading (as compared to previously analyzed sdAb proteins), we observed over 20,000 HQRs for these three libraries with relatively high (∼2000) TA and HQA numbers. We noticed, however, that the 2-h, K-linker-incubated library compared to the 3-h incubated one shows a decrease in HQA number by 30% while the HQR is lower by only 15%, suggesting a significant deterioration of the read quality. For the 1-h, K-linker-incubated library, the CL value dropped below 1%, HQR below 100, and zero alignments to the reference sequence were found. This indicates a complete failure of the library. Similar values were observed in four runs without the library loaded onto the chip (Supporting Information, Table S2). These results, therefore, indicate the minimum K-linker incubation time for SEBv without significant deterioration is 2–3 h.
Shortened Library Protocol
On the basis of the results presented above, we designed a modified library preparation protocol for use in subsequent parts of this study. The protocol included the use of Lys-C for protein digestion, 2-h Lys-C digestion, and 3-h K-linker incubation. The elimination of overnight incubations means that this modified protocol reduces the total time needed to prepare a high-quality sequencing library from 37 h to merely 10 h.
Sequencing Run Duration
To test the impact of run duration on sequencing quality, we conducted a series of tests with two proteins: SEBv and V2C3. Five sets of runs were conducted with run times decreasing from the standard 10 h. to 2 h. CL, HQR, TA, and HQA obtained for these runs are shown in Table , Figures C, and S2C. CL values for SEBv ranged from 24.1% to 30.6% and from 21% to 68.7% for V2C3. In both cases, the decrease in the runtime resulted in reduced HQRs, which were generally mirrored by the decrease of alignments (both TA and HQA). The decrease of the runtime to 8 h. resulted in approximately 20% decrease in HQRs. Both the 6-h and 4-h sequencing runs resulted in approximately 80% reduced HQR numbers compared to the 10-h run. For the 2-h runs, HQRs were reduced by more than 90% for both samples. These results indicate that, despite a high reduction of obtained reads, 2-h runs may still be capable of producing enough reads to allow for protein identification. This outcome, however, will be strongly dependent on the protein’s sequencability and its concentration in the analyzed sample.
5. Sequencing Run Time.
| sample | run time | CL (%) | HQR | TA | HQA |
|---|---|---|---|---|---|
| SEBv | 10 | 26.5 | 25159 | 2962 | 2617 |
| 8 | 28.1 | 20651 | 2066 | 1787 | |
| 6 | 24.1 | 5804 | 670 | 607 | |
| 4 | 30.6 | 5124 | 689 | 637 | |
| 2 | 27 | 1387 | 161 | 128 | |
| V2C3 | 10 | 68.7 | 110034 | 32359 | 32307 |
| 8 | 47.8 | 87597 | 35244 | 35155 | |
| 6 | 21 | 26326 | 11743 | 11720 | |
| 4 | 35.8 | 28909 | 13417 | 13407 | |
| 2 | 32.7 | 9602 | 5202 | 5202 |
Testing Optimized Protocol for Samples of Varying Composition
sdAbs with Diverse Sequences
Sixteen different sdAbs were sequenced to test the sequencer’s performance for samples with diverse amino acid compositions. See Table for the detailed information about the sdAbs used. The sequencing libraries were prepared using the shortened protocol described above and loaded at 0.2 nM except for three samples loaded at 0.4 nM (V2B3, V2C3, and V3A8f) for which an earlier version of the library protocol was used (with 2-h Lys-C digestions and 16-h K-linker incubation). The results of the sequencing, including CL values and HQR, TA, and HQA numbers are shown in Table . While multiple samples were sequenced only once, many selected samples were run multiple times (from 2 to 19 replicates) to explore the reproducibility of the sequencing system. Variation in replicate numbers across sample types simply reflects the availability of sequencing chips, reagents, as well as prioritization of samples used for method optimization. We analyzed run-to-run reproducibility of three samples (two sdAbs: a16, WE11f and SEBv) by analyzing variability of CL, HQRs, and both types of alignments (Figure , Tables , and S3). The variability appears to be relatively high and specific to sample type (higher for sdAbs vs SEBv). To investigate whether library degradation from repeated freeze–thaw cycles contributed to this variability, we analyzed the effects of freeze–thaw cycles and sequencing quality for the sdAbs a16 and WE11f (Figure S4). Remarkably, we found no substantial correlation between freeze–thaw cycle and either HQR or HQA (a16: R 2 = 0.014 and 0.105; WE11f: R 2 = 0.157 and 0.217, for HQR and HQA, respectively). This suggests that freeze–thaw cycles did not appear to substantially contribute to the variability observed. The variability observed, therefore, may plausibly be a reflection of other confounding factors such as changes in ambient temperature or simply inherent stochasticity in the sequencing architecture.
6. Sequencing of Diverse sdAbs.
|
|
CL |
HQR |
TA |
HQA |
|||||
|---|---|---|---|---|---|---|---|---|---|
| sample | replicates | % | std dev | n | std dev | n | std dev | n | std dev |
| V2B3 | 1 | 62.1 | n/a | 21962 | n/a | 649 | n/a | 428 | n/a |
| V2C3 | 1 | 68.7 | n/a | 110034 | n/a | 32331 | n/a | 32307 | n/a |
| V3A8f | 1 | 34.8 | n/a | 8521 | n/a | 478 | n/a | 470 | n/a |
| WD11f | 2 | 71.7 | 0.2 | 4041 | 1159 | 19 | 1 | 0 | 0 |
| WE11f | 19 | 64.3 | 14 | 18237 | 12735 | 3213 | 2494 | 3190 | 2485 |
| WB9 | 1 | 52.5 | n/a | 1400 | n/a | 170 | n/a | 168 | n/a |
| WF4 | 1 | 46.5 | n/a | 10639 | n/a | 33 | n/a | 0 | n/a |
| WC10 | 1 | 75.8 | n/a | 10928 | n/a | 1235 | n/a | 1222 | n/a |
| WH11 | 1 | 75.6 | n/a | 26645 | n/a | 2181 | n/a | 2163 | n/a |
| WE10 | 1 | 71.3 | n/a | 15879 | n/a | 1594 | n/a | 1540 | n/a |
| a16 | 9 | 61.7 | 17 | 30388 | 17894 | 3884 | 2650 | 3864 | 2643 |
| a18 | 2 | 63.4 | 10.2 | 3632 | 231 | 35.0 | 6 | 14.5 | 21 |
| a19 | 1 | 46.7 | n/a | 5031 | n/a | 143 | n/a | 121 | n/a |
| a86 | 1 | 52.9 | n/a | 4287 | n/a | 20 | n/a | 0 | n/a |
| a155 | 1 | 65.1 | n/a | 6578 | n/a | 1038 | n/a | 1031 | n/a |
| ACVE | 3 | 62.1 | 2.4 | 30263 | 1445 | 5225 | 2096 | 5215 | 2095 |
Library of all sdAbs was prepared using the optimized/shortened protocol, including 2-h Lys-C digestion and 3-h K-linker incubation, and was loaded at a concentration of 0.2 nM, except for V2B3, V2C3, and V3A8f for which 16-h K-linker incubation was performed, loaded at 0.4 nM. All libraries were sequenced using a standard 10-h QSi protocol. Italic font was used for four samples (WD11f, WF4, a18, and a86) with no or very low HQA numbers, indicating low efficiency of sequencing.
2.
Variability among QSI sequencing runs across three proteins: a16, SEBv, and WE11f. (A) Box and whisker plots of high-quality reads for a16 (N = 8, x̅ = 30356, std = 20289), SEBv (N = 9, x̅ = 15055, std = 6567), and WE11f (N = 19, x̅ = 18237, std = 12735). (B) Box and whisker plots of chip loading percentage for a16 (x̅ = 61, std = 20), SEBv (x̅ = 31, std = 31), WE11f (x̅ = 64, std = 14). (C) Box and whisker plots of total alignments to each respective protein for a16 (x̅ = 3794, std = 2991), SEBv (x̅ = 1655, std = 809), and WE11f (x̅ = 3213, std = 2494). (D) Box and whisker plots of high-quality alignments to each respective protein for a16 (x̅ = 3777, std = 2984), SEBv (x̅ = 1447, std = 745), and WE11f (x̅ = 3190, std = 2485).
9. Run-to-Run Variability–Summary.
| CL |
HQR |
TA |
HQA |
||||||
|---|---|---|---|---|---|---|---|---|---|
| sample | replicates (n) | % | std dev | n | std dev | n | std dev | n | std dev |
| a16 | 8 | 61 | 20 | 30356 | 20289 | 3794 | 2991 | 3777 | 2984 |
| WE11f | 19 | 64 | 14 | 18237 | 12735 | 3213 | 2494 | 3190 | 2485 |
| SEBv | 9 | 31 | 11 | 15055 | 6567 | 1655 | 809 | 1447 | 745 |
CL for all analyzed sdAb runs varied between 34.8% and 75.8%, with the majority in the 50%–70% range. HQRs ranged from 1400 (WB9) to 110034 (V2C3) with a median value of 10783. The alignment numbers in general correlate with HQR apart from several exceptions. Despite a significant number of HQRs (ranging from 3632 to 10639), the alignment numbers for WD11f, WF4, a18, and a86 were within levels associated with experimental noise. The TA for all four samples was below 40 and there were no HQAs for three of them. Multiple runs of some of these samples have shown similar low alignment numbers, making a technical run failure an unlikely explanation of these results. More probably, an explanation of the low number of alignments is, therefore, the QSi Platinum sequencing system’s inability to sequence peptides comprising certain amino acid sequences (e.g., proline-rich peptides).
Diluted Samples
To explore the QSi Platinum system’s ability to sequence samples containing lower initial concentrations of target proteins, we sequenced a series of SEBv library dilutions (Table ). For the 10x dilution library, both the chip loading and read numbers decreased significantly, with CL dropping below 5% and HQRs and alignments to approximately 10% of what is typical in a 0.2 nM library. For the 100× diluted library, CL, HQRs, and alignments dropped to levels observed for runs without samples loaded on the chip.
7. Sequencing of Diluted Libraries.
| sample | library concentration (nM) | CL (%) | HQR | TA | HQA |
|---|---|---|---|---|---|
| SEBv | 0.2 | 29.4 | 22871 | 2496 | 2231 |
| 0.02 | 4.5 | 2460 | 264 | 217 | |
| 0.002 | 0.4 | 77 | 9 | 0 |
Mixed Proteins and Complex Backgrounds
To explore the performance of the QSi Platinum sequencing system’s ability to sequence both mixed protein samples and those within complex backgrounds, we sequenced two types of samples: (i) a mixture of libraries containing SEBv and V2C3 and (ii) a series of libraries with raw bacterial lysate background including: libraries prepared using crude lysates of Escherichia coli cells overexpressing two different proteins (SEBv and ACVE), libraries prepared from partially purified and pure preparations of SEBv and ACVE and libraries prepared from the same strain of E. coli cells without protein expression plasmids. The results of these experiments are summarized in Table , where CL ranged from 22.7% to 62.3%.
8. Sequencing of Mixed Libraries and Samples in Complex Matrices.
| CL |
HQR |
SEBv |
V2C3 |
||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TA | HQA | TA | HQA | ||||||||||
| sample | replicates | n | std dev | n | std dev | n | std dev | n | std dev | n | std dev | n | std dev |
| SEBv+V2C3 | 1 | 30.0 | 29277 | 916 | 811 | 10601 | 10574 | ||||||
| SEBv | ACVE | ||||||||||||
| SEBv-1 | 1 | 32.2 | n/a | 6541 | n/a | 381 | n/a | 179 | n/a | 13 | n/a | 0 | n/a |
| SEBv-2 | 1 | 22.7 | n/a | 6577 | n/a | 458 | n/a | 390 | n/a | 4 | n/a | 0 | n/a |
| SEBv-3 | 1 | 47.1 | n/a | 9796 | n/a | 764 | n/a | 683 | n/a | 6 | n/a | 0 | n/a |
| ACVE-1 | 2 | 34.9 | 17.3 | 8018 | 4723 | 26 | 15 | 0 | 0 | 435 | 243 | 427 | 233 |
| ACVE-2 | 2 | 49.7 | 1.7 | 20505 | 12262 | 35 | 8 | 0 | 0 | 2295 | 1583 | 2285 | 1579 |
| ACVE-3 | 2 | 61.0 | 1.8 | 29633 | 1337 | 29 | 1 | 0 | 0 | 4019 | 221 | 4009 | 214 |
| E. coli | 2 | 62.3 | 5.8 | 11264 | 5093 | 88 | 31 | 0 | 0 | 36 | 6 | 0 | 0 |
-
(i)
Mixed Library. Sequencing the library created by mixing previously prepared libraries of SEBv and V2C3 in a 1:1 proportion produced a high number of HQRs (n = 29277) and alignments with SEBv and V2C3 reference sequences, resulting in high numbers of TA and HQA for both references. The number of alignments for each reference was reduced by approximately 60% when compared to running each library individually: for HQA, the alignments dropped from 2617 to 811 in the case of SEBv and from 32307 to 10574 in the case of V2C3. This reduction in HQA is plausibly a reflection of the lower concentrations of peptides from each protein in the library. The results, though, demonstrate that both proteins may be easily detected within a mixed sample.
-
(ii)
Libraries with complex backgrounds To investigate the ability of the QSi Platinum system to obtain sequences of proteins present in a background of complex mixtures, we prepared libraries using the following samples: (i) three libraries of crude E. coli lysates expressing SEBv, ACVE, or the same E. coli cells not containing any expression vectors as a negative control, (ii) two libraries of enriched SEBv and ACVE preparations; the first step of protein purification by nickel affinity enrichment, and (iii) two libraries of SEBv and ACVE preparations fully purified from the aforementioned lysates. All libraries were sequenced using the standard 10-h sequencing protocol. Libraries containing SEBv were sequenced once, while the ACVE and nonexpressing E. coli libraries were prepared and run in duplicate. For each of the sequencing runs, alignment analysis was performed using SEBv and ACVE reference sequences. These results are summarized in Table and Figure . We observed high HQRs for all libraries with generally lower HQR (6541–9796) for SEBv libraries and higher HQR (8017–29632) for libraries containing ACVE and the controls not expressing any extrinsic proteins (Table ).
3.
Sequencing SEBv on the QSi Instrument after various stages of protein purification. SEBv high-quality reads (SEBv + ACEV -; blue bars) are plotted at multiple purification stages after expression in E. coli: (i) E. coli lysate, (ii) Nickel -purification, (iii) FPLC-purification (see methods). E. coli expressing ACVE (SEBv–ACVE +; orange bars) are plotted at multiple purification stages after expression in E. coli: (i) E. coli lysate, (ii) Nickel purification, and (iii) FPLC purification. Nonexpressing E. coli lysate number of high-quality reads (gray bar) shown as the negative control. Blue horizontal lines and values represent the number of high-quality alignments to enterotoxin SEBv, whereas orange horizontal lines and values equal high-quality alignments to sdAb ACEV.
When comparing runs from raw lysate through partially purified to pure protein, we observed an increase in HQRs with increasing purity and a decrease in the complexity of the library. For both SEBv- and ACVE-containing lysates, we observed alignment numbers (both TA and HQA) significantly above the background for each specific reference sequence (e.g., SEBv for SEBv-containing libraries), while only background level numbers of alignments for the other reference sequence (e.g., ACVE for SEBv-containing libraries). The number of alignments increased with increasing purity of the protein in the library, with numbers of alignments for pure proteins around an order of magnitude higher than the number for crude cell lysates. We observed background noise level numbers of alignments for both reference sequences for the libraries containing controls that did not express any extrinsic proteins. Notably, these runs, especially libraries prepared from crude E. coli lysates expressing or not expressing extrinsic proteins, are complex mixtures containing hundreds of different proteins. Results here, in context of the original use case for QSi sequencing - which is recommended for mixtures of no more than 10 proteins, demonstrate that even low HQAs are remarkable, as this potentially indicates the ability to detect proteins in complex mixtures, at least for the proteins tested in this study.
Conclusions
Our revised sequencing protocol demonstrated a significant reduction in the sample-to-answer timeline, enabling a complete sequencing workflow in less than 24 h. Results presented here furthermore suggest that this protocol could still be shortened more by optimizing sequencing runtime without a substantial loss in sequencing quality. We additionally confirmed the applicability of this approach for a range of proteins with diverse sequences, including single-domain antibodies and a staphylococcal enterotoxin. However, proteins with sizes and compositions different from those tested here could perform differently. Additionally, future validation of the relationship between CL and sequencing output over a range of proteins would substantiate the applicability of this relationship to proteins beyond SEBv, therefore, strengthening run quality assessment criteria. Regardless, this method will significantly benefit from the improved sequencing chemistry offered in the V4 QSi sequencing kit, which allows for improved detection of alanine and serine, permits for the detection of glycine, and expands the sequencable range of proline-containing peptides.
Beyond the optimized sequencing protocol presented here, we have shown that it is possible to obtain sequences for diluted target libraries and target proteins in the presence of backgrounds more complex than currently indicated by the manufacturer. While the current QSi protocol does not recommend sequencing of mixtures of more than 10 proteins, we were successful in sequencing and detecting targets in matrices of significantly higher complexity. This should also be improved with the introduction of the V4 sequencing kit.
However, while these results certainly advance the field of single-molecule protein sequencing, especially within time-sensitive applications, a clear limitation lies at the reliance on reference sequences during the analysis of primary data from the instrument. In other words, due to the nature of the current (V3) sequencing chemistry, obtaining de novo sequences for unknown targets is very limited.
To advance the QSi technology on its progress toward truly de novo protein sequencing, significantly improved chemistry is needed. In addition, the analysis of the raw sequencing output will certainly benefit from advancements in data processing technologies that take advantage of rapidly improving machine learning tools. Even with progress here, significant further advancements are required to sequence modified amino acids and nonproteinogenic amino acids at the heart of rapidly advancing contemporary biotechnology. , Overall, the QSi Platinum technology has the potential for application in many areas of proteomics that require rapid high-resolution characterization of complex biological samples to significantly improve protein-based diagnostic technologies.
Methods
Protein Samples
The protein preparations used in this study included inactivated staphylococcal enterotoxin B (SEBv) and 16 single-domain antibodies (sdAbs). See Table for the detailed protein information. All of the protein samples were synthesized and purified in-house, as described below.
SEBv
The coding sequence for SEBv was PCR amplified from the previously described pET15-based expression vector to introduce flanking Nco I and Not I restriction sites. The PCR product was purified using a QIAquick PCR purification kit (Qiagen) prior to digesting with Nco I and Not I. The digested fragment was ligated into pET22b that had been similarly digested and treated with CIP. The SEBv expression vector was transformed into the Tuner (DE3) strain of E. coli, and a fresh colony was used to start a 50 mL overnight culture in Terrific Broth supplemented with ampicillin at 100 μg/mL (TB-amp). All growth steps in TB-amp, including the overnight incubation, were conducted at 25 °C. The overnight culture was used to inoculate 450 mL TB-amp in a 2L baffled flask and grown for a further ∼8 h before induction with IPTG (0.5 mM final concentration). Induced cultures were grown overnight. The next morning, the cells were centrifuged, and cell pellets were subjected to an osmotic shock protocol; cells were kept on ice throughout the osmotic shock procedure. Briefly, the pellets were resuspended in 14 mL of cold Tris-Sucrose buffer (100 mM Tris, 0.75 M sucrose pH 7.5), and 1 mL of lysozyme (1 mg/mL made up in the Tris-Sucrose buffer) was added to the homogenized cells. As the cells were shaking on a rotating platform, 28 mL of 1 mM EDTA was added dropwise. After 15 min, 1 mL of 0.5 M MgCl2 was added and the mix was incubated for a further 15 min, prior to pelleting the spheroplasts. The supernatant (termed the lysate) was poured into a 50 mL conical tube that contained 5 mL of 10× IMAC buffer (0.2 M Na2HPO4, 4 M NaCl, 0.2 M imidazole, pH 7.5) and 0.5 mL of Ni Sepharose (GE Healthcare) that had been equilibrated with 1× IMAC buffer. The sample was tumbled for between 1 and 2 h at 4 °C on a rotisserie. Afterward, resin was washed twice in batch with 25 mL 1× IMAC buffer. The resin was poured into a small column, washed with a further ∼10 mL 1× IMAC buffer, and eluted with 1 mL of 1× IMAC buffer containing 250 mM imidazole. Protein was then further purified into PBS by size exclusion chromatography using a Bio-Rad Enrich SEC70 10 300 and a Bio-Rad Duo-Flow System.
sdAbs
The sdAbs used in this study were purified as described previously. Briefly, the expression plasmid was transformed into the Tuner (DE3) and a single colony was used to start an overnight culture as described above. The overnight culture was used to inoculate 450 mL TB-amp, grown for 2 h followed by induction with IPTG (final concentration 0.5 mM) and grown for a further 2 h. Cells were pelleted and processed using our standard periplasmic production protocol and purified through a combination of immobilized metal affinity chromatography (IMAC), followed by size exclusion chromatography as described for SEBv above. All protein preparations were aliquoted and stored frozen at −80 °C until use.
Library Preparation
The purified protein preparations were quantified using Qubit Protein Assay Kit or Qubit Protein BR Assay Kit and Qubit 4 Fluorometer (Thermo Fisher Scientific, Waltham, MA). Then they were processed to prepare peptide libraries ready for loading on sequencing chips according to the Quantum Si Inc. (QSi) library preparation protocol using the Quantum-Si Library Preparation Kit–Lys-C.
Briefly, buffer exchange was performed to remove buffer components incompatible with library preparation or sequencing. The buffer replacement was accomplished using Amicon Ultra-0.5 3 kDa Centrifugal Filter Units (Sigma-Aldrich, Inc., St. Louis, MO) following the QSi library preparation protocol. Sample buffer provided in the QSi library preparation reagent kit was used as the replacement buffer. Subsequently, the concentration of the samples was measured again as described earlier and diluted to obtain 100 μL at 5 μM concentration. Diluted sample was subjected to cysteine reduction by 30 min incubation with TCEP (Tris(2-carboxyethyl) phosphine hydrochloride) at 37 °C to disrupt disulfide bridges. Next, exposed thiol groups on cysteine side chains were alkylated by 30 min incubation with CAA (Chloroacetamide) at room temperature (RT) to prevent the disulfide bridges from reforming. The reduced proteins were then digested using Lys-C endopeptidase (or trypsin). The endopeptidase digestion was conducted at 37 °C overnight/16 h (standard protocol) or for 2 h (shortened protocol). As a result of endopeptidase digestion, variable-length peptides (with a lysine (K) residue at the C-terminus in case of Lys-C and with a lysine (K) or arginine (R) residue at the C-terminus in case of trypsin) were obtained. In the next step, the pH of the peptide mixture was adjusted by adding potassium carbonate (K2CO3), and lysines were converted to azido-terminated lysines by 90 min incubation with ISA (imidazole-1-sulfonyl azide) and copper(II) sulfate (CuSO4) at RT. Then, the obtained mixture was incubated with polyamide beads with shaking for 30 min at RT to quench the excess of ISA. Next, the polyamide beads were removed by filtration, and the pH of the derivatized peptide mixture was adjusted by the addition of acetic acid. In the next step, CTAB and K-linker were added to the peptide preparation and K-linker was conjugated to the peptides during incubation at 37 °C overnight/16 h (standard protocol) or for a shorter period (as short as 2 h). Finally, quantification of the K-linked peptide library was conducted using a Qubit 4 Fluorometer using the fluorometer function of the instrument. Measurements were performed using the Blue (470 nm) excitation setting, and RFU values from the Green (510–580 nm) emission were used for standard curve generation and sample concentration measurements. The libraries were stored at −80 °C until use. The libraries were diluted to 0.2 nM (or in some cases to 0.4 nM) before loading on the sequencing chips.
Protein Sequencing
All protein sequencing in this work followed the QSi Platinum Sequencing Kit V3 Protocol using sequencing V3 chips and sequencing chemistry. We optimized this workflow, however, for shortened runs (see Figure and Supporting Information) as compared to the original QSi method. The QSi chip was first washed three times by pipetting 50 μL QSi Wash Buffer into the upper flow cell ports and pipette-mixed ten times. The chip was then loaded into the QSi sequencer for a chip-check with 50 μL of QSi Wash Buffer loaded. Following a successful chip-check, each side of the chip was then loaded with 30 μL of loading solution containing properly diluted libraries prepared according to QSi sequencing protocol and was mixed by pipette ten times on each side. Following a 15 min incubation at room temperature, the loading solution was removed and 30 μL of imaging solution (nuclease-free water and kit components: Seq Buffer 3, Quencher, GOx, and Trolox) was loaded to the upper flow cell ports by pipette mixing ten times, and the chip was then loaded into the sequencer for a loading check. After a successful loading check, the imaging solution was removed from the chip and 27 μL of recognizer solution (nuclease-free water and kit components: Seq Buffer 3, Quencher, GOx, Trolox, COT, and Recognizer Mix 3) was loaded into the upper flow cell ports by pipette mixing ten times and the chip was reinserted into the QSi instrument for initial sequencing. Following 15 min of sequencing, 3 uL of aminopeptidase mix (AP mix 3) was added to the chip’s upper reservoir, mixed three times with the existing Recognizer Solution previously loaded onto the chip with a pipette set to 12 μL, and then loaded into the flow cell. The aminopeptidases were mixed by pipetting ten times in both the upper and lower flow cell ports. Subsequently the chip was sealed with a chip plug before reloading it onto the QSi instrument for a total of a remaining 9.75-h sequencing run, 10 h in total. The total run length was adjusted to shorter times for length of run experiments. For run-to-run variability see Tables and S3.
Sequencer Configuration and Sequencing Data Analysis
A Quantum-Si Platinum sequencer can be set up in two different configurations, which differ in the way the data analysis is performed. In the online configuration the data are sent to a remote server where it is processed and analyzed using a web-based QSi Platinum Analysis Software. However, for this study, we used an offline configuration in which the data is transferred from the sequencing unit to a local server. In the offline system, no Internet access is needed for data analysis. The data analysis was performed locally using a web browser-based user interface and involved running analysis scripts provided by QSi. Upon completion of sequencing, all runs described in this study were analyzed by executing the “peptide_alignment_v2.5.2” script, which performed both primary data analysis and alignment to a reference sequence.
Primary data analysis provided high-level run metrics, including chip loading (CL) and number of high-quality reads (HQR), and generated sequencing reads subsequently used for alignment. CL is defined as the percentage of analyzed apertures (chip wells) that were loaded, including single-loaded and multiloaded apertures, and HQRs are defined as reads with recognizer read lengths ≥ 4 and unique active recognizers ≥ 3.
Alignment analysis required a reference sequence, which for most analyses in this study, was the sequence of the analyzed protein. Alignment analysis produced a list of reference sequence-derived peptides with the number of sequencing reads aligning with each of these peptides and a false discovery rate (FDR) for each of these alignment sets. FDR is an alignment quality metric developed by QSi using a decoy generation method adapted from methods used in peptide identification by mass spectrometry. The alignments with FDR < 0.05 were considered high-quality alignments (HQA) and were a subset of the total alignment (TA) number, which included all observed alignments.
Statistical and Graphics Software
Figures and statistical calculations were made using Microsoft Excel (v. 2502) and Python (v. 3.11.11, https://www.python.org) with the following packages: seaborn (v. 0.13.2, https://seaborn.pydata.org), pandas (v. 2.2.3, https://pandas.pydata.org), matplotlib (v. 3.10.0, https://matplotlib.org), and numpy (v. 2.2.5, https://numpy.org).
Supplementary Material
Acknowledgments
This work was supported by the Defense Threat Reduction Agency (CB11498) and the National Academies of Sciences Engineering and Medicine’s NRC Research Associateship Program. The opinions and assertions contained herein are those of the authors and are not to be construed as those of the U.S. Navy, military service at large, or the U.S. Government. T.A.L., Z.T.J., S.N.D., E.R.G., and D.A.S. are employees of the US Government.
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsomega.6c01140.
Correlation between various sequencing metrics and their ranges for all runs; impact of varying incubation times during library preparation and sequencing run duration on sequencing performance; impact of protease selection on sample sequencability; impact of sample freeze-thaw cycles on sequencing performance (Figures S1–S4); and QSi sequencing result summaries for all samples; no library control run results; and detailed results of run to run variability study for three selected samples (Tables S1–S3) (PDF)
The authors declare the following competing financial interest(s): The authors completed this work while in collaborative efforts with QSi. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.
References
- Alfaro J. A., Bohländer P., Dai M.. et al. The emerging landscape of single-molecule protein sequencing technologies. Nat. Methods. 2021;18:604–617. doi: 10.1038/s41592-021-01143-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Restrepo-Pérez L., Joo C., Dekker C.. Paving the way to single-molecule protein sequencing. Nat. Nanotechnol. 2018;13:786–796. doi: 10.1038/s41565-018-0236-6. [DOI] [PubMed] [Google Scholar]
- Edman P., Högfeldt E., Sillén L. G.. et al. Method for Determination of the Amino Acid Sequence in Peptides. Acta Chem. Scand. 1950;4(7):283–293. doi: 10.3891/acta.chem.scand.04-0283. [DOI] [Google Scholar]
- Miyashita M., Presley J. M., Buchholz B. A.. et al. Attomole level protein sequencing by Edman degradation coupled with accelerator mass spectrometry. Proc. Natl. Acad. Sci. U.S.A. 2001;98(8):4403–4408. doi: 10.1073/pnas.071047998. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Deol H., Raeisbahrami A., Ngo P. H.. et al. After 75 Years, an Alternative to Edman Degradation: A Mechanistic and Efficiency Study of a Base-Induced Method for N-Terminal Peptide Sequencing. J. Am. Chem. Soc. 2025;147(16):13973–13982. doi: 10.1021/jacs.5c03385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reed B. D., Meyer M. J., Abramzon V.. et al. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device. Science. 2022;378(6616):186–192. doi: 10.1126/science.abo7651. [DOI] [PubMed] [Google Scholar]
- Roberts D. S., Loo J. A., Tsybin Y. O.. et al. Top-down proteomics. Nat. Rev. Methods Primers. 2024;4(1):38. doi: 10.1038/s43586-024-00318-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bordin N., Dallago C., Heinzinger M.. et al. Novel machine learning approaches revolutionize protein knowledge. Trends Biochem. Sci. 2023;48(4):345–359. doi: 10.1016/j.tibs.2022.11.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gallo E.. Revolutionizing Synthetic Antibody Design: Harnessing Artificial Intelligence and Deep Sequencing Big Data for Unprecedented Advances. Mol. Biotechnol. 2025;67:410–424. doi: 10.1007/s12033-024-01064-2. [DOI] [PubMed] [Google Scholar]
- Reed B. D., Meyer M. J., Abramzon V.. et al. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device. Science. 2022;378(6616):186–192. doi: 10.1126/science.abo7651. [DOI] [PubMed] [Google Scholar]
- Edman P.. A method for the determination of amino acid sequence in peptides. Arch. Biochem. Biophys. 1949;22(3):475–476. doi: 10.1016/j.abb.2022.109297. [DOI] [PubMed] [Google Scholar]
- Swaminathan J., Boulgakov A. A., Hernandez E. T.. et al. Highly parallel single-molecule identification of proteins in zeptomole-scale mixtures. Nat. Biotechnol. 2018;36:1076–1082. doi: 10.1038/nbt.4278. [DOI] [PMC free article] [PubMed] [Google Scholar]
- O’Bryon I., Jenson S. C., Merkley E. D.. Flying blind, or just flying under the radar? The underappreciated power of de novo methods of mass spectrometric peptide identification. Protein Sci. 2020;29(9):1864–1878. doi: 10.1002/pro.3919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cao C., Cirauqui N., Marcaida M. J.. et al. Single-molecule sensing of peptides and nucleic acids by engineered aerolysin nanopores. Nat. Commun. 2019;10(1):4918. doi: 10.1038/s41467-019-12690-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wilson J., Sarthak K., Si W.. et al. Rapid and Accurate Determination of Nanopore Ionic Current Using a Steric Exclusion Model. ACS Sens. 2019;4(3):634–644. doi: 10.1021/acssensors.8b01375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Di Muccio G.. et al. Insights into protein sequencing with an alpha-Hemolysin nanopore by atomistic simulations. Sci. Rep. 2019;9(1):6440. doi: 10.1038/s41598-019-42867-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ryu J., Komoto Y., Ohshiro T.. et al. Direct biomolecule discrimination in mixed samples using nanogap-based single-molecule electrical measurement. Sci. Rep. 2023;13(1):9103. doi: 10.1038/s41598-023-35724-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Taniguchi M.. Combination of Single-Molecule Electrical Measurements and Machine Learning for the Identification of Single Biomolecules. ACS Omega. 2020;5(2):959–964. doi: 10.1021/acsomega.9b03660. [DOI] [PMC free article] [PubMed] [Google Scholar]
- MacCoss M. J., Alfaro J. A., Faivre D. A.. et al. Sampling the proteome by emerging single-molecule and mass spectrometry methods. Nat. Methods. 2023;20(3):339–346. doi: 10.1038/s41592-023-01802-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Quantum Si Inc . Library Preparation Kit - Protein Protocol.; Quantum Si Inc: Branford, CT, 2024. [Google Scholar]
- Sittipongpittaya N., Skinner K. A., Jeffery E. D.. et al. Protein Sequencing with Single Amino Acid Resolution Discerns Peptides That Discriminate Tropomyosin Proteoforms. J. Proteome Res. 2025;24(8):3798–3807. doi: 10.1021/acs.jproteome.4c00978. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Skinner K. A., Fisher T. D., Lee A.. et al. Next-generation protein sequencing and individual ion mass spectrometry enable complementary analysis of interleukin-6. Anal. Bioanal. Chem. 2025;417(28):6291–6299. doi: 10.1007/s00216-025-06120-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu J. L., Zabetakis D., Gardner C. L.. et al. Bivalent single domain antibody constructs for effective neutralization of Venezuelan equine encephalitis. Sci. Rep. 2022;12(1):700. doi: 10.1038/s41598-021-04434-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu J. L., Shriver-Lake L. C., Zabetakis D.. et al. Selection of Single-Domain Antibodies towards Western Equine Encephalitis Virus. Antibodies. 2018;7(4):44. doi: 10.3390/antib7040044. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liu J. L., Bayacal G. C., Alvarez J. A. E.. et al. Generative Deep Learning Design of Single-Domain Antibodies Against Venezuelan Equine Encephalitis Virus. Antibodies. 2025;14(2):41. doi: 10.3390/antib14020041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Anderson G. P., Liu J. L., Shriver-Lake L. C.. et al. Oriented Immobilization of Single-Domain Antibodies Using SpyTag/SpyCatcher Yields Improved Limits of Detection. Anal. Chem. 2019;91(15):9424–9429. doi: 10.1021/acs.analchem.9b02096. [DOI] [PubMed] [Google Scholar]
- Ulrich R. G., Olson M. A., Bavari S.. Development of engineered vaccines effective against structurally related bacterial superantigens. Vaccine. 1998;16(19):1857–1864. doi: 10.1016/S0264-410X(98)00176-5. [DOI] [PubMed] [Google Scholar]
- Anderson G. P., Legler P. M., Zabetakis D.. et al. Comparison of immunoreactivity of staphylococcal enterotoxin B mutants for use as toxin surrogates. Anal. Chem. 2012;84(12):5198–5203. doi: 10.1021/ac300864j. [DOI] [PubMed] [Google Scholar]
- Quantum Si Inc Platinum Sequencing Kit V3 Protocol. 2024, Quantum Si Inc,: Branford, CT. [Google Scholar]
- Sigma-Aldrich . Product information: Endoproteinase Lys-C from Lysobacter enzymogenes suitable for Protein Sequencing; Sigma-Aldrich Co. LLC: St. Louis, 2014. [Google Scholar]
- Glatter T., Ludwig C., Ahrné E.. et al. Large-scale quantitative assessment of different in-solution protein digestion protocols reveals superior cleavage efficiency of tandem Lys-C/trypsin proteolysis over trypsin digestion. J. Proteome Res. 2012;11(11):5145–5156. doi: 10.1021/pr300273g. [DOI] [PubMed] [Google Scholar]
- Quantum Si Inc . Unlock Proteome Depth with the Platinum Pro Instrument and Sequencing Kit V4. https://www.quantum-si.com/wp-content/uploads/Q423-DS-Sequencing-Kit-V4-090825.pdf.
- Brown S. M., Mayer-Bacon C., Freeland S.. Xeno Amino Acids: A Look into Biochemistry as We Do Not Know It. Life. 2023;13(12):2281. doi: 10.3390/life13122281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brown S. M., Allgair E., Krystufek R.. Mapping the Edges of Mass Spectral Prediction: Evaluation of Machine Learning EIMS Prediction for Xeno Amino Acids. Anal. Chem. 2025;97(19):10282–10288. doi: 10.1021/acs.analchem.5c00286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldman E. R., Sugiharto V. A., Shriver-Lake L. C.. et al. A single domain antibody-based Luminex assay for the detection of SARS-CoV-2 in clinical samples. Front. Immunol. 2024;15:1446095. doi: 10.3389/fimmu.2024.1446095. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Quantum Si Inc . Platinum Next-Generation Protein Sequencing Advanced Data Analysis Technical Note; Quantum Si Inc. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.





