Skip to main content
Nature Communications logoLink to Nature Communications
. 2026 Feb 28;17:3320. doi: 10.1038/s41467-026-70061-7

Data storage and retrieval with unnatural proteins expressed via E. coli

Yin Zhou 1,2,3,4,5,#, Cheuk Chi A Ng 1,2,3,4,5,#, Chengxi Liu 1,2,3,4,5, Wai Man Tam 6, Francis C M Lau 6, Zhong-Ping Yao 1,2,3,4,5,✉
PMCID: PMC13066418  PMID: 41764153

Abstract

Data storage using proteins offers high capacity and stability, enabling utilization of protein techniques for data storage and retrieval. However, expressing unnatural proteins with random sequences for data storage and sequencing them for accurate data retrieval remain challenging. In this study, by encoding digital data into amino acid sequences and incorporating them into collagen-like protein templates, we achieve successful expression of the proteins via E. coli for data storage; the data-bearing proteins containing selective amino acids and arginine intervals can be sequenced through tryptic digestion followed by LC-MS/MS analysis to achieve complete data recovery, even for protein mixtures encoding multiple datasets. We further demonstrate much higher stability of the data-bearing protein than DNA, and random access and cryptographic data protection using affinity-tagged proteins. This work establishes a robust framework for protein-based data storage, opening up avenues for data storage and retrieval, protein engineering and chemistry, synthetic biology, proteomics, and beyond.

Subject terms: Proteins, Protein design, Proteomics, Chemical biology, Expression systems


Data storage using proteins offers high capacity and stability, however, expressing unnatural proteins with random sequences often fails. Here the authors encode digital data into amino acid sequences based on collagen-like protein templates to allow stable data storage and retrieval.

Introduction

We are in the era of big data, with data generated at an exponentially increasing rate1. Methods that can store data with high capacity and long duration are highly desirable. Molecular data storage, as a promising solution, is actively explored. In those methods, digital data are typically stored by using sequences of polymers such as peptides2,3, deoxyribonucleic acid (DNA)4–6, polyamides7 and dendrimers8, or by using combinations of small molecules9–11. Among these molecules, peptides can be comprised of 20 canonical amino acids and also many non-canonical amino acids12, meaning that peptides can achieve a high storage capacity, as well as having the capability to optimize the stability and chemical properties. By using peptides composed of D-amino acids, proteolysis by conventional L-proteases could be avoided, enhancing the stability13. Archeology14–16 and recent studies3 have also shown that peptides and proteins could achieve much more stability than DNA. Compared to peptides, proteins have longer sequences of amino acids and thus are more efficient for data storage. Preparation of long sequences of amino acids by protein expression is typically more convenient than by chemical synthesis of peptides. Data storage using proteins can thus bring new advantages, and new possibilities, e.g., utilization of the well-developed protein techniques to perform specific functions for data storage and retrieval. However, there is no successful experimental work reported yet.

Proteins can be readily expressed by biological systems such as E. coli, yeast, insect cells and mammalian cells, and cell-free systems. It has been reported that 4256 E. coli proteins could be expressed and purified at once17, and even de novo designed proteins could be expressed in this way18. This will significantly reduce the cost of storing and duplicating data using amino acid sequences. More importantly, based on the basic procedure of de novo design19, i.e., backbone sampling, sequence optimization, functional site design and scoring, proteins with designed functions, e.g., self-assembling biomaterials20,21, inhibitors for targeted therapeutics22,23 and enzymes24, can be generated. Deep-learning algorithms such as RFdiffusion25 can generate functional protein designs based on simple molecular specifications. In addition, deep-learning protein structure prediction algorithms such as AlphaFold26 and DeepTMHMM27 can be used to assist protein design by confirming the protein structures and avoiding the undesired regions. Nevertheless, most of the engineered proteins reported so far are variations of existing sequences or have strict backbone sequences, which is not compatible with the requirement of data storage, where the required protein designs should allow for highly variable amino acid sequences significantly different from existing protein sequences to accommodate the random nature of bit sequences of actual data. While it is relatively straightforward to design and express engineered proteins with definite sequence requirements, the design and expression of such unnatural proteins with random sequences are much more challenging.

In the past decades, protein sequencing has been developed along with the development of proteomics, where proteins are typically digested, followed by sequencing of the resulting peptides using tandem mass spectrometry (MS/MS)28, allowing rapid sequencing of thousands of proteins in very low concentrations29. This can be utilized to sequence the data-bearing proteins for data retrieval. However, while relatively low coverages are sufficient to identify proteins in conventional proteomics by searching against a database, full sequence coverages and accurate de novo sequencing are required to retrieve all data stored in proteins correctly, which is also challenging. In addition, how to utilize protein techniques to perform special functions for data storage and retrieval, e.g., selective data retrieval and cryptography, remains exploration.

For data storage and retrieval with proteins, as shown in Fig. 1, raw data is first converted to binary digits. By assigning amino acids as specific sequences of digital bits, the binary digits are translated into amino acid sequences, which are incorporated into protein sequences with pre-designed templates. The data-bearing proteins are then expressed by cell-based systems or cell-free systems for data storage. For data retrieval, the protein sequences are read out with sequencing techniques, converted back to binary digits, and then the raw data is recovered. In this study, proteins with the pre-designed templates were expressed in E. coli and then affinity-purified. These data-bearing proteins were then stored in the form of lyophilized powder. To retrieve the data, these proteins were digested and analyzed by liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS), with the acquired MS/MS spectra processed by our developed software for assignment of amino acid sequences. The feasibility of data storage using proteins was demonstrated by overcoming the challenges of the expression and sequencing of de novo designed data-bearing proteins, together with the innovative utilization of protein chemistry for special data storage and retrieval functions.

Fig. 1. The process of data storage and retrieval with proteins.

Fig. 1

In data storing process, the raw data is encoded into protein sequences with encoding scheme and pre-designed templates, and the data-bearing proteins are expressed for data storage. In data retrieving process, proteins are sequenced, and the acquired sequences are decoded to obtain the original data.

Results

Direct fusion of encoding peptides

In the first protein template, i.e., Template I (Supplementary Fig. 1 and Supplementary Table 1), binary digit sequences were encoded into peptide segments with a length of 14 amino acids. Similar to our previous design2, each peptide segment carried an address code and its N- and C-termini were fixed as phenylalanine (F) and arginine (R), respectively, to facilitate de novo sequencing by MS/MS. These peptide segments were concatenated to form the protein sequences, which could be digested to the peptide segments by using trypsin that could cleave specifically at the C-terminus of R.

By employing Template I, we encoded the English texts of “The Hong Kong Polytechnic University” and its motto (Fig. 2a), a total of 690 bits, into one protein PolyU-T1a (Supplementary Data 1). Eight data-bearing amino acids, i.e., tyrosine (Y), threonine (T), glutamic acid (E), valine (V), alanine (A), serine (S), leucine (L) and phenylalanine (F), were selected among the 20 canonical amino acids following the selection rules we proposed in the peptide-based method2, with 3 bits of the digital information represented by one amino acid (Supplementary Table 1b). However, after the expression, the intact protein of PolyU-T1a (theoretical mass: 40,824 Da) could not be detected by SDS-PAGE (Fig. 3), and no peptide fragments of PolyU-T1a were identified by LC-MS/MS after trypsin digestion (Supplementary Fig. 2a), indicating that PolyU-T1a was not successfully expressed.

Fig. 2. Overview of storing the text file into and retrieving data from data-bearing proteins, with PolyU-T1b as an example.

Fig. 2

a The English text message to be stored. b The mapping of bit sequences to amino acids is used for encoding. c The data-bearing region sequence for storing the text message. d The LC chromatogram of the resulting tryptic peptides after digestion of the data-bearing proteins. e A typical MS/MS spectrum obtained for determining the peptide sequence with the home-made software.

Fig. 3. SDS-PAGE for analysis of PolyU-T1a, PolyU-T1b, PolyU-T2a, and PolyU-T2b expressed via E. coli.

Fig. 3

Proteins were repeatedly expressed and purified using Ni spin columns three times. Source data are provided as a Source Data file.

Considering the positive correlation between the expression level and protein solubility30, hydrophobic data-bearing amino acids V, A, L, F in Template I were substituted by more hydrophilic amino acids asparagine (N), aspartic acid (D), glycine (G) and glutamine (Q) (Supplementary Table 1b), and protein PolyU-T1b (Supplementary Data 1) was thus constructed and expressed to store the same text information. Still no intact protein (theoretical mass: 40,130 Da) could be observed in the SDS-PAGE (Fig. 3). However, LC-MS/MS analysis of the tryptic digestion products after the purification revealed significant signals of tryptic peptides and amino acid sequences (see Fig. 2d for the total ion chromatogram and Fig. 2e for a MS/MS spectrum for the sequence assignment). Analysis of all the acquired data showed that most tryptic peptides from PolyU-T1b could be detected, with 90% of the data-bearing amino acids correctly recovered for PolyU-T1b (Supplementary Table 2). These results suggested that the designed amino acid sequences of PolyU-T1b could be expressed with E. coli, but the expression level was low and the whole sequence might be subjected to proteolytic degradation during expression such that the intact protein was undetectable by SDS-PAGE. Compared to PolyU-T1a, the use of more hydrophilic amino acids in PolyU-T1b improved the protein expression. However, for the Template I used for the design of these two proteins, data-bearing peptide segments were simply joined together to form proteins, with the protein structures and stabilities totally dependent on the raw data. Such totally unnatural proteins might be incompatible with the E. coli system and be readily digested by the housekeeping enzymes in the system. Most likely, sequences that are designed with this template may become disordered proteins with low expression level and unsatisfactory stability.

Use of a collagen-like template

Inspired by the paleoproteomics studies, we emulated the sequence pattern of collagen, which was the abundant protein type found in fossils from millions of years ago31,32, in another template, Template II (Supplementary Fig. 1). This template followed the pattern of typical collagen-like (GXX)n sequences33 (where X represents an amino acid), which might prevent data-bearing proteins from enzymatic degradation. The data-bearing amino acids (Supplementary Table 1b) were selected among the 20 canonical amino acids according to the amino acid composition of collagen34,35. Binary digit sequences were encoded into peptide segments with a length of 3n amino acids (Supplementary Table 1a), where n is a natural number. The N-terminus of each peptide segment was fixed as F and the C-terminus was fixed as R, and G was fixed on certain positions (Supplementary Table 1a). Amino acids other than the data-bearing amino acids, i.e., V and D, were used to indicate the order of peptide segments. These peptide segments were concatenated to form the amino acid sequences in the data-bearing region. The data-bearing region was then fused with the registration domain (V domain) of the Scl2 collagen-like protein from S. pyogenes to facilitate the protein expression35,36. PolyU-T2a (Supplementary Data 1) was constructed using Template II. By applying the same conventional expression and purification methods used for PolyU-T1a and PolyU-T1b, PolyU-T2a (theoretical mass: 56,657 Da) was successfully expressed, and the yields were significantly higher than those of PolyU-T1a and PolyU-T1b (Fig. 3). The mass of PolyU-T2a was confirmed by denaturing MS analysis (Fig. 4a). The protein samples could be further purified by using affinity chromatography with gradient elution (Supplementary Fig. 3) or other purification methods such as size-exclusion chromatography. The bicinchoninic acid (BCA) assay of the purified protein indicated that the yields of PolyU-T2a could reach over 10 mg/L via E. coli expression. It should be noted that a very high purity was not required for successful data retrieval, and a lower purity (Fig. 3) obtained by the simple purification method using Ni spin column or agarose beads was sufficient. After sequencing by LC-MS/MS (Supplementary Fig. 2c), the correct rate of data-bearing amino acids retrieved from PolyU-T2a was 93% (Supplementary Table 3). The native MS analysis of PolyU-T2a and comparison of the spectral results with the spectrum obtained by denaturing MS (Fig. 4) indicated that PolyU-T2a was predominantly very unfolded under the native conditions, which could facilitate the protein sequencing, while the stability might be impacted. To investigate the stability of the Template II data-bearing protein and compare with that of DNA data storage, a representative protein PolyU-T2a and synthesized DNA sequences (Supplementary Table 4) encoding the same information using the Wukong algorithm37 were incubated under the same conditions and tested for stability. The protein and DNA were quantified using label-free quantification (Supplementary Fig. 4) and quantitative polymerase chain reaction (qPCR) (Supplementary Table 5), respectively. The results (Supplementary Table 6) demonstrated that the stability of data-bearing protein PolyU-T2a was much higher than that of the corresponding data-bearing DNA in high temperature conditions and highly acidic conditions. In DI water, ~75% PolyU-T2a remained after 14 days in 70 °C, while the quantity of DNA dropped to <0.1% only after 1 day at 70 °C. Although dried DNA was more stable than in DI water, its quantity dropped to <3% after 14 days in 70 °C, while there was still ~95% PolyU-T2a remaining. In strongly acidic conditions (pH 1), ~87% PolyU-T2a remained after 1 day, while the quantity of DNA dropped to <0.1%. In all these treatments, the correct rates of data-bearing amino acids retrieved from PolyU-T2a were around 90% (Supplementary Tables 7-9), enabling full recovery of the stored data with the application of error-correction scheme. Overall, by employing Template II, we improved the production of data-bearing proteins, which could be used as a potential scheme for data that needs large amounts of copies and long-term storage.

Fig. 4. The MS spectra of PolyU-T2a.

Fig. 4

Spectra obtained with (a) denaturing MS and (b) native MS.

Storage and retrieval of a larger amount of data

As Template II provided a higher protein yield while preserving the integrity of the expressed protein, we attempted to apply this template with some optimizations to encode a larger amount of data. We started the test with a single protein PolyU-T2b (Supplementary Data 1), which encoded the same text as PolyU-T2a but with different amino acid compositions (Supplementary Table 1b) to avoid isomeric amino acid pairs, and a fixed position of the 2-amino-acid long address (Supplementary Table 1a). PolyU-T2b (theoretical mass: 55,776 Da) was successfully expressed (Fig. 3), and the correct rate of data-bearing amino acids sequenced from it was 95% (Supplementary Table 10). After that, more proteins were designed to store a larger amount of data with further modified Template II, where for each peptide, in addition to a 2-amino-acid-long internal address, a 1-amino-acid protein ID, unique in a protein mixture, was included to identify the precursor protein (Supplementary Table 1a), and the total length of each peptide was changed to 18 amino acids. An error-correction scheme was designed such that at most 4 incorrect or missed sequences out of 30 (13%) could be corrected. We have stored 6 famous quotes (Table 1) into proteins (Supplementary Data 1) which were expressed via E. Coli. A protein mixture M8, containing 8 proteins D1_P, D1_Q, D1_A, D2_L, D2_V, D3_S, D4_D and D4_T, storing the quotes D1-4 with a total of 3,136 bits, was analyzed using our workflow of data retrieval from His-tag labeled proteins. 225 out of 240 peptide sequences (94%) were sequenced correctly (Supplementary Data 2), with no individual protein containing more than 3 incorrect sequences. This was within the designed error-correction capability, and therefore all four quotes could be correctly retrieved. This result demonstrated the robustness of Template II, which allowed many such proteins with vastly different sequences to achieve high expression yield, such that storing and retrieving a larger dataset with multiple proteins would be feasible.

Table 1.

The details of the datasets D1-D6 and the proteins storing these datasets

Dataset Contenta Size (bits) No. of proteins required Encoded protein Mass (Da) Protein ID C-terminal tag No. of data-bearing peptides
D1 “Both in thought and in feeling, even though time be real, to realise the unimportance of time is the gate of wisdom” - Bertrand Russell 1088 3 D1_P 64,946 P V5 30
D1_Q 65,538 Q 30
D1_A 63,060 A 30
D2 “Was vernünftig ist, das ist wirklich; und was wirklich ist, das ist vernünftig” - G. W. F. Hegel 792 2 D2_L 65,200 L VSV-G 30
D2_V 64,266 V 30
D3 “人法地, 地法天, 天法道, 道法自然” - 老子 448 1b D3_S 64,417 S HA 30
D3_T 64,865 T 30
D4 “ὁ δὲ ἀνεξέταστος βίος οὐ βιωτὸς ἀνθρώπῳ” - Σωκράτης 808 2 D4_D 65,255 D Flag 30
D4_T 64,145 T 30
D5 “cogito, ergo sum” - René Descartes 288 1 D5_Dc 62,165 D c-Myc 30
D6 Hello World! 96 1 D6_Sc 35,512 S Strep-tag-II 14

aThe quotes D2 to D5 are in German, Chinese, Greek and Latin, respectively, and their English translations are “What is reasonable is real; and what is real is reasonable—G. W. F. Hegel”, “Man follows the earth, the earth follows the sky, the sky follows the Tao, the Tao follows nature—Lao Zi”, “But an unexamined life is not livable to man—Socrates”, and “I think, therefore I am - René Descartes”, respectively.

bThe content was encoded repeatedly in two proteins with different ID to avoid addressing conflict when mixed with other proteins.

cThe protein does not contain an N-terminal His tag.

Random access and cryptography

After establishing the method of data storage using proteins, we have designed functionalized proteins, such that using established protein chemistry techniques would allow these proteins to perform data storage with special functions. The first such function is random access38. Without specially designed proteins, in order to access a selected part of data from the entire dataset, the entire dataset from all data-bearing proteins must be retrieved first, then select the data of interest afterwards. Using proteins with specific C-terminal affinity tags, the proteins containing the data of interest could be selectively purified and have the data retrieved, saving many efforts. We utilized the proteins encoding the quotes, where a unique C-terminal affinity tag was added to all proteins encoding the same quote (Table 1), to provide a pathway for selective purification. A mixture M6, containing 6 proteins D1_P, D1_Q, D1_A, D3_T, D5_D and D6_S, storing the quotes D1, D3, D5 and D6, with a total of 1,920 bits, were analyzed by affinity purification using the magnetic beads with the corresponding antibodies, i.e., anti-V5 for D1 (Supplementary Table 11), anti-HA for D3 (Supplementary Table 12), anti-c-Myc for D5 (Supplementary Table 13), and anti-Strep-tag-II for D6 (Supplementary Table 14). All the proteins with the corresponding C-terminal tag were selectively purified from the mixture. After sequencing, it was found that at most 4 incorrect data-bearing peptides were found in each protein, which was within the error-correction capability. Therefore, each individual quote could be retrieved selectively and correctly from the proteins encoding that quote and other quotes.

The second specific function is cryptography. Protein data storage is inherently suitable for this purpose, as it is hard to edit proteins after production, and information integrity can thus be protected. In addition, proteins cannot be replicated in a similar fashion as DNA and each MS/MS analysis would consume some proteins, therefore if a potential attacker did not know the correct affinity tag on the proteins storing the secret message, they only have finite chances to guess the correct antibody and retrieve the data correctly39. To demonstrate protein cryptography, the ‘secret messages’ were denoted by the 288-bit quote D5 and the 96-bit quote D6, where the proteins encoding them (D5_D and D6_S, respectively) only contained the ‘secret key’ C-terminal tags but not the ‘universal’ His-tag (Supplementary Data 1). Four His-tag-containing decoy proteins D1_P, D1_Q, D1_A and D3_T, encoding the 1,536-bit ‘decoy messages’ D1 and D3, were included in the mixture M6. Using the previous data retrieval workflow for the His-tag labeled proteins, the peptides from proteins D1_P, D1_Q, D1_A and D3_T were sequenced with ≥90% correct rate which were within the designed error-correction capability, while the peptides from proteins D5_D and D6_S were only sequenced with 29% and 63% correct rate respectively, which exceed the designed error-correction capability (Supplementary Data 3). Since the latter proteins did not contain a His-tag, only a very small amount was non-specifically adsorbed onto the beads and co-eluted with the specifically adsorbed His-tag labeled proteins, leading to a high error rate in sequencing. Therefore, if the attacker did not know the ‘secret key’, only the ‘decoy messages’ but not the ‘secret messages’, could be correctly decoded. Correct antibodies (‘keys’) were needed to retrieve the ‘secret messages’ correctly. When the cell lysate was treated with the correct ‘key’, anti-c-Myc antibody, the protein D5_D could be purified. After digestion and LC-MS/MS analysis, 93% of the sequences from this protein were correctly obtained (Supplementary Table 13), allowing correct retrieval of the ‘secret message’ D5 after the error correction. When the cell lysate was treated with another ‘key’, anti-Strep-tag-II antibody, another protein D6_S could be purified. Following the same workflow, another ‘secret message’ D6 could be correctly retrieved (Supplementary Table 14). These results successfully demonstrated the feasibility of protein cryptography.

Discussion

In this study, Template I was simply data-bearing peptide segments jointed together to form the proteins, which allowed relatively high data capacity but without the consideration of stability and survival of proteins. Depending on the original data files, the corresponding sequences of amino acids may have huge variations, which lead to various structures and properties of the resulting proteins, many of which may not be successfully expressed. For example, proteins with transmembrane structures are usually highly hydrophobic, which can largely reduce the soluble expression level. By letting the sequence pattern be commensurate with that of collagen, fusing protein with V domain and optimizing the expression conditions, the yields of Template II proteins were significantly increased, reaching over 10 mg/L. The Template II data-bearing protein PolyU-T2a was very stable in high-temperature conditions and highly acidic conditions. Even compared to DNA (7 days at 70 °C40) or peptides (3.5 days at 70 °C41) with additional protection, PolyU-T2a showed readability after 70 °C treatment for a much longer time (14 days), demonstrating the potential of Template II proteins for long-term data storage. While the stability of a Template II protein has been established, a study with more data-bearing proteins, potentially with different structures, and a larger variety of incubating conditions, would provide a more comprehensive analysis of the stability of the data-bearing proteins. Furthermore, we stored 7 datasets into 13 data-bearing proteins with Template II, demonstrating the feasibility of larger dataset storage and retrieval, random access, and cryptography.

The feasibility of data storage and retrieval using unnatural proteins has been demonstrated in this study, but there is still much progress to be made. The sizes of data-bearing proteins in this study were in the range of 35 to 65 kDa. Theoretically, a larger protein would be more efficient in storing information. However, a protein that is too large would be difficult to express, and therefore a balance of expression yield and storage efficiency must be considered. To explore the limit of protein length, we have successfully expressed a data-bearing protein larger than 100 kDa via E. coli (Supplementary Fig. 5). The theoretical data capacity can also be increased by increasing sequence diversity. For example, if we use 16 different amino acids, we could encode log216=4 bits per amino acid. As the average molecular mass of canonical amino acids is ∼110 Da, a 100 kDa data-bearing protein can theoretically store up to 4 × (100,000/110) = 3636 bits of data, However, due to the requirement of storing addressing bits and error-correction codes, the actual data capacity would always be less than the theoretical value, and designs that reduce this overhead could bring the actual value closer to the theoretical one. Practically, for example, PolyU-T2a stored 690 bits of data, and for each nanoLC-MS/MS analysis, ~1.33 ng of the data-bearing protein was injected, so the readable data storage density was 690 / (1.33 × 10−9) = 5.18 × 1011 bits/g, which was about 30 times that of peptide-based data storage in our previous work2. This data density is still lower than that of DNA-based methods which can apply PCR for amplification before sequencing, allowing the detection of much fewer number of molecules. Theoretically, the data storage density of protein-based data storage using 16 canonical amino acids could reach about 6 times that of DNA data storage, considering that DNA data storage typically uses 4 nucleotides, which encode 2 bits per nucleotide and have an average molecular mass of ~330 Da. With the further development of protein sequencing techniques in the future, we believe that the readable data density of protein-based data storage could be much more competitive.

The cost for each amino acid in this study was ~US$0.4, which was comparable to commercial DNA synthesis and much lower than that of peptide synthesis ( ~ US$3 per amino acid). Although the cost of storing a large amount of data is still expensive compared to conventional data storage methods, the cost of reproducing the same data is low, and hopefully the cost could be further reduced with the development of synthetic techniques. For example, by employing high-throughput cell-free protein microarrays which can simultaneously express thousands of proteins by cell-free expression systems42, and robotic-assisted expression and purification systems43, the throughput of storing data could be increased with the cost further reduced. For data retrieval, we would expect that with further method improvement and possibly novel protein sequencing techniques in the future, the cost of sequencing could be further reduced, similar to the drastic reduction of cost of sequencing the human genome in the past 20 years. The storage and retrieval speeds of our method are still low, which is still one of the main challenges for molecular data storage. Storing data that would not be retrieved for a long time would mitigate the issue of low speed as well. In particular, the development of single-molecule protein sequencing techniques44, e.g., nanopore45–47 and real-time dynamic single-molecule protein sequencing48, would allow the data capacity, density and accessibility to be significantly improved, by sequencing entire proteins at the single-molecule level with a resolution that could identify post-translational modifications49, as well as designing protein templates with reduced addressing and error-correction overheads. With the development of protein expression and sequencing techniques, expressing and analyzing tens of thousands of proteins in short time may become possible in the foreseeable future, and data storage using proteins may become practically usable. In general, data storage using proteins was one of the promising approaches to storing huge amounts of data.

This is the first study to successfully store data into and retrieve data from de novo unnatural proteins. Using a collagen-like sequence pattern, we successfully tackled the challenges of expressing data-bearing proteins with highly variable sequences while ensuring their integrity and high stability, and demonstrated successful data retrieval by LC-MS/MS sequencing with high accuracy, overcoming the obstacles towards protein-based data storage. By storing and retrieving different data requiring multiple proteins, the robustness of our method was established, on which more data-bearing proteins could be built. Furthermore, random access and cryptography with data-bearing proteins were successfully demonstrated as well. Our method couples data storage with protein science, synthetic biology, and proteomics, and can create additional possibilities for these important fields. The features of protein make it a good carrier of information, particularly for long-term storage of big data, which may have applications in data archival and data storage in long-term space missions. Furthermore, as proteins are relatively biocompatible, it is possible to store digital data in living organisms with premise of careful consideration of potential biosafety risks and ethical concerns, which may facilitate our future life.

Methods

Encoding scheme and sequence recovery

The structures of the sequences used for encoding, sequencing, and decoding for data-bearing proteins are shown in Supplementary Table 1. We mapped 3-bit symbols (000 to 111, or symbols from 0 to 7 corresponding to the three bits) to the 8 amino acids to obtain the peptide sequences.

The sequence recovery was achieved by converting the raw data files outputted directly from the mass spectrometer using MSConvert (Version 3.0.25172-38fa11a9) to MS1 and MS2 formats, followed by analyzing with the updated version of the in-house software developed in our previous study2. The maximum error for each peak was set to 25 ppm (in line with the experimental parameters), and the masses were corrected to at least five decimal places. In the updated version of the software, the function allowing recovery of varied lengths of sequences in one run was developed. Additional filters were employed to sift out sequence candidates that did not conform to sequence patterns and address codes of pre-designed templates (Supplementary Table 1). This allowed distinguishing isomeric amino acid combinations such as GA and Q (both are 128.05858 Da) in case of incomplete fragmentation, as G only occurred at specific locations in Template II. Sequence candidates were scored according to the length of consecutive amino acids, the number of amino acids retrieved, the number of occurrences of same sequence, mass match error, and ion intensity.

Error-correction code design

We used low-density parity-check (LDPC) code as the error-correction code. The parity-check matrix structure, the encoding and decoding methods of the LDPC code, adopted from our previous work50, are described in details in Supplementary Information. In the construction of the LDPC code, we used the progressive edge-growth (PEG) algorithm to generate the code such that the code had a girth of 6 (without cycles of length of 4) based on the dual-diagonal structure for encoding purpose.

We assumed that 13% of peptide sequences (i.e., 4 sequences out of 30 sequences for a protein) cannot be retrieved due to the missing or incorrect estimation of the amino acids during storage and retrieval. For the generated code, we erased any combination of 4 sequences in the decoding for testing to ensure that total 30 sequences could be recovered for all combinations.

The design of a LDPC code can be adjusted easily. For example, it can be designed to generate more parity bits and hence to provide more protection to the data bits. Such designs can be applied to more complex mixtures where the error rate might be much higher than 10%.

Moreover, we considered the case that multiple proteins were used to store the data bits in one or multiple datasets, and the sequences have fixed length and pattern (GXX)n. We also assumed that 13% of peptide sequences cannot be retrieved in each protein. Suppose that there were multiple proteins for one dataset, each stores a subset of data bits. Then we needed to design a common LDPC code to ensure that 13% errors can be recovered for each protein when different subsets of data bits were used as input. When there were multiple datasets, we considered all input sets of data bits in the design of the LDPC code. Since both encoding and decoding were independent for each protein, a decoding failure in one of the proteins would not affect the decoding of other proteins, making this suitable for random access.

As an example, the 1088-data-bits dataset D1 (encoded in 3 proteins D1_P, D1-Q and D1_A), first, the total 1088 data bits were divided into 3 portions, except the last portion with 188 data bits, each portion had 450 data bits. Then the data bits in each portion were encoded into 630 bits with 180 parity bits by a LDPC encoder to generate 30 sequences of length 10 (Supplementary Table 15). Next, the 450 data bits (i.e., b1, b2, …, b450) were filled into Columns 1–3 and 9–10, while the 180 parity bits (i.e., P1, P2, …, P180) were arranged in Columns 4 and 8. Also, the address symbol pairs (i.e., {A1,1, A1,2}, {A2,1, A2,2}, …, {A30,1, A30,2}) and the protein ID symbols (i.e., p1, p2, …, p30) were located in Columns 5, 6 and 7, respectively.

If the number of bits in the last portion and some datasets is smaller than 450, then zero bits were appended as the input information bits for the LDPC code to generate the parity bits. If the number of data bits was larger than 150, then all bits, including zero bits, were arranged in the sequences to form the protein with the same length as other portions. Otherwise, only the data bits and the parity bits were stored in the protein (protein D6-S for dataset D6). Thus, the protein had a shorter length.

Finally, according to the sequence pattern (GXX)n, six amino acids G were inserted in the sequences; N-terminal F and C-terminal R were added to the sequences. Hence, the total length of the sequence was 18.

Protein expression

The corresponding coding sequences for proteins were generated and codon-optimized. These customized gene inserts (Supplementary Data 4) were cloned into the pET28a(+) vector via the NdeI/XhoI site to generate recombinant plasmids (Sangon Biotech, Shanghai, China). These plasmids were transformed into E. coli BL21(DE3) strain individually.

For each recombinant protein, production was performed by inoculating a single colony into 10 mL LB media containing kanamycin and cultured at 37 °C, 250 rpm overnight. The overnight culture was used to inoculate 200 mL culture that was incubated at 37 °C, 250 rpm until the cells were grown to an OD600 nm of approximately 0.6. Protein expression was induced with the addition of 1 mM isopropyl β-D-1- thiogalactopyranoside (IPTG) (BBI, Shanghai, China), and the cells were grown at 20 °C, 250 rpm for a further 18 h. Cells were harvested by centrifugation (4000 × g, 4 °C, 20 min). Cell pellets were washed with ice-cold PBS and then lysed with CelLytic B Plus Kit (Sigma, St. Louis, USA). N-terminal His-tag labeled proteins were purified from the lysate using Ni spin columns (NEB, Ipswich, USA) or Ni-NTA agarose beads (NEB, Ipswich, USA) or Cytiva ÄKTA Pure 25 M1 equipped with a HisTrap™ HP His tag protein purification column (Cytiva, Marlborough, USA), and C-terminal affinity tag labeled proteins were purified from the lysate using antibody-coated magnetic beads (V5: Sigma, St. Louis, USA; c-Myc and HA: Beyotime, Shanghai, China; Strep-tag-II: BBI, Shanghai, China). Ultrafiltration tubes (Millipore, Billerica, USA) were used to perform buffer exchange and to concentrate the purified proteins. BCA assay was performed using Pierce BCA Protein Assay Kit (Thermo Fisher Scientific, Waltham, USA) and Varioskan LUX Multimode Microplate Reader (Thermo Fisher Scientific, Waltham, USA) to measure the protein concentration. Proteins were dried using refrigerated CentriVap vacuum concentrator (Labconco, Kansas City, USA) at 4 °C overnight.

Trypsin digestion

The frozen protein pellets were dissolved with 50 mM ammonium bicarbonate. Trypsin (Promega, Madison, USA) was added to the protein samples (protein to trypsin ratio = 50:1), and the solutions were incubated overnight at 37 °C.

LC-MS/MS analysis

The tryptic peptide mixtures were separated using a Waters Acquity UPLC system equipped with a C18 column (AdvanceBio Peptide Map, 2.1 × 150 mm, 2.7 µm particle size, 120 Å pore size, Agilent, Santa Clara, USA) held at 55 °C. Mobile phase A was 0.2% formic acid in water and B was 0.2% formic acid in acetonitrile. The system flow rate was set to 0.3 mL/min. The gradient stayed at 1% B at 0 to 2 min, changed linearly from 1% B to 9% B at 2 to 8 min, from 9% B to 23% B at 9 to 53 min, from 23% B to 80% B at 53 to 65 min, and remained at 80% B from 65 to 68 min.

MS/MS analysis was performed on an Orbitrap Fusion Lumos mass spectrometer (Thermo Fisher Scientific, Waltham, USA) in positive ion mode. The spray voltage was set at 3.6 kV. The ion transfer tube temperature and vaporizer temperature were set at 280 °C. In each cycle (3 s), an MS scan was performed with the resolution of 30,000, AGC target of 400,000, and scan range of 450–2000 Da. Ions were selected for MS/MS using the quadrupole, with charge states from +2 to +4, a dynamic exclusion window of 4 s, isolation window of 1.6 Da. The fragmentation method was high-energy collision dissociation (HCD) with stepped collision energy: 23, 28, and 33. MS/MS spectra were obtained with the resolution of 15,000, AGC target of 50,000, and scan range of 140 to 3500 Da. One technical repeat was performed for the digests of PolyU-Txx, and three technical repeats were performed for the digests of D1-6.

Denaturing and native MS analysis

Denaturing MS analysis was performed using LC-MS on a Waters Acquity UPLC system (mobile phase solvent A: 0.1% formic acid in water; solvent B: 0.1% formic acid in acetonitrile) coupled with a Synapt G2-Si ion mobility quadrupole time-of-flight mass spectrometer (Waters, Milford, USA). The protein samples ( ~ 5 pmol in 50 mM ammonium bicarbonate each) were trapped on a Waters C4 column and washed with solvent A for 3 min and then eluted by a 10 min gradient (20% B to 80% B at 0 to 5 min, 80% B from 5 to 8 min, and 80% B to 20% B at 8–10 min) through a Waters C4 analytical column at 100 μL/min. Eluted proteins were measured in positive ion mode, with the settings as follows: capillary voltage, 3 kV; cone voltage, 30 V; ion source temperature, 150 °C; desolvation temperature, 350 °C; desolvation gas flow, 600 L/h; nebulizer gas flow, 6.5 bar; scan range, 100–5000 Da. Data were analyzed with the MassLynx v4.2 software.

Native MS analysis was performed using nano-electrospray ionization (nano-ESI) on the same mass spectrometer with the NanoLock ion source. The protein samples (2-10 μM in 200 mM ammonium acetate) were loaded into the home-made gold-coated glass nano-ESI emitters for the analysis, with the settings as follows: capillary voltage, 1.5 kV; cone voltage, 30 V; ion source temperature, 30 °C; nano-flow gas pressure, 0.3 bar; scan range, 500–7000 Da. Data were analyzed with the MassLynx v4.2 software.

Stability tests

The concentrations of protein and DNA were adjusted to 0.2 µg/µL. Aliquots of 10 µL (2 µg) samples were stored in 1.5 mL Eppendorf tubes and dried using a refrigerated CentriVap vacuum concentrator (Labconco, Kansas City, USA) at 4 °C overnight. The dried samples were stored in −80 °C as a control group. Protein and DNA samples were either kept dried or dissolved in 50 µL DI water, and stored at 70 °C for 14 days with a constant temperature metal bath (Jingxin, Shanghai, China). For stability under the acidic condition, dried samples were dissolved in pH 1 (HCl) and pH 7 solutions, and stored at 20 °C for 1 day. Samples were vacuum dried and analyzed by LC-MS/MS or qPCR. Three technical repeats were performed.

Protein quantification

Label-free protein quantification with data-dependent acquisition (DDA) was performed to measure the amount of protein remaining in the stability study. Three technical repeats were performed unless otherwise stated. After the trypsin digestion as described above, 0.5 µL of each tryptic peptide mixture sample was injected, and separated using an ACQUITY nanoLC system (Waters, Milford, USA) equipped with an Aurora Elite C18 column (1.7 µm, 25 cm×75 um, IonOpticks, Collingwood, Australia). Mobile phase A was 0.1% formic acid in water and B was 0.1% formic acid in acetonitrile. The system flow rate was set to 0.3 µL/min. The gradient stayed at 1% B at 0–2 min, changed linearly from 1% B to 9% B at 2–8 min, from 9% B to 23% B at 9–53 min, from 23% B to 80% B at 53–65 min, and remained at 80% B from 65–68 min.

MS/MS analysis was performed on an Orbitrap Fusion Lumos Mass Spectrometer (Thermo Fisher Scientific, Waltham, USA) equipped with a nanoelectrospray ionization source in positive ionization mode. The spray voltage was set at 2.2 kV. The ion transfer tube temperature was set at 300 °C. In each cycle (3 s), an MS scan was performed with the resolution of 30,000, AGC target of 400,000, and scan range of 350 to 1800 Da. Ions were selected for MS/MS using the quadrupole, with charge states from +2 to +4, a dynamic exclusion window of 10 s, isolation window of 1.6 Da. The fragmentation method was HCD with stepped collision energy: 23, 28, and 33. MS/MS spectra were obtained with the resolution of 15,000, AGC target of 50,000, and scan range defined by the first m/z of 110 Da.

Peak areas of data-bearing peptides were analyzed by PEAKS Studio 13 (Bioinformatics Solutions Inc., Waterloo, Canada) and summed up to quantify the protein.

Quantitative polymerase chain reaction (qPCR)

DNA and primers were purchased from GenScript Biotech Corporation (Nanjing, China). The qPCR of dilute data-bearing DNA was performed with the primer pair (5’-ACACGACGCTCTTCCGATCT-3’ and 5’-AGACGTGTGCTCTTCCGATCT-3’) in a 20 µL PCR volume using PowerTrack™ SYBR Green Master Mix (Thermo Fisher Scientific, Waltham, USA) and Applied Biosystems QuantStudio 5 Real-Time PCR System (Thermo Fisher Scientific, Waltham, USA) with the following thermal profile: (1) 95 °C for 2 min, (2) 95 °C for 15 s, (3) 60 °C for 1 min. The total number of cycles of steps 2–3 was 40.

Statistics and reproducibility

No statistical method was used to predetermine the sample size. No data were excluded from the analyses. The experiments were not randomized. The Investigators were not blinded to allocation during experiments and outcome assessment.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Supplementary Information (826.3KB, pdf)
Peer Review File (915.4KB, pdf)
41467_2026_70061_MOESM3_ESM.pdf (88.7KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (84.8KB, pdf)
Supplementary Data 2 (219.7KB, pdf)
Supplementary Data 3 (169.7KB, pdf)
Supplementary Data 4 (89.8KB, pdf)
Reporting Summary (1.7MB, pdf)

Source data

Source data (5.6MB, zip)

Acknowledgements

We would like to thank Prof. Yanxiang Zhao and Prof. Clarence Chun Ting Wong (PolyU) for their help and useful discussions on this project. This work was supported by National Key Research and Development Program of China (Grant No. 2024YFF0725800, Z.P.Y.), Hong Kong Research Grants Council (Grant Nos. R5013-19F, C5026-24GF, C4002-20WF, C4014-23G, CRS_CUHK405/23 and AoE/M-402/25-N, Z.P.Y.), Faculty of Science (Grant No. 1-WZA2, Z.P.Y.), the University Research Facility in Chemical and Environmental Analysis, and the University Research Facility in Life Sciences of The Hong Kong Polytechnic University.

Author contributions

Y.Z., C.C.A.N., and Z.P.Y. designed the experiments. W.M.T. and F.C.M.L. encoded and decoded the data with the error-correction schemes. Y.Z. and C.C.A.N. performed the LC-MS/MS analysis. C.L. and Y.Z. performed the native and denatured MS and stability analysis. Y.Z., C.C.A.N., C.L., and W.M.T. drafted the manuscript. Z.P.Y. and F.C.M.L. revised the manuscript. Z.P.Y. initiated and coordinated the whole project.

Peer review

Peer review information

Nature Communications thanks Chunhai Fan, and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Data availability

The spectral data generated in this study for LC-MS/MS analysis of the data-bearing proteins have been deposited in the MassIVE database under accession code MSV000098849 [https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?task=7f1745678fc846728d68438e2f742071]. Source data are provided with this paper.

Code availability

The codes of the home-made sequencing program are protected due to patent restrictions, but may be available for academic exchange and collaboration purposes by sending email requests to the corresponding author, with the expected response time of around 1 week.

Competing interests

The authors declare the following competing interests: Y.Z., C.C.A.N., C.L., and Z.P.Y. are inventors for a related patent entitled “Data storage using proteins” (Zhongping Yao, Yin Zhou, Cheuk Chi Ng, Chengxi Liu, PCT patent application No. PCT/CN2023/108347, filed on 20 July 2023; CN patent application No. 202380100287.2, filed on 7 Jan 2026; US Non-Provisional patent application No. 19/501,196, filed on 12 January 2026; EP patent application No. 23945486.1, filed on 16 January 2026). W.M.T. and F.C.M.L. declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yin Zhou, Cheuk Chi A. Ng.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-026-70061-7.

References

  • 1.Wright, A. Worldwide IDC Global DataSphere Forecast, 2025–2029. IDChttps://my.idc.com/getdoc.jsp?containerId=US53363625 (2025).
  • 2.Ng, C. C. A. et al. Data storage using peptide sequences. Nat. Commun.12, 4242 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Rössler, S. L., Grob, N. M., Buchwald, S. L. & Pentelute, B. L. Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis. Science379, 939–945 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Extance, A. How DNA could store all the world’s data. Nature537, 22–24 (2016). [DOI] [PubMed] [Google Scholar]
  • 5.Organick, L. et al. Random access in large-scale DNA data storage. Nat. Biotechnol.36, 242 (2018). [DOI] [PubMed] [Google Scholar]
  • 6.Yaniv, E. & Dina, Z. DNA Fountain enables a robust and efficient storage architecture. Science355, 950–954 (2017). [DOI] [PubMed] [Google Scholar]
  • 7.Roy, R. K. et al. Design and synthesis of digitally encoded polymers that can be decoded and erased. Nat. Commun.6, 7237 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Huang, Z. et al. Binary tree-inspired digital dendrimer. Nat. Commun.10, 1918 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Cafferty, B. J. et al. Storage of information using small organic molecules. ACS Cent. Sci.5, 911–916 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Nagarkar, A. A. et al. Storing and reading information in mixtures of fluorescent molecules. ACS Cent. Sci.7, 1728–1735 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Rosenstein, J. K. et al. Principles of information storage in small-molecule mixtures. IEEE T. Nanobiosci.19, 378–384 (2020). [DOI] [PubMed] [Google Scholar]
  • 12.Zhang, H. et al. Rational incorporation of any unnatural amino acid into proteins by machine learning on existing experimental proofs. Comput. Struct. Biotechnol. J.20, 4930–4941 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Zhang, G. & Zhu, T. F. Mirror-image trypsin digestion and sequencing of D-proteins. Nat. Chem.16, 592–598 (2024). [DOI] [PubMed] [Google Scholar]
  • 14.Hendy, J. Ancient protein analysis in archaeology. Sci. Adv.7, eabb9314 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Service, R. F. Protein power. Science349, 372–373 (2015). [DOI] [PubMed] [Google Scholar]
  • 16.Welker, F. et al. Ancient proteins resolve the evolutionary history of Darwin’s South American ungulates. Nature522, 81 (2015). [DOI] [PubMed] [Google Scholar]
  • 17.Chen, C.-S. et al. A proteome chip approach reveals new DNA damage recognition activities in Escherichia coli. Nat. Methods5, 69–74 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Fisher, M. A., McKinley, K. L., Bradley, L. H., Viola, S. R. & Hecht, M. H. De novo designed proteins from a library of artificial sequences function in Escherichia Coli and enable cell growth. PLoS ONE6, e15364 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pan, X. & Kortemme, T. Recent advances in de novo protein design: principles, methods, and applications. J. Biol. Chem.296, 100558 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Shen, H. et al. De novo design of self-assembling helical protein filaments. Science362, 705–709 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chen, Z. et al. Self-assembling 2D arrays with de novo protein building blocks. J. Am. Chem. Soc.141, 8891–8895 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Chevalier, A. et al. Massively parallel de novo protein design for targeted therapeutics. Nature550, 74–79 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Cao, L. et al. De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science370, 426–431 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol.41, 1099–1106 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature620, 1089–1100 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hallgren, J. et al. DeepTMHMM predicts alpha and beta transmembrane proteins using deep neural networks. Preprint at BioRxivhttps://www.biorxiv.org/content/10.1101/2022.04.08.487609v1 (2022).
  • 28.Hughes, C., Ma, B. & Lajoie, G. A. De novo sequencing methods in proteomics. Methods Mol. Biol.604, 105–121 (2010). [DOI] [PubMed] [Google Scholar]
  • 29.Sun, B., Kovatch, J. R., Badiong, A. & Merbouh, N. Optimization and modeling of quadrupole orbitrap parameters for sensitive analysis toward single-cell proteomics. J. Proteome Res.16, 3711–3721 (2017). [DOI] [PubMed] [Google Scholar]
  • 30.Price, W. N. et al. Large-scale experimental studies show unexpected amino acid effects on protein expression and solubility in vivo in E. coli. Microb. Inform. Exp.1, 1–20 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Torres, J. M., Borja, C., Gibert, L., Ribot, F. & Olivares, E. G. Twentieth-century paleoproteomics: lessons from Venta Micena Fossils. Biology11, 1184 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Warinner, C., Korzow Richter, K. & Collins, M. J. Paleoproteomics. Chem. Rev.122, 13401–13446 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Yu, Z., An, B., Ramshaw, J. A. & Brodsky, B. Bacterial collagen-like proteins that form triple-helical structures. J. Struct. Biol.186, 451–461 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Qiu, Y., Zhai, C., Chen, L., Liu, X. & Yeo, J. Current insights on the diverse structures and functions in bacterial collagen-like proteins. ACS Biomater. Sci. Eng.9, 3778–3795 (2021). [DOI] [PubMed] [Google Scholar]
  • 35.Xu, C., Yu, Z., Inouye, M., Brodsky, B. & Mirochnitchenko, O. Expanding the family of collagen proteins: recombinant bacterial collagens of varying composition form triple-helices of similar stability. Biomacromolecules11, 348–356 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Peng, Y. Y. et al. Towards scalable production of a collagen-like protein from Streptococcus pyogenes for biomedical applications. Microb. Cell Fact.11, 1–8 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Huang, X. et al. Storage-D: A user-friendly platform that enables practical and personalized DNA data storage. Imeta3, e168 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Organick, L. et al. Random access in large-scale DNA data storage. Nat. Biotechnol.36, 242–248 (2018). [DOI] [PubMed] [Google Scholar]
  • 39.Tirmazi, H. & Tran, T. P. An introduction to protein cryptography. Preprint at https://eprint.iacr.org/2025/089 (2025).
  • 40.Chen, W. D. et al. Combining data longevity with high storage capacity—layer-by-layer DNA encapsulated in magnetic nanoparticles. Adv. Funct. Mater.29, 1901672 (2019). [Google Scholar]
  • 41.Luo, B. et al. High-capacity information storage using peptide-encapsulated hydrogels for long-term data preservation. Commun. Mater.6, 183 (2025). [Google Scholar]
  • 42.He, M., Stoevesandt, O. & Taussig, M. J. In situ synthesis of protein arrays. Curr. Opin. Biotechnol.19, 4 (2008). [DOI] [PubMed] [Google Scholar]
  • 43.Wiesler, S. C. & Weinzierl, R. O. Robotic high-throughput purification of affinity-tagged recombinant proteins. Methods Mol. Biol.1286, 97–107 (015). [DOI] [PubMed]
  • 44.Restrepo-Pérez, L., Joo, C. & Dekker, C. Paving the way to single-molecule protein sequencing. Nat. Nanotechnol.13, 786–796 (2018). [DOI] [PubMed] [Google Scholar]
  • 45.Brinkerhoff, H., Kang, A. S. W., Liu, J., Aksimentiev, A. & Dekker, C. Multiple rereads of single proteins at single–amino acid resolution using nanopores. Science374, 1509–1513 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Afshar Bakshloo, M. et al. NanoporE-based Protein Identification. J. Am. Chem. Soc.144, 2716–2725 (2022). [DOI] [PubMed] [Google Scholar]
  • 47.Jiang, J. et al. Protein nanopore reveals the renin–angiotensin system crosstalk with single-amino-acid resolution. Nat. Chem.15, 578–586 (2023). [DOI] [PubMed] [Google Scholar]
  • 48.Reed, B. D. et al. Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device. Science378, 186–192 (2022). [DOI] [PubMed] [Google Scholar]
  • 49.Niu, H. et al. Direct mapping of tyrosine sulfation states in native peptides by nanopore. Nat. Chem. Biol.21, 716–726 (2025). [DOI] [PubMed] [Google Scholar]
  • 50.YYao, Z., Ng, C. C., Lau, C. M. & Tam, W. M. Data storage using peptides. US Patent No. 11,315,023 (2022).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (826.3KB, pdf)
Peer Review File (915.4KB, pdf)
41467_2026_70061_MOESM3_ESM.pdf (88.7KB, pdf)

Description of Additional Supplementary Files

Supplementary Data 1 (84.8KB, pdf)
Supplementary Data 2 (219.7KB, pdf)
Supplementary Data 3 (169.7KB, pdf)
Supplementary Data 4 (89.8KB, pdf)
Reporting Summary (1.7MB, pdf)
Source data (5.6MB, zip)

Data Availability Statement

The spectral data generated in this study for LC-MS/MS analysis of the data-bearing proteins have been deposited in the MassIVE database under accession code MSV000098849 [https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?task=7f1745678fc846728d68438e2f742071]. Source data are provided with this paper.

The codes of the home-made sequencing program are protected due to patent restrictions, but may be available for academic exchange and collaboration purposes by sending email requests to the corresponding author, with the expected response time of around 1 week.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES