Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 May 22.
Published in final edited form as: Chem Rev. 2024 May 1;124(10):6592–6642. doi: 10.1021/acs.chemrev.4c00110

Genetic encoding of phosphorylated amino acids into proteins

Michael C Allen 1, P Andrew Karplus 1, Ryan A Mehl 1, Richard B Cooley 1,*
PMCID: PMC11658404  NIHMSID: NIHMS2031424  PMID: 38691379

Abstract

Reversible phosphorylation is a fundamental mechanism for controlling protein function. Despite the critical roles phosphorylated proteins play in physiology and disease, our ability to study individual phospho-proteoforms has been hindered by a lack of versatile methods to efficiently generate homogeneous proteins with site-specific phosphoamino acids or with functional mimics that are resistant to phosphatases. Genetic code expansion (GCE) is emerging as a transformative approach to tackle this challenge, allowing direct incorporation of phosphoamino acids into proteins during translation in response to amber stop codons. This genetic programming of phospho-protein synthesis eliminates the reliance on kinase-based or chemical semi-synthesis approaches, making it broadly applicable to diverse phospho-proteoforms. In this comprehensive review, we provide a brief introduction to GCE and trace the development of existing GCE technologies for installing phosphoserine, phosphothreonine, phosphotyrosine and their mimics, discussing both their advantages as well as their limitations. While some of the technologies are still early in their development, others are already robust enough to greatly expand the range of biologically relevant questions that can be addressed. We highlight new discoveries enabled by these GCE approaches, provide practical considerations for the application of technologies by non-GCE experts, and also identify avenues ripe for further development.

Keywords: Genetic Code Expansion, phosphoserine, phosphothreonine, phosphotyrosine, non-canonical amino acids, amber suppression

Graphical Abstract

graphic file with name nihms-2031424-f0017.jpg

1. Introduction

1.1. Scope and purpose.

Phosphorylation is the most abundant protein post-translational modification (PTM), serving to regulate protein function across all three kingdoms of life.1 Understanding the structure, function and biological role of specific phospho-proteins is key to understanding how life works, the molecular basis of disease and how to create novel therapeutics. Progress in this area has been slow because it has been difficult to efficiently produce proteins with one or more specific sites of phosphorylation, and especially to make such proteins in sufficient quantity and homogeneity for downstream in vitro and in vivo characterization. Recently, “genetic code expansion” (GCE) technologies have emerged as a promising generalizable solution to this problem that can revolutionize our ability to study phosphorylated protein function. This is because GCE – in the ideal case – enables in living cells the efficient and quantitative co-translational installation of phosphorylated amino acids and their mimics into proteins at genetically programmable sites.2,3 In 2022, NIH acknowledged GCE’s potential impact on biological research by funding the first national center dedicated to optimizing and enhancing access to GCE technologies for the broader research community.4

As with many nascent technologies, the development of GCE tools for studying phosphoproteins has been and continues to be an incremental process, with many innovative researchers “standing on each other’s shoulders” as they have explored various approaches and developed increasingly effective and robust tools. This has resulted in a rather complex literature that is not easy to parse. Our purpose in this review is to help those desiring to use GCE in their phosphoprotein research, as they seek to sort through the literature and find the tools most suitable for use in their studies. We do this by: first providing a high-level overview of protein phosphorylation and approaches to make phosphorylated proteins; second, giving a brief introduction to GCE technologies; third, in the bulk of the review, comprehensively reviewing the developments, strengths, weaknesses and uses of existing phosphoserine (pSer), phosphothreonine (pThr) and phosphotyrosine (pTyr) GCE tools, with discussions on strategies for improving them; fourth, touching on opportunities for encoding “less studied” phosphoamino acids; and fifth, providing some closing thoughts.

1.2. Protein phosphorylation as a regulatable protein post-translational modification.

Protein phosphorylation occurs when a phosphoryl group is transferred from ATP to a protein by kinases. The localized -2 charge and tetrahedral geometry of the phosphoryl group provides a chemical signature that is distinct from the canonical 20 amino acids and other PTMs, and its presence leads to alterations in electrostatic interactions within the modified protein and with molecules in its environment. As a consequence, phosphorylated proteins undergo structural, conformational and dynamics changes that can be highly specific and can modify protein functions, such as enzyme activity and its interactions with other proteins and biomolecular components including membranes, organelles, DNA, RNA, co-factors and metals (Fig. 1). Over 500 kinases are encoded in the human proteome, each with a different substrate specificity and distinct mechanisms of activation and regulation.5-8 Of these, ~90 are tyrosine kinases and ~430 are Ser/Thr kinases.9 They generally bind their substrates through an induced fit/conformational selection mechanism based on the recognition of specific amino acid sequences flanking the intended phosphorylation site.10,11 Because kinase substrates need to be both accessible and structurally conform to the kinase active site binding groove, the large majority of phosphorylation sites lie within intrinsically disordered regions (IDRs) of proteins.12-14

Figure 1. Diverse functional consequences of protein phosphorylation.

Figure 1.

Simple schematic of a generic protein (purple) phosphorylation/dephosphorylation cycle (central blue box) is surrounded by some examples of physical properties and biological activities that can be altered by changes in the protein’s phosphorylation state. Here and in other figures in this review, the “P” in the yellow circle denotes a phosphoryl group attached to a protein or other molecule. Cartoon images here and in Figures 4, 5, 7, 8, 12 and 14 were created with BioRender.com.

Proper regulation of phospho-protein function requires they also be specifically dephosphorylated by phosphatases at select sites with spatial and temporal control (Fig. 1). While there are roughly as many protein tyrosine phosphatases as tyrosine kinases (~100), there are much fewer Ser/Thr phosphatase catalytic domains than kinases (~30).9 This led to the hypothesis that Ser/Thr phosphatases are non-specific for their substrates and act as so called “housekeeping” proteins.15 However, more recently it has become clear that phosphatases are highly specific for their substrates, with substrate identification mediated through a plethora of accessory or adaptor proteins that bind phosphatases and recruit phosphorylated substrates, placing their phospho-sites within proximity of the phosphatase active site.16 The importance of properly regulating kinase and phosphatase function is underscored by how often imbalances in their actions are the molecular basis of cancers and other diseases.17

1.2.1. Types of phosphorylation sites and their abundance in biology.

Nine residue types – Ser (1), Thr (3), Tyr (5), His (7), Asp (10), Glu (12), Cys (14), Lys (16) and Arg (18) (Fig. 2) – are known to undergo biologically-relevant phosphorylation, and here we designate their phosphorylated forms by adding a “p” in front of the three-letter code (e.g. pSer). The first three of these are by far the best studied, and in fact, the largest active database, PhosphoSitePlus currently only tracks pSer (2), pThr (4) and pTyr (6) (with annotations for 294,716 phospho-sites across 20,208 proteins).18 Among these, the majority are on Ser residues (60%), followed by Thr (25%) and then Tyr (15%).18 The phospho-ester bonds formed by the phosphorylation of Ser, Thr and Tyr are stable at neutral and acidic pH, and therefore are readily detected by standard mass spectrometry proteomics work-flows.19 Though not annotated yet on PhosphoSitePlus, His, Lys and Arg can be phosphorylated via phosphoramidate bonds (8, 9, 17, and 19, Fig. 2). These phosphoamino acids are much less commonly observed, but this need not mean they are at a much lower abundance in the cell; it could instead be because they are more transient and/or more susceptible to acid hydrolysis during isolation and analysis. Reports vary depending on the methodology and organism of study, but pHis (8, 9), pArg (19) and pLys (17) might be at least as abundant, if not more abundant than pTyr (6) in mammalian cells.19 Lastly, pCys (15), pAsp (11) and pGlu (13) (Fig. 2) have also been detected in notable quantities in mammalian cells,19 and like phosphoramidates, the phosphoanhydride bonds of pAsp/pGlu and the phosphothioester bond of pCys are much more labile than phospho-ester bonds.20 Whereas the biological roles of pGlu are poorly defined, pAsp and pCys have known roles as catalytic intermediates of several enzymes, and pAsp also plays a well-established regulatory role in bacterial two-component histidine kinase signaling systems [reviewed previously in 21,22].

Figure 2. Biologically-relevant phosphorylated amino acids.

Figure 2.

Each box shows the side chain of one of the nine canonical amino acid residue types known to be phosphorylated in proteins, providing both its unmodified (left-hand image) and its phosphorylated (right-hand image) form or forms. Ser, Thr and Tyr (blue boxes) are the three most studied phosphorylated residue types.

1.2.2. Vastness of the unexplored “Dark Proteome”.

It has been estimated that as much as 80% of the human proteome is phosphorylated at some point.23 Futhermore, because the majority of these phospho-proteins can be modified at more than one site, there is a vast array of unique proteo-forms that can exist at any given time in a cell. For example, 29 pSer, pThr and pTyr sites have been reported on the essential hub protein 14-3-3ζ,18 creating a possible 537 million (i.e. 229) potential unique phospho-forms of 14-3-3ζ. Even if one assumes only a small fraction of these exist, it is sobering to consider what it would take to define their individual roles in cellular function. Nevertheless, identifying which of these phospho-forms exist in a cell at a given time is a key first step in this process. But even this is a challenge because the common “bottom-up” mass spectrometry (MS) approach to characterize phosphoprotein modifications identifies modified peptide fragments of a protein, rather than detecting whole proteins, leaving open the question of which sets of modifications are simultaneously present in a given protein molecule.24 These challenges in both the detection and functional characterization of phospho-proteins has, in part, led to the phospho-proteome being called the “Dark Proteome”.7

Considering the importance of protein phosphorylation in biology, substantial effort has been invested in understanding how these PTMs alter protein function, yet despite these efforts fewer than 3% of phospho-sites have a reported function.7 Indeed, the in vivo study of phospho-proteins is non-trivial, often requiring a priori knowledge of the kinase activation cascade involved in phosphorylating a protein of interest, at the site(s) of interest, as well as an ability to acutely control these signaling processes (e.g. by using selective inhibitors/activators or modulation of metabolites and environmental factors). Antibodies are commonly used to track the presence or formation of phosphorylated proteins, but reliable phospho-specific antibodies can be resource intensive to generate and often have varying levels of specificity or sensitivity, depending on the target protein. While these systems-level studies of phospho-proteins can be highly informative, future development of therapeutics targeting specific phospho-proteoforms requires a molecular and mechanistic understanding for how specific phosphorylation events alter protein structure, dynamics and protein-protein interactions, yet insufficient progress has been made on this front.

To highlight this deficiency, we turned to the Protein Data Bank (PDB)25 where experimentally resolved protein structures are deposited and we asked how many unique protein structures are known that contain at least one instance of pSer (2), pThr (4) or pTyr (6). As of October 1, 2023, the PDB contained 35,184 unique eukaryotic protein entities larger than 10 kDa in size. Of these unique proteins, only 177 proteins contain a pSer (0.5%), 140 a pThr (0.4%) and 72 a pTyr (0.2%) (Fig. 3A). In other words, only ~1% of the 35,000 eukaryotic proteins of known structures have been determined in a phosphorylated form, despite the expectation that the majority of them are phosphorylated at some point as part of their biological function. It is also worth pointing out that of the phospho-protein structures available, about half are kinases (Fig. 3B), even though kinases represent only ~1-2% of the total human proteome. We suspect that this high proportion of kinases among phosphoprotein structures is because their phosphorylated forms are readily produced either through auto-phosphorylation or phosphorylation by endogenous kinases in the recombinant protein expression host (e.g. insect cells). In fact, for some of these kinases it takes a special effort to produce the non-phosphorylated form.26-28 This underscores that a difficulty in producing authentic, specifically and quantitatively phosphorylated proteins is a major hurdle that must be overcome in order to shed light on the vast dark phospho-proteome.

Figure 3. Structurally-known phosphorylated proteins containing pSer, pThr and pTyr.

Figure 3.

(A) Histogram showing the number of eukaryotic protein structures (in PDB as of October 1, 2023) which are ≥10 kDa and non-redundant at a 95% sequence identity threshold, and how many of these have at least one pSer, pThr or pTyr residue, respectively. The size cutoff was used to ensure that nearly all selections would be proteins made via recombinant protein expression rather than chemical synthesis. (B) Pie charts showing the PANTHER29 classification of the functions of the structurally-known pSer-, pThr- and pTyr-containing phosphoproteins, respectively. Among the fifteen groups represented (numbered 1-15 and differing colors), percentages are given for the overall most common 4 (numbered 1 - 4). For group 1 (“Protein modifying enzymes”, PC00260) which is dominant, a breakdown is given of the four enzyme types present (lower pie charts with key). Remaining groups are: 2 – metabolite interconversion enzymes (PC00262); 3 – transmembrane signal receptors (PC00197); 4 – transcriptional regulators (PC00264); 5 – transporters (PC00227); 6 – scaffold/adaptor protein (PC00226); 7 – membrane traffic proteins (PC00150); 8 – chaperones (PC00072); 9 – cell adhesion molecules (PC00069), 10 – protein-binding activity modulators (PC00095); 11 – defense/immunity proteins (PC00090); 12 – RNA metabolism proteins (PC00031), 13 – cytoskeletal proteins (PC00085); 14 – translational proteins (PC00263); and 15 – chromatin/chromatin-binding proteins (PC00077). For proteins in group 3 (transmembrane signal receptors, PC00197), we note that most of the pSer/pThr and all the pTyr proteins are in fact kinases but are classified as transmembrane signal receptors based on PATHER priority assignment rules. Therefore, the true percentage of phospho-proteins in the PDB that are kinases is higher than even represented here.

1.3. Methods to generate site-specifically phosphorylated proteins.

Methods developed to make site-specifically phosphorylated proteins include (i) chemical synthesis of polypeptides, (ii) chemical ligation strategies that covalently link synthetic peptides and/or recombinant proteins together, (iii) enzymatic phosphorylation (Fig. 4), and (iv) GCE, which involves the translational encoding of phosphorylated amino acids into recombinant proteins. Before delving deeper into the translational encoding techniques – the primary focus of this review – we will in the following sections provide brief descriptions of the first three approaches, along with their respective advantages and disadvantages. For a more comprehensive treatment of all three of these alternative strategies, we recommend an excellent, in-depth review by Bilbrough et al.30

Figure 4. Three non-GCE approaches to make site-specifically phosphorylated proteins.

Figure 4.

(A) In solid-phase peptide synthesis, an initial amino acid (AA in blue circle) with protecting groups (PG in grey triangle) is covalently linked to a solid resin support (grey box). Further residues are added in a stepwise fashion with suitable removal and addition of protecting groups (including on side-chains but not shown here for simplicity), and a phosphorylated residue can be added at any position with various methods used to prevent hydrolysis of the phosphoryl group. The final product of up to about 70 residues long (i.e. n+m+2) is released from the resin and deprotected. (B) In protein chemical ligation, a polypeptide that has a special C-terminal modification (X) reacts with a second polypeptide that has a special N-terminal modification (Y) to form an intermediate linkage (Z) that rearranges to yield a conventional amide bond joining the two polypeptides. Not shown are byproducts produced, which vary depending on the method used. While the phosphorylated residue shown here is part of the second peptide, it can be on either. (C) In enzymatic phosphorylation, an active form of a specifically selected protein kinase (which may itself require phosphorylation to be active) catalyzes the transfer of phosphoryl group from ATP to a protein of interest (POI), producing ADP as a byproduct.

1.3.1. Chemical synthesis of phospho-peptides.

Several methods have emerged for generating peptides with genuine phosphorylated amino acids and their mimics. Solid-phase peptide synthesis (Fig. 4A) offers direct incorporation of αN-protected phosphoamino acids, albeit with protection groups on the phosphate that can be removed after completion of peptide chain elongation.31 Alternatively, free (unprotected) hydroxyl groups of Ser, Thr or Tyr can be chemically phosphorylated post-synthesis.32 Similarly, pCys (15) peptides can be produced by converting Cys (14) or pSer (2) to dehydroalanine, allowing the introduction of a thiophosphate group.33 Strategies also exist to create pHis-containing peptides, but these may result in mixtures of different phosphorylation states (e.g. 1-pHis (8) and 3-pHis (9), Fig. 2).34-36 One particularly important advantage of peptide synthesis is it allows for installation of stable (i.e. non-hydrolyzable) phosphorylated amino acid analogs, including C-P bond-containing phosphonates and fluoro-phosphonates and photo-protected variants.37 Also, the synthesis of peptides with up to three phosphoamino acids is relatively straightforward, and even though including more sites becomes increasingly challenging, peptides with as many as nine sites of phosphorylation have been reported in reasonable yield and purity.38 Despite its strengths, peptide synthesis is generally limited to peptides of approximately 50-70 residues, making it unsuitable to directly generate phospho-proteins.

1.3.2. Chemical ligation for combining synthetic peptides and/or recombinant proteins.

Chemical ligation strategies (Fig. 4B) offer a solution to the size constraints associated with phospho-peptide synthesis methods, enabling the production of larger proteins. “Native chemical ligation” (NCL) is the common approach to selectively couple two peptides.39 Typically, one peptide harbors a C-terminal thioester, while the other features an N-terminal cysteine, which can be combined together to form a native amide bond that links the two peptides.30 NCL can be executed in aqueous solutions at near neutral pH levels and is compatible with denaturing agents like guanidinium, enhancing peptide solubility. This approach represents a powerful means to access more extensive polypeptides with enhanced chemical diversity compared to peptide synthesis alone, although it still imposes limitations on the final product’s size.

To overcome the size limitations of NCL, “expressed protein ligation” (EPL) offers a method for chemically fusing synthetic phospho-peptides to folded proteins that were expressed and purified from a recombinant host.40,41 In EPL, the target protein is expressed from a genetic construct in which its C-terminal segment containing the site of phosphorylation has been replaced with an intein protein. After purification of this target protein-intein fusion, the addition of a thiol-containing molecule, such as mercapto-ethanesulfonate, triggers intein-mediated thiolysis, resulting in the production of the target protein with a free C-terminal thioester group. Using the NCL chemistry introduced above, this C-terminal thioester group can be linked to a phospho-peptide produced via chemical synthesis (and containing an N-terminal Cys residue) that corresponds to the protein's native, phosphorylated C-terminal segment. This generates a full-length target protein with a C-terminal phosphorylation site. Conversely, target proteins with an N-terminal phosphorylated segment can be formed by an analogous approach in which a target protein without its native N-terminal segment is expressed and processed via proteolytic cleavage to yield a Cys residue as its N-terminus, and then ligating this with a chemically synthesized version of the N-terminal phospho-peptide segment which has a thioester at its C-terminus.41

EPL’s versatility stems from the range of chemistries that can be incorporated into peptides during chemical synthesis, along with the efficiency of modern recombinant protein expression and purification techniques. These allow EPL to provide access to a fairly wide variety of biologically relevant phosphorylated proteins, especially because many natural phosphorylation sites are located in flexible regions at a protein's N- or C-terminus.42 One elegant study that highlights the utility of EPL is Chu et al.43, who sought to understand how AKT kinase was regulated via phosphorylation of its C-terminal tail. Since all three phospho-regulatory sites in question were within 20 residues of the C-terminus of AKT, they could use EPL to ligate singly, doubly and triply phosphorylated chemically synthesized peptides to AKT. This revealed that including either (i) pSer473 or (ii) pSer477 and pThr479 led to AKT activation, but interestingly they did so via very different molecular mechanisms. In terms of limitations, generating phospho-proteins with phosphorylation sites located more than ~50 amino acids away from either terminus remains challenging with EPL techniques, and not all potential target proteins exhibit sufficient stability for expression and purification without their N- or C-termini. Also, EPL chemistries result in the presence of a vestigial Cys residue at the linkage site that may deviate from the native protein sequence.41 Finally, since EPL cannot be performed inside cells, in vivo investigations of the functions of target proteins made using EPL necessitate the use of protein transfection and/or cell-penetrating delivery strategies.

1.3.3. Enzymatic phosphorylation.

Historically, for phosphoproteins which are not predominantly found in their phosphorylated form in nature, relevant phosphorylated forms have been synthesized “naturally” by exposing a target protein to a specific activated kinase, magnesium ions (Mg2+), and adenosine triphosphate (ATP) (Fig. 4C). These reactions can be performed (i) using purified proteins and kinases in vitro, (ii) by co-expressing the target protein alongside the kinase or (iii) by expressing the target protein in its native eukaryotic host, followed by purification of the phosphorylated proteins. While these approaches can be highly effective in favorable cases, they are still of limited general applicability because the kinases responsible for the vast majority of human phosphoylation sites are unknown7 – though ongoing efforts to characterize the substrate specificities of broad selections of kinases8 are extending the utility of these approaches. A further limitation is that even for many cases in which the relevant kinase is known, obtaining an active form of the kinase for in cell or in vitro use is not trivial, because its activation involves regulatory mechanisms such as auto-phosphorylation, phosphorylation by another kinase or interaction with activator proteins.27,44-46 Recent evidence has also shown that specificity of kinases can change depending activation status47, and that the broad substrate specificity8 of many kinases can result in multiple off-target sites being inadvertently phosphorylated on a given substrate protein. Consequently, generating site-specifically phosphorylated proteins using kinases, without modifications to other sites in the protein (e.g. Ser/Thr to Ala to prevent off-target phosphorylation) is challenging at best, and often unfeasible.

1.3.4. Mimicking protein phosphorylation with aspartate and glutamate substitutions.

When genuine phosphorylated protein production is unattainable, researchers often resort to “phosphomimicking” point mutations. In the most common approach, Ser (1), Thr (3), or Tyr (5) sites are mutated to Asp (10) or Glu (12) residues. The carboxylates of Asp/Glu, each with a negative charge, are clearly the only canonical amino acids that can be considered similar to phosphorylated Ser, Thr and Tyr, and so it makes sense they have been the go-to residues for phosphomimetic mutations. These phosphomimicking mutations have been widely employed because they are technically straight-forward to make and the carboxylate side chains are stable in a cellular environment. Furthermore, introducing the point mutations at functionally relevant phosphorylation sites often results in functional alterations relative to the unmodified protein, and in many cases these alterations at least partially resemble the effects of phosphorylation.

Nevertheless, this simplified approach overlooks the reality that the planar geometry and single negative charge of the carboxylates do not provide a very close mimic for the native −2 charge and the tetrahedral structure of phosphoryl groups found in pSer (2), pThr (4), and pTyr (6).48 Furthermore, even though the lengths of the Asp/Glu side chains are similar to the lengths of the pSer and pThr side chains, they are both much shorter than the side chain of pTyr, making them even worse mimics for pTyr than for pSer and pThr. For these reasons, it should be expected that, in general, Asp and Glu will not faithfully mimic details of the regulation of protein function and signaling provided by phosphorylation. Consistent with this expectation, the more systems for which researchers have directly compared the functions of Asp- or Glu-phosphomimicking protein variants versus authentically phosphorylated forms, the more evidence there is of discrepancies (e.g. 49-57). Thus, it is becoming very evident that conclusions drawn solely from studies using Asp/Glu phosphomimetic protein variants should be interpreted with caution and be reevaluated when possible using authentic phosphorylated protein and/or more faithful phosphomimetics.

2. Genetic Code Expansion as tool to generate phosphorylated proteins.

2.1. Origins of GCE.

In 2001, Schultz and colleagues demonstrated an archaeal tyrosine amino-acyl tRNA synthetase (Tyr-RS)/tRNATyr pair from the Methanocaldococcus jannaschii (Mj) was orthogonal to Escherichia coli translational machinery and it could be evolved to encode non-canonical amino acids (ncAAs, i.e. amino acids that are not among the standard 20) into proteins during translation in response to the TAG (i.e. UAG in mRNA) amber stop codon in living cells.58 This established the groundbreaking framework by which living organisms can be programmed to encode 21 amino acids into proteins, whereby the ncAA is structurally, chemically and functionally unique from the 20 canonical amino acids found in nature.59 Since this pivotal discovery in 2001, researchers have successfully encoded several hundred ncAAs into proteins, working with various RS/tRNA pairs that are compatible with not only E. coli but also eukaryotic and mammalian translational machineries.3,60 In addition to the Mj Tyr-RS/tRNATyr pair, pairs now commonly used to expand the genetic code have originated from the pyrrolysyl-(Pyl) RS/tRNAPyl of Methanosarcina barkeri, Methanosarcina mazei, Methanomethylophilus alvus, as well as the tyrosyl-RS/tRNATyr and leucyl-tRNA RS/tRNALeu pairs from E. coli, and the tryptophanyl-RS/tRNATrp pair from S. cerevisiae.61 Recent advances in genomic sequencing and bioinformatic analyses of translational systems have unveiled further pairs, many of which exhibit orthogonality to one another, providing the capacity for multi-ncAA encoding.62-64

Initially termed “unnatural amino acid incorporation” this technique has now transitioned to the more inclusive terminology of “genetic code expansion” (GCE). This shift in nomenclature reflects the broader notion that GCE encompasses encoding of amino acids that are not found in nature as well as amino acids that in fact are naturally occurring, but are not directly built into proteins during translation.65 A prominent group of such naturally occurring ncAAs are the many post-translationally modified forms of canonical amino acids, including phosphoamino acids which are the subject of this review.

2.2. Leveraging GCE to encode phosphorylated amino acids.

In response to the above challenges associated with generating phosphorylated proteins, GCE was quickly recognized as a viable approach to incorporate phosphorylated amino acids into proteins during translation within living cells. A significant milestone in this journey was reached in 2011 when Söll and colleagues achieved the first successful encoding of free phosphoserine amino acid into proteins at UAG amber stop codons in E. coli.66 This accomplishment introduced an entirely novel methodology for producing phosphorylated proteins, sidestepping the needs for any chemical peptide syntheses, ligation chemistries, or kinases. By genetically programming the sites of phosphorylation through the introduction of a TAG codon at the desired location within the target gene, the GCE approach eliminated off-target phosphorylation events. Moreover, the versatility of phospho-GCE allows for phosphorylation at virtually any site on the protein, regardless of its position or proximity to the termini, and it enables the encoding of phosphoamino acid mimics that are resistant to dephosphorylation by phosphatases (discussed further below).

However, the genetic encoding of phosphorylated amino acids is not without its challenges and shortcomings, and extensive efforts have been and are still being directed towards the development and optimization of GCE systems for various phosphorylated amino acids and their mimics. Before we provide an overview of these pSer (2), pThr (4) and pTyr (6) GCE systems (Section 3), in the remainder of this section we first introduce the fundamental mechanics of installing phosphorylated non-canonical amino acids into proteins using GCE (Section 2.3), as well as terms and concepts common to the field (Section 2.4). Meaningful evaluation of these GCE technologies also requires knowledge of methods to assess the phosphorylation status of proteins, and the strengths and weakness of each (Section 2.5).

2.3. General mechanics of phosphoamino acid GCE systems.

2.3.1. Steps of phosphoamino acid encoding.

Production of phosphorylated proteins through GCE involves a series of carefully orchestrated processes which must all function in concert. We find it helpful to think about it as a series of five discrete steps, which we illustrate in Fig. 5, and briefly discuss below.

Figure 5. Key steps in GCE-based phosphoryl-ncAA incorporation into proteins in E. coli.

Figure 5.

Highlighted are both key steps (numbered circles) and processes to be avoided or minimized (arrows with red "x”). Step 1. The phosphoryl-ncAA is made available in the cytosol, either by: (1a) diffusion into cells of an exogenously added ncAA; or (1b) endogenous production via a biosynthetic pathway that is either introduced or results from altering natural metabolic pathways in E. coli. Factors to be minimized are ncAA degradation or dephosphorylation and ncAA toxicity. Step 2. An orthogonal aminoacyl tRNA synthetase (RS) charges its cognate tRNACUA with the ncAA. To be minimized are breakdowns in orthogonality due to either the evolved RS charging its cognate tRNA with canonical amino acids or charging endogenous tRNAs with the ncAA. Step 3. The ncAA-loaded tRNACUA is then delivered to the ribosome by elongation factor Tu (EF-Tu), which in the case of GCE systems for phosphoryl-ncAAs is an engineered form to better transport these negatively charged ncAAs. Step 4. The ribosome binds the charged tRNACUA in response to the UAG codon to incorporate the ncAA into the protein of interest (POI). The charged tRNACUA must outcompete both endogenous tRNAs to avoid canonical amino acid (cAA) misincorporation (i.e. near-cognate suppression), and release factor 1 (RF1), if present, to avoid premature termination that produces truncated peptides. Step 5. Finally, the phosphorylated POI must properly fold and retain functionality, so it can be purified or studied in vivo. To be minimized are both misfolding, often producing insoluble protein aggregates (i.e. inclusion bodies), and dephosphorylation of the POI by phosphatases.

  1. The first step is to ensure that the free phosphoamino acid is available in sufficient concentrations inside the cell to feed the GCE machinery; the needed concentrations of ncAAs typically range from high micromolar to low millimolar levels due to relatively high Km values of the tRNA synthetases.67-70 In most GCE systems, the ncAAs to be incorporated are chemically synthesized and then introduced into the culture media from which they can diffuse into the cell (1a in Fig. 5). This approach, however, does not generally work well for phospho-ncAAs, which because of their high negative charge, poorly traverse the lipid bilayer.71-73 Phosphate transporters may assist in the import of some phosphorylated amino acids from the media, however transporter activities can be unpredictable and contingent on culturing conditions.71,74,75 The most successful ways that this challenge has been addressed is the engineering of cells to biosynthesize high levels of the phospho-ncAA from central metabolites (1b in Fig. 5).

  2. In the second step, an orthogonal amino-acyl tRNA synthetase (RS) amino-acylates its cognate tRNACUA with the free phospho-ncAA (step 2 in Fig. 5). The orthogonality of the RS/tRNACUA pair ensures that the RS does not amino-acylate tRNAs native to the expression host, and that the tRNACUA is not amino-acylated by any endogenous RSs. While orthogonality is not absolute as it depends on the relative concentrations of these components within the host, an initial level of orthogonality is often established by selecting RS/tRNA pairs from evolutionarily distant organisms, such as archaea. If necessary, further engineering is employed to enhance orthogonality, which may include modifications to the variable loop or acceptor stem regions of the tRNACUA to minimize mis-acylation by non-cognate RSs.61,62,76,77 The goal in these cases isn’t necessarily to achieve perfect “absolute” orthogonality, but instead a sufficient level of “relative” or “functional” orthogonality so that when ncAA is present, any propensity of the RS/tRNACUA pair to encode of canonical amino acids is outcompeted and the ncAA it is quantitively incorporated.78

  3. In the third step, the amino-acylated tRNACUA binds to a GTPase called EF-Tu, which both protects the tRNACUA from deacylation and delivers it to the ribosome (step 3 in Fig. 5). While the synthesis of standard proteins and of most ncAA-proteins is well-served by a single EF-Tu, the delivery of tRNAs amino-acylated with phosphorylated amino acids is less efficient using native EF-Tu proteins.66 Hence, specially engineered EF-Tu variants, as further discussed below, have been developed to enhance this process.66,79

  4. In step 4, the amino-acylated tRNACUA pairs its anticodon with the UAG codon of the target gene and the phospho-ncAA is added to the growing nascent peptide chain (step 4 in Fig. 5). It is important to note that when the ribosome encounters the UAG codon, various undesired outcomes are in competition with ncAA-tRNACUA delivery and encoding. In standard expression hosts, Release Factor-1 (RF1) is responsible for terminating translation at UAG codons and thus truncating the expression of the target protein. Moreover, endogenous amino-acylated tRNAs can suppress UAG codons through wobble-base pairing, leading to the unwanted encoding of natural amino acids at the intended phosphorylation site—referred to as near-cognate suppression (NCS, discussed more in section 2.4.2).80 This competition between ncAA encoding, RF1-mediated truncation, and NCS favors intended ncAA encoding in efficient GCE systems that generate high concentrations of ncAA amino-acylated tRNACUA.81-84

  5. In the final step, translation proceeds past the UAG codon, ultimately yielding full-length, properly folded fully phosphorylated protein (step 5 in Fig. 5). This stage faces two distinct challenges: ensuring the proper folding of the phospho-protein and the prevention of dephosphorylation by host cell phosphatases before purification or analysis. With respect to protein folding, biologically relevant phosphorylated proteins are evolved to accomadate stable phosphorylation post-translation, not during translation. Consequently, the folding of the target protein may be compromised, potentially leading to degradation or precipitation in inclusion bodies. Phosphorylation often induces protein-protein interactions in natural systems as well,85 a feature that may not be achievable when phospho-proteins are produced in isolation within a heterologous expression host. In such cases, strategies promoting phospho-protein folding and stability during expression, including the use of solubility tags (e.g., SUMO, MBP, GST), co-expression with folding chaperones, or binding partners, may be helpful.86 With regard to unwanted dephosphorylation of target proteins, E. coli expresses a variety of putative phosphatases.87 Most of these have unconfirmed functions and undefined substrate specificities, but many can act on phosphorylated metabolites, as well as phosphorylated proteins. The tendency of a target protein to be dephosphorylated in E. coli depends on the protein of interest and site of encoding, and is likely higher when the site of phosphorylation resides within a disordered or flexible region of a protein easily accessible to phosphatases. Shorter expression times,84 lower expression temperatures88 or use of non-hydrolyzable mimics (discussed more below) can mitigate unwanted dephosphorylation.49 In mammalian cells, target proteins with pSer encoded into them were purified with serine at the site of encoding, caused by the elaborate network of endogenous protein phosphatases, necessitating the need to encode a non-hydrolyzable mimic of pSer.89

2.4. Evaluating GCE system utility.

Efficiency, fidelity and permissivity are three important properties of GCE systems that define their utility. Although characterizations of GCE systems do not always encompass all three of these qualities, they should all be kept in mind when evaluating and comparing them. In this review, as we delve into the history and development of GCE tools for making biologically relevant proteins with phosphoamino acids and their mimics, we will use these terms with their technical meanings as defined in this section.

2.4.1. Efficiency.

The efficiency of a GCE system is quantified by comparing the amount of ncAA-containing protein produced to that of the wild-type protein with the same expression system (Fig. 6). It is typically reported as a percent of the wild type protein level - GCE systems efficiencies are highly variable but usually range from 10 to 90%. Lower efficiency of GCE systems can often be attributed to the fact that RS/tRNACUA pairs evolved for ncAA encoding exhibit lower catalytic activity compared to natural RS/tRNA pairs, resulting in reduced production of amino-acylated tRNACUA.67-70 Not surprisingly, efficiency tends to decrease as the number of encoded ncAA sites increases.90,91 While RS/tRNACUA pair catalytic activity is often the limiting factor in ncAA-protein production with GCE,68 the encoding of ncAAs can face bottlenecks at other steps in the translational process shown in Fig. 5, including phospho-ncAA bioavailability [step 1]71-73, EF-Tu delivery to the ribosome [step 3]92, and proper protein folding [step 5].93

Figure 6. Conceptual illustration distinguishing efficiency, fidelity and permissivity outcomes in GCE systems.

Figure 6.

Expression efficiency is plotted for expressions in bacterial strains with RF1 present (e.g. BL21) or absent (e.g. B95), and as a function of ncAA added, with ncAA1 being the ncAA for which the RS was selected and ncAA2 and ncAA3 being distinct ncAAs for which the RS, respectively, has low and moderate permissivity. Wild type (WT) expression levels (black bar; no ncAA and no TAG codon) are a positive control defining 100% efficiency. A single black bar is shown, because even if the RF+ and RF- strains yield different amounts of WT protein, each result defines the 100% standard to which other expressions in that strain are normalized. For other experiments, the apparent total level of full-length folded protein (full bar height) is divided into portions roughly assignable as: (i) “Orthogonality Breakdown” (dotted green bar; seen in the absence of ncAA in an RF+ strain); (ii) “Near Cognate Suppression” (striped green bar; typically maximal in the absence of ncAA in an RF− strain); and (iii) “Accurate ncAA Encoding” (solid green bar; requires mass spec or a functionality test to assess). Fidelity is defined as the amount of accurate ncAA encoding (solid green bar) divided by the total full-length folded protein produced (full bar height). For sfGFP test protein expression, the levels of properly folded protein produced are easily estimated by fluorescence measurements, but for real POIs this is more difficult to do without some protein purification.

For measuring GCE efficiency, fluorescent reporter proteins like super-folder GFP94 with a TAG codon (sfGFP-TAG) are commonly used.95 These are used because full-length, fluorescent sfGFP is produced only when the UAG codon is suppressed, allowing for the quantification of ncAA encoding efficiency through a simple measurement of total fluorescence from a culture and normalized by the cell density. Because no protein purification is needed, this approach enables the high-throughput evaluation of numerous conditions and expression parameters for ncAA encoding and is very useful for optimizing GCE systems. sfGFP is an ideal reporter compared to other fluorescent proteins due to its fast folding and chromophore maturation kinetics (t1/2 ~13 min at 37 °C in cells96), allowing for real-time readout of protein production. Nevertheless, important to note is that the efficiency measured for any GCE system is sensitive to the protein being expressed and the position in the protein into which the ncAA is incorporated. Such variation is commonly seen, but not often discussed, and the factors leading to this variation are not generally understood.

2.4.2. Fidelity.

Fidelity measures the proportion of the expressed target protein that contains the desired ncAA (Fig. 6). Various methods can be used to assess fidelity of phospho-ncAA encoding (as discussed in section 2.5), including whole-protein mass spectrometry (MS). While the ideal is 100% fidelity, most methods to assess fidelity are not always effective at revealing low level impurities. Thus, even for GCE systems that appear to generate pure protein by whole-protein MS analysis, the level of fidelity that can be claimed is realistically in the 90 - 95% range. GCE systems are typically optimized for achieving this range of fidelity because lower levels of fidelity lead to protein preparations that have a mixture of protein species, some of which contains the desired ncAA and some of which do not. And this impurity complicates subsequent analyses unless the ncAA-containing form can be purified away from the contaminants.

Mis-encoding in GCE systems most often arises from two main processes and both of these reflect inadequacies of the GCE system. The first process is mis-acylation in which the tRNACUA is inappropriately amino-acylated with canonical amino acids due to a breakdown in orthogonality (step 2 in Fig. 5). This can depend on many factors including the expression host (e.g., E. coli vs. eukaryotic cells) and the relative concentrations of the GCE components compared to endogenous translational elements.76 The second process is near-cognate suppression (NCS), which takes place when endogenous tRNAs outcompete the GCE-derived amino-acylated tRNACUA in suppressing the UAG codons (step 3 in Fig. 5).80 This mis-recognition can occur because the anticodons of some endogenous tRNAs (e.g. tRNAUUCGlu, tRNACUCGlu, tRNACCATrp, tRNAGUATyr, tRNAUUULys, and tRNACUGGln) can wobble base-pair with the UAG codon,81 leading to the erroneous incorporation of a canonical amino acid at the intended site of ncAA inclusion. NCS is most prevalent in low-efficiency GCE systems when paired with RF1-deficient cell lines, because low concentrations of the ncAA-acylated tRNACUA do not effectively outcompete near-cognate suppressing tRNAs.81,97 Conversely, in RF1-containing expression hosts, translational termination at the UAG codon typically outcompetes NCS, resulting in a mixture of full-length ncAA protein and truncated protein (Fig. 6).98

2.4.3. Permissivity.

Permissivity99, also referred to as poly-specificity100, is the capacity of an RS to encode multiple ncAAs with similar structural characteristics. Thus, a highly permissive RS is one that can efficiently charge its cognate tRNACUA with a wide variety of ncAAs in addition to the one for which it was selected (Fig. 6). In some cases, an RS can be more efficient at encoding an alternate ncAA than it is at encoding the ncAA for which it was selected.101,102 Important is to recognize that high permissivity is perfectly compatible with high fidelity (the ability to discriminate against and not encode any of the canonical 20 amino acid types).

Permissivity can be quite an advantageous property, because discovering that an existing RS efficiently encodes a novel ncAA of interest saves the time and effort it takes to develop or select a new RS variant for that ncAA. As one example, the well-used RSs for installing p-azido-Phe into proteins were an RS originally selected to install p-cyano-Phe.103,104 Also, assessing permissivity is rather easy, with the typical approach being to culture cells in a 96-well plate with various wells supplemented with a different ncAA to be tested and using sfGFP-TAG expression levels to quantify the efficiency of incorporation of each of the tested ncAAs. As we delve into further below, this approach has proven valuable in constructing phospho-GCE systems, because some of the RSs that encode authentic phospho-PTMs can also efficiently encode similarly structured mimics as permissive substrates (e.g. a p-carboxy-methyl-phenylalanine RS can also install phosphotyrosine72).

2.5. Methods for evaluating protein phosphorylation status.

Given the potential for mis-encoding and the susceptibility of phosphorylated proteins to hydrolyze, it is very important to evaluate phosphorylation status of phospho-proteins produced via GCE at various stages of expression, purification and characterization. In the following subsections, we briefly review four commonly employed methods for this purpose (Fig. 7).

Figure 7. Conceptual illustration of four methods to evaluate protein phosphorylation status.

Figure 7.

Proteins 1 – 4 (in center box) are four possible phosphorylation states of a protein with two potential phosphorylation sites. Expected phosphorylation status outcomes for the four protein forms are shown for analysis by: mass spectrometry (arrow A; shown as a composite result overlaying all 4 spectra); Phos-tag gel electrophoresis (arrow B; uninformative SDS gel result shown for comparison); Western blot using residue-specific phospho-sensitive (e.g. a pSer-, pThr- or pTyr-specific) antibodies (arrow C); and protein gel staining using phospho-specific dyes (arrow D).

2.5.1. Mass spectrometry.

MS provides a powerful strategy to characterize phosphorylated proteins made with GCE. These methods take advantage of the characteristic +80 Da mass increase associated with each added phosphate group, allowing for clear discrimination between phosphorylated and non-phosphorylated proteins (Fig. 7A). When using whole-protein MS methods to evaluate fidelity, important to consider is that low level impurities, i.e. ~5-10% of the total protein, can easily go unnoticed because the collective impurities often constitute a variety of off-target species resulting from near-cognate suppression, mis-acylation and, in the case of phospho-GCE, dephosphorylation. In addition, different variants of the same protein may ionize with different efficiencies, especially when evaluating phosphorylated proteins, and therefore without the use of internal standards relative MS peak heights may not reflect the true proportion of each species.105,106 Nevertheless, careful assessment of the non-phosphorylated species of the phospho-protein produced via GCE can provide clues to their origin (i.e. either via dephosphorylation or mis-encoding), which can guide protein expression optimization strategies. For example, a non-phosphorylated species with a mass consistent with Gln incorporation at the UAG codon likely arose from NCS, whereas a mass consistent with Ser likely arose from dephosphorylation (when encoding pSer (2)).

The identification of specific sites of phosphorylation is possible with MS, and requires fragmentation of the target protein into peptides using “bottom-up” (e.g., enzymatic digestion) or “top-down” (e.g., electron-capture or electron-transfer dissociation) approaches.107 Either way, the resulting peptides are fragmented within the MS instrument and, if successful, the sequence information can localize the exact site(s) of phosphorylation. Similar to whole protein MS, the ionization and detection efficiencies of phospho-peptides may differ from their non-phosphorylated counterparts, again making it challenging to precisely quantify the extent of target protein phosphorylation.105,106 In some cases, unless careful testing and optimization of MS methods are carried out, phospho-peptides may even go undetected.97 Constraints related to the size of the target protein, access to specialized equipment, and the need for expertise, along with the relatively low throughput nature of MS data collection and analysis, can limit its routine application. But given its ability to give residue-specific sequence information, when both whole protein and fragmentation MS approaches are feasible, MS remains the “gold standard” and most informative strategy for assessing protein phosphorylation status.

2.5.2. Phos-tag gel electrophoresis.

Phos-tag gel electrophoresis is a versatile technique using a metal chelator – called “Phos-tag” – that forms transient complexes with phosphorylated amino acid residues.108-110 When a Phos-tag acrylamide derivative is co-polymerized into SDS-PAGE gels (with added Mn2+ or Zn2+), the migration of phosphorylated protein forms is slowed and proteins are separated based on their phosphorylation status as well as size (Fig. 7B). Importantly, the extent to which migration is slowed depends on (i) the number of phosphorylation sites, (ii) the specific context of these sites within the protein and (iii) the type of phosphorylation (e.g., pSer vs pThr).111 Phos-tag is known to interact with pSer (2), pThr (4), and pTyr (6)110, as well as pHis (8,9) and pAsp (11)111, and fortuitously, it can even be used to differentiate between proteins with authentic phosphoamino acids and their non-hydrolyzable counterparts.49,89

This method is therefore invaluable for the separation of individual phospho-species across multiple samples or within complex mixtures, and in doing so allows for the semi-quantitative assessment of the relative amounts of each species present. However, if a target protein exists in a mixture of phosphorylation states, several bands with attenuated electrophoretic mobility will be resolved in the Phos-tag gel and it is not possible to assign the identity of each band to a specific phospho-protein variant without standards. For example, when evaluating a protein in which two phosphorylated residues are encoded into it (e.g. protein 4 in Fig. 7), running additional lanes in the Phos-tag gel with the wild-type and the two singly phosphorylated species alongside it (proteins 1, 2, and 3 in Fig. 7B) allows the user to confirm the identity of any sub-stoichiometrically phosphorylated protein species. A key advantage of Phos-tag electrophoresis compared to MS is in estimating the relative amounts of phosphorylated and non-phosphorylated protein forms present in a sample, since all non-phosphorylated species regardless of their origin collapse down into a single band. MS would be required to identify the nature of the non-phosphorylated species and whether they arose from NCS, dephosphorylation or other breakdowns in the GCE system. Another important consideration is to run a standard SDS-PAGE gel alongside the Phos-tag gel to confirm the electrophoretic mobility shifts observed in the latter are due to phosphorylation status and not protein size. Lastly, treating the phospho-protein with a non-specific phosphatase (e.g. λ-phosphatase) should cause it to migrate with the same mobility as wild-type protein, providing an additional layer of confidence that the Phos-tag gel shifts are due to authentic, reversible phosphorylation.

The technique is relatively straightforward, requiring only a traditional SDS-PAGE setup. As a result, Phos-tag electrophoresis is now routinely used for rapidly assessing GCE-encoded phosphoamino acids and monitoring the phosphorylation status of target proteins during purification and characterization. Phos-tag gels are compatible with western blotting, making it possible to analyze the phosphorylation status of proteins in complex mixtures such as cell-lysates. One caveat is that not all phosphorylated proteins exhibit distinctive electrophoretic mobility shifts with Phos-tag.111 In such cases, MS strategies may be necessary to confirm phosphorylation status. Additionally, mobility shifts can be challenging to distinguish for high-molecular-weight phospho-proteins, especially those exceeding 100 kDa, which already migrate slowly in SDS-PAGE.112

2.5.3. Western Blotting.

With the development of numerous phospho-protein and phosphoamino acid specific antibodies, western blotting techniques have become valuables tools for the detection of protein phosphorylation events (Fig. 7C).113 The efficacy of these techniques is intricately tied to the specificity and efficiency of antibody binding to the phospho-antigen, a variable factor that hinges on the quality and design of the antibodies themselves.114 These antibodies serve as a reliable means to confirm the presence of phosphorylated proteins within a sample, and while they are invaluable for the qualitative assessment of phosphorylation, they are of limited use for quantifying the amount of phosphoprotein present unless standards of the same protein with identical phosphorylation sites are available for side-by-side comparative analysis. Also, since non-phosphorylated protein forms are not visualized with the same antibody, the approach provides no information about the relative amounts of phospho- and unmodified-protein present and therefore is unsuitable for evaluating homogeneity of phospho-proteins expressed via GCE.

2.5.4. Phospho-specific gel stains.

Phosphorylation-specific gel stains are specialized dyes or chemical compounds used to visualize phosphorylated proteins in polyacrylamide gels (Fig. 7D).115,116 These stains are designed to fluoresce when bound to phosphorylated amino acid residues, including pSer (2), pThr (4), and pTyr (6). They provide a straightforward and rapid means of assessing the phosphorylation status of proteins within a complex mixture and, like western blotting, they cannot provide information about the proportion of protein that is phosphorylated in a sample without known standards for comparison.

3. Development and applications of phosphoamino acid GCE systems

In this section, we dive into the history of GCE development for encoding phosphorylated amino acids pSer ((2), section 3.1), pThr ((4), section 3.2), and pTyr ((6), section 3.3), as well as their mimics. We discuss the evolution of these recombinant expression systems, focusing on the particular challenges that were overcome with each development while highlighting the differences and unique characteristics of each system. We also provide practical considerations for non-experts on how to adopt these systems, and also explore noteworthy studies that have harnessed these technologies to uncover novel insights into protein function, cellular processes and signaling pathways. We offer our perspective as well on the future of these phospho-GCE technologies, where new directions and exciting opportunities promise to unlock mysteries in our understanding of phosphorylated protein function.

3.1. GCE systems for phosphoserine (2) and its mimic

Phosphoserine, as the most prevalent phospho-PTM in the eukaryotic proteome,32 naturally emerged early on as a focal point for GCE system development. In this section, we present how the GCE systems for encoding pSer and a non-hydrolyzable mimic have developed over time. Since the first pSer encoding system was published in 2011,66 a variety of research groups have contributed to its refinement, all with the aim of enhancing access to biologically relevant phosphorylated proteins and insights into their functions. When tracing this history, two largely independent paths of development emerge: one for what is referred to as “SepOTS” (the phosphoserine orthogonal translational system), and the other for a system that lacks a formal designator but we will refer to here as the “pKW2” system to reflect the assigned name of the original machinery plasmid developed to drive pSer encoding.

3.1.1. Development of the “SepOTS” phosphoserine GCE platforms

3.1.1.1. The first GCE encoding of pSer.

In 2011, Park et al. reported the first successful GCE-based encoding of pSer (2, Fig. 8A) into proteins in E. coli.66 This system took advantage of an earlier discovery by the same group that there are natural pSer-specific tRNA synthetases found in some methanogenic archaea.117 These archaea do not encode Cys into proteins by directly loading cysteine onto a tRNA but instead they use a pSer-specific RS (SepRS) to amino-acylate tRNAGCACys with free pSer.117 Once amino-acylated, the Sep-tRNAGCACys undergoes sulfhydryl transfer to generate Cys-tRNAGCACys, and this Cys is then incorporated into proteins by the ribosome at UGC codons. These naturally occurring RSs specific for pSer provided an ideal starting point for encoding pSer into proteins. To identify a SepRS/tRNAGCACys pair amenable to pSer encoding, in vitro experiments by Hohn et al.118 found that the SepRS from Methanococcus maripaludis (Mma) could reasonably efficiently charge tRNAGCACys from Methanocaldococcus jannaschii (Mj), whereas the tRNAGCACys from Mma folded poorly and could not be charged by its cognate Mma SepRS.

Figure 8. pSer, its nhpSer mimics and metabolic alterations of specialized E. coli strains used for pSer and nhpSer GCE systems.

Figure 8.

(A) Shown are the side-chain structures of pSer (left, blue box) and three nhpSer mimics (right, gray box) which have the pSer bridging γ-oxygen replaced with either CH2, CHF, or CF2 (highlighted red). (B) In the “ΔserB” E. coli strains used for pSer incorporation, the cytosolic pSer concentration is increased by knocking out the pSer phosphatase SerB (red “X”), enhancing pSer incorporation by GCE machinery into the protein of interest (POI). (C) In the “ΔserC” E. coli strains used for nhpSer incorporation, to minimize competing pSer (upper pathway), intracellular pSer is decreased by knocking out the pSer aminotransferase SerC (red “X”). In addition, in the “Permaphos” system, overexpression of the pSer phosphatase SerB (⇑) is achieved by including a SerB gene on the pERM2 plasmid, and cytosolic nhpSer is produced (lower pathway) using five enzymes (Frb A/B/C/D/E from S. rubellomurinus) encoded by the pCDF-Frb-v1 plasmid plus the action of an E. coli transaminase (?). The nhpSer is then incorporated by GCE machinery into the POI. No GCE systems yet exist for F1-nhp-Ser or F2-nhpSer.

Building on this foundation, Park et al.66 created an amber suppressing “Sep-tRNACUA” by making three base changes in Mj Sep-tRNAGCACys (changing the GCA anticodon to CUA, and changing a C→U at position 20 to improve amino-acylation activity, Fig. 9). This Sep-tRNACUA was amino-acylated by SepRS about 40% as well as the Mj tRNAGCACys. The next challenge was to achieve sufficient concentrations of pSer inside the cell to feed the GCE system. Noting that free pSer amino acid is the biosynthetic precursor to serine,119 Park et al. circumvented the issues associated with poor uptake of exogenously adding pSer, as well as preventing its breakdown once inside the cell, by knocking out the phosphatase responsible for converting pSer amino acid into serine (ΔserB, Fig. 8B). This succeeded at elevating intracellular pSer levels sufficiently to charge the Mj tRNACUACys, but the system failed to encode pSer at UAG codons. Interestingly, addition of the archaeal enzyme responsible for converting Sep-tRNACys to Cys-tRNACys, led to Cys encoding at UAG codons, proving that Sep-tRNACUA was being generated but was not efficiently delivered to the ribosome by the native E. coli EF-Tu.

Figure 9. Mutations producing Sep-tRNA variants used in pSer encoding systems.

Figure 9.

Shown is a cloverleaf depiction of M. jannaschii tRNACys along with the mutations present in each of the four Sep-tRNA variants used in pSer/pThr-encoding systems. Mutation sites are indicated by arrows pointing from the wild-type nucleotide (gray circles) to the new nucleotide identity (circles with colors corresponding to each tRNA variant per the key in the figure).

To solve this problem, the authors identified “EF-Sep” from a library of E. coli EF-Tu variants that carried six mutations in its amino acid binding pocket allowing it to accommodate the −2 charged pSer amino acid at the 3’ end of Sep-tRNACUA (Fig. 10A). Using EF-Sep improved the delivery of Sep-tRNACUA to the ribosome with a ten-fold enhancement of pSer incorporation in response to UAG stop codons. This first detectable phosphorylated amino acid encoding into a protein – using both myoglobin and MEK1 as test cases – was an impressive tour de force that established the basic SepOTS platform, later termed SepOTSα, that comprised the Mma SepRS (wild-type)/Mj Sep-tRNACUA pair and “EF-Sep” in an E. coli ΔserB strain (system pSer-1, Table 1). While this was a major achievement, the yield of only 25 μg of full-length MBP-MEK1(pSer218) per liter culture left plenty of room for refinements and improvements to make this technology more effective.

Figure 10. Rationale behind engineering EF-Tu, SepRS and Sep-tRNA for pSer and nhpSer encoding.

Figure 10.

(A) Shown are surface views of the amino acid binding pocket of E. coli EF-Tu (left, PDB: 1OB2) and EF-Sep (right, model generated by mutating residues in PyMOL) colored by electrostatic potentials (scale on the right; calculated with PYMOL default values). The pSer ligand shown (stick figure, atom coloring) was converted from Phe (in the original PDB entry) using PyMOL. Also visible is a semitransparent ladder cartoon of the tRNA backbone and bases. (B) Structure of the A. fulgidus Sep-tRNA (pink ribbon) complex with SepRS (light blue ribbon) is shown (PDB: 2DU3) focused on the RS-tRNA binding interface and highlighting the four protein residues (stick models with atom coloring and green carbons) corresponding to sites in the M. maripaludis SepRS for which variants were screened to select the SepRS9/Sep-tRNA pair.79 Two tRNA base pairs in the anticodon (G34C and C35U; yellow stick models) were mutated directly without screening. No tRNA base pairs were screened as noted in the left-hand table. For the RS, the positions screened, their wild-type identities and what they were mutated to are listed in the right-hand table. (C) Structure of the A. fulgidus Sep-tRNA complex with SepRS, as in panel B, but highlighting the six tRNA bases (pink stick models) and the six protein residues corresponding to sites in the M. jannaschii Sep-tRNAGCA and the M. maripaludis SepRS for which variants were screened during the selection of the SepRS2/Sep-tRNA(B4) pair.120 As in panel B, two tRNA base pairs in the anticodon (G34C and C35U; yellow stick models) were mutated directly without screening. For the tRNA and the RS, respectively, the positions screened, their wild-type identities and what they were mutated to are listed in the left- and right-hand tables. This figure, as well as Figures 13 and 15, were created using PyMOL.

TABLE 1:

E. coli in-cell GCE systems for encoding phosphoserine and its non-hydrolyzable mimic, with relevant Addgene IDsa

Year
[ref]
ncAA System Plasmid 1 Plasmid 2 Plasmid 3 E. coli
Expression
host
Name
(ori)
ORF 1
Promoter/Gene(s)
ORF 2
Promoter/Gene
Additional ORFs
Promoter/Gene
Name
(ori)
ORF 1
Promoter/Gene
ORF 2
Promoter/Gene
Name
(ori)
ORF 1
Promoter/Gene
2011 [66] pSer pSer-1 (OTS-α) pKD (pBR322) trcb SepRS/EFSepc - (p15a) lpp POI lpp Sep-tRNAc BL21(DE3) ΔserB, Top10 ΔserB
2012 [121] pSer-2 pKD (pBR322) trcb SepRS/EFSepd - (p15a) lpp POI lpp Sep-tRNAc EcAR7 (RF1 KO)
2013 [79] pSer-3 pKD (pBR322) trcb SepRS9e/EFSep21d pETduet (pBR322) T7 SepRS9e T7 Sep-tRNAc pCDFduet (CloDF) T7 POI BL21(DE3)
2015 [81] pSer-4 (OTS-μ) pKD (pBR322) trcb SepRS/EFSepd lpp Sep-tRNAc (5 copies) Not indicated (N/I) N/I POI EcAR7 (RF1 KO)
2015 [120] pSer-5f (pKW2) pKW2-EFSep (pBR322) glnS SepRS2e tac EFSepd lpp Sep-tRNA(B4) c pNHD (p15a) T7 POI BL21(DE3) ΔserB
2015 [97] pSer-6g (OTS-λ) SepOTSλg (pBR322) trcb SepRS9e/EFSep21d lpp Sep-tRNAc G37A (4 copies) PCRT7 (pBR322) PLtetO-1 POI C321 ΔA ΔserB (RF1 KO)
2019 [84] pSer-7h (pKW2-ΔA) pKW2-EFSep (pBR322) glnS SepRS2e tac EFSepd lpp Sep-tRNA(B4)c pRBC (p15a) T7 POI B-95(DE3) ΔA ΔfabR ΔserB (RF1 KO)
2023 [122] pSer-8i pSerOTS-C1* (V70) (pBR322) glnSb SepRS9e/EFSep21d proK Sep-tRNAc G2:C71 Unspecified plasmid (p15a) rEcoliXpS
2015 [120] nhpSer nhpSer-1 pKW2-EFSep (pBR322) glnS SepRS2e tac EFSepd lpp Sep-tRNA(B4)c pNHD (p15a) T7 POI pCDFduet (CloDF) T7 SerB/EFSepd Keio ΔserC, BL21(DE3) ΔserC, DH10b ΔserC
2021 [123] nhpSer-2 pEVOL (p15a) glnS SepRS2e araC SepRS2e proK Sep-tRNAc (B4) pRSF (RSF) tacb EFSepd - serB pKS (pBR322) tac POI B834 (DE3)
2023 [49] nhpSer-3j pERM2 (pUC) glnS SepRS2e tac EFSepd lpp Sep-tRNA v2c pRBC (p15a) T7 POI pCDF-Frb-v1 T7 FrbABCDE BL21(DE3) ΔserC
OXB 20 SerB
a

For components available from Addgene, IDs are provided in a footnotes below for that GCE System number

b

bi-cistronic expression

c

Sep-tRNA mutations (C20U, G34C, C35U); Sep-tRNA(B4) mutations (G34C, C35U,G29A, A31U, G37A, U39A, C40U, C41U); Sep-tRNA G2:C71 mutations (Sep-tRNA + G37A, C2G, G71C); Sep-tRNA-v2 mutations (Sep-tRNA(B4) + C2U, G4C, G6C, C67G, C69G, G71A). See Fig. 9 for structural representation of Sep-tRNA mutants.

d

EFSep mutations (F219Y, D217G, T229S, E216N, H67R, N274W); EFSep21 mutations (F219Y, D217G, T229S, E216V, H67R, N274W)

e

SepRS9 mutations (K347E, N352D, E412S, E414I, P495R, I496R, L512I); SepRS2 mutations (E412P, E414F, P495M, I496W, F529S);

f

Component Addgene IDs are 173897 (plasmid 1), 174075 and 174076 (control sfGFP forms and for POI), 34929 (E. coli strain). The sfGFP/POI plasmids are not the ones originally used with this system, but work well with it.

g

Component Addgene IDs are 68292 (plasmid 1), 68295 (GFP-TAG control and POI), 68306 (E. coli strain). Note: the control/POI plasmid has the same origin of replication as the machinery plasmids, so we do not recommend using this reporter plasmid. We also do not recommend using the pRBC plasmid (see system 7). Even though it has a compatible p15a origin, the proteins are expressed via a T7 promoter, and the C321.ΔA and rEcoliXpS strains do not have a T7 polymerase.

h

Component Addgene IDs are 173897 (plasmid 1), 174075 and 174076 (plasmids for control sfGFP forms and for POI), 197655 (E. coli strain)

i

Component Addgene IDs are 188537 (plasmid 1) and 192872 (E. coli strain).

j

Available as Addgene kit #1000000226. Component Addgene IDs are 201922 (plasmid 1), 174075 and 174076 (plasmids for control sfGFP forms and for POI), 201923 (plasmid 3), and 197656 (E. coli strain).

3.1.1.2. SepOTSμ: deleting RF1 and increasing tRNA gene copy number enhances phosphoprotein expression.

As for all E. coli GCE systems based on amber codon suppression, a key factor limiting pSer encoding using SepOTS was the competition between the amber suppressing tRNA and RF1. In typical E. coli strains, the deletion of RF1 is lethal, however Mukai et al. in 2010 demonstrated that the mutation of just seven chromosomal stop codons from TAG to TAA allowed for RF1 deletion, thereby enhancing UAG codon suppression efficiency.124 Using this strategy, Heinemann et al. in 2012 were able to delete RF1 in the ΔserB background strain of E. coli, creating a strain they called EcAR7.121 When paired with the pSer GCE machinery from Park et al.,66 UAG codon suppression was improved 10-fold in EcAR7 compared to that the original RF1-containing ΔserB strain of E. coli (system pSer-2, Table 1). One downside of the EcAR7 strain was that it exhibited slow growth rates (doubling time ~2 h), and these worsened (to ~6 hours) when expressing pSer GCE components. Also, even though some successful encoding of pSer was confirmed via mass spectrometry, it remained unclear how much of the UAG codon suppression resulted from pSer encoding versus that arising from near-cognate suppression.

A 2014 follow-up study resolved this question using a more comprehensive top-down MS analysis of GFP expressed with a TAG site at position 17 in the EcAR7 host.71 These analyses showed that only ~25% of GFP produced was phosphorylated, with the rest having either Gln (~30%), Lys (~15%), Tyr (~15%) or one of a few other natural amino acids incorporated through near-cognate suppression. As a strategy to improve the fidelity of pSer incorporation, they tested increasing the Sep-tRNACUA concentration by including 5 copies of the tRNA gene in the expression plasmid (later termed the “SepOTSμ” system, system pSer-4 Table 1), and this did help, increasing to ~40% the fidelity of pSer-incorporation into GFP protein produced using the EcAR7 strain.71

3.1.1.3. SepRS9 and EFSep21: next generation SepRS and EF-Sep variants.

RSs have evolved intricate mechanisms to ensure they aminoacylate only their cognate tRNA(s). In the context of the SepRS/tRNAGCACys pair, the anticodon loop of tRNAGCACys makes extensive contact with SepRS indicating its anticodon is a key recognition motif for cognate SepRS/tRNA pairing (Fig. 10B).125 Altering the tRNAGCACys anticodon loop to CUA for amber codon suppression compromised interactions between SepRS and Sep-tRNACUA leading to poor amino-acylation kinetics. To address this issue, Lee et al. in 2013 identified four mutations in the anticodon recognition domain of SepRS (to create the SepRS9 variant) that enhanced SepRS and Sep-tRNACUA interactions (Fig. 10B), resulting in a roughly two-fold increase in overall pSer encoding efficiency (as defined by its ability to suppress a TAG codon in a chloramphenicol resistance gene and confer survival at increasing concentrations of chloramphenicol).79

Lee et al. also sought to improve Sep-tRNACUA delivery to the ribosome, by probing the contributions of the six sites changed in EF-Sep.79 They found that an EF-Sep variant having an N216V mutation, dubbed “EF-Sep21”, showed 4-fold increased chloramphenicol tolerance, and the combination of SepRS9 with EF-Sep21 provided a six-fold improvement in chloramphenicol tolerance when compared to the original SepOTSα introduced by Park et al. in 2011. With these improvements (system pSer-3, Table 1), Lee et al. generated milligram quantities Histone H3 uniformly phosphorylated at Ser10 in the RF1-containing E. coli BL21(DE3), avoiding use of the RF1-deficient strain EcAR7 because “cell viability was critically impaired”.79 With pure histone H3-pSerl0 in hand, they were able to demonstrate that SAGA-mediated acetylation of H3 was stimulated by phosphorylation. This improved SepOTS was the first to show that pSer GCE technologies could be used to answer biologically relevant questions.

3.1.1.4. SepOTSλ: adapting SepOTS in a genomically recoded E. coli expression host.

The enhanced suppression efficiency of RF1-deficient expression hosts provided an attractive strategy for improving pSer encoding, but having observed “the partially recoded EcAR7.ΔA… exhibited severe growth impairment and suppression with natural amino acids97, engineering a healthier RF1-deficient host was seen as key to leveraging their advantages. Seminal work from Lajoie et al. in 2013 showed that by first mutating all 321 instances of chromosomal TAG stop codons to TAA in the MG1655 strain of E. coli (instead of only 7), cellular viability could be maintained upon RF1 deletion.126 This C321.ΔA strain marked a pivotal achievement as the first fully “genomically recoded” organism, with only 63 codons in its genetic code and leaving UAG unassigned. The C321.ΔA strain displayed much improved health compared to EcAR7.ΔA (doubling times ~90 min vs ~140 min, respectively), though its growth rate was still slower than standard E. coli protein expression strains like BL21(DE3), which typically double in ~30 minutes. The reduced fitness of the C321.ΔA compared to its parent MG1655 strain is attributed to “hitchhiker mutations that accumulated during recoding” as well as the web of overlapping genes in E. coli, where the mutation of some chromosomal TAG codons to TAA inadvertently mutated other genes and negatively impacted their function.127,128

In 2015, Pirman et al. hypothesized that with its improved fitness compared to EcAR7, the fully-recoded C321.ΔA strain containing the ΔserB mutation could be used to improve pSer encoding efficiency and fidelity.97 Using the pSer GCE machinery SepOTSμ from Steinfeld et al.,71 containing 5 expression cassettes of the tRNACUA, they found no substantial improvements in pSer encoding efficiency using the C321.ΔA strain. Subsequent refinement in GCE machinery included (i) the introduction of the G37A point mutation in Sep-tRNACUA to enhance amber suppression (Fig. 9), (ii) reducing the number of Sep-tRNACUAA37 genes from five to four, (iii) and adopting the SepRS9 and EFSep21.97 These modifications led to “SepOTSλ” (system pSer-6, Table 1), providing (i) a ~2 fold increase in overall sfGFP production compared to the “SepOTSμ” from Aerni et al.71 in the C321.ΔA strain, and (ii) an increase in single site pSer encoding fidelity into GFP from ~50% to ~80%. With these developments, the successful production of fully activated, doubly phosphorylated MEK1 (pSer218/222) represented the first GCE expression of a multiply phosphorylated protein. However, the majority of MEK1 (pSer218/222) was not phosphorylated, suggesting “that near-cognate suppression of the amber codon can lead to natural amino-acid incorporations that interfere with SepOTS activity and phosphoprotein purity”. 97 Thus, challenges related to near-cognate suppression still persisted, and pSer encoding fidelity varied quite notably based on the target protein and site of pSer encoding, ranging from below 50% and up to about 80%.

3.1.1.5. pSerOTS-C1* (V70): Improving expression host fitness for systems-level biology applications.

GCE components can impact expression host growth and viability due to breakdowns in orthogonality that lead to mis-acylation of endogenous tRNAs and the amber suppressing tRNA. Cells experiencing such stresses associated with proteome-wide mis-encoding become less tolerable to extended passages or growth cycles needed to perform systems-level studies or the screening of libraries for phosphorylation-dependent phenotypes. Mohler et al. sought to address these challenges and evaluated the physiological effects of expressing pSer GCE components in wild-type BL21(DE3) cells and in RF1-deficient cells, including rEcoliXpS (a derivative of the genomically recoded C321.ΔA).122 They found that higher expression levels of SepRS and Sep-tRNACUA were associated with widespread proteome dysregulation and slower growth rates, although these were less severe in the genomically recoded host.

In probing for underlying molecular mechanism for the dysregulations, the authors found issues with orthogonality. Specifically, they observed overlap in RS identity elements between Sep-tRNACUA, tRNAGly and tRNAThr, causing the SepRS to mis-acylate endogenous tRNAGly (leading to mis-encoding of pSer at Gly codons), and the Sep-tRNACUA to be mis-acylated by Gly and Thr RSs (leading to Gly/Thr mis-encoding at UAG codons).122 By lowering the number of Sep-tRNACUA expression cassettes back down to one (instead of four as in SepOTSλ), optimizing transcriptional promoters, and introducing a C2:G71→G2:C71 mutation in the acceptor stem of Sep-tRNACUA (Fig. 9), pSer encoding fidelity was improved from ~50% to ~75% while the overall phospho-protein yield increased ~5-fold compared to SepOTS. Also, the host cells experienced less systemic dysregulation.

While the resulting pSer GCE system, dubbed “pSerOTS-C1* (V70)” (pSer-system 8, Table 1), did not achieve homogenous encoding of pSer into target proteins, these studies importantly highlighted the negative effects of introducing non-natural GCE translational components into living cells, and also provided some strategies for how to systematically mitigate the negative effects. Having such a system in which the expression hosts experiences lower levels of dysregulation while synthesizing pSer proteins sets the stage for more extensive systems-biology studies probing phospho-protein interaction networks. Such studies would not be feasible in protein over-expression strains such as BL21(DE3) and its derivatives, where plasmid instability and toxic effects of GCE systems would complicate longer term cell-based studies.129

3.1.2. Development of the “pKW2” phosphoserine GCE systems

3.1.2.1. Engineering the “pKW2” system with a much-improved SepRS/Sep-tRNACUA pair.

In the same year that Pirman et al.97 reported adopting the genomically recoded C321.ΔA strain to create SepOTSλ, Rogerson et al.120 reported a substantially enhanced pSer encoding system using an alternate strategy to evolve a much more efficient Mma SepRS/Mj Sep-tRNACUACys pair. Their approach was predicated on the hypotheses that (i) the Sep-tRNACUA amber codon suppression could be improved by evolution of its anticodon loop region (pink atoms, Fig. 10C), (ii) the interactions between the SepRS and Sep-tRNACUA could be improved by evolving the SepRS anticodon recognition loop (green amino acids, Fig. 10C) with a more extensive library than was previously screened, and (iii) the ideal SepRS and Sep-tRNACUA variants would be identified as a pair via co-evolution, rather in separated, sequential evolutions.

First, Rogerson et al. generated a Sep-tRNACUA library containing all possible combinations of the mutations in the ten nucleotides flanking the CUA anticodon (G29 to U33 and G37 to C41, Fig. 9).120 They also generated a library of SepRS variants in which all possible amino acid combinations at 6 sites within its anticodon recognition domain were generated (E412, E414, K417, P495, I496 and F529) (Fig. 10C). While unfeasible to screen both together (104 x 206 = ~1013 combinations), the Sep-tRNACUA library was first screened for variants with improved ability to suppress UAG codons when co-expressed with wild-type (WT) SepRS and EF-Sep. These screens identified 13 Sep-tRNACUA variants conferring a ~10-fold increase in chloramphenicol resistance when compared to the original Sep-tRNACUA. Then each of these 13 functional Sep-tRNACUA variants was used to screen the six-site SepRS library, and the top SepRS/Sep-tRNACUA pairs from all 13 selections were evaluated and compared against each other for encoding efficiency as measured by chloramphenicol resistance.

From this process, SepRS(2)/Sep-tRNA(B4)CUA emerged as the most efficient SepRS/Sep-tRNACUA pair (the “pKW2” system, system pSer-5, Table 1, Fig. 9 and 10C), and it exhibited an impressive 18-fold improvement in pSer encoding efficiency compared to SepRS9/Sep-tRNACUA. Notably, SepRS(2)/Sep-tRNA(B4)CUA even drove pSer encoding using wild-type EF-Tu, although using EF-Sep (or EF-Sep21) improved encoding efficiency by an additional ~10-fold confirming its importance in pSer encoding. The power of the SepRS(2)/Sep-tRNA(B4)CUA pair with EF-Sep, when expressed in the RF1-containing strain BL21(DE3) ΔserB, was showcased by achieving homogeneous (i.e. >95%) pSer encoding into Ubiquitin at site S65 and into Nek7 kinase at site S195, as confirmed by whole-protein MS analysis. Reported yields were 3 and 2 mg per liter culture, respectively. Successful encoding of pSer into calmodulin and mitigation of truncation was also demonstrated using the RF1-deficient C321.ΔA strain, although neither the encoding fidelity nor yields were specifically assessed.

3.1.2.2. pKW2-ΔA: high-fidelity pSer encoding in a healthy, truncation-free expression host.

To generate a robust RF1-deficient strain of E. coli designed for high-level protein expression, Mukai et al. first sought to identify the minimal set of genes that naturally end in TAG codons that are either needed for maintaining cell viability or are needed to maintain full cellular fitness.127 They scoured 9 E. coli gene databases and identified candidate genes that had both a TAG stop codon and were considered necessary for vigorous growth; these included genes involved in the maintenance of DNA, RNA, and proteins, as well as key genes for healthy metabolism, energy regeneration, cell division, and starvation/stress responses. Importantly, they also checked for instances when mutating the TAG stop codon of one gene to TAA/TGA would inadvertently mutate another key overlapping gene, and employed clever strategies to ensure neither protein’s function was changed. Collectively, 95 chromosomal TAG codons were mutated to TAA or TGA prior to RF-1 deletion and impressively, the resulting strain, called B95(DE3) ΔA, grew with the same doubling time as its parent strain BL21(DE3) (~30-40 min) in standard rich media.127 Slower growth was observed in minimal media, however, and so after 360 generations of selective pressure a variant with a frameshift mutation in the fabR gene was identified that displayed improved growth in minimal media, as well as lower temperatures. In this way, Mukai et al. created a BL21(DE3) derivative called B95(DE3) ΔA ΔfabR that lacked RF1 and displayed full cellular fitness in rich media plus the ability to grow in minimal media and at lower temperatures. With this strain they successfully employed GCE systems to encode sulfo-tyrosine into hirudin at up to 3 sites, and p-azido-Phe into the Fab antibody fragment of Herceptin also at up to 3 sites.127

Having recognized the enhanced vitality and utility of the B95(DE3) ΔA ΔfabR strain for GCE applications, Zhu et al. adapted this RF1 deficient expression host for pSer encoding to produce the first instance of homogeneous, single and multi-site pSer encoding without truncation.84 To do this, they first deleted the serB gene to make the strain B95(DE3) ΔA ΔfabR ΔserB. Coupling this strain with the high efficiency pKW2 GCE machinery plasmid and the target protein expression plasmid from Rogerson et al.120 (pNHD, conferring tetracycline resistance) resulted in slow cell growth. Cell growth was restored by changing the antibiotic resistance of pNHD to ampicillin to produce the pRBC plasmid. After optimizing expression parameters, use of this strain with the improved pKW2/pRBC plasmid combination created the pKW2-ΔA system (pSer-7, Table 1), which provided >95% pSer encoding fidelity in an sfGFP-150TAG reporter protein, with high efficiency (~200-300 mg/L culture of protein) and without truncated protein. The overall yields for single and double pSer encoding into sfGFP were 2- and 3-times greater in the B95(DE3) ΔA ΔfabR ΔserB strain, respectively, compared to the RF1-containing parental strain BL21(DE3) ΔserB, highlighting its improved utility. Furthermore, this system allowed the faithful encoding of pSer into an sfGFP reporter protein at up to 5 simultaneous sites, into dimeric STING (at 2 different sites), and into the Bcl2-associated agonist of cell death (BAD) at up to 3 sites at once, all without truncation nor notable mis-encoding.

3.1.2.3. Making isotopically labeled pSer-proteins for NMR studies.

Nuclear magnetic resonance spectroscopy (NMR) is an ideal technique for studying the structural features and dynamics of disordered proteins.130 This means that NMR is particularly valuable for the study for phosphoproteins, because as noted earlier, the majority of biologically relevant phosphorylation sites reside within disordered or unstructured regions of proteins.12-14 The preparation of proteins for most kinds of NMR studies involves isotopic enrichment with 15N, 13C, and/or deuterium, and this enrichment is typically achieved by expressing proteins in minimal media containing 13C-sugars, 15NH4Cl, and/or D2O. However, E. coli strains used for encoding pSer into proteins with GCE that contain the ΔserB mutation cannot grow in these minimal media formulations, because they are serine auxotrophs (i.e. unable to biosynthesize serine, Fig. 8B).

The first report of generating isotopically-labeled protein with pSer encoded was in 2020 in which 15N-labeled pSer111-Rab1B was produced to study phosphorylation-dependent structural conformations.131 Shortly after in 2021, pSer was encoded into 15N/13C labeled ubiquitin, GB1 and Hsp90-N to generate lanthanide metal binding sites for long-range distance measurements using paramagnetic NMR analysis.132 Puzzlingly, however, both studies reported target protein expression in standard minimal M9 media with BL21(DE3) ΔserB. Perhaps trace amounts of serine impurities in the phosphoserine amino acid used to supplement the media supported cell growth in these cases. Nevertheless, around the same time, Stuber et al. generated 15N/13C-labeled phospho-ubiquitin and NEDD8 in a BL21(DE3) strain having an intact serB gene.123 In this case the cells could grow in fully-defined minimal media and high amounts of unlabeled pSer supplementation to the media was sufficient to drive incorporation. The expressed ubiquitin protein purified as a mixture of phosphorylated and unphosphorylated forms that had to be separated, but such separation of phosphorylated and unphosphorylated forms of NEDD8 was not feasible.

To provide a more generalizable strategy—one in which pSer is biosynthesized ensuring it is isotopically labeled prior to encoding—Vesely et al. demonstrated that E. coli BL21(DE3) ΔserB growth in minimal media could be restored by supplementing with isotopically-enriched Celtone, a commercially available algal extract.88 After optimization of growth and expression parameters, Vesely et al. used the pKW2 system to produce 15N-labeled pSer-containing proteins – sfGFP-150TAG and the SARS-CoV-2 nucleocapsid protein – with ~80-90% fidelity and yields of ~10 mg/L and ~3 mg/L, respectively. Optimized conditions included expression at lower temperatures (18 °C) to minimize phosphatase activity, use of the Sep-tRNACUAv2.0 to minimize mis-acylation events (Fig. 9, see section 3.2.1 below on pThr GCE encoding systems), and high-density culturing conditions to enhance yields.

Later, Buchko et al.133 attempted to use the strategy from Vesely et al.88 with BL21(DE3) ΔserB to produce 15N, 13C, and deuterium labeled phosphorylated amelogenin, a protein involved in tooth enamel formation, but insufficient protein yields hindered progress. Efforts to improve pSer encoding efficiency with the RF1-deficient B95(DE3) ΔA ΔfabR ΔserB strain using the low-temperature expression methods by Vesely et al.88 led to unacceptable amounts of near-cognate suppression. However, after a variety of refinements, including expression at higher temperature (37 °C) and use of low-density cultures, pSer-amelogenin was expressed and purified with yields of ~10 mg per liter of culture with ~90% fidelity. That little dephosphorylation was observed at these higher temperatures, in contrast to reports from Vesely et al.,88 could be attributed to the fact that amelogenin expresses in the insoluble fraction thereby protecting it from phosphatases. These results collectively highlight that E. coli ΔserB serine auxotrophic strains cannot grow in conventional minimal media used for isotopic enrichment of pSer proteins, but that growth and protein expression can be restored by supplementing the media with isotopically-enriched supplements such as Celtone. Also, these results are a reminder that for any given protein or site of phosphorylation, tweaking parameters such as temperature, media and culture density are often needed for improving yields and purity of recombinant phosphorylated proteins.

3.1.3. GCE systems for encoding non-hydrolyzable phosphoserine (nhpSer, 20)

3.1.3.1. nhpSer as a faithful phosphoserine mimic.

An advantage of GCE is its adaptability in installing phosphorylated amino acid mimics that are stable to hydrolysis and are more faithful than Asp and Glu in mimicking the function of authentic pSer. In this section, we review recent GCE advances for encoding a non-hydrolyzable mimic of pSer in which the bridging γ-oxygen is replaced with a methylene (CH2) group (phosphono-methyl-alanine, referred to here as non-hydrolyzable phosphoserine or nhpSer for simplicity; 20, Fig. 8A). This change renders nhpSer impervious to phosphatase hydrolysis while maintaining the tetrahedral geometry of the phospho-group.

Given the loss of a potential hydrogen-bond acceptor in the bridging oxygen position, and an elevated pKa2 of ~7.0-8.0 compared to pSer at ~5.6-6.0,134 researchers sought to make better pSer mimics by substituting the methylene hydrogen atoms of nhpSer with one or two electron-withdrawing fluorine atoms (21 and 22, Fig. 8A). Substitution with a single fluorine (21, F1-nhpSer) lowers the phosphonate pKa2 to match that of phosphoserine, while double fluorine substitution (22, F2-nhpSer) drops the pKa2 to be lower than pSer at ~5.0 but it may mimic the hydrogen bonding properties of pSer better.134 To gain some insight into the extent to which the loss of a hydrogen bonding γ-oxygen in nhpSer would impact its ability to mimic pSer (2) in proteins, we analyzed all 554 unique structures in the PDB (<95% sequence identity) containing pSer within a peptide (i.e. non-ligands) with HBPlus.135 Of these 554 structures, only in 92 (17%) does the γ-oxygen of pSer hydrogen bond with a protein donor atom (Supporting Table 1). Though a more thorough analysis would require a separate study, we speculate based on this cursory analysis that the γ-oxygen of pSer is generally an uncommon structural recognition element in pSer-dependent protein interactions. Indeed, well-established pSer-binding domains such as 14-3-3, BRCT and WW domains do not make interactions with the γ-oxygen.49 We could find only one case in which pSer (2), nhpSer (20) and F2-nhpSer (22) were directly compared side-by-side, and no substantive differences in their ability to promote 14-3-3/client protein interactions were reported.136,137

Nevertheless, this analysis also makes clear that there will be scenarios in which F2-nhpSer (22) could prove beneficial over nhpSer (20) for its hydrogen bonding properties. In one such example, recent work by Patskovsky et al. showed that the dimeric complex composed of a cancer-associated pSer-peptide and the major histocompatibility complex (MHC) bound the T-cell receptor-27- LC13 (TCR27) with high affinity (1.9 μM) to activate T-cells.138 Crystal structures of this MHC/pSer-peptide complex showed the γ-oxygen of pSer made important H-bonds with MHC, and indeed when pSer (2) was replaced with nhpSer (20), the phospho-peptide/MHC complex bound TCR27 with 20-fold weaker affinity (47 μM) and activated T-cells less well than the pSer peptide/MHC complex.138 While nhpSer (20) mimicry was imperfect in this context, important to note is that the nhpSer-peptide was still able to bind TCR27 and activate T-cells, whereas peptides containing Asp/Glu or sulfo-serine mimics were unable to do so. Thus, both nhpSer and F2-nhpSer appear to be viable mimics of pSer, particularly in contexts where Asp/Glu are not. To our knowledge, mimicry of F1-nhpSer (21, either diastereomer) in context of a peptide or protein has not been tested. In this section, we focus on systems for encoding nhpSer (20), because no GCE systems have yet been reported for installing the fluorinated nhpSer variants (21, 22).

3.1.3.2. Adapting pSer GCE components to encode nhpSer.

In 2015, Rogerson et al. demonstrated that nhpSer (20) was a permissive substrate for the SepRS2/Sep-tRNA(B4)CUA pair.120 Recognizing that high-fidelity installation of nhpSer would not be possible if competing pSer (2) amino acid were present in the cell, they created an E. coli BL21(DE3) derivative in which the endogenous pSer was depleted by deleting serC. SerB was also overexpressed to facilitate removal of residual pSer that might enter the cell from the media, or be made via promiscuous transaminases that can substitute for SerC function (Fig. 8C). Thus, by co-expressing (i) the SepRS2/Sep-tRNA(B4)CUA pair, (ii) EF-Sep and (iii) SerB in BL21(DE3) ΔserC using media supplemented with exogenously added nhpSer (2 mM) (system nhpSer-1, Table 1), Rogerson et al. successfully incorporated nhpSer into myoglobin. Although the yields were not reported, it marked the first success in translationally incorporating a stable phosphonate mimic of pSer into proteins.

Later in 2021, Stuber et al. adapted this system to encode nhpSer into Ubiquitin and NEDD8 at Ser65 to eliminate concerns of pSer hydrolysis when isolating phosphorylation-dependent interactors from cell lysates, and knowing that Asp/Glu mimics do not recapitulate pSer function in these systems.123 The encoding of nhpSer was performed in the B834 (DE3) expression strain having an intact serC gene (system nhpSer-2, Table 1), but even with the over-expression of SerB and the addition of 8 mM nhpSer in the media, trace levels of contaminating pSer were reported. The contamination was low enough that they were able to separate the pSer (2) vs nhpSer (20) forms of Ubiquitin, and they successfully showed nhpSer65-Ubiquitin adopted the same phosphorylation induced conformation states as pSer65-Ubiquitin though in a lower proportion, implying some imperfection in nhpSer’s mimicry of pSer in this context.123 Similar purification attempts to separate unmodified vs pSer vs nhpSer65-NEDD8 were not successful. Nevertheless, by leveraging the resistance of nhpSer65-Ubiquitin and nhpSer65-NEDD8 to dephosphorylation, Stuber et al. conducted pulldown assays using them as bait to isolate phosphorylation-dependent interactomes from cell lysates and discover novel interactions such as those between HSP70 and phosphorylated NEDD8.121

3.1.3.3. Biosynthesis of nhpSer improves GCE encoding efficiency and scalability of expressions.

Charged amino acids do not cross the cellular membranes well, explaining why it is important to use ΔserB expression strains that generate high concentrations of intracellular pSer amino acid via biosynthesis for pSer protein expression (Fig. 8B). Thus, to address the low nhpSer encoding efficiency that resulted when nhpSer (20) was supplemented to the media, Zhu et al. sought to engineer its biosynthesis.49 To do this, they built on work by Johannes et al. who predicted that nhpSer (20) was an intermediate along the 10-step Streptomyces rubellomurinus pathway that converted phosphoenolpyruvate (24) into the antimalarial natural product FR-900098.139 With the nhpSer GCE machinery in BL21(DE3) ΔserC cells, they found an optimal set of transcriptional promoters for the five genes producing nhpSer (FrbABCDE) by using the nhpSer-dependent production of sfGFP-150TAG as a read-out for nhpSer biosynthesis. The result was an E. coli expression system biosynthesizing this stable mimic of pSer at concentrations high enough that, when coupled to nhpSer GCE machinery in BL21(DE3) ΔserC cells, accurately encoded it into sfGFP with improved fidelity and an over 40-fold improved efficiency compared to media supplementation (Fig. 8C; system nhpSer-3 Table 1). The use of this system – dubbed “PermaPhos” – also allowed expressions to be scaled more effectively and economically since expensive media supplementation of the nhpSer amino acid, or any other precursor molecule, was no longer required.

With this PermaPhos system, an informative set of head-to-head comparisons were done using sets of equivalent proteins with Ser, pSer and Asp/Glu encoded at the same sites.49 Studies of the proteins Heat Shock Protein B6, SARS-CoV-2 nucleocapsid protein, and 14-3-3ζ all showed that nhpSer (20) imparted similar structure/function changes as pSer (2) compared to the wild-type forms, whereas Asp/Glu did not. Using 31P NMR, they also measured the pKa2 of nhpSer on a protein to be 7.0, confirming that the majority of nhpSer moieties are −2 charged at physiologic pH. Also, Zhu et al. used PermaPhos to study how 14-3-3ζ monomerized by phosphorylation at Ser58 altered its interactome, something not easily done using authentic pSer58-containing protein because the pSer58 modification is hydrolyzed rapidly when exposed to the HEK293 cell extract.49 Among the newly discovered interactions was one between phosphorylated 14-3-3ζ and the E3 ubiquitin ligase adaptor protein Cereblon. PermaPhos encoding of nhpSer (20) was also similarly used by Tugaeva et al. to circumvent extensive hydrolysis of phosphorylated SARS-CoV-2 nucleocapsid protein that occurred during protein expression in E. coli, and this allowed them to perform structure/function analysis on its interactions with 14-3-3.140 Attempts to adopt PermaPhos in an RF1-decificient strain of E. coli B95(DE3) ΔA ΔfabR ΔserC resulted in too much mis-encoding of natural amino acids due to near-cognate suppression, so a truncation free nhpSer GCE system is not yet available.49

3.1.3.4. Encoding nhpSer in HEK293 cells.

In 2018, Beránek et al. demonstrated the orthogonality of the Mma SepRS/Mj Sep-tRNACUACys pair in HEK293 cells, enabling the first encoding of pSer into proteins in a human cell.89 This was achieved through a combination of factors, including: (i) co-expressing of the SepRS(2)/Sep-tRNA(B4)CUA pair with the human EF-Tu ortholog (EF-1α) having equivalent mutations to EF-Sep, (ii) knocking out the endogenous phosphoserine phosphatase (PSPH, the human ortholog of SerB), and (iii) co-expressing a dominant negative eRF1 mutant141 to attenuate endogenous eRF1 activity responsible for translation termination at UAG codons.89 However, the expressed target proteins contained serine instead of pSer due to rapid hydrolysis by endogenous phosphatases. To address this, they established a system for encoding nhpSer in HEK293 cells instead. This involved the expression of the same SepRS(2)/Sep-tRNA(B4)CUA/EF-1α-Sep/eRF1-E55D proteins but in cells devoid of native pSer, achieved by deleting the phosphoserine amino transferase gene (PSAT, the human ortholog of SerC) and overexpressing PSPH, while supplementing the culture media with exogenous nhpSer (20).89 With this system, they showed nhpSer (20) could be encoded into MEK1 at site S218, which activated it to phosphorylate ERK. Greater than 10 mM nhpSer was needed in the media to achieve ~50% accurate nhpSer incorporation. The other 50% of the protein contained Gln at the site of nhpSer encoding, highlighting similar issues of nhpSer bioavailability and near-cognate suppression. Nevertheless, this work provides a foundational system for probing the function of specific phospho-protein variants in their native cellular context.

3.1.4. Cell-free protein expression systems for phosphoserine encoding.

Cell-free protein expression provides a controlled environment for in vitro protein synthesis, allowing precise modulation of SepRS, Sep-tRNACUA, EF-Sep and pSer (2) concentrations while potentially eliminating interfering factors such as RF1 and near-cognate suppressing tRNAs. Oza et al. demonstrated the utility of this approach in 2015 using soluble extracts derived from C321.ΔA cells expressing the SepOTSλ pSer GCE system.142 They successfully transcribed and translated both singly and doubly phosphorylated MEK1 (S218/222), yielding up to 1 mg of protein from a 3 mL reaction. Characterization of these phosphorylated MEK1 variants revealed that both mono-phosphorylated and the di-phosphorylated MEK1 forms were functionally active in phosphorylating its substrate ERK.142 While near-cognate suppression remained a challenge, with roughly half of the MEK1 protein being non-phosphorylated, the flexibility to fine-tune individual translational components in cell-free extract expression systems underscores the potential of these systems for helping advance future technologies. Indeed, by eliminating near-cognate amber suppressing tRNAs in cell extracts lacking RF1, Gan and Fan demonstrated a 5-fold improvement in the fidelity of pSer (2) encoding.143 Recently, a panel of SepRS and Sep-tRNA variants were screened that provided enhanced incorporation of pSer (2) and nhpSer (20) using cell-free expression systems, though the fidelities were not reported.144 Overall, cell-free expression systems for pSer encoding have seen limited usage, though they have helped uncover how the phosphorylation of hydroxyacyl-coenzyme A regulates its import into the mitochondria,145 as well as how phosphorylation can extend the catalytic properties of triosephosphate isomerase beyond the limits of diffusion.146

3.1.5. First fruits: biological insights already gained from pSer GCE technology

With the collective efforts invested into their development, pSer (2)/nhpSer (20) GCE systems are increasingly being used to address important biological questions about phospho-protein function, as can be seen in Table 2 and Fig. 11 that summarize aspects of each of the 73 studies to date that have developed or used these technologies. As adopting these GCE technologies has not always been trivial, motivating factors are often articulated and these largely include: (i) technical challenges associated with using the needed kinase such as incomplete and/or off-target phosphorylation (16 reports),47,49,84,123,140,147-157 (ii) inadequate knowledge about the required kinase (7 reports),98,133,158-162 and (iii) uncertainty about the ability of Asp/Glu to mimic pSer function (31 reports).47,49,50,55,66,84,123,148,149,154,157,159,161,163-180 In at least 13 of the latter cases, GCE encoding of pSer was used because Asp/Glu mutations failed to recapitulate pSer function.47,49,50,55,66,123,149,157,159,171,174,175,178 Although the phosphorylation-dependent pathways and processes studied span many areas – as one would expect – four notably well-represented foci are the areas of (i) ubiquitin and protein degradation, (ii) kinase activation, (iii) 14-3-3/client complexes and (iv) histone function. In the following sections, we briefly summarize the studies in these three areas.

TABLE 2:

Users of GCE systems for encoding phosphoserine

YearRef POI: site(s) encoded a GCE System b
201166 Human MEK1 kinase: 218,222 1c
2012121 Mouse WNK4 kinase: 172,199 2c
201379 Histone H3: 10 3c
2013126 Many peptides 1
201471 GFP: E17 1,3
2014150 α-synuclein: 87,129 1,3
2014181 TRIM9: 76 1
2014182 TPR2A-MESVD peptide complex: ME-pSer-VD 2
201581 GFP: Q157,E17 4c
2015183 Mouse cGAS: 291 1
2015120 Ubiquitin: 65
Nek7 kinase: 195
Myoglobin: 127
5c,nhpSer-1c
201597 Human MEK1 kinase: 218,222 6c
2015142 Human MEK1 kinase: 218,222 4d,6d
201555 Ubiquitin: 20,57,65 4
2015184 OPTN: 473,513 6
201698 Ubiquitin: 7,12,20,57,65 1,3
2016154 OPTN: 177,473,513 2
2016158 Ubiquitin: 20,57,65 5
2016185 MDH: 280 3
2017143 sfGFP: 151 1d
2017159 Ubiquitin: 12,57 1,3
2017163 Cit1p: 462 6
2017164 Human UBE2T: 4 6
2018149 TACC3: 558 5
2018155 Caspase-9: 99,183,195 6
2018165 Histone H3: 10,28 3
201850 Akt1 kinase: 308,473 3
2018166 Human MEK1 kinase: 218 5
2018167 TRIM21: 80 5
2018168 Hsp 72: T66 5
2018186 Akt1 kinase: 473 3
2018187 Many peptides 6
2018188 sfGFP: 2 3
2018189 MDH: 280 3
2018190 Ubiquitin: 65
Sperm whale myoglobin: D127
5
2018191 DCNL5: 41 2
2018192 TAB1: 423 5
201984 STING: 358,366
BAD: 112,136,155
7c
2019152 Human WNK1: 382 6
2019157 Rat nNOS holoprotein: 1412 6
2019160 PKL kinase: 12,113 6
2019170 Human HADH: 13 6d
2019171 ACBD6: 106,108 5
2020131 SF3 motif of Rab GTPase: 111 5
2020151 MEI-1 subunit of Katanin: 90,92,113,137 5
2020172 Rat GlyT2: 157 6
2020173 Human STARD3 FFAT motif: 209 6
2020175 Mouse cGAS: 420 6
2020176 MDM2: 429 4
2020193 Histone H2A: T120 6
2020194 Cdc8: 125 6
2020195 Rpn1: 361 6
2021132 Ubiquitin: E18,T22,T66,H68
GB1: T11,K10,A24,K28
Hsp90-N: 36,D40
6
2021123 Ubiquitin: 65
NEDD8: 65
nhpSer-2c
2021148 DEP domain: 435 5
202147 Akt 1 kinase: 473 3,4
2021156 PSD-95: many sites 5
2021177 PDZ2 domain of NHERF1: 162 5
2021178 WNK1 kinase: 382,378 6
2021179 PDZ2 domain of NHERF1: 162 5
2021196 Hst2: 320,324 5
2021197 CENP-A: 68 6
202288 Nucleocapsid protein of SARS-CoV-2: 188 5,7
2022162 TPI: 20 4d,6d
2022198 Many peptides 8
202349 Nucleocapsid protein of SARS-CoV-2: 188,206
HSPB6: 16
14-3-3ζ: 58
7,nhpSer-3c
2023122 GFP: E17 8c
2023133 Mouse amelogenin: 16 7
2023140 Nucleocapsid protein of SARS-CoV-2: 197,T205 7,nhpSer-3
2023180 HSPB6/HSP20: 16 7
2023199 MITF-A: 5 8
2023200 METTL3: 43 6
2023144 MEK1: 217,221 Variedd
a

WT sites are serine unless otherwise indicated

b

System numbers represent the corresponding GCE system found in Table 1

c

GCE system was created in this report

d

Cell-free protein synthesis

Figure 11. Time course of publications using pSer GCE technologies.

Figure 11.

Shown is the time course of the numbers of peer-reviewed publications (see Table 2) using pSer GCE technologies (systems numbered as in Table 1) since the 2011 creation of the original SepOTS system (i.e. “pSer-1”). Publications using a given pSer system are represented by color-coded segments (“SepOTS” systems in orange hues, “pKW2” systems in blue hues per the key provided in the figure), with each segment also labeled with the number corresponding to the system used.

3.1.5.1. Regulation of ubiquitin function and protein degradation systems.

The largest group of studies using pSer GCE have focused on uncovering the roles of phosphorylation in regulating ubiquitin function and protein degradation pathways. For example, GCE was used to encode pSer (2) into ubiquitin at sites S12, S20, S57 and S65 by several groups, collectively revealing how the location of phosphorylation influenced (i) the extent of Parkin (an E3 ubiquitin ligase) activation,55,98,159,184 (ii) the linkage patterns of poly-ubiquitination,184 and (iii) ubiquitin susceptibility to deubiquitinases.158 Importantly, at the time only the kinase for phosphorylating S65 was known (PINK1),201 and in some scenarios it led to incomplete phosphorylation, making GCE essential for these studies. Of note, the first crystal structure of a phosphorylated protein synthesized with GCE, reported in 2016, showed how ubiquitin’s surface charge distribution was perturbed by phosphorylation at site S20.158

In later work NEDD8 (a ubiquitin-like protein) was shown to activate Parkin when phosphorylated at site S65, just as pSer65-ubiquitin did; and NMR analyses of 13C/15N-labeled pSer65-NEDD8 revealed it adopted a similar set of phosphorylation-dependent conformations as did pSer65-Ub.123 Then by using nhpSer65-NEDD8 as bait in a pull-down experiment, they were able to hinder its dephosphorylation in cell-lysates and isolate proteins that specifically bound phosphorylated NEDD8, including HSP70.123 In another study, the phosphorylation at S473 of the autophagy adaptor optineurin dramatically increased its ability to bind K48- and K63-linked poly-ubiquitinated chains at the surface of damaged mitochondria to recruit the autophagosome and activate mitophagy.154,184 These works showcase the power of GCE approaches to reveal specific ways that phosphorylation of ubiquitin and its receptor proteins work to coordinate protein degradation.

GCE has also be instrumental in revealing how phosphorylation regulates protein degradation at the level of E2/E3 ligase function. For example, GCE was used to solve the crystal structure of the E3 ligase MDM2 phosphorylated at Ser429, and to show how this phosphorylation promoted MDM2 binding to an E2-ubiqitin ligase, ultimately leading to MDM2 auto-degradation.176 In separate work, TRIM9 phosphorylated at Ser79 was able to bind to and block the activity of the E2/E3 ligase β-TrCP-SCF, preventing it from ubiquitinating target proteins IκBα and p100.181 Also, pSer GCE aided the discovery of several pSer-degrons that recruit target proteins to E2/E3 ligase complexes in a phosphorylation-dependent manner: the target proteins discovered include centromeric protein A (CENP-A), master transcriptional regulators TFE3 and MITF, and potentially 14-3-3ζ.49,199 In the case of 14-3-3ζ, nhpSer encoded at Ser58 pulled down the E3 ligase adaptor protein Cereblon from cell-extracts, implying that phosphorylation serves to promote ubiquitination of 14-3-3 as well as the bound clients of 14-3-3ζ.49

3.1.5.2. Mechanisms of kinase activation.

Many kinases are regulated by phosphorylation, often at multiple sites, and GCE can greatly aid in the challenging process of deconvoluting how the sites contribute to their activation. Using cell-free pSer GCE, for example, it was shown that single site phosphorylation of MEK1 at either S218 or S222 was sufficient for activation, with the latter variant being nearly as active as the doubly phosphorylated MEK1 form.142 Akt1 (also known as Protein Kinase B) is also activated by phosphorylation at two sites: Thr308 and Ser473. Expression of activated Akt1 in insect cells results in a mixture of singly and doubly phosphorylated species,202 and so an E. coli expression system to make both the singly and doubly phosphorylated variants was devised which used co-expression of PDK1 kinase to install pThr at residue 308 and GCE to encode pSer at residue 473.147 With this system each singly phosphorylated species could also easily be generated without the need to use kinase-null mutants (e.g. Ser/Thr to Ala). The results showed that pThr308 was the primary driver of Akt1 activation, though the doubly phosphorylated form was the most active.147 Akt1 activation by Thr308 phosphorylation could not be recapitulated with Asp/Glu, and the unphosphorylated form of Akt1 could not be mimicked with the T308A mutation.147 Akt1 substrate specificity also changed depending on which site or if both sites were phosphorylated,47 with for instance, the Akt-1/2 inhibitor VIII being highly effective against the pThr308 variant, whereas phosphorylation at Ser473 – to make the doubly phosphorylated form – attenuated inhibition 4-fold.186

Another study illustrates how a clever GCE-aided generation of simple but effective kinase networks in E. coli can make it possible to produce phospho-proteins that would be inaccessible using only current GCE technologies or other approaches alone. The SPAK/OSR1 kinases have an activating pThr site that could in theory be generated by co-expressing the WNK1 kinase. However, the complication is that WNK1 itself required phosphorylation to be active, and phospho-mimetics are not functional in this regard. Therefore, to phosphorylate SPAC/OSR1 kinases, a two-step co-expression system was created in which GCE was used to encoded pSer into WNK1 at S382 (or S378 and S382), and this phospho-WNK1 then phosphorylated the SPAK/OSR1 at the required threonine site.152 With pThr-activated SPAK and OSR1 kinases in hand, substrate profiling was possible for the first time, and inhibitors against SPAK could be screened.178

3.1.5.3. Studies of 14-3-3 function and client regulation.

Common pSer/pThr binding domains important in eukaryotic biology include 14-3-3 proteins, WW domains, Polo-box domains, WD40 repeats, BRCA1 carboxy-terminal (BRCT) domains and forkhead-associated (FHA) domains.203 Among these domains, the 14-3-3 proteins constitute a family of essential hub proteins that collectively bind many hundreds of client proteins when they are phosphorylated at specific pSer/pThr motifs,204-206 and it is well established that Asp/Glu are poor mimics in this system, unable to promote client/14-3-3 complexation.207-209 Even though GCE has thus far only been adopted in a handful of cases to study phospho-dependent client complexes with 14-3-3 proteins, it has already brought advances. In one example, pSer GCE was employed to successfully express phosphorylated Bcl2-associated agonist of cell death (BAD),84 a protein that in its native (non-phosphorylated form) is notoriously unstable and does not express well in E. coli.210 By co-expressing singly (pSer136, mouse numbering), doubly (pSer112/136) and triply (pSer112/136/155) phosphorylated BAD along with its stabilizing 14-3-3β binding partner protein multi-milligram quantities of complex were obtained.84 This was an interesting case in which encoding pSer into a protein with GCE actually increased its expression yields.

In other work, GCE was used to pinpoint the exact locations of phosphorylation on the histone deacetylase Hst2 that enabled stable complexation with BMH1 (a yeast ortholog of 14-3-3), with the results implying that there is crosstalk between histone lysine acetylation and serine phosphorylation PTMs.196 Similarly, in studies of the SARS-CoV-2 nucleocapsid protein, nhpSer (20) incorporation at sites S197 and T205 was shown to be sufficient for recruiting and binding 14-3-3, revealing how viral proteins could hijack host 14-3-3 proteins to enhance viral survival.140 Implementing GCE was essential for these nucleocapsid protein studies for two reasons: (i) since kinases phosphorylated an extensive number of off-target sites, site-specific phosphorylation was needed to localize the exact nucleocapsid phospho-sites needed for 14-3-3 interactions; and (ii) high susceptibility of these nucleocapsid pSer residues to hydrolysis during expression and purification necessitated the use of nhpSer encoding to produce homogenously modified protein.140 Later work confirmed that phospho-client protein containing pSer and nhpSer bound to 14-3-3 with indistinguishable affinities,49 highlighting the utility the nhpSer GCE systems for studying client/14-3-3 interactions in cellular-like environments. Lastly, GCE was used to install pSer (2) and nhpSer (20) into site Ser58 of 14-3-3ζ, confirming this modification caused 14-3-3ζ to shift from being a dimer to a monomer, and to release one subset of clients and engage in new interaction with others, including the ubiquitin ligase adaptor protein Cereblon.49

An innovative genomic-level approach to identify new phosphorylation-dependent 14-3-3 binding partners involved screening peptide libraries made with GCE.187 This study took advantage of advances in high-throughput gene synthesis, and an adaptation of the split-mCherry fluorescent protein reporter system designed so that it would report formation of phosphorylation dependent protein-protein interactions with 14-3-3.182 First, a library of oligonucleotides was synthesized encoding a series of peptides representing all 110,139 sites of known serine phosphorylation in the human proteome.187 Each oligonucleotide contained a TAG codon in the middle to direct pSer encoding and was fused to a gene encoding one half of the split-mCherry reporter protein. Then, by fusing 14-3-3β to the second half of split mCherry, cells expressing pSer-peptides that bound 14-3-3β could be isolated by fluorescence sorting. As a key control, by replacing the GCE machinery with a system to encode Ser at UAG codons in the same pool of peptides, truly phosphorylation-dependent interactions could be distinguished from those driven by other factors.187 The isolated peptide sequences conformed well with established 14-3-3 pSer binding motifs, but also revealed nine novel phospho-client proteins that had not yet been established as 14-3-3 binders, including Repressor of RNA polymerase III transcription MAF1 and Spermatogenesis-associated protein 19.187

3.1.5.4. Studies of histone regulation by phosphorylation.

Histone proteins play critical roles in packaging DNA into nucleosomes and regulating gene transcription, and their function is controlled by a variety of PTMs, including phosphorylation at Ser, Thr, Tyr and His residues.211 With four different histone chains (H2A, H2B, H3 and H4) each having distinct sets of phosphorylation sites, studying the functional consequences of histone phosphorylation has not been trivial.212 As mentioned briefly in section 3.1.1.3, Lee et al.79 encoded pSer into a histone H3 and showed that the SAGA-mediated acetylation at residues K9, K14, K18, and K23 was inhibited when H3 was phosphorylated at Ser10. However, this inhibition of SAGA acetyl-transferase activity was relieved when phosphorylated H3 was assembled into octamers with H2A, H2B and H4. Interestingly, SAGA-mediated acetylation of H3-pSer10 was enhanced nearly 3-fold at K14, K18, and K23 compared to unmodified H3 when the histone octamer was assembled into nucleosomal arrays. These results highlight the complex nature of phosphorylation and acetylation cross-talk in regulating histone function.

In other work, Ahn et al.165 in 2018 showed through in vitro assays that histone phosphorylation by mitogen- and stress-activated kinase 1 (MSK1) stimulated transcription at p53-dependent gene promoters. To pinpoint the phospho-sites involved, they used GCE to encode pSer at sites Ser10 and Ser28 of histone H3, and showed that while both sites enhanced p53-dependent gene transcription, phosphorylation at Ser10 had a greater impact. In 2020, Zhang et al.193 sought to understand how cells achieve timely separation of sister DNA strands at mitotic centromeres via the activity of topoisomerase IIα (TOP2A). They revealed that phosphorylation at residue Thr120 of histone H2A by the Bub1 kinase served to recruit TOP2A to centromeres in cells. At the time, pThr GCE tools were not readily available and so to confirm this interaction in vitro, they installed pSer in place of Thr120 and by pull-down assays confirmed that TOP2A interacted more strongly with H2A-pSer120 than with unmodified H2A. With more robust pThr GCE tools available now (section 3.2), it would be interesting to assess the extent to which this pSer installation faithfully mimics the effects of H2A Thr120 phosphorylation.

3.1.6. Practical considerations for pSer/nhpSer GCE

In Table 1, we delineate 8 systems designed for encoding pSer and 3 for nhpSer within E. coli. Many of these can be viewed as key stepping stones in the evolution of pSer/nhpSer GCE technologies (e.g. pSer systems 1-4), with a select few of these having been used for addressing biologically relevant questions by non-GCE experts (e.g. pSer systems 5-7). As all of these systems share the same base pSer GCE chassis (Mma SepRS, Mj Sep-tRNACUA, EF-Sep), their distinction lies in which variants are used, how the components are expressed and which E. coli strain is employed – crucial factors impacting pSer/nhpSer encoding. Extracting comparisons from the literature is not always straightforward due to varying reporter protein constructs, TAG site placements, and expression methodologies across different studies. However, the effective application of pSer/nhpSer GCE technologies to address biological questions hinges on selecting the most suitable system for the intended experiment. In this section, we outline a few key considerations to assist users in navigating these choices.

3.1.6.1. The importance of controls: phosphoamino acid encoding into sfGFP reporter proteins.

Translational encoding of phosphoamino acids requires co-expression of many different components simultaneously for efficient and faithful installation of ncAAs. Media, temperature, aeration and induction methods are among many of the variables that can affect the success of target protein expression, and these often need to be optimized to identify conditions in which the GCE machinery functions appropriately and the phosphorylated target protein expresses well and is stable. Before attempting to encode any phosphoamino acid into a protein of interest, we emphasize the importance of first confirming the GCE tools are functioning properly using established conditions by evaluating sfGFP-150TAG expression efficiency (using in-cell fluorescence) and fidelity (via Phos-tag electrophoresis or whole-protein mass spectrometry of purified protein). Afterward, expression of pSer-target proteins can be optimized, and we recommend this also be done side-by-side with sfGFP-150TAG control expressions; this way if pSer encoding into the target proteins is not successful, it can be easily determined whether the issue resides with the GCE system being incompatible in those expression conditions, or with the target protein itself. Multiple rounds of optimizing target protein construct design and expression conditions may be needed to express phosphorylated target protein in suitable quantity and quality.

3.1.6.2. Multi-plasmid expression chassis for pSer GCE.

Most E. coli GCE systems adopt a two-plasmid configuration, where the GCE components are expressed on one plasmid, and the protein(s) of interest are expressed from a second plasmid. These plasmids must confer distinct antibiotic resistances and possess different origins of replication to ensure their stable propagation in E. coli.213 An important aspect of current pSer/nhpSer GCE systems (and pThr as well, discussed in Section 3.2) is that machinery component expression was optimized on plasmids featuring the pBR322 origin of replication, necessitating the expression of the target protein from a plasmid with an alternate origin (e.g., p15a, CloDF, or RSF). Commonly used E. coli target protein expression plasmids like pET, pGEX, and pBAD carry the pBR322 origin, rendering them incompatible with pSer/nhpSer GCE machineries. This is in contrast to most other GCE platforms in which machinery components are expressed from a p15a or CDF origin of replication plasmid so that the protein of interest can be expressed from standard pBR322 origin plasmids like pET, pGEX, and pBAD. Therefore, a key first step in adopting current pSer/nhpSer expression systems is typically cloning target gene(s) onto an expression plasmid compatible with the machinery vector.

3.1.6.3. pSer expression hosts.

All E. coli pSer GCE expression hosts have the ΔserB genomic mutation (Fig. 8B), but they vary in whether RF1 has been deleted. BL21(DE3) ΔserB is a prevalent choice for pSer protein expression due to its robust fitness and the benefits associated with a BL21 derivative designed for T7 transcriptional promoter-based protein overexpression. Having an intact RF1 gene however means target pSer-proteins will be co-expressed with truncated protein that must be purified away. Alternatively, the B95(DE3) ΔA ΔfabR ΔserB strain is a healthy derivative of BL21(DE3) which lacks RF1. With this expression host, premature truncation is mitigated or eliminated so that users can leverage the benefits of BL21 strains as well as N-terminal solubility and affinity purification tags (e.g. SUMO, GST, MBP), while also often achieving higher yields of full-length protein. However, BL21 strains are recA+ and may lose expression fitness with multiple passages due to plasmid recombination, or inactivation of the T7 polymerase,214 limiting their use in systems-level biology applications or screening of libraries. In these cases, the genomically recoded and RF1-deficient strain C321.ΔA and its derivatives (e.g rEcoliXPS) are better suited.

3.1.6.4. Pairing of pSer GCE machinery, target protein plasmid and expression host.

For high-level pSer protein expression where high yields and faithful pSer encoding is paramount, we recommend using the pKW2-EFSep machinery. It has been well-vetted for its ability to ensure homogenous pSer encoding both in the RF+ BL21(DE3) ΔserB strain (system pSer-5 in Table 1) and the RF B95(DE3) ΔA ΔfabR ΔserB strain (system pSer-7 in Table 1) when it is paired with a T7-based target protein expression plasmid with a p15a origin. The pKW2-EFSep machinery has not been assessed extensively in C321.ΔA strains, so for applications demanding stability in systems-biology studies, where maximal target protein expression and clean pSer encoding is not critical, recent evidence supports the optimized pSerOTS-C1* (V70) platform (system pSer-8 in Table 1) when paired with C321.ΔA strains as the most effective available system. A key advantage of the pSerOTS-C1* (V70) system over SepOTSλ is it having only 1 copy of the Sep-tRNACUA gene, making it more stable than the SepOTSλ machinery plasmid having 4 Sep-tRNACUA gene copies. We only found one instance in which the pKW2-EFSep machinery plasmid was used in the C321.ΔA strain120 and while pSer encoding fidelity wasn’t reported we suspect that, given its higher efficiency compared to the SepOTS systems, the pKW2 system would have less mis-encoding caused by near cognate suppression.

3.1.6.5. Practical considerations for encoding nhpSer in E. coli.

GCE systems to encode nhpSer in E. coli have not yet been used extensively. Factors limiting its use could include low encoding efficiency, the high cost of supplementing media with nhpSer amino acid, and concerns about nhpSer mimicry of pSer. As discussed above, the work by Zhu et al.49 published last year addressed these technical challenges with development of the PermaPhos system (system nhpSer-3 in Table 1). In the PermaPhos system, three plasmids are required: one for expressing the nhpSer GCE machinery, a second for expressing proteins involved in nhpSer biosynthesis, and a third for the target protein. The same p15a origin pRBC plasmid used for truncation-free expression of target proteins with pSer (system pSer-7 in Table 1) can be used with the PermaPhos system. Important to remember is that for faithful nhpSer encoding, an expression host with the ΔserC mutation and a machinery plasmid that overexpresses SerB are important to eliminate intracellular pSer that would compete for nhpSer encoding (Fig. 8C).120 While the BL21(DE3) ΔserC expression host grows robustly, prematurely truncated protein needs to be purified away from full-length, nhpSer-containing protein.

3.1.6.6. Practical considerations for encodins nhpSer in mammalian cells.

For nhpSer encoding in HEK293 cells,89 an analogous two plasmid system was used: one expresses the GCE machinery components (EF1α-Sep, PSPH, SepRS(2), and eRF1 E55D) and the other expresses the protein of interest, while both express four copies of Sep-tRNA(B4)CUA. Just as with the E. coli nhpSer (20) encoding systems utilizing BL21(DE3) with the ΔserC mutation to avoid mis-encoding with pSer (2), the HEK293 system employed a cell line with the mammalian homolog of SerC, phosphoserine amino-transferase (PSAT), knocked out. It is not clear why this system has not yet been more widely adopted but reasons could include low expression efficiency of target protein, low bioavailability of media-supplemented nhpSer and mis-encoding of natural amino acids. While site-specific nhpSer encoding in eukaryotic cells can have many transformative applications, caution should be exercised when doing so to study cell physiology, as the introduction of GCE components and the alteration of metabolic pathways may introduce artifacts and off-target cellular effects that impact the results. For example, GCE machinery may suppress endogenous TAG codons such that translation continues past the intended site of termination causing some natural host proteins to have artifactual C-termini. Genomically recoded eukaryotic cells lacking endogenous TAG codons have not yet been produced to mitigate this concern, however developing systems that encode phosphoamino acids at naturally unused (e.g quadruplet) codons could be helpful (discussed in section 3.1.7).

3.1.6.7. Availability of pSer/nhpSer protein expression reagents and protocols.

These pSer/nhpSer GCE technologies can only be transformative for the research community if they are widely available. Thankfully, researchers who have developed the above-described E. coli pSer (2) and nhpSer (20) systems have in many cases graciously deposited the relevant GCE plasmids and strains in the non-profit repository Addgene as described in the footnotes of Table 1. Also, protocols for expressing pSer-proteins with the pKW2215 and SepOTSλ216 systems, as well for nhpSer-proteins using the PermaPhos kit,217 have been published. And finally, in addition to the GCE reagents, libraries for expressing and screening phospho-peptides for protein-protein interactions are also available (Addgene #111704-111708 and #188526-188530).

3.1.7. Future opportunities in pSer and nhpSer GCE technology development.

pSer/nhpSer GCE systems have been slow to catch on, understandably so since high efficiency/fidelity systems and optimized protocols have only recently become available, making them still largely untapped tools to probe biological questions that are otherwise unapproachable. With these E. coli systems now achieving >50% efficiency and >90% fidelity for single site encoding, and installation of as many as 5 pSer residues into a single protein proven feasible, pSer GCE systems are on par with some of the top GCE systems available. Still needed, however, are accessible strategies to encode pSer/nhpSer into proteins in other organisms, both to expand the number of proteins that can be expressed and to study phospho-protein function in biologically relevant contexts. The Mma SepRS/Mj Sep-tRNACUA pair was shown orthogonal in another bacterium (Salmonella)185 and eukaryotic cells (HEK293)89 and this provides good reason to believe it can be transplanted into other model recombinant expression hosts as well, including other bacteria (e.g. Vibrio natriegens), yeast (e.g. Saccharomyces cerevisiae and Pichia pastoris) or insect cells (e.g. Sf9). Adapting a pSer GCE system in V. natriegens could provide an attractive alternative with its 10-minute doubling time and ability to recombinantly express proteins that E. coli cannot.218 In the case of eukaryotic expression hosts, the encoding of non-hydrolyzable mimics is likely the only tractable strategy due to extensive phosphatase activity. Improvements in the reported nhpSer system for HEK293 cells will be important, as up to 20 mM nhpSer was needed to reach only 50% encoding fidelity.89 Strategies that improve nhpSer bioavailability would help in this regard.

Regarding non-hydrolyzable mimics of pSer, nhpSer (20) is clearly much more effective than either Asp or Glu, and recent advances in nhpSer encoding should lead to a phasing out of the use of Asp/Glu mimics for proteins produced in E. coli. It is not yet clear how much better the fluorinated derivates of nhpSer (21, 22) would be than nhpSer (20), but it is reasonable to assume there would be instances where they would be and thus creating systems to encode them is a potential area of development. With no clear leads in terms of biosynthetic enzymes to take advantage of though, an important hurdle to overcome for fluorinated nhpSer encoding is how best to make them bioavailable. Possible ways to improve uptake could involve chemically masking the phosphonate -2 charge with a group that is spontaneously removed intracellularly,219 or identifying a mechanism that promotes their active import.220

Finally, worth noting is that current systems to install pSer/nhpSer (as well as pThr and pTyr, discussed below) do so at all available TAG codons, including those found in endogenous genes that serve as natural translational stop signals.182 To achieve “true site-specific” encoding of phosphoamino acids requires they be incorporated at a codon not used anywhere else in the genome, and for E. coli this has already been demonstrated with the genomically reduced C321 strain.97 The recently engineered E. coli strain Syn61 that uses 61 codons (lacking TAG and serine codons TCG and TCA) for protein synthesis provides two additional unique codons available for specific phospho-ncAA encoding.221 For eukaryotic expression hosts, genomically reduced organisms are not yet available. A generalizable system that encodes phospho-ncAAs at codons that are neither available for natural amino acids nor translational termination in eukaryotic hosts would be of great value, as it would minimize inadvertent changes to the cell’s physiology that could cause artifacts. In this sense, encoding systems that encode phospho-ncAAs at quadruplet codons222,223, or that target protein synthesis to orthogonal ribosomes224 could prove highly valuable. Developing such systems for pSer would require altering the anticodon of Sep-tRNA, which in turn would require engineering SepRS to recognize these changes to ensure efficient amino-acylation of Sep-tRNA with pSer.

3.2. GCE systems for phosphothreonine and its mimics

Threonine phosphorylation plays key roles in many cellular processes.17 The functional differences between pSer (2) and pThr (4), and the circumstances in which they are interchangeable, are not well-defined, though sufficient evidence is emerging that pSer and pThr have discrete functions despite differing by only a methyl group. For example, FHA (“forkhead-associated”) domains are well established as phospho-dependent binding modules but are unique in binding exclusively to pThr.225 Further, threonine is phosphorylated by kinases more slowly than pSer, and pThr is more rapidly degraded by phosphatases, explaining why pThr is much less commonly observed than pSer and implying that pThr is a more transient modification than pSer.11 Recently, Pandey et al. concluded that the side chain of pThr adopts a more conformationally limited and rigid structure than pSer that mimics the backbone cyclization of prolines.226 Yet, in the few cases where pSer (2) (and nhpSer (20)) were encoded at threonine phosphorylation sites as a pThr (4) substitute, they provided better mimicry than Asp/Glu.140,168,193 In 2017 the first pThr GCE system was reported in E. coli73 but overall, less effort has gone into pThr systems development compared to pSer, and the pThr systems have not yet seen as much use in addressing biological questions. Also, to date, no systems have been reported for encoding pThr in eukaryotic cells or for encoding non-hydrolyzable mimics of pThr (25-27, Fig. 12A) in any cell system. In this section, we will highlight the developments for encoding pThr (4) in E. coli recombinant systems and then discuss opportunities ripe for future work.

Figure 12. pThr, its nhpThr mimics and metabolic alterations of specialized E. coli strains used for pThr GCE systems.

Figure 12.

(A) Shown are the side-chain structures of pThr (left, blue box) and three nhpThr mimics (right, gray box) which have the bridging γ-oxygen replaced with either CH2, CHF, or CF2 (highlighted red). (B) In the “ΔserC ΔycdX” and the “ΔserC ΔycdX ΔpphA” E. coli strains used for pThr incorporation (see Table 3), to minimize competing pSer (upper pathway), intracellular pSer is decreased by knocking out SerC (red “X”). In addition, pThr is produced (lower pathway) using the threonine kinase PduX (from S. enterica) encoded by either the pUC-pThr plasmid or pThrOTS-Zeus plasmid (see Table 3). pThr is then incorporated by GCE machinery into the protein of interest (POI). POI dephosphorylation is minimized by knocking out endogenous phosphatases (red X on arrow going up from POI), either YcdX alone (pThr-1 system in Table 3), or both YcdX and PphA (pThr-2 and pThr-3 systems in Table 3). No GCE systems yet exist for F1-nhp-Thr or F2-nhpThr.

3.2.1. Evolving the pSer pKW2 GCE system to encode pThr (4).

A major challenge with developing a pThr GCE system is avoiding mis-incorporation of pSer given their similarities, and that pThr intracellular concentrations are very low to undetectable, while pSer is a natural metabolite at relatively high concentrations.71 Also, since pThr is highly charged, it does not get into cells well when added to the media. One source of encouragement that a system could be developed that distinguishes pThr from pSer is that many natural amino-acyl RSs distinguish between amino acids that differ by only a single methyl group, and indeed the SepRS itself is specific for pSer over pThr by a factor of 104.70

In 2018, the first instance of a GCE pThr system was reported by Zhang et al., and it was based on evolving the pSer pKW2 GCE machinery to selectively encode pThr instead of pSer.73 First, they overcame the issue of low bioavailability by developing a method to biosynthesize the pThr amino acid inside E. coli. Serendipitously, some bacteria including Salmonella enterica harbor a threonine kinase (PduX) that converts free threonine into pThr during the de novo biosynthesis of vitamin B12,228 and so expression of PduX in E. coli elevated pThr intracellular concentrations to levels sufficient for developing a pThr GCE system.73 Next, the authors noted that when the Sep-tRNACUA(B4) developed for pSer encoding was expressed in the absence of SepRS, it was mis-charged at low levels by the E. coli RSs for glycine and valine. Mitigating this low-level background encoding was needed in order to effectively screen for SepRS mutants with enhanced specificity for pThr, so a library of acceptor stem tRNACUA(B4) variants was generated. The acceptor stem mutations that were selected in the new “Sep-tRNAv2.0CUA” variant (Table 3, Fig. 9) disrupted mis-acylation by endogenous Gly- and Val-RSs by 4000-fold, but not cognate-acylation by SepRS2.

TABLE 3:

E. coli GCE systems for encoding phosphothreonine, with relevant Addgene IDsa

Year
(ref)
ncAA System Plasmid 1 Plasmid 2 E. coli Expression host
Name
(ori)
ORF 1
Promoter/Gene(s)
ORF 2
Promoter/Gene
ORF 3
Promoter/Gene
Orf 4
Promoter/Gene
Name
(ori)
ORF 1
Promoter/Gene
2018 [73] pThr pThr-1b pUC-pThr (pUC) glnS pThrRSc tac EFSepd Oxb20 PduX lpp Sep-tRNA v2e pNHD (p15a) T7 POI BL21(DE3) ΔserC ΔycdX
2022 [227] pThr-2 pThrOTS-Zeus (pBR322) trc f pThrRSc/EFSep21d/PduX proK Sep-tRNA v2e Varies (pRSF/p15a) araC/PLtetO-1 POI BL21(DE3) ΔpphA ΔserC ΔycdX
2022 [227] pThr-3 pThrOTS-Zeus (pBR322) trc f pThrRSc/EFSep21d/PduX proK Sep-tRNA v2 e Varies (pRSF/p15a) araC/PLtetO-1 POI C321 ΔA ΔpphA ΔserC ΔycdX ΔrelA
a

For components available from Addgene, IDs are provided in a footnote for that GCE System number

b

Component Addgene IDs are 173899 (plasmid 1), 174075 and 174076 (control sfGFP forms and for POI). The sfGFP/POI plasmids are not the ones originally used with this system, but work well with it. Competent cells of BL21(DE3) ΔserC ΔycdX can be purchased commercially from Amid Biosciences (Catalog # BLPCY-201)

c

pThrRS2 mutations (SepRS2 + G320A, L321Y)

d

EFSep mutations (F219Y, D217G, T229S, E216N, H67R, N274W); EFSep21 mutations (F219Y, D217G, T229S, E216V, H67R, N274W)

e

tRNA-v2 mutations (Sep-tRNA(B4) + C2U, G4C, G6C, C67G, C69G, G71A)

f

multi-cistronic expression cassette

Next, Zhang et al.73 developed a novel screening approach to select – from a library varying 6 active site residues – SepRS2 variants with selectivity for pThr over pSer. Using an E. coli ΔserC strain devoid of endogenous pSer, parallel positive selections were performed in the absence of pThr amino acid and in the presence of pThr (created by expressing the PduX enzyme). The relative abundance of SepRS2 variants that survived the parallel positive and negative selections were evaluated via deep sequencing. Then, a set of 16 variants that were heavily enriched in the positive selections compared to negative selections were characterized for their ability to selectively encode pThr. Among these, a SepRS(2) variant with G320A and L321Y mutations (Fig. 13), now dubbed pThrRS, yielded quantitative site-specific incorporation of pThr into sfGFP-TAG when co-expressed with tRNAv2.0CUA, EF-Sep and PduX in the BL21(DE) ΔserC host.

Figure 13. Rationale behind engineering SepRS for pThr encoding.

Figure 13.

Structure of A. fulgidus SepRS, as in Figure 10B and 10C, but focused on the amino acid binding pocket and highlighting the five A. fulgidus residues (stick models with atom coloring and green carbons) equivalent to those varied in the M. maripaludis SepRS library used for selecting a pThr-specific RS variant.73 Also shown is a bound pSer amino acid (stick model with orange carbon atoms) and select hydrogen bonds (yellow dashes) involving side chains included in the library. The table lists the library sites and the mutations in the selected M. maripaludis pThrRS.

When evaluating the generality of their approach by encoding pThr in ubiquitin at sites T12 and T66, a large proportion of the ubiquitin was purified with Thr.73 Hypothesizing that the proteins were being dephosphorylated during expression, the authors knocked out each 5 putative endogenous E. coli phosphatases and found that deletion of the phosphatase YcdX allowed the purification of homogenously phosphorylated pThr12-ubiquitin and pThr66-ubiquitin. Collectively, co-expression of pThrRS, tRNAv2.0CUA, EFSep and PduX in BL21(DE3) ΔserC ΔycdX (system pThr-1 in Table 3; Fig. 12B) allowed them to generate milligram quantities of homogeneously phosphorylated pThr12-ubiquitin and they solved its crystal structure at near 1 Å resolution. The authors further showed pThr could be encoded in the activation loop of cyclin-dependent kinase 2 (Cdk2) to produce catalytically active Cdk2 that was able to phosphorylate Histone H1.

3.2.2. Optimizing a pThr GCE system for reconstructing phospho-regulatory networks.

Seeking to improve access to pThr-containing proteins, Moen et al.227 altered the expression plasmid architecture of the pThrRS/tRNAv2.0CUA pair developed by Zhang et al.73, and tested pThr (4) encoding in E. coli strains derived from either C321.ΔA or BL21(DE3) with various additional genes knocked out that might enhance the proportion of phosphorylated protein expressed. The resulting machinery plasmid, dubbed pThrOTSZeus, expressed the pThrRS, EF-Sep21 and PduX under control of a poly-cistronic trc promoter, and the tRNAv2.0CUA via a proK promoter. Overall, the highest pThr encoding fidelity was observed when a ΔpphA chromosomal mutation was introduced into the ΔserC ΔycdX background previously identified by Zhang et al.73, where pphA is a putative phosphatase homologous to the promiscuous λ-phosphatase.227 The proportion of expressed sfGFP containing pThr from the BL21(DE3) ΔserC ΔycdX ΔpphA (i.e. triple knockout) strain (system pThr-2 in Table 3) or the C321.ΔA triple knockout strain varied depending on the evaluation technique, with MS methods reporting ~70% and Phos-tag indicating ~40% pThr content from both expression hosts. Moen et al. observed higher fidelity pThr encoding in the BL21(DE3) triple knockout strain than with in the C321.ΔA triple knockout strain. The latter had higher rates of Gly and Gln mis-incorporation consistent with near-cognate suppression in RF1-deficient cell lines.

Moen et al. then sought to demonstrate how these systems could be used to screen for pThr-dependent protein-protein interactions of 14-3-3β.227 They expressed a library of pThr-containing peptides encompassing the 57,536 annotated pThr sites in the human proteome. Similar to prior work identifying pSer-dependent peptide-protein interactions,187 these pThr-peptides were fused to one half of a split-mCherry protein, while the other half of mCherry was fused to 14-3-3β. Unfortunately, neither the BL21(DE3) or C321.ΔA triple knock out strains with the pThrOTSZeus was able to reproducibly enrich phospho-dependent 14-3-3β/pThr-peptide interactions. But by knocking out relA (a tRNA-mediated stringent response effector) in the C321.ΔA triple knock out strain (making pThr-system 3 in Table 3), they successfully enriched of ~200 peptides sequences that conformed to the established phospho-peptide binding motif recognized by 14-3-3 proteins.

Lastly, in a serendipitous development, Sep-tRNACUA modified at the acceptor stem allowed it to be charged by the endogenous Ala-RS for facile encoding of Ala at UAG codons. Moen et al. then encoded either pThr or Ala into sites Thr383 and Thr387 of the activation loop of CHK2 kinase in response to amber stop codons.227 Interestingly, only the pThr383/Thr387 variant phosphorylated its substrate CDC25C. The doubly phosphorylated pThr383/pThr387 and pThr383/Ala387 were not active, revealing important mechanistic data about CHK2 activation while also highlighting that a phosphorylation-null mutation (e.g. Thr387 to Ala) need not be neutral and in this case rendered the kinase inactive. With activated CHK2 kinase containing pThr383 in hand, the authors were able to characterize its substrate specificity by screening the phosphorylation of ~167,000 peptides corresponding to pSer and pThr sites in the human proteome.

3.2.3. First fruits: biological insights gained from pThr GCE technology.

Unlike pSer GCE, which saw a quick surge in use immediately after high-efficiency systems were established in 2015 (Fig. 11), we only found a single paper using a pThr GCE system to address a biologically-relevant question aside from the papers reporting GCE system developments (Table 4). Reasons for this are not clear, but may include the lower efficiency of the system compared to the pSer GCE systems (~30% that of pSer in sfGFP-TAG control expressions; unpublished observations by the authors of this review), as well as a greater susceptibility of the expressed target proteins to dephosphorylation, complicating downstream characterizations.

TABLE 4:

Users of GCE systems for encoding phosphothreonine

YearRef POI: site(s) encoded GCE System a
201773 Ubiquitin: 12,66
Cdk2: 160
1b
2020229 Ubiquitin: 66 1
2022227 CHK2: 383,387 2b,3b
a

System numbers represent the corresponding pThr GCE system found in Table 3

b

pThr GCE system was created in this report

The single paper applying a pThr GCE system, Deng et al.,229 encoded pThr into Ubiquitin at Thr66 to better understand how this modification impacts ubiquitination and protein degradation pathways. With the purified pThr66 ubiquitin, in vitro ubiquitination assays showed that phosphorylated ubiquitin blocked the poly-ubiquitin chain assembly catalyzed by the UbcH5C/RNF8 E2-E3 ligase complex. Similar results were observed when RNF8 was coupled to other E2 ligases, including Ubc13, UBE2E1, UBE2E2, and UBE2E3. When probing which stage in the poly-ubiquitination process was blocked, they found that the E2 ligases were successfully charged with pThr66-ubiquitin, and it was the discharging by RNF8 that was inhibited. This phosphorylation-dependent blocking was apparently specific to RNF8, however, as other E3 ligases MDM2 or X-linked inhibitor of apoptosis were able to promote chain assembly with pThr66-ubiquitin.229

3.2.4. Practical considerations for expressing proteins with pThr.

In Table 3, we delineate 3 systems available for pThr encoding in E. coli. All utilize the same pThr GCE components developed by Zhang et al.,73 but differ in the transcriptional promoters used and the expression hosts. While side-by-side comparisons of each system are not available, below we provide general considerations when adopting pThr GCE technologies for expressing proteins of interest.

3.2.4.1. Multi-plasmid expression systems for pThr GCE.

The standard two-plasmid architecture is the same for pThr GCE as it is for pSer GCE systems in which the machinery plasmid uses a pBR322 origin of replication (or pUC origin, which is the same compatibility group as pBR322), while target proteins are typically expressed on a p15a origin. In this way, the same plasmids used for pSer (2) and nhpSer (20) encoding can also be used for pThr (4) encoding (e.g. pRBC; Table 1), provided a BL21(DE3) derivative is the expression host. Plasmids with a CDF or RSF origin of replications can also be used for expressing the target protein.

3.2.4.2. pThr expression host considerations.

pThr GCE expression hosts have the ΔserC genomic mutation to eliminate pSer inside the cell (which can compete for pThr encoding), and the ΔycdX mutation to mitigate pThr hydrolysis. Zhang et al.73 demonstrated that proteins homogenously modified with pThr could be purified when expressing them in BL21(DE3) ΔserC ΔycdX with the pUC-pThr machinery plasmid (system pThr-1, Table 3), and we have independently corroborated these results using the pRBC-sfGFP-150TAG reporter plasmid (unpublished results). However, we anticipate there may well be scenarios in which the ΔpphA mutation helps to further minimize pThr hydrolysis of other target proteins. Where pre-mature truncation is a concern, and when consistent target protein expression needs to be maintained over multiple generations (e.g. systems biology applications or directed evolution experiments), expression in the C321.ΔA ΔpphA ΔserC ΔycdX ΔrelA may prove most useful (system pThr-3, Table 3). Concerns about slower cell growth compared to BL21(DE3) strains and mis-encoding natural amino acids due to near-cognate suppression should be considered in these cases.

3.2.4.3. Accessibility of pThr GCE systems.

Based on currently available information, reagents and protocols, our opinion is that the most accessible strategy to produce milligram quantities of target protein with homogenous pThr is to pair the pUC-pThr plasmid from Zhang et al.73 with the pRBC plasmid in the BL21(DE3) ΔserC ΔycdX expression host, and follow expression protocols used for pSer encoding.215 The pUC-pThr machinery and pRBC-sfGFP control plasmids are available on Addgene, and the BL21(DE3) ΔserC ΔycdX strain can be purchased commercially (see Table 3). As we learn more about these systems and more reagents become available, we anticipate that options to employ ΔpphA and RF1-deficient strains will expand the scope of proteins accessible with pThr GCE systems.

3.2.5. Future opportunities in pThr GCE technology development

3.2.5.1. Encoding non-hydrolyzable mimics of phosphothreonine in E. coli and eukaryotic cells.

Synthesis routes and peptide incorporation methods have been described for both nhpThr (25) (with a bridging CH2 group) as well as its difluorinated analog, F2-nhpThr (27) (Fig. 12A).136,230,231 To date no systems have been reported for encoding either ncAA in any cell type. Based on how the Sep-RS/Sep-tRNACUA pairs are permissive for nhpSer, we expect that the pThr GCE system components will be similarly permissive for nhpThr (25) and F2-nhpThr (27). And since the pThr GCE system is derived from the pSer system that is orthogonal in HEK293 cells, we would expect nhpThr machinery to also be compatible is these mammalian cells. Regardless, the main challenge we see as standing in the way of incorporating nhpThr (25) is overcoming bioavailability issues associated with poor uptake of the ncAAs. We are not aware of any biosynthetic pathway for making nhpThr (25) which could be used to overcome this limitation, and in the absence of the discovery or development of such a pathway, progress may depend on strategies to facilitate the import of nhpThr into the cell from the media.

While it is intuitive that nhpThr or F2-nhpThr would mimic pThr function much better than Asp/Glu mutations, limited evidence exists regarding their ability to mimic the function of pThr. In one example, peptides containing pThr (4) and nhpThr (25) were developed against Polo-like kinase 1 (Plk1) and shown to have similar binding affinities, while the F2-nhpThr (27)-containing peptide was reported to bind slightly weaker.232,233 In another example, nhpThr and F2-nhpSer containing peptides bound 14-3-3ζ with similar affinity, and surprisingly they both bound 10-fold tighter than the same peptide with native pThr.136,234 To shed additional light on this question about mimicry, we performed a similar PDB survey as above with pSer in which we asked how frequently the γ-oxygen of pThr engages in protein hydrogen bonds among all unique (<95% sequence identity) pThr-containing structures in the PDB. This analysis revealed that of 284 unique structures, only in 51 (18%) does the pThr residue fit this criterion (Supporting Table 1); thus, like was found for pSer, the bridging γ-oxygen of pThr is infrequently used as a hydrogen bond acceptor. Collectively these data indicate non-hydrolyzable phosphonate derivatives of pThr are effective pThr mimics, and are certainly better than Asp/Glu.

3.2.5.2. Dual encoding of pSer and pThr into the same protein.

Many proteins undergo simultaneous phosphorylation at both serine and threonine residues. One current advantage of EPL over GCE, as alluded to in section 1.3.2, is that peptides with both pSer and pThr (and/or pTyr) can be chemically synthesized and ligated to the same protein to generate phospho-proteoforms having multiple types of phosphorylation.43 In contrast, methods for dual encoding pSer and pThr into the same protein via GCE do not yet exist, and several challenges must be overcome to accomplish this. First, orthogonality between SepRS and pThrRS must be established so that one RS encodes a phosphoamino acid at TAG codons, and the other RS encodes the other phosphoamino acid at either TAA or TGA codons. This may be possible via directed evolution of the current Mma SepRS/Mj tRNA pairs or by identifying an orthologous SepRS/tRNA pair from an archaeon sufficiently diverse from Mma and Mj. Second the anticodon loop of the SepRS must be re-engineered to effectively amino-acylate an ochre/opal stop codon suppressing tRNA (i.e. tRNAUUA/tRNAUCA). Third, because pSer and pThr amino acids would both need to be present in the cell at high concentrations, the SepRS and pThrRS would need to be sufficiently selective for their respective phosphoamino acids so that they don’t incorrectly mis-acylate their cognate tRNAs. Another strategy would be to synthesize and encode photo-protected phospho-serine and threonine derivatives, since the photo-protecting groups could be sufficiently different in structure to more easily generate orthogonal RSs with the appropriate specificity.235 Needless to say, notable effort would be required to engineer such powerful dual phosphoamino acid GCE systems.

3.3. GCE systems for phosphotyrosine and its mimics

pTyr (6) is the third phosphoamino acid that has been encoded into proteins with GCE. While pSer (2) and pThr (4) systems have become robust enough and are starting to be adopted to answer biologically relevant questions, no GCE system installing authentic pTyr has yet been developed to that level of utility. Among four challenges that make GCE with pTyr more challenging, the first is making the free pTyr amino acid bioavailable; it does not enter the cell efficiently from the media, and unlike pSer, nhpSer and pThr, no biosynthetic pathways for the pTyr amino acid are known, so intracellular biosynthesis has not been possible. The second challenge is making a highly efficient pTyr-RS and EF-Tu, since there is no known natural pTyr-RS/tRNA pair as there was for pSer and native EF-Tu may not accommodate pTyr-tRNA well just like Sep-tRNA. Also, until the bio-availability issues can be solved, standard life/death cell-based selections to evolve a pTyr-RS or an EF-Tu for pTyr are not tractable. The third challenge is blocking the rapid dephosphorylation of pTyr by endogenous phosphatases. A fourth challenge, as discussed below, is that the ribosome appears to poorly accommodate pTyr amino-acylated tRNAs, so some engineering of ribosomal components may be needed for efficient pTyr encoding.

As a consequence of these challenges, the earliest published GCE efforts to study pTyr focused not on pTyr itself, but on encoding pTyr mimics using ncAAs that could enter cells, so that bioavailability was not a problem. The first two ncAAs used were para-carboxy-methyl-Phe (pCMF (32); Fig. 14A) and para-azido-Phe (34). pCMF mimics pTyr similarly to how Glu mimics pSer, and para-azido-Phe can be post-translationally chemically modified to make para-phosphoramidate-Phe (amido-pTyr, 31) replacing the bridging oxygen of pTyr with an NH group (Fig. 14A and 14B). Noteworthy is that sulfo-tyrosine (sTyr (33); Fig. 14A) is also a plausible pTyr mimic – and GCE systems for sulfo-Tyr installation are well established.237 Although sTyr has not generally been considered a viable mimic of pTyr, the available data imply that it is roughly as good a mimic as is pCMF (32) (see below), and so we include it in this review.

Figure 14. pTyr, its mimics and approaches to incorporate pTyr or amido-pTyr into proteins.

Figure 14.

(A) Shown are the side-chain structures pTyr (left, blue box) and six pTyr mimics (right, gray box). The top row shows three nhpTyr mimics with the bridging γ-oxygen of pTyr replaced with either CH2, CHF, or CF2 (highlighted red). (B) Amido-pTyr (panel A, top right) can be incorporated into a protein of interest (POI) by first encoding azido-Phe ncAA into the protein using GCE in E. coli. Then, after protein purification the azide moiety is reacted with a water-soluble phosphite to produce a nitrobenzyl ester, which is saponified by UV light to produce amido-pTyr.236 (C) For authentic pTyr incorporation into a POI via the chemically protected pTyr precursor (NpY), NpY is encoded using GCE in E. coli. Then, after the POI purification, the phosphoramidate protecting groups are cleaved by mild HCl acid treatment.258 (D) In the dipeptide approach for incorporating pTyr or nhpTyr, Lys-pTyr or Lys-nhpTyr is transported into the cell via the E. coli dipeptide transporter (DppA).72 In the cytosol, non-specific peptidases hydrolyze the dipeptide to make pTyr/nhpTyr which is then incorporated by GCE into the POI. In the case of pTyr, to minimize POI dephosphorylation, E. coli phosphatase activity is inhibited (red X) by sodium orthovanadate added to the media.

About a decade after the development of systems for encoding these pTyr mimics, the first GCE systems for installing authentic pTyr (6) were published, with some of these also being also able to install the non-hydrolyzable pTyr mimic with a CH2 replacing that bridging oxygen, that we will refer to as nhpTyr (28, Fig. 14A). Due to the timeline of developments for pTyr relevant GCE systems, in this section, we will not start with the systems for authentic pTyr, but will first provide an overview regarding pCMF (32), sTyr (33), and amido-pTyr (31), before moving on to describe the GCE systems for encoding pTyr (6, both directly and indirectly) and how they could also be used to encode the isosteric mimic nhpTyr (28). After that, as we did for the pSer and pThr sections, we will provide a brief overview of biological questions addressed using these systems, and future opportunities for GCE developments in this area.

3.3.1. pCMF, sTyr and amido-pTyr as replacements for pTyr function.

The utility of GCE systems that install pTyr mimics hinges on how well the mimic can recapitulate native pTyr function. In contrast to pSer and pThr, notable effort has gone into developing and characterizing pTyr mimics, particularly testing for their ability to promote stable peptide complexes with the pTyr-specific binding domains SH2 and PTB (for an extensive review on this topic, see Burke et al238). These domains form an intricate network of hydrogen bonds with the tetrahedral pTyr phosphoryl group, including the bridging η-oxygen of pTyr, making them a stringent test case for evaluating pTyr mimicry.239 Whereas neither Asp nor Glu can drive complexation between peptides and SH2 and PTB domains, pCMF can do so, albeit with affinities up to 1000-fold weaker than pTyr.240-242 Sulfo-tyrosine (sTyr) (as well as its sulfonate derivative having a Cη-S bond) is not commonly considered a mimic of pTyr, yet peptide ligands with sTyr bound SH2 domains 2 to 3-fold tighter than the pCMF-containing peptides.241,242 In additional studies, pTyr-dependent STAT1 dimerization and DNA binding could be induced by replacing pTyr at position 701 with pCMF and sTyr, albeit 20 and 60-fold worse, respectively, than STAT1 with pTyr, whereas Asp/Glu did not promote any detectable binding.243,244 Interestingly, pTyr has also been shown to be a reasonable mimic for sTyr. For example, both pTyr and sTyr modifications increased the affinity of the chemokine receptor CCR7 for the chemokine ligand CCL21.245 In another example, the affinity of recombinant hirudin for α-thrombin was fully maintained when replacing sTyr with pTyr, whereas unmodified hirudin bound 10-fold worse.246 Collectively, these data support the notion that sTyr is at least as good a pTyr mimic as pCMF is, even though neither would be considered a high-fidelity mimic of pTyr.

While amido-pTyr (31; Fig. 14A) can be installed onto proteins using Staudinger-phosphite reactions with translationally encoded para-azido-Phe236 (34; Fig. 14B), it is not clear how well amido-pTyr mimics pTyr. Phosphoramidates have tetrahedral geometry like pTyr and hydrogen bonding ability at the bridging η-nitrogen atom as well, and the pKa2 of the oxygen atoms of phosphoramidates is ~6.9, indicating the majority of molecules will be doubly deprotonated at physiologic pH like native pTyr.247 However, protonation at the Nη atom leads to rapid acid-catalyzed hydrolysis even at pH 7, raising concerns about its stability. Even so, phosphoramidate pro-drugs were able to inhibit the SH2 domain of Lck in in-cell assays equally as well as the phosphate-containing pro-drug derivatives.247

3.3.2. GCE encoding of the pTyr mimic pCMF

3.3.2.1. GCE systems for encoding pCMF in E. coli.

pCMF gets into cells well enough that it doesn’t suffer from the same bioavailability issues associated with pTyr, and in 2007, it become the first ncAA encoded into proteins in E. coli as a pTyr mimic.244 A novel M. jannaschii Tyr-tRNA synthetase/tRNACUATyr pair was evolved for pCMF (32) (Fig. 15), and used in the RF1-containing BL21(DE3) E. coli strain (system pCMF-1 in Table 5), to express a Staphylococcus aureus Z-domain protein containing pCMF at position Lys7 (at ~1.2 mg/L culture) without detectable mis-encoding of natural amino acids.244 To test how well pCMF served to mimic pTyr function, they expressed STAT1 with pCMF position at Y701, a site that when phosphorylated causes STAT1 to homodimerize and bind tightly to certain DNA sequences. Even though STAT1-pCMF701 bound the DNA with an affinity about 20-fold weaker than STAT1-pTyr701 (Kapp ~21 nM vs. Kapp ~1 nM), it bound ~6-fold tighter than unphosphorylated STAT1 (Kapp of WT is ~125 nM) leading to the conclusions that pCMF does partially mimic pTyr function in this case, and that how well it mimics pTyr in other cases is expected to be context dependent.

Figure 15. Rationale behind engineering TyrRS for pTyr encoding.

Figure 15.

Structure of the active site of M. jannaschii TyrRS (PDB: 1J1U) with Tyr bound, highlighting the six side chains varied in the library used to select a CMFRS (coloring and hydrogen bonds as in Figure 13). As recommended by Molprobity303, the Gln109 side chain amide has been flipped to better fit its hydrogen bonding environment. The table lists the six library sites and their mutations in the selected CMFRS #1.244

TABLE 5:

GCE systems for encoding phosphotyrosine and its mimics, with relevant Addgene IDsa

Year
[ref]
ncAA System Plasmid 1 Plasmids 2/3 E. coli Expression
host
Name
(ori)
ORF 1
Promoter/Gene
ORF 2
Promoter/Gene
ORF 3
Promoter/Gene
Name
(ori)
ORF 1
Promoter/Gene
ORF 2
Promoter/Gene
2007 [244] pCMF pCMF-1 pLEI (p15a) T5 POI lpp Mj tRNACUATyr b pBK (pBR322) GlnS pCMFRS #1 c BL21(DE3)
2011 [248] pCMF-2 N/I N/I Mj tRNACUATyr b N/I JX33 (RF1 KO)
2018 [249] pCMF-3 pEVOL (p15a) araC/GlnS d CMFRS #1 c lpp Mj tRNACUATyr b pAMH390 (pBR322) tac POI C321.ΔA.exp
2022 [250] pCMF-4 (HEK293) pB3 UbiC pCMF-RS e U6/H1 Ec tRNACUATyr (16 copies) CMV POI HEK293
2022 [251] pCMF-5 (HEK293) pCMFRS CMV CMFRS-1f U6/H1 B. subtilis - tRNACUATyr (9 copies) pcDNA3.1 CMV POI HEK293
2009 [236] p-azido-Phe Amido-Tyr-1 N/I N/I AzPhe RS-1 g N/I Mj tRNACUATyr b N/I N/I
2006 [237] sTyr sTyr-1 pSup (p15a) GlnS STyrRS h proK Mj tRNACUATyr b (six copies) pBAD (pBR322) araC POI DH10B
2009 [252] sTyr-2 pSUPAR6-L3-3SY (p15a) araC/GlnSd STyrRS2 i proK Mj tRNACUATyr b (six copies) pBAD and pET (pBR322) T7 or araC POI DH10B/BL21(DE3)
2016 [253] sTyr-3j pUltra (CDF) tac STyrRS h proK Mj tRNACUATyr b pBAD araC POI C321.ΔA.exp
2022 [254] sTyr-4 k pUltra (CDF) tac STyrRS h proK Mj tRNACUATyr b pEvol (p15a) araC NnSULT1C1 GlnS cysDNCQ BW25113 ΔcysH
pET22b-T5 (pBR322) T5 POI
2020 [255] sTyr-5 (HEK) pB3 UbiC sTyrRS- A1 (VGL) l U6/H1 Ec tRNACUATyr (16 copies) CAG POI HEK293
2020 [256] sTyr-6 (HEK) psTyrRS CMV sTyrRS c2 m U6 B. subtilis - tRNACUATyr (1 copy) pcDNA3.1 CMV POI HEK293
2022 [254] sTyr-7 (HEK) pAcBac2 CMV sTyrRS- A1 (VGL) l U6/H1 Ec and Bs tRNACUATyr (2 copy ea) pAcBac2 CMV POI U6/H1 Ec and Bs tRNACUATyr (2 copy ea) HEK293 (with integrated CAG driven NnSULT1C1)
Year
[ref]
ncAA System Plasmid 1 Plasmids 2/3 E. coli Expression
host
Name
(ori)
ORF 1
Promoter/Gene
ORF 2
Promoter/Gene
ORF 3
Promoter/Gene
Name
(ori)
ORF 1
Promoter/Gene
2016 [87] pTyr/nhpTyr pTyr-1 pTECH (p15a) lpp b pYRS1 n/EF-pY o proK Mj tRNACUATyr b pBAD (pBR322) araC POI TOP10 ΔserB ΔpgpA ΔpphB ΔphoA ΔaphA
2017 [257] pTyr-2 (CFPS) pUC18-rrnBmut p (pBR322) lac mutated rrnB operon N/I Sc tRNACUAPhe pET16b (pBR322) T7 POI BL21(DE3)
2017 [72] pTyr-3 pEVOL (p15a) araC/GlnS d CMFRS #1 c lpp Mj tRNACUATyr b pBAD (pBR322) araC POI DH10B
2017 [72] pTyr-4 pUltra (CDF) tac CMFRS #1 c proK Mj tRNACUATyr b pET-T5 (pBR322) T5 POI DH10B
2017 [258] (Me2N)2-pTyr (NpY) pTyr-5 pTAK (p15a) N/I POI N/I Mm tRNACUAPyl pBK (pBR322) GlnS MmNpY-RS q BL21(DE3)
a

For components available from Addgene, IDs are provided in a footnote for that GCE System number

b

Mj tRNACUATyr: M. jannaschii tRNAGUATyr with mutations C17A, U17aG, U20C, G34C, G37A, U47G

c

CMFRS #1: M. jannaschii TyrRS with mutations Y32S, L65A, F108K, Q109H, D158G, L162K

d

Two independently transcribed reading frames, one with araC promoter and the other with GlnS promoter

e

pCMF-RS E. coli TyrRS with mutations Y37H, L71V, D182G, L186M

f

CMFRS-1 E. coli TyrRS with mutations Y37H, L71V, D182G

g

AzPheRS-1: M. jannaschii TyrRS with mutations Y32T, E107N, D158P, I159L, L162Q

h

STyrRS: M. jannaschii TyrRS with mutations Y32L, L65P, D158G, I159C, L162K

i

STyrRS2: M. jannaschii TyrRS with mutations Y32L, L65P, D158G, I159T, L162K

j

Component Addgene IDs are 82417 (plasmid 1), 85482 and 85483 (control pBAD-sfGFP forms and for POI if expressing in DH10b or C321.ΔA.exp), 85492 and 85493 (control pET28-sfGFP forms and for POI if expressing in BL21(DE3) or B95(DE3) ΔA ΔfabR), 49018 (E. coli strain C321.ΔA.exp), and 197934 (E. coli strain B95(DE3) ΔA ΔfabR).

k

Component Addgene IDs are the same as those in footnote j, plus the sTyr biosynthesis plasmid 188983. Note that the pET22b-T5 vector used with this system for POI expression is available on Addgene (e.g. # 188998) but not with an sfGFP reporter protein. The pBAD-sfGFP vectors in footnote j are compatible with the machinery plasmid and expression host. However in the original publication only 15 mg/L arabinose was suggested for pEVOL biosynthetic machinery plasmid expression, which is about 20-fold lower than what would normally be used for pBAD-sfGFP protein expression.

l

sTyrRS-A1 (VGL): E. coli TyrRS – L71V, D182G

m

sTyrRS-c2: E. coli TyrRS – L71V, W129F, D182G

n

pYRS1: M. jannaschii TyrRS with mutations Y32L, L65R, D158G, I159C, L162K, D286R

o

EF-pY: E. coli EF-Tu with mutations E216V, D217G, F219G

p

rrnB operon with 23S rRNA mutations : 2057GAAAGAC20632057AGCGTGA2063 and 2600ACAGTT26012600GTTCGG2605

q

MmNpYRS: Methanosarcina mazei PylRS with mutations A302S, L309M, I322L, N346A, C348G, W417T

In 2011 and in 2018, this expression system was adopted for use in RF1-deficient strains, first in JX33248 and then in C321.ΔA.exp249 ( systems pCMF-2 and pCMF-3, respectively, in Table 5). Since the fidelity of encoding was not noted in either report, it is not clear the extent to which the efficiency of pCMF encoding was sufficient to out-compete near-cognate suppression or other ways of mis-encoding. Although the pCMF amino acid is not commercially available yet, the E. coli pCMF-1 and the pCMF-3 systems (Table 5) have been used several times to study pTyr dependent signaling systems (Table 6).

TABLE 6:

Users of GCE systems for encoding phosphotyrosine

YearRef POI: site(s) encoded a GCE System b
2007244 STAT1: 701 pCMF-1c
2009236 SecB: 156 Amido-Tyr-1c,d
2011273 RAD52: 104 pCMF-1
2011248 EGFP: 39,151,182 pCMF-2c
2012274 FGF2: 82 pCMF-1
2013243 STAT1: 701 sTyr-1
2014275 PRMT1: 291 pCMF-1
2015276 Cytochrome c: 48 pCMF-1
201687 sfGFP: 143 1c
2016277 RAD51: 54,315 pCMF-1
2017257 DHFR: 10
IκB-α: 42
2c,d
201772 Myoglobin: K99
Abl1 SH3 domain: 30,52
3c,e,4e
2017258 CaM: M76
GFP: 182
Ubiquitin: 59
5c
2017278 Cytochrome c: 48 pCMF-1
2018279 NCK1 SH3–1: 55 3
2018249 Cra: 47 pCMF-3c
2020280 Ubiquitin: 59
ZAP70: 221,248
3
2021281 NF-κB: 44,60,82,90 2d
2022282 Cytochrome c: 48 pCMF-1
2022250 STAT3: 705 pCMF-4c
2022251 STAT1: 701 pCMF-5c
2023272 NF-κB: 7,44,60,82,90,241,270 2e,d
2023283 Many peptides pCMF-1
2023284 Parkin: 143 pCMF-1
a

WT sites are tyrosine unless otherwise shown

b

System numbers represent the corresponding pTyr GCE system unless otherwise shown; all systems can be found in Table 5

c

GCE system created in this report

d

Cell-free protein synthesis

e

nhpTyr is incorporated

3.3.2.2. GCE systems for encoding pCMF in mammalian cells.

In 2022, two groups independently engineered GCE systems to encode the pCMF mimic into proteins in eukaryotic cells. A challenge was that the most commonly adopted GCE chassis compatible with eukaryotic ncAA encoding – the pyrrolysine-RS/tRNACUAPyl pair – tends to not tolerate ncAAs with positive or negative charges in the side chains, such as pCMF.250 In solving this problem, both groups chose to work with the E. coli TyrRS/tRNACUATyr pair, which is orthogonal in eukaryotic hosts and tolerant of negatively charged ncAAs. Evolving this Ec TyrRS/tRNACUATyr cannot be performed in standard strains of E. coli, so that problem also needed to be solved.

The approach of Grasso et al.250 was to use their previously developed “altered translational machinery tyrosyl” (ATMY) E. coli strains in which the endogenous EcTyrRS/tRNA pair was functionally substituted with the archaeal M. jannaschii TyrRS/tRNA pair counterpart.259,260 In these strains, the EcTyrRS/tRNA has been functionally “liberated” from its cellular role, allowing it to be independently evolved using well-established E. coli-based life/death selection strategies. After first engineering a tRNACUATyr variant with increased orthogonality to endogenous E. coli synthetases (specifically, the GlnRS), Grasso et al. generated a library of EcTyrRS mutants and carried out a selection for pCMF encoding into sfGFP.250 The top performing hit (called “pCMF-RS”) was then transferred to a eukaryotic compatible expression plasmid (pB3) that also expressed 16 copies of the tRNACUATyr and the protein-of-interest (system pCMF-4 in Table 5). Fidelity of pCMF encoding into eGFP-39TAG in HEK293T was assessed as >95% by mass spectrometry. This system was then used to encode pCMF at Tyr705 of STAT3, a site of phosphorylation known to cause homodimerization and transcriptional activation of STAT3-regulated promoters. Using a luciferase reporter expressed from a STAT3-regulated promoter, the STAT3-705-pCMF mutant triggered significantly elevated levels of luciferase expression over the wild-type STAT3 and the Y705F mutant, suggesting some level of pTyr mimicry by pCMF. However, a direct comparison to STAT3 with authentic pTyr at position 705 was not feasible.

The approach of He et al. was to use Saccharomyces cerevisiae instead of E. coli to select a mutant EcTyrRS/tRNACUATyr pair for pCMF encoding.251 In a testament to the robustness of these alternative selection platforms, He et al. identified nearly the identical variant of the EcTyrRS as Grasso et al. that they called CMFRS-1 (Table 5). Optimization of HEK293T plasmids expressing CMFRS-1 and 9 copies of B. subtilis-tRNACUATyr (generating system pCMF-5 in Table 5) afforded faithful encoding of pCMF when it was supplemented in the media at 5 mM.251 Similar to other work showing pCMF can substitute for pTyr function, He et al. installed pCMF at Y701 of STAT1 (equivalent to Y705 of STAT3) and showed it induced homodimerization and binding to STAT1-regulated DNA promoters at 120 nM concentration. This concentration was just like STAT1 with pTyr at Y701 (installed by a kinase), whereas WT and Asp/Glu mimetics showed no binding even at 240 nM, the highest concentration tested.251 They also demonstrated that STAT1-Y701pCMF stimulated expression of interferon regulatory factor-1 (IRF-1), a protein whose transcription is regulated by STAT1 Tyr705 phosphorylation, at 1.15-fold higher than when wild-type STAT1 was expressed. Collectively, these works demonstrate pCMF encoding in HEK293 cells is now feasible and is a valuable tool to study pTyr-dependent signaling systems in a eukaryotic context.

3.3.3. GCE encoding of the pTyr mimic amido-pTyr (31) in E. coli.

To circumvent encoding challenges of native pTyr, Serwa et al.236 in 2009 showed that Staudinger-type reactions of azides with phosphites can occur in high yields and at room temperature and pH ~8 to produce phosphorimidates, which can be hydrolyzed to form phosphoramidates. By reacting a phosphite derivative containing photo-reactive 2-nitrobenzyl ester substituents, a photo-protected phosphoramidate pTyr analog was formed, that upon UV irradiation formed amido-Tyr (31). To showcase this methodology, Serwa et al. encoded para-azido-Phe (34) into SecB using an M. jannaschii AzF-RS/tRNACUA pair261, which was then quantitatively converted to amido-pTyr (31) after a 24 h reaction with phosphite and deprotection via UV irradiation at pH 8.2.236 The extent to which this amido-Tyr analog mimics pTyr is not clear, but Serwa et al. did confirm their Sec B derivative with amido-pTyr was readily detected by an anti-pTyr antibody in western blots.236 At an elevated pH of ~8, the photoprotected phosphoramidate on SecB was stable for at least 72 h. The stability of amido-pTyr after deprotection was not evaluated, and so it is not clear how fast it would undergo hydrolysis at physiologic pH.

3.3.4. GCE encoding of the pTyr mimic sulfo-Tyr (sTyr, 33).

sTyr (33) is a natural PTM formed by tyrosyl-protein sulfotransferases during protein maturation in the Golgi apparatus.262 Proteins with sTyr are found exclusively in secreted proteins, including peptide hormones, chemokine receptors and the thrombin inhibitor hirudin. Interestingly, no known “eraser” proteins are known for sTyr, making it a long-lived PTM.262 Like phosphorylation, tyrosine sulfation adds a tetrahedral functional group that is fully ionized at neutral pH however it makes weaker hydrogen bonds due its −1 charge compared to the −2 charge of pTyr. Despite fewer studies having tested its ability to mimic pTyr as compared with pCMF, its structural features more closely resemble pTyr, and indeed published data support it being a mimic on par with pCMF (as discussed above).

3.3.4.1. GCE systems for encoding sTyr in E. coli.

In 2006, Liu et al. employed standard life/death selection strategies to engineer a M. jannaschii Tyr-RS/tRNACUA pair, enabling selective encoding of sTyr into proteins in E. coli (system sTyr-1, Table 5).237 Faithful incorporation of sTyr into a Z-domain variant from S. aureus was confirmed through mass spectrometry. Subsequent optimization of this GCE system by incorporating 6 copies of the M. jannaschii tRNACUATyr gene and adding 10 mM sTyr in the media, enabled expression of hirudin—the most potent natural thrombin inhibitor found in leeches—with sTyr at position 63. The yield of 5 mg/L culture compared with the 12 mg/L yield of wild-type hirudin represented ~40% efficiency.237 Mass spectrometry analysis revealed two peaks—one corresponding to hirudin with sTyr63 and the other unmodified. The unmodified peak was attributed to sulfate loss during ionization and MS analysis. Consistent with previous findings, the sulfated hirudin exhibited an ~10-fold increase in potency against thrombin activity compared to unmodified hirudin.237 That 10 mM sTyr was used in the media suggested sTyr might suffer from similar bioavailability concerns of phosphoamino acids. Nevertheless, since its original inception, adaptations of this sTyr GCE machinery system have been made with improved plasmid design and expression hosts (systems sTyr-2 and sTyr-3, Table 5).252-254

3.3.4.2. GCE systems for encoding sTyr in mammalian cells.

In 2020, two groups independently engineered GCE systems to encode sTyr into proteins in eukaryotic cells. Like with pCMF encoding in mammalian cells, both chose to work with the E. coli TyrRS/tRNACUATyr pair, which is orthogonal in eukaryotic hosts and tolerant of negatively charged ncAAs.

The approach of Italia et al. was to use their ATMY E. coli strains (see section 3.3.2.2) to evolve the EcTyrRS for selective sTyr encoding using well-established E. coli-based life/death selection strategies.255 The top performing RS hit (L71V and D182G) enabled expression of 8–10 mg per liter culture of sfGFP with sTyr in E. coli with only 1 mM sTyr in the media. This novel sTyr-RS was then transferred to a eukaryotic compatible expression plasmid (pB1U) that also expressed 16 copies of the tRNACUATyr (system sTyr-5 in Table 5), enabling sTyr encoding into GFP and human heparin cofactor II with 2 mM sTyr in the media.255

The approach of He et al. was to use S. cerevisiae to select a mutant EcTyrRS/tRNACUATyr pair for sTyr encoding.256 These selections identified the same EcTyrRS variant of the as Italia et al.255 as well as one with improved fidelity containing an additional W129F mutation (sTyrRS c2). Transfer of this “sTyrRS c2” to a eukaryotic expression vector having one copy of tRNACUATyr (making system sTyr-6 in Table 5) enabled faithful sTyr encoding into GFP in HEK293T cells with 1 mM sTyr in the media at ~70% the efficiency of wild-type expression.256 sTyr was then encoded with ~40% efficiency into the chemokine receptor CXCR4, thereby demonstrating that sTyr could be incorporated into a naturally sulfated target protein. To our knowledge, neither mammalian system has been used to explore pTyr protein function in a eukaryotic context.

3.3.4.3. In-cell biosynthesis of sTyr improves encoding efficiency in E. coli and mammalian cells.

Chen et al. in 2022 noted the previously described sTyr GCE systems often required up to 20 mM sTyr for efficient encoding due to its poor bioavailability, and pursued engineering a biosynthesis of sTyr to improve these systems.254 Using bioinformatic approaches, they tested 27 proteins with similarity to rat and human cytosolic sulfotransferases, finding one from Nipponia nippon that transferred the sulfate moiety from 3′-phosphoadenosine-5′-phosphosulfate (PAPS) to the free tyrosine amino acid, thus producing sTyr. After modifying E. coli with a ΔcysH (generating the BW25113 ΔcysH strain) to minimize PAPS degradation, and co-expressing an M. jannaschii sTyr-RS/tRNACUA from Liu et al.237 with the N. nippon sulfotransferase (making system sTyr-4 in Table 5), sTyr encoding into GFP was improved ~4-fold compared to supplementing media with 1 mM sTyr.254 The N. nippon sulfotransferase was then stably integrated into the genome of HEK293T, enabling biosynthesis of sTyr in a mammalian cell and its encoding into proteins using the Ec sTyrRS/tRNACUATyr pair (system sTyr-7 in Table 5). Using this system, the efficiencies of sTyr-containing GFP production were ~3-fold higher than when media was supplemented with 3 mM sTyr, thus demonstrating the advantages of sTyr biosynthesis. If sTyr is indeed as effective a pTyr mimic as pCMF, these systems provide an additional and potentially very useful strategy to probe pTyr function both in vitro and in vivo.

3.3.5. GCE systems for encoding native phospho-tyrosine in E. coli.

As noted in the introduction to Section 3.3, four challenges to be addressed for efficient GCE encoding of pTyr are (i) making the free pTyr amino acid bioavailable, (ii) making a highly efficient pTyr RS and EF-Tu, (iii), blocking the rapid dephosphorylation of pTyr by endogenous phosphatases, and (iv) ensuring that the ribosome can efficiently handle a pTyr-loaded tRNA. Although none of the pTyr GCE systems developed thus far address all of these challenges at once, each one includes innovative approaches to address one or more of these challenges. Also, the focus here is on E. coli pTyr GCE systems, because to our knowledge, no systems working in eukaryotic cells have yet been developed.

3.3.5.1. Addressing the phosphatase, synthetase and EF-Tu challenges to encoding of pTyr.

The first instance of pTyr (6) encoding in E. coli was reported by Fan et al. in 2016.87 Noting that pTyr was quickly hydrolyzed in E. coli lysates, they identified the top 5 phosphatases responsible for pTyr dephosphorylation (SerB, PgpA, PphB, PhoA and AphA) and removed them all from the E. coli TOP10 genome to create the “TOP10 Δ5P” strain. Next, they purified and tested the in vitro abilities of 76 rationally designed variants of an M. jannaschii sTyr-RS to amino-acylate the Mj tRNACUATyr with pTyr, calling the highest activity variant “pYRS1”. Then, by expressing this pYRS1 in cells and monitoring GFP reporter protein production, they screened a set of rationally designed EF-Sep variants having larger amino acid binding pockets, and identified a variant dubbed “EF-pY” as the best at bringing a charged pTyr-tRNA to the ribosome. By co-expressing pYRS1, EF-pY and Mj tRNACUATyr in the TOP10 Δ5P strain (system pTyr-1 in Table 5) along with including 10 mM pTyr in the media, Fan et al. expressed sfGFP-pTyr143 at ~4-fold above background and 20 mg/L culture yields.87 Though neither pTyr encoding fidelity nor pTyr stability post- purification were quantified, bottom-up MS/MS fragmentation analysis indicated pTyr was successfully encoded into sfGFP.

3.3.5.2. A cell-free and ribosomal engineering approach to pTyr encoding.

In 2017, Chen et al.257 used an E. coli cell-free expression system to overcome issues associated with bio-availability, and also chemically acylated their amber suppressing tRNACUA with pTyr (as well as a photo-protected form of pTyr) to circumvent the lack of an efficient pTyr-RS. This cell-free system allowed them to show that when expressing dihydrofolate reductase, the ribosome extended nascent peptides with pTyr at only about 3% the efficiency of extending peptides with Tyr, prompting the authors to modify the ribosome itself, though poor EF-Tu transport could have also contributed to the low efficiency. To improve ribosome compatibility with pTyr-amino acylated tRNAs, a rather elegant strategy was devised in which 23S rRNA variants were selected for sensitivity to phospho-puromycin.257 Puromycin is an antibiotic that mimics the 3’ end of a tRNA amino-acylated with Tyr, and binds to the ribosome A site to inhibit peptide elongation.263 Phospho-puromycin, in which the tyrosine moiety is modified to phospho-tyrosine, does not bind to wild-type ribosomes, and Chen et al. hypothesized that if 23S rRNA variants could be found that were sensitive to phospho-puromycin, they would be better at encoding pTyr.257 From a library of 23S rRNA mutants, they identified a subset sensitive to phospho-puromycin, and the top performing ribosomal mutant completed their cell-free GCE system (system pTyr-2 in Table 5). Using this system, Chen et al. encoded pTyr into dihydrofolate reductase in response to an amber codon at 22% efficiency, compared to just 3% for the native ribosome. These studies highlight that pTyr encoding in E. coli may well be bottlenecked by the inability of the ribosome to efficiently transfer pTyr to the growing peptide chain. In these studies, no modified EF-Tu was used to facilitate delivery of the pTyr amino-acylated tRNA to the ribosome, and so encoding efficiency may be further improved by using EF-pY.

3.3.5.3. Encoding a chemically protected analog of pTyr.

To circumvent both the bioavailability and hydrolysis issues of pTyr, Hoppmann et al.258 in 2017 encoded the pTyr precursor bis(dimethylamino)phospho-tyrosine (NpY (35); Fig. 14C) by evolving a novel Methanosarcina mazeii pyrrolysine RS/tRNACUAPyl pair. Being uncharged, NpY (35) is expected to more easily get into the cell than does pTyr and also to be better accommodated by the native EF-Tu and ribosome. It is not itself a pTyr mimic, but it can be chemically deprotected in vitro in 36 to 48 h under mild acid conditions (0.1 - 0.4M HCl, pH ~1) at 4 °C to yield authentic pTyr (6, Fig. 14C).258 Using this system (system pTyr-5 in Table 5), Hoppman et al. encoded NpY into calmodulin at site M76 with yields of 1 mg/L culture, and mass spectrometry confirmed quantitative incorporation of NpY as well as conversion to pTyr after acid treatment.258 NpY was also encoded into GFP with yield of 1.25 mg/L culture (corresponding to ~30% efficiency), as well as into ubiquitin at Tyr59 in isotopically labeled media. Given that the MmNpY-RS/tRNACUAPyl pair is also orthogonal in eukaryotic cells, this system could be used in other expression hosts for preparing proteins difficult to express in bacteria. That the Fmoc-protected form of NpY is commercially available and only requires a single synthesis step to form the free NpY amino acid is advantageous, but major limitations are that this system (i) cannot be effectively used for proteins that do not tolerate acid treatment or cannot be re-folded after acid treatment, and (ii) cannot produce authentic pTyr-protein inside cells.

3.3.5.4. Using di-peptides to improve pTyr bioavailability.

In 2017, Luo et al.72 pursued a different approach to overcome bioavailability issues as well as the lack of an existing efficient pTyr-RSs. To overcome bioavailability bottlenecks, they built on prior work264,265 showing that poorly bioavailable amino acids were better transported into cells – via the DppA dipeptide transporter – if they were conjugated to another amino acid to make a dipeptide (Fig. 14D). A 6-step synthetic route was devised to make a Lys-pTyr dipeptide (36), and by including it in the media at 2 mM they tested the permissivity of Mj TyrRS variants previously evolved to encode pCMF (32), p-borono-phenylalanine (BoroF) and sTyr (33) for their ability to encode pTyr (6). No modified EF-Tu variants were used.72 These screens identified the RS previously isolated as “CMFRS #1” (Fig. 15) as the most permissive for pTyr, and using this RS (system pTyr-3 in Table 5) pTyr was encoded into myoglobin-K99TAG with yields of ~60 mg/L culture. As MS data showed that only Tyr was present at site 99, the authors concluded that pTyr was successfully encoded during translation, but was dephosphorylated afterward. Addition to the expression media of the phosphatase inhibitor sodium orthovanadate allowed for detection by whole protein MS of some purified pTyr99-myoglobin, albeit at less than 10% of the total protein. The fidelity of the encoding step itself was not quantified.

3.3.6. GCE systems for encoding the isosteric pTyr mimic nhpTyr in E. coli

3.3.6.1. nhpTyr mimics (28-30) of pTyr.

As described extensively earlier in this review, an isosteric and effective non-hydrolyzable mimic for pSer (2) is the nhpSer (20) molecule in which the bridging γ-oxygen is replaced with a CH2 group (Fig 8A). The equivalent isosteric mimic of pTyr, in which the bridging η-oxygen is replaced with a CH2 group, is commonly called phosphono-methyl-phenylalanine or Pmp, but we will refer to it here by the more intuitive moniker “nhpTyr” (28, Fig 13A). nhpTyr (28) has been shown to be an effective mimic of pTyr in many situations, and a much better mimic than pCMF (32) in all situations in which they have been compared, while the mono- or di-fluorinated versions of nhpTyr (F1- and F2-nhpTyr, 29 and 30, respectively) have been shown in most situations to be better mimics than nhpTyr (28).241,266-269 Two such situations are the SH2 and PTB protein families mentioned above which in binding pTyr peptides and proteins make key interactions with the bridging η-oxygen. In these cases, the nhpTyr (28) substitution generally resulted in modest (20 to 40-fold) reductions in affinity compared to the same peptides with pTyr (6), while F2-nhpTyr (30) substitutions resulted in only ~5-fold reduction. The weaker affinities of nhpTyr peptides were attributed to the loss of pTyr η-oxygen interactions in nhpTyr rather than differences in the pKa2.270 Given the similarity of nhpTyr to pTyr, two of the GCE systems developed for encoding pTyr have also been used to install nhpTyr (28) in to proteins. To our knowledge none of the pTyr GCE systems have yet been tested for how well they incorporate either of the fluorinated derivatives 29 and 30.

3.3.6.2. GCE systems for encoding nhpTyr.

In the same publication reporting their di-peptide approach for encoding for pTyr (system pTyr-3 in Table 5), Luo et al.72 showed nhpTyr (28) could be encoded by the same system if the media was supplemented with 2 mM Lys-nhpTyr (37). Using the M. jannaschii based CMFRS#1 RS in a slightly modified plasmid architecture (system pTyr-4 in Table 5), nhpTyr was encoded into sites Tyr30 and Tyr52 of the SH3 domain of ABL1 with yields of ~7 mg/L culture, and MS spectra supported that the encoding of nhpTyr was homogenous. ABL1-nhpTyr30 and ABL1-nhpTyr52 bound the 3BP2 peptide, a native ligand from the adaptor protein 3BP2, with similar affinities to the ABL2 variants containing pTyr at those positions (made via chemical ligation strategies271). While this provides an attractive strategy to encode nhpTyr into proteins, the synthesis of the Lys-nhpTyr (37) is not trivial and large quantities of it (~600 mg per liter culture for 2 mM concentration) were needed for the expression.

In 2023, Chen et al.272 showed that the same cell-free translational system they developed in 2017 to install pTyr (system pTyr-2 in Table 5; containing the modified 23S ribosomal subunits) could also be used to install nhpTyr (28). The only change was that nhpTyr rather than pTyr was chemically ligated to an amber suppressing tRNA. Using this approach, they encoded nhpTyr (28) into site Tyr60 of the p50 subunit of NF-kB transcription factor. Being impervious to phosphatase activity, the NF-kB-nhpTyr60 protein was slower to release from IL2-promoter DNA than was the protein with pTyr encoded at this site. Because this system leverages cell-free translational systems as well as chemical methods to amino-acylate the amber codon suppressing tRNA, the bottlenecks associated with both ncAA bioavailability and pTyr RS engineering were avoided.

3.3.7. Practical considerations for expressing proteins with pTyr mimics.

Currently, there is a lack of robust GCE systems to directly produce proteins homogeneously and stably modified with pTyr. The adoption of pTyr (and nhpTyr) systems 1-4 (Table 5) by laboratories beyond their respective developers has been virtually nonexistent, leaving uncertainty regarding their practical usability. The system to install NpY (35) (system pTyr-5, Table 5, Fig. 14C) offers an elegant solution to acquire proteins with native pTyr, provided the proteins can be refolded after acid-mediated deprotection, but we found only two publications using it since 2017. Consequently, encoding mimics stands out as the currently most tractable strategy for investigating pTyr proteins using GCE. The successful encoding of pCMF (32) and sTyr (33) by multiple independent laboratories into various biologically significant proteins (Table 6) leading to noteworthy insights into pTyr protein structure and function, underscores the effectiveness of these analogs as superior mimics compared to Asp/Glu, though it is important to consider that results still must be interpreted in context of the imperfect mimicry of pCMF and sTyr.

3.3.7.1. Multi-plasmid E. coli GCE expression systems for pTyr mimics and NpY.

pCMF (32) and sTyr (33) GCE systems follow the standard multi-plasmid architecture in which proteins of interest are expressed from more classical vectors such as pET and pBAD having pBR322 origins, while the machinery plasmids contain either a p15a or CDF origin of replication. For the E. coli GCE system employing biosynthesis of sTyr (system sTyr-4, Table 5), a third plasmid is used to express the enzymes required for sTyr biosynthesis. The published plasmid architecture for NpY (35) encoding is less standard, where the RS is expressed on one plasmid (pBK) while the tRNACUA and protein of interest are expressed on a second plasmid (pTAK) (system pTyr-5, Table 5). Therefore, proteins of interest must be cloned into this pTAK plasmid specifically for the purpose of NpY encoding.

3.3.7.2. Expression host considerations for pTyr mimics in E. coli.

pCMF (32) and sTyr (33) GCE systems do not require any unique modifications to the expression hosts and therefore can be used with standard, commercially available protein expression strains such as BL21(DE3) and DH10B. The RF1-deficient strain C321.ΔA.exp has also been used to encode pCMF and sTyr for truncation free expression, while B95(DE3) ΔA ΔfabR has been successfully employed for sTyr encoding.127 Encoding of biosynthesized sTyr does require the expression host have the ΔcysH genomic modification. While the strain BW25113 ΔcysH was used with this system, it is reasonable to predict a ΔcysH derivative of BL21(DE3) would be compatible as well, if it were available.

3.3.7.3. pCMF encoding in mammalian cells.

The two-plasmid architecture of mammalian GCE expression systems typically mirrors that of E. coli systems in which one plasmid houses the machinery components while the second plasmid houses the protein of interest. However, one key difference from E. coli systems is that both mammalian plasmids have multiple expression cassettes of the tRNACUA in order to maximize tRNACUA expression, and also to allow users to optimize the ratios of the two transfected plasmids to maximize target protein expression while maintaining constant tRNACUA levels. System sTyr-7 adopts this standard plasmid architecture, and by transfecting the plasmids into HEK293 cells that are stably expressing NnSULT1C1, it leverages the advantages of sTyr biosynthesis. For pCMF (32) encoding, published plasmid architectures deviate from the traditional setup: system pCMF-4 expresses the protein of interest on the machinery plasmid for a single plasmid expression setup, while system pCMF-5 uses two plasmids but only the machinery plasmid has tRNACUA expression cassettes. Studies to compare advantages and disadvantages of these two pCMF mammalian systems will help clarify how to best use these technologies and possibly how to improve them further.

3.3.7.3. Availability of sTyr GCE systems.

Among the machinery plasmids needed for NpY (35), pCMF (32) and sTyr (33) encoding, Addgene appears to have the machinery plasmids for sTyr encoding in E. coli, as well as the plasmid for the sTyr biosynthetic pathway (see footnotes of Table 5). pBAD and pET vectors for expressing the sfGFP-150TAG reporter protein are available that can be paired with the sTyr systems to verify efficiency and fidelity. However, we note that for the biosynthetic pathway approach of sTyr encoding, the pBAD reporter plasmids may not be ideal (see footnote in Table 5).

3.3.8. First fruits of pTyr GCE.

Despite the unique challenges of pTyr encoding, GCE technologies are seeing increasing usage to address biological questions related to pTyr functionality (Table 6). The majority of these studies (14 out of 24) employed pCMF (32) as a pTyr mimic because of its stability in cellular environments, the relatively straight-forward methodology to encode it into proteins, and its ability to better mimic pTyr compared to Asp/Glu. Six studies successfully generated biologically relevant proteins with authentic pTyr, using either the NpY (35) encoding/deprotection strategy or the cell-free expression strategy with modified ribosomes. Below, we provide a few highlights of these studies.

3.3.8.1. pTyr-dependent DNA repair mechanisms and transcriptional activation.

Tyrosine phosphorylation plays crucial roles in the regulation of transcription factors as well as proteins contributing to DNA repair mechanisms. As noted above, several studies have shown the encoding of pCMF into STAT1 (Y701) and STAT3 (Y705) induces pTyr-dependent homodimerization and binding to STAT1/3-regulated promoters, both in vitro and in vivo.244,250,251 In other studies, pCMF was encoded into RAD52 at Y104 and RAD51 at Y54 and Y315 to understand how phosphorylation of these proteins regulates their function in repairing single-stranded DNA (ssDNA). Using the RAD52-Y104pCMF mimic, Honda et al.273 in 2011 showed that pCMF at position Y104 confined RAD52 to binding ssDNA regions in order to recruit RAD51 to sites of DNA damage, whereas unmodified RAD52 readily diffused into double-stranded DNA. Later in 2016, Subramanyam et al.277 showed that pCMF at Y54 enhances RAD51 recombinase activity by modifying RAD51 nucleoprotein filament formation, and allowing RAD51 to efficiently displace RPA from the ssDNA. In these cases, the functions of phosphorylated RAD51/52 could not be recapitulated with Asp/Glu mimetics.

Using the cell-free encoding system with the modified 23S rRNA, Chen et al.257,272,281 in a series of three studies sought to understand how Tyr phosphorylation of the IκB-α/NF-κB complex regulated gene transcription. When unmodified, IκB-α binds to the heterodimeric transcription factor NF-κB, masking the nuclear localization sequence of NF-κB and inhibiting its ability to bind promoter DNA. First, in 2017257, they showed that when phosphorylated at position Y42, IκB-α still interacted with NF-κB but only transiently, allowing NF-κB to form overall more stable complexes with DNA while also mediating the exchange of exogenously added DNA with pre-formed NF-κB/DNA complexes. Then, in 2021, Chen et al.281 encoded pTyr into the p50 subunit of NF-κB at positions Y60 and Y82, revealing that phosphorylation at these sites facilitated its binding to the CD40 promoter, resulting in increased CD40 expression. But interestingly, in 2023,272 Chen et al. showed this same phosphorylated form of NF-κB (with the p50 subunit phosphorylated at Y60) did not bind the IL-2 promoter, revealing how phosphorylation can tune which genes NF-κB activates for transcription. Instead, binding to the IL-2 promoter could be achieved if NF-κB was first phosphorylated at Y60 in the p50 subunit and after that was further phosphorylated by an unknown kinase in Jurkat cell extract. Since phosphatases were present in the cell extract, the binding of NF-κB to the IL-2 promoter was transient, and notably could be extended when nhpTyr was encoded at Y60 of the p50 subunit. Collectively, these works showcase the power of this cell-free encoding system for revealing biological pTyr signaling mechanisms.

3.3.8.2. pTyr-dependent regulation of signaling systems.

SH2 domains, which recognize pTyr-containing sequences, and SH3 domains, which recognize PxxP-containing segments, often exist as interconnected modules within multi-domain proteins that can integrate phosphorylation-dependent signal recognition with protein-protein interactions. Adding complexity is that some SH3 domains are phosphorylated at tyrosine residues, and how such PTMs affect SH3 function has been addressed in a couple of GCE-enabled studies. Dionne et al.279 in 2018 encoded the chemically protected pTyr analog NpY into site Tyr55 of the SH3 domain of NCK1, and after chemical deprotection and refolding, they found the authentically phosphorylated SH3-NCK1 protein no longer bound partner proteins. These observations showed that Tyr phosphorylation of SH3 domains can serve to release binding clients and collapse SH3-mediated signaling networks. However, this is not universal, as in a study of the SH3 domain of Abl1, when Luo et al.72 installed nhpTyr at positions Tyr30 and Tyr52 via the di-peptide approach (system pTyr-4 in Table 5), they did not observe notable differences in binding to its native ligand, 3BP2.

In 2014, Rust et al.275 encoded pCMF at position Y291 of the protein arginine methyltransferase 1 (PRMT1) to show that phosphorylation alters its substrate specificity. It caused little effect on the methylation of full-length histone H4 but led to weakened interactions with another known substrate, hnRNP A1, suggesting a regulatable connection between Tyr phosphorylation and arginine methylation. In other work, Li et al.283 in 2023 sought to understand how a pTyr residue in a peptide substrate of the c-Src kinase influenced the ability of a nearby tyrosine to be phosphorylated. They found that when pCMF was encoded at the −2 and −3 positions of the substrate peptide, phosphorylation was enhanced. As similar results were obtained when Asp and Glu were installed at these sites, follow up studies are needed to further refine c-Src specificity determinants.

Tyrosine phosphorylation also plays an important role in regulating bacterial protein signaling systems, including those involved in pathogenesis.285 In 2018, by installing pCMF at position Y47 into the E. coli metabolite-responsive LacI/GalR family regulator Cra, Robertson et al.249 showed that phosphorylation of Cra diminishes its ability to target DNA as a means to fine-tune the expression of virulence-associated genes. Such insights provide valuable leads for ongoing efforts to develop new antibiotics against bacterial pathogens that pose threats to human health.

3.3.8.3. pTyr-dependent regulation of protein degradation systems.

As with Ser and Thr, Tyr phosphorylation plays important roles in the regulation of ubiquitin function and protein degradation systems. When Hoppmann et al.258 developed the NpY encoding system (system pTyr-5, Table 5), they demonstrated via NMR that ubiquitin adopted altered conformations within the E51-Y59 loop when Tyr59 was phosphorylated. Furthermore, these changes decreased its ability to be conjugated to the E2 enzyme UBE2D3, implying that Tyr59 phosphorylation could play a negative regulatory role in protein ubiquitination processes.

Given this role that Tyr59 phosphorylation on ubiquitin played in regulating protein degradation, and the additional observation that this modification it is exclusively observed in cancer tissues, Zhou et al.280 sought to develop antibodies against it. Antibodies that recognize pTyr motifs are commonly generated using peptides, since generating authentic pTyr proteins is so challenging. But Zhou et al.280 surveyed all natural pTyr sites and found that pTyr occurs within an α-helix or β-sheet segment ~70% of the time, and hypothesized that antibodies raised against unstructured pTyr containing peptides will not be as effective against such pTyr sites in a folded protein. To overcome this, ubiquitin with authentic pTyr59 (located in a short helical segment) was first generated by encoding NpY followed by deprotection in acid and refolding. Then, starting with the Fab domain of commercially available pan-specific anti-pTyr antibody, they used phage display to find variants with improved affinity to ubiquitin-pTyr59 (called B1). Then a second antibody Fab fragment (B2) was engineered that bound the B1/ubiquitin-pTyr59 complex, but not B1 or ubiquitin-pTyr59 alone. Genetic fusion of the B1 scFV/B2 Fab proteins created a bi-specific chimeric antibody-derived protein that bound ubiquitin-pTyr59 with 0.5 nM affinity. Such an approach, when coupled with GCE to make pTyr-containing protein antigens, offers a power strategy to generate antibodies that can be highly specific for the native form of targeted pTyr proteins.

In other work using GCE to uncover roles that pTyr plays in protein degradation pathways, Stevens et al.284 sought to understand the how tyrosine phosphorylation regulates Parkin E3 ligase activity. Being unable to phosphorylate Parkin with a kinase at Tyr143, they instead encoded pCMF and found that the pCMF-containing Parkin had higher auto-ubiquitination activity caused by the release of its autoinhibitory domain from the catalytic domain. In this case, Parkin Tyr143 phosphorylation may serve to promote polyubiquitination of target proteins and thus the onset of mitophagy.

3.3.8.4. pTyr-dependent regulation of electron transport systems

Cytochrome c is a critical electron-transport chain protein which transfers electrons from the cytochrome bc1 complex to cytochrome c oxidase. Phosphorylation at Tyr48 was shown to impair oxidative phosphorylation,286 and because the kinase responsible for Tyr48 phosphorylation was not known, GCE became the key to understanding the effects of this phosphorylation through three studies spanning 2015 to 2022 in which pCMF was encoded into cytochrome C at Tyr48. In the initial experiments, pCMF at position 48 was shown to modulate the redox properties of cytochrome C and to lower the pKa value of the “alkaline transition” to physiological pH (i.e. the replacement of Met80 as an axial ligand with a neighboring Lys).276 Then, in 2018, an NMR study of the structure of pCMF48 cytochrome c revealed conformational shifts and enhanced dynamics around the pCMF that helped explain how phosphorylation impairs cytochrome c diffusion between respiratory complexes, enhances reactive oxygen species scavenging, and hinders caspase-dependent apoptosis.278 In 2022, Gomila et al.282 showed that the impaired diffusion of pCMF48 cytochrome C is caused by a strengthening of its interaction with cytochrome bc1. This stabilized interaction disrupts the “Gouy-Chapman conduit” that facilitates long-distance electron transfer between the two proteins, and sustains the high turnover rate of the electron transport chain by avoiding the need to establish a stable protein complex.

3.3.9. Future opportunities for pTyr GCE.

Among the current GCE systems for encoding pTyr (Table 5), each addressed one or more of the four primary issues discussed above – bioavailability, hydrolysis, engineering a pTyr-RS and EF-Tu, and ribosome compatibility – but no single system has solved all of these to produce a robust, facile strategy to encode pTyr in living cells or in a cell free system. More approachable pTyr GCE encoding strategies are needed to achieve a broader impact akin to the success observed with pSer systems.

3.3.9.1. Combining components from existing pTyr GCE systems.

An obvious first step toward a better pTyr encoding system is to simply combine each of the individual advancements made in previous systems and see how it works. For initial tests and optimizations, supplementing media with high concentrations of pTyr (6) or nhpTyr (28) (e.g. >10 mM) would be easier than synthesizing the di-peptide with lysine (36), since the free amino acids appear to get inside the cell, albeit inefficiently.71,87 Encoding could then be tested by expressing the Mj tRNACUATyr with either pYRS187 or the permissive CMFRS#172 alongside the pTyr specific EF-Tu variant (EF-pY87) and the modified 23S rRNA that better accommodates tRNAs amino-acylated with pTyr in E. coli cells257 that are void of the five known pTyr phosphatase genes (serB, pgpA, pphB, phoA, aphA)87.

Provided such a system displays efficiencies of pTyr encoding over background, each of the steps can be then individually optimized using standard, in-cell GFP fluorescence or chloramphenicol resistance assays. Standard life/death screening of RS and EF-Tu active site mutant libraries would likely be worth doing to generate more efficient variants that could function well at low intracellular pTyr concentrations, which would lower the burden of finding a highly efficient pTyr import method. Also, a recently published approach for an in vitro screening of RS libraries could be used to isolate novel variants capable of charging barcoded tRNA with pTyr, as this would decouple RS activity from other steps including cell import and ribosome compatibility.287 Use of orthogonal ribosomes would also permit exploring more extensive ribosomal modifications that improve pTyr tolerance without affecting the cell’s natural translational systems as well.288

3.3.9.2. Improving bioavailability of pTyr.

Increasing pTyr bioavailability can be achieved by (i) improving import mechanisms for free pTyr, (ii) developing ncAA analogs that can be converted to pTyr, or (iii) biosynthesizing pTyr. With regard to improving pTyr import, the di-peptide approach72 offers in-principle an attractive strategy, though challenges associated with the dipeptide synthesis may put people off as we found no further publications that have reported using it. Exploring alternative approaches such as using and evolving organo-phosphate transporters could prove helpful,71 as could developing active ncAA scavenging and import mechanisms as has been done for other ncAAs.220

With regard to developing protected forms of pTyr that could be useful, chemical modification of phosphate groups serves the dual purpose of facilitating cellular uptake by masking the negative charge and protecting against phosphatase-mediated hydrolysis once inside the cell.289 When selecting chemically protected pTyr derivatives that could be used for GCE encoding, two important considerations come to the forefront. First, the structural diversity of ncAAs that tRNA synthetases and the EF-Tu can accommodate is constrained by the need to preserve their backbone architecture and catalytic activity. This imposes limitations on the size and shape of encodable pTyr analogs. Second, the method chosen for deprotecting the pTyr analog must be carefully considered, particularly in context to the ncAA size and shape constraints imposed by limitations of the RSs and EF-Tu engineerability. Finding protected pTyr analogs that are both easily encoded and controllably deprotected under physiological conditions to yield native pTyr (or phosphonate isosteres) remains an outstanding challenge. Photo-protected pTyr analogs, with sufficiently small photo-protecting groups, would offer an ideal solution. Esterifying the phosphoryl group to mask its negative charge is a time-tested method to increase cellular uptake of phosphate-containing prodrugs.289 Depending on the nature of the phospho-ester substituents, they can be removed via non-specific cellular esterases for direct pTyr encoding,289 or they could be designed so that removal is achieved only via specific phospho-di/tri-esterases, such as those recently identified in diverse marine bacteria.290 In the latter case, deprotection after translational encoding might be feasible to avoid ribosome compatibility issues with native pTyr. Nonetheless, resolving this complex question demands a unique integration of chemistry, enzyme engineering, and GCE expertise.

Biosynthesis of pTyr would also offer an elegant solution to pTyr bioavailability issues, which would be possible by co-expressing a kinase for free tyrosine amino acid, however no natural kinases for free tyrosine are known. The diversity of biosynthetic transformations in plants, fungi, and bacteria, such as Streptomyces,65 offer promising avenues for identifying such enzymes, as demonstrated by the discoveries of biosynthetic pathways for nhpSer (20)49 and sTyr (33).254 Perhaps the specificity of existing small-molecule kinases, like phenol phosphorylase291, could be evolved to accept tyrosine as a substrate.

3.3.9.3. Improving stability of pTyr.

Perhaps the most difficult challenge is that of addressing the intracellular stability of pTyr, which would involve identifying all E. coli phosphatases with activity toward free pTyr amino acid and/or pTyr proteins. In the Fan et al.87 work, the screening of just 14 of the 90 annotated phosphatases in E. coli identified 5 with activity toward the free amino acid, so there may be quite a few more to be identified. After all the relevant phosphatases are identified (potentially using genetic or biochemical approaches), the next challenge would be to delete or alter problematic pTyr phosphatases while maintaining cell viability. Solving this challenge in E. coli would be quite useful, as it would make possible the expression of authentic pTyr containing proteins for in vitro studies. Nevertheless, this advance is not the key hurdle standing in the way of studies of pTyr function in eukaryotic cells because it would not solve the problem of pTyr hydrolysis that will occur during experiments inside eukaryotic cells and will confound the results. We expect that the advance needed for in cell-studies of pTyr function is the efficient encoding of a validated, faithful pTyr mimic that is hydrolysis-resistant.

3.3.9.4. Encoding stable and functional mimics of pTyr.

With the pCMF (32) and sTyr (33) amino acids being readily encodable and outperforming Asp/Glu as mimics, we consider them the best option for pTyr mimics currently available using GCE, even though they appear to be rather mediocre mimics. Available data indicate sTyr is a mimic that is on par with pCMF and may even be better in some contexts. Since no systematic studies have yet been done on either pCMF or sTyr as mimics for pTyr in a variety of contexts, such studies would be helpful for guiding when and how these two encodable mimics might be most fruitfully employed for gaining insight into pTyr functionality.

Many alternative mimics of pTyr have been evaluated, including non-phosphonate di-anionic, mono-anionic and uncharged amino acids (as previously reviewed238), and the evidence makes clear that the phosphonate mimics nhpTyr (28) and F2-nhpTyr (30) are the most reliable. But the tradeoff between mimicry and encodability is evident: encoding the more faithful pTyr mimetics like F2-nhpTyr and nhpTyr faces the same challenges as encoding pTyr except for the problem of hydrolysis, while more easily encoded but less faithful mimics introduce uncertainties regarding their biological relevance. Given the isosteric nature of the phosphonate mimics to pTyr, efforts focused on how best to improve pTyr GCE components as discussed above would be readily transferrable for nhpTyr/F2-nhpTyr encoding, and vice versa. It seems to us that just as for pTyr, overcoming the bioavailability hurdle for nhpTyr/F2-nhpTyr – whether through developing a biosynthetic pathway for the ncAA or a way to transport it or a precursor into the cell – is worth investing in, as that would allow all the other GCE components to be optimized.

4. Outlook on the future of GCE technologies for “less studied” phosphoamino acids

4.1. Challenges of encoding less studied phosphoamino acids

Increasing evidence underscores the biological significance of the phosphorylation of amino acids His, Lys, Arg, Cys, Asp, and Glu. While methods for incorporating most of these phosphorylated residues (all except pAsp and pGlu) into peptides have been reported, their encoding into proteins via GCE remains unexplored. Overcoming challenges similar to those encountered in pSer, pThr, and pTyr encoding, such as bioavailability issues due to their negative charge, is anticipated. Additionally, the need for novel orthogonal RS/tRNA pairs for each target ncAA, the inherent instability of phosphorylated His, Lys, Arg, and Cys under acidic conditions and the transient nature of pAsp and pGlu even at neutral pH, coupled with their susceptibility to phosphatase activity, pose significant hurdles. Given these challenges, we see the best path to successful GCE tools for the study of these phosphorylated residues will be through the use of stable mimics. To lay a foundation for any such efforts, in the next sections we briefly summarize for each of these phosphoamino acids what is known about its stability and about stable analogs previously evaluated for their ability to act as a mimic and that could, in principle, be encodable via GCE. In this section, we use the nomenclature “ncAA-mX” to denote each mimic, e.g. pLys-m1 in reference to pLys mimic 1. For an excellent, much more in-depth review on these phosphoamino acids, their mimics, synthetic routes, and their incorporation into peptides, see Marmelstein et al.20

4.2. Overview of less studied phosphoamino acids and their mimics

4.2.1. Phospho-histidine.

Phosphorylation on histidine can occur in two different isomeric forms: 1-pHis and 3-pHis (8 and 9, Fig. 16A) (also called, π-His and τ-His, respectively). Both isomers hydrolyze in seconds in 1 M HCl, but are semi-stable between pH 4 and 6. At physiologic pH ~7, 1-pHis (8) has a reported half-life of only 1 min, while 3-pHis (9) is more stable with a half-life of ~80 min.292 Stability of the pHis amino acid can differ from that of pHis in peptides and proteins. For example, the half-life of a pHis (unknown isomer) in histone H4 was estimated to be 2 h.293 One attempt to make a stable 3-pHis (8) mimic was through replacing the N-P bond with C-P, generating 4’-phosphono-2’-pyrrolyl-alanine analog (40, 3-pHis-m1, Fig. 16A). However, the nitrogen atom of the pyrrolyl ring became protonated, making it an obligatory hydrogen bond donor rather than an acceptor and indeed antibodies raised against this analog bound only to it and not pHis.294 A more electronically similar mimic, preserving the hydrogen bonding characteristics of pHis, has a furan ring in place of the pyrrolyl ring (41, 3-pHis- m2, Fig. 16A).294 Phosphoryltriazolalanine analogs of 1-pHis and 3-pHis have also been synthesized (38, 1-pHis-m1 and 42, 3-pHis-m3, respectively, Fig. 16A), and were successfully used to generate antibodies specific for pHis.295 More recently, with the aid of electrostatic surface potential calculations, a pyridine analog of 1-pHis (39, 1-pHis-m2, Fig. 16A) and a pyrazole analog of 3-pHis (43, 3-pHis-m4, Fig. 16A) were chosen as mimics, synthesized and used successfully to create isomer selective antibodies.296 Encoding chemically protected forms of these analogs (to increase bioavailability) may be feasible given the plasticity of the pyrrolysine tRNA-RS/tRNACUA pairs for encoding histidine analogs (e.g. methyl-histidine),297-299 as well as recent discoveries of natural histidine RS/tRNA pairs that are orthogonal in E. coli.62

Figure 16. Structures of “less studied” phosphoamino acids and their analogues.

Figure 16.

(A) Shown are the side-chain structures of two isomers of pHis (left, blue boxes) along with their mimics that have been studied (right, gray boxes). (B through E) The same as panel A, but for pArg, pLys, pAsp and pGlu, respectively.

4.2.2. Phospho-arginine.

Phosphorylation of arginine changes its net charge from +1 to −1, with the guanidinium group remaining protonated at physiologic pH.300 pArg (19) hydrolyzes extremely fast at low pH (1-3), but is kinetically stable at physiologic pH. Arginine phosphorylation was first discovered in bacteria, and bacterial Arg-specific kinases and pArg-specific phosphatases have been identified.300,301 Only recently has evidence supported the relevance of pArg in eukaryotes as well, with at least 139 unique phosphosites found on 116 proteins in the human proteome.19,302 To mitigate pArg hydrolysis by phosphatases, mimics in which the labile N-P bond is replaced with a stable C-P bond in the phosphoguanidinium moiety have been developed (44, pArg-m1, Fig. 14B), and it has successfully been used to raise antibodies that recognize pArg-containing proteins and peptides.303,304 These successes suggest reasonable mimicry of pArg by its phosphonate analog, in which case GCE methods to encode it (or its chemically protected forms) might be a tractable approach to advance studies of arginine phosphorylated proteins. Pyrrolysine-tRNA-RS/tRNACUA pairs could be engineered to encode these mimics, and alternate RSs to base selections on are some recently discovered natural Arg RS/tRNA pairs that are orthogonal in E. coli.62

4.2.3. Phospho-lysine.

Lysine kinase and pLys (17) phosphatase activities were discovered in rat liver almost 50 years ago, but almost nothing is known about the identity of these enzymes, their prevalence and sequence specificities.20,300 Only through recent advances in MS-based proteomics have studies demonstrated specific examples of lysine phosphorylation, with at least 140 sites of lysine phosphorylation on 125 human proteins identified.19 Just as for pArg (19), the phosphorylation of lysine changes its net charge from +1 to −1 at physiologic pH because the Nε remains protonated at physiologic pH (pKa ~10),20,305 and such charge reversals can readily impact protein structure and function. pLys (17) is rapidly hydrolyzed at acidic pH like other amino acids with phosphoramidate (N-P) bonds (t1/2 <1 min for N-(n-butyl) phosphoramidate), and is moderately stable at physiologic pH.306 Recently, a phosphonate analog of pLys with a stable C-P bond replacing the labile N-P bond was synthesized and incorporated into peptides (45, pLys-m1, Fig. 16C).305 The pKa2 of this phosphonate pLys analog was ~7.1 suggesting an overall net charge at physiologic pH that more closely mimics the −1 charge of authentic pLys compared to pLys-m2 (46) having a pKa2 of ~6.5, making it fully deprotonated and −2 charged at physiologic pH.305 The ability to site-specifically install pLys or its stable analogs into full-length proteins would open many new opportunities to study these signaling systems, including through generating pLys antibodies. And given the structural similarities of pLys and pyrrolysine, selecting Pyl-RS/tRNACUA pairs to install the phosphonate analog of pLys or its chemically protected forms may be relatively easy.

4.2.4. Phospho-aspartate and glutamate.

pAsp (11) and pGlu (13) are phosphoanhydrides that are unstable at acidic, neutral and basic pH’s; at neutral pH, 30% of pAsp is hydrolyzed in 30 min.20 pAsp (11) occurs as an enzymatic intermediate in a number of enzyme-catalyzed reactions as well as in two-component bacterial signaling systems,21 and recent proteomics work identified 410 unique sites of pAsp in 342 unique proteins from human cells.19 The biological function of pGlu (13) is more obscure, though it is apparently roughly as abundant in humans as pAsp (427 unique pGlu sites across 364 proteins).19 Little work has been done to develop stable analogs of either pAsp and pGlu, and to our knowledge they have not yet been incorporated into peptides. However, some work has been done to develop small molecule inhibitors of certain enzymes where pAsp is an enzymatic intermediate (47-57, pAsp-m1 through -m6, Fig. 16D), such as aspartate-semialdehyde dehydrogenase – a target for antibiotic development. For example, by replacing the mixed anhydride oxygen of pAsp with CF2 (48) or NH (50), stable pAsp analogs with low micromolar inhibition constants for aspartate-semialdehyde dehydrogenase.307,308 Similarly, pGlu (13) is an enzymatic intermediate in the conversion of glutamate to glutamine by glutamine synthetase, and inhibitors of E. coli glutamine synthetase with inhibition constant of ~3 mM and 0.3 mM were developed by replacing the mixed anhydride oxygen with CH2 (58) and CF2 (59), respectively (Fig. 16D).21 However, at physiologic pH the fluorinated pGlu analog 59 existed predominantly in the cyclic iminium form.309 Given these successes with inhibitor development, developing GCE tools to install these as well as other previously proposed stable analogs of pAsp and pGlu (47-60, Fig. 16D and 16E), and/or chemically protected mimics of pAsp and pGlu seems within reach. However, any efforts to encode these or other analogs will be more easily justified once the biological roles these PTMs play in human health and disease are established.

4.2.5. Phospho-cysteine.

pCys (15) is found as an enzymatic intermediate in a variety of enzymes, including protein phosphatases and the phosphoenolpyruvate-dependent phosphotransferase system.20 It is also involved in bacterial signaling systems as a regulatory PTM, including in transcriptional regulators that mediate virulence.310,311 It too is rapidly hydrolyzed at low pH, with a half-life of ~15 min at pH 3-4, but it is stable above pH 7.20 With only 55 identified sites of Cys phosphorylation on human proteins, it is the least commonly observed phosphoamino acid in the human proteome.19 Some work has been done to install pCys chemically into peptides and proteins, both to study its biological role as well as to use pCys as a mimetic for pSer and pThr.312-314 Though other strategies exist,315 the primary method to install pCys into proteins involves the rapid and selective conversion (at pH ~8 and 4 °C) of Cys to dehydroalanine, which can then undergo nucleophilic addition by sodium thiophosphate to produce pCys.33,316 Though elegant in that it can be performed on folded proteins, the method requires that all off-target cysteines be mutated, and the reaction results in a mixture of D- and L-pCys stereoisomers. The encoding of pCys would offer the opportunity to install pCys without these concerns and it may well be feasible to achieve by evolving current pSer GCE systems, provided that sufficient pCys can get inside in the cell and remain stable. We could not find any reports describing development of pCys-specific stable analogs, in which case the most useful mimics might well be pSer (2) and nhpSer (20). Nevertheless, efforts to develop pCys GCE technologies will be boosted as its physiologic roles are better understood.

5. Concluding Remarks

We hope this review makes a compelling case that GCE technologies can provide a general solution for accelerating research to advance our understanding phospho-protein function. Further, we hope that the highlighted “first-fruits” and the summaries of the GCE systems make clear that some of the systems available today are highly effective, and that this will inspire more researchers to adopt these systems. In particular, we consider that GCE systems producing pSer (2), nhpSer (20) and pThr (4) proteins in E. coli are ready for widespread use, while there is a need for further development of systems for pThr and pTyr mimics. In terms of pTyr mimics, the GCE systems for encoding pCMF (32) and sTyr (33) seem effective in both E. coli and mammalian cells, but work is needed to assess the fidelity of their mimicry in various contexts. Technologies for encoding stable phosphonate mimics of pSer (20-22), pThr (25-27) and pTyr (28-30) in eukaryotic cells are still ripe for development and optimization.

By eliminating the requirement of protein kinase “writers” or complex semi-synthetic approaches to install phosphorylated amino acids onto proteins, in principle any full-length phospho-protein variant or a high-quality stable phospho-protein mimic can be generated and studied. We anticipate these advances will spawn a new era of research into how specific phospho-protein variants are involved in physiology and disease, as well as stimulate work to therapeutically target specific phospho-proteins with small molecules, synthetic peptides or antibodies. Recombinant proteins and peptides with phospho-mimics can themselves be therapeutics.

Finally, we gratefully acknowledge that the GCE tools we have today for phospho-proteins are the result of work from many innovative researchers who have invested substantial efforts and synergistically built on one another’s advances. The creative work involved includes both conceiving and developing the pioneering proof-of-concept technologies, and also the many efforts that go into refining and optimizing these translational systems and their expression methods to yield robust, detailed protocols so that recombinant, biologically relevant phospho-proteins can be accessible to all researchers, including those who are not experts in GCE.

Supplementary Material

Supp Material
Supp Table 1

ACKNOWLEDGEMENTS

This work was supported by the GCE4All Biomedical Technology Optimization and Dissemination Center supported by National Institute of General Medical Science grant RM1-GM144227 (to RAM). ChatGPT was used at early stages of manuscript writing to improve sentence structure.

ABBREVIATIONS

GCE

Genetic Code Expansion

RS

amino-acyl tRNA-synthetase

SepRS

phosphoserine amino-acyl tRNA-synthetase

Pyl

pyrrolysine

pSer

phosphoserine

pThr

phosphothreonine

pTyr

phosphotyrosine

pLys

phospholysine

pHis

phosphohistidine

pArg

phosphoarginine

pAsp

phosphoaspartate

pGlu

phosphoglutamate

pCys

phosphocysteine

nhpSer

non-hydrolyzable phosphoserine

nhpThr

non-hydrolyzable phosphothreonine

nhpTyr

non-hydrolyzable phosphotyrosine

Pmp

phosphono-methyl-phenylalanine

POI

protein of interest

ncAA

non-canonical amino acid

RF1

release factor 1

pCMF

para-carboxymethyl-phenylalanine

sTyr

sulfo-tyrosine

azido-Phe

para-azido-phenylalanine

NpY

bis(dimethylamino)phospho-tyrosine

Amido-Tyr

para-phosphoramidate-phenylalanine

PDB

Protein Data Bank

MS

mass spectrometry

NCS

near-cognate suppression

NCL

Native chemical ligation

EPL

Expressed protein ligation

Biographies

Michael C. Allen Michael C. Allen graduated from the University of Washington with his B.S. in biochemistry in 2018, where he participated in undergraduate research under Dr. Edith Wang, focusing on the function of the transcription factor complex TFIID, and how mutations in its TAF subunits contribute to cancer, neuronal developmental disease and degenerative disorders. In 2023, after a 5-year period of working in a different field, he joined the GCE4All NIH National Center under Drs. Ryan Mehl and Rick Cooley where he works as a Faculty Research Assistant. Michael’s current research is specifically focused on using and optimizing genetic code expansion tools to incorporate non-canonical amino acids into antibody fragments, enhancing their capabilities for use in various research applications.

P. Andrew Karplus Andy Karplus is Distinguished Professor Emeritus of Biochemistry & Biophysics at Oregon State University and Director of Communications of the GCE4All NIH National Center. He received his Ph.D. in 1984 from the University of Washington in Seattle and after post-doctoral work with Dr. Georg Schulz in Germany, he was on the faculty at Cornell University (1988-1998) and then Oregon State University (1998-2023). His research expertise is in the area of structural biology, with one main focus being solving detailed protein structures by X-ray crystallography, and interpreting those structures in light of functional and evolutionary information to figure out how they work.

Ryan A. Mehl Ryan Mehl is Professor of Biochemistry & Biophysics at Oregon State University and Director of the GCE4All NIH National Center. Dr. Mehl received his Ph.D. from Cornell University in 2001 in organic synthesis and mechanistic enzymology under Tadhg Begley, and was trained in chemical biology and Genetic Code Expansion under Peter Schultz at Scripps Research Institute. He joined the Chemistry Faculty at Franklin & Marshall College (2002-2011) and focused on studying protein function with GCE tools. In 2011, he joined Biochemistry & Biophysics Department at Oregon State University to increase researcher access to GCE technology, understand the biological role of tyrosine nitration, and develop ultra-fast and quantitative genetically encoded bioorthogonal ligations that function in live cells. As the GCE4All Center Director, he focused on optimizing GCE technologies and disseminating them to the broader community.

Richard B. Cooley Rick Cooley is an Assistant Professor (Sr Research) and Assistant Director of the GCE4All NIH National Center in the Biochemistry & Biophysics Department at Oregon State University. He graduated from Middlebury College (VT) with a B.S. in Chemistry in 2004 and received his Ph.D. degree in structural biology and protein biochemistry in 2011 from OSU under the supervision of Drs. Andy Karplus and Dan Arp. He joined Ryan Mehl’s lab at OSU in 2011 as post-doctoral fellow where he was trained in Genetic Code Expansion techniques and in 2012 he joined Holger Sondermann’s lab at Cornell University for a second post-doctoral fellowship studying molecular mechanisms of bacterial biofilm formation. In 2016 he returned to OSU as faculty where he works with Dr. Mehl to develop GCE tools for installing post-translational modifications. His research program centers on the development and application of these GCE tools to study structural mechanisms by which post-translational modifications regulate protein function.

Footnotes

CONFLICTS OF INTEREST

The authors declare no conflicts of interest.

SUPPORTING INFORMATION AVAILABLE: The following files are available free of charge.

Supporting Table 1_gamma-oxygen interactions.xlsx. Table listing unique structures in the Protein Data Bank as of October 1, 2023 with <95% sequence identity containing pSer or pThr within a poly-peptide chain (i.e. non-ligands) in which the γ-oxygen of pSer/pThr forms a hydrogen bond with a protein donor atom.

REFERENCES

  • (1).Ramazi S; Zahiri J Posttranslational modifications in proteins: resources, tools and prediction methods. Database (Oxford) 2021, 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (2).Mukai T; Lajoie MJ; Englert M; Soll D Rewriting the genetic code. Annu. Rev. Microbiol 2017, 71, 557. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (3).Chin JW Expanding and reprogramming the genetic code. Nature 2017, 550, 53. [DOI] [PubMed] [Google Scholar]
  • (4).Mehl RA The GCE4All Center: Unleashing the potential of genetic code expansion for biomedical Rresearch. NIH National Institute of General Medical Sciences, 2022. RM1–GM144227 [Google Scholar]
  • (5).Manning G; Whyte DB; Martinez R; Hunter T; Sudarsanam S The protein kinase complement of the human genome. Science 2002, 298, 1912. [DOI] [PubMed] [Google Scholar]
  • (6).Fabbro D; Cowan-Jacob SW; Moebitz H Ten things you should know about protein kinases: IUPHAR Review 14. Br. J. Pharmacol 2015, 172, 2675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (7).Needham EJ; Parker BL; Burykin T; James DE; Humphrey SJ Illuminating the dark phosphoproteome. Sci. Signal 2019, 12. [DOI] [PubMed] [Google Scholar]
  • (8).Johnson JL; Yaron TM; Huntsman EM; Kerelsky A; Song J; Regev A; Lin TY; Liberatore K; Cizin DM; Cohen BM et al. An atlas of substrate specificities for the human serine/threonine kinome. Nature 2023, 613, 759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (9).Shi Y Serine/threonine phosphatases: mechanism through structure. Cell 2009, 139, 468. [DOI] [PubMed] [Google Scholar]
  • (10).Masterson LR; Cheng C; Yu T; Tonelli M; Kornev A; Taylor SS; Veglia G Dynamics connect substrate recognition to catalysis in protein kinase A. Nat. Chem. Biol 2010, 6, 821. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (11).Ubersax JA; Ferrell JE Jr. Mechanisms of specificity in protein phosphorylation. Nat Rev Mol. Cell. Biol 2007, 8, 530. [DOI] [PubMed] [Google Scholar]
  • (12).Iakoucheva LM; Radivojac P; Brown CJ; O'Connor TR; Sikes JG; Obradovic Z; Dunker AK The importance of intrinsic disorder for protein phosphorylation. Nucleic Acids Res. 2004, 32, 1037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (13).Bah A; Forman-Kay JD Modulation of intrinsically disordered protein function by post-translational modifications. J. Biol. Chem 2016, 291, 6696. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (14).Koike R; Amano M; Kaibuchi K; Ota M Protein kinases phosphorylate long disordered regions in intrinsically disordered proteins. Protein Sci. 2020, 29, 564. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (15).Honkanen RE; Golden T Regulators of serine/threonine protein phosphatases at the dawn of a clinical era? Curr. Med. Chem 2002, 9, 2055. [DOI] [PubMed] [Google Scholar]
  • (16).Sacco F; Perfetto L; Castagnoli L; Cesareni G The human phosphatase interactome: An intricate family portrait. FEBS Lett. 2012, 586, 2732. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (17).Ardito F; Giuliani M; Perrone D; Troiano G; Lo Muzio L The crucial role of protein phosphorylation in cell signaling and its use as targeted therapy (Review). Int. J. Mol. Med 2017, 40, 271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (18).Hornbeck PV; Zhang B; Murray B; Kornhauser JM; Latham V; Skrzypek E PhosphoSitePlus, 2014: mutations, PTMs and recalibrations. Nucleic Acids Res. 2015, 43, D512. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (19).Hardman G; Perkins S; Brownridge PJ; Clarke CJ; Byrne DP; Campbell AE; Kalyuzhnyy A; Myall A; Eyers PA; Jones AR et al. Strong anion exchange-mediated phosphoproteomics reveals extensive human non-canonical phosphorylation. EMBO J. 2019, 38, e100847. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (20).Marmelstein AM; Moreno J; Fiedler D Chemical Approaches to studying labile amino acid phosphorylation. Top Curr. Chem. (Cham.) 2017, 375, 22. [DOI] [PubMed] [Google Scholar]
  • (21).Attwood PV; Besant PG; Piggott MJ Focus on phosphoaspartate and phosphoglutamate. Amino Acids 2011, 40, 1035. [DOI] [PubMed] [Google Scholar]
  • (22).Piggott MJ; Attwood PV Focus on O-phosphohydroxylysine, O-phosphohydroxyproline, N (1)-phosphotryptophan and S-phosphocysteine. Amino Acids 2017, 49, 1309. [DOI] [PubMed] [Google Scholar]
  • (23).Sharma K; D'Souza RC; Tyanova S; Schaab C; Wisniewski JR; Cox J; Mann M Ultradeep human phosphoproteome reveals a distinct regulatory nature of Tyr and Ser/Thr-based signaling. Cell Rep. 2014, 8, 1583. [DOI] [PubMed] [Google Scholar]
  • (24).Aebersold R; Agar JN; Amster IJ; Baker MS; Bertozzi CR; Boja ES; Costello CE; Cravatt BF; Fenselau C; Garcia BA et al. How many human proteoforms are there? Nat. Chem. Biol 2018, 14, 206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (25).Berman HM; Westbrook J; Feng Z; Gilliland G; Bhat TN; Weissig H; Shindyalov IN; Bourne PE The Protein Data Bank. Nucleic Acids Res. 2000, 28, 235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (26).Seeliger MA; Young M; Henderson MN; Pellicena P; King DS; Falick AM; Kuriyan J High yield bacterial expression of active c-Abl and c-Src tyrosine kinases. Protein Sci. 2005, 14, 3135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (27).Albanese SK; Parton DL; Isik M; Rodriguez-Laureano L; Hanson SM; Behr JM; Gradia S; Jeans C; Levinson NM; Seeliger MA et al. An open library of human kinase domain constructs for automated bacterial expression. Biochemistry 2018, 57, 4675. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (28).Bhoir S; Shaik A; Thiruvenkatam V; Kirubakaran S High yield bacterial expression, purification and characterisation of bioactive Human Tousled-like Kinase 1B involved in cancer. Sci. Rep 2018, 8, 4796. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (29).Thomas PD; Ebert D; Muruganujan A; Mushayahama T; Albou LP; Mi H PANTHER: Making genome-scale phylogenetics accessible to all. Protein Sci. 2022, 31, 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (30).Bilbrough T; Piemontese E; Seitz O Dissecting the role of protein phosphorylation: a chemical biology toolbox. Chem. Soc. Rev 2022, 51, 5691. [DOI] [PubMed] [Google Scholar]
  • (31).McMurray JS; Coleman D. R. t.; Wang W; Campbell ML The synthesis of phosphopeptides. Biopolymers 2001, 60, 3. [DOI] [PubMed] [Google Scholar]
  • (32).Perich JW; Johns RB Di-tert-butyl N,N-eiethylphosphoramidite and eibenzyl N,N-eiethylphosphoramidite - highly reactive reagents for the phosphite-triester phosphorylation of serine-containing peptides. Tetrahedron Letters 1988, 29, 2369. [Google Scholar]
  • (33).Bernardes GJ; Chalker JM; Errey JC; Davis BG Facile conversion of cysteine and alkyl cysteines to dehydroalanine on protein surfaces: versatile and switchable access to functionalized proteins. J. Am. Chem. Soc 2008, 130, 5052. [DOI] [PubMed] [Google Scholar]
  • (34).Attwood PV; Ludwig K; Bergander K; Besant PG; Adina-Zada A; Krieglstein J; Klumpp S Chemical phosphorylation of histidine-containing peptides based on the sequence of histone H4 and their dephosphorylation by protein histidine phosphatase. Biochim. Biophys. Acta 2010, 1804, 199. [DOI] [PubMed] [Google Scholar]
  • (35).Medzihradszky KF; Phillipps NJ; Senderowicz L; Wang P; Turck CW Synthesis and characterization of histidine-phosphorylated peptides. Protein Sci. 1997, 6, 1405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (36).Hohenester UM; Ludwig K; Konig S Chemical phosphorylation of histidine residues in proteins using potassium phosphoramidate -- a tool for the analysis of acid-labile phosphorylation. Curr. Drug Deliv. 2013, 10, 58. [DOI] [PubMed] [Google Scholar]
  • (37).Kafarski P Phosphonopeptides containing free phosphonic groups: recent advances. RSC Adv. 2020, 10, 25898. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (38).Grunhaus D; Friedler A; Hurevich M Automated synthesis of heavily phosphorylated peptides. Eur. J. Org. Chem 2021, 2021, 3737. [Google Scholar]
  • (39).Dawson PE; Muir TW; Clark-Lewis I; Kent SB Synthesis of proteins by native chemical ligation. Science 1994, 266, 776. [DOI] [PubMed] [Google Scholar]
  • (40).Muir TW; Sondhi D; Cole PA Expressed protein ligation: a general method for protein engineering. Proc. Natl. Acad. Sci. U S A 1998, 95, 6705. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (41).Thompson RE; Muir TW Chemoenzymatic Semisynthesis of Proteins. Chem. Rev 2020, 120, 3051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (42).Jimenez JL; Hegemann B; Hutchins JR; Peters JM; Durbin R A systematic comparative and structural analysis of protein phosphorylation sites based on the mtcPTM database. Genome Biol. 2007, 8, R90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (43).Chu N; Salguero AL; Liu AZ; Chen Z; Dempsey DR; Ficarro SB; Alexander WM; Marto JA; Li Y; Amzel LM et al. Akt Kinase Activation Mechanisms Revealed Using Protein Semisynthesis. Cell 2018, 174, 897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (44).Diaz Galicia ME; Aldehaiman A; Hong S; Arold ST; Grunberg R Methods for the recombinant expression of active tyrosine kinase domains: Guidelines and pitfalls. Methods Enzymol. 2019, 621, 131. [DOI] [PubMed] [Google Scholar]
  • (45).Shrestha A; Hamilton G; O'Neill E; Knapp S; Elkins JM Analysis of conditions affecting auto-phosphorylation of human kinases during expression in bacteria. Protein Expr. Purif 2012, 81, 136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (46).Reinhardt R; Leonard TA A critical evaluation of protein kinase regulation by activation loop autophosphorylation. Elife 2023, 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (47).McKenna M; Balasuriya N; Zhong S; Li SS; O'Donoghue P Phospho-form specific substrates of protein kinase B (AKT1). Front. Bioeng. Biotechnol 2020, 8, 619252. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (48).Hunter T Why nature chose phosphate to modify proteins. Philos. Trans. R Soc. Lond. B Biol. Sci 2012, 367, 2513. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (49).Zhu P; Stanisheuski S; Franklin R; Vogel A; Vesely CH; Reardon P; Sluchanko NN; Beckman JS; Karplus PA; Mehl RA et al. Autonomous synthesis of functional, permanently phosphorylated proteins for defining the interactome of monomeric 14-3-3zeta. ACS Cent. Sci 2023, 9, 816. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (50).Balasuriya N; Kunkel MT; Liu XG; Biggar KK; Li SSC; Newton AC; O'Donoghue P Genetic code expansion and live cell imaging reveal that Thr-308 phosphorylation is irreplaceable and sufficient for Akt1 activity. J. Biol. Chem 2018, 293, 10744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (51).Kozelekova A; Naplavova A; Brom T; Gasparik N; Simek J; Houser J; Hritz J Phosphorylated and phosphomimicking variants may differ-a case study of 14-3-3 protein. Front. Chem 2022, 10, 835733. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (52).Pedersen SW; Albertsen L; Moran GE; Levesque B; Pedersen SB; Bartels L; Wapenaar H; Ye F; Zhang M; Bowen ME et al. Site-specific phosphorylation of PSD-95 PDZ domains reveals fine-tuned regulation of protein-protein interactions. ACS Chem. Biol 2017, 12, 2313. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (53).Herhaus L; Dikic I Expanding the ubiquitin code through post-translational modification. Embo Reports 2015, 16, 1071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (54).Alessi DR; Cuenda A; Cohen P; Dudley DT; Saltiel AR PD 098059 is a specific inhibitor of the activation of mitogen-activated protein kinase kinase in vitro and in vivo. J. Biol. Chem 1995, 270, 27489. [DOI] [PubMed] [Google Scholar]
  • (55).Ordureau A; Heo JM; Duda DM; Paulo JA; Olszewski JL; Yanishevski D; Rinehart J; Schulman BA; Harper JW Defining roles of PARKIN and ubiquitin phosphorylation by PINK1 in mitochondrial quality control using a ubiquitin replacement strategy. Proc. Natl. Acad. Sci. U S A 2015, 112, 6637. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (56).Gogl G; Jane P; Caillet-Saguy C; Kostmann C; Bich G; Cousido-Siah A; Nyitray L; Vincentelli R; Wolff N; Nomine Y et al. Dual Specificity PDZ- and 14-3-3-Binding Motifs: A Structural and Interactomics Study. Structure 2020, 28, 747. [DOI] [PubMed] [Google Scholar]
  • (57).Somale D; Di Nardo G; di Blasio L; Puliafito A; Vara-Messler M; Chiaverina G; Palmiero M; Monica V; Gilardi G; Primo L et al. Activation of RSK by phosphomimetic substitution in the activation loop is prevented by structural constraints. Sci. Rep 2020, 10, 591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (58).Wang L; Brock A; Herberich B; Schultz PG Expanding the genetic code of Escherichia coli. Science 2001, 292, 498. [DOI] [PubMed] [Google Scholar]
  • (59).Schultz P Expanding the genetic code. Protein Sci. 2023, 32, e4488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (60).Dumas A; Lercher L; Spicer CD; Davis BG Designing logical codon reassignment - Expanding the chemistry in biology. Chem. Sci 2015, 6, 50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (61).Krahn N; Tharp JM; Crnkovic A; Soll D Engineering aminoacyl-tRNA synthetases for use in synthetic biology. Enzymes 2020, 48, 351. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (62).Cervettini D; Tang S; Fried SD; Willis JCW; Funke LFH; Colwell LJ; Chin JW Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol 2020, 38, 989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (63).Andrews J; Gan Q; Fan C "Not-so-popular" orthogonal pairs in genetic code expansion. Protein Sci. 2023, 32, e4559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (64).Bednar RM; Karplus PA; Mehl RA Site-specific dual encoding and labeling of proteins via genetic code expansion. Cell Chem. Biol 2023, 30, 343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (65).Hedges JB; Ryan KS Biosynthetic pathways to nonproteinogenic alpha-amino Acids. Chem. Rev 2020, 120, 3161. [DOI] [PubMed] [Google Scholar]
  • (66).Park HS; Hohn MJ; Umehara T; Guo LT; Osborne EM; Benner J; Noren CJ; Rinehart J; Soll D Expanding the genetic code of Escherichia coli with phosphoserine. Science 2011, 333, 1151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (67).Guo LT; Wang YS; Nakamura A; Eiler D; Kavran JM; Wong M; Kiessling LL; Steitz TA; O'Donoghue P; Soll D Polyspecific pyrrolysyl-tRNA synthetases from directed evolution. Proc. Natl. Acad. Sci. U S A 2014, 111, 16724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (68).Rauch BJ; Porter JJ; Mehl RA; Perona JJ Improved incorporation of noncanonical amino acids by an engineered tRNA(Tyr) suppressor. Biochemistry 2016, 55, 618. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (69).Kamtekar S; Hohn MJ; Park HS; Schnitzbauer M; Sauerwald A; Soll D; Steitz TA Toward understanding phosphoseryl-tRNACys formation: the crystal structure of Methanococcus maripaludis phosphoseryl-tRNA synthetase. Proc. Natl. Acad. Sci. U S A 2007, 104, 2620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (70).Hauenstein SI; Hou YM; Perona JJ The homotetrameric phosphoseryl-tRNA synthetase from Methanosarcina mazei exhibits half-of-the-sites activity. J. Biol. Chem 2008, 283, 21997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (71).Steinfeld JB; Aerni HR; Rogulina S; Liu Y; Rinehart J Expanded cellular amino acid pools containing phosphoserine, phosphothreonine, and phosphotyrosine. ACS Chem. Biol 2014, 9, 1104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (72).Luo X; Fu G; Wang RE; Zhu X; Zambaldo C; Liu R; Liu T; Lyu X; Du J; Xuan W et al. Genetically encoding phosphotyrosine and its nonhydrolyzable analog in bacteria. Nat. Chem. Biol 2017, 13, 845. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (73).Zhang MS; Brunner SF; Huguenin-Dezot N; Liang AD; Schmied WH; Rogerson DT; Chin JW Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat. Methods 2017, 14, 729. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (74).Wanner BL Gene regulation by phosphate in enteric bacteria. J. Cell Biochem 1993, 51, 47. [DOI] [PubMed] [Google Scholar]
  • (75).Wanner BL; Metcalf WW Molecular genetic studies of a 10.9-kb operon in Escherichia coli for phosphonate uptake and biodegradation. FEMS Microbiol. Lett 1992, 100, 133. [DOI] [PubMed] [Google Scholar]
  • (76).Melnikov SV; Soll D Aminoacyl-tRNA synthetases and tRNAs for an expanded genetic code: what makes them orthogonal? Int. J. Mol. Sci 2019, 20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (77).Willis JCW; Chin JW Mutually orthogonal pyrrolysyl-tRNA synthetase/tRNA pairs. Nat. Chem 2018, 10, 831. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (78).Cooley RB; Feldman JL; Driggers CM; Bundy TA; Stokes AL; Karplus PA; Mehl RA Structural basis of improved second-generation 3-nitro-tyrosine tRNA synthetases. Biochemistry 2014, 53, 1916. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (79).Lee S; Oh S; Yang A; Kim J; Soll D; Lee D; Park HS A facile strategy for selective incorporation of phosphoserine into histones. Angew. Chem. Int. Ed. Engl 2013, 52, 5771. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (80).O'Donoghue P; Prat L; Heinemann IU; Ling J; Odoi K; Liu WR; Soll D Near-cognate suppression of amber, opal and quadruplet codons competes with aminoacyl-tRNAPyl for genetic code expansion. FEBS Lett. 2012, 586, 3931. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (81).Aerni HR; Shifman MA; Rogulina S; O'Donoghue P; Rinehart J Revealing the amino acid composition of proteins within an expanded genetic code. Nucleic Acids Res 2015, 43, e8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (82).Beyer JN; Hosseinzadeh P; Gottfried-Lee I; Van Fossen EM; Zhu P; Bednar RM; Karplus PA; Mehl RA; Cooley RB Overcoming near-cognate suppression in a release factor 1-deficient host with an improved nitro-tyrosine tRNA synthetase. J. Mol. Biol 2020, 432, 4690. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (83).Eddins AJ; Bednar RM; Jana S; Pung AH; Mbengi L; Meyer K; Perona JJ; Cooley RB; Karplus PA; Mehl RA Truncation-free genetic code expansion with tetrazine amino acids for quantitative protein ligations. Bioconjug. Chem 2023, 34, 2243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (84).Zhu P; Gafken PR; Mehl RA; Cooley RB A highly versatile expression system for the production of multiply phosphorylated proteins. ACS Chem. Biol 2019, 14, 1564. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (85).Nishi H; Hashimoto K; Panchenko AR Phosphorylation in protein-protein binding: effect on stability and function. Structure 2011, 79, 1807. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (86).Costa S; Almeida A; Castro A; Domingues L Fusion tags for protein solubility, purification and immunogenicity in Escherichia coli: the novel Fh8 system. Front. Microbiol 2014, 5, 63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (87).Fan C; Ip K; Soll D Expanding the genetic code of Escherichia coli with phosphotyrosine. FEBS Lett. 2016, 590, 3040. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (88).Vesely CEL; Reardon PN; Yu Z; Barbar E; Mehl RA; Cooley RB Accessing isotopically labeled proteins containing genetically encoded phosphoserine for NMR with optimized expression conditions. J. Biol. Chem 2022, 298, 102613. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (89).Beranek V; Reinkemeier CD; Zhang MS; Liang AD; Kym G; Chin JW Genetically encoded protein phosphorylation in mammalian Cells. Cell Chem. Biol 2018, 25, 1067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (90).Zheng Y; Lajoie MJ; Italia JS; Chin MA; Church GM; Chatterjee A Performance of optimized noncanonical amino acid mutagenesis systems in the absence of release factor 1. Mol. Biosyst 2016, 12, 1746. [DOI] [PubMed] [Google Scholar]
  • (91).Amiram M; Haimovich AD; Fan C; Wang YS; Aerni HR; Ntai I; Moonan DW; Ma NJ; Rovner AJ; Hong SH et al. Evolution of translation machinery in recoded bacteria enables multi-site incorporation of nonstandard amino acids. Nat. Biotechnol 2015, 33, 1272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (92).Hohsaka T; Kajihara D; Ashizuka Y; Murakami H; Sisido M Efficient incorporation of nonnatural amino acids with large aromatic groups into streptavidin in in vitro protein synthesizing systems. J. Am. Chem. Soc 1999, 121, 34. [Google Scholar]
  • (93).Makarov M; Rocha ACS; Krystufek R; Cherepashuk I; Dzmitruk V; Charnavets T; Faustino AM; Lebl M; Fujishima K; Fried SD et al. Early selection of the amino acid alphabet was adaptively shaped by biophysical constraints of foldability. J. Am. Chem. Soc 2023, 145, 5320. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (94).Pedelacq JD; Cabantous S; Tran T; Terwilliger TC; Waldo GS Engineering and characterization of a superfolder green fluorescent protein. Nat. Biotechnol 2006, 24, 79. [DOI] [PubMed] [Google Scholar]
  • (95).Stokes AL; Miyake-Stoner SJ; Peeler JC; Nguyen DP; Hammer RP; Mehl RA Enhancing the utility of unnatural amino acid synthetases by manipulating broad substrate specificity. Mol. Biosyst 2009, 5, 1032. [DOI] [PubMed] [Google Scholar]
  • (96).Balleza E; Kim JM; Cluzel P Systematic characterization of maturation time of fluorescent proteins in living cells. Nat. Methods 2018, 15, 47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (97).Pirman NL; Barber KW; Aerni HR; Ma NJ; Haimovich AD; Rogulina S; Isaacs FJ; Rinehart J A flexible codon in genomically recoded Escherichia coli permits programmable protein phosphorylation. Nat. Commun 2015, 6, 8130. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (98).George S; Aguirre JD; Spratt DE; Bi Y; Jeffery M; Shaw GS; O'Donoghue P Generation of phospho-ubiquitin variants by orthogonal translation reveals codon skipping. FEBS Lett. 2016, 590, 1530. [DOI] [PubMed] [Google Scholar]
  • (99).Miyake-Stoner SJ; Refakis CA; Hammill JT; Lusic H; Hazen JL; Deiters A; Mehl RA Generating permissive site-specific unnatural aminoacyl-tRNA synthetases. Biochemistry 2010, 49, 1667. [DOI] [PubMed] [Google Scholar]
  • (100).Young DD; Young TS; Jahnz M; Ahmad I; Spraggon G; Schultz PG An evolved aminoacyl-tRNA synthetase with atypical polysubstrate specificity. Biochemistry 2011, 50, 1894. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (101).Cooley RB; Karplus PA; Mehl RA Gleaning unexpected fruits from hard-won synthetases: probing principles of permissivity in non-canonical amino acid-tRNA synthetases. Chembiochem 2014, 15, 1810. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (102).Yanagisawa T; Kuratani M; Seki E; Hino N; Sakamoto K; Yokoyama S Structural basis for genetic-code expansion with bulky lysine derivatives by an engineered pyrrolysyl-tRNA synthetase. Cell Chem. Biol 2019, 26, 936. [DOI] [PubMed] [Google Scholar]
  • (103).Schultz KC; Supekova L; Ryu Y; Xie J; Perera R; Schultz PG A genetically encoded infrared probe. J. Am. Chem. Soc 2006, 128, 13984. [DOI] [PubMed] [Google Scholar]
  • (104).Miyake-Stoner SJ; Miller AM; Hammill JT; Peeler JC; Hess KR; Mehl RA; Brewer SH Probing protein folding using site-specifically encoded unnatural amino acids as FRET donors with tryptophan. Biochemistry 2009, 48, 5953. [DOI] [PubMed] [Google Scholar]
  • (105).Wu Z; Tiambeng TN; Cai W; Chen B; Lin Z; Gregorich ZR; Ge Y Impact of phosphorylation on the mass spectrometry quantification of intact phosphoproteins. Anal. Chem 2018, 90, 4935. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (106).Steen H; Jebanathirajah JA; Rush J; Morrice N; Kirschner MW Phosphorylation analysis by mass spectrometry: myths, facts, and the consequences for qualitative and quantitative measurements. Mol. Cell Proteomics 2006, 5, 172. [DOI] [PubMed] [Google Scholar]
  • (107).Chait BT Chemistry. Mass spectrometry: bottom-up or top-down? Science 2006, 314, 65. [DOI] [PubMed] [Google Scholar]
  • (108).Kinoshita E; Kinoshita-Kikuta E; Koike T History of Phos-tag technology for phosphoproteomics. J. Proteomics 2022, 252, 104432. [DOI] [PubMed] [Google Scholar]
  • (109).Kinoshita E; Kinoshita-Kikuta E; Takiyama K; Koike T Phosphate-binding tag, a new tool to visualize phosphorylated proteins. Mol. Cell Proteomics 2006, 5, 749. [DOI] [PubMed] [Google Scholar]
  • (110).O'Donoghue L; Smolenski A Analysis of protein phosphorylation using Phos-tag gels. J Proteomics 2022, 259, 104558. [DOI] [PubMed] [Google Scholar]
  • (111).Kinoshita E; Kinoshita-Kikuta E; Matsubara M; Yamada S; Nakamura H; Shiro Y; Aoki Y; Okita K; Koike T Separation of phosphoprotein isotypes having the same number of phosphate groups using phosphate-affinity SDS-PAGE. Proteomics 2008, 8, 2994. [DOI] [PubMed] [Google Scholar]
  • (112).Kinoshita E; Kinoshita-Kikuta E; Koike T Separation and detection of large phosphoproteins using Phos-tag SDS-PAGE. Nat. Protoc 2009, 4, 1513. [DOI] [PubMed] [Google Scholar]
  • (113).Kaufmann H; Bailey JE; Fussenegger M Use of antibodies for detection of phosphorylated proteins separated by two-dimensional gel electrophoresis. Proteomics 2001, 1, 194. [DOI] [PubMed] [Google Scholar]
  • (114).Mandell JW Phosphorylation state-specific antibodies: applications in investigative and diagnostic pathology. Am. J. Pathol 2003, 163, 1687. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (115).Steinberg TH Protein gel staining methods: an introduction and overview. Guide to Protein Purification, Second Edition 2009, 463, 541. [DOI] [PubMed] [Google Scholar]
  • (116).Sundaram RK; Balasubramaniyan N; Sundaram P Protein stains and applications. Methods Mol. Biol 2012, 869, 451. [DOI] [PubMed] [Google Scholar]
  • (117).Sauerwald A; Zhu W; Major TA; Roy H; Palioura S; Jahn D; Whitman WB; Yates JR 3rd; Ibba M; Soll D RNA-dependent cysteine biosynthesis in archaea. Science 2005, 307, 1969. [DOI] [PubMed] [Google Scholar]
  • (118).Hohn MJ; Park HS; O'Donoghue P; Schnitzbauer M; Soll D Emergence of the universal genetic code imprinted in an RNA record. Proc. Natl. Acad. Sci. U S A 2006, 103, 18095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (119).Umbarger HE; Umbarger MA The biosynthetic pathway of serine in Salmonella typhimurium. Biochim. Biophys. Acta 1962, 62, 193. [DOI] [PubMed] [Google Scholar]
  • (120).Rogerson DT; Sachdeva A; Wang K; Haq T; Kazlauskaite A; Hancock SM; Huguenin-Dezot N; Muqit MM; Fry AM; Bayliss R et al. Efficient genetic encoding of phosphoserine and its nonhydrolyzable analog. Nat. Chem. Biol 2015, 11, 496. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (121).Heinemann IU; Rovner AJ; Aerni HR; Rogulina S; Cheng L; Olds W; Fischer JT; Soll D; Isaacs FJ; Rinehart J Enhanced phosphoserine insertion during Escherichia coli protein synthesis via partial UAG codon reassignment and release factor 1 deletion. FEBS Lett. 2012, 586, 3716. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (122).Mohler K; Moen JM; Rogulina S; Rinehart J System-wide optimization of an orthogonal translation system with enhanced biological tolerance. Mol Syst Biol 2023, 19, e10591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (123).Stuber K; Schneider T; Werner J; Kovermann M; Marx A; Scheffner M Structural and functional consequences of NEDD8 phosphorylation. Nat. Commun 2021, 12, 5939. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (124).Mukai T; Hayashi A; Iraha F; Sato A; Ohtake K; Yokoyama S; Sakamoto K Codon reassignment in the Escherichia coli genetic code. Nucleic Acids Res. 2010, 38, 8188. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (125).Fukunaga R; Yokoyama S Structural insights into the first step of RNA-dependent cysteine biosynthesis in archaea. Nat. Struct. Mol. Biol 2007, 14, 272. [DOI] [PubMed] [Google Scholar]
  • (126).Lajoie MJ; Rovner AJ; Goodman DB; Aerni HR; Haimovich AD; Kuznetsov G; Mercer JA; Wang HH; Carr PA; Mosberg JA et al. Genomically recoded organisms expand biological functions. Science 2013, 342, 357. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (127).Mukai T; Hoshi H; Ohtake K; Takahashi M; Yamaguchi A; Hayashi A; Yokoyama S; Sakamoto K Highly reproductive Escherichia coli cells with no specific assignment to the UAG codon. Sci. Rep 2015, 5, 9699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (128).Wannier TM; Kunjapur AM; Rice DP; McDonald MJ; Desai MM; Church GM Adaptive evolution of genomically recoded Escherichia coli. Proc. Natl. Acad. Sci. U S A 2018, 115, 3090. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (129).Du F; Liu YQ; Xu YS; Li ZJ; Wang YZ; Zhang ZX; Sun XM Regulating the T7 RNA polymerase expression in E. coli BL21 (DE3) to provide more host options for recombinant protein production. Microb. Cell Fact 2021, 20, 189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (130).Konrat R NMR contributions to structural dynamics studies of intrinsically disordered proteins. J. Magn. Reson 2014, 241, 74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (131).Vieweg S; Mulholland K; Brauning B; Kachariya N; Lai YC; Toth R; Singh PK; Volpi I; Sattler M; Groll M et al. PINK1-dependent phosphorylation of Serine111 within the SF3 motif of Rab GTPases impairs effector interactions and LRRK2-mediated phosphorylation at Threonine72. Biochem J. 2020, 477, 1651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (132).Mekkattu Tharayil S; Mahawaththa MC; Loh CT; Adekoya I; Otting G Phosphoserine for the generation of lanthanide-binding sites on proteins for paramagnetic nuclear magnetic resonance spectroscopy. Magn. Reson. (Gott) 2021, 2, 1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (133).Buchko GW; Zhou M; Vesely CH; Tao J; Shaw WJ; Mehl RA; Cooley RB High-yield recombinant bacterial expression of (13) C-, (15) N-labeled, serine-16 phosphorylated, murine amelogenin using a modified third generation genetic code expansion protocol. Protein Sci. 2023, 32, e4560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (134).Elliott TS; Slowey A; Ye Y; Conway SJ The use of phosphate bioisosteres in medicinal chemistry and chemical biology. Medchemcomm 2012, 3, 735. [Google Scholar]
  • (135).McDonald IK; Thornton JM Satisfying hydrogen bonding potential in proteins. J. Mol. Biol 1994, 238, 777. [DOI] [PubMed] [Google Scholar]
  • (136).Chen HX; Kang J; Chang R; Zhang YL; Duan HZ; Li YM; Chen YX Synthesis of alpha,alpha-difluorinated phosphonate pSer/pThr mimetics via rhodium-catalyzed asymmetric hydrogenation of beta-difluorophosphonomethyl alpha-(acylamino)acrylates. Org Lett. 2018, 20, 3278. [DOI] [PubMed] [Google Scholar]
  • (137).Kang J; Chen H-X; Huang S-Q; Zhang Y-L; Chang R; Li F-Y; Li Y-M; Chen Y-X Facile synthesis of Fmoc-protected phosphonate pSer mimetic and its application in assembling a substrate peptide of 14-3-3 ζ. Tetrahedron Letters 2017, 58, 2551. [Google Scholar]
  • (138).Patskovsky Y; Natarajan A; Patskovska L; Nyovanie S; Joshi B; Morin B; Brittsan C; Huber O; Gordon S; Michelet X et al. Molecular mechanism of phosphopeptide neoantigen immunogenicity. Nat. Commun 2023, 14, 3763. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (139).Johannes TW; DeSieno MA; Griffin BM; Thomas PM; Kelleher NL; Metcalf WW; Zhao H Deciphering the late biosynthetic steps of antimalarial compound FR-900098. Chem. Biol 2010, 17, 57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (140).Tugaeva KV; Sysoev AA; Kapitonova AA; Smith JLR; Zhu P; Cooley RB; Antson AA; Sluchanko NN Human 14-3-3 proteins site-selectively bind the mutational hotspot region of SARS-CoV-2 nucleoprotein modulating its phosphoregulation. J. Mol. Biol 2023, 435, 167891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (141).Schmied WH; Elsasser SJ; Uttamapinant C; Chin JW Efficient multisite unnatural amino acid incorporation in mammalian cells via optimized pyrrolysyl tRNA synthetase/tRNA expression and engineered eRF1. J. Am. Chem. Soc 2014, 136, 15577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (142).Oza JP; Aerni HR; Pirman NL; Barber KW; ter Haar CM; Rogulina S; Amrofell MB; Isaacs FJ; Rinehart J; Jewett MC Robust production of recombinant phosphoproteins using cell-free protein synthesis. Nat. Commun 2015, 6, 8168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (143).Gan Q; Fan C Increasing the fidelity of noncanonical amino acid incorporation in cell-free protein synthesis. Biochim. Biophys. Acta 2017, 1861, 3047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (144).Liu D; Liu Y; Duan HZ; Chen X; Wang Y; Wang T; Yu Q; Chen YX; Lu Y Customized synthesis of phosphoprotein bearing phosphoserine or its nonhydrolyzable analog. Synth. Syst. Biotechnol 2023, 8, 69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (145).Niemi NM; Wilson GM; Overmyer KA; Vogtle FN; Myketin L; Lohman DC; Schueler KL; Attie AD; Meisinger C; Coon JJ et al. Pptc7 is an essential phosphatase for promoting mammalian mitochondrial metabolism and biogenesis. Nat. Commun 2019, 10, 3197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (146).Schachner LF; Soye BD; Ro S; Kenney GE; Ives AN; Su T; Goo YA; Jewett MC; Rosenzweig AC; Kelleher NL Revving an engine of human metabolism: activity enhancement of triosephosphate isomerase via hemi-phosphorylation. ACS Chem. Biol 2022, 17, 2769. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (147).Balasuriya N; Kunkel MT; Liu X; Biggar KK; Li SS; Newton AC; O'Donoghue P Genetic code expansion and live cell imaging reveal that Thr-308 phosphorylation is irreplaceable and sufficient for Akt1 activity. J. Biol. Chem 2018, 293, 10744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (148).Beitia GJ; Rutherford TJ; Freund SMV; Pelham HR; Bienz M; Gammons MV Regulation of Dishevelled DEP domain swapping by conserved phosphorylation sites. Proc. Natl. Acad. Sci. U S A 2021, 118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (149).Burgess SG; Mukherjee M; Sabir S; Joseph N; Gutierrez-Caballero C; Richards MW; Huguenin-Dezot N; Chin JW; Kennedy EJ; Pfuhl M et al. Mitotic spindle association of TACC3 requires Aurora-A-dependent stabilization of a cryptic alpha-helix. EMBO J. 2018, 37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (150).Ha Y; Yang A; Lee S; Kim K; Liew H; Suh YH; Park HS; Churchill DG Facile "stop codon" method reveals elevated neuronal toxicity by discrete S87p-alpha-synuclein oligomers. Biochem. Biophys. Res. Commun 2014, 443, 1085. [DOI] [PubMed] [Google Scholar]
  • (151).Joly N; Beaumale E; Van Hove L; Martino L; Pintard L Phosphorylation of the microtubule-severing AAA+ enzyme Katanin regulates C. elegans embryo development. J. Cell. Biol 2020, 219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (152).Miller CJ; Lou HJ; Simpson C; van de Kooij B; Ha BH; Fisher OS; Pirman NL; Boggon TJ; Rinehart J; Yaffe MB et al. Comprehensive profiling of the STE20 kinase family defines features essential for selective substrate targeting and signaling output. Plos Biol. 2019, 17, e2006540. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (153).Mukherjee M; Sabir S; O'Regan L; Sampson J; Richards MW; Huguenin-Dezot N; Ault JR; Chin JW; Zhuravleva A; Fry AM et al. Mitotic phosphorylation regulates Hsp72 spindle localization by uncoupling ATP binding from substrate release. Sci. Signal 2018, 11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (154).Richter B; Sliter DA; Herhaus L; Stolz A; Wang C; Beli P; Zaffagnini G; Wild P; Martens S; Wagner SA et al. Phosphorylation of OPTN by TBK1 enhances its binding to Ub chains and promotes selective autophagy of damaged mitochondria. Proc. Natl. Acad. Sci. U S A 2016, 113, 4039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (155).Serrano BP; Hardy JA Phosphorylation by protein kinase A disassembles the caspase-9 core. Cell Death Differ. 2018, 25, 1025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (156).Vistrup-Parry M; Chen X; Johansen TL; Bach S; Buch-Larsen SC; Bartling CRO; Ma C; Clemmensen LS; Nielsen ML; Zhang M et al. Site-specific phosphorylation of PSD-95 dynamically regulates the postsynaptic density as observed by phase separation. iScience 2021, 24, 103268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (157).Zheng H; He J; Li J; Yang J; Kirk ML; Roman LJ; Feng C Generation and characterization of functional phosphoserine-incorporated neuronal nitric oxide synthase holoenzyme. J. Biol. Inorg. Chem 2019, 24, 1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (158).Huguenin-Dezot N; De Cesare V; Peltier J; Knebel A; Kristaryianto YA; Rogerson DT; Kulathu Y; Trost M; Chin JW Synthesis of isomeric phosphoubiquitin chains reveals that phosphorylation controls deubiquitinase activity and specificity. Cell Rep. 2016, 16, 1180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (159).George S; Wang SM; Bi YM; Treidlinger M; Barber KR; Shaw GS; O'Donoghue P Ubiquitin phosphorylated at Ser57 hyper-activates parkin. Biochim. Biophys. Acta 2017, 1861, 3038. [DOI] [PubMed] [Google Scholar]
  • (160).Gassaway BM; Cardone RL; Padyana AK; Petersen MC; Judd ET; Hayes S; Tong SL; Barber KW; Apostolidi M; Abulizi A et al. Distinct hepatic PKA and CDK signaling pathways control activity-independent pyruvate kinase phosphorylation and hepatic glucose production. Cell Rep. 2019, 29, 3394. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (161).Vieweg S; Mulholland K; Bräuning B; Kachariya N; Lai YC; Toth R; Singh PK; Volpi I; Sattler M; Groll M et al. PINK1-dependent phosphorylation of Serine111 within the SF3 motif of Rab GTPases impairs effector interactions and LRRK2-mediated phosphorylation at Threonine72. Biochem. J 2020, 477, 1651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (162).Schachner LF; Des Soye B; Ro S; Kenney GE; Ives AN; Su TJF; Goo YA; Jewett MC; Rosenzweig AC; Kelleher NL Revving an engine of human metabolism: activity enhancement of triosephosphate isomerase via hemi-phosphorylation. ACS Chem. Biol 2022, 17, 2769. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (163).Guo X; Niemi NM; Hutchins PD; Condon SGF; Jochem A; Ulbrich A; Higbee AJ; Russell JD; Senes A; Coon JJ et al. Ptc7p dephosphorylates select mitochondrial proteins to enhance metabolic function. Cell Rep. 2017, 18, 307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (164).Lv Z; Rickman KA; Yuan L; Williams K; Selvam SP; Woosley AN; Howe PH; Ogretmen B; Smogorzewska A; Olsen SKS pombe Uba1-Ubc15 structure reveals a novel regulatory mechanism of ubiquitin E2 activity. Mol. Cell 2017, 65, 699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (165).Ahn J; Lee JG; Chin C; In S; Yang A; Park HS; Kim J; Park JH MSK1 functions as a transcriptional coactivator of p53 in the regulation of gene expression. Exp. Mol. Med 2018, 50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (166).Beránek V; Reinkemeier CD; Zhang MS; Liang AD; Kym G; Chin JW Genetically encoded protein phosphorylation in mammalian cells. Cell Chem. Biol 2018, 25, 1067. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (167).Dickson C; Fletcher A; Vaysburd M; Yang JC; Mallery DL; Zeng JW; Johnson CM; McLaughlin SH; Skehel M; Maslen S et al. Intracellular antibody signalling is regulated by phosphorylation of the Fc receptor TRIM21. Elife 2018, 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (168).Mukherjee M; Sabir S; O'Regan L; Sampson J; Richards MW; Huguenin-Dezot N; Ault JR; Chin JW; Zhuravleva A; Fry AM et al. Mitotic phosphorylation regulates Hsp72 spindle localization by uncoupling ATP binding from substrate release. Sci. Signal 2018, 11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (169).Serrano BP; Hardy JA Phosphorylation by protein kinase A disassembles the caspase-9 core. Cell Death and Differentiation 2018, 25, 1025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (170).Niemi NM; Wilson GM; Overmyer KA; Vögtle FN; Myketin L; Lohman DC; Schueler KL; Attie AD; Meisinger C; Coon JJ et al. Pptc7 is an essential phosphatase for promoting mammalian mitochondrial metabolism and biogenesis. Nat. Comm 2019, 10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (171).Soupene E; Kuypers FA ACBD6 protein controls acyl chain availability and specificity of the N-myristoylation modification of proteins. J. Lipid Res 2019, 60, 624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (172).Baliova M; Jursky F Phosphorylation of serine 157 protects the rat glycine transporter GlyT2 from calpain cleavage. J. Mol. Neurosci 2020, 70, 1216. [DOI] [PubMed] [Google Scholar]
  • (173).Di Mattia T; Martinet A; Ikhlef S; McEwen AG; Nominé Y; Wendling C; Poussin-Courmontagne P; Voilquin L; Eberling P; Ruffenach F et al. FFAT motif phosphorylation controls formation and lipid transfer function of inter-organelle contacts. Embo J. 2020, 39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (174).Joly N; Beaumale E; Van Hove L; Martino L; Pintard L Phosphorylation of the microtubule-severing AAA plus enzyme Katanin regulates embryo development. J. Cell Biol 2020, 219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (175).Li M; Shu HB Dephosphorylation of cGAS by PPP6C impairs its substrate binding activity and innate antiviral response. Prot. Cell 2020, 11, 584. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (176).Magnussen HM; Ahmed SF; Sibbet GJ; Hristova VA; Nomura K; Hock AK; Archibald LJ; Jamieson AG; Fushman D; Vousden KH et al. Structural basis for DNA damage-induced phosphoregulation of MDM2 RING domain. Nat. Comm 2020, 11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (177).Mamonova T; Friedman PA Noncanonical sequences involving NHERF1 interaction with NPT2A govern hormone-regulated phosphate transport: binding outside the box. Int. J. Mol. Sci 2021, 22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (178).Schiapparelli P; Pirman NL; Mohler K; Miranda-Herrera PA; Zarco N; Kilic O; Miller C; Shah SR; Rogulina S; Hungerford W et al. Phosphorylated WNK kinase networks in recoded bacteria recapitulate physiological function. Cell Rep. 2021, 36, 109416. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (179).Vistrup-Parry M; Sneddon WB; Bach S; Stromgaard K; Friedman PA; Mamonova T Multisite NHERF1 phosphorylation controls GRK6A regulation of hormone-sensitive phosphate transport. J. Biol. Chem 2021, 296, 100473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (180).Zhu P; Nguyen KT; Estelle AB; Sluchanko NN; Mehl RA; Cooley RB Genetic encoding of 3-nitro-tyrosine reveals the impacts of 14-3-3 nitration on client binding and dephosphorylation. Protein Sci. 2023, 32, e4574. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (181).Shi M; Cho H; Inn KS; Yang A; Zhao Z; Liang Q; Versteeg GA; Amini-Bavil-Olyaee S; Wong LY; Zlokovic BV et al. Negative regulation of NF-kappaB activity by brain-specific TRIpartite Motif protein 9. Nat. Commun 2014, 5, 4820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (182).Sawyer N; Gassaway BM; Haimovich AD; Isaacs FJ; Rinehart J; Regan L Designed phosphoprotein recognition in Escherichia coli. ACS Chem. Biol 2014, 9, 2502. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (183).Seo GJ; Yang A; Tan B; Kim S; Liang Q; Choi Y; Yuan W; Feng P; Park HS; Jung JU Akt Kinase-mediated checkpoint of cGAS DNA sensing pathway. Cell Rep. 2015, 13, 440. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (184).Heo JM; Ordureau A; Paulo JA; Rinehart J; Harper JW The PINK1-PARKIN mitochondrial ubiquitylation pathway drives a program of OPTN/NDP52 recruitment and TBK1 activation to promote mitophagy. Mol. Cell 2015, 60, 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (185).Gan Q; Lehman BP; Bobik TA; Fan C Expanding the genetic code of Salmonella with non-canonical amino acids. Sci. Rep 2016, 6, 39920. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (186).Balasuriya N; McKenna M; Liu X; Li SSC; O'Donoghue P Phosphorylation-dependent inhibition of Akt1. Genes (Basel) 2018, 9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (187).Barber KW; Muir P; Szeligowski RV; Rogulina S; Gerstein M; Sampson JR; Isaacs FJ; Rinehart J Encoding human serine phosphopeptides in bacteria for proteome-wide identification of phosphorylation-dependent interactions. Nat. Biotechnol 2018, 36, 638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (188).Crnkovic A; Vargas-Rodriguez O; Merkuryev A; Soll D Effects of heterologous tRNA modifications on the production of proteins containing noncanonical amino acids. Bioengineering (Basel) 2018, 5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (189).Venkat S; Sturges J; Stahman A; Gregory C; Gan Q; Fan C Genetically incorporating two distinct post-translational modifications into one protein simultaneously. ACS Synth. Biol 2018, 7, 689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (190).Marmelstein AM; Morgan JAM; Penkert M; Rogerson DT; Chin JW; Krause E; Fiedler D Pyrophosphorylation via selective phosphoprotein derivatization. Chem. Sci 2018, 9, 5929. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (191).Thomas Y; Scott DC; Kristariyanto YA; Rinehart J; Clark K; Cohen P; Kurz T The NEDD8 E3 ligase DCNL5 is phosphorylated by IKK alpha during Toll-like receptor activation. PLoS One 2018, 13, e0199197. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (192).De Nicola GF; Bassi R; Nichols C; Fernandez-Caggiano M; Golforoush PA; Thapa D; Anderson R; Martin ED; Verma S; Kleinjung J et al. The TAB1-p38alpha complex aggravates myocardial injury and can be targeted by small molecules. JCI Insight 2018, 3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (193).Zhang M; Liang C; Chen Q; Yan H; Xu J; Zhao H; Yuan X; Liu J; Lin S; Lu W et al. Histone H2A phosphorylation recruits topoisomerase IIalpha to centromeres to safeguard genomic stability. EMBO J. 2020, 39, e101863. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (194).Palani S; Koester D; Balasubramanian MK Phosphoregulation of tropomyosin-actin interaction revealed using a genetic code expansion strategy. Wellcome Open Res. 2020, 5, 161. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (195).Liu X; Xiao W; Zhang Y; Wiley SE; Zuo T; Zheng Y; Chen N; Chen L; Wang X; Zheng Y et al. Reversible phosphorylation of Rpn1 regulates 26S proteasome assembly and function. Proc. Natl. Acad. Sci. U S A 2020, 117, 328. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (196).Jain N; Janning P; Neumann H 14-3-3 Protein Bmh1 triggers short-range compaction of mitotic chromosomes by recruiting sirtuin deacetylase Hst2. J. Biol. Chem 2021, 296, 100078. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (197).Wang KH; Liu YT; Yu ZL; Gu B; Hu J; Huang L; Ge X; Xu LY; Zhang MY; Zhao JC et al. Phosphorylation at Ser68 facilitates DCAF11-mediated ubiquitination and degradation of CENP-A during the cell cycle. Cell Rep. 2021, 37. [DOI] [PubMed] [Google Scholar]
  • (198).Gassaway BM; Li J; Rad R; Mintseris J; Mohler K; Levy T; Aguiar M; Beausoleil SA; Paulo JA; Rinehart J et al. A multi-purpose, regenerable, proteome-scale, human phosphoserine resource for phosphoproteomics. Nat. Methods 2022, 19, 1371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (199).Nardone C; Palanski BA; Scott DC; Timms RT; Barber KW; Gu X; Mao A; Leng Y; Watson EV; Schulman BA et al. A central role for regulated protein stability in the control of TFE3 and MITF by nutrients. Mol. Cell 2023, 83, 57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (200).Perez-Pepe M; Desotell AW; Li H; Li W; Han B; Lin Q; Klein DE; Liu Y; Goodarzi H; Alarcon CR 7SK methylation by METTL3 promotes transcriptional activity. Sci. Adv 2023, 9, eade7500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (201).Koyano F; Okatsu K; Kosako H; Tamura Y; Go E; Kimura M; Kimura Y; Tsuchiya H; Yoshihara H; Hirokawa T et al. Ubiquitin is phosphorylated by PINK1 to activate parkin. Nature 2014, 510, 162. [DOI] [PubMed] [Google Scholar]
  • (202).Fabbro D; Batt D; Rose P; Schacher B; Roberts TM; Ferrari S Homogeneous purification of human recombinant GST-Akt/PKB from Sf9 cells. Protein Expr. Purif 1999, 17, 83. [DOI] [PubMed] [Google Scholar]
  • (203).Yaffe MB; Elia AE Phosphoserine/threonine-binding domains. Curr. Opin. Cell Bio.l 2001, 13, 131. [DOI] [PubMed] [Google Scholar]
  • (204).Sluchanko NN Association of multiple phosphorylated proteins with the 14-3-3 regulatory hubs: problems and perspectives. J. Mol. Biol 2018, 430, 20. [DOI] [PubMed] [Google Scholar]
  • (205).Pennington KL; Chan TY; Torres MP; Andersen JL The dynamic and stress-adaptive signaling hub of 14-3-3: emerging mechanisms of regulation and context-dependent protein-protein interactions. Oncogene 2018, 37, 5587. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (206).Obsilova V; Obsil T Structural insights into the functional roles of 14-3-3 proteins. Front. Mol. Biosci 2022, 9, 1016071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (207).Chernik IS; Seit-Nebi AS; Marston SB; Gusev NB Small heat shock protein Hsp20 (HspB6) as a partner of 14-3-3gamma. Mol. Cell. Biochem 2007, 295, 9. [DOI] [PubMed] [Google Scholar]
  • (208).Johnson C; Crowther S; Stafford MJ; Campbell DG; Toth R; MacKintosh C Bioinformatic and experimental survey of 14-3-3-binding sites. Biochem. J 2010, 427, 69. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (209).Kast DJ; Dominguez R Mechanism of IRSp53 inhibition by 14-3-3. Nat. Commun 2019, 10, 483. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (210).Hinds MG; Smits C; Fredericks-Short R; Risk JM; Bailey M; Huang DC; Day CL Bim, Bad and Bmf: intrinsically unstructured BH3-only proteins that undergo a localized conformational change upon binding to prosurvival Bcl-2 targets. Cell Death Differ. 2007, 14, 128. [DOI] [PubMed] [Google Scholar]
  • (211).Banerjee T; Chakravarti D A peek into the complex realm of histone phosphorylation. Mol. Cell Biol 2011, 31, 4858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (212).Shogren-Knaak MA; Fry CJ; Peterson CL A native peptide ligation strategy for deciphering nucleosomal histone modifications. J. Biol. Chem 2003, 278, 15744. [DOI] [PubMed] [Google Scholar]
  • (213).Rosano GL; Ceccarelli EA Recombinant protein expression in Escherichia coli: advances and challenges. Front. Microbiol 2014, 5, 172. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (214).Vethanayagam JG; Flower AM Decreased gene expression from T7 promoters may be due to impaired production of active T7 RNA polymerase. Microb. Cell Fact 2005, 4, 3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (215).Zhu P; Mehl RA; Cooley RB Site-specific incorporation of phosphoserine into recombinant proteins in Escherichia coli. Bio Protoc. 2022, 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (216).Mohler K; Rinehart J Expression of authentic post-translationally modified proteins in organisms with expanded genetic codes. Methods Enzymol. 2019, 626, 539. [DOI] [PubMed] [Google Scholar]
  • (217).Zhu P; Mehl RA; Cooley RB Biosynthesis and genetic encoding of non-hydrolyzable phosphoserine into recombinant oroteins in Escherichia coli. Bio Protoc. 2023, 13, e4861. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (218).Gonzalez SS; Ad O; Shah B; Zhang Z; Zhang X; Chatterjee A; Schepartz A Genetic code expansion in the engineered organism Vmax X2: High Yield and Exceptional Fidelity. ACS Cent. Sci 2021, 7, 1500. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (219).Arrendale A; Kim K; Choi JY; Li W; Geahlen RL; Borch RF Synthesis of a phosphoserine mimetic prodrug with potent 14-3-3 protein inhibitory activity. Chem. Biol 2012, 19, 764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (220).Ko W; Kumar R; Kim S; Lee HS Construction of bacterial cells with an active transport system for unnatural amino acids. ACS Synth. Biol 2019, 8, 1195. [DOI] [PubMed] [Google Scholar]
  • (221).Fredens J; Wang K; de la Torre D; Funke LFH; Robertson WE; Christova Y; Chia T; Schmied WH; Dunkelmann DL; Beranek V et al. Total synthesis of Escherichia coli with a recoded genome. Nature 2019, 569, 514. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (222).Niu W; Schultz PG; Guo J An expanded genetic code in mammalian cells with a functional quadruplet codon. ACS Chem. Biol 2013, 8, 1640. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (223).Anderson JC; Wu N; Santoro SW; Lakshman V; King DS; Schultz PG An expanded genetic code with a functional quadruplet codon. Proc. Natl. Acad. Sci. U S A 2004, 101, 7566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (224).Neumann H; Wang K; Davis L; Garcia-Alai M; Chin JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 2010, 464, 441. [DOI] [PubMed] [Google Scholar]
  • (225).Durocher D; Taylor IA; Sarbassova D; Haire LF; Westcott SL; Jackson SP; Smerdon SJ; Yaffe MB The molecular basis of FHA domain:phosphopeptide binding specificity and implications for phospho-dependent signaling mechanisms. Mol. Cell 2000, 6, 1169. [DOI] [PubMed] [Google Scholar]
  • (226).Pandey AK; Ganguly HK; Sinha SK; Daniels KE; Yap GPA; Patel S; Zondlo NJ An inherent difference between serine and threonine phosphorylation: phosphothreonine strongly prefers a highly ordered, compact, cyclic conformation. ACS Chem. Biol 2023, 18, 1938. [DOI] [PubMed] [Google Scholar]
  • (227).Moen JM; Mohler K; Rogulina S; Shi X; Shen H; Rinehart J Enhanced access to the human phosphoproteome with genetically encoded phosphothreonine. Nat. Commun 2022, 13, 7226. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (228).Fan C; Bobik TA The PduX enzyme of Salmonella enterica is an L-threonine kinase used for coenzyme B12 synthesis. J. Biol. Chem 2008, 283, 11322. [DOI] [PubMed] [Google Scholar]
  • (229).Deng M; Lin J; Nowsheen S; Liu T; Zhao Y; Villalta PW; Sicard D; Tschumperlin DJ; Lee S; Kim J et al. Extracellular matrix stiffness determines DNA repair efficiency and cellular sensitivity to genotoxic agents. Sci. Adv 2020, 6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (230).Otaka A; Mitsuyama E; Kinoshita T; Tamamura H; Fujii N Stereoselective synthesis of CF(2)-substituted phosphothreonine mimetics and their incorporation into peptides using newly developed deprotection procedures. J. Org. Chem 2000, 65, 4888. [DOI] [PubMed] [Google Scholar]
  • (231).Duan HZ; Chen HX; Yu Q; Hu J; Li YM; Chen YX Stereoselective synthesis of a phosphonate pThr mimetic via palladium-catalyzed gamma-C(sp(3))-H activation for peptide preparation. Org. Biomol. Chem 2019, 17, 2099. [DOI] [PubMed] [Google Scholar]
  • (232).Liu F; Park JE; Lee KS; Burke TR Jr. Preparation of orthogonally protected (2S, 3R)-2-amino-3-methyl-4-phosphonobutyric acid (Pmab) as a phosphatase-stable phosphothreonine mimetic and its use in the synthesis of Polo-box domain-binding peptides. Tetrahedron 2009, 65, 9673. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (233).Hymel D; Burke TR Jr. Phosphatase-stable phosphoamino acid mimetics that enhance binding affinities with the Polo-box domain of Polo-like kinase 1. ChemMedChem 2017, 12, 202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (234).Duan H-Z; Chen H-X; Yu Q; Hu J; Li Y-M; Chen Y-X Stereoselective synthesis of a phosphonate pThr mimetic via palladium-catalyzed γ-C(sp3)–H activation for peptide preparation. Org. Biomol.Chem 2019, 17, 2099. [DOI] [PubMed] [Google Scholar]
  • (235).Goguen BN; Aemissegger A; Imperiali B Sequential activation and deactivation of protein function using spectrally differentiated caged phosphoamino acids. J. Am. Chem. Soc 2011, 133, 11038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (236).Serwa R; Wilkening I; Del Signore G; Muhlberg M; Claussnitzer I; Weise C; Gerrits M; Hackenberger CP Chemoselective Staudinger-phosphite reaction of azides for the phosphorylation of proteins. Angew. Chem. Int. Ed. Engl 2009, 48, 8234. [DOI] [PubMed] [Google Scholar]
  • (237).Liu CC; Schultz PG Recombinant expression of selectively sulfated proteins in Escherichia coli. Nat. Biotechnol 2006, 24, 1436. [DOI] [PubMed] [Google Scholar]
  • (238).Burke TR Jr.; Yao ZJ; Liu DG; Voigt J; Gao Y Phosphoryltyrosyl mimetics in the design of peptide-based signal transduction inhibitors. Biopolymers 2001, 60, 32. [DOI] [PubMed] [Google Scholar]
  • (239).Makukhin N; Ciulli A Recent advances in synthetic and medicinal chemistry of phosphotyrosine and phosphonate-based phosphotyrosine analogues. RSC Med. Chem 2020, 12, 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (240).Tong L; Warren TC; Lukas S; Schembri-King J; Betageri R; Proudfoot JR; Jakes S Carboxymethyl-phenylalanine as a replacement for phosphotyrosine in SH2 domain binding. J. Biol. Chem 1998, 273, 20238. [DOI] [PubMed] [Google Scholar]
  • (241).Gilmer T; Rodriguez M; Jordan S; Crosby R; Alligood K; Green M; Kimery M; Wagner C; Kinder D; Charifson P et al. Peptide inhibitors of src SH3-SH2-phosphoprotein interactions. J. Biol. Chem 1994, 269, 31711. [PubMed] [Google Scholar]
  • (242).Beaulieu PL; Cameron DR; Ferland JM; Gauthier J; Ghiro E; Gillard J; Gorys V; Poirier M; Rancourt J; Wernic D et al. Ligands for the tyrosine kinase p56lck SH2 domain: discovery of potent dipeptide derivatives with monocharged, nonhydrolyzable phosphate replacements. J. Med. Chem 1999, 42, 1757. [DOI] [PubMed] [Google Scholar]
  • (243).Ju T; Niu W; Cerny R; Bollman J; Roy A; Guo J Molecular recognition of sulfotyrosine and phosphotyrosine by the Src homology 2 domain. Mol. Biosyst 2013, 9, 1829. [DOI] [PubMed] [Google Scholar]
  • (244).Xie J; Supekova L; Schultz PG A genetically encoded metabolically stable analogue of phosphotyrosine in Escherichia coli. ACS Chem. Biol 2007, 2, 474. [DOI] [PubMed] [Google Scholar]
  • (245).Phillips AJ; Taleski D; Koplinski CA; Getschman AE; Moussouras NA; Richard AM; Peterson FC; Dwinell MB; Volkman BF; Payne RJ et al. CCR7 Sulfotyrosine enhances CCL21 binding. Int. J. Mol. Sci 2017, 18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (246).Hofsteenge J; Stone SR; Donella-Deana A; Pinna LA The effect of substituting phosphotyrosine for sulphotyrosine on the activity of hirudin. Eur. J. Biochem 1990, 188, 55. [DOI] [PubMed] [Google Scholar]
  • (247).Garrido-Hernandez H; Moon KD; Geahlen RL; Borch RF Design and synthesis of phosphotyrosine peptidomimetic prodrugs. J. Med. Chem 2006, 49, 3368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (248).Johnson DB; Xu J; Shen Z; Takimoto JK; Schultz MD; Schmitz RJ; Xiang Z; Ecker JR; Briggs SP; Wang L RF1 knockout allows ribosomal incorporation of unnatural amino acids at multiple sites. Nat. Chem. Biol 2011, 7, 779. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (249).Robertson CD; Hazen TH; Kaper JB; Rasko DA; Hansen AM Phosphotyrosine-mediated regulation of enterohemorrhagic Escherichia coli virulence. mBio 2018, 9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (250).Grasso KT; Singha Roy SJ; Osgood AO; Yeo MJR; Soni C; Hillenbrand CM; Ficaretta ED; Chatterjee A A facile platform to engineer Escherichia coli tyrosyl-tRNA synthetase adds new chemistries to the eukaryotic genetic code, including a phosphotyrosine mimic. ACS Cent. Sci 2022, 8, 483. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (251).He X; Ma B; Chen Y; Guo J; Niu W Genetic encoding of a nonhydrolyzable phosphotyrosine analog in mammalian cells. Chem. Commun. (Camb.) 2022, 58, 5897. [DOI] [PubMed] [Google Scholar]
  • (252).Liu CC; Cellitti SE; Geierstanger BH; Schultz PG Efficient expression of tyrosine-sulfated proteins in E. coli using an expanded genetic code. Nat. Protoc 2009, 4, 1784. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (253).Schwessinger B; Li X; Ellinghaus TL; Chan LJ; Wei T; Joe A; Thomas N; Pruitt R; Adams PD; Chern MS et al. A second-generation expression system for tyrosine-sulfated proteins and its application in crop protection. Integr. Biol. (Camb.) 2016, 8, 542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (254).Chen Y; Jin S; Zhang M; Hu Y; Wu KL; Chung A; Wang S; Tian Z; Wang Y; Wolynes PG et al. Unleashing the potential of noncanonical amino acid biosynthesis to create cells with precision tyrosine sulfation. Nat. Commun 2022, 13, 5434. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (255).Italia JS; Peeler JC; Hillenbrand CM; Latour C; Weerapana E; Chatterjee A Genetically encoded protein sulfation in mammalian cells. Nat. Chem. Biol 2020, 16, 379. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (256).He X; Chen Y; Beltran DG; Kelly M; Ma B; Lawrie J; Wang F; Dodds E; Zhang L; Guo J et al. Functional genetic encoding of sulfotyrosine in mammalian cells. Nat. Commun 2020, 11, 4820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (257).Chen S; Maini R; Bai X; Nangreave RC; Dedkova LM; Hecht SM Incorporation of phosphorylated tyrosine into proteins: in vitro translation and study of phosphorylated IkappaB-alpha and its Interaction with NF-kappaB. J. Am. Chem. Soc 2017, 139, 14098. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (258).Hoppmann C; Wong A; Yang B; Li S; Hunter T; Shokat KM; Wang L Site-specific incorporation of phosphotyrosine using an expanded genetic code. Nat. Chem. Biol 2017, 13, 842. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (259).Iraha F; Oki K; Kobayashi T; Ohno S; Yokogawa T; Nishikawa K; Yokoyama S; Sakamoto K Functional replacement of the endogenous tyrosyl-tRNA synthetase-tRNATyr pair by the archaeal tyrosine pair in Escherichia coli for genetic code expansion. Nucleic Acids Res. 2010, 38, 3682. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (260).Italia JS; Latour C; Wrobel CJJ; Chatterjee A Resurrecting the bacterial tyrosyl-tRNA synthetase/tRNA pair for expanding the genetic code of both E. coli and eukaryotes. Cell Chem. Biol 2018, 25, 1304. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (261).Chin JW; Schultz PG In vivo photocrosslinking with unnatural amino Acid mutagenesis. Chembiochem 2002, 3, 1135. [DOI] [PubMed] [Google Scholar]
  • (262).Stewart V; Ronald PC Sulfotyrosine residues: Interaction specificity determinants for extracellular protein-protein interactions. J. Biol. Chem 2022, 298, 102232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (263).Aviner R The science of puromycin: From studies of ribosome function to applications in biotechnology. Comput. Struct. Biotechnol. J 2020, 18, 1074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (264).Fickel TE; Gilvarg C Transport of impermeant substances in E. coli by way of oligopeptide permease. Nat. New Biol 1973, 241, 161. [DOI] [PubMed] [Google Scholar]
  • (265).Payne JW; Gilvarg C The role of the terminal carboxyl group on peptide transport in Escherichia coli. J. Biol. Chem 1968, 243, 335. [PubMed] [Google Scholar]
  • (266).Mikol V; Baumann G; Keller TH; Manning U; Zurini MG The crystal structures of the SH2 domain of p56lck complexed with two phosphonopeptides suggest a gated peptide binding site. J. Mol. Biol 1995, 246, 344. [DOI] [PubMed] [Google Scholar]
  • (267).Charifson PS; Shewchuk LM; Rocque W; Hummel CW; Jordan SR; Mohr C; Pacofsky GJ; Peel MR; Rodriguez M; Sternbach DD et al. Peptide ligands of pp60(c-src) SH2 domains: a thermodynamic and structural study. Biochemistry 1997, 36, 6283. [DOI] [PubMed] [Google Scholar]
  • (268).Burke TR Jr.; Smyth MS; Otaka A; Nomizu M; Roller PP; Wolf G; Case R; Shoelson SE Nonhydrolyzable phosphotyrosyl mimetics for the preparation of phosphatase-resistant SH2 domain inhibitors. Biochemistry 1994, 33, 6490. [DOI] [PubMed] [Google Scholar]
  • (269).Giorgetti-Peraldi S; Ottinger E; Wolf G; Ye B; Burke TR Jr.; Shoelson SE Cellular effects of phosphotyrosine-binding domain inhibitors on insulin receptor signaling and trafficking. Mol. Cell Biol 1997, 17, 1180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (270).Chen L; Wu L; Otaka A; Smyth MS; Roller PP; Burke TR Jr.; den Hertog J; Zhang ZY Why is phosphonodifluoromethyl phenylalanine a more potent inhibitory moiety than phosphonomethyl phenylalanine toward protein-tyrosine phosphatases? Biochem. Biophys. Res. Commun 1995, 216, 976. [DOI] [PubMed] [Google Scholar]
  • (271).Zitterbart R; Seitz O Parallel chemical protein synthesis on a surface enables the rapid analysis of the phosphoregulation of SH3 domains. Angew. Chem. Int. Ed. Engl 2016, 55, 7252. [DOI] [PubMed] [Google Scholar]
  • (272).Chen S; Ji X; Dedkova LM; Potuganti GR; Hecht SM Site-selective tyrosine phosphorylation in the activation of the p50 Subunit of NF-kappaB for DNA binding and transcription. ACS Chem. Biol 2023, 18, 59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (273).Honda M; Okuno Y; Yoo J; Ha T; Spies M Tyrosine phosphorylation enhances RAD52-mediated annealing by modulating its DNA binding. EMBO J. 2011, 30, 3368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (274).Steringer JP; Bleicken S; Andreas H; Zacherl S; Laussmann M; Temmerman K; Contreras FX; Bharat TA; Lechner J; Muller HM et al. Phosphatidylinositol 4,5-bisphosphate (PI(4,5)P2)-dependent oligomerization of fibroblast growth factor 2 (FGF2) triggers the formation of a lipidic membrane pore implicated in unconventional secretion. J. Biol. Chem 2012, 287, 27659. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (275).Rust HL; Subramanian V; West GM; Young DD; Schultz PG; Thompson PR Using unnatural amino acid mutagenesis to probe the regulation of PRMT1. ACS Chem. Biol 2014, 9, 649. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (276).Guerra-Castellano A; Diaz-Quintana A; Moreno-Beltran B; Lopez-Prados J; Nieto PM; Meister W; Staffa J; Teixeira M; Hildebrandt P; De la Rosa MA et al. Mimicking tyrosine phosphorylation in human cytochrome c by the evolved tRNA synthetase technique. Chemistry 2015, 21, 15004. [DOI] [PubMed] [Google Scholar]
  • (277).Subramanyam S; Ismail M; Bhattacharya I; Spies M Tyrosine phosphorylation stimulates activity of human RAD51 recombinase through altered nucleoprotein filament dynamics. Proc. Natl. Acad. Sci. U S A 2016, 113, E6045. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (278).Moreno-Beltran B; Guerra-Castellano A; Diaz-Quintana A; Del Conte R; Garcia-Maurino SM; Diaz-Moreno S; Gonzalez-Arzola K; Santos-Ocana C; Velazquez-Campoy A; De la Rosa MA et al. Structural basis of mitochondrial dysfunction in response to cytochrome c phosphorylation at tyrosine 48. Proc. Natl. Acad. Sci. U S A 2017, 114, E3041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (279).Dionne U; Chartier FJM; Lopez de Los Santos Y; Lavoie N; Bernard DN; Banerjee SL; Otis F; Jacquet K; Tremblay MG; Jain M et al. Direct phosphorylation of SRC homology 3 domains by tyrosine kinase receptors disassembles ligand-induced signaling networks. Mol. Cell 2018, 70, 995. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (280).Zhou XX; Bracken CJ; Zhang K; Zhou J; Mou Y; Wang L; Cheng Y; Leung KK; Wells JA Targeting phosphotyrosine in native proteins with conditional, bispecific antibody traps. J. Am. Chem. Soc 2020, 142, 17703. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (281).Chen S; Ji X; Dedkova LM; Hecht SM Site-selective incorporation of phosphorylated tyrosine into the p50 subunit of NF-kappaB and activation of its downstream gene CD40. Chem. Commun. (Camb.) 2021, 57, 12651. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (282).Gomila AMJ; Perez-Mejias G; Nin-Hill A; Guerra-Castellano A; Casas-Ferrer L; Ortiz-Tescari S; Diaz-Quintana A; Samitier J; Rovira C; De la Rosa MA et al. Phosphorylation disrupts long-distance electron transport in cytochrome c. Nat. Commun 2022, 13, 7100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (283).Li A; Voleti R; Lee M; Gagoski D; Shah NH High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display. Elife 2023, 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (284).Stevens MU; Croteau N; Eldeeb MA; Antico O; Zeng ZW; Toth R; Durcan TM; Springer W; Fon EA; Muqit MM et al. Structure-based design and characterization of Parkin-activating mutations. Life Sci. Alliance 2023, 6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (285).Hansen AM; Chaerkady R; Sharma J; Diaz-Mejia JJ; Tyagi N; Renuse S; Jacob HK; Pinto SM; Sahasrabuddhe NA; Kim MS et al. The Escherichia coli phosphotyrosine proteome relates to core pathways and virulence. PLoS Pathog. 2013, 9, e1003403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (286).Yu H; Lee I; Salomon AR; Yu K; Huttemann M Mammalian liver cytochrome c is tyrosine-48 phosphorylated in vivo, inhibiting mitochondrial respiration. Biochim. Biophys. Acta 2008, 1777, 1066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (287).Chintan S; Noam P; Matthew H; David FS; Alanna S; Abhishek C A translation-independent directed evolution strategy to engineer aminoacyl-tRNA synthetases. bioRxiv 2023, 12.13.571473. [Google Scholar]
  • (288).Wang K; Neumann H; Peak-Chew SY; Chin JW Evolved orthogonal ribosomes enhance the efficiency of synthetic genetic code expansion. Nat. Biotechnol 2007, 25, 770. [DOI] [PubMed] [Google Scholar]
  • (289).Wiemer AJ; Wiemer DF Prodrugs of phosphonates and phosphates: crossing the membrane barrier. Top. Curr. Chem 2015, 360, 115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (290).Despotovic D; Aharon E; Trofimyuk O; Dubovetskyi A; Cherukuri KP; Ashani Y; Eliason O; Sperfeld M; Leader H; Castelli A et al. Utilization of diverse organophosphorus pollutants by marine bacteria. Proc. Natl. Acad. Sci. U S A 2022, 119, e2203604119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (291).Narmandakh A; Gad'on N; Drepper F; Knapp B; Haehnel W; Fuchs G Phosphorylation of phenol by phenylphosphate synthase: role of histidine phosphate in catalysis. J. Bacteriol 2006, 188, 7815. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (292).Kalagiri R; Hunter T The many ways that nature has exploited the unusual structural and chemical properties of phosphohistidine for use in proteins. Biochem. J 2021, 478, 3575. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (293).Chen CC; Bruegger BB; Kern CW; Lin YC; Halpern RM; Smith RA Phosphorylation of nuclear proteins in rat regenerating liver. Biochemistry 1977, 16, 4852. [DOI] [PubMed] [Google Scholar]
  • (294).Attwood PV; Piggott MJ; Zu XL; Besant PG Focus on phosphohistidine. Amino Acids 2007, 32, 145. [DOI] [PubMed] [Google Scholar]
  • (295).Kee JM; Villani B; Carpenter LR; Muir TW Development of stable phosphohistidine analogues. J. Am. Chem. Soc 2010, 132, 14327. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (296).Makwana MV; Dos Santos Souza C; Pickup BT; Thompson MJ; Lomada SK; Feng Y; Wieland T; Jackson RFW; Muimo R Chemical tools for studying phosphohistidine: generation of selective tau-phosphohistidine and pi-phosphohistidine antibodies. Chembiochem 2023, 24, e202300182. [DOI] [PubMed] [Google Scholar]
  • (297).Xiao H; Peters FB; Yang PY; Reed S; Chittuluru JR; Schultz PG Genetic incorporation of histidine derivatives using an engineered pyrrolysyl-tRNA synthetase. ACS Chem. Biol 2014, 9, 1092. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (298).Cheung JW; Kinney WD; Wesalo JS; Reed M; Nicholson EM; Deiters A; Cropp TA Genetic encoding of a photocaged histidine for light-control of protein activity. Chembiochem 2023, 24, e202200721. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (299).Burke AJ; Lovelock SL; Frese A; Crawshaw R; Ortmayer M; Dunstan M; Levy C; Green AP Design and evolution of an enzyme with a non-canonical organocatalytic mechanism. Nature 2019, 570, 219. [DOI] [PubMed] [Google Scholar]
  • (300).Besant PG; Attwood PV; Piggott MJ Focus on phosphoarginine and phospholysine. Curr. Protein Pept. Sci 2009, 10, 536. [DOI] [PubMed] [Google Scholar]
  • (301).Attwood PV P-N bond protein phosphatases. Biochim. Biophys. Acta 2013, 1834, 470. [DOI] [PubMed] [Google Scholar]
  • (302).Fu SS; Fu C; Zhou Q; Lin RC; Ouyang H; Wang MN; Sun Y; Liu Y; Zhao YF Widespread arginine phosphorylation in human cells-a novel protein PTM revealed by mass spectrometry. Sci. China Chem 2020, 63, 341. [Google Scholar]
  • (303).Ouyang H; Fu C; Fu S; Ji Z; Sun Y; Deng P; Zhao Y Development of a stable phosphoarginine analog for producing phosphoarginine antibodies. Org. Biomol. Chem 2016, 14, 1925. [DOI] [PubMed] [Google Scholar]
  • (304).Fuhrmann J; Subramanian V; Thompson PR Synthesis and use of a phosphonate amidine to generate an anti-phosphoarginine-specific antibody. Angew. Chem. Int. Ed. Engl 2015, 54, 14715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (305).Hauser A; Poulou E; Muller F; Schmieder P; Hackenberger CPR Synthesis and evaluation of non-hydrolyzable phospho-lysine peptide mimics. Chemistry 2021, 27, 2223. [DOI] [PubMed] [Google Scholar]
  • (306).Benkovic SJ; Sampson EJ Structure-reactivity correlation for hydrolysis of phosphoramidate monoanions. J. Am. Chem. Soc 1971, 93, 4009. [Google Scholar]
  • (307).Cox RJ; Hadfield AT; Mayo-Martin MB Difluoromethylene analogues of aspartyl phosphate: the first synthetic inhibitors of aspartate semi-aldehyde dehydrogenase. Chem. Commun. (Camb.) 2001, 18, 1710. [DOI] [PubMed] [Google Scholar]
  • (308).Cox RJ; Gibson JS; Mayo Martin MB Aspartyl phosphonates and phosphoramidates: the first synthetic inhibitors of bacterial aspartate-semialdehyde dehydrogenase. Chembiochem 2002, 3, 874. [DOI] [PubMed] [Google Scholar]
  • (309).Hiratake J; Inoue M; Sakata K gamma-Glutamyltranspeptidase and gamma-glutamyl peptide ligases: fluorophosphonate and phosphonodifluoromethyl ketone analogs as probes of tetrahedral transition state and gamma-glutamyl-phosphate intermediate. Methods Enzymol. 2002, 354, 272. [DOI] [PubMed] [Google Scholar]
  • (310).Sun F; Ding Y; Ji Q; Liang Z; Deng X; Wong CC; Yi C; Zhang L; Xie S; Alvarez S et al. Protein cysteine phosphorylation of SarA/MgrA family transcriptional regulators mediates bacterial virulence and antibiotic resistance. Proc. Natl. Acad. Sci. U S A 2012, 109, 15461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (311).Meins M; Jeno P; Muller D; Richter WJ; Rosenbusch JP; Erni B Cysteine phosphorylation of the glucose transporter of Escherichia coli. J. Biol. Chem 1993, 268, 11604. [PubMed] [Google Scholar]
  • (312).Chooi KP; Galan SR; Raj R; McCullagh J; Mohammed S; Jones LH; Davis BG Synthetic phosphorylation of p38alpha recapitulates protein kinase activity. J. Am. Chem. Soc 2014, 136, 1698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (313).Chalker JM; Lercher L; Rose NR; Schofield CJ; Davis BG Conversion of cysteine into dehydroalanine enables access to synthetic histones bearing diverse post-translational modifications. Angew. Chem. Int. Ed. Engl 2012, 51, 1835. [DOI] [PubMed] [Google Scholar]
  • (314).Rowan FC; Richards M; Bibby RA; Thompson A; Bayliss R; Blagg J Insights into Aurora-A kinase activation using unnatural amino acids incorporated by chemical modification. ACS Chem. Biol 2013, 8, 2184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (315).Bertran-Vicente J; Penkert M; Nieto-Garcia O; Jeckelmann JM; Schmieder P; Krause E; Hackenberger CP Chemoselective synthesis and analysis of naturally occurring phosphorylated cysteine peptides. Nat. Commun 2016, 7, 12703. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • (316).Chalker JM; Gunnoo SB; Boutureira O; Gerstberger SC; Fernández-González M; Bernardes GJL; Griffin L; Hailu H; Schofield CJ; Davis BG Methods for converting cysteine to dehydroalanine on peptides and proteins. Chem. Sci 2011, 2, 1666. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp Material
Supp Table 1

RESOURCES