Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2025 Oct 7;122(41):e2506036122. doi: 10.1073/pnas.2506036122

Deciphering the determinants of recombinant protein expression across the human secretome

Helen O Masson a,1, Pablo Di Giusto b,c,1, Chih-Chung Kuo a, Magdalena Malm d, Magnus Lundqvist d, Åsa Sievertsson d, Anna Berling d, Hanna Tegel d, Sophia Hober d,e, Mathias Uhlen e,f,g, Luigi Grassi h, Kimberly Robasky c,i,j, Chen-Lin Hsieh c, Diane Hatton h, Johan Rockberg d,1,2, Nathan E Lewis a,b,c,1,2
PMCID: PMC12541331  NIHMSID: NIHMS2109099  PMID: 41055974

Significance

Chinese hamster ovary (CHO) cells are the standard host for therapeutic protein production, yet some proteins remain challenging to express. By profiling a diverse panel of human secretome proteins expressed in CHO cells, we find that transgene mRNA abundance explains <1% of titer variability, while physicochemical traits (e.g., molecular weight, cysteine content) account for ~15%. Meanwhile, host‐cell transcriptomes reveal co‐varying secretory, metabolic, and stress‐response signatures, upregulated ER‐associated degradation and reticulophagy in poor producers, versus enhanced lipid metabolism and oxidative‐stress resilience in top performers. These insights quantify the contributions of known factors and highlight host-cell transcriptional differences as major sources of variability in protein expression, informing rational CHO cell-line and bioprocess engineering strategies.

Keywords: recombinant protein, protein secretion, Chinese hamster ovary cells, transcriptomics, machine learning

Abstract

Protein secretion is an essential process of mammalian cells. In biomanufacturing, this process can be optimized to enhance production yields and biotherapeutic quality. While cell line engineering and bioprocess optimization have yielded high protein titers for some recombinant proteins, many remain difficult to express. Here, we investigated factors influencing protein expression in Chinese hamster ovary (CHO) cells, expressing 2,135 Human Secretome Project proteins. While the abundance of mRNA from recombinant proteins explained less than 1% of observed variation in secretion titers, analysis of 218 biochemical and biophysical descriptors uncovered intrinsic protein features that account for ~15% of secretion variability, pinpointing key drivers such as molecular weight, cysteine content, and N-linked glycosylation, and establishing a roadmap for rational design of difficult-to-express proteins. We subsequently analyzed RNA-Seq data from 95 CHO cell cultures, each expressing a distinct recombinant protein, spanning a wide range of titers. Host cell transcriptomic signatures showed strong correlations with titer, thereby providing insights into cellular processes that covary with expression. Cells failing to produce proteins exhibited increased ubiquitin-mediated proteasomal degradation, including ER-associated degradation; whereas high-producing cells demonstrated enhanced lipid metabolism and a stronger response to oxidative stress, suggesting these factors may support successful recombinant protein productions. Together, using this resource, we quantified the contributions of various protein and cellular factors that correlate with the expression of diverse recombinant human proteins in a heterologous host, thereby providing insights for next-generation CHO cell engineering.


Protein secretion is a fundamental biological process essential for cellular function, responsible for the synthesis, processing, and delivery of proteins encoded by approximately one-third of all protein-coding genes in mammals, including both secreted and membrane-bound proteins (1, 2). These proteins are pivotal in mediating cellular interactions with the extracellular environment and regulating a broad spectrum of biological processes (3, 4), underscoring their critical role in maintaining cellular homeostasis, cell signaling, and intercellular communication.

Understanding the factors that influence protein secretion is critical for studying diverse diseases and advancing the production of life-altering biotherapeutics. Thus, optimizing protein secretion in host cells, particularly mammalian cells, is crucial for enhancing production yields and protein quality (5, 6). Chinese hamster ovary (CHO) cells are the leading mammalian expression platform for producing recombinant biopharmaceuticals in part due to their scalability and capacity to perform human-like posttranslational modifications (PTMs) (7). To systematically assess the potential of CHO cells to produce secreted proteins, the Human Secretome Project (HSP) was launched, aiming to study the entire set of secreted proteins (8). Although the HSP has extensively characterized this critical subset of the human proteome as a resource for drug discovery, it remains unclear which factors correlate with successful versus failed protein secretion (8, 9).

To elucidate variability in recombinant protein expression in CHO cells, we quantified the contributions of transgene transcript abundance and protein-specific attributes and explored the transcriptomics differences between high and low producers. Transgene mRNA levels explained less than 1% of the variation. Using a dataset of 2,135 HSP proteins, we curated 218 features, covering amino acid composition (AAC), biophysical parameters, and PTMs and used correlation and machine learning (ML) analyses to uncover factors potentially influencing protein secretion. While molecular weight (MW), disulfide bonds, cysteine composition, and N-linked glycosylation emerged as key predictors, these accounted for only about 15% of the variability. Host‐cell transcriptomic signatures in a panel of 95 CHO cultures elucidated genes and pathways whose expression correlated with secretion titers. Low‐producing cultures exhibited enrichment of ubiquitin–proteasome and ER-associated degradation pathways, whereas high producers showed upregulation of lipid metabolism and oxidative stress response pathways. These results highlight the dominant role of host cell physiology and open avenues for engineering improved CHO-based biomanufacturing processes.

Results

Recombinant Protein Expression in CHO Varies Extensively.

Among the 2,135 human proteins expressed in CHO in the aforementioned study, we observed considerable variability in total protein expression. Only 59% of the human secretome (1,257 proteins) were successfully expressed in CHO cells above the established quality threshold (8). Furthermore, among the proteins that met quality standards, titers varied by several orders of magnitude depending on the protein (Fig. 1A).

Fig. 1.

Fig. 1.

Production of the human secretome in CHO. (A) Histogram and cumulative distribution of total protein produced for 1,257 (59%) recombinant human secretome proteins successfully expressed using the high-throughput production pipeline in CHO cells. The amount of recovered protein for expressed proteins ranged between 44 and 5,388 µg. (B) Relationship between transgene mRNA abundance (TPM) and amount of secreted protein (µg). The CHO cell line was unable to produce any recoverable product for 4 of the selected recombinant proteins (blue), while cells with the top 15 highest and lowest yields are colored in pink and green, respectively. Cells expressing the remaining proteins are shown in gray.

From this panel of 1,257 proteins, we selected a representative subset of 95 for transfection and RNA-Seq analysis, ensuring that this group spanned a broad range of cell counts and viabilities at harvest, encompassing both high and low values in each category. The selection also aimed to capture different growth profiles by including proteins produced by fast-growing and slow-growing cell lines. This stratified approach ensured that the subset included diverse conditions related to cell growth and protein production. For comparison, we also included the wild-type (WT) Icosagen QMCF CHO-S host cell line. A total of 96 single-measurement cell cultures were subjected to RNA-Seq (Datasets S1 and S7). Notably, the transgenes, as defined by their recombinant sequences, consistently accounted for approximately 3% of the entire transcriptome, making them among the most highly expressed genes in most samples.

Variation in Recombinant Protein Expression Cannot be Explained by Transgene mRNA Abundance.

Some studies report that low transgene mRNA levels can limit secreted protein titers (10, 11). To evaluate whether the variation in protein production in our panel of cells can be explained by transgene mRNA levels, we modeled the relationship between transgene levels and protein expression using linear regression. Across the 95 RNA-sequenced recombinant protein-expressing cell cultures, we found that transgene mRNA abundance explained less than 1% of the variance in protein titer (Fig. 1B). This finding is markedly lower than the approximately 40% correlation reported for endogenous genes in other mammalian cell studies (5, 1214). This contrast is likely due to the high mRNA expression levels achieved in the transient Icosagen QMCF CHO cell system, which may saturate the translational machinery and minimize the impact of mRNA abundance on protein production. We conclude that adequate transgene mRNA is produced in these cells, and transgene mRNA abundance is likely not the limiting factor in our expression system. These results allow us to identify alternative bottlenecks in the production of difficult-to-express proteins studied here.

A Large Set of 218 Features Describing the HSP Proteins.

Since transgene mRNA levels could not explain the variance in recombinant protein production in our system, we wondered how protein-specific features contribute to the variability in protein expression. To test this, we curated a set of 218 protein features as potential predictors of abundance of the 2,135 HSP proteins. These features fall within three main classifications: i) experimental abundance, ii) sequence features, and iii) biophysical features (Table 1). Experimental abundance features measure the expression of the protein in other systems including various human tissues, other species, and the endogenous expression in CHO. Sequence features encompass attributes linked to the nucleotide and amino acid sequence of the protein such as MW, AAC, and PTMs. Finally, biophysical features cover metrics related to protein stability, solubility, secondary structure, etc. A detailed description of all features can be found in Dataset S2. The influence of these protein features on protein expression was investigated using correlation and ML methods.

Table 1.

Protein features and their sources

Feature type Feature group # Features Description Source/Software packages
Experimental abundance Production in mouse 18 Protein and mRNA copy numbers, half-lives, transcription rates, and translation rate constants in mouse fibroblasts 10.1038/nature10098 (5)
Production in Escherichia coli 1 Production yield of fusion proteins with fractions of human secretome in E. coli. 10.1093/bioinformatics/btx207 (16)
Human tissue expression 17 Secretome expression in various human tissues GTEx (17)
Human tissue protein level 19 Protein level across different human tissues HPA (18)
Endogenous expression of CHO ortholog 6 Endogenous expression of CHO ortholog under various conditions

This study

10.1038/srep40388 (19)

Sequence features MW 1 MW of protein This study
PTMs 29 Number of PTMs normalized with respect to sequence length

10.1371/journal.pone.0063284 (20)

iPTMnet (21)

ScanProsite (22)

AA composition 20 AAC This study
AA composition correlation with CHO 22 Correlation of AAC with AAC in native CHO cells
AA class composition 30 Global percentage of various AA classes

Peptides (23)

protr (24)

AA class transition 21 Percent frequency of transitions between pairs of AA classes protr (24)
RNA secondary structure 3 RNA minimum free energy (MFE), normalized ensemble free energy (EFE), and MFE normalized with respect to sequence length RNAfold (25)
Biophysical features Stability 4 Stability, instability, and aliphatic indices

Peptides (23)

ProtParam (26)

ProTstab (27)

Solubility 7 Isoelectric point, net charge, percent solubility, and grand average of hydrophobicity (GRAVY)

Peptides (23)

ProtParam (26)

Protein-Sol (28)

PPI potential 1 Potential protein–protein interaction index. Peptides (23)
Secondary structure 11 3- and 8-category predictions of protein secondary structure Scratch (29)
Relative solvent accessibility 8 Solvent-accessible fraction, percent hydrophobic and hydrophilic solvent-accessible residues, mean accessibility score, and GRAVY of inner and outer residues Scratch (29)

Protein Features Account for ~15% of the Variability in Recombinant Protein Expression.

Correlation analysis revealed that MW is by far the strongest determinant of protein expression (SI Appendix and Dataset S3). However, each cell has a limited amount of resources (e.g., amino acids, secretory pathway machinery, etc.), so to uncouple these known resource constraints from protein production, we analyzed how protein features affect the total producible mass of protein (µg) in CHO. To this end, we used ML to generate descriptive regression and pass/fail classification models of recombinant protein production in CHO (SI Appendix, SI Materials and Methods). Predictor variable (i.e., protein feature) importance for each model was ranked, and the consensus among the top 10 predictors for each model was evaluated (Fig. 2 A and B). All eight regression models ranked MW and AAC of cysteine among the top 10 most important features affecting total protein expression (µg). Furthermore, the best-performing regression model ranked these predictors as the most important features affecting expression (Fig. 2C). The pass/fail classification models showed increased consensus among important protein features. Among the consensus features were the presence of disulfide bonds and N-linked glycans, which are critical for folding and quality control of glycoproteins, specifically through the calnexin/calreticulin cycle (29, 30). A simple Welch’s t test indicated that the proteins that failed to produce have significantly lower N-linked glycosylation (P < 2.2e-16). Together this suggests that proteins with greater N-linked glycosylation tend to express better.

Fig. 2.

Fig. 2.

Protein-specific features affect recombinant protein expression (µg). A compilation of the top 10 most important features identified in the 8 regression (A) and 8 classification (B) models. A consensus of 8 indicates that the feature was ranked among the top 10 most important features in all 8 models. Regression models showed lower consensus, highlighting a total of 32 features, only 2 of which showed up in the top 10 features of all 8 models (consensus = 8). However, the classification models showed higher consensus highlighting a total of 15 features, wherein 4 of them have been deemed highly important in all 8 models (consensus = 8). (C) Bar graph showing the most influential protein features identified in the best-performing regression model. Variable importance measures have been scaled to have a maximum value of 100, and their directional effect on protein expression has been inferred and colored based on feature correlation with protein titer (µg). (D) Variability in protein titers explained by protein features was determined by sequentially adding protein features to a linear regression model and calculating the percent variability explained by the set of features. AA comp: amino acid composition; MW: molecular weight. A detailed description of each protein feature can be found in Dataset S2.

Shapley additive explanations (SHAP) analysis further supported these findings, identifying the presence of disulfide bonds, protein net charge, N-linked glycosylation, and MW as the primary features driving the model’s predictions of whether a protein is secreted or not (SI Appendix, Fig. S1A). Additionally, this analysis highlighted factors associated with failures in protein production, including low disulfide bond content, reduced N-linked glycosylation, insufficient signal peptides, and excessively high net charge or MW (SI Appendix, Fig. S1B).

To quantify the variability in total protein expression that can be explained by protein features, we sequentially added the ranked features of the best-performing regression model to a linear model fit and calculated the fraction of variance explained by the model (Fig. 2D). The explained variance peaks at approximately 15% when 32 protein features are included. While significantly greater than the variability explained by transgene mRNA abundance, protein features only account for a fraction of the variability in protein titers.

Given that industrial bioprocesses often focus on maximizing overall product yield, it is critical to account for the potential impact of protein features on cellular growth and viability, as these factors ultimately influence productivity. To explore this, we conducted an analysis examining the correlation between protein features of the expressed protein with cell viability and cell growth. We observed that alanine and leucine composition, net charge, and physicochemical “transition” descriptors (small-medium/large volume shifts and helix-coil flips) were positively associated with both metrics. In contrast, aliphatic residue content and predicted α‐helical propensity showed a stronger relationship with cell growth alone, whereas disulfide‐bond density and positive-negative charge transitions correlated specifically with viability (SI Appendix, Fig. S3).

To further explore how protein features relate to viable cell specific productivity, we normalized protein expression by both cell growth and viability. In both analyses, the top positive correlates were features characterized by frequent alternating physicochemical transitions (i.e., transition polarizability: low-medium and transition secondary structure: helix-coil) together with elevated proline and glycine content, and MW; when normalizing by viability, avoiding large‐volume or heavily charged/aromatic stretches seems to preserve cell health, whereas when normalizing by growth rate, increases in O-glycosylation sites and balanced hydrophobic-neutral motifs helps are associated with stable cell division rate under secretory load (SI Appendix, Fig. S4).

Transcriptomic Profiles Distinguish Producing from Nonproducing CHO Cells.

Next, we sought to characterize the transcriptomic signatures across host‐cell clones and evaluate their association with recombinant protein expression. Principal component analysis (PCA) of the 96 RNA-Seq samples clearly shows that the nonproducing cells are transcriptional outliers compared to the cells producing recombinant protein (Fig. 3A). The first principal component (PC1), which accounts for approximately 19% of transcriptome variability, seemingly separates successfully producing cells from those that failed to produce any recombinant protein. These transcriptomic profiles reflect the cellular response to protein production, they are covarying outputs of the biological system rather than direct drivers of protein expression differences.

Fig. 3.

Fig. 3.

Nonproducing cell lines are transcriptional outliers. (A) PCA of transcriptomics data. Top five genes with a positive and negative loading to the PC1 are shown in light gray. The dashed red line shows a clear division between the cells capable of producing recombinant proteins (pink, green, and gray) and cells that failed to produce any detectable protein (blue). (B) Results from a gene set enrichment analysis (GSEA) performed between the failed producers and the cells that successfully produced protein. Terms with a positive normalized enrichment score (NES) are enriched in the nonproducers, while terms with a negative NES are enriched in the producers.

LOC100754005 (Prpf8) is one of the genes with the strongest negative loading on PC1. This protein functions as a component of the spliceosome complex which is critical for pre-mRNA processing. We found that higher expression of this gene characterizes the protein-producing cells. In line with this finding, previous work comparing the proteome of various CHO host cells revealed an up-regulation of Prpf8 in the high-producing cell lines and alluded to its contribution to the high production of biopharmaceuticals in CHO (31).

To gain additional insights into biological pathways and processes characteristic of the nonproducers, we conducted GSEA (32, 33) between the failed producers and the cells that successfully produced protein (Fig. 3B). Unsurprisingly, results showed signs of cell stress in both groups, likely due to the burden of overexpressing a foreign protein. Additionally, we found that the failed producers upregulated genes involved in translation and oxidative phosphorylation and showed signs of amino acid deficiency. We also observed increased activity in the early stages of protein secretion (i.e., targeting to the ER) in the failed producers and depletion in later portions of the secretory pathway (i.e., intra-Golgi traffic, vesicle targeting, protein processing, and secretory granules) compared to the producers. Furthermore, the successful producers showed increased transmembrane transport, potentially alleviating the burden of amino acid deficiency.

Cells that Fail to Produce Protein Show an Inadequate Response to ER Stress.

The secretory pathway is often implicated as a bottleneck during recombinant protein production (3436). To better understand the differences in secretory pathway signatures within our panel of cells, we calculated activity scores (see Materials and Methods for details) for 13 secretory pathway functions. Activity scores for the 95 recombinant protein-expressing CHO cells were normalized to express the change in pathway activity with respect to the WT host cell (Fig. 4A).

Fig. 4.

Fig. 4.

Secretory pathway cell signatures. (A) Clustered heatmap of the normalized change in secretory pathway activity compared to WT for each of the 95 recombinant protein expressing cells. Highlighted here are four clusters that show distinct secretory pathway footprints. Rows/cells are annotated according to the amount of protein they produce. (B) Chord diagram of differentially expressed ER stress, UPR, and ERAD genes (|FC| > 1.5; FDR ≤ 0.01) between the failed and successful producers. A positive LFC indicates higher expression in the failed producers and vice versa. (C) Lollipop plot showing the significant correlations (|r| ≥ 0.6; FDR ≤ 0.1) between secretory pathway genes and protein abundance among the cells in Cluster 3.

ER calcium homeostasis seemed to be increased across all recombinant protein-expressing cells regardless of productivity, suggesting that overexpression of heterologous proteins in CHO is associated with a general imbalance in ER calcium homeostasis, leading to an activation of the unfolded protein response (UPR) (37). An in-depth analysis of cellular response to stress showed activation of many ER stress response genes among the panel of cells, with the majority of genes showing depletion in the failed producers (Fig. 4B).

Protein folding, in particular, is a common bottleneck in recombinant protein production (38, 41, 42), and the accumulation of improperly folded proteins can trigger ER stress and UPR (4345). Consequently, many studies have found that the upregulation of protein folding genes enhances protein production (46, 48, 49). Our results showed a mild, yet significant, correlation between protein folding activity and total protein production (r = 0.21, P-value = 0.05) among the 91 successfully producing cells. Additionally, we observed upregulation of several genes regulating protein folding, including ER stress-induced genes hypoxia up-regulated 1 (Hyou1), protein disulfide isomerase family A member 3 (Pdia3), and endoplasmic reticulum oxidoreductase 1 alpha (Ero1a) in the successfully producing cells compared to the failed cells (Fig. 4B). Overall, this modest correlation suggests an underlying complexity likely driven by an interplay of mechanisms that ensure efficient folding and proper function of glycoproteins, potentially influenced by varying folding demands among different proteins.

N-Linked Glycosylation and ERAD Correlate with Protein Secretion.

We identified four distinct groups when clustering the RNA-Seq data from 95 recombinant protein-expressing cell culture samples based on secretory pathway activity (Fig. 4A). Interestingly, Cluster 4 consists of a single failed producer (Ccl20) that shows considerable decreases across all secretory pathway functions. The cells in the remaining clusters show a range of productivity, suggesting that these secretory pathway fingerprints do not define a cell’s ability to successfully produce and secrete recombinant protein. Cluster 3 was of particular interest since it included the remaining samples that failed to produce the human proteins of interest. The samples in this cluster showed low activity across secretory functions. However, along with the failed producers, it also contained some of the highest producers. This suggests that high secretory pathway activity alone is not sufficient for achieving high protein production. This motivates us to explore other adjacent pathway genes within Cluster 3 that may correlate with high production.

We calculated correlations between individual secretory pathway genes and protein abundance in the cells belonging to Cluster 3. We identified 43 secretory machinery genes that showed significant correlation (|r| ≥ 0.6; FDR ≤ 0.1) with protein abundance (Fig. 4C and Dataset S8). Notably, these signatures are unique to Cluster 3 as the other two clusters do not show significant correlation with the expression of secretory pathway genes. We observed a set of N-linked glycosylation genes, such as Alg12, Rpn1, Rpn2, and Ddost that positively correlated with protein expression. Alg12, encodes glycosyltransferase involved in the assembly of the dolichol-PP-oligosaccharide precursor required for N-linked glycosylation. Rpn1-2 and Ddost encode proteins of the oligosaccharide transferase complex (OST complex), which catalyzes the first step of N-linked glycosylation: transfer of the preassembled N-linked glycan from the dolichol lipid carrier to the client protein. Since ER N-linked glycosylation is important for protein folding: stabilizing proteins, preventing aggregation, and providing quality control through interactions with chaperones (50), this role may be one of the mechanisms explaining why N-linked glycosylation signals show a stronger correlation with protein production compared to protein folding signals.

Only a single gene, Derlin 2 (Derl2), negatively correlated with protein expression. The Derlin family of genes encode components of ERAD machinery, where they participate in the retrotranslocation of unfolded and misfolded proteins from the ER to the cytosol for proteasomal degradation (51, 52).

To evaluate how variation in secretory‐machinery gene expression associates with heterogeneity in Cluster 3, we fitted a linear regression model. We began with five candidates (Alg12, Rpn1, Rpn2, Ddost, and Derl2) but, to mitigate multicollinearity, we ultimately retained only Alg12 and Derl2. Alg12 stands out among other N-linked glycosylation genes due to its dual role in the glycosylation process and protein folding within the ER. In a linear regression of protein titer on Alg12 and Derl2 expression, we observed an R2 of 0.87 across Cluster 3 clones (SI Appendix, Fig. S2), indicating a strong correlation between these two genes’ expression levels and titer variability. However, this high R2 largely reflects the collinearity inherent in our RNA-Seq dataset. We highlight N-linked glycosylation (Alg12, Rpn1/2, Ddost) together with ERAD (Derl2) because these statistically significant Cluster-3 signals represent opposing arms of ER quality control (glycan transfer/maturation vs. retro-translocation), with Derl2 being the only significant negative correlate, and because Alg12 and Derl2 were retained as representative, noncollinear markers in our parsimonious regression (Fig. 4C and Dataset S8).

High Producers Are Characterized by Increased Lipid Metabolism.

Recombinant protein production is energy intensive with increased raw material demands, thus inducing significant alterations in host cell metabolism. Consequently, many cell line engineering efforts have targeted metabolism to enhance recombinant protein production (53). To identify metabolic variation within our panel of cells, we implemented the CellFie tool (54, 55), which quantifies metabolic task activity from omics data (Dataset S4). We identified 79 core metabolic tasks active in all cells, 27 tasks inactive across all cells, and 79 tasks with differential activation (Fig. 5A). Many differentially active tasks are involved in amino acid, carbohydrate, and lipid metabolism. Of the 79 tasks showing differential activation across the panel of CHO cells, the nonproducers showed on average 33% active metabolic tasks, while the highest and lowest producers showed 66% and 58%, respectively. The CHO cell cultures used in this study are identical prior to transfection, with the transgene being the only experimental variable. This design ensures that observed differences in protein production and metabolic responses are attributable to the expressed gene and its associated demands on cellular resources. The observed metabolic changes represent a cellular response to the demands of protein production, reflecting an adaptive reprogramming process where the cell reallocates resources to support the expression of recombinant proteins. Other studies have reported similar metabolic restructuring when comparing cells producing secreted vs. intracellular proteins, implicating increased energy demand of the secretory pathway during recombinant protein production (56).

Fig. 5.

Fig. 5.

Metabolic cell signatures. (A) Pie chart showing the proportion of metabolic tasks that are active, inactive, and differentially active among the 95 recombinant protein-expressing cells. Boxplot to the right shows the percentage of active metabolic tasks for the different productivity groups among the 79 differentially activated tasks. (B) Treemap of CellFie metabolic tasks organized into systems and subsystems. Each square represents a single metabolic task which is colored according to significant correlation (FDR ≤ 0.1) with protein expression among the high- and low-producing cell lines. Tasks with nonsignificant correlations are colored in gray. (C) Volcano plot of differentially expressed stress response genes between the failed and successful producers. A positive LFC indicates higher expression in the failed producers and vice versa. Genes specifically involved in response to oxidative stress have been highlighted in red.

To further understand the metabolic differences differentiating the high and low producers, we calculated the correlation between individual metabolic tasks and protein expression among the subset of high and low producers (Fig. 5B). The metabolic tasks with the largest and most significant correlation with protein expression are involved in fatty acid (FA) metabolism. FAs have a diverse range of important cellular functions, including critical structural components of cell membranes and fundamental energy supplies for the cell. In particular, we observed a strong positive correlation with synthesis of several FAs: palmitate synthesis (r = 0.62), synthesis of palmitoyl-CoA (r = 0.59), arachidonate synthesis (r = 0.59), and synthesis of malonyl-CoA (r = 0.51). Additionally, the low-producing cells show greater conversion of aspartate to beta-alanine (r = −0.43), a precursor of Coenzyme A (CoA) and acyl carrier protein (ACP) involved in FA metabolism. This negative correlation could suggest a depletion of these important precursors in the low-producing cells.

Strong and Dynamic Response to Oxidative Stress Enhances Recombinant Protein Production.

Our analysis shows that proteins with high cysteine content are poorly expressed, consistent with reports that excess cysteines drive disulfide overformation, aggregation, instability, and yield loss (57). In addition, suboptimal cysteine supplementation perturbs redox homeostasis, thereby reducing titer, productivity, and product quality in CHO cultures (58). To explore adjacent pathways that might be related to a specific imbalance in cysteine metabolism under our experimental conditions, we performed additional metabolic analyses using CellFie, which revealed a negative correlation (r = −0.51) between cysteine depletion (via taurine synthesis) and protein expression (Fig. 5B). Diverting cysteine into taurine synthesis may further burden protein production by activating amino acid deprivation pathways (59) and impairing mitochondrial function, thereby reducing oxidative phosphorylation (58). While cysteine levels in the spent medium were not directly measured here, we recognize this as a limitation. Measuring cysteine levels would provide direct evidence to validate or refute the hypothesis of external cysteine depletion. Moreover, in intensified processes, supplementation of cysteine at high concentrations is a challenge due to its limited solubility and instability in solution (60). This could be a result of increased oxidative stress and subsequent damage in the low-producing cells.

The prevalence of oxidative stress within our panel of cells was evident in the analysis of cellular response to stress. Oxidative stress occurs when there is an imbalance between antioxidant defenses and the accumulation of oxygen reactive species (ROS), which are generated during recombinant protein production in CHO (61). We noticed that the successfully producing cells show a more profound response to oxidative stress, upregulating 1.6x as many oxidative stress response genes compared to the nonproducing cells (Fig. 5C). Interestingly, we observed that three of the genes depleted in the failed producers encode proteins belonging to the solute carrier (Slc) superfamily. This supports the negative enrichment in Slc transmembrane transport observed in the preliminary GSEA (Fig. 3B). As a major family of transmembrane proteins responsible for the transport of essential nutrients and metabolites, Slc proteins are critical in many essential physiological functions including oxidative stress (62). Solute carrier family 7 member 11 (Slc7a11) shows the greatest depletion among oxidative stress genes in the failed cells (LFC = −1.85; FDR = 5.19E-07). Slc7a11 is involved in the specific transport of cysteine and glutamate, which could help alleviate the stress of cysteine depletion (63).

The failed producers showed a significant increase (LFC = 1.46; FDR = 2.46E-67) in the oxidative stress sensor Kelch Like ECH Associated Protein 1 (Keap1). Keap1 acts as a substrate-specific adapter of the E3 ubiquitin ligase complex BCR (BTB-CUL3-RBX1) which mediates ubiquitination and degradation of the transcription factor (NFE2-like bZIP transcription factor 2) Nrf2. In response to oxidative stress, modifications of the highly reactive cysteine residues in Keap1 result in inactivation of the ubiquitin ligase activity of the BCR complex and subsequent expression of Nrf2-induced phase 2 detoxifying enzymes (64, 66, 67). Interestingly, the Nrf2/Keap1 pathway also regulates the expression of many genes involved in lipid metabolism (6870) as well as cysteine metabolism and transport (7173). Increased expression of the negative regulator Keap1 could be hindering the cell’s ability to mount an adequate Nrf2-induced transcriptional response to oxidative stress.

To further elucidate the link between cysteine composition and oxidative stress, we performed a correlation analysis of oxidative stress gene expression with both total protein production and cysteine composition (SI Appendix, Fig. S5). This analysis shows that Keap1, Endog, and Trex1 genes significantly upregulated in nonproducing cells are positively correlated with cysteine content and negatively correlated with protein expression, whereas genes enriched in high-producing cells display the opposite trend. These results suggest that high cysteine content had a negative impact on the cellular antioxidant defenses, contributing to the low-producing phenotype. Together these results suggest that, within the panel of proteins used in this study, the ability to properly respond to oxidative stress positively correlates with recombinant protein production.

Discussion

The continual development of new biologics is accompanied by pressure to establish methods and technologies for enhancing product quality and host productivity. However, many proteins struggle to express well or at all in the nonnative environment of CHO-based biomanufacturing. The HSP demonstrated that even standard human proteins can be difficult to produce in CHO. In this study, we leverage this resource to explore why CHO cells produce some proteins better than others. While many studies have focused on identifying biological drivers of protein production in CHO, the primary aim of this study was to quantify the variability explained by the inputs in this analysis, represented by the recombinantly expressed proteins in the CHO cells, and to explore the transcriptomics signatures associated with the high producing and low producing phenotypes. Together these factors may help guide the experimental approaches and rational design of protein-producing CHO cell platforms.

Here, we found that transgene mRNA levels were consistently high, explaining less than 1% of the variability in protein expression. While such abundant transgene expression rules out mRNA availability as a limiting factor in this study, it may still be a concern in other systems. Notably, low transgene mRNA expression in CHO cells has become increasingly rare (10, 11).

Using statistical and ML methods, we quantified how 218 protein features affect the expression and secretion of proteins in CHO. Overall, we found the protein features explored in this study account for only a fraction of the observed variability in protein expression (~15%). Among the 218 features, MW, the presence of disulfide bonds, cysteine composition, and N-linked glycosylation had the strongest effect on protein expression. Previous studies found that protein size is the primary factor in determining folding rates and protein stability (74), two critical factors affecting protein secretion. Our investigation also brought attention to cysteine composition as a notable protein attribute. Here, we noted a significant negative correlation with protein expression. Given that we found strong transcriptional signatures of oxidative stress between failed and successful producers in our panel of cells, high cysteine composition could be int roducing destabilizing nonnative disulfide bonds (75, 76). In fact, studies attempting to stabilize proteins by introducing artificial disulfide bridges have found that it can lead to overall protein destabilization (5761). Additionally, while this study only looked at AAC, codon utilization can affect protein yield in CHO (77, 79, 80) and could be an area of future exploration for quantifying protein variability.

Protein misfolding and aggregation are an increasingly important area of study in protein chemistry and molecular medicine (81). In addition, the biochemical properties of the transgene itself may also affect cell growth and viability, which in turn impact overall productivity (82). We therefore examined how individual protein features correlate with viability and growth rate. We found that higher alanine and leucine content, overall net charge, and certain transition descriptors (small-medium/large volume shifts and helix-coil transitions) all track positively with both metrics. By contrast, the fraction of aliphatic residues and predicted α-helix content correlated more strongly with growth, whereas disulfide-bond density and positive-negative charge transitions correlated primarily with viability. After normalizing protein titers by either cell viability or cell growth, the most strongly positive correlates were low-medium polarizability and frequent helix-coil transitions, together with elevated proline and glycine content [residues known to promote local disorder and helix-coil interconversion (83)] and MW. These patterns likely reflect a balance between secretory burden and maintenance of cell homeostasis. Notably, a higher overall fraction of polar residues has been shown to enhance recombinant protein expression in E. coli (15).

Most of the variability in protein production could not be accounted for by intrinsic protein features or by transgene mRNA levels. We therefore turned to host-cell transcriptomic profiles to identify signatures that distinguish high- from low-producing clones. It is important to emphasize that protein‐specific attributes represent defined experimental inputs, whereas the transcriptome is a downstream readout of the cell’s phenotype. Because we measured transcriptomic changes after transient transfection rather than perturbing them directly, their tight correlation with product titer reflects co‐variation to the perturbation, not proof that these pathways drive protein expression differences.

We identified coordinated changes in secretory-pathway, metabolic, and stress-response genes that distinguish successful from failed producers. Successful producers upregulate UPR and folding machinery, whereas failed producers upregulate ERAD/ERpQC (ER protein quality control) and reticulophagy, as previously reported (84, 85), along with a downregulation of the UPR. Interestingly, others have observed a similar increase in proteasomal degradation without triggering UPR in difficult to express protein production (85). Furthermore, we observed a group of genes associated with N-linked glycosylation whose expression covaried with protein titer. To illustrate how one pair of these markers captures much of that shared signal, we built a linear regression using Alg12 and Derl2 expression and achieved an R2 of 0.87 across Cluster 3 clones (SI Appendix, Fig. S2). It is important to note, however, that this high R2 primarily reflects the redundancy inherent in our highly correlated omics dataset. Thus, Alg12 and Derl2 are representative markers of a broader group of genes that covary with protein production (Dataset S8). Future functional studies will be necessary to dissect the individual contributions of these and other correlated genes to recombinant protein expression and to identify the most promising candidates for engineering interventions.

While many efforts to improve recombinant protein production in CHO have focused on manipulating metabolism and/or components of the secretory pathway, it is surprising that, to date, there have been few studies exploring the effects of manipulating lipid metabolism in relation to bioproduction (8688). Lipids are the major component of cell membranes and are critical for a wide range of fundamental bioprocesses including energy metabolism, organelle formation and containment, vesicle trafficking and secretion, and cell signaling (86, 89). Here, we observed strong positive correlations between protein expression and metabolic tasks related to the synthesis of several FAs. In contrast, the conversion of aspartate to beta-alanine, a precursor of CoA and ACP, emerged as a negative predictor of recombinant protein production. Furthermore, we found that Keap1, an important regulator of lipid metabolism (90), exhibited elevated expression levels in failed producers compared to successful ones. We hypothesize that increased lipid metabolism enhances recombinant protein production by dynamically expanding the endomembrane system, as supported by prior work showing that overexpression of lipid‐biosynthesis genes expands the ER and boosts protein production (86).

Given these findings, variation in CHO expression of native human secreted proteins appears driven primarily by host-cell physiology and global cellular state rather than intrinsic protein sequence or physicochemical traits. To address the remaining unexplained variance, future work should apply causal perturbations that directly test mechanism. Barcoded secretion reporters coupled to pooled perturbations (e.g., CRISPRi/a or ORF overexpression) can nominate effectors whose modulation measurably changes protein secretion; for instance, a recent FcBAR-guided study reported that overexpressing Agpat4, Ephx1, and Nsdhl increased recombinant antibody secretion, implicating lipid remodeling and ER-localized activities as leverage points for productivity (91). Consistent with this, BAR-coupled RNA-Seq in glycoengineered CHO identified and validated host interactors Cul4a and Ywhah whose overexpression increased secretion of a difficult-to-express soluble HCV E1E2 vaccine candidate (92). In parallel, programmable tuning of ER capacity and proteostasis (e.g., UPR/ERAD and folding chaperones), as well as trafficking nodes (COPII/COPI and Golgi processing), offers orthogonal causal inputs to test secretory throughput limits (93). Finally, controlled bioprocess perturbations (temperature shift, chemical chaperones, feed composition) remain practical levers to induce interpretable changes in secretion phenotypes and can be integrated with multiomics readouts to attribute variance to specific mechanisms (82). Overall, this work provides a resource for future studies to tease apart the specific mechanisms limiting secretion of individual recombinant proteins, ultimately impacting the vast biologics industry. As such, biomanufacturing platforms seeking to improve product yield could leverage the expanding power of omics technologies to distill actionable insight for enhanced expression system designs (94).

Materials and Methods

Human Secretome Production Data.

Protein titers for the human secretome transiently expressed in the Icosagen QMCF cell line were obtained from a previously published dataset (8). We removed samples whose status is “Ongoing,” as well as samples that passed QC (Status = “Pass”) yet were missing titer information (Dataset S5). We note that as previously reported (8), the titers were estimated upon purification, which could influence the results if different proteins purified differently. However, all purifications relied upon the same peptide tag, thus minimizing potential biases. Because the original publication reported titers only as total protein expression (µg) rather than a concentration or specific productivity, we have retained the same units here to ensure consistency with the previously published dataset.

RNA-Seq Analysis.

The subset of 95 samples selected for RNA-seq was chosen based on an initial pilot study and is meant to represent a range of the highest and lowest growers and producers. Cells were grown and sampled in log phase following successive propagations. Cell pellets were resuspended in RNAlater Stabilization Solution (Invitrogen) according to the manufacturer’s recommendations until RNA extraction. Total RNA was extracted from three replicates of each cell line using Qiagen’s RNeasy plus Mini Kit according to the manufacturer’s instructions. Concentrations were determined with a NanoDrop ND-1000 spectrophotometer and RNA quality was assessed on a 2100 Bioanalyzer (Agilent Technologies) using RNA 6000 Nano chips (Agilent Technologies). All samples had an RNA integrity number of at least 9.9. Library preparation and sequencing was carried out at Scilifelab national genomics center using polyA-positive enrichment.

Additional experimental procedures and methods are listed in SI Appendix, SI Materials and Methods.

Supplementary Material

Appendix 01 (PDF)

Dataset S01 (XLSX)

pnas.2506036122.sd01.xlsx (13.7KB, xlsx)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2506036122.sd03.xlsx (28.1KB, xlsx)

Dataset S04 (XLSX)

Dataset S05 (XLSX)

pnas.2506036122.sd05.xlsx (806.9KB, xlsx)

Dataset S06 (XLSX)

pnas.2506036122.sd06.xlsx (479.4KB, xlsx)

Dataset S07 (XLSX)

pnas.2506036122.sd07.xlsx (22.3MB, xlsx)

Dataset S08 (XLSX)

pnas.2506036122.sd08.xlsx (34.5KB, xlsx)

Acknowledgments

This work was supported by generous funding from National Institute of General Medical Sciences (R35 GM119850), National Institute of Allergy and Infectious Diseases (UH2AI153029), the Novo Nordisk Foundation (NNF10CC1016517 and NNF20SA0066621), Swedish Foundation for Strategic Research (SB16-0017), the Swedish innovation agency Vinnova (GeneNova 2021-02640, CellNova 2017-02105, and AdBIOPRO 2022-03170), Swedish Research Council (NAISS, 2022-06725), AstraZeneca, and Knut and Alice Wallenberg foundation (Wallenberg Center for Protein Research).

Author contributions

H.O.M., P.D.G., J.R., and N.E.L. designed research; H.O.M., P.D.G., C.-C.K., M.M., M.L., Å.S., A.B., H.T., S.H., M.U., L.G., K.R., and D.H. performed research; J.R. and N.E.L. contributed new reagents/analytic tools; H.O.M., P.D.G., C.-C.K., M.M., M.L., Å.S., A.B., H.T., S.H., M.U., L.G., K.R., C.-L.H., D.H., J.R., and N.E.L. analyzed data; and H.O.M., P.D.G., C.-L.H., and N.E.L. wrote the paper.

Competing interests

N.E.L. is a co-founder of Augment Biologics and NeuImmune, Inc. with equity and stock. He is also a scientific advisor for CHO Plus, Inc. and Neion. L.G. and D.H. are employees of AstraZeneca and may own AstraZeneca stock or stock options.

Footnotes

This article is a PNAS Direct Submission.

Contributor Information

Johan Rockberg, Email: johanr@biotech.kth.se.

Nathan E. Lewis, Email: natelewis@uga.edu.

Data, Materials, and Software Availability

All data and code supporting the findings and analyses of this study are publicly available in the following GitHub repository (https://github.com/LewisLabUCSD/CHO96) (95) and GEO (GSE225989) (96). All other data are included in the manuscript and/or supporting information.

Supporting Information

References

  • 1.Kondylis V., Pizette S., Rabouille C., The early secretory pathway in development: A tale of proteins and mRNAs. Semin. Cell Dev. Biol. 20, 817–827 (2009). [DOI] [PubMed] [Google Scholar]
  • 2.Vázquez-Martínez R., et al. , Revisiting the regulated secretory pathway: From frogs to human. Gen. Comp. Endocrinol. 175, 1–9 (2012). [DOI] [PubMed] [Google Scholar]
  • 3.Stefan C. J., et al. , Membrane dynamics and organelle biogenesis-lipid pipelines and vesicular carriers. BMC Biol. 15, 102 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Barlowe C. K., Miller E. A., Secretory protein biogenesis and traffic in the early secretory pathway. Genetics 193, 383–410 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Schwanhäusser B., et al. , Global quantification of mammalian gene expression control. Nature 473, 337–342 (2011). [DOI] [PubMed] [Google Scholar]
  • 6.Hou J., Tyo K., Liu Z., Petranovic D., Nielsen J., Engineering of vesicle trafficking improves heterologous protein secretion in Saccharomyces cerevisiae. Metab. Eng. 14, 120–127 (2012). [DOI] [PubMed] [Google Scholar]
  • 7.Park S.-Y., et al. , Driving towards digital biomanufacturing by CHO genome-scale models. Trends Biotechnol. 42, 1192–1203 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Tegel H., et al. , High throughput generation of a resource of the human secretome in mammalian cells. New Biotechnol. 58, 45–54 (2020). [Google Scholar]
  • 9.Uhlén M., et al. , The human secretome. Sci. Signal. 12, eaaz0274 (2019). [DOI] [PubMed] [Google Scholar]
  • 10.Jiang Z., Huang Y., Sharfstein S. T., Regulation of recombinant monoclonal antibody production in Chinese hamster ovary cells: A comparative study of gene copy number, mRNA level, and protein expression. Biotechnol. Prog. 22, 313–318 (2006). [DOI] [PubMed] [Google Scholar]
  • 11.Eisenhut P., et al. , Systematic use of synthetic 5’-UTR RNA structures to tune protein translation improves yield and quality of complex proteins in mammalian cell factories. Nucleic Acids Res. 48, e119 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Liebermeister W., et al. , Visual account of protein investment in cellular functions. Proc. Natl. Acad. Sci. U.S.A. 111, 8488–8493 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Wilhelm M., et al. , Mass-spectrometry-based draft of the human proteome. Nature 509, 582–587 (2014). [DOI] [PubMed] [Google Scholar]
  • 14.Bhandari B. K., et al. , Analysis of 11,430 recombinant protein production experiments reveals that protein yield is tunable by synonymous codon changes of translation initiation sites. PLoS Comput. Biol. 17, e1009461 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Sastry A., et al. , Machine learning in computational biology to accelerate high-throughput protein expression. Bioinformatics 33, 2487–2495 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Carithers L. J., et al. , A novel approach to high-quality postmortem tissue procurement: The GTEx project. Biopreserv. Biobank. 13, 311–319 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Uhlén M., et al. , Proteomics. Tissue-based map of the human proteome. Science 347, 1260419 (2015). [DOI] [PubMed] [Google Scholar]
  • 18.Kallehauge T. B., et al. , Ribosome profiling-guided depletion of an mRNA increases cell growth rate and protein secretion. Sci. Rep. 7, 40388 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Feizi A., Österlund T., Petranovic D., Bordel S., Nielsen J., Genome-scale modeling of the protein secretory machinery in yeast. PLoS ONE 8, e63284 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Huang H., et al. , iPTMnet: An integrated resource for protein post-translational modification network discovery. Nucleic Acids Res. 46, D542–D550 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.de Castro E., et al. , Scanprosite: Detection of PROSITE signature matches and prorule-associated functional and structural residues in proteins. Nucleic Acids Res. 34, W362–W365 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Osorio D., Rondón-Villarreal P., Torres R., Peptides: A package for data mining of antimicrobial peptides. R J. 7, 4–14 (2015). [Google Scholar]
  • 23.Xiao N., Cao D.-S., Zhu M.-F., Xu Q.-S., Protr/ProtrWeb: R package and web server for generating various numerical representation schemes of protein sequences. Bioinformatics 31, 1857–1859 (2015). [DOI] [PubMed] [Google Scholar]
  • 24.Gruber A. R., Lorenz R., Bernhart S. H., Neuböck R., Hofacker I. L., The Vienna RNA websuite. Nucleic Acids Res. 36, W70–W74 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Gasteiger E., et al. , ExPASy: The proteomics server for in-depth protein knowledge and analysis. Nucleic Acids Res. 31, 3784–3788 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Yang Y., et al. , Protstab—Predictor for cellular protein stability. BMC Genomics 20, 804 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Hebditch M., Carballo-Amador M. A., Charonis S., Curtis R., Warwicker J., Protein-Sol: A web tool for predicting protein solubility from sequence. Bioinformatics 33, 3098–3100 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Cheng J., Randall A. Z., Sweredoski M. J., Baldi P., SCRATCH: A protein structure and structural feature prediction server. Nucleic Acids Res. 33, W72–W76 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ferris S. P., Kodali V. K., Kaufman R. J., Glycoprotein folding and quality-control mechanisms in protein-folding diseases. Dis. Model. Mech. 7, 331–341 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Lamriben L., Graham J. B., Adams B. M., Hebert D. N., N-glycan-based ER molecular chaperone and protein quality control system: The calnexin binding cycle. Traffic 17, 308–326 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Xu N., et al. , Comparative proteomic analysis of three Chinese hamster ovary (CHO) host cells. Biochem. Eng. J. 124, 122–129 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Subramanian A., et al. , Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. Proc. Natl. Acad. Sci. U.S.A. 102, 15545–15550 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Mootha V. K., et al. , PGC-1alpha-responsive genes involved in oxidative phosphorylation are coordinately downregulated in human diabetes. Nat. Genet. 34, 267–273 (2003). [DOI] [PubMed] [Google Scholar]
  • 34.Pérez-Rodriguez S., et al. , Compartmentalized proteomic profiling outlines the crucial role of the classical secretory pathway during recombinant protein production in Chinese hamster ovary cells. ACS Omega 6, 12439–12458 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Malm M., et al. , Harnessing secretory pathway differences between HEK293 and CHO to rescue production of difficult to express proteins. Metab. Eng. 72, 171–187 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Mathias S., et al. , Visualisation of intracellular production bottlenecks in suspension-adapted CHO cells producing complex biopharmaceuticals using fluorescence microscopy. J. Biotechnol. 271, 47–55 (2018). [DOI] [PubMed] [Google Scholar]
  • 37.Groenendyk J., Agellon L. B., Michalak M., Calcium signaling and endoplasmic reticulum stress. Int. Rev. Cell Mol. Biol. 363, 1–20 (2021). [DOI] [PubMed] [Google Scholar]
  • 38.Khan S. U., Schröder M., Engineering of chaperone systems and of the unfolded protein response. Cytotechnology 57, 207–231 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Shusta E. V., Raines R. T., Plückthun A., Wittrup K. D., Increasing the secretory capacity of Saccharomyces cerevisiae for production of single-chain antibody fragments. Nat. Biotechnol. 16, 773–777 (1998). [DOI] [PubMed] [Google Scholar]
  • 40.Reinhart D., Sommeregger W., Debreczeny M., Gludovacz E., Kunert R., In search of expression bottlenecks in recombinant CHO cell lines–a case study. Appl. Microbiol. Biotechnol. 98, 5959–5965 (2014). [DOI] [PubMed] [Google Scholar]
  • 41.Mead E. J., Chiverton L. M., Smales C. M., von der Haar T., Identification of the limitations on recombinant gene expression in CHO cell lines with varying luciferase production rates. Biotechnol. Bioeng. 102, 1593–1602 (2009). [DOI] [PubMed] [Google Scholar]
  • 42.Nishimiya D., Mano T., Miyadai K., Yoshida H., Takahashi T., Overexpression of CHOP alone and in combination with chaperones is effective in improving antibody production in mammalian cells. Appl. Microbiol. Biotechnol. 97, 2531–2539 (2013). [DOI] [PubMed] [Google Scholar]
  • 43.Schröder M., Kaufman R. J., The mammalian unfolded protein response. Annu. Rev. Biochem. 74, 739–789 (2005). [DOI] [PubMed] [Google Scholar]
  • 44.Kaufman R. J., Stress signaling from the lumen of the endoplasmic reticulum: Coordination of gene transcriptional and translational controls. Genes Dev. 13, 1211–1233 (1999). [DOI] [PubMed] [Google Scholar]
  • 45.Lam M., Marsters S. A., Ashkenazi A., Walter P., Misfolded proteins bind and activate death receptor 5 to trigger apoptosis during unresolved endoplasmic reticulum stress. eLife 9, e52291 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Berger A., et al. , Overexpression of transcription factor Foxa1 and target genes remediate therapeutic protein production bottlenecks in Chinese hamster ovary cells. Biotechnol. Bioeng. 117, 1101–1116 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Mohan C., Park S. H., Chung J. Y., Lee G. M., Effect of doxycycline-regulated protein disulfide isomerase expression on the specific productivity of recombinant CHO cells: Thrombopoietin and antibody. Biotechnol. Bioeng. 98, 611–615 (2007). [DOI] [PubMed] [Google Scholar]
  • 48.Hsu T. A., Betenbaugh M. J., Coexpression of molecular chaperone BiP improves immunoglobulin solubility and IgG secretion from Trichoplusia ni insect cells. Biotechnol. Prog. 13, 96–104 (1997). [DOI] [PubMed] [Google Scholar]
  • 49.Ku S. C. Y., Ng D. T. W., Yap M. G. S., Chao S.-H., Effects of overexpression of X-box binding protein 1 on recombinant protein production in Chinese hamster ovary and NS0 myeloma cells. Biotechnol. Bioeng. 99, 155–164 (2008). [DOI] [PubMed] [Google Scholar]
  • 50.Schoberer J., Shin Y.-J., Vavra U., Veit C., Strasser R., Analysis of protein glycosylation in the ER. Methods Mol. Biol. 1691, 205–222 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Lilley B. N., Ploegh H. L., A membrane protein required for dislocation of misfolded proteins from the ER. Nature 429, 834–840 (2004). [DOI] [PubMed] [Google Scholar]
  • 52.Oda Y., et al. , Derlin-2 and Derlin-3 are regulated by the mammalian unfolded protein response and are required for ER-associated degradation. J. Cell Biol. 172, 383–393 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Richelle A., Lewis N. E., Improvements in protein production in mammalian cells from targeted metabolic engineering. Curr. Opin. Syst. Biol. 6, 1–6 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Richelle A., et al. , Model-based assessment of mammalian cell metabolic functionalities using omics data. Cell Rep. Methods 1, 100040 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Masson H. O., et al. , ImmCellFie: A user-friendly web-based platform to infer metabolic function from omics data. STAR Protoc. 4, 102069 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Saghaleyni R., et al. , Enhanced metabolism and negative regulation of ER stress support higher erythropoietin production in HEK293 cells. Cell Rep. 39, 110936 (2022). [DOI] [PubMed] [Google Scholar]
  • 57.Kamionka M., Engineering of therapeutic proteins production in Escherichia coli. Curr. Pharm. Biotechnol. 12, 268–274 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Ali A. S., et al. , Multi-omics reveals impact of cysteine feed concentration and resulting redox imbalance on cellular energy metabolism and specific productivity in CHO cell bioprocessing. Biotechnol. J. 15, e1900565 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Lee J.-I., et al. , HepG2/C3A cells respond to cysteine deprivation by induction of the amino acid deprivation/integrated stress response pathway. Physiol. Genomics 33, 218–229 (2008). [DOI] [PubMed] [Google Scholar]
  • 60.Greenfield L., et al. , Metabolic engineering of CHO cells towards cysteine prototrophy and systems analysis of the ensuing phenotype. Metab. Eng. 84, 128–144 (2024). [DOI] [PubMed] [Google Scholar]
  • 61.Chevallier V., Andersen M. R., Malphettes L., Oxidative stress-alleviating strategies to improve recombinant protein production in CHO cells. Biotechnol. Bioeng. 117, 1172–1186 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Hediger M. A., et al. , The ABCs of solute carriers: Physiological, pathological and therapeutic implications of human membrane transport proteinsIntroduction. Pflugers Arch. 447, 465–468 (2004). [DOI] [PubMed] [Google Scholar]
  • 63.Jyotsana N., Ta K. T., DelGiorno K. E., The role of cystine/glutamate antiporter SLC7A11/xCT in the pathophysiology of cancer. Front. Oncol. 12, 858462 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Eggler A. L., Small E., Hannink M., Mesecar A. D., Cul3-mediated Nrf2 ubiquitination and antioxidant response element (ARE) activation are dependent on the partial molar volume at position 151 of Keap1. Biochem. J. 422, 171–180 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Eggler A. L., Liu G., Pezzuto J. M., van Breemen R. B., Mesecar A. D., Modifying specific cysteines of the electrophile-sensing human Keap1 protein is insufficient to disrupt binding to the Nrf2 domain Neh2. Proc. Natl. Acad. Sci. U.S.A. 102, 10070–10075 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Furukawa M., Xiong Y., BTB protein Keap1 targets antioxidant transcription factor Nrf2 for ubiquitination by the Cullin 3-Roc1 ligase. Mol. Cell. Biol. 25, 162–171 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Baird L., Yamamoto M., The molecular mechanisms regulating the KEAP1-NRF2 pathway. Mol. Cell. Biol. 40, e00099-20 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Knatko E. V., et al. , Downregulation of Keap1 confers features of a fasted metabolic state. iScience 23, 101638 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Xu J., Donepudi A. C., Moscovitz J. E., Slitt A. L., Keap1-knockdown decreases fasting-induced fatty liver via altered lipid metabolism and decreased fatty acid mobilization from adipose tissue. PLoS ONE 8, e79841 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Ludtmann M. H. R., Angelova P. R., Zhang Y., Abramov A. Y., Dinkova-Kostova A. T., Nrf2 affects the efficiency of mitochondrial fatty acid oxidation. Biochem. J. 457, 415–424 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Song M.-Y., Lee D.-Y., Chun K.-S., Kim E.-H., The role of NRF2/KEAP1 signaling pathway in cancer metabolism. Int. J. Mol. Sci. 22, 4376 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Sasaki H., et al. , Electrophile response element-mediated induction of the cystine/glutamate exchange transporter gene expression. J. Biol. Chem. 277, 44765–44771 (2002). [DOI] [PubMed] [Google Scholar]
  • 73.Panieri E., Telkoparan-Akillilar P., Suzen S., Saso L., The NRF2/KEAP1 axis in the regulation of tumor metabolism: Mechanisms and therapeutic perspectives. Biomolecules 10, 791 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.De Sancho D., Doshi U., Muñoz V., Protein folding rates and stability: How much is there beyond size? J. Am. Chem. Soc. 131, 2074–2075 (2009). [DOI] [PubMed] [Google Scholar]
  • 75.Ying J., Clavreul N., Sethuraman M., Adachi T., Cohen R. A., Thiol oxidation in signaling and response to stress: Detection and quantification of physiological and pathophysiological thiol modifications. Free Radic. Biol. Med. 43, 1099 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Betz S. F., Disulfide bonds and the stability of globular proteins. Protein Sci. 2, 1551–1558 (1993). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Fath S., et al. , Multiparameter RNA and codon optimization: A standardized tool to assess and enhance autologous mammalian gene expression. PLoS ONE 6, e17596 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Chung B.K.-S., Yusufi F. N. K., Mariati Y., Yang D.-Y. L., Enhanced expression of codon optimized interferon gamma in CHO cells. J. Biotechnol. 167, 326–333 (2013). [DOI] [PubMed] [Google Scholar]
  • 79.You M., et al. , Efficient mAb production in CHO cells with optimized signal peptide, codon, and UTR. Appl. Microbiol. Biotechnol. 102, 5953–5964 (2018). [DOI] [PubMed] [Google Scholar]
  • 80.Mauro V. P., Codon optimization in the production of recombinant biotherapeutics: Potential risks and considerations. BioDrugs 32, 69–81 (2018). [DOI] [PubMed] [Google Scholar]
  • 81.Stefani M., Protein misfolding and aggregation: New examples in medicine and biology of the dark side of the protein world. Biochim. Biophys. Acta 1739, 5–25 (2004). [DOI] [PubMed] [Google Scholar]
  • 82.Li Z.-M., Fan Z.-L., Wang X.-Y., Wang T.-Y., Factors affecting the expression of recombinant protein and improvement strategies in Chinese hamster ovary cells. Front. Bioeng. Biotechnol. 10, 880155 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Pace C. N., Scholtz J. M., A helix propensity scale based on experimental studies of peptides and proteins. Biophys. J. 75, 422–427 (1998). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Harreither E., et al. , Microarray profiling of preselected CHO host cell subclones identifies gene expression patterns associated with increased production capacity. Biotechnol. J. 10, 1625–1638 (2015). [DOI] [PubMed] [Google Scholar]
  • 85.Mathias S., et al. , Unraveling what makes a monoclonal antibody difficult-to-express: From intracellular accumulation to incomplete folding and degradation via ERAD. Biotechnol. Bioeng. 117, 5–16 (2020). [DOI] [PubMed] [Google Scholar]
  • 86.Budge J. D., et al. , Engineering of Chinese hamster ovary cell lipid metabolism results in an expanded ER and enhanced recombinant biotherapeutic protein production. Metab. Eng. 57, 203–216 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Zhang Y., et al. , High-throughput lipidomic and transcriptomic analysis to compare SP2/0, CHO, and HEK-293 mammalian cell lines. Anal. Chem. 89, 1477–1485 (2017). [DOI] [PubMed] [Google Scholar]
  • 88.Budge J. D., et al. , Data for engineering lipid metabolism of Chinese hamster ovary (CHO) cells for enhanced recombinant protein production. Data Brief 29, 105217 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Marichal-Gallardo P. A., Alvarez M. M., State-of-the-art in downstream processing of monoclonal antibodies: Process trends in design and validation. Biotechnol. Prog. 28, 899–916 (2012). [DOI] [PubMed] [Google Scholar]
  • 90.Dewanjee S., et al. , Nrf2/Keap1/ARE regulation by plant secondary metabolites: A new horizon in brain tumor management. Cell Commun. Signal. 22, 1–31 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Wu M. Y. M., et al. , Improving recombinant antibody production using FcBAR: An in situ approach to detect and amplify protein-protein interactions. Metab. Eng. 92, 174–184 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Wu M. Y. M., et al. , Enhanced production of HCV E1E2 subunit vaccine candidates via protein-protein interaction identification in glycoengineered CHO cells. Biotechnol. J. 20, e70112 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Zhang J., et al. , Using endoplasmic reticulum engineering to improve recombinant protein production in CHO cells. Int. J. Biol. Macromol. 315, 144695 (2025). [DOI] [PubMed] [Google Scholar]
  • 94.Masson H. O., Karottki K. J. la C., Tat J., Hefzi H., Lewis N. E., From observational to actionable: Rethinking omics in biologics production. Trends Biotechnol. 41, 1127–1138 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Masson H. O., Lewis Lab of Systems Biology & Cell Engineering (UGA). CHO96: Code for “Deciphering the determinants of recombinant protein yield across the human secretome.” GitHub. https://github.com/LewisLabUCSD/CHO96. Deposited 25 May 2023.
  • 96.Lewis N. E., Rockberg J., Deciphering the determinants of recombinant protein yield across the human secretome. Gene Expression Omnibus. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE225989. Deposited 23 February 2023.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Dataset S01 (XLSX)

pnas.2506036122.sd01.xlsx (13.7KB, xlsx)

Dataset S02 (XLSX)

Dataset S03 (XLSX)

pnas.2506036122.sd03.xlsx (28.1KB, xlsx)

Dataset S04 (XLSX)

Dataset S05 (XLSX)

pnas.2506036122.sd05.xlsx (806.9KB, xlsx)

Dataset S06 (XLSX)

pnas.2506036122.sd06.xlsx (479.4KB, xlsx)

Dataset S07 (XLSX)

pnas.2506036122.sd07.xlsx (22.3MB, xlsx)

Dataset S08 (XLSX)

pnas.2506036122.sd08.xlsx (34.5KB, xlsx)

Data Availability Statement

All data and code supporting the findings and analyses of this study are publicly available in the following GitHub repository (https://github.com/LewisLabUCSD/CHO96) (95) and GEO (GSE225989) (96). All other data are included in the manuscript and/or supporting information.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES