Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2019 Dec 1.
Published in final edited form as: AIChE J. 2018 Aug 9;64(12):4247–4254. doi: 10.1002/aic.16378

Bioinformatic analysis of Chinese hamster ovary host cell protein lipases

Madolyn L MacDonald 1,2, Nathaniel Hamaker 1,3, Kelvin H Lee 1,3
PMCID: PMC6430115  NIHMSID: NIHMS986122  PMID: 30911190

Abstract

Complete, accurate genome assemblies are necessary to design targets for genetic engineering strategies. Successful gene knockdowns and knockouts in Chinese hamster ovary (CHO) cells may prevent the expression of difficult-to-remove host cell proteins (HCPs). HCPs, if not removed, can cause problems in stability, safety, and efficacy of the biotherapeutic. A significantly improved Chinese hamster (CH) reference genome was used to identify new knockout targets with similar predicted functions and characteristics as the difficult-to-remove host cell lipases, LPL, PLBL2, and LPLA2. The CHO-K1 gene and protein sequences of several of these lipases were corrected using the updated CH genome. Sequence alignments were then used to identify conserved regions that may serve as possible targets for multiple simultaneous gene knockouts. Finally, comparison of the CHO-K1 lipase protein sequences to their human orthologs provided insight into which lipases, if persistent in the drug product, could possibly cause immunogenic responses in patients.

Keywords: lipase, polysorbate, host cell protein, Chinese hamster ovary (CHO), genome

Introduction

Chinese hamster ovary (CHO) cells are the preferred platform for biotherapeutic protein production. Monoclonal antibodies (mAbs) alone are predicted to reach global sales of 125 billion USD in 20201 and are used to treat many oncological, immunological and cardiovascular diseases. During the production of therapeutic proteins by CHO cells, host cell proteins (HCPs) are also secreted by the cells. Certain HCPs, if not removed during subsequent purification processes, have been shown to cause immunogenic responses in patients2 and others can shorten the shelf life of the final drug product through a variety of mechanisms including polysorbate degradation.3,4,5,6 HCPs, therefore, need be reduced to minimal levels, typically 1–100 ppm, in final mAb formulations.7 While most HCPs are removed from the therapeutic product during downstream purification steps, certain ‘difficult-to-remove’ HCPs can remain.8 Several types of lipases have been identified as problematic HCPs, especially regarding the stability of the mAb product.

Lipoprotein lipase (LPL) has been identified as a particularly difficult-to-remove impurity in CHO cell mAb production that possesses polysorbate 20 (PS-20) and polysorbate 80 (PS-80) degradation activity.3,9 PS20 and PS80 are surfactants often added to the drug product as protection from degradation during storage.10,11 It has been hypothesized that LPL is able to degrade PS20 and PS80 because polysorbates share structural similarities to triglycerides, the natural substrate of LPL. In particular, they share an ester bond which LPL hydrolyzes within triglycerides to form fatty acids and alcohol molecules.12 Part of the reason LPL may be especially difficult to remove in a variety of processes producing a variety of products is that LPL has been shown to associate with multiple mAbs in protein A affinity chromatography and also to co-elute in non-affinity polishing columns used in subsequent steps of protein purification.8,13

Two other lipases have also displayed polysorbate degrading activity and have been identified in CHO cell-derived drug products. Group XV lysosomal phospholipase A2 (LPLA2 or PLA2G15) was found in the drug product of several mAb-producing cell lines at less than 1 ppm. Even at these low levels, LPLA2 was associated with the hydrolysis of PS20 and PS80.14 The rate of polysorbate hydrolysis was shown to be both time and concentration dependent.

Putative phospholipase B-like 2 (PLBL2 or PLBD2) is another difficult-to-remove HCP that has been shown to co-elute with several biotherapeutic antibodies during the protein A chromatography purification process.15 PLBL2 has been associated with the degradation of PS-20 in a sulfatase drug product.4 In addition, drug material used in Lebrikizumab clinical trials was found to contain 34–328 ng of CHO PLBL2 per mg of product, and approximately 90% of patients in the clinical trial developed an immune response against PLBL2.2 PLBL2 also displayed variable expression during an extended culture of 136 days,9 a characteristic of difficult-to-remove HCPs because purification processes may not adequately remove the wide range of expression levels reached over time.

Genome editing techniques have been used to knockout a variety of different genes in CHO cell lines. For instance, clustered regularly interspaced short palindromic repeats (CRISPR)/CRISPR-associated protein 9 (Cas9) has been used to knockout methyltransferase genes in CHO cells to stabilize therapeutic protein productivity16 and a fucosyltransferase gene to prevent the fucosylation of the target biotherapeutic.17 CRISPR/Cas9 has also been used to knockout a difficult-to-remove HCP impurity. A successful knockout of Lpl using CRISPR/Cas9 was shown to decrease PS80 degradation by 41–47% percent and PS20 degradation by 44–57%.3 Other genome editing techniques such as transcription activator-like effector nucleases (TALENs) and zinc-finger nucleases (ZFNs) have also been used to effectively knockout genes in CHO cells.18,19,20

While unknown, it is possible that other lipases with similar enzymatic activity to LPL, LPLA2, and PLBL2 could result in polysorbate degradation and/or immunogenic responses if they exist in the final drug product. Here, we identified potentially problematic lipases based on an analysis of the CHO-K1 and Chinese hamster (CH) genomes, and protein sequence similarity to LPL, LPLA2, and PLBL2. Several misassemblies and/or misannotations in the sequences of CHO-K1 lipases were identified and corrected using the most recent CH genome,21 highlighting the importance of accurate and complete reference genomes. The corrected sequences were then examined to identify conserved regions that could be targeted to knockout multiple lipases simultaneously. We also compared the newly corrected CH/CHO-K1 lipase protein sequences to their human orthologs to understand the extent to which any of the lipases may be immunogenic in humans.

Materials and Methods

Protein and CDS alignments of LPL, PLBL2, and LPLA2 from various CH and CHO assemblies

Protein and mRNA sequences of LPL, PLBL2, and LPLA2 were extracted from CHO-K1 Refseq,22 CH RefSeq,23 and the updated PICR CH assembly,21 to compare sequence differences between CHO-K1 and CH, and to examine changes among different CH assembly versions. Protein alignment was done using MUSCLE24 and mRNA alignment was done using ClustalO25 using the default parameters. For mRNA alignments, only the coding sequence (CDS) regions from each transcript were used because untranslated regions (UTRs) are difficult to annotate correctly.26,27 An error in the CHO-K1 LPL protein sequence was identified and corrected using the MUSCLE alignment to the CH PICR, mouse, rat, and human orthologs. The corrected CHO-K1 LPL sequence and the original CHO-K1 sequences for PLBL2 and LPLA2 were used in further analyses. The RefSeq IDs for the CHO-K1 transcripts and proteins used in this project are listed in Supplementary Table SI.

Identification of lipases similar to LPL, PLBL2, and LPLA2

An extensive list of lipase enzymes was compiled from searching EMBL-EBI’s QuickGO database28 with the GO term, lipase activity (GO:0016298). Corresponding protein sequences for the identified lipases were extracted from the PICR and CHO-K1 assemblies. BLASTP29 was used to query LPL, PLBL2, and LPLA2 against this list to identify the most similar proteins with lipase activity. Hits with an E-value < 0.001 were further examined by full sequence alignment with the query (LPL, PLBL2, or LPLA2). For proteins with more than one hit with an E-value < 0.001, a phylogenetic tree of the protein sequences was created using the neighbor-joining algorithm within JalView30 with PAM250 as the position specific matrix.

Correction of lipase protein and gene sequences

Errors in the CHO-K1 protein sequences were detected by aligning each protein sequence with their orthologs in human, mouse, rat, and CH PICR. The human, mouse, and rat sequences were extracted from UniProt31 release 2018_1. Once an error was identified, the type and location of the error was characterized by examining the transcript alignment against mouse and PICR using SnapGene (www.snapgene.com). Most errors involved a missing or incomplete exon at the 5’ end of the gene. In these cases, the ‘correct’ exon from mouse that corresponded with the erroneous exon in CHO-K1 was realigned to the CHO-K1 gene to correct the CHO-K1 gene annotation. Realignment of the newly modified CHO-K1 protein sequence against human, mouse, rat, and CH PICR orthologs was done to validate the correction.

The correction for CHO-K1 PNLIPRP2 was more complicated and benefited significantly from the updated PICR genome. Alignment to PNLIPRP2 mouse, rat, and human protein orthologs showed that the protein sequence for CHO-K1 PNLIPRP2 was missing a segment of amino acids at the 5’ end. Visualization of Pnliprp2 on its scaffold, NW_003617188.1, showed that the gene was incomplete because it was located directly on the end of the scaffold. PICR was used to find the neighboring gene, Pnliprp1 (XM_007655147.2), which was located on CHO-K1 scaffold, NW_003617412.1. The two scaffolds were realigned to the longer PICR scaffold, picr_24, to confirm that these two scaffolds should be merged in the CHO-K1 genome. The mouse transcript (NM_011128.2) was then aligned to determine that the entire first exon of the gene was located in the NW_003617412.1 scaffold. Finally, the sequence in CHO-K1 that aligned to the mouse exon was used to correct the gene and protein sequence for PNLIPRP2. Exact boundaries were identified using the mouse and golden hamster (XM_005085339.3) exons.

Determination of expression levels of lipases of interest

Each gene was checked for expression in CHO-K1 cells using data from GEO: GSE75094.32 The expression levels were visualized on the ‘CHO-K1 mRNA expression data’ browser on CHOgenome.org33 using the FPKM (Fragments Per Kilobase of transcript per Million mapped reads) and SAM (Sequence Alignment/Map) coverage tracks. Genes that appeared with any amount of expression in CHO cells were examined further.

Determination of conserved regions in each grouping of lipases

Conserved regions among each of the three groups of BLAST hits (one group per LPL, LPLA2, and PLBL2) were located from the protein alignments. The DNA sequences underlying the conserved regions were examined as well and used as the query in BLASTn to search against the non-redundant sequence set for CH species ID (10029) to identify possible off-target effects of using the conserved regions as knockout targets. The PICR assembly was also queried to ensure there were no additional hits when assembly improvements were considered.

Examining the immunogenicity potential of similar lipases

Human protein sequences for LPL, LPLA2, and PLBL2 were extracted from UniProt release 2018_1 and aligned against the corresponding CHO-K1 and PICR protein sequences to calculate percent identity. The percent identity of PLBL2, which is known to be immunogenic, was then used to select a threshold of 80% identity to assess the possible immunogenicity of the BLAST hits of LPL, LPLA2, and PLBL2. The percent identity with their human orthologs were determined, and if the percent identity was lower than that of PLBL2 and its human ortholog, it was flagged as having the potential to cause immunogenic responses.

Results

Sequence differences among the CH and CHO-K1 Assemblies

Alignments of the protein sequences across the different CH and CHO-K1 assemblies for LPLA2 and PLBL2 showed very little difference. However, the protein sequence for LPL isoform X1 from CH RefSeq has 11 additional amino acids at the 3’ end, but the X2 isoform is the same length as the PICR protein (Figure 1). The protein sequence of LPL from CHO-K1 is missing one amino acid as shown by the gap in the alignment at position 24, which is then followed by an unknown amino acid at position 25 (Figure 1). These errors are also reflected in the mRNA coding sequence alignments (Supplementary Figure S1), which show that the CHO-K1 LPL sequence is missing four guanines. The existence of these nucleotides in CHO-K1 LPL were confirmed by Sanger sequencing. If an sgRNA for a CRISPR/Cas9 knockout was designed to target this region of Lpl based on the CHO-K1 genome sequence alone, it would be missing four nucleotides. This would greatly decrease the binding affinity of the sgRNA to this region in the gene and thus, the knockout efficiency.

Figure 1.

Figure 1.

Alignment of LPL protein sequence from CHO-K1 RefSeq, CH RefSeq (isoforms X1 and X2) and the updated CH genome, PICR. Positions 24–25 are in red to highlight the error in the CHO-K1 LPL sequence. The difference between CH RefSeq LPL isoform X1 and X2 is shown in purple.

There were no differences in the protein sequences for LPLA2 among the different assemblies (Supplementary Figure S2). Three unknown amino acids exist in the PLBL2 CH RefSeq sequence at base positions 43–46, but alignment to the CHO-K1 and CH PICR sequences suggest that these do not actually exist (Supplementary Figure S3).

Identification and sequence correction of HCPs related to LPL

BLASTP hits with significant alignment (E-value < 0.001) to LPL from CH/CHO-K1 included LIPC, LIPG, LIPH, LIPI, PLA1A, PNLIP, PNLIPRP1, and PNLIPRP2. All of these lipases belong to the pancreatic lipase gene family,34 which is composed of members with triglyceride lipase activity (EC 3.1.1.3) and the closely related lipoprotein lipase (EC 3.1.1.34).35 LPL is most related to the LIPG (endothelial lipase) and LIPC (hepatic lipase) proteins, and then to the PNLIP proteins (pancreatic lipases) (Figure 2). The similarity of these eight proteins to LPL at the sequence level suggests that they could potentially degrade PS20 and PS80.

Figure 2.

Figure 2.

Phylogenetic tree derived from the multiple sequence alignment of LPL to its significant BLASTP hits using the neighbor joining algorithm (PAM250) in JalView. Distances of each branch are labeled.

Five of these genes (Lipi, Liph, Pla1a, Pnliprp2, and Pnliprp1) had evidence that supported their expression in CHO-K1 cells from the CHOgenome.org browser. Pnliprp1 and LipH have also previously been identified as differentially expressed in sodium butyrate treated CHO cells when compared to non-treated cells.36 In addition, a higher than 1.5 fold change in expression of Pnliprp1 was observed between a low-producing and a high-producing cell line.37 Three of the five genes (Pnliprp2, Pnliprp1, and Lipi) had errors in their annotations for the CHO-K1 genome, which could be seen in the multiple sequence alignment against the corresponding protein in mouse, rat, human, and CH PICR. All three genes had a missing exon at the 5’ end. Lipi was missing the first exon, most likely because the 5’ end of the gene overlapped with another gene, Rbm11, located on the antisense DNA strand, which may have complicated the annotation. The CHO-K1 Pnliprp2 did not contain the first exon because the gene was split over two scaffolds in the CHO-K1 genome. The longer scaffold length in the PICR assembly allowed the two CHO-K1 scaffolds to be merged and the gene to be resolved (Figure 3). Not only was PICR able to correct the Pnliprp2 gene, but the longer scaffold length in PICR enabled the PNLIP family of lipases to be ordered within the CHO-K1 genome. This section of genes was originally split over three scaffolds (Figure 4). It is unclear why Pnliprp1 had a missing exon in its annotation. Sanger sequencing confirmed the joining of the scaffolds.

Figure 3.

Figure 3.

View of Pnliprp2 gene split over two scaffolds, NW_003617188 and NW_003617412, in the CHO-K1 assembly. This is confirmed with alignment of the mouse Pnliprp2 gene (NM_011128.2) to the scaffolds. Alignment to the first exon in golden hamster Pnliprp2 (XM_005085339.3) helped to determine the exact boundaries of the exon in CHO-K1.

Figure 4.

Figure 4.

View of the Pnlip lipase genes, positioned and ordered in a single superscaffold in CHO-K1. Part of the PICR scaffold, picr_24 (2,501,648 – 2,752,160), is aligned above in red, showing the overlap across the three CHO-K1 scaffolds. The LOC100762115 gene is an uncharacterized relative of the Pnlip gene.

Once the sequence errors were resolved, the five genes similar to Lpl and expressed in CHO-K1 cells were aligned (Figure 5). The alignment shows that the six proteins all share the same active site residues which make up the well-known catalytic triad.38 Regions around the first two active site residues are well conserved, particularly the ‘RITGLDP’ peptide (highlighted in Figure 5). This peptide could provide a target location to simultaneously knockdown, knockout, or purify the potentially troublesome HCPs. The underlying DNA sequence, however, is not well conserved (Supplementary Figure 4) and therefore, multiple targets will need to be designed and tested for their off-target effects for knockdown and knockout studies.

Figure 5.

Figure 5.

Protein sequence alignment of LPL, positions 1 to 318, to the five similar lipases expressed in CHO-K1 cells. Active sites are highlighted in green (LPL positions: 159 S, 183 D, and 268 H). The potential target conserved peptide, ‘RITGLDP’, is highlighted within the red box.

Identification of HCPs related to PLBL2

PLBL2 only had a single significant BLASTp hit which was PLBL1. They share 36.33% identity and have 51.76% positive scoring amino acid replacements. Alignment against mouse, rat, and human suggested that there were no errors in either CDS or protein sequence. Alignment of PLBL2 and PLBL1 show that they share the same set of active site residues where five of six align exactly (Figure 6). The active sites were identified from the annotation of human PLBL2 and PLBL1 proteins (http://genomewiki.ucsc.edu/index.php/Phospholipases_PLBD1_and_PLBD2). However, the exact function or substrates of PLBL2 and PLBL1 are unknown.

Figure 6.

Figure 6.

Protein sequence alignment of PLBL2 and PLBL1. Active sites (PLBL2 positions: 240 C, 257 H, 260 W, 301 T, 423 N, 454 R) are highlighted in green. The four conserved peptides ‘FSSYPG’, ‘DDFYIL’, ‘NSGTYNNQ’, and ‘SYNIPF’ are highlighted within the red boxes.

Four conserved potential target regions between PLBL2 and PLBL1 are the peptides ‘FSSYPG’, ‘DDFYIL’, ‘NSGTYNNQ’, and ‘SYNIPF’ (Figure 6). The ‘DDFYIL’ is the most conserved on the DNA level (Supplementary Figure 5a) and can be extended to contain the ‘NGG’ PAM spacer (Supplementary Figure 5b). This PAM spacer is necessary in the single guide RNA (sgRNA) target site for the most commonly used type of CRISPR/Cas9, Streptococcus pyogenes. The extended target site, N’-GATGACTTCTACATCCTNNGCAG-C’, also appears to have no off-target hits with less than seven base mismatches when querying the CHO/CH RefSeq/PICR genomes. However, it should be noted that no evidence to date has been found that suggests the Plbl1 gene is expressed in CHO-K1 cells.

Identification of HCPs related to LPLA2

The CHO-K1 LPLA2 protein shared significant similarity with the CHO-K1 LCAT (Lecithin-Cholesterol Acyltransferase) protein. This similarity has been described previously: LPLA2 and LCAT are closely related acyltransferases39 and are members of the αβ-hydrolase family40,41. LPLA2 transfers fatty acids from glycerophospholipids to lipophilic alcohols, while LCAT transfers fatty acids from glycerophospholipids to cholesterol39. The alignment between LPLA2 and LCAT reflects this functional similarity as they share 48.8% protein sequence similarity (68.6% positive scoring amino acid replacements) and the same active site residues making up the catalytic triad (Figure 7). LCAT appears to be expressed in CHO-K1 cells on the CHOgenome.org RNA browser and has been previously described as differentially expressed between sodium butyrate treated and non-treated CHO cells.36 The peptide ‘LEAKLDKP’ shared between LPLA2 and LCAT could be used as a gene editing target to knockout the expression of both (highlighted in Figure 7). This peptide is also well conserved at the DNA level (Supplementary Figure 6) and no off-target effects were found in the CHO-K1, CH, and PICR genomes, using the sequence N’-CTNGAAGCNAAGCTGGANAAACCA-C’ as the BLASTn query. Another potential target is the ‘FISLGAPWGG’ peptide, but this is not as well conserved at the DNA level.

Figure 7.

Figure 7.

Protein sequence alignment of LPLA2 and LCAT. Active sites (LPLA2 positions: 198 S, 360 D, 392 H) are highlighted in green. The conserved peptides ‘LEAKLDKP’ and ‘FISLGAPWGG’ are highlighted within the red boxes.

CHO-K1 lipase similarity to their human orthologs

It has been hypothesized that the more dissimilar the protein sequences of CH are to their human orthologs, the more likely the CH protein can cause an immunogenic response in a patient.42,43 LPL and LPLA2 have an identity of 93.47% and 88.59%, respectively, with their human orthologs. There has been no evidence of immunogenic responses against either LPL or LPLA2. PLBL2 is more different from its human ortholog with an identity of 78.95% and it has been shown to cause immunogenic responses in patients.2 Using 80% identity as the threshold cutoff, we compared the sequences of the other related lipases identified here to their human orthologs. Three proteins did not meet the threshold: LIPI (62.18% identity), PNLIPRP2 (76.97% identity), and PLBL1 (76.35% identity). These differences suggest that these proteins may be more likely to cause immunogenic responses than the other lipases if not removed during the purification process. PLA1A and LIPH isoform X2 were just above the cutoff of at 80.92% and 81.92% identity, respectively. LCAT (89.32%) and PNLIPRP1 (86.30%) were above the threshold.

Discussion

Sequence errors are a reality for all draft genome assemblies based on existing technology. De novo assemblies of human genomes built from short sequencing alone have been shown to be missing millions of bases of duplicated sequence and common repeats, and missing thousands of coding exons.44 Many of these errors are corrected in later rounds of resequencing which reach an adequate coverage depth, producing more finished assemblies.45 Reference genomes need to be improved beyond the draft status to avoid making incorrect inferences in reference-guided studies. For instance, accurate and complete assemblies and annotations are key to perform effective genetic engineering techniques.

Here, we show the advantage of having a significantly higher quality reference genome for CHO cell lines. The new PICR assembly enabled the identification and the correction of errors in the sequence of the difficult-to-remove host cell protein, LPL, and the similar PNLIPRP2 protein. The PICR assembly, along with the accurate mouse genome, enabled us to correct two other proteins similar to LPL, PNLIPRP1 and LIPI. The sequences underlying these annotation errors have been corrected in the PICR sequence, indicating that a new NCBI RefSeq annotation for PICR will have the correct sequences and coordinate boundaries for the LPL, PNLIPRP2, PNLIPRP1, and LIPI genes/proteins.

Correct gene and protein sequences allowed us to identify significantly similar lipases to three known difficult-to-remove HCPs: LPL, PLBL2, and LPLA2. Our findings are summarized in Table 1. Functional and sequence similarity suggest that these related lipases have the potential to cause similar issues if present in the final drug product. Within each grouping of lipases, conserved regions were identified that could serve as targets for the mitigation of the negative impacts of these lipases. An sgRNA could be designed to target the DNA that codes for the ‘LEAKLDKP’ conserved peptide in LPLA2 and LCAT, knocking out both genes simultaneously. Even if the DNA sequences underlying the conserved peptides are not identical among the genes of interest, multiple different guide RNAs can be applied to target the same location. Multiplexing of CRISPR/Cas9 has previously been successful in CHO cells to knockout three genes simultaneously.17,46 It has also been able to target up to 62 retroviral elements in porcine kidney cells47 and primary purified porcine cells.48 Targeting the same region would provide consistency in the knockouts, removing any positional impacts of where the inserted or deleted nucleotide(s) occur. For instance, targeting the LPL group of lipases at the ‘RITGLDP’ peptide identified here would knock out all lipases directly near their second active site residue.

Table 1.

Summary of Findings for LPL, PLBL2, LPLA2 including the Related Lipases (ones known to be Expressed in CHO-K1 Cells are Highlighted in Bold), Conserved Peptides, and whether Comparison to the Corresponding Human Ortholog suggests the Similar Lipases could cause an Immunogenic Response in Patients

Known HCP Similar Lipases Conserved Peptides
(N’ to C’)
Protein sequence identity to
human ortholog suggests
immunogenicity?
LPL LIPC, LIPG, LIPH, LIPI,
PLA1A, PNLIP,
PNLIPRP1, PNLIPRP2
‘RITGLDP’ Yes for LIPI, PNLIPRP2
PLBL2 PLBL1 ‘DDFYIL’,
‘FSSYPG’,
‘NSGTYNNQ’,
‘SYNIPF’
Yes
LPLA2 LCAT ‘LEAKLDKP’,
‘FISLGAPWGG’
No

While this approach will prevent the catalytic activity of the targeted lipases, it is important to note frameshift mutations from knocking out the lipases can still result in expression of some peptides of the target protein. In theory, these fragments could still have the ability to bind to the product and/or cause immunogenic effects if not purified from the final drug product. A full gene deletion approach, as described in Q. Zheng et al,49 could mitigate this concern. This approach consists of targeting two sequences on either side of the gene of interest, which when cut and repaired can remove the entire sequence between the targets. Multiple, simultaneous gene deletions using a combination of CRISPR/Cas9 and CRISPR/Cpf1 systems have been carried out in CHO cells with no deletion size limitations within 2–150 kb50. Thus, it is feasible that sets of several problematic lipase genes could be fully deleted using this multiplex approach.

In addition, alignment of protein sequences to human orthologs can provide insights into the possible immunogenicity of an HCP. Here, we discovered that LIPI, PNLIPRP2, and PLBL1 (if expressed) are the most different from their human orthologs and may have the most potential to be immunogenic if not removed during purification processes. This work shows the immense value of having access to the highest quality reference genomes and annotations when carrying out genetic engineering studies.

Supplementary Material

Supp info

Acknowledgments

We are grateful for financial support from the NSF IGERT SBE2 grant 1144726 and NSF grants 1412365 and 1736123. The authors would also like to acknowledge that use of the Biomix computing cluster was made possible through funding from Delaware INBRE (NIH GM103446), the State of Delaware, and the Delaware Biotechnology Institute.

Footnotes

Topical heading:

Biomolecular Engineering, Bioengineering, Biochemicals, Biofuels, and Food

Literature Cited

  • 1.Ecker DM., Jones SD, Levine HL. The therapeutic monoclonal antibody market. mAbs 2015;7(1):9–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Fischer SK, Cheu M, Peng K, Lowe J, Araujo J, Murray E, McClintock D, Matthews J, Siguenza P, Song A. Specific Immune Response to Phospholipase B-Like 2 Protein, a Host Cell Impurity in Lebrikizumab Clinical Material. The AAPS Journal 2017;19(1):254–263. [DOI] [PubMed] [Google Scholar]
  • 3.Chiu J, Valente KN, Levy NE, Min L, Lenhoff AM, Lee KH. Knockout of a difficult-to-remove CHO host cell protein, lipoprotein lipase, for improved polysorbate stability in monoclonal antibody formulations. Biotechnology and Bioengineering 2017;114(5):1006–1015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Dixit N, Salamat-Miller N, Salinas PA, Taylor KD, Basu SK. Residual Host Cell Protein Promotes Polysorbate 20 Degradation in a Sulfatase Drug Product Leading to Free Fatty Acid Particles. Journal of Pharmaceutical Sciences 2016;105(5):1657–1666. [DOI] [PubMed] [Google Scholar]
  • 5.Dorai H, Santiago A, Campbell M, Tang QM, Lewis MJ, Wang Y, Lu QZ, Wu SL, Hancock W. Characterization of the proteases involved in the N-terminal clipping of glucagon-like-peptide-1-antibody fusion proteins. Biotechnology Progress 2011;27(1):220–231. [DOI] [PubMed] [Google Scholar]
  • 6.Gao SX, Zhang Y, Stansberry-Perkins K, Buko A, Bai S, Nguyen V, Brader ML. Fragmentation of a highly purified monoclonal antibody attributed to residual CHO cell protease activity. Biotechnology and Bioengineering 2011;108(4):977–982. [DOI] [PubMed] [Google Scholar]
  • 7.Eaton LC. Host cell contaminant protein assay development for recombinant biopharmaceuticals. Journal of Chromatography A 1995;705(1):105–114. [DOI] [PubMed] [Google Scholar]
  • 8.Levy NE, Valente KN, Choe LH, Lee KH, Lenhoff AM. Identification and Characterization of Host Cell Protein Product-Associated Impurities in Monoclonal Antibody Bioprocessing. Biotechnology Bioengineering 2014;111:904–912. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Valente KN, Lenhoff AM, Lee KH. Expression of difficult-to-remove host cell protein impurities during extended Chinese hamster ovary cell culture and their impact on continuous bioprocessing. Biotechnology Bioengineering 2015;112(6):1232–1242. [DOI] [PubMed] [Google Scholar]
  • 10.Marichal-Gallardo PA, Álvarez MM. State-of-the-art in downstream processing of monoclonal antibodies: Process trends in design and validation. Biotechnology Progress 2012;28(4):899–916. [DOI] [PubMed] [Google Scholar]
  • 11.Kerwin BA. Polysorbates 20 and 80 Used in the Formulation of Protein Biotherapeutics: Structure and Degradation Pathways. Journal of Pharmaceutical Sciences 2008;97(8):2924–2935. [DOI] [PubMed] [Google Scholar]
  • 12.Mead J, Irvine S, Ramji D. Lipoprotein lipase: structure, function, regulation, and role in disease. Journal of Molecular Medicine 2002;80(12):753–769. [DOI] [PubMed] [Google Scholar]
  • 13.Levy NE, Valente KN, Lee KH, Lenhoff AM. Host cell protein impurities in chromatographic polishing steps for monoclonal antibody purification. Biotechnology Bioengineering, 2016;113(6):1260–1272. [DOI] [PubMed] [Google Scholar]
  • 14.Hall T, Sandefur SL, Frye CC, Tuley TL, Huang L. Polysorbates 20 and 80 Degradation by Group XV Lysosomal Phospholipase A2 Isomer X1 in Monoclonal Antibody Formulations. Journal of Pharmaceutical Sciences 2016;105:(5),1633–1642. [DOI] [PubMed] [Google Scholar]
  • 15.Tran B, Grosskopf V, Wang X, Yang J, Walker D, Yu C, McDonald P. Investigating interactions between phospholipase B-Like 2 and antibodies during Protein A chromatography. Journal of Chromatography A 2016;1438:31–38. [DOI] [PubMed] [Google Scholar]
  • 16.Aga M, Yamano N, Kumamoto T, Frank J, Onitsuka M, Omasa T. Construction of a gene knockout CHO cell line using a simple gene targeting method. BMC Proceedings 2015;9(Suppl 9):P2. [Google Scholar]
  • 17.Grav LM, Lee JS, Gerling S, Kallehauge TB, Hansen AH, Kol S, Lee GM, Pedersen LE, Kildegaard HF. One-step generation of triple knockout CHO cell lines using CRISPR/Cas9 and fluorescent enrichment. Biotechnology Journal 2015;10(9):1446–1456. [DOI] [PubMed] [Google Scholar]
  • 18.Chan KF, Shahreel W, Wan C, Teo G, Hayati N, Tay SJ, Tong WH, Yang Y, Rudd PM, Zhang P, Song Z. Inactivation of GDP-fucose transporter gene (Slc35c1) in CHO cells by ZFNs, TALENs and CRISPR-Cas9 for production of fucose-free antibodies. Biotechnology Journal 2016;11(3):399–414. [DOI] [PubMed] [Google Scholar]
  • 19.Santiago Y, Chan E, Liu PQ, Orlando S, Zhang L, Urnov FD, Holmes MC, Guschin D, Waite A, Miller JC, Rebar EJ, Gregory PD, Klug A, Collingwood TN. Targeted gene knockout in mammalian cells by using engineered zinc-finger nucleases. PNAS 2008;105(15):5809–5814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Yang Z, Wang S, Halim A, Schulz MA, Frodin M, Rahman SH, Clausen H. Engineered CHO cells for production of diverse, homogeneous glycoproteins. Nature Biotechnology 2015;33(8):842–844. [DOI] [PubMed] [Google Scholar]
  • 21.Rupp O, MacDonald ML, Li S, Dhiman H, Polson S, Griep S, Heffner K, Hernandez I, Brinkrolf K, Jadhav V, Samoudi M, Hao H, Kingham B, Goesmann A, Betenbaugh MJ, Lewis NE, Borth N, Lee KH. A reference genome of the Chinese hamster based on a hybrid assembly strategy. Biotechnology Bioengineering 2018;In Revision. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Xu X, Nagarajan H, Lewis NE, Pan S, Cai Z, Liu X, Chen W, Xie M, Wang W, Hammond S, Andersen MR, Neff N, Passarelli B, Koh W, Fan HC, Wang J, Gui Y, Lee KH, Betenbaugh MJ, Quake SR, Famili I, Palsson BO, Wang J. The genomic sequence of the Chinese hamster ovary (CHO)-K1 cell line. Nature Biotechnology 2011;29(8):735–741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Lewis NE, Liu X, Li Y, Nagarajan H, Yerganian G, O’Brien E, Bordbar A, Roth AM, Rosenbloom J, Bian C, Xie M, Chen W, Li N, Baycin-Hizal D, Latif H, Forster J, Betenbaugh MJ, Famili I, Xu X, Wang J, Palsson BO. Genomic landscapes of Chinese hamster ovary cell lines as revealed by the Cricetulus griseus draft genome. Nature Biotechnology 2013;31(8):759–767. [DOI] [PubMed] [Google Scholar]
  • 24.Edgar RC. MUSCLE: Multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research 2004;32(5):1792–1797. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Sievers F, Wilm A, Dineen D, Gibson TJ, Karplus K, Li W, Lopez R, McWilliam H, Remmert M, Söding J, Thompson JD, Higgins DG. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Molecular Systems Biology 2011;7(1):539. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Schurch NJ, Cole C, Sherstnev A, Song J, Duc C, Storey KG, McLean WH, Brown SJ, Simpson GG, Barton GJ. Improved Annotation of 3′ Untranslated Regions and Complex Loci by Combination of Strand-Specific Direct RNA Sequencing, RNA-Seq and ESTs. PLoS ONE 2014;9(4):e94270. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Shenker S, Miura P, Sanfilippo P, Lai EC. IsoSCM: improved and alternative 3’ UTR annotation using multiple change-point inference. RNA 2015;21(1):14–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Binns D, Dimmer E, Huntley R, Barrell D, O’Donovan C, Apweiler R. QuickGO: a web-based tool for Gene Ontology searching. Bioinformatics 2009;25(22):3045–3046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Altschul SF., Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. Journal of Molecular Biology 1990;215(3):403–410. [DOI] [PubMed] [Google Scholar]
  • 30.Waterhouse AM, Procter JB, Martin DM, Clamp M, Barton GJ. Jalview Version 2-A multiple sequence alignment editor and analysis workbench. Bioinformatics 2009;25(9):1189–1191. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Apweiler R, Bairoch A, Wu CH, Barker WC, Boeckmann B, Ferro S, Gasteiger E, Huang H, Lopez R, Magrane M, Martin MJ, Natale DA, O’Donovan C, Redaschi N, Yeh LL. UniProt : the Universal Protein knowledgebase. Nucleic Acids Research 2004;32:115–119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Lee N, Shin J, Park JH, Lee GM, Cho S, Cho BK. Targeted Gene Deletion Using DNA-Free RNA-Guided Cas9 Nuclease Accelerates Adaptation of CHO Cells to Suspension Culture. ACS Synthetic Biology/ 2016;5(11):1211–1219. [DOI] [PubMed] [Google Scholar]
  • 33.Kremkow BG, Baik JY, MacDonald ML, Lee KH. CHOgenome.org 2.0: Genome resources and website updates. Biotechnology Journal 2015;10(7):931–938. [DOI] [PubMed] [Google Scholar]
  • 34.Carrière F, Withers-Martinez C, van Tilbeurgh H, Roussel A, Cambillau C, Verger R. Structural basis for the substrate selectivity of pancreatic lipases and some related proteins. Biochimica et Biophysica Acta (BBA) - Reviews on Biomembranes 1998;1376(3):417–432. [DOI] [PubMed] [Google Scholar]
  • 35.Persson B, Bengtsson-Olivecrona G, Enerback S, Olivecrona T, Jornvall H. Structural features of lipoprotein lipase. Lipase family relationships, binding interactions, non-equivalence of lipase cofactors, vitellogenin similarities and functional subdivision of lipoprotein lipase. European Journal of Biochemistry 1989;179(1):39–45. [DOI] [PubMed] [Google Scholar]
  • 36.Birzele F, Schaub J, Rust W, Clemens C, Baum P, Kaufmann H, Weith A, Schulz TW, Hildebrandt T. Into the unknown: expression profiling without genome sequence information in CHO by next generation sequencing. Nucleic Acids Res 2010;38(12):3999–4010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Orellana CA, Marcellin E, Palfreyman RW, Munro TP, Gray PP, Nielsen LK. RNA-Seq Highlights High Clonal Variation in Monoclonal Antibody Producing CHO Cells. Biotechnol J 2018;13(3):1700231. [DOI] [PubMed] [Google Scholar]
  • 38.Emmerich J, Beg OU, Peterson J, Previato L, Brunzell JD, Brewer HB, Santamarina-Fojo S. Human lipoprotein lipase. Analysis of the catalytic triad by site-directed mutagenesis of Ser-132, Asp-156, and His-241. The Journal of Biological Chemistry 1992;267(6):4161–4165. [PubMed] [Google Scholar]
  • 39.Glukhova A, Hinkovska-Galcheva V, Kelly R, Abe A, Shayman JA, Tesmer JJ. Structure and function of lysosomal phospholipase A2 and lecithin:cholesterol acyltransferase. Nature Communications 2015;6:6250. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Hiraoka M, Abe A, Shayman JA. Structure and function of lysosomal phospholipase A2: identification of the catalytic triad and the role of cysteine residues. Journal of Lipid Research 2005;46(11):2441–2447. [DOI] [PubMed] [Google Scholar]
  • 41.Shayman JA, Kelly R, Kollmeyer J, He Y, Abe A. Group XV phospholipase A2, a lysosomal phospholipase A2. Progress in Lipid Research 2011;50(1):1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Bailey-Kellogg C, Gutiérrez AH, Moise L, Terry F, Martin WD, De Groot AS. CHOPPI: a web tool for the analysis of immunogenicity risk from host cell proteins in CHO-based protein production. Biotechnology and Bioengineering 2014;111(11):2170–2182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Gutiérrez AH, Moise L, De Groot AS. Of [Hamsters] and men: a new perspective on host cell proteins. Human Vaccines & Immunotherapeutics 2012;8(9):1172–1174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Alkan C, Sajjadian S, Eichler EE. Limitations of next-generation genome sequence assembly. Nature Methods 2011;8(1):61–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Meader S, Hillier LW, Locke D, Ponting CP, Lunter G. Genome assembly quality: assessment and improvement using the neutral indel model. Genome Research 2010;20(5):675–684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Shin J, Lee N, Song Y, Park J, Kang TJ, Kim SC, Lee GM, Cho BK. Efficient CRISPR/Cas9-mediated multiplex genome editing in CHO cells via high-level sgRNA-Cas9 complex. Biotechnology and Bioprocess Engineering 2015;20(5):825–833. [Google Scholar]
  • 47.Yang L, Güell M, Niu D, George H, Lesha E, Grishin D, Aach J, Shrock E, Xu W, Poci J, Cortazio R, Wilkinson R, Fishman J, Church G. Genome-wide inactivation of porcine endogenous retroviruses (PERVs). Science 2015;350(6264):1101–1104. [DOI] [PubMed] [Google Scholar]
  • 48.Niu D, Wei HJ, Lin L, George H, Wang T, Lee IH, Zhao HY, Wang Y, Kan Y, Shrock E, Lesha E, Wang G, Luo Y, Qing Y, Jiao H, Zhou X, Wang S, Wei H, Güell M, Church GM, Yang L. Inactivation of porcine endogenous retrovirus in pigs using CRISPR-Cas9. Science 2017;357(6357):1303–1307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Zheng Q, Cai X, Tan MH, Schaffert S, Arnold CP, Gong X, Chen CZ, Huang S. Precise gene deletion and replacement using the CRISPR/Cas9 system in human cells. Biotechniques 2014;57(3). [DOI] [PubMed] [Google Scholar]
  • 50.Schmieder V, Bydlinski N, Strasser R, Baumann M, Kildegaard HF, Jadhav V, Borth N. Enhanced Genome Editing Tools For Multi-Gene Deletion Knock-Out Approaches Using Paired CRISPR sgRNAs in CHO Cells. Biotechnol J 2018;13(3):1700211. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp info

RESOURCES