Abstract
Carcinoembryonic antigen (CEA) is used as a biomarker for colorectal cancer. It is expressed during fetal development but in healthy adult cells the expression is low. Due to its size and the high degree of glycosylation, there are no structures available for mature CEA. By employing novel structure prediction methods, we aim to investigate CEA tertiary structure and interactions.
Alphafold 3 server has increased the accuracy of structure predictions and allows for modelling of glycans in proteins and complexes. Models were created for a monomeric CEA, dimeric CEA and for CEA in complex with the antibody Tusamitamab. The structure of the monomeric glycosylated CEA exhibit two bends, one in the domain interface B1–A2 and one in the domain interface B2-A3. The dimer structure pairs in a parallel manner, with direct contacts in the N and the A2 domains of the two chains. The complex of CEA with Tusamitamab closely resembles the EM structure of the complex that was released after the training of Alphafold 3 was completed.
Overall, the investigations give new angles to investigate for CEA. The predicted bend, primarily in the B2 and A3 domain interface, would allow for dimer formation of CEA from both the same cell as from adjacent cells and could help to explain the outstanding issue on how it can fulfil both tasks. The prediction of the antibody binding to CEA was accurate, the all-atom RMSD was 1.3 Å. This is encouraging for other antibody – protein complexes predictions as the complex structure was not part of the training set for Alphafold 3.
Keywords: Carcinoembryonic antigen, Structure prediction, Protein complex, Colorectal cancer
Graphical abstract
Highlights
-
•
Prediction of the glycosylated tertiary structure of Carcinoembryonic antigen (CEA).
-
•
Prediction of homodimer of CEA and possible orientation in cell-cell interactions.
-
•
Accurate prediction of CEA and Tusamitamab complex with an all atom RMSD of 1.3 Å.
1. Introduction
Carcinoembryonic antigen (CEA) is a glycoprotein that participates in cell adhesion and is produced in the gastrointestinal tract during fetal development. In healthy tissues, CEA expression is limited and primarily found in the cells of the digestive system. CEA contains a variable (V)-like domain, named the N domain, followed by six constant C2-like domains named: A1, B1, A2, B2, A3, and B3, that are organized in A-B pairs. CEA is associated with the membrane through a glycosylphosphatidylinositol (GPI) linkage in the C-terminus of the protein [19]. The most common cause of CEA increase in plasma is gastrointestinal malignancy, such as pancreatic and colorectal cancer (CRC) [1,2]. However, plasma CEA may also increase in inflammation, e.g. pancreatitis and Crohn's disease as well as in elderly people and in people who smokes.
Colorectal cancer is the third most common form of cancer worldwide, after lung and breast cancer [3,4]. Emerging evidence highlights that chronic inflammation that promotes a local environment that is favourable for tumor survival and systemic immune response are two key factors in the oncogenesis in colon tissue. One of the pathways that drives the chronic inflammation in CRC are the Toll-like Receptors of the innate immune system, which currently is an active research area [5]. Once the tumor is established it has been shown that the high metabolic demand of cancer cells can cause imbalance in the mitochondria, leading to enhanced proliferation and metastasis [6]. CEA is used to facilitate the detection of CRC or to detect tumor recurrence. It has a certain value for following the treatment effect of palliative care. CEA has too low sensitivity to be used in the normal population or for high-risk groups but can be used in follow-up as a marker for early detection of recurrence [7].
The advent of novel, faster, and more accurate prediction methods for protein structures and protein complexes prompts the investigation of common disease-related proteins. Even though Alphafold has generated the structures of all human proteins and made them available for all in the Alphafold Protein Structure Database [8], they are based on the complete protein sequence, including eventual signal and propeptide that are processed out of the mature protein. It also omits post-translational modifications, such as glycosylations. Even though glycosylation often does not affect the fold of a protein, it might affect the dynamics of protein [9]. Furthermore, in the process of forming a protein complex, the glycosylation will impose steric hindrance and may even shape which areas of the protein surface are available as interactions sites.
The aim of the study was to investigate if novel structure prediction tools may shed light on the structure and function of CEA.
2. Material and methods
The CEA sequence was collected from Uniprot entry P06731 [10]. After omitting the signal and the pro-peptide from the sequence, residues 35–685 were used as input for the modelling (Fig. 1).
Fig. 1.
Sequence of CEA (CEAM5_HUMAN, P06731). Annotated with signal and propeptide and domains. The 21 glycosylation sites that are occupied in the predictions are marked with an arrow.
The structure of the mature CEA protein was predicted using the Alphafold 3 server [11]. N-glycans were added based on the occupied glycosylation sites identified by Pont et al. [12]. To each of the following asparagines at position 104, 152, 197, 204, 208, 256, 274, 330, 351, 375, 432, 466, 480, 508, 529, 553, 560, 580, 612, 650 and 665 were a glycan added. Each glycan was modeled as the common core structure of all eukaryotic N-glycans, Manα1-3(Manα1-6) Manβ1-4GlcNAcβ1–4GlcNAcβ1–Asn-X-Ser/Thr [13]. The notation used in Alphafold for the above core structure is NAG(NAG(MAN(MAN)(MAN))).
Dimerization of CEA was modeled by the inclusion of two chains of monomeric CEA in Alphafold 3, described above.
Lastly, modeling was performed with monomeric CEA and the full sequence of the fab of the antibody Tusamitamab, whose ScFv binding to CEA is described in a recent EM study [14].
Result and output files are stored in Nordling [15].
3. Results
3.1. Monomer structure
The predicted structure for CEA with N-glycans forms an overall elongated structure with the seven Ig-like domains, with two bends formed, the first one between the domain B1 and A2 and the second one between B2 and A3 domain (Fig. 2A). The bend between B1 and A2 domains is caused by a glycan at asparagine 351 that infers steric hindrance between the two domains. This is evident by visual inspection and verified by modeling of the structure without the aforementioned glycan present, where the two domains are positioned more linear (data not shown).
Fig. 2.
A. Predicted structure of CEA monomer with 21 glycans. The N-terminal is collored blue and the c-terminal is colored red. The attachment point to the cellmembrane is located in the c-terminal where an alanine is lapidated and anchors the protein to the cell. Asparagines with glycans close to the bends in the structure are labeled in blue with position and one letter amino acid abbreviation. B. Monomer colored according to regions that are predicted with high likelihood of relative position.
The bend between domains B2 and A3 is not obviously caused by a glycan as removal of all nearby glycans does not abolish the bend; although there are several glycans present in the two domains, they are on the opposite side of the bend. Interestingly, the sequence of the region between these two domains share similarities with each other, and the bend itself is located at a centrally positioned alanine in the linker region between the two domains (Fig. 1). The score of the model is pTM = 0.51 indicating that the overall fold is reasonably correct predicted. The ipTM of 0.5 indicates that some of the relative positions within the model could be erroneous. A plot that depicts which portion of CEA that are predicted with higher certainty is shown in Fig. 2B, where the molecule is divided into two regions that consist of domains 1–3 and 4–7. The relative position of the individual domains within these regions are likely to be predicted accurately. The position of the two regions in relation to each other is however of less certain. This indicates that the more N-terminal bend is predicted with lower probability than the more c-terminal bend, and this is most likely the region that causes the lower ipTM score.
3.2. Dimer structure
Two plausible models are predicted for the dimer complex. The top-scoring model details interactions between the A2-B2-A3 domains and corresponding domain in the other chain. The C-terminus that is anchored to the plasma membrane with an GPI-anchor, is directed in such that it could imply that it is a dimer protruding from the same cell. However, this conformation of the dimer does not contain any interactions between the variable N domain where point mutations have been shown to disrupt dimer-formation [[16], [17]]. The second-best scoring dimer model includes direct interactions between the N domains of the two chains (Fig. 3A). The glycans of the A1 and B1domains are pointing into a vacant area in between the protein chains, causing the dimer to form an extended version of a Fc domain of an antibody that is ended by direct interactions of the A2 domain, which in turn forms extensive contacts through the edge of one of the beta-sheets of the Ig-fold (Table 1, Fig. 3B). The mutated residues that have been shown to disrupt dimer formation Korotkova et al., 2008, [17] are positioned in the interface region of the two N domains (Table 1, Fig. 3C). The membrane attachment points are directed toward opposite sides, indicating that this dimer formation possibility is compatible with accommodate cell-cell interactions. The third-best model forms a similar structure as the second-scoring model, but the N domains are rotated 180° and form a dimer with an interface that do not contain the mutationally verified residues that influence dimer formation.
Fig. 3.
A. Predicted CEA dimer, where the first five Ig-like domains form an elongated structure with primary interactions between the n-terminal Ig-like V type domain and the third Ig-like C2 type domain. Residues in the interface regions of the Ig-like V type that have been verified to influence dimer formation are marked in orange. B. Dimer rotated 90° into the plane of view to show the interface region more clearly. The central cavity of the dimer that is partly filled with glycans from the first and second of the Ig-like C2 type domains. C. Close-up of the n-terminal Ig-like V type domain with residues that are known to influence dimer formation marked with blue labels.
Table 1.
Detailed list of interactions between domains in the dimer structure. Residues that are reported to interrupt dimer formation are marked with bold text and an asterisk.
| Chain 1: Domain | Residues | Type of interaction | Residues | Chain 2: Domain |
|---|---|---|---|---|
| N | Phe63∗ sidechain | hydrophobic | Phe63∗ sidechain | N |
| Phe63∗ sidechain | hydrophobic | Leu129 sidechain | ||
| Gly64 backbone Ca | hydrophobic | Ile125∗ sidechain | ||
| Ser66∗ sidechain | hydrogen bond | Leu129 backbone O | ||
| Tyr68∗ sidechain | hydrogen bond | Asn131 sidechain | ||
| Glu71 sidechain | salt bridge | Arg72 sidechain | ||
| Arg72 sidechain | salt bridge | Glu71 sidechain | ||
| Val73∗ sidechain | hydrophobic | Val73∗ sidechain | ||
| Gln78∗ sidechain | hydrogen bond | Leu129 backbone O | ||
| Val83 sidechain | hydrophobic | Leu129 sidechain | ||
| Thr90 sidechain | hydrophobic | Val130 sidechain | ||
| His123 sidechain | hydrogen bond | His123 sidechain | ||
| Ile125∗ sidechain | hydrophobic | Leu129 sidechain | ||
| Leu129 sidechain | hydrophobic | Phe63∗ sidechain | ||
| Glu133∗ sidechain | hydrogen bond | Gly75 backbone N | ||
| A1 | Asn152 Glycan | hydrophilic | Asn152 Glycan | A1 |
| Asn204 Glycan | hydrophilic | Asn204 Glycan | ||
| B1 | no contacts | B1 | ||
| A2 | Phe326 sidechain | hydrophobic | Asn330 sidechain | A2 |
| Ile327 backbone N | hydrogen bond | Asn330 backbone O | ||
| Ile327 backbone O | hydrogen bond | Asn330 backbone N | ||
| Ser329 backbone N | hydrogen bond | Ser329 backbone O | ||
| Ser329 backbone O | hydrogen bond | Ser329 backbone N | ||
| Asn330 sidechan | hydrophobic | Phe326 sidechain | ||
| B2 | no contacts | B2 | ||
| A3 | no contacts | A3 | ||
| B3 | no contacts | B3 |
3.3. CEA and FAb complex structure
The complex of the full-length CEA with glycans and the FAb of Tusamitamab captures the information from the EM structure, including both the protein interactions and the interaction with the Glycan at Asn612. (Fig. 4A). Interestingly, the overall model quality score pTM is higher than for the monomer itself at 0.52, while the score that relates to domain relations, ipTM are the same, 0.5. Overlay with the EM structure in 8bw0 show that the all-atom RMSD is 1.3 Å for the fragment and the two domains of CEA included in the EM structure (Fig. 4B). This structure was released on 2024-01-24 which is after the cutoff date of 2021-09-30 for the training set of Alphafold 3 [14].
Fig. 4.
A. Structure of CEA, colored in light-blue, bound to the FAb of Tusimatimab, colored in green and yellow. B. Comparison of EM structure of complex of ScFv of Tusimatimab (colored in khaki and light-green) and CEA Ig-like C2 type domain 5 and 6 (colored in cyan) and the corresponding portion of the complex of the FAb and CEA, the predicted structure is colored as in A.
4. Discussion
The structure of the monomer with added glycans depart from the structure that is predicted without glycans present and the one of the sequence of the immature chain, including signal and propeptide that is deposited in the Alphafold database (https://alphafold.ebi.ac.uk/entry/P06731). The current prediction has a bend in the structure in the linker region between B1 and A2 domains and that creates a sharp angle between B2 and A3 domains. The structure of the first bend is predicted with low certainty, it is influenced strongly by the glycan at position 351, as the removal of that glycan predicts a more elongated orientation at this linker region. The other region is less influenced by local interactions with glycans and the driver of the formation of the sharp bend is to hide a hydrophobic patch of residues, including Tyr424, Tyr426, Leu500, Pro501, Pro525, Ala527. These two regions are predicted with lower certainty in the unglycosylated models these regions might be flexible and adopt several positions in relation to each other. Another possible explanation is that the algorithm is skewed towards trying to hide these hydrophobic regions when more hydrophilic groups are present in the simulation.
Previous modeling based on EM data suggests a more elongated form of length 27–33 nm and a cylindrical shape of dimension 8 nm × 1 nm. In these experiments, dimers might form and what is actually measured could be the homodimer of CEA. Upon inspection of the dimer prediction, the first five domains of CEA form an elongated structure that is 8.5 nm wide and 3 nm high, which is in agreement with the published EM data. However, the length of the portion is only 21 nm. The two c-terminal domains are approximately 8 nm long and angled out of the plane. But, if these were more in line with the other domains, the distance would match the EM data. The bend might be a flexible hinge in CEA allowing the CEA dimer formation to accommodate from both the same cell and two separate cells [18]. In this manner, the dimer formation that is observed for the N-terminal domain can be satisfied by both trans and cis dimer formation of CEA [16].
The modeling of the complex with the Fab of the Tusimatimab antibody bound to CEA monomer creates a similar overall structure for CEA that is consistent with the previously reported structure of the ScFv CEA complex. The RMSD of all atoms 1.3 Å. This is a remarkable accomplishment of the algorithm, as the structure of the ScFv-CEA complex was released 18 months after the training set for Alphafold 3 was compiled. Even the influence of the glycan at position 612 is predicted.
5. Conclusion
Overall, the predictions present a possible interaction model that allows for dimer formation of CEA in both cis and trans orientation regarding position on cells, i.e. CEA located on the same cell or on different cells. The bent conformation will allow for the binding of CEA molecules from different cells as the interaction area is presented facing upwards from the cell surface. As the regions between the Ig-like domain are flexible, we assume that the same orientation driven by the interaction of the N-terminal Ig-like V-type domain can achieve cis-dimerization. This allows to the design of directed experiments to verify the proposed complex.
CRediT authorship contribution statement
Ivan Shabo: Writing – review & editing, Visualization, Investigation. Erik Nordling: Writing – review & editing, Visualization, Validation, Methodology. Mirna Abraham-Nordling: Writing – review & editing, Writing – original draft, Project administration, Methodology, Conceptualization.
Funding information
No funding was received for the current study.
Declaration of competing interest
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Co-author Erik Nordling is currently employed by Swedish Orphan Biovitrum AB. The company has no financial or scientific interest in the published work. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- 1.Cao W., Tang Q., Zeng J., Jin X., Zu L., Xu S. A review of biomarkers and Their clinical impact in resected early-stage non-small-cell lung cancer. Cancers (Basel) 2023 Sep 14;15(18):4561. doi: 10.3390/cancers15184561. PMID: 37760531; PMCID: PMC10526902. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Desai S., Guddati A.K. Carcinoembryonic antigen, carbohydrate antigen 19-9, cancer antigen 125, prostate-specific antigen and other cancer markers: a primer on commonly used cancer markers. World J. Oncol. 2023 Feb;14(1):4–14. doi: 10.14740/wjon1425. Epub 2023 Feb 26. PMID: 36895994; PMCID: PMC9990734. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Bray F., Ferlay J., Soerjomataram I., Siegel R.L., Torre L.A., Jemal A. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin. 2018 Nov;68(6):394–424. doi: 10.3322/caac.21492. Epub 2018 Sep 12. Erratum in: CA Cancer J Clin. 2020 Jul;70(4):313. doi: 10.3322/caac.21609. PMID: 30207593. [DOI] [PubMed] [Google Scholar]
- 4.Siegel R.L., Wagle N.S., Cercek A., Smith R.A., Jemal A. Colorectal cancer statistics, 2023. CA Cancer J. Clin. 2023 May-Jun;73(3):233–254. doi: 10.3322/caac.21772. Epub 2023 Mar 1. PMID: 36856579. [DOI] [PubMed] [Google Scholar]
- 5.Mukherjee S., Patra R., Behzadi P., Masotti A., Paolini A., Sarshar M. Toll-like receptor-guided therapeutic intervention of human cancers: molecular and immunological perspectives. Front. Immunol. 2023 Sep 26;14 doi: 10.3389/fimmu.2023.1244345. PMID: 37822929; PMCID: PMC10562563. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Midde A., Arri N., Kristian T., Mukherjee S., Sen Gupta P.S., Zhang Y., Karbowski M., Waddell J., Maharajan N., Hassan M.S., O'Hagan H.M., Zalzman M., Banerjee A. Targeting mitochondrial ribosomal protein expression by andrographolide and melatonin for colon cancer treatment. Cancer Lett. 2025 Jun 1;619 doi: 10.1016/j.canlet.2025.217647. Epub 2025 Mar 22. PMID: 40127816. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Wang R., Wang Q., Li P. Significance of carcinoembryonic antigen detection in the early diagnosis of colorectal cancer: a systematic review and meta-analysis. World J. Gastrointest. Surg. 2023 Dec 27;15(12):2907–2918. doi: 10.4240/wjgs.v15.i12.2907. PMID: 38222002; PMCID: PMC10784816. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Varadi M., Bertoni D., Magana P., Paramval U., Pidruchna I., Radhakrishnan M., Tsenkov M., Nair S., Mirdita M., Yeo J., Kovalevskiy O., Tunyasuvunakool K., Laydon A., Žídek A., Tomlinson H., Hariharan D., Abrahamson J., Green T., Jumper J., Birney E., Steinegger M., Hassabis D., Velankar S. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024 Jan 5;52(D1):D368–D375. doi: 10.1093/nar/gkad1011. PMID: 37933859; PMCID: PMC10767828. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Lee H.S., Qi Y., Im W. Effects of N-glycosylation on protein conformation and dynamics: protein Data Bank analysis and molecular dynamics simulation study. Sci. Rep. 2015 Mar 9;5:8926. doi: 10.1038/srep08926. PMID: 25748215; PMCID: PMC4352867. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.UniProt Consortium UniProt: the universal protein knowledgebase in 2023. Nucleic Acids Res. 2023 Jan 6;51(D1):D523–D531. doi: 10.1093/nar/gkac1052. PMID: 36408920; PMCID: PMC9825514. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A.J., Bambrick J., Bodenstein S.W., Evans D.A., Hung C.C., O'Neill M., Reiman D., Tunyasuvunakool K., Wu Z., Žemgulytė A., Arvaniti E., Beattie C., Bertolli O., Bridgland A., Cherepanov A., Congreve M., Cowen-Rivers A.I., Cowie A., Figurnov M., Fuchs F.B., Gladman H., Jain R., Khan Y.A., Low C.M.R., Perlin K., Potapenko A., Savy P., Singh S., Stecula A., Thillaisundaram A., Tong C., Yakneen S., Zhong E.D., Zielinski M., Žídek A., Bapst V., Kohli P., Jaderberg M., Hassabis D., Jumper J.M. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024 Jun;630(8016):493–500. doi: 10.1038/s41586-024-07487-w. Epub 2024 May 8. PMID: 38718835; PMCID: PMC11168924. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Pont L., Kuzyk V., Benavente F., Sanz-Nebot V., Mayboroda O.A., Wuhrer M., Lageveen-Kammeijer G.S.M. Site-specific N-linked glycosylation analysis of human carcinoembryonic antigen by sheathless capillary electrophoresis-tandem mass spectrometry. J. Proteome Res. 2021 Mar 5;20(3):1666–1675. doi: 10.1021/acs.jproteome.0c00875. Epub 2021 Feb 9. PMID: 33560857; PMCID: PMC8023805. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Stanley P., Moremen K.W., Lewis N.E., Taniguchi N., Aebi M. In: Essentials of Glycobiology [Internet] fourth ed. Varki A., Cummings R.D., Esko J.D., Stanley P., Hart G.W., Aebi M., Mohnen D., Kinoshita T., Packer N.H., Prestegard J.H., Schnaar R.L., Seeberger P.H., editors. Cold Spring Harbor Laboratory Press; Cold Spring Harbor (NY): 2022. N-Glycans. Chapter 9. PMID: 35536965. [Google Scholar]
- 14.Kumar A., Duffieux F., Gagnaire M., Rapisarda C., Bertrand T., Rak A. Structural insights into epitope-paratope interactions of a monoclonal antibody targeting CEACAM5-expressing tumors. Nat. Commun. 2024 Oct 30;15(1):9377. doi: 10.1038/s41467-024-53746-9. PMID: 39477960; PMCID: PMC11525548. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Nordling Erik. Mendeley Data, V1; 2024. CEA Predicted Structures and Complexes. [DOI] [Google Scholar]
- 16.Korotkova N., Yang Y., Le Trong I., Cota E., Demeler B., Marchant J., Thomas W.E., Stenkamp R.E., Moseley S.L., Matthews S. Binding of Dr adhesins of Escherichia coli to carcinoembryonic antigen triggers receptor dissociation. Mol. Microbiol. 2008 Jan;67(2):420–434. doi: 10.1111/j.1365-2958.2007.06054.x. Epub 2007 Dec 11. PMID: 18086185; PMCID: PMC2628979. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Taheri M., Saragovi U., Fuks A., Makkerh J., Mort J., Stanners C.P. Self recognition in the Ig superfamily. Identification of precise subdomains in carcinoembryonic antigen required for intercellular adhesion. J. Biol. Chem. 2000 Sep 1;275(35):26935–26943. doi: 10.1074/jbc.M909242199. PMID: 10864933. [DOI] [PubMed] [Google Scholar]
- 18.Boehm M.K., Mayans M.O., Thornton J.D., Begent R.H., Keep P.A., Perkins S.J. Extended glycoprotein structure of the seven domains in human carcinoembryonic antigen by X-ray and neutron solution scattering and an automated curve fitting procedure: implications for cellular adhesion. J. Mol. Biol. 1996 Jun 21;259(4):718–736. doi: 10.1006/jmbi.1996.0353. PMID: 8683578. [DOI] [PubMed] [Google Scholar]
- 19.Beauchemin N., Arabzadeh A. Carcinoembryonic antigen-related cell adhesion molecules (CEACAMs) in cancer progression and metastasis. Cancer Metastasis Rev. 2013 Dec;32(3–4):643–671. doi: 10.1007/s10555-013-9444-6. PMID: 23903773. [DOI] [PubMed] [Google Scholar]





