Abstract
The HUGO Gene Nomenclature Committee (HGNC) assigns unique symbols and names to human genes and its sister project, the Vertebrate Gene Nomenclature Committee (VGNC), names genes across selected vertebrates (chimp, macaque, horse, cattle, pig, dog, cat) in line with their human orthologs. The A2M gene family, a subfamily of the thioester-containing protein (TEP) superfamily, is well conserved across vertebrates and several members of this family have already been characterized as non-specific peptidase inhibitors. Chicken ovostatin, originally termed “ovomacroglobulin”, is an A2M family member that was first identified as being one of the most abundant proteins found in egg white. Two uncharacterized ovostatin homologs have also been reported in human. We wanted to assign standardized nomenclature to the multiple members of the A2M family across a wide range of vertebrate species, to capture the variation within this complex gene family. We constructed a maximum likelihood phylogenetic tree based on a multiple alignment of A2M protein sequences to help assign new nomenclature to previously unnamed A2M family genes, including the ovostatins, in human and across selected vertebrate species. This resulted in the naming of 4 human A2M family pseudogenes and 14 protein coding genes and 4 pseudogenes across VGNC species. An additional 48 genes were also named in model organisms (mouse, rat, xenopus, zebrafish, chicken) by their nomenclature committees based on this phylogenetic analysis.
Keywords: Phylogenetics, Alpha-2-macroglobulins, Ovostatins, Vertebrate evolution, Gene nomenclature, HGNC, VGNC, Gene family
Background
Gene nomenclature: the HGNC and the VGNC
The HUGO Gene Nomenclature Committee (HGNC) is the only group worldwide responsible for assigning unique symbols and names to human genes [1, 2]. Its sister project, the Vertebrate Gene Nomenclature Committee (VGNC), was established to name genes in a standardized way and in line with their human orthologs, across a selected set of vertebrate species that lacked their own nomenclature committees (currently chimpanzee, rhesus macaque, cat, dog, cattle, horse, pig) [1]. Although many high-confidence vertebrate orthologs can be named via an automated VGNC pipeline, others require manual curation involving steps such as assessing synteny, comparing gene models, phylogenetic analysis, and reviewing the literature [3].
The VGNC aims to name orthologs across vertebrates in a consistent manner. When there is a clear 1:1 orthology relationship between a vertebrate gene and a human gene, the VGNC routinely assigns the same symbol to both genes, e.g. A2M. New members of a gene family that have arisen through duplication receive new but related gene symbols, e.g. chicken A2MB. In this study we constructed a maximum likelihood (ML) phylogenetic tree to help us assign appropriate new nomenclature to previously unnamed alpha-2-macroglobulin (A2M) family genes in human and across VGNC species. We collaborated with other nomenclature committees including the MGNC (mouse) [4, 5] the RGNC (rat) [6, 7] and the CGNC (chicken) [8, 9] to ensure that any currently unnamed A2M family genes are named in these model organisms.
Human alpha-2-macroglobulin domain containing proteins
The alpha-2-macroglobulin (A2M) family in vertebrates includes homologs of ovostatin, which was first identified in chicken, and the murinoglobulins, which are predominantly found in rodents. The A2Ms are a subfamily of the thioester-containing protein superfamily (TEPs), which includes the complement genes C3, C4A, C4B, C5 and their closely related paralogs CPAMD8 and CD109; Fig. 1 illustrates the domain structure of the TEP superfamily proteins. The complement system is a key part of the innate immune system, which is thought to have evolved in early eukaryotes. The HGNC gene group “C3 and PZP like, alpha-2-macroglobulin domain containing (CPAMD)” [10] has been named after the shared domain and represents the TEP superfamily. It consists of the six complement genes mentioned above and a subfamily of the three currently characterized human members of the alpha-2-macroglobulin family. This superfamily is also known as the “I39 protease inhibitor family” in the MEROPS database of peptidases and their inhibitors [11]. All TEP superfamily members are thought to have originally evolved from an ancestral core of genes encoding proteins that contain a set of eight homologous domains [12].
Fig. 1.
TEP superfamily domain structure. Schematic created using SMART [15] to show domain structure in the A2M family and related complement family proteins. Note that we have not included the complement family in our phylogeny because their nomenclature is already well established
The protein products of all human TEP superfamily genes, with the sole exception of C5, contain an A2M thiol ester-containing domain (TED, shown in Fig. 1 as a green rectangle). The A2M TED domain features a conserved GCGEQ amino acid sequence, and the internal thioester bond between the glutamine and cysteine residues within this motif is highly reactive. This thioester bond is key to the complement factors C3b and C4b being able to bind hydroxyl and amine groups present on the surfaces of foreign cells, tagging them as targets for phagocytosis [13, 14].
The alpha-2-macroglobulin family
The A2M gene family is conserved across vertebrates and some invertebrates, including lancelets [16], ticks [17], shrimp [18], sea snails [19], sea urchins, oysters, flatworms [20], sea cucumbers [21] and horseshoe crabs [22, 23]. There is an A2M-like protein (ACY74611.2) in GenBank for the Spongia officinalis (sea sponge), but this is a partial sequence and is not currently linked to a gene. Some gram-negative bacteria including E.coli also have genes that encode A2M-like proteins - these may have been acquired via lateral gene transfer from their metazoan hosts [24].
The A2M family proteins all contain several alpha-2-macroglobulin domains and have largely been characterized as non-specific peptidase inhibitors that play an important role in regulating the innate immune system [25–27]. The proteins encoded by some family members may also have additional functions, including modulating the activity of T-helper cells, promoting immune cell migration and acting as chaperones enabling protein folding [28–30].
The human genome encodes three characterized A2M family genes that are located on the p-arm of chromosome 12 in the 12p13.31 region [31]. These are A2M (alpha-2-macroglobulin) (HGNC:7), PZP (PZP alpha-2-macroglobulin like) (HGNC:9750) and A2ML1 (alpha-2-macroglobulin like 1) (HGNC:23336). Two further human A2M family genes have been identified on chromosome 12 as homologs of the chicken genes that encode ovostatin proteins [32]; these genes - OVOS1P (Gene ID: 408186/ENSG00000214776) and OVOS2P (Gene ID: 144203/ENSG00000177359) - have been annotated as pseudogenes in human by NCBI RefSeq [33] and Havana/Ensembl-GENCODE [34]. However, OVOS2P has been annotated as protein coding on an alternate genome patch (GRCh38.p13) (ENSG00000275428) and has been referred to as being protein coding under certain conditions such as in ovarian cancer and melanoma [32, 35].
Characterized A2M family members: A2M, PZP and A2ML1
A2M (alpha-2-macroglobulin, HGNC:7) is the most well characterized member of the family. Its encoded protein forms tetramers and uses a unique “bait and trap” method: a cavity within the tetramer contains “bait” stretches of peptides (the A2M_N_2 bait region domain, shown in Fig. 1 as a pink rectangle) that are highly susceptible to cleavage by most peptidases. When the bait region is bound by a peptidase, it changes its conformation into an activated form, exposing the GCGEQ motif in the A2M TED domain (shown in Fig. 1 as a green rectangle) that forms a thioester bond to covalently bind and “trap” the enzyme. This reconfiguration also releases the A2M receptor binding domain (shown in Fig. 1 as a purple pentagon) which then allows the protein to act as a ligand and bind its receptor. This mechanism enables it to inhibit a wide range of peptidases, including members of all four peptidase classes (aspartic, cysteine, metallo and serine peptidases) [27, 39–41]. The A2M protein inhibitor cavity may fit peptidases of approximately 20–100 kDa and can trap more than one peptidase at once [36]. Steric hindrance then prevents the activity of these peptidases towards larger substrates, although they may remain active against substrates with a lower molecular weight that can still enter the A2M tetramer “trap” [37, 42].
The A2M family proteins including A2M, A2ML1 and PZP can interact with the multifunctional endocytic receptor encoded by LRP1 (LDL receptor related protein 1, HGNC:6692) [38]. LRP1 has also been identified and characterized as a receptor for apolipoprotein E (encoded by APOE (apolipoprotein E), HGNC:613) [39], β1-integrin (encoded by ITGB1 (integrin subunit beta 1), HGNC:6153) [40] and the neuronally expressed protein tau (encoded by MAPT (microtubule associated protein tau), HGNC:6893) [41]. A2M-LRP1 signaling has been shown to upregulate the activity of RhoA-GTPase (encoded by RHOA, HGNC:667) which plays a role in regulating cell cytoskeleton dynamics [42]. This interaction may also be involved in wound healing, as A2M has been shown to regulate microvesicle shedding through the activation of LRP1 on myofibroblasts [43]. Receptor-mediated endocytosis can then take the A2M-peptidase complex into the cell where it can be degraded in the lysosome [44].
The A2M protein also has a cytokine binding site that can interact with a wide range of signaling molecules, including multiple interleukins, growth factors, neurotrophins, apolipoproteins and amyloid beta [45–49]. It has further been reported to play a role as a cytokine carrier [50–53]), although its exact role in cytokine regulation in vivo still awaits further elucidation [53].
The human A2M gene has two characterized paralogs: A2ML1 (alpha-2-macroglobulin like 1) and PZP (PZP alpha-2-macroglobulin like). A2M displays ~ 72% similarity to PZP and ~ 39% similarity to A2ML1 at the amino acid level.
The gene PZP was previously named “pregnancy-zone protein”, as it was initially discovered in the blood serum of pregnant women [54, 55]. The HGNC updated its name in 2016 to “PZP alpha-2-macroglobulin like” to reflect its evolutionary history as part of the A2M family and because the name “pregnancy-zone protein” could be misleading, as it is expressed in multiple tissues and in males as well as females. Although it has been reported to be highly expressed in the human placenta during pregnancy [29, 56], PZP is also expressed in the liver [57], blood serum [58, 59] and immune cells [60], and may be particularly highly expressed during conditions of inflammation [28, 58].
The PZP protein usually forms dimers, although it can also associate into a tetrameric form [61, 62]. Its alternative bait region attracts a more restricted range of peptidases than A2M [30]. Both the A2M and PZP proteins can also act as chaperones, although dimeric PZP has been reported to be more efficient than the A2M tetramer at preventing protein misfolding [30]. PZP acts as an immune system modulator and interacts with several proteins known to be upregulated during pregnancy, including those encoded by PGF (placental growth factor, HGNC:8893), PAEP (progestagen associated endometrial protein, HGNC:8573) and VEGFA (vascular endothelial growth factor, HGNC:12680) [29, 38]. The PZP and PAEP proteins form a complex in the blood plasma during pregnancy and act synergistically to regulate T-cell activation, proliferation, and cytokine production [38]. These immunosuppressive effects may help prevent fetal rejection.
A2ML1 (alpha-2-macroglobulin like 1) has also been reported to be highly expressed in human and other species during pregnancy when it accumulates in the fetal amnion [29, 56]. It was first identified as a member of the A2M family in Xenopus [63] and its expression was subsequently reported in human epidermis as a monomeric protein [64]. The A2ML1 protein may be important for desquamation (skin peeling) through its inhibition of a range of peptidases, and it has also been confirmed to be one of several ligands that interact with the LRP1 (LDL receptor related protein 1) encoded receptor [44, 64].
Ovostatins
Ovostatin, originally termed “ovomacroglobulin”, was first identified as one of the most abundant proteins in the white of chicken eggs [65, 66]. Hen egg white ovostatin is an active enzyme that inhibits the peptidases trypsin, papain and thermolysin [67]. Ovostatin is also found in egg white from several other species, including crocodile [68, 69], duck [70, 71] and green sea turtle [72].
Chicken is a model organism for epithelial ovarian cancer studies because chickens develop ovarian tumors spontaneously, ovulate profusely, and the metastatic development of the disease seems to follow a similar course to that in humans [32, 73]. Nepomuceno et al. found six putative A2M gene family members in chicken (see [32]): A2ML1, A2ML2, A2M, LOC425756, OVOS1/OVST and OVOS2/OVSTL (LOC425756 has since been withdrawn by NCBI due to reannotation), and also refer to human OVOS1 (LOC408186/ENSG00000214776) and OVOS2 (LOC144203/ENSG00000177359). They looked at the role of ovostatins in chicken ovarian cancer development and also reported an overexpression of human OVOS2 in primary human ovarian cancers relative to non-cancerous tissues [32, 35, 74]). They generated an antibody against the human OVOS2 N-terminal amino acid sequence EGAKASKQGVLDLP present at residues 99–128 in the protein (highlighted in Fig. 2 in green). An alignment of all known human protein coding A2M family sequences plus the OVOS1 and OVOS2 sequences shows that this antibody should specifically identify OVOS2, and cross-reactivity with protein sequences encoded by its paralogs should be minimal.
Fig. 2.
A section of a Clustal Omega Multiple Alignment [75] of human A2M family proteins highlighting the sequence targeted by the antibody generated against human OVOS2 in Nepomuceno et al. in green, the GCGEQ motif that enables the formation of a thiol-ester bond that covalently traps peptidases in yellow, and the similar GGGEQ and GSGEQ motifs predicted to be present in the human ovostatin proteins by Nepomuceno et al. in pink
In 2018, Huang et al. reported that the knockdown of human OVOS2 has an effect on cell proliferation in melanoma cells [35]. They also found two potential mouse ovostatin homologs, one of 360 kDa and one of 88 kDa, which they speculated were probably protein isoforms from the same gene [76]. Their data showed that the gene they refer to as “mOH” encodes a serine peptidase inhibitor that inhibits type I collagen-degrading enzymes, and thus plays a role in remodelling the extracellular matrix (ECM). They speculated that this mouse ovostatin ortholog may be involved in regulating ECM changes during the estrous cycle and might also play a role in fetal development [76].
Studies suggest that, like A2M, ovostatins form tetrameric structures [77, 78]. Chicken ovostatin was reported to lack the CGEQ motif that can form a thiol-ester bond, and unlike chicken A2M it was not recognized by the mammalian A2M receptor and was most active towards metallopeptidases [78]. However, duck ovostatin was reported to have the thiol ester group [70]. The predicted human ovostatin sequences obtained from Nepomuceno et al. have GSGEQ (OVOS1) and GGGEQ (OVOS2) motifs respectively (see Fig. 2).
Murinoglobulins
The murinoglobulins, also referred to as “Mugs”, are monomeric proteins mainly found in rodents and are closely related to the A2M family. They have previously been characterized in mouse [79], rat [80], guinea pig [81] and hamster [82], but also in hedgehog, which is not a rodent, but a member of the Erinaceidae family [83]. The murinoglobulins have been suggested to be a product of a tandem duplication of an A2m/Pzp protogene [84].
CPAMD8 (C3 and PZP like alpha-2-macroglobulin domain containing 8) (HGNC:23228), a member of the TEP superfamily, has been lost in the rodent lineage; although the murinoglobulins may functionally replace the lost rodent Cpamd8, they are not orthologous [84].
The murinoglobulins are active at a different pH to the macroglobulins [79]. Iwasaki et al. suggest that the murinoglobulins are not simply an early form of the tetrameric alpha-macroglobulins, but may have evolved under different selective pressures [81]. Rat murinoglobulin has been shown to act as an inhibitor of viper venom [85].
Mug1, also published using the alias name “macroglobulin alpha (1)-inhibitor III”, has been described as the main macroglobulin of rodent blood plasma and its encoded protein is also referred to as ‘transcuprein’. Liu et al. state that “transcupreins are specific macroglobulins that not only carry zinc but also transport copper in the blood.” [86]. While the Mug1 protein acts as a copper transporter in rat, the A2M protein may be carrying out this function in humans [87]. Prior to this study, the mouse and rat genomes contained two genes named as murinoglobulins, Mug1 (MGI:99837, RGD:621366) and Mug2 (MGI:99836,RGD:1302962).
A2M family links to phenotypes and conditions of medical interest
A2M has been tentatively linked to the neurological condition Alzheimer disease (AD). The APP gene (amyloid beta precursor protein, HGNC:620) encodes a protein that undergoes proteolytic processing to reach its mature form. During this cleavage step, forms of amyloid beta peptides with different C-termini can be generated. A particular form, Aβ42, is prone to clumping and has been found as the main component of amyloid plaques in the brains of Alzheimer patients [88]. As the A2M protein can help degrade amyloid beta, it may play an important role in preventing pathology. However, despite this link - and although some studies have suggested that A2M variants may be associated with AD risk [89]—many larger studies have not shown this [90]. This is still a very active area of research and may lead to the development of new treatments for Alzheimer disease in the future.
A2M has also been linked to several cardiovascular conditions due to its regulatory role in hemostasis and fibrinolysis. It affects blood coagulation as it binds and inhibits activated protein C (encoded by PROC, HGNC:9451) [91], thrombin (encoded by F2, HGNC:3535) [92], factor III (also known as factor Xa, encoded by F3, HGNC:3541), plasmin (encoded in a zymogen form by PLG, plasminogen, HGNC:9071), and plasminogen activators such as urokinase encoded by PLAU (HGNC:9052). The A2M protein was found to be at a higher concentration in the blood plasma of ischemic stroke patients than in healthy controls, and A2M levels in stroke patients were found to positively correlate with patient age and the severity of areas of abnormal myelination in the brain, known as white matter lesions [93]. The A2M protein has also been identified in early lesions associated with atherosclerosis, where it may be having a protective effect, as it is expressed at a lower level in later stage plaques where the levels of elastase (encoded by ELANE, HGNC:3309) and interleukin 1 alpha (encoded by IL1A, HGNC:5991) are much higher. An A2M variant has been associated with an increase in the presence and severity of atherosclerotic plaques [94, 95].
PZP is more highly expressed in human blood plasma during pregnancy. Its encoded protein plays an important role as a chaperone. It has been suggested that high levels of PZP inhibit the aggregation of misfolded proteins including amyloid beta peptide, high levels of which are linked with the pregnancy associated condition preeclampsia [96] as well as with Alzheimer disease [97]. The A2M protein is also thought to be able to act as a chaperone, although it is not as efficient as PZP in this role; however, its chaperone activity can be increased by inducing its dissociation into PZP-like dimers [96]. A2M levels may also play a role in the development of pre-eclampsia in pregnancy [42]. Blocking the A2M-LRP1 interaction in a rat model was shown to slow the development of early onset pre-eclampsia, whereas smooth muscle cell expression of A2M was shown to promote its pathological progression by upregulating RhoA-GTPase [42].
An A2ML1 variant has been linked to a phenotype that is associated with similar characteristics to the RASopathy Noonan syndrome, including specific craniofacial and skeletal features [98]. However, there is also a study that disputes this and suggests these A2ML1 variants may not be causative for this syndrome [99]. The interaction of the variant A2ML1 protein with the protein encoded by LRP1 may affect its regulation of the MAPK/ERK cascade, hence the putative association with a Noonan like phenotype.
In addition, a duplication variant of A2ML1 has been associated with increased susceptibility to otitis media [100–102]. This could be due to an effect of the encoded variant peptidase inhibitor on the makeup of the microbiome of the outer and middle ear, affecting the prevalence of several pathogenic bacteria [100].
The A2ML1 protein has also been published (sometimes using the alias p170) as acting as an autoantigen in patients with the blistering disease paraneoplastic pemphigus (PNP) [103–107].
Higher levels of A2M may protect children against COVID-19 due to its role in maintaining hemostatic balance and mediating inflammation [108]. The A2M protein may play a role in restricting the peptidase-mediated entry of SARS-CoV-2 into human cells [109].
Results
There are three major subgroups of A2M family genes in vertebrates
The ML phylogeny in Fig. 3 shows that there are three major clades of A2M family genes in amphibians, reptiles, birds and mammals that are highly supported in this phylogenetic analysis (with strong ultrafast bootstrap support > 90%). These three clades are (1) the A2M/PZP/PZP2/Mugs/MUGL clade, (2) the ovostatins (OVOS) clade and (3) the A2ML gene clade. Neighbor-Joining (NJ) tree and Minimum Evolution (ME) trees constructed based on the same multiple alignment were both largely congruent with the ML tree with all three major groupings and all the same subclades being supported. Minor differences included slight rearrangements in the branching order for a few zebrafish genes in the NJ and ME trees in comparison to the ML tree; however zebrafish a2m1 remained at the base of the clade and the two subclades a2m2 and a2m3 were still strongly supported using all three tree building methods. There was also a small difference in the branching order in the xenopus a2ml3 clade with a2ml3a.3 branching before a2ml3a.2 in the NJ and ME trees. We recommend that ZFIN and Xenbase follow the naming shown in the ML tree as this method is generally considered to have an advantage over distance based and parsimony methods when looking to understand sequence evolution [110]; if they decide to investigate further, including more fish and amphibian species and comparing the synteny might help better guide the numbering within these clades. The branching pattern in the bird/reptile A2M clade also slightly differed between the trees as the ML tree grouped all the bird A2M genes together (100% support) and all the bird A2MB genes together (100% support), whereas the NJ and ME trees split these (A2M ostrich and crow together, A2MB ostrich and crow together, then chicken A2M and A2M together). However, this difference does not change the proposed naming for these genes. The ML and pruned ML subtrees corresponding to the Figures and the NJ and ME trees can be viewed interactively via a shared project key at the iTOL website (https://itol.embl.de/shared/bbraschi) [111].
Fig. 3.

Maximum Likelihood Tree built in IQ-TREE based on an edited Clustal Omega Multiple Alignment of A2M family amino acid sequences across vertebrates. Ultrafast bootstrap values (%) are shown below the nodes. The tree is rooted with the single A2M family gene in lamprey. Nodes have been collapsed and the fully expanded tree with all taxa labels can be seen on the iTOL (interactive tree of life) website at https://itol.embl.de/shared/bbraschi
Conserved gene blocks and synteny across species
The gene order of the majority of A2M family members has been fairly well conserved across vertebrates. Most A2M family genes are found together with their paralogs in a cluster of genes that we refer to as the “A2M family cluster” which shows some degree of shared synteny in most vertebrates (Fig. 4).
Fig. 4.
Synteny of A2M family genes. The arrows representing the A2M family members are coloured to match the corresponding subgroups in the phylogeny in Fig. 3. All non-A2M family gene arrows are shown in black. The grey blocks represent conserved flanking genes. A zigzag line represents a break in synteny where adjacent genes are separated by one or more loci. * denotes a gene currently annotated in NCBI as protein coding but the CDS is short and so it has been excluded from the phylogeny in Fig. 3. ‘P’ designates a pseudogene. ✝ denotes the position of the mouse and rat murinoglobulins which are shown in the light grey box titled “Murinoglobulins (Mugs)”. Although gene symbols should usually be italicized we have chosen to use a Roman font here for clarity
There are however some differences in the exact gene order of the A2M family members between species; these microsyntenic differences are likely due to gene duplications in select species and also some gene transposition events. The zig-zag icon in Figs. 4 and 5 shows where synteny is broken and genes that are shown as adjacent in the diagram have one or more non-A2M family member genes lying between them. OVOS1, OVOS3 and PZP2 in both marmoset and squirrel have been translocated further downstream but remain on the same chromosome (Chr 9 and Chr 4 respectively). A pseudogenized copy of marmoset OVOS1 remains in the cluster (OVOS1BP) and is syntenic to other primate OVOS1 and OVOS1B genes, some of which are protein coding. However, we have chosen to name the translocated single protein coding marmoset copy of this gene OVOS1; although it is not syntenic with other primate ovostatin genes as shown in Fig. 5, it groups with other primate OVOS1 sequences with 100% bootstrap support. In birds, the A2ML3 genes have been transposed onto different chromosomes, away from the main A2M family cluster.
Fig. 5.
Synteny of A2M family genes for selected primate species. a The primate A2M gene cluster. b Diagram to show the synteny of OVOS2. An asterisk * designates a gene that is currently annotated as protein coding, but only encodes a short amino acid sequence relative to other sequences in the alignment and so has not been included in the phylogeny shown in Fig. 3. All non-A2M family gene arrows are shown in black. The grey blocks represent conserved flanking genes. A zigzag line represents a break in synteny where adjacent genes are separated by one or more loci. Although gene symbols should usually be italicized we have chosen to use a Roman font here for clarity. The marmoset genes OVOS3, PZP2 and OVOS1 have been translocated out of the A2M family cluster further along chromosome 9
Conserved gene blocks of non-A2M family members can be seen across amphibians, reptiles, birds and mammals, which can act as useful markers when assessing the clusters of A2M family genes. The synteny diagrams in Fig. 4 show the conserved non-A2M gene block M6PR > < PHC1 (dark grey) in the same orientation across birds, reptiles and mammals. There is a second conserved non-A2M gene block of RIMKLB-MFAP5-AICDA (light grey), but the order of these marker genes and of the A2M family genes differs among species, likely due to gene inversion events.
Alpha-2-macroglobulin family genes in fish
The zebrafish D.rerio appears to have undergone several duplications of A2M family members, with at least sixteen genes within this family. The fish sequences group together in a single clade, suggesting that the ancestral amniote diverged from fish before an expansion of the ancestral A2M family genes took place in both sets of species. We have shared this data with ZFIN [112]. Gene ID:100006972 was previously named as a2ml (ZDB-GENE-090212-1), with Gene ID:100006993 named as a2ml2 (ZDB-GENE-030131-1342) and Gene ID:100006947 as a2ml3 (ZDB-GENE-030131-9800). However, the symbols A2ML1, 2, 3 and 4 are already assigned to other vertebrate genes in the A2M family and the zebrafish genes are not 1:1 orthologs of these genes. Therefore, as a result of this study Gene ID: 407643 (branching at the base of the zebrafish A2M clade) was updated to a2m1. All other zebrafish sequences fell into two subclades with one being named as a2m2a, a2m2b, a2m2c etc. and the second being named as a2m3a, a2m3b, a2m3c etc. This avoids both reusing any symbols or clashing with nomenclature in other vertebrates.
A2M in amphibians, reptiles and birds
The pruned tree shown in Fig. 6 shows that all the amphibian A2M genes group together in a monophyletic clade with high bootstrap support (99%). There are two X.tropicalis genes grouping with a R.bivittatum A2M gene and a M.unicolor gene in the A2M clade. These X.tropicalis genes are located on either side of the ‘conserved block’ of phc1 and m6pr genes and are both adjacent to other X.tropicalis A2M family genes. The gene LOC100489060 had already been assigned the name “alpha-2-macroglobulin” by Xenbase, whereas LOC619586 had been assigned the official symbol a2m alongside the same name “alpha-2-macroglobulin”. It therefore made sense for LOC619586 to remain as a2m (which could be aliased as a2ma) and for LOC100489060 to be assigned the new symbol a2m.2 by Xenbase and renamed as “alpha-2-macroglobulin gene 2”.
Fig. 6.
Pruned ML tree showing expanded A2M amphibian, reptile and bird clades. Ultrafast bootstrap values (%) are shown below the nodes, or connected to the node by a dashed line if moved to improve legibility. Taxon labels in red text denote genes that have been named by other nomenclature committees based on this study. NCBI Gene IDs are given after species names
The snake Python bivittatus and the lizard species Anolis carolinensis A2M sequences group together in a clade with 100% bootstrap support (Fig. 7). Crocodilians, represented here by Alligator mississippiensis, are more closely related to birds than lizards and snakes [113], and so the clade grouping alligator with birds is as expected.
Fig. 7.
Pruned ML tree showing expanded PZP2, A2M, PZP and MUGL/Mug mammalian clades. Ultrafast bootstrap values (%) are shown below the nodes, or connected to the node by a dashed line if moved to improve legibility. Taxon labels in blue text denote genes that have been named by the HGNC/VGNC based on this study. Taxon labels in red text denote genes that have been named by other nomenclature committees based on this study. NCBI IDs are given after species names
There appears to have been a pre-speciation duplication of A2M in birds; chicken, crow and ostrich all have two very closely related A2M paralogs that lie next to each other and fall into distinct subclades. As the root symbol A2ML has already been assigned to another set of genes, the second A2M-like gene adjacent to chicken A2M (Gene ID: 418251, CGNC:50716) was named as A2MB (Gene ID: 100857394, CGNC:65621) following discussion with the CGNC. Crow and ostrich also have orthologs of both of these genes lying in the same orientation.
A2M/PZP/PZP2/Mugs in mammals
The mammalian genes A2M, PZP, PZP2 and the murinoglobulins are a set of closely related paralogs that group together with 100% bootstrap support (Fig. 7).
The HGNC aims to name genes across vertebrates based on orthology and evolutionary relationships. All mammalian A2M/A2m protein sequences included in our analysis have a 1:1 ortholog relationship with human A2M. All VGNC species genes, plus A2m in rat and mouse, had already been named to reflect this by the VGNC/MGNC/RGNC prior to this analysis.
Of the species we selected, apes, monkeys, squirrel and cat have annotated protein coding copies of PZP which are around 1480 aa in length. Ring-tailed and mouse lemur, cattle, horse, rabbit, mouse and rat are missing PZP. At the time of this analysis, Canis lupus (dog) had a predicted PZP protein which was comparatively small at 1136 aa in length (but was annotated in NCBI as being a low quality protein sequence). An expert NCBI RefSeq annotator re-examined the dog PZP model and reannotated this locus as a pseudogene (NG_244666.1). We therefore updated the nomenclature of the dog PZP gene to PZPP to reflect the change in biotype and did not include the partial dog PZP protein sequence in our subsequent phylogenetic analysis.
Marmoset (a New World monkey) and squirrel are the only species included in this analysis that have copies of both PZP2 and PZP that do not lie adjacent to each other. Marmoset PZP and A2M lie next to each other in the A2M cluster region, syntenic to these genes in most other primate species considered here. Marmoset also has a second copy of OVOS1 that has been pseudogenized (OVOS1BP) in this region; OVOS3, PZP2 and OVOS1 have been translocated further along chromosome 9 and the gene order has slightly changed.
Squirrel is similar to marmoset in that its A2M family cluster is split into two separate smaller clusters of microsynteny. Squirrel has PZP and A2M lying in the A2M family cluster region with OVOS1, OVOS3 and PZP2 grouped together further along the same chromosome.
All the placental mammals we looked at have a single copy of PZP2 although this seems to have become pseudogenized in great apes (see Fig. 5 for synteny diagram). The non-coding human ortholog of PZP2 was named as PZP2P (HGNC:56245, Gene ID: 124903067). The mouse and rat Pzp2 genes were previously both named as Pzp, but as they group within the mammalian PZP2 clade they have now been renamed. Mouse and rat have a synteny breakpoint involving an insertion of around 7 Mb between Pzp2 and A2m and the mouse/rat Pzp ortholog may have been lost during this rearrangement event. Alternatively, it is possible that the murinoglobulin family has evolved from an ancestral copy of PZP/Pzp in some species as some, but not all, of the species that lack PZP have murinoglobulins or murinoglobulin like genes.
Murinoglobulins (Mugs)
The mouse genes Mug1 and Mug2 had already been named by the MGNC as murinoglobulins prior to this study, as well as two pseudogenes: Mug-ps1 (MGI:99838 Gene ID: 17835) (aliased as Mug3) and Mug4-ps (MGI:101843 Gene ID: 434083). Following consultation with the MGNC they agreed to rename Mug-ps1 as Mug3-ps (MGI:99838 Gene ID: 17835) to establish this as part of the murinoglobulin cluster. We also identified a previously unnamed mouse murinoglobulin encoding gene, Gm7298, which the MGNC have now assigned as Mug5 (MGI:3648717 Gene ID: 640530). A further unnamed predicted pseudogene in mouse, Gm10319, was named as Mug7-ps (MGI:3643520 Gene ID: 381806). It appears likely that the mouse murinoglobulin genes have evolved via multiple duplications of an ancestral A2m like gene. A Tmsb10 (thymosin, beta 10) pseudogene has also been duplicated in tandem in mouse alongside the murinoglobulin genes.
The rat orthologs of mouse Mug1 and Mug2 had already been named by the RGNC prior to this study. Two further rat murinoglobulin genes, previously named as Mug1l1 (RGD:1584999 Gene ID: 297568) and Cpamd8 (RGD:1566313 Gene ID: 297572), have been renamed based on this study. Our phylogeny and syntenic analyses suggest that mouse Mug5 (Gm7298) and rat Mug1l1 are likely orthologs and hence RGNC renamed this rat gene as Mug5. The rat gene previously called Cpamd8 is not the ortholog of human CPAMD8 and was renamed as Mug6.
Some other rodents that we looked at appear to have only a single Mug gene, such as Heterocephalus glaber (naked molerat) and Mesocricetus auratus (hamster). There are three adjacent MUGL gene models in the lagomorph Oryctolagus cuniculus (rabbit) but only two of these are included in the tree in Fig. 4 - the third gene model (LOC138843649) that lies in between these two paralogs appears to be partial and is currently predicted to encode a short 158aa protein, so may represent a pseudogene. These three rabbit MUGL genes may have resulted from an in-species duplication event of an ancestral MUG-like gene. Surprisingly, there is also a murinoglobulin-like gene predicted to encode a full length 1482 aa protein present in Equus caballus (horse).
Ovostatins: amphibians, reptiles and birds
All amphibian OVOS-like genes cluster together in a monophyletic clade with 100% bootstrap support (Fig. 8). The amphibian ovostatin genes were assigned the root symbol OVOS5 in our tree and so the five X.tropicalis genes which all grouped together in a clade were named as ovos5a-e. There seems to have been an expansion of ovostatins in X.tropicalis when compared to the caecilians. The reptile and bird ovostatin genes fall into two sister clades which we have designated as OVOS4A and OVOS4B. Chicken OVOS4A was previously referred to as “OVSTL” (ovostatin like) and chicken OVOS4B as “OVST” (ovostatin) in Nepomuceno et al. [32]. Updating the nomenclature to use the OVOS root reflects the homology of these genes to other ovostatin genes across vertebrates.
Fig. 8.

Pruned ML tree showing expanded ovostatin (OVOS) clades in all species. Ultrafast bootstrap values (%) are shown below the nodes, or connected to the node by a dashed line if moved to improve legibility. Taxon labels in blue text denote genes that have been named by the HGNC/VGNC based on this study. Taxon labels in red text denote protein coding genes that have been named by other nomenclature committees based on this study. NCBI IDs are given after species names with the exception of OVOS1P and OVOS2P in human which list UniProt IDs and correspond to the sequences from Nepomuceno et al. [32]
Copies of OVOS4A and OVOS4B are adjacent to each other in all of the reptile and bird species we have looked at, and there appears to be a many: many orthology relationship between reptiles/birds and the mammalian ovostatin genes (i.e. OVOS4A and OVOS4B are co-orthologs of the mammalian OVOS1, OVOS2 and OVOS3 genes).
Ovostatins: mammals
OVOS1 and OVOS2P group together in the same clade with 100% bootstrap support (Fig. 8). Human OVOS1P groups with the great ape OVOS2P pseudogenes rather than with the other primate protein coding OVOS1 genes. Predicted protein sequences were used for loci we now regard as pseudogenes.
As Fig. 8 shows, the bootstrap support for grouping the OVOS1 and OVOS2 proteins with OVOS3 is low (only 47.8%). As HGNC/VGNC naming is human-centric and OVOS1 and OVOS2 had already been used in the literature [32, 35, 114], the human genes were assigned the symbols OVOS1P, OVOS1BP and OVOS2P, with OVOS2P being used to name the great ape specific duplication. OVOS3 was assigned to genes encoding predicted protein sequences grouping in a clade with the macaque Gene ID:722294. Another option could have been to name OVOS3 as OVOS1C, but we felt that as the exact evolutionary history of these genes is difficult to ascertain that assigning a new number would be justified in this case. The mammalian OVOS3 gene is currently predicted to be protein coding in diverse species including elephant, macaque, marmoset, squirrel, and molerat, but is missing from most other mammals.
Human has an additional ovostatin transcribed pseudogene lying between DDX12P and DDX12B. It was initially annotated as a lncRNA, but when we contacted NCBI a RefSeq annotator confirmed that it is a likely pseudogene which we have now named as OVOS1BP (ovostatin 1B, pseudogene)(HGNC:58687, Gene ID:728715). Chimp also has a pseudogenized ortholog of this gene (Gene ID: LOC465369). The gorilla gene LOC115930232 lying between OVOS1P and PZP2 is currently annotated as protein coding. A TBLASTN of this sequence against human and chimp suggests that this gorilla predicted protein sequence shows greater similarity to human OVOS1BP (89% cover, 88% identity) and OVOS1P (84% cover, 92% identity) than to orangutan OVOS3 (62% cover, 86% identity). As the predicted protein is considerably shorter (207 aa) than full length ovostatins, we think this is likely to be a pseudogene. Given this gene also lies in the same orientation as Chimp OVOS1BP we are confident this is the gorilla OVOS1BP ortholog.
There is good transcript and CDS conservation evidence that OVOS1 and OVOS3 are protein coding in monkeys (represented in this analysis by rhesus macaque and marmoset) and gibbons (represented by Siamang) where they encode proteins of ~ 1450 aa. Transcript data is very limited in orangutan but it is likely that their OVOS1 and OVOS3 genes are protein coding, as the ORFs encoding > 1445 aa of Siamang models OVOS3 (Gene ID: 129483393) and OVOS1 (Gene ID:129483384) are maintained by Pongo abelli OVOS1 and OVOS3, respectively.
OVOS2 and DDX11: a tandem duplication and translocation in great apes
The pseudogene status of OVOS1 in apes (with the exception of orangutan) and its protein coding status in monkeys is largely consistent with the hypothesis that African apes had a duplication and translocation of DDX11 + OVOS2 coincident with the pseudogenization of their parental DDX12 + OVOS1 genes in the primate A2M/PZP/PZP2 gene block. This event may have taken place after the ancestor of gorillas, bonobos, chimps and humans diverged from their shared ancestor with orangutans. Orangutan, siamang and macaque all lack a copy of OVOS1B, but have a copy of OVOS3 instead. OVOS1B and OVOS3 are close paralogs and it is possible that they could be orthologous; however, as they fall into distinct clades we have chosen to assign them different numbers. The synteny of the primate OVOS1, OVOS1B and OVOS3 genes can be seen in Fig. 5a and that of the great ape OVOS2 genes in Fig. 5b.
A2ML cluster genes
The A2ML1, A2ML2 and A2ML3 genes fall into three separate clades and represent a set of closely related paralogs that have evolved from a common ancestral gene.
A2ML1: amphibians, reptiles and birds
X.tropicalis currently has four annotated models that may represent A2ML1 genes, but at present only one is predicted to encode a full length protein (Gene ID:108644844, 1468 aa). We have named this gene as a2ml1c in line with its caecilian orthologs in the same clade (100% bootstrap support). The other three Xenopus a2ml1 genes are currently predicted to encode very short proteins and may be pseudogenes, or require further annotation. More work is needed to fully annotate and name this family of genes in the closely related species X. laevis. All data has been shared with Xenbase.
The caecilians have five A2ML1 genes that are annotated as encoding full length proteins (~ 1457 aa). Four of these genes lie adjacent to each other in R.bivittatum and M.unicolor; two have been flipped in orientation in M.unicolor in comparison to R.bivittatum (A2ML1E and A2ML1F, Fig. 6, genes represented by light green arrows). A2ML1C in both caecilian species has been named as an ortholog of A2ML1C in birds and reptiles. The other three caecilian A2ML1 genes in the cluster have been named as A2ML1D, A2ML1E and A2ML1F as these are not 1:1 orthologs of either of the bird and reptile A2ML1 or A2ML1B genes. Both representative caecilian genomes also have a copy of an A2ML1 gene that lies outside the A2M family gene cluster, that we denote here as A2ML1G. These caecilian A2ML1G genes may have arisen from the duplication and subsequent translocation of either an ancestral amphibian A2ML1 gene, or resulted from the duplication/translocation of one of the A2ML1 genes that are currently found within the A2M family gene cluster.
The ground boa has a single copy of A2ML1 that groups with alligator A2ML1 and branches at the base of the mammalian A2ML1 clade with a 100% bootstrap value (Fig. 9). The green anole appears to have lost this gene completely. Alligator has three A2ML1 copies in its genome, as do some birds, including the New Caledonian crow. Chicken has a single A2ML1 gene encoding a full length protein. It had already been named as A2ML1 by the CGNC prior to this study, as the 1:1 ortholog of human A2ML1. We denote the crow and ostrich orthologs found in the same clade in the tree in line with this and assign the two additional A2ML1 genes in alligator, crow and ostrich as A2ML1B and A2ML1C. The chicken A2ML1B gene is currently predicted to encode a protein that is only 100 residues long, which is far shorter than a typical A2M family protein, so it has likely been pseudogenized. The chicken ortholog of the A2ML1C crow and ostrich genes has been lost completely.
Fig. 9.
Pruned tree showing expanded A2ML1 clades. Ultrafast bootstrap values (%) are shown below the nodes, or connected to the node by a dashed line if moved to improve legibility. These genes had already been named in VGNC species as 1:1 orthologs of human A2ML1 via the VGNC pipeline. Taxon labels in red text denote protein coding genes that have been named by other nomenclature committees based on this study
A2ML1: mammals
All the mammalian species included in our analysis have a single A2ML1 gene that lies between two conserved blocks of non-A2M family genes (shown in Fig. 4 as a light green arrow between two gray bars). The human gene had already been named as A2ML1 prior to this study and all 1:1 orthologs in VGNC species were already named in line with this by the VGNC pipeline [3]. Hamster has a full length protein coding copy of A2ML1, but this gene appears to have become pseudogenized in mouse (Gene ID: LOC141566927/ENSMUSG00000118448/Gm50492) and rat (Gene ID: LOC141566928); some remnants of these genes remain, but there is no transcriptional evidence for them and so they are considered non-transcribed pseudogenes and are now named as A2ml1-ps in both these species.
A2ML2
The A2ML2 genes are found in some amphibians, reptiles and birds, and are strongly supported as a clade with a 93.7% bootstrap value (Fig. 10). The CGNC had already named the chicken gene adjacent to chicken A2ML1 (CGNC:13800 Gene ID: 418254) as A2ML2 (CGNC:53231 Gene ID: 427942) prior to this study. We have therefore named all genes in the same clade as chicken A2ML2 in line with this. X.tropicalis appears to be the only amphibian genome containing this gene out of our selected species - it is not present in the two caecilian species included. Constructing a NJ tree using a Poisson model rather than a Jones-Taylor-Thornton (JTT) model [115] grouped the Xenopus a2ml2 sequence with the Xenopus ovostatins instead of with the A2ML2s. It is likely that the ML tree and NJ tree with an underlying JTT model are a better fit to the data and that this gene is part of the A2ML2 clade and is not an ovostatin.
Fig. 10.
Pruned ML tree showing expanded A2ML2 and A2ML3 clades. Ultrafast bootstrap values (%) are shown below the nodes, or connected to the node by a dashed line if moved to improve legibility. Chicken A2ML2 had been previously named by the CGNC. Taxon labels in red text denote protein coding genes that have been named by other nomenclature committees based on this study
The three alligator genes LOC102577230, LOC132243187 and LOC106737331 (not shown in the tree) were all annotated as being protein coding at the time of this analysis, but all are predicted to encode proteins of less than 500 residues, so either represent A2ML pseudogenes, or should be merged into a single A2ML2 gene model. Following discussion with a RefSeq annotator the decision was made to merge these three loci into a single putatively protein coding model, LOC102577230 (show with a dark green arrow and an asterisk in Fig. 4), though there is still a lack of transcript evidence for this locus. This new alligator A2ML2 protein sequence could not be included in our phylogeny at this time.
There has been a gene duplication and inversion event for A2ML2 in boa and green anole; we have named these genes as A2ML2A and A2ML2B (Fig. 10).
A2ML3
The amphibian, reptile and bird A2ML3 sequences group together with high bootstrap support (Fig. 10). There have been multiple duplication events in Xenopus leading to an expanded A2ML3 repertoire. The caecilian A2ML3 genes are found within the A2M gene cluster and are surrounded by A2ML1 genes. This arrangement probably best represents the ancestral location of this gene, which is a very close paralog of the A2ML1 genes; A2ML3 could have evolved from a duplication of an A2ML1 gene, or vice versa. There appears to have been a subsequent translocation event for this gene in an ancestor of reptiles and birds - their A2ML3 orthologs are not found within their A2M gene cluster regions and instead lie between PRR29 and FTSJ3 in birds and COL1A1 and FTSJ3 in alligator.
Discussion
In this study we constructed a phylogenetic tree for the well conserved A2M gene family using a multiple alignment containing representative vertebrate taxa. This has revealed the highly variable repertoire of A2M family genes across key species and enabled us to elucidate the relationships between them. We have determined that there are three major clades within this family and have used the scientific literature as a basis to assign nomenclature for the genes within each clade - clade 1: A2M, PZP, PZP2, Mug, and MUGL root symbols; clade 2: the OVOS root symbol; and clade 3: the A2ML root symbol. Phylogeny and synteny (conserved gene order) were further used to determine the precise gene numbering within each clade. Note that due to pseudogenization of the ovostatin (OVOS) genes in the human genome the A2M family lacked a common nomenclature for any of the members of this major clade.
The HGNC has named the human ovostatin orthologs as OVOS1P (HGNC:34045), OVOS1BP (HGNC:58687) and OVOS2P (HGNC:34046), and the human PZP2 ortholog, which is a unitary pseudogene, as PZP2P (HGNC:56245). It remains possible that human OVOS2 could be protein coding under certain conditions or in some individuals (i.e. it could be a segregating pseudogene) as previously reported [32]. A tandem duplication of an ovostatin and a DDX family gene in great apes, and the fact that orangutan still has a coding OVOS1 gene in line with monkeys such as macaque, helps support the hypothesis that orangutan is not as closely related to human as the other great apes are.
We chose to name the pseudogenized human ortholog of the protein coding mammalian PZP2 genes as PZP2P (HGNC:56245, Gene ID: 124903067) rather than creating a new symbol using the root “A2M”. PZP2 is a closely related paralog of both PZP and A2M and these three genes could all have been named using the root symbol A2M (e.g. PZP could have been assigned the symbol “A2M2” and PZP2 could have been “A2M3”). However, there are already multiple genes named with an A2ML root, as well as there now being an A2MB gene in chicken, and so we chose to use the symbol PZP2 to name this very close paralog of both PZP and A2M. In addition, the HGNC is pragmatic in its gene naming and aims to retain symbols that have already been published in the literature where possible, including A2M and PZP, in order to avoid confusion.
The murinoglobulins were, as the name suggests, first named in the mouse and rat based on publications [45, 79, 80, 116]. With hindsight, the term “murinoglobulin” is not ideal, as even some placental mammals including rabbit and horse appear to have a single copy of a murinoglobulin like gene. However, as the ‘Mug’ nomenclature has already been well published for rodents it makes sense to retain it and use the root symbol MUGL (Mug like) to refer to the closely related non-murine genes. The MUGL/Mugs are close paralogs of A2M, PZP and PZP2.
We looked for murinoglobulins in other vertebrates besides horse and rabbit and also found gene models representing likely MUGL orthologs in some Perissodactyla (odd-toed ungulate) species including zebra and rhino, and also in some Artiodactyl (even-toed ungulates) species such as hippo. There is also a probable Mug-like gene in pangolins, and less surprisingly also in other non-murine rodents.The horse MUGL (Mug like) gene (Gene ID: 100061656) lies next to A2M, syntenic to the murinoglobulins in mouse and rat that lie adjacent to the mouse and rat A2m genes.
The predicted human OVOS1 protein sequence groups with the great ape OVOS2P sequences rather than with the protein coding primate OVOS1 sequences (Fig. 8). In addition, the OVOS3 sequences from squirrel, molerat, macaque, siamang and orangutan only group with OVOS1 + OVOS2 with a low bootstrap value (47.8%). Some of the currently predicted mammalian ovostatin sequences may not ultimately remain annotated as protein coding. Pseudogenes degrade and their sequences diverge from those of their parent genes as time passes, although the rate of this degradation varies. While amino acid sequences derived from pseudogenes would still be expected to group with their protein coding parent genes, this degradation may result in noise in a dataset [117] and a loss of phylogenetic signal that could cause tree building artifacts and/or low bootstrap support values.
Indeed, although [32, 35] reported identifying human ovostatin proteins, an NCBI RefSeq annotator studied the gene models and decided that they are most likely to be transcribed pseudogenes, i.e. gene copies that were previously functional but that have undergone a series of mutations to become non-protein coding, although they are still transcribed into RNA. There is a lot of transcript data in human (for OVOS1P and OVOS2P) with plenty of exon variation, but it is currently not possible to assemble these transcripts to make a plausible protein model of the length typical of an ovostatin. In all the African apes, the 5’ end encoding roughly residues 1–190 seems to be missing. Remnant partial open reading frames (ORFs) for the ovostatins can be observed. For example, human OVOS2P (Gene ID: 144203) has an ovostatin-like ORF of 950 codons that contains the sequence EGAKASKQGVLDLP, as detected by the Nepomuceno et al. antibody, but translation would have to bypass upstream ORFs of 225, 283, and 59 codons and evade nonsense mediated decay. The OVOS1 and OVOS2 loci also have a complex rearrangement, and perhaps local duplication including some of DDX11, in the African apes. It is possible that proteomic processes are so highly dysregulated in ovarian cancer cells that these obstacles are overcome, but there is currently no transcriptional evidence for normal ovostatin protein production in African apes and humans.
The A2ML2 and A2ML3 genes could have also been assigned A2ML1 symbols with letter suffixes. However, their protein sequences fall into distinct, supported clades (See Fig. 10) and follow different evolutionary histories to the A2ML1 genes despite their undoubted shared ancestry (see Fig. 4 for synteny diagrams). For example, the caecilian A2ML3 genes lie in the opposite orientation to their A2ML1 genes, the A2ML3 gene in birds and reptiles has been translocated from the A2ML cluster, and A2ML2 has been duplicated in reptiles. Assigning these loci separate A2ML numbers can be justified based on phylogeny (Fig. 10) and simplifies communication, enabling clearer discussion of this set of paralogs.
Conclusions
This study has enabled us to clarify the relationships between the members of the A2M family in vertebrates and to name 4 human pseudogenes in this family (OVOS1P, OVOS2P, OVOS1BP and PZP2P). Following our analysis and recommendations, a further 14 protein coding genes (PZP2 in macaque, dog, cat, horse, cattle and pig; MUGL in horse; OVOS1 in macaque, dog, cat, horse, cattle and pig; OVOS3 in macaque) and 4 pseudogenes (OVOS1BP, OVOS2P and PZP2P in chimp and PZPP in dog) across seven species within the A2M family have been classified and named by the VGNC. The nomenclature committees for other vertebrate model organisms have been able to name an additional 48 A2M family genes. Collaborations with RefSeq have also enabled improvements to the annotations of this gene family in the NCBI Gene resource. As automated gene predictions rely on the validity of current annotations, these updates and refinements will likely result in more comprehensive and accurate annotation of this gene family in further species in the future.
In summary, this study provides a greatly improved understanding of this interesting gene family and its complex evolutionary history in vertebrates, including many key model organisms. This could also prove particularly insightful when investigating how variants of A2M family members are linked to human phenotypes.
Methods
Sequence selection—picking a range of vertebrate species
The phylogenetic analysis includes representative taxa from the five main classes of vertebrates: fish, amphibians, reptiles, birds and mammals. We chose to include all VGNC species (chimp, macaque, cattle, horse, pig, cat and dog) so that genes could be named in these species, and several model organism genomes (zebrafish, green anole, chicken, mouse, rat) as the assembly and annotation of their genomes are of comparatively high quality. Xenopus tropicalis was chosen rather than the model organism Xenopus laevis because it has a diploid rather than a tetraploid genome, which we hoped would simplify the analysis. The New Caledonian crow, Corvus moneduloides, and the ostrich Struthio camelus were included as these both represent a different group of birds to the Galliformes, which chicken (G.gallus) is classified as. In addition these two genomes look to be of fairly high quality, and their genes have been assigned to specific chromosomes, which aids in synteny analysis. We chose Alligator mississippiensis as a representative crocodilian because most of the saltwater crocodile Crocodylus porosus genome is still on unplaced scaffolds whereas the alligator genes have been assigned to specific chromosomes. The Papuan ground boa (Candoia aspera), also known as the viper boa, represents snakes, as a BLAST search with the human OVOS2 protein sequences against snakes suggested that the ground boa appeared to have a large repertoire of A2M family genes assigned to specific chromosomes rather than unplaced scaffolds.
As naming human genes is the key HGNC priority, we added several primate species as well as the VGNC species chimpanzee and macaque: gorilla (Gorilla gorilla) and the Sumatran orangutan (Pongo abelli) as additional great ape species, and the lesser ape siamang gibbon (Symphalangus syndactylus), the New World monkey marmoset (Callithrix jacchus), and the wet-nosed primates ring-tailed lemur (Lemur catta) and the mouse lemur (Microcebus murinus) to help us better understand the evolution of the A2M gene family in primates. We also added in some further rodent species including gray squirrel (Sciurus carolinensis), woodchuck (Marmota monax also known as groundhog or marmot), molerat (Heterocephalus glaber) and hamster (Mesocricetus auratus), to help in investigating the evolution of murinoglobulins, which were initially thought to be murine specific. We also included rabbit (Oryctolagus cuniculus) to represent the lagomorphs.
Multiple alignment construction and phylogenetic analysis
Predicted human OVOS1 and OVOS2 sequences were taken from Nepomuceno et al. [32]. The predicted human protein sequence for OVOS2 was BLASTed against selected species (one at a time) using the NCBI tool BLASTP [118, 119] with the default settings. All BLAST hits were downloaded and then sequences with < 60% query cover and < 25% percentage identity were removed from the FASTA file along with any duplicate sequences. The remaining sequences were then examined and any that were already annotated as being from more distantly related genes (such as those from the TEP superfamily, e.g. CPAMD8 or C3) were briefly assessed to check that they were not A2M family genes that had been incorrectly labelled. If we were certain that they were not A2M family sequences they were then removed from the file. If any sequences left in the FASTA file appeared to be partial or were labelled as a “low quality protein” we checked Ensembl [120] to see if there were any longer protein sequences available to replace those taken from NCBI. When there were multiple protein isoforms available from a single gene the longest isoform (typically X1) was retained.
A multiple alignment was constructed using ClustalOmega [75] (213 sequences, 2648 positions/data columns). The alignment was then automatically edited using TrimAl [121] (https://vicfero.github.io/trimal/) using the -automated1 setting. The manual states that the automated heuristic is optimized for trimming alignments to be analysed by maximum likelihood phylogenetic analyses. As there were still some columns with a high % of gaps trimal was run using the options -in < infile> _edit.fasta -out < outfile> -gt 0.8. This stripped out any columns for which at least 20% of the sequences had a gap. The alignment was then reviewed using AliView [122]. The resulting alignment had 213 sequences with 880 columns remaining.
The online server hosted at the Center for Integrative Bioinformatics in Vienna (CIBIV) was used to run IQ-TREE to build a ML tree [123]. The sequence type was selected as Protein, the substitution model was autoselected and JTT + R6 was used as the best fit according to the Bayesian Information Criterion. Free rate heterogeneity was selected and no ascertainment bias correction was applied. Ultrafast bootstrap support analysis [124] was selected and run for 1000 iterations. SH-aLRT branch tests and approximate Bayes testing was carried out to test single branch support in the ML tree. A Neighbor-Joining tree and a Minimum Evolution tree were constructed using MEGA12 using a JTT model based on the same edited multiple alignment that was used to build the ML tree. All trees were visualized and annotated using iTOL [125] and can be viewed online using the shared project key [111].
Rooting the tree
The ML tree (Fig. 3) was rooted using the single Petromyzon marinus (lamprey) A2M family protein sequence as an outgroup. P.marinus is considered a model organism representative of an early jawless vertebrate and as such is a good choice as an outgroup for rooting the tree.
Acknowledgements
The HGNC would like to thank David Webb from the NCBI for his genome annotation work improving A2M family gene models in vertebrate species and for helpful discussions about this manuscript. Thanks also to Susan Tweedie for constructive discussions.
Author contributions
Study concept by E.B, initial draft, phylogenetic analyses and figures by B.B, editing by B.B, R.S and E.B.
Funding
B.B, R.S and E.B are supported by the National Human Genome Research Institute of the National Institutes of Health [U24HG003345]. The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH.
Data availability
HGNC services are freely available from https://www.genenames.org/. HGNC code is available at the GitHub repository https://github.com/HGNC. The Fig. 3 ML phylogenetic tree and the pruned subtrees shown in Figs. 6, 7, 8, 9 and 10 can be viewed online using the shared project key https://itol.embl.de/shared/bbraschi in iTOL. All trees can be fully expanded in iTOL. The edited multiple alignment used to construct this phylogeny is available at https://storage.googleapis.com/public-download-files/supplemental_data/A2M_paper/A2M_FINAL_clustalo_ready.fasta.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
EB serves as an Executive Associate Editor of the Human Genomics journal.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Tweedie S, Braschi B, Gray K, Jones TEM, Seal RL, Yates B, et al. Genenames.org: the HGNC and VGNC resources in 2021. Nucleic Acids Res. 2021;49:D939–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Bruford EA, Braschi B, Denny P, Jones TEM, Seal RL, Tweedie S. Guidelines for human gene nomenclature. Nat Genet. 2020;52:754–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Jones TEM, Yates B, Braschi B, Gray K, Tweedie S, Seal RL, et al. The VGNC: expanding standardized vertebrate gene nomenclature. Genome Biol. 2023;24:115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.MGI-Mouse Nomenclature Home Page. Available from: https://www.informatics.jax.org/mgihome/nomen/
- 5.Ringwald M, Richardson JE, Baldarelli RM, Blake JA, Kadin JA, Smith C, et al. Mouse genome informatics (MGI): latest news from MGD and GXD. Mamm Genome. 2022;33:4–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Laulederkind SJF, Hayman GT, Wang S-J, Kaldunski ML, Vedi M, Demos WM, et al. The rat genome database: Genetic, genomic, and phenotypic data across multiple species. Curr Protoc. 2023;3:e804. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Available from: https://rgd.mcw.edu/nomen/cnt_nomen.shtml
- 8.Burt DW, Carrë W, Fell M, Law AS, Antin PB, Maglott DR, et al. The chicken gene nomenclature committee report. BMC Genomics. 2009;10(Suppl 2):S5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Available from: http://birdgenenames.org/cgnc/index.jsp
- 10.Gene group. Available from: https://genenames.org/data/genegroup/#!/group/1234
- 11.Rawlings ND, Waller M, Barrett AJ, Bateman A. MEROPS: the database of proteolytic enzymes, their substrates and inhibitors. Nucleic Acids Res. 2014;42:D503–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Janssen BJC, Huizinga EG, Raaijmakers HCA, Roos A, Daha MR, Nilsson-Ekdahl K, et al. Structures of complement component C3 provide insights into the function and evolution of immunity. Nature. 2005;437:505–11. [DOI] [PubMed] [Google Scholar]
- 13.Mortensen S, Kidmose RT, Petersen SV, Szilágyi Á, Prohászka Z, Andersen GR. Structural basis for the function of complement component C4 within the classical and lectin pathways of complement. J Immunol. 2015;194:5488–96. [DOI] [PubMed] [Google Scholar]
- 14.Foley JH, Peterson EA, Lei V, Wan LW, Krisinger MJ, Conway EM. Interplay between fibrinolysis and complement: plasmin cleavage of iC3b modulates immune responses. J Thromb Haemost. 2015;13:610–8. [DOI] [PubMed] [Google Scholar]
- 15.Letunic I. SMART: Select the running mode. Available from: http://smart.embl-heidelberg.de/
- 16.Pathirana A, Diao M, Huang S, Zuo L, Liang Y. Alpha 2 macroglobulin is a maternally-derived immune factor in amphioxus embryos: new evidence for defense roles of maternal immune components in invertebrate chordate. Fish Shellfish Immunol. 2016;50:21–6. [DOI] [PubMed] [Google Scholar]
- 17.Buresova V, Hajdusek O, Franta Z, Sojka D, Kopacek P. IrAM-an alpha2-macroglobulin from the hard tick Ixodes ricinus: characterization and function in phagocytosis of a potential pathogen Chryseobacterium indologenes. Dev Comp Immunol. 2009;33:489–98. [DOI] [PubMed] [Google Scholar]
- 18.Ma H, Wang B, Zhang J, Li F, Xiang J. Multiple forms of alpha-2 macroglobulin in shrimp Fenneropenaeus Chinesis and their transcriptional response to WSSV or vibrio pathogen infection. Dev Comp Immunol. 2010;34:677–84. [DOI] [PubMed] [Google Scholar]
- 19.Borisova EA, Gorbushin AM. Molecular cloning of α-2-macroglobulin from hemocytes of common periwinkle littorina littorea. Fish Shellfish Immunol. 2014;39:136–7. [DOI] [PubMed] [Google Scholar]
- 20.Qian J, Ren C, Xia J, Chen T, Yu Z, Hu C. Discovery, structural characterization and functional analysis of alpha-2-macroglobulin, a novel immune-related molecule from holothuria Atra. Gene. 2016;585:205–15. [DOI] [PubMed] [Google Scholar]
- 21.Jiang D, Shao Y, Zhang S, Li C. A2M possesses anti-bacterial functions by recruiting and enhancing phagocytosis through GRP78 in an echinoderm. Int J Biol Macromol. 2024;265:131016. [DOI] [PubMed] [Google Scholar]
- 22.Armstrong PB, Quigley JP. Limulus alpha 2-macroglobulin. First evidence in an invertebrate for a protein containing an internal thiol ester bond. Biochem J. 1987;248:703–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Melchior R, Quigley JP, Armstrong PB. Alpha 2-macroglobulin-mediated clearance of proteases from the plasma of the American horseshoe crab, Limulus polyphemus. J Biol Chem. 1995;270:13496–502. [DOI] [PubMed] [Google Scholar]
- 24.Doan N, Gettins PGW. alpha-Macroglobulins are present in some gram-negative bacteria: characterization of the alpha2-macroglobulin from Escherichia coli. J Biol Chem. 2008;283:28747–56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Garcia-Ferrer I, Marrero A, Gomis-Rüth FX, Goulas T. α2-macroglobulins: structure and function. Subcell Biochem. 2017;83:149–83. [DOI] [PubMed] [Google Scholar]
- 26.Rehman AA, Ahsan H, Khan FH. α-2-macroglobulin: a physiological guardian. J Cell Physiol. 2013;228:1665–75. [DOI] [PubMed] [Google Scholar]
- 27.Armstrong PB, Quigley JP. Alpha2-macroglobulin: an evolutionarily conserved arm of the innate immune system. Dev Comp Immunol. 1999;23:375–90. [DOI] [PubMed] [Google Scholar]
- 28.Wyatt AR, Cater JH, Ranson M. PZP and PAI-2: structurally-diverse, functionally similar pregnancy proteins? Int J Biochem Cell Biol. 2016;79:113–7. [DOI] [PubMed] [Google Scholar]
- 29.Tayade C, Esadeg S, Fang Y, Croy BA. Functions of alpha 2 macroglobulins in pregnancy. Mol Cell Endocrinol. 2005;245:60–6. [DOI] [PubMed] [Google Scholar]
- 30.Cater JH, Wilson MR, Wyatt AR. Alpha-2-Macroglobulin, a hypochlorite-regulated chaperone and immune system modulator. Oxid Med Cell Longev. 2019;2019:5410657. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Devriendt K, Zhang J, van Leuven F, van den Berghe H, Cassiman JJ, Marynen P. A cluster of alpha 2-macroglobulin-related genes (alpha 2 M) on human chromosome 12p: cloning of the pregnancy-zone protein gene and an alpha 2M pseudogene. Gene. 1989;81:325–34. [DOI] [PubMed] [Google Scholar]
- 32.Nepomuceno AI, Shao H, Jing K, Ma Y, Petitte JN, Idowu MO, et al. In-depth LC-MS/MS analysis of the chicken ovarian cancer proteome reveals conserved and novel differentially regulated proteins in humans. Anal Bioanal Chem. 2015;407:6851–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Goldfarb T, Kodali VK, Pujar S, Brover V, Robbertse B, Farrell CM, et al. NCBI refseq: reference sequence standards through 25 years of curation and annotation. Nucleic Acids Res. 2025;53:D243–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Mudge JM, Carbonell-Sala S, Diekhans M, Martinez JG, Hunt T, Jungreis I, et al. GENCODE 2025: reference gene annotation for human and mouse. Nucleic Acids Res. 2025;53:D966–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Huang Y-X, Song H, Tao Y, Shao X-B, Zeng X-S, Xu X-L, et al. Ovostatin 2 knockdown significantly inhibits the growth, migration, and tumorigenicity of cutaneous malignant melanoma cells. PLoS One. 2018;13:e0195610. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Vandooren J, Itoh Y. Alpha-2-macroglobulin in inflammation, immunity and infections. Front Immunol. 2021;12:803244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Marrero A, Duquerroy S, Trapani S, Goulas T, Guevara T, Andersen GR, et al. The crystal structure of human α2-macroglobulin reveals a unique molecular cage. Angew Chem Int Ed Engl. 2012;51:3340–4. [DOI] [PubMed] [Google Scholar]
- 38.Skornicka EL, Kiyatkina N, Weber MC, Tykocinski ML, Koo PH. Pregnancy zone protein is a carrier and modulator of placental protein-14 in T-cell growth and cytokine production. Cell Immunol. 2004;232:144–56. [DOI] [PubMed] [Google Scholar]
- 39.Strickland DK, Ashcom JD, Williams S, Burgess WH, Migliorini M, Argraves WS. Sequence identity between the alpha 2-macroglobulin receptor and low density lipoprotein receptor-related protein suggests that this molecule is a multifunctional receptor. J Biol Chem. 1990;265:17401–4. [PubMed] [Google Scholar]
- 40.Theret L, Jeanne A, Langlois B, Hachet C, David M, Khrestchatisky M, et al. Identification of LRP-1 as an endocytosis and recycling receptor for β1-integrin in thyroid cancer cells. Oncotarget. 2017;8:78614–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Rauch JN, Luna G, Guzman E, Audouard M, Challis C, Sibih YE, et al. LRP1 is a master regulator of tau uptake and spread. Nature. 2020;580:381–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Huang Z, Zhang P, Chen R, Sun L, Wang J, Yan R, et al. Targeting A2M-LRP1 reverses uterine spiral artery remodeling disorder and alleviates the progression of preeclampsia. Cell Commun Signal. 2025;23:107. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Laberge A, Ayoub A, Arif S, Larochelle S, Garnier A, Moulin VJ. α-2-Macroglobulin induces the shedding of microvesicles from cutaneous wound myofibroblasts. J Cell Physiol. 2019;234:11369–79. [DOI] [PubMed] [Google Scholar]
- 44.Galliano M-F, Toulza E, Jonca N, Gonias SL, Serre G, Guerrin M. Binding of alpha2ML1 to the low density lipoprotein receptor-related protein 1 (LRP1) reveals a new role for LRP1 in the human epidermis. PLoS One. 2008;3:e2729. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Umans L, Serneels L, Overbergh L, Stas L, Van Leuven F. Alpha2-macroglobulin- and murinoglobulin-1- deficient mice. A mouse model for acute pancreatitis. Am J Pathol. 1999;155:983–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Hochepied T, Ameloot P, Brouckaert P, Van Leuven F, Libert C. Differential response of a(2)-macroglobulin-deficient mice in models of lethal TNF-induced inflammation. Eur Cytokine Netw. 2000;11:597–601. [PubMed] [Google Scholar]
- 47.Westwood M, Aplin JD, Collinge IA, Gill A, White A, Gibson JM. Alpha 2-macroglobulin: a new component in the insulin-like growth factor/insulin-like growth factor binding protein-1 axis. J Biol Chem. 2001;276:41668–74. [DOI] [PubMed] [Google Scholar]
- 48.Krimbou L, Tremblay M, Davignon J, Cohn JS. Association of apolipoprotein E with alpha2-macroglobulin in human plasma. J Lipid Res. 1998;39:2373–86. [PubMed] [Google Scholar]
- 49.Mettenburg JM, Webb DJ, Gonias SL. Distinct binding sites in the structure of alpha 2-macroglobulin mediate the interaction with beta-amyloid peptide and growth factors. J Biol Chem. 2002;277:13338–45. [DOI] [PubMed] [Google Scholar]
- 50.Borth W, Luger TA. Identification of alpha 2-macroglobulin as a cytokine binding plasma protein. Binding of interleukin-1 beta to F alpha 2-macroglobulin. J Biol Chem. 1989;264:5818–25. [PubMed] [Google Scholar]
- 51.Chensue SW, Terebuh PD, Remick DG, Scales WE, Kunkel SL. In vivo biologic and immunohistochemical analysis of interleukin-1 alpha, beta and tumor necrosis factor during experimental endotoxemia. Kinetics, Kupffer cell expression, and glucocorticoid effects. Am J Pathol. 1991;138:395–402. [PMC free article] [PubMed] [Google Scholar]
- 52.Webb DJ, Gonias SL. A modified human alpha 2-macroglobulin derivative that binds tumor necrosis factor-alpha and interleukin-1 beta with high affinity in vitro and reverses lipopolysaccharide toxicity in vivo in mice. Lab Invest. 1998;78:939–48. [PubMed] [Google Scholar]
- 53.Gourine AV, Gourine VN, Tesfaigzi Y, Caluwaerts N, Van Leuven F, Kluger MJ. Role of alpha(2)-macroglobulin in fever and cytokine responses induced by lipopolysaccharide in mice. Am J Physiol Regul Integr Comp Physiol. 2002;283:R218-26. [DOI] [PubMed] [Google Scholar]
- 54.Smithies O. Zone electrophoresis in starch gels and its application to studies of serum proteins. Adv Protein Chem. 1959;14:65–113. [DOI] [PubMed] [Google Scholar]
- 55.Lin TM, Halbert SP, Spellacy WN, Gall S. Human pregnancy-associated plasma proteins during the postpartum period. Am J Obstet Gynecol. 1976;124:382–7. [DOI] [PubMed] [Google Scholar]
- 56.Kashiwagi H, Ishimoto H, Izumi S-I, Seki T, Kinami R, Otomo A, et al. Human PZP and common marmoset A2ML1 as pregnancy related proteins. Sci Rep. 2020;10:5088. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Lin J, Jiang X, Dong M, Liu X, Shen Q, Huang Y, et al. Hepatokine pregnancy zone protein governs the diet-induced thermogenesis through activating brown adipose tissue. Adv Sci. 2021;8:e2101991. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Shao J, Jin Y, Shao C, Fan H, Wang X, Yang G. Serum exosomal pregnancy zone protein as a promising biomarker in inflammatory bowel disease. Cell Mol Biol Lett. 2021;26:36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Ijsselstijn L, Dekker LJM, Stingl C, van der Weiden MM, Hofman A, Kros JM, et al. Serum levels of pregnancy zone protein are elevated in presymptomatic Alzheimer’s disease. J Proteome Res. 2011;10:4902–10. [DOI] [PubMed] [Google Scholar]
- 60.Stimson WH. Identification of pregnancy-associated alpha-macroglobulin on the surface of peripheral blood leucocyte populations. Clin Exp Immunol. 1977;28:445–52. [PMC free article] [PubMed] [Google Scholar]
- 61.Christensen U, Sottrup-Jensen L, Simonsen M. Kinetics and mechanism of proteinase-binding of pregnancy zone protein (PZP). Appearance of sulfhydryl groups in reactions with proteinases. J Enzyme Inhib. 1992;5:269–79. [DOI] [PubMed] [Google Scholar]
- 62.Saidi N, Samel M, Siigur J, Jensen PE. Lebetase, an alpha(beta)-fibrin(ogen)olytic metalloproteinase of Vipera lebetina snake venom, is inhibited by human alpha-macroglobulins. Biochim Biophys Acta. 1999;1434:94–102. [DOI] [PubMed] [Google Scholar]
- 63.Pineda-Salgado L, Craig EJ, Blank RB, Kessler DS. Expression of Panza, an alpha2-macroglobulin, in a restricted dorsal domain of the primitive gut in xenopus laevis. Gene Expr Patterns. 2005;6:3–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Galliano M-F, Toulza E, Gallinaro H, Jonca N, Ishida-Yamamoto A, Serre G, et al. A novel protease inhibitor of the alpha2-macroglobulin family expressed in the human epidermis. J Biol Chem. 2006;281:5780–9. [DOI] [PubMed] [Google Scholar]
- 65.Miller HT, Feeney RE. The physical and chemical properties of an immunologically cross-reacting protein from avian egg whites. Biochemistry. 1966;5:952–8. [DOI] [PubMed] [Google Scholar]
- 66.Nielsen KL, Sottrup-Jensen L, Nagase H, Thøgersen HC, Etzerodt M. Amino acid sequence of Hen ovomacroglobulin (ovostatin) deduced from cloned cDNA. DNA Seq. 1994;5:111–9. [DOI] [PubMed] [Google Scholar]
- 67.Kitamoto T, Nakashima M, Ikai A. Hen egg white ovomacroglobulin has a protease inhibitory activity. J Biochem. 1982;92:1679–82. [DOI] [PubMed] [Google Scholar]
- 68.Ikai A, Kitamoto T, Nishigai M. Alpha-2-macroglobulin-like protease inhibitor from the egg white of Cuban crocodile (Crocodylus rhombifer). J Biochem. 1983;93:121–7. [DOI] [PubMed] [Google Scholar]
- 69.Arakawa H, Osada T, Ikai A. Unusual properties of crocodilian ovomacroglobulin shown in its methylamine treatment and sulfhydryl titration. Arch Biochem Biophys. 1986;244:447–53. [DOI] [PubMed] [Google Scholar]
- 70.Nagase H, Harris ED Jr, Brew K. Evidence for a thiol ester in duck ovostatin (ovomacroglobulin). J Biol Chem. 1986;261:1421–6. [PubMed] [Google Scholar]
- 71.Nagase H, Brew K. Amino acid sequence of a 32-residue region around the thiol ester site in Duck ovostatin. FEBS Lett. 1987;222:83–8. [DOI] [PubMed] [Google Scholar]
- 72.Osada T, Sasaki T, Ikai A. Purification and characterization of alpha-macroglobulin and ovomacroglobulin of the green turtle (Chelonia mydas japonica). J Biochem. 1988;103:212–7. [DOI] [PubMed] [Google Scholar]
- 73.Bernardo ADEM, Thorsteinsdóttir S, Mummery CL. Advantages of the avian model for human ovarian cancer. Mol Clin Oncol. 2015;3:1191–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Andrews Kingon GL, Petitte JN, Muddiman DC, Hawkridge AM. Multi-peptide nLC-PC-IDMS-SRM-based assay for the quantification of biomarkers in the chicken ovarian cancer model. Methods. 2013;61:323–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Madeira F, Madhusoodanan N, Lee J, Eusebi A, Niewielska A, Tivey ARN, et al. The EMBL-EBI job dispatcher sequence analysis tools framework in 2024. Nucleic Acids Res. 2024;52:W521–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Huang H-L, Li S-C, Wu J-F. A complex of novel protease inhibitor, ovostatin homolog, with its cognate proteases in immature mice uterine luminal fluid. Sci Rep. 2019;9:4973. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Li Z, Huang X, Tang Q, Ma M, Jin Y, Sheng L. Functional properties and extraction techniques of chicken egg white proteins. Foods. 2022;11:2434. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Nielson KL, Sottrup-Jensen NL, Nagase H, Etzerodt M. The primary structure of ovomacroglobulin. Ann N Y Acad Sci. 1994;737:476–9. [DOI] [PubMed] [Google Scholar]
- 79.Saito A, Sinohara H. Murinoglobulin, a novel protease inhibitor from murine plasma. Isolation, characterization, and comparison with murine alpha-macroglobulin and human alpha-2-macroglobulin. J Biol Chem. 1985;260:775–81. [PubMed] [Google Scholar]
- 80.Saito A, Sinohara H. Rat plasma murinoglobulin: isolation, characterization, and comparison with rat alpha-1- and alpha-2-macroglobulins. J Biochem. 1985;98:501–16. [DOI] [PubMed] [Google Scholar]
- 81.Iwasaki H, Suzuki Y, Sinohara H. Cloning and sequencing of cDNAs encoding plasma alpha-macroglobulin and murinoglobulin from Guinea pig: implications for molecular evolution of alpha-macroglobulin family. J Biochem. 1996;120:1167–75. [DOI] [PubMed] [Google Scholar]
- 82.Miyake Y, Shinomura M, Ito T, Yamamoto K, Abe K, Amemiya S, et al. Hamster alpha-macroglobulin and murinoglobulin: comparison of chemical and biological properties with homologs from other mammals. J Biochem. 1993;114:513–21. [DOI] [PubMed] [Google Scholar]
- 83.de Wit CA, Weström BR. Further studies of plasma protease inhibitors in the hedgehog, Erinaceus europaeus; collagenase, papain and plasmin inhibitors. Comp Biochem Physiol A Comp Physiol. 1987;86:1–5. [DOI] [PubMed] [Google Scholar]
- 84.Cheong S-S, Hentschel L, Davidson AE, Gerrelli D, Davie R, Rizzo R, et al. Mutations in CPAMD8 cause a unique form of autosomal-recessive anterior segment dysgenesis. Am J Human Genet. 2016. 10.1016/j.ajhg.2016.09.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Ribeiro Filho W, Sugiki M, Yoshida E, Maruyama M. Inhibition of hemorrhagic and edematogenic activities of snake venoms by a broad-spectrum protease inhibitor, murinoglobulin; the effect on venoms from five different genera in Viperidae family. Toxicon. 2003;42:173–81. [DOI] [PubMed] [Google Scholar]
- 86.Liu N, Lo LS-L, Askary SH, Jones L, Kidane TZ, Trang T, et al. Transcuprein is a macroglobulin regulated by copper and iron availability. J Nutr Biochem. 2007;18:597–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Moriya M, Ho Y-H, Grana A, Nguyen L, Alvarez A, Jamil R, et al. Copper is taken up efficiently from albumin and α2-macroglobulin by cultured human cells by more than one mechanism. Am J Physiol Cell Physiol. 2008;295:C708-21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Gu L, Guo Z. Alzheimer’s Aβ42 and Aβ40 peptides form interlaced amyloid fibrils. J Neurochem. 2013;126:305–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Kovacs DM. Alpha2-macroglobulin in late-onset alzheimer’s disease. Exp Gerontol. 2000;35:473–9. [DOI] [PubMed] [Google Scholar]
- 90.Chen H, Li Z, Liu N, Zhang W, Zhu G. Influence of alpha-2-macroglobulin 5 bp I/D and Ile1000Val polymorphisms on the susceptibility of alzheimer’s disease: a systematic review and meta-analysis of 52 studies. Cell Biochem Biophys. 2014;70:511–9. [DOI] [PubMed] [Google Scholar]
- 91.Hoogendoorn H, Toh CH, Nesheim ME, Giles AR. Alpha 2-macroglobulin binds and inhibits activated protein C. Blood. 1991;78:2283–90. [PubMed] [Google Scholar]
- 92.Lagrange J, Lecompte T, Knopp T, Lacolley P, Regnault V. Alpha-2-macroglobulin in hemostasis and thrombosis: an underestimated old double-edged sword. J Thromb Haemost. 2022;20:806–15. [DOI] [PubMed] [Google Scholar]
- 93.Nezu T, Hosomi N, Aoki S, Deguchi K, Masugata H, Ichihara N, et al. Alpha2-macroglobulin as a promising biomarker for cerebral small vessel disease in acute ischemic stroke patients. J Neurol. 2013;260:2642–9. [DOI] [PubMed] [Google Scholar]
- 94.Larionov S, Dedeck O, Birkenmeier G, Thal DR. Expression of alpha2-macroglobulin, neutrophil elastase, and interleukin-1alpha differs in early-stage and late-stage atherosclerotic lesions in the arteries of the circle of Willis. Acta Neuropathol. 2007;113:33–43. [DOI] [PubMed] [Google Scholar]
- 95.Larionov S, Dedeck O, Birkenmeier G, Orantes M, Ghebremedhin E, Thal DR. The intronic deletion polymorphism of the Alpha2- macroglobulin gene modulates the severity and extent of atherosclerosis in the circle of Willis. Neuropathol Appl Neurobiol. 2006;32:451–4. [DOI] [PubMed] [Google Scholar]
- 96.Cater JH, Kumita JR, Zeineddine Abdallah R, Zhao G, Bernardo-Gancedo A, Henry A, et al. Human pregnancy zone protein stabilizes misfolded proteins including preeclampsia- and Alzheimer’s-associated amyloid beta peptide. Proc Natl Acad Sci U S A. 2019;116:6101–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.Verbeek MM, Ruiter DJ, de Waal RM. The role of amyloid in the pathogenesis of Alzheimer’s disease. Biol Chem. 1997;378:937–50. [DOI] [PubMed] [Google Scholar]
- 98.Vissers LELM, Bonetti M, Paardekooper Overman J, Nillesen WM, Frints SGM, de Ligt J, et al. Heterozygous germline mutations in A2ML1 are associated with a disorder clinically related to Noonan syndrome. Eur J Hum Genet. 2015;23:317–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Brinkmann J, Lissewski C, Pinna V, Vial Y, Pantaleoni F, Lepri F, et al. The clinical significance of A2ML1 variants in Noonan syndrome has to be reconsidered. Eur J Hum Genet. 2021;29:524–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Mittal R, Sanchez-Luege SV, Wagner SM, Yan D, Liu XZ. Recent perspectives on gene-microbe interactions determining predisposition to otitis media. Front Genet. 2019;10:1230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Geng R, Wang Q, Chen E, Zheng QY. Current understanding of host genetics of otitis media. Front Genet. 2019;10:1395. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Larson ED, Magno JPM, Steritz MJ, Llanes EGDV, Cardwell J, Pedro M, et al. A2ML1 and otitis media: novel variants, differential expression, and relevant pathways. Hum Mutat. 2019;40:1156–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Schepens I, Jaunin F, Begre N, Läderach U, Marcus K, Hashimoto T, et al. The protease inhibitor alpha-2-macroglobulin-like-1 is the p170 antigen recognized by paraneoplastic pemphigus autoantibodies in human. PLoS One. 2010;5:e12250. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Numata S, Teye K, Tsuruta D, Sogame R, Ishii N, Koga H, et al. Anti-α-2-macroglobulin-like-1 autoantibodies are detected frequently and may be pathogenic in paraneoplastic pemphigus. J Invest Dermatol. 2013;133:1785–93. [DOI] [PubMed] [Google Scholar]
- 105.Poot AM, Diercks GFH, Kramer D, Schepens I, Klunder G, Hashimoto T, et al. Laboratory diagnosis of paraneoplastic pemphigus. Br J Dermatol. 2013;169:1016–24. [DOI] [PubMed] [Google Scholar]
- 106.Ohzono A, Sogame R, Li X, Teye K, Tsuchisaka A, Numata S, et al. Clinical and immunological findings in 104 cases of paraneoplastic pemphigus. Br J Dermatol. 2015;173:1447–52. [DOI] [PubMed] [Google Scholar]
- 107.Bazzini C, Begré N, Favre B, Hashimoto T, Hertl M, Schlapbach C, et al. Detection of autoantibodies against alpha-2-macroglobulin-like 1 in paraneoplastic pemphigus Sera utilizing novel green fluorescent protein-based immunoassays. J Dermatol Sci. 2020;98:173–8. [DOI] [PubMed] [Google Scholar]
- 108.Seitz R, Gürtler L, Schramm W. Thromboinflammation in COVID-19: can α2 -macroglobulin help to control the fire? J Thromb Haemost. 2021;19:351–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Oguntuyo KY, Stevens CS, Siddiquey MN, Schilke RM, Woolard MD, Zhang H et al. In plain sight: the role of alpha-1-antitrypsin in COVID-19 pathogenesis and therapeutics. bioRxiv. 2020; 10.1101/2020.08.14.248880.
- 110.Yang Z, Rannala B. Molecular phylogenetics: principles and practice. Nat Rev Genet. 2012;13:303–14. [DOI] [PubMed] [Google Scholar]
- 111.Letunic I, iTOL. Shared projects for user bbraschi. Available from: https://itol.embl.de/shared/bbraschi
- 112.ZFIN The Zebrafish Information Network. Available from: https://zfin.org/
- 113.Shine R. Reptiles. Curr Biol. 2013;23:R227-31. [DOI] [PubMed] [Google Scholar]
- 114.Huang Y-X, Qi J, Wang H-S, Shao X-B, Zeng X-S, Li A-M, et al. Expression analysis of ovostatin 2 reveals its involvement in proliferation, invasion and angiogenesis of cutaneous malignant melanoma. J Dermatol. 2013;40:901–10. [DOI] [PubMed] [Google Scholar]
- 115.Jones DT, Taylor WR, Thornton JM. The rapid generation of mutation data matrices from protein sequences. Comput Appl Biosci. 1992;8:275–82. [DOI] [PubMed] [Google Scholar]
- 116.Esadeg S, He H, Pijnenborg R, Van Leuven F, Croy BA. Alpha-2 macroglobulin controls trophoblast positioning in mouse implantation sites. Placenta. 2003;24:912–21. [DOI] [PubMed] [Google Scholar]
- 117.Porter TM, Hajibabaei M. Profile hidden Markov model sequence analysis can help remove putative pseudogenes from DNA barcoding and metabarcoding datasets. BMC Bioinformatics. 2021;22:256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 118.Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215:403–10. [DOI] [PubMed] [Google Scholar]
- 119.Boratyn GM, Camacho C, Cooper PS, Coulouris G, Fong A, Ma N, et al. BLAST: a more efficient report with usability improvements. Nucleic Acids Res. 2013;41:W29–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Dyer SC, Austine-Orimoloye O, Azov AG, Barba M, Barnes I, Barrera-Enriquez VP, et al. Ensembl 2025. Nucleic Acids Res. 2025;53:D948–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. Trimal: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25:1972–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 122.Larsson A. AliView: a fast and lightweight alignment viewer and editor for large datasets. Bioinformatics. 2014;30:3276–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 123.Trifinopoulos J, Nguyen L-T, von Haeseler A, Minh BQ. W-IQ-TREE: a fast online phylogenetic tool for maximum likelihood analysis. Nucleic Acids Res. 2016;44:W232–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Minh BQ, Nguyen MAT, von Haeseler A. Ultrafast approximation for phylogenetic bootstrap. Mol Biol Evol. 2013;30:1188–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 125.Letunic I, Bork P. Interactive tree of life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool. Nucleic Acids Res. 2024;52:W78–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 126.HUGO Gene Nomenclature Committee. Available from: https://www.genenames.org/
- 127.HUGO Gene Nomenclature Committee. Github; Available from: https://github.com/HGNC
- 128.Available from: https://storage.googleapis.com/public-download-files/supplemental_data/A2M_paper/A2M_FINAL_clustalo_ready.fasta
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
HGNC services are freely available from https://www.genenames.org/. HGNC code is available at the GitHub repository https://github.com/HGNC. The Fig. 3 ML phylogenetic tree and the pruned subtrees shown in Figs. 6, 7, 8, 9 and 10 can be viewed online using the shared project key https://itol.embl.de/shared/bbraschi in iTOL. All trees can be fully expanded in iTOL. The edited multiple alignment used to construct this phylogeny is available at https://storage.googleapis.com/public-download-files/supplemental_data/A2M_paper/A2M_FINAL_clustalo_ready.fasta.








