Abstract
Homo erectus remains have been found in Africa, Eurasia and Southeast Asia1–3, dating back around two million years; however, owing to their age and state of preservation, obtaining informative molecular data from them has proved challenging. Here we successfully extracted and analysed ancient enamel proteins from five male and one female Middle Pleistocene H. erectus specimens from approximately 0.4 million years ago, from the Zhoukoudian, Hexian and Sunjiadong sites. All specimens from all three sites share two amino acid variants. Of these, A253G in AMBN is previously unknown and has not been identified in other human lineages, including H. erectus from Dmanisi (Georgia), Homo antecessor from Atapuerca (Spain), Denisovans, Neanderthals and modern humans. The other variant, AMBN(M273V), has previously been identified in Denisovans, and our evidence now indicates it may have been introduced through populations related to these Middle Pleistocene H. erectus. The regions in the Denisovan genome attributed to super-archaic introgression, some of which later passed to modern humans, are likely to have originated from H. erectus. Late Middle Pleistocene H. erectus may have coexisted with Denisovans in parts of East Asia, where these interactions are presumed to have occurred.
Subject terms: Proteomics, Palaeontology
Palaeoproteomic analysis of ancient enamel proteins extracted from Middle Pleistocene Homo erectus specimens from the Zhoukoudian, Hexian and Sunjiadong sites in China suggests that they are a new genetic monogroup, and super-archaic introgression in Denisovans is likely to have originated from H. erectus.
Main
Broadly defined H. erectus (H. erectus sensu lato) was widely distributed across three continents and was geologically long-lived1–3, lasting nearly 2 million years (Myr). Homo erectus is thought to have spread from East Africa into western Asia (1.8 Myr ago (Ma), Dmanisi, Georgia), marking the earliest generally accepted evidence of the genus Homo outside of Africa4,5. The earliest H. erectus specimen found in Europe is from the TE7 level of the Sima del Elefante site (Sierra de Atapuerca, Spain)6, dated between 1.4 Ma and 1.1 Ma. After dispersing to China and Indonesia, H. erectus persisted in these regions for a considerable geological duration. In China, the record of H. erectus dates back to about 2.1–1.6 Ma (refs. 7–9) and disappears roughly 0.4–0.3 Ma (ref. 10). In Java, Indonesia, the fossils of H. erectus range from 1.5 Ma to as recent as 0.1 Ma. Overall, H. erectus has had an important role in human evolution. Understanding the molecular characteristics of this lineage is essential to help us to better understand hominin biology throughout the genus Homo. Nevertheless, the only molecular data previously recovered from H. erectus comprise peptides extracted from a 1.77-Myr-old tooth originating from Dmanisi in Georgia, West Asia. However, these sequences lack any single amino acid polymorphisms (SAPs) that can distinguish H. erectus from other human lineages11.
The H. erectus specimen represented by the Zhoukoudian fossils dates from approximately 0.78 to 0.3 Ma during the Middle Pleistocene in northern China12,13. It is frequently used as the primary reference for the overall morphological characteristics of all H. erectus or Asian H. erectus14. This group has had a crucial role in research on the origin and evolution of H. erectus. Studies of the Zhoukoudian fossils have settled debates over whether the Indonesian Java fossils are apes or humans, confirming the systematic position of H. erectus within the human lineage14,15. In addition to the Zhoukoudian fossils, several other Middle Pleistocene H. erectus fossils have been discovered in China, including remains from Hexian, dated to about 0.41 Ma (ref. 16), from Nanjing, dated to at least 0.58 Ma (ref. 17), and from Sunjiadong18, dated around 0.4 Ma (ref. 19). Nanjing and Sunjiadong fossils show similar morphology to the Zhoukoudian fossils, whereas Hexian remains are morphologically closer to Indonesian Java fossils than to Zhoukoudian fossils14,15,18,20. Additionally, some affinities are shared with a Middle–Late Pleistocene mandible from the Penghu Channel in Taiwan, which is purported to be a Denisovan21. Retrieving proteins from more than one of these fossils would help to determine whether they share common molecular characteristics that can distinguish them from others and provide insights into Middle Pleistocene H. erectus genetic material.
The fossils used in the protein experiments and analyses of this study are from Zhoukoudian, Hexian and Sunjiadong sites (Fig. 1 and Supplementary Note 1). Although the human fossils excavated at Zhoukoudian22 in Beijing, northern China between 1927 and 1937 were lost, some human fossils were discovered at this site between 1949 and 1951, and again in 1966. The human tooth (PA69)23 used in this study is from layers 8–9 (0.42 Ma (refs. 24,25)), found between 1949 and 1951. We also studied two human teeth from the Hexian site26 in Anhui Province, southern China (0.41 Ma (ref. 16)) and three teeth from Sunjiadong (around 0.4 Ma (ref. 19)) in Henan Province, northern China18. Two of the teeth (12SJD1#10-42 and 12SJD1#10-45) at Sunjiadong are from layer 4, and the other tooth (12SJD1#14-22) is from layer 5 (ref. 18). Altogether, the dating of all 6 samples from these 3 sites yielded an approximate age of 0.4 Myr, making them roughly contemporary and representing individuals from both northern and southern China. Additionally, we included a tooth from an earlier-reported Denisovan skull (minimum age 0.15 Myr) from Harbin, China27,28.
Fig. 1. The geographic locations and samples of the Middle Pleistocene H. erectus sites used in this study.
a, The geographic locations of Pleistocene H. erectus sites in China, with sites in the present study highlighted in red. Background raster: Natural Earth, Natural Earth II with Shaded Relief and Water (1:50 m), v3.2.0; public domain (https://www.naturalearthdata.com/downloads/10m-natural-earth-2/10m-natural-earth-ii-with-shaded-relief-and-water/). b, Stratigraphy of the Zhoukoudian site22. The Zhoukoudian tooth, ZKD (PA69) is from layers 8–9. c, Stratigraphy of the Sunjiadong site18. The Sunjiadong 12SJD1#10-42 (SJD10M) and Sunjiadong 12SJD1#10-45 (SJD10A) teeth are from layer 4, and Sunjiadong 112SJD1#14-22 (SJD14M) is from layer 5. d, Stratigraphy of the Hexian site26. Hexian sample 1 (HX-S1) and Hexian sample 2 (HX-S2) are from layer 4 (ref. 16).
Preservation of animal fossil proteins
By first examining protein preservation in animal fossils from Zhoukoudian, Hexian and Sunjiadong, we gained initial insights into the feasibility of the analysis without damaging the ancient human fossils. Tests using matrix-assisted laser desorption/ionization-time-of-flight (MALDI-TOF) mass spectrometry revealed no evidence of ancient proteins in animal dentin or bones from 8 samples at Zhoukoudian, 40 samples from Hexian, and 37 samples from Sunjiadong (Supplementary Data 1). However, enamel from 3, 14 and 5 animal teeth from these respective sites did show protein signals (Supplementary Data 1). The successful retrieval of proteomic data from animal enamel suggests that extracting enamel proteins from human teeth would be feasible.
The endogenous enamel proteome
Ancient enamel proteins were successfully extracted using minimally invasive sampling techniques from six H. erectus teeth from Zhoukdoudian, Hexian and Sunjiadong sites, and a tooth from Harbin, China (Methods). The database used to search for endogenous peptide sequences included enamel proteins from translated genomes and published proteomes of present-day humans and other hominids11,29–31 with spectrum data from liquid chromatography–tandem mass spectrometry (LC–MS/MS) analysed using MaxQuant32, PEAKS Online33, and pFind34 (Methods, Supplementary Note 2). These approaches generated between 4,097 and 22,107 peptide–spectrum matches (PSMs), and between 652 and 3,476 peptides (Extended Data Table 1 and Supplementary Tables 1 and 2) across 7 specimens. Peptides matching multiple genes were considered non-unique and excluded from further analysis.
We then evaluated the endogeneity of these enamel proteomes based on several criteria (Supplementary Note 3): (1) the protein and amino acid composition profiles across all seven human samples were similar and consistent with other published ancient enamel proteins35, with minimal exogenous contamination (Fig. 2a, Extended Data Table 2 and Supplementary Tables 3 and 4); (2) peptide lengths and levels of peptide bond hydrolysis were typical of other ancient enamel proteins11,35 (Extended Data Table 1, Supplementary Figs. 1–3 and Supplementary Table 5); and (3) elevated deamidation levels (Fig. 2b and Supplementary Tables 6 and 7) suggest more extensive protein degradation and help to authenticate the ancient proteome11,36. These criteria resulted in 650–3,457 peptides recovered from 6–11 enamel-related proteins (for example, AMEL, AHSG, ALB, AMBN, AMTN, COL17A1, ENAM, KLK4, MMP20, ODAM and SERPINC1) for the 7 samples (Extended Data Table 1 and Supplementary Table 1). The PSM coverage of proteins is shown in Fig. 3a and Supplementary Fig. 4. In total, the consensus sequences of endogenous proteins from each sample were constructed by retaining the overlapping amino acid coverage in PEAKS Online, pFind and MaxQuant (Methods and Supplementary Note 4), which covered 269 to 903 amino acid positions across the 6 H. erectus samples and the Harbin individual.
Fig. 2. Ancient enamel protein preservation and sex identification of six Middle Pleistocene H. erectus specimens and the Harbin individual.
a, Number of endogenous peptides after filtering. b, Overall deamidation rates of N and Q residues in the seven new samples and published samples (Penghu 1 and Atapuerca_ATD6–92). MH represents the enamel results of modern humans43. c, The observed ratio (RY) of peptides covering the AMELY-specific position relative to all peptides from both AMELY and AMELX in seven samples of this study (red circles), 4 Paranthropus specimens and 4 known archaic humans (light blue circles), and 16 ancient and 4 present-day modern humans of known sex (orange circles)38,39 labelled as F_ (female) or M_ (male).
Extended Data Table 2.
Protein composition and coverage details for all six H. erectus samples, the Harbin tooth, and blank controls (_BK)
Total proteins recovered from all samples, including possible contaminants, along with the percentage of endogenous proteins. The peptide count for AMEL includes peptides from both AMELX and AMELY. The main proteins in “Others” are typical lab contaminants such as keratins and trypsin.
Extended Data Table 1.
Enamel data information for all six H. erectus samples and the Harbin tooth
Data is provided before and after application of the filtering criteria described in the text. Consensus positions are listed for filtered data.
Fig. 3. PSM coverage and supporting peptides for SAP AMBN(A253G) and AMBN(M273V) in the six H. erectus samples and the Harbin individual.
a, PSM coverage of AMBN. b, The alignments of multiple high coverage human sequences, and the small subset of supporting peptides and fragment ions for the SAPs at AMBN 253 and AMBN 273 in the new samples. e, N-terminal pyroglutamic acid (pyro-Glu) from E; m, oxidated M in ZKD and SJD10M, SJD10A, SJD14M, or dioxidated M in HX-S1 and HX-S2; n, deamidated N; q, deamidated Q (not N-terminal Q) or pyro-Glu from Q (N-terminal Q); y, dichlorinated Y in HX-S1, or chlorinated Y in HX-S2. Peptides marked with an asterisk were from pFind, and the others were from PEAKS. The alignment presents a small subset of the supporting peptides.
Sex determination
The male-specific marker amelogenin Y isoform (AMELY) is a gene located on the Y chromosome that is involved in tooth enamel development, and its presence is used to determine male sex37. However, owing to the restricted database for enamel and the high sensitivity of the searching software, spurious detection of AMELY-specific peptides may result in the false-positive assignment of samples as male38,39. This outcome is not unexpected, as previous research has demonstrated cross-species proteomic effects resulting from search database bias40,41. To address this issue, we used rigorous filtering criteria to ensure the reliability of all identified SAPs in the endogenous proteome (details below), including those used to determine sex in amelogenin.
To further validate the sex determination, we established protSexInferer42, a method based on the ratio of peptide counts covering the AMELY-specific positions relative to all peptides from both AMELY and AMELX (RY). Using 16 ancient and 4 present-day modern humans38,39, along with 4 Paranthropus specimens (SK14132, SK830, SK835 and SK850)2 and 4 archaic hominins (Penghu 1 (ref. 35), Dmanisi (D4163) and Atapuerca (ATD6-92)11 and Laos_TNH243) with confirmed sex, samples can be classified as male if RY exceeds 0.058, and female if RY is below 0.024 (Fig. 2c). Using this approach, the Zhoukoudian sample, the two Sunjiadong, two Hexian samples and the Harbin sample with RY from 0.187–0.295 were identified as males (Fig. 2c). The remaining Sunjiadong SJD10A (12SJD1#10-45, an adult), with RY of 0.019, was classified as female (Fig. 2c). The sex determination results were confirmed by the intensity method described in ref. 31 (Extended Data Fig. 1).
Extended Data Fig. 1. The intensity of AMELY-59M site as a function of the intensity of AMELX-60 site.

The Y-axis and X-axis represent the log2-transformed sum of intensities of all peptides covering the AMELY-59M and AMELX-60 sites, respectively. Black dots denote reference samples of known sex; red dots denote new samples.
SAPs
To explore the relationships between these specimens and other humans and primates, we identified peptide counts with amino acid substitutions that are specific to modern humans, Neanderthals, Denisovans, chimpanzees, gorillas and orangutans across all nine endogenous proteins. These required at least two peptides per amino acid allele, derived from different mass spectrometry instruments or laboratories. For all SAPs recovered from the endogenous proteomes, the tandem mass (MS/MS) spectra linked to the relevant PSMs were manually inspected to confirm the presence of typical fragmentation patterns of high-energy collision dissociation, the fragmentation method used in our mass spectrometry. Furthermore, at least three out of four y/b ions bracketing the SAP site were required to ensure high-quality confirmation. SAPs at the C or N terminus of the peptide were accepted if one related PSM contained a single y or b ion or more, as these could only produce one y ion and one b ion at most. All SAPs were confirmed in this way using PEAKS Online, pFind and MaxQuant (Methods), which resulted in two confidently identified SAPs.
A new SAP in the ameloblastin protein (AMBN) (A253>G), was identified in all 6 human samples from Zhoukoudian, Hexian, and Sunjiadong, with PSM counts of 124, 3, 9, 191, 24 and 145, and peptide counts of 21, 2, 4, 39, 9 and 35 supporting each sample, respectively (Table 1 and Supplementary Note 5). The highest-quality peptide for this SAP in each of the six samples has at least three y/b ions (Fig. 3b and Extended Data Figs. 2 and 4). Comparative samples with available data covering this position—including mammals, all primates, Homo antecessor from Atapuerca11, H. erectus from Dmanisi11, Denisovan 3, Penghu 1 (ref. 35), all Neandertals and Human Genome Diversity Project (HGDP) modern humans—nearly all carry AMBN A253, with a few mammals having T (horse), S (pig), V (lesser mole-rats) or missing (bovines) residues. Thus, AMBN(A253G), identified in all six human samples from three distinct sites across northern and southern China, representing vastly different geographic regions, has not been previously detected in any other primate species, suggesting that it is likely to be a specific variant for these Middle Pleistocene H. erectus. This mutation causes these six Middle Pleistocene H. erectus samples from Zhoukoudian, Hexian and Sunjiadong to cluster together in the phylogenetic tree (Supplementary Note 6) (posterior probability 100%; Extended Data Fig. 3). These results also indicate that the Hexian teeth can be attributed to H. erectus, and are not Denisovan as has been previously proposed based on morphology21.
Table 1.
The details of SAPs within Homo identified in six Middle Pleistocene H. erectus specimens and other individuals
| SAP assignment | East Asian H. erectus (new) | Denisovan-related | Confirmation | |||||
|---|---|---|---|---|---|---|---|---|
| Protein (UniProt accession) | AMBN (Q9NP70) | |||||||
| Position | 253 | 273 | PEAKS | pFind | MaxQuant | |||
| Data | Species | Sample | Age (Myr) | |||||
| New | H. erectus | ZKD | 0.42 | Ga (124;21)b | Va (252;52) | ✓ | ✓ | ✓ |
| New | H. erectus | SJD10M | 0.4 | Ga (191;39) | Va (219;75) | ✓ | ✓ | ✓ |
| New | H. erectus | SJD10A | 0.4 | Ga (24;9) | Va (87;46) | ✓ | ✓ | ✓ |
| New | H. erectus | SJD14M | 0.4 | Ga (145;35) | Va (211;69) | ✓ | ✓ | ✓ |
| New | H. erectus | HX-S1 | 0.41 | Ga (3;2) | Va (4;3) | ✓ | ✓ | M273V only (1;1) |
| New | H. erectus | HX-S2 | 0.41 | Ga (9;4) | Va (8;6) | ✓ | ✓ | M273V only (1;1) |
| Published11 | H. erectus | D4163 | 1.77 | A | M | ✓ | × | ✓ |
| Published11 | H. antecessor | ATD6-92 | 0.772–0.949 | A | M | ✓ | × | ✓ |
| Published43 | Possible Denisovan | TNH-2 | 0.131–0.164 | A | M | ✓ | × | ✓ |
| New | Denisovan | Harbin | 0.146–0.309 | A (276;58) | Va (211;79) M (49;38) | ✓ | ✓ | ✓ |
| Published35 | Denisovan | Penghu 1 | 0.01–0.07 or 0.13–0.19 | A | Va | ✓ | ✓ | ✓ |
| Preprint47 | Denisovan | Denisova 25 | 0.2 | A | M/Va | Genome | ||
| Published52 | Denisovan | Denisova 3 | 0.072 | A | Va | Genome | ||
| Published45 | Neanderthal | Altai | 0.122 | A | M | Genome | ||
| Published53 | Neanderthal | Vindija33.19 | 0.052 | A | M | Genome | ||
| Published54 | Early modern human | Ust’-Ishim | 0.045 | A | M | Genome | ||
| Database30 | Modern human | HGDP | Modern | A | M 99.6%, Va 0.4%c | Genome | ||
| Database55 | Modern human | gnomAD | Modern | A (>99.999%) | M 99.993%, Va 0.007%c | Genome and exome | ||
| Database56 | Hominids | Great apes | Modern | A | M | Genome, proteome and so on | ||
Note that peptides shorter than 8 amino acids were discarded in SAP confirmation.
aDerived amino acid.
bThe first number in the parentheses is the number of PSMs, and the second number is the number of peptides in the PEAKS result.
cMost individuals with AMBN(M273V) are in Southeast Asia and Oceania Islands, which have high Denisovan introgression.
Extended Data Fig. 2. The mirror plot for the fragment ions of observed and predicted spectra for SAP AMBN−253G in the six H. erectus samples and AMBN-253A in the Harbin specimen.
n = deamidated N, m = oxidated M (in SJD10M, SJD10A, SJD14M and ZKD) or dioxidated M (in HX-S1 and HX-S2).
Extended Data Fig. 4. The mirror plot for the fragment ions of observed and predicted spectra for SAP AMBN-273V in the six H. erectus samples and AMBN-273V/M in the Harbin specimen.
y = dichlorinated Y (in HX-S1), or chlorinated Y (in HX-S2).
Extended Data Fig. 3. Phylogenetic tree of six Middle Pleistocene H. erectus specimens, one Denisovan, two Neanderthals (Vindija33.19 and Altai), an African (MH_San987), an Oceanian (MH_Bougainville1027), a chimpanzee, and a gorilla, using Bayesian analysis.

Numbers on the tree indicate posterior probabilities for each clade, and those below 0.2 were not shown. The scale bar represents substitution per site.
Another SAP, AMBN(M273V), is also present in all six Middle Pleistocene H. erectus samples (Table 1, Fig. 3 and Extended Data Fig. 4). The corresponding single-nucleotide polymorphism (SNP allele (rs564905233, A>G)) is found in Denisovans and appears to have contributed to modern humans through Denisovan introgression. It exhibits a frequency of 21% in the Philippines, 1.17% in India, 0.71% in Papua New Guinea, and is absent from most other modern human populations44. Notably, genomic studies reveal that Denisovans received 0.5–8% gene flow from a hominin whose ancestors diverged more than 1 Ma from the common lineage ancestral to Neanderthals, Denisovans and modern humans45, and about 15% of these ‘super-archaic’ DNA regions introgressed from Denisovans into Asian and Oceanian individuals46. This situation is similar to what we observe with AMBN(M273V). Consequently, this variant is not exclusive to Denisovans but appears to have been introduced into them through a population linked to Middle Pleistocene H. erectus from Zhoukoudian, Hexian, and Sunjiadong (Fig. 4).
Fig. 4. A possible model of gene flow related to AMBN(M273V) among H. erectus associated with the populations of Zhoukoudian, Hexian and Sunjiadong, Denisovans and modern humans.

Neanderthal introgressions are not displayed.
Additionally, the DNA region containing rs564905233, which corresponds to the protein at position 273 of AMBN, shows greater sequence divergence between Africans and Denisova 3 than between Africans and the Altai Neanderthal, as identified through sliding window analysis (Methods, Extended Data Fig. 5). This supports the idea that this region, including rs564905233, may originate from a more diverged super-archaic group. In addition, although the AMBN(M273V) variant is homozygous in the more recent Denisovans, Denisova 3 and Penghu 1, it is heterozygous in both known ‘earlier’ Denisovans, Harbin (less than 0.15 Ma) and Denisova 25 (around 0.2 Ma)47 (Table 1, Fig. 3b and Extended Data Fig. 4), indicating that early Denisovans had a lower allele frequency of the derived allele. Furthermore, the morphology of Zhoukoudian and Hexian is distinct from that of the Harbin cranium48. Therefore, because of differences in age, morphology and known genetic variants, Zhoukoudian, Hexian and Sunjiadong can be excluded from the early Denisovans.
Extended Data Fig. 5. Permutation test p-values for DNA sequence divergence between modern Africans compared with a Denisovan and a Neanderthal.

Significance calculated in sliding 20 kb windows with 2 kb steps along a 70 kb segment of chromosome 4 containing rs564905233 (SNP position marked in red). Higher negative log p-values indicate more significant sequence divergence between Africans and Denisova 3 than between Africans and the Altai Neanderthal. The red dotted line shows that SNP rs564905233 is found within a contiguous 34 kb window of sequence divergence p < 0.05. The X-axis shows the chromosome 4 coordinates aligning to the center of each window (full results in Supplementary Data 3).
Discussion and conclusions
Recent research has demonstrated the potential of ancient protein analysis to uncover human evolutionary history11,35,36,49,50. However, genetic data about H. erectus, a key hominin and the first to travel to and inhabit geographically diverse regions, are scarce. The enamel proteins obtained from six Middle Pleistocene H. erectus fossils from Zhoukoudian, Hexian and Sunjiadong—three well-known sites spanning northern and southern China—permit a rare insight into the genetic makeup of H. erectus.
Morphological variability among these three sites, especially the Hexian samples, which differ from those of Zhoukoudian and Sunjiadong, prompted discussion about whether they are more closely related to H. erectus from Indonesia or possibly Denisovans14,15,18,20,21. Protein evidence clearly shows that the two Hexian specimens possess the newly identified AMBN(A253G) mutation, as do the H. erectus specimens from Zhoukoudian and Sunjiadong, and that this mutation is not found in Denisova 3, Penghu 1 or Harbin, placing them within H. erectus. Further palaeoproteomic research is needed to explore the molecular diversity within H. erectus.
The AMBN(A253G) and AMBN(M273V) variants are potentially specific to the populations to which these six Middle Pleistocene H. erectus specimens belonged. Notably, AMBN(A253G) has not been observed in any humans or primates before, making this finding highly notable. These H. erectus specimens come from three sites across East Asia, all dating to around 0.4 Ma, and Denisovans diverged from Neanderthals around 0.38–0.47 Ma (ref. 45). The known range of Denisovans includes Siberia, the Tibetan Plateau, Harbin in northeast Asia and the Penghu Channel in Southeast Asia, and these three H. erectus sites are from both northern and southern China. Their shared habitats create opportunities for interactions between populations related to these Middle Pleistocene H. erectus individuals and Denisovans. The identification of the derived AMBN(M273V) variant in all East Asian H. erectus specimens of this study and its occurrence in Denisovans, makes H. erectus a candidate source for the super-archaic introgressed DNA identified in the Denisova 3 genome. This variant is likely to originate from populations related to Late Middle Pleistocene H. erectus that interacted with Denisovans during the periods when these groups coexisted51 in East Asia. Further research on H. erectus, including molecular data across different periods and regions, will help to clarify their microevolution, population diversity and interactions with Denisovans.
Methods
Protein extraction
Animal fossils
Powder was drilled from the dentin (Supplementary Data 1), then mixed with 1.5 ml of 0.6 M HCl for decalcification. The precipitate was washed with ultrapure water at least 3 times until the pH reached approximately 7. Then, 200 µl of 50 mM NH4HCO3 was added, and the mixture was incubated at 65 °C for 3 h to extract soluble proteins. The supernatant containing the proteins was then transferred into a new centrifuge tube, and 1 µg of trypsin (Promega) was added. The mixture was incubated at 37 °C for 18 h. A 2.5% trifluoroacetic acid (TFA) solution (final concentration 0.1%) was added to stop the reaction. Subsequently, the trypsinized peptides were desalted and purified using C18 ZipTips57. For enamel samples, the powder was mixed with 1 ml of 5% HCl37 for decalcification, and the acid was replaced daily until the reaction ceased. The acid solution containing dissolved enamel peptides was concentrated using a vacuum concentrator, and the peptides were then desalted and purified using C18 ZipTips58.
Peptides were eluted into a solution of 80% acetonitrile with 0.1% TFA for MALDI-TOF mass spectrometry analysis. All sample preparations were performed in the dedicated clean room at the Molecular Paleontology Laboratory, IVPP of the Chinese Academy of Sciences in Beijing.
Hominin fossils
An acid etching method was used to extract protein from tooth enamel, modified from the process described in ref. 37. Disposable toothbrushes were used to remove surface contaminants from a small area of the enamel for etching. At the same time, the remaining teeth were wrapped with parafilm to prevent contact with any liquids. Before etching, the small enamel area was initially washed with 3% H2O2 for 30 s, followed by a rinse with ultrapure water. Approximately 100 µl of 5% (v/v) HCl was placed in the cap of a 1.5-ml microcentrifuge tube. A 2-min etch was performed by immersing the etching region in the HCl solution, and the initial etch solution was discarded. A second etch, lasting 15 min, was carried out in the cap of another separate microcentrifuge tube, and the etch solution was retained. This second etch was repeated, and the etch solutions were combined. After etching, the etched area was treated with 100 µl of 50 mM ammonium bicarbonate solution for 1 min to neutralize the acid. It was then rinsed with ultrapure water for 30 s and dried. The combined etch solution was then desalted using C18 ZipTips (Thermo Fisher Scientific) and eluted into a solution of 0.1% TFA and 80% acetonitrile (ACN). The peptide mixture (50 µl) was further divided into three aliquots; one aliquot was composed of 16 µl, among which 3 μl was used for the MALDI-TOF mass spectrometry test and 13 µl was retained as backup in our laboratory; two aliquots (each composed of 17 µl) were dried for LC–MS/MS analysis in two independent laboratories. All sample preparation for the experiment was conducted in the dedicated clean room at the Molecular Paleontology Laboratory, IVPP of the Chinese Academy of Sciences in Beijing.
MALDI-TOF mass spectrometry analysis
The peptide mixture was analysed on a Bruker autoflex maX MALDI-TOF mass spectrometer. In detail, 1 µl of peptide mixture was spotted onto a MTP384 Bruker ground-steel MALDI target plate, and 1 µl of α-cyano-4-hydroxycinnamic acid matrix solution (1% in 50% ACN/0.1% TFA (v/v/v)) was added on top. They were mixed, dried, and analysed on the mass spectrometer with a m/z range of 700–3,500. Each sample was analysed in triplicate. The raw data files were processed by mMass (v5.5.0)59.
LC–MS/MS analysis
For each hominin fossil, the eluted peptides from the enamel extraction were analysed under DDA mode in triplicate, including two runs on an Orbitrap Fusion Lumos mass spectrometer (Thermo Fisher Scientific) at Capital Medical University, Beijing, and one run on an Orbitrap Exploris 480 mass spectrometer (Thermo Fisher Scientific) at Fudan University. Both devices were coupled to an Easy nLC 1200 HPLC system (Thermo Fisher Scientific).
For the Orbitrap Fusion Lumos at Capital Medical University, the peptides were initially loaded onto a 100 μm internal diameter × 2 cm trap column and then separated on a 150 μm internal diameter × 15 cm analytical column. Both columns were packed in-house using 3 μm reversed-phase silica (Reprosil-Pur C18 AQ, Dr. Maisch). The peptides were eluted using a 120 min linear gradient programme (0–8 min, 7–11% B; 8–96 min, 11–28% B; 96–108 min, 28–40% B; 108–113 min, 40–90% B; 113–120 min, 90% B) at a flow rate of 500 nl min−1. Buffer A was 0.1% formic acid in water, and buffer B was 80% acetonitrile and 0.1% formic acid. The MS1 data were acquired across 375–1,400 m/z, with a resolution of 120k at m/z 200, a 250% AGC target, and a maximum injection time of 50 ms. The MS2 scans were performed with a resolution of 15k at m/z 200, an AGC target of 100%, a 35% normalized collision energy, and a maximum injection time of 22 ms.
For the Orbitrap Exploris 480 at Fudan University, Shanghai, the peptides were separated on a 75 μm internal diameter × 25 cm analytical column, which was packed in-house using reversed-phase silica of 1.9 μm (Reprosil-Pur C18 AQ, Dr. Maisch). Buffer A was 0.1% formic acid in water, and buffer B was 80% acetonitrile and 0.1% formic acid. An 80 min gradient was used with the following profile: 5–8% B, 2 min, at a flow rate of 200 nl min−1; 8–44% B, 38 min, 200 nl min−1; 44–70% B, 8 min, 200 nl min−1; 70–100% B, 2 min, 200 nl min−1; 100% B, 10 min, 200 nl min−1; 100–5% B, 2 min, 200 nl min−1; 5% B, 2 min, 300 nl min−1; 5–100% B, 6 min, 300 nl min−1; 100% B, 10 min, 300 nl min−1. Full mass spectrometry scans were acquired for the first 65 min, after which the column was washed and re-equilibrated for 15 min without data acquisition. The full mass spectrometry data acquisition was conducted across the range of m/z 350–1,600, with a resolution of 60k at m/z 200. The AGC target was set to ‘standard’, and the maximum injection time mode was set to ‘auto’. The MS/MS spectra were acquired with a resolution of 15k at m/z 200, a maximum injection time of 30 ms and a normalized collision energy of 30%. The AGC target was also set to standard.
For each animal fossil, the eluted peptides from enamel extractions were analysed for one run under DDA mode on the Orbitrap Fusion Lumos mass spectrometer (Thermo Fisher Scientific), either at Capital Medical University or Fudan University. For the Orbitrap Fusion Lumos at Capital Medical University, the liquid chromatography gradient and mass spectrometry parameters were the same as those of the hominin fossils. The Orbitrap Fusion Lumos at Fudan University was also interfaced with an Easy nLC 1200 HPLC system (Thermo Scientific). The peptides were separated on a 75 μm internal diameter × 20 cm analytical column packed with 1.9 μm reversed-phase silica. Mobile phase A consisted of 0.1% formic acid, and mobile phase B consisted of 80% acetonitrile and 0.1% formic acid. An 80 min gradient was used with the following profile: 2–5% B, 3 min, at a flow rate of 200 nl min−1; 5–35% B, 40 min, 200 nl min−1; 35–44% B, 5 min, 200 nl min−1; 44–100% B, 2 min, 200 nl min−1; 100% B, 10 min, 200 nl min−1; 100–5% B, 2 min, 200 nl min−1; 5% B, 2 min, 300 nl min−1; 5–100% B, 6 min, 300 nl min−1; 100% B, 10 min, 300 nl min−1. Full mass spectrometry scans were acquired for the first 65 min, after which the column was washed and re-equilibrated for 15 min without data acquisition. The full mass spectrometry data acquisition was conducted across the m/z range 350–1,600, with a resolution of 60k at m/z 200, a 100% AGC target, and a maximum injection time of 50 ms. The MS/MS spectra were acquired with a resolution of 15k at m/z 200, a 100% AGC target, a 30% normalized collision energy, and a maximum injection time of 30 ms. Blank extractions were processed concurrently to monitor the exogenous contaminants during the procedure.
Data search strategy
MaxQuant (v2.6.0.0)32, PEAKS Online (v12)33 and pFind (v3.2.1)34 were used to search the raw data. The H. erectus raw files were searched with the corresponding laboratory blanks and modern H. sapiens raw files43 against the ‘Hominidae enamel database’, supplemented with the contaminant database. Unspecific digestion was selected in each software. The animal raw files were searched against the ‘mammal enamel database’.
Database composition
The ‘Hominidae enamel database’ comprised 13 selected enamel proteins from Hominidae. Besides the commonly used 12 proteins (AHSG, ALB, AMBN, AMELX, AMELY, AMTN, COL17A1, ENAM, KLK4, MMP20, ODAM and TUFT1), SERPINC1 was also added because we identified this protein in Harbin with abundant peptides and elevated deamidation rates (nearly 100%), and this protein was also reported in modern enamel60,61. The sequences were retrieved from UniProt, downloaded from the ‘Hominid Palaeoproteomic Reference Dataset’ (https://zenodo.org/records/7728060), translated from the genomes of public projects30, and specific sequences from published palaeoproteomes11. The ‘mammal enamel database’ was composed of mammal sequences of the same enamel proteins above, retrieved from UniProt using the gene names and ‘Mammalia (mammals) [40674]’. The contaminant database was composed of a previously published contaminant database62, the cRAP database (https://www.thegpm.org/crap/), and the contaminant database from MaxQuant (v2.6.0.0). The ‘Hominidae enamel database’ was accessible through the ProteomeXchange Consortium (Data availability).
PEAKS search
The precursor ion (MS1) mass tolerance was set to 10 ppm, with a fragment ion (MS2) mass tolerance of 0.02 Da for all PEAKS searches, with unspecific digestion. Variable modifications included deamidation (NQ), oxidation (M), hydroxylation (P), phosphorylation (STY), N-terminal pyro-Glu from E, and N-terminal pyro-Glu from Q, with no fixed modifications and up to three modifications allowed per peptide. The peptide length was set to 6–45. PSMs were filtered using a false discovery rate (FDR) of 1%, and proteins were filtered with criteria of −10logP ≥ 20 and average local confidence (ALC) ≥ 50% (de novo only).
After our initial search with PEAKS, additional variable modifications were included in the second-round search for Hexian H. erectus samples: chlorination and dichlorination of tyrosine residues, dehydration, dioxidation (W), carbonyl E, dioxidation (M), oxidation (HW), ornithine derived from arginine, tryptophan oxidation to kynurenine, tryptophan oxidation to oxolactone, and proline oxidation to pyroglutamic acid. Up to five modifications per peptide were permitted. The peptide length range was set to 6–30.
After removing peptides from the contaminant database and those detected in the extraction blank samples or matching multiple genes, the deamidation rates of glutamine (Q) and asparagine (N) were calculated for each sample based on PSM counts. We used auxiliary tools for PSM prediction with PEAKS.
pFind search
We included Deamidation [N], Deamidation [Q], Oxidation [M], Oxidation [P], Oxidation [W], Gln->pyro-Glu[AnyN-termQ], Glu->pyro-Glu[AnyN-termE], Pro->pyro-Glu[P], Phospho[S], Phospho[T], Phospho[Y], Dehydrated[S], Dehydrated[T], Dehydrated[Y], Arg->Orn[R], Dioxidation[M], Dioxidation[W], Thiazolidine[W], Trp->Kynurenin[W], Trp->Oxolactone[W], Ammonia-loss[N], His->Asp[H], Pro->HAVA[P], and Amidated[AnyC-term] as variable modifications, with no fixed modifications included. Spectra FDR was set at 1%, and protein FDR was 10%. The mass range of each peptide was set from 350 to 4,000 Da. Open search was enabled. Within pFind’s modification configuration, the default mass ‘X’ is preset to that of isoleucine/leucine. To avoid incorrect identification of the X residue, we set its mass to 6,228.71 Da (Sm41), which substantially exceeds the mass of any natural amino acid. Therefore, any in silico peptide with an X produces abnormal theoretical precursor and fragment ion masses, effectively preventing its matching to experimental spectra during database searches and thus reducing false positives from these ambiguous sequence regions. The other parameters were the same as the settings in PEAKS.
After our first-round search with pFind, some additional variable modifications were selected for inclusion in the second-round search for Hexian H. erectus samples: Chlorination[Y], dichlorination[Y], Carbonyl[E], and Dioxidation[P](Pro->Glu[P]).
MaxQuant search
No fixed modifications were specified. Variable modifications included Deamidation (NQ), Phosphorylation (STY), Gln to N-terminal pyro-Glu, Glu to N-terminal pyro-Glu, Dioxidation (MW), Oxidation (M), Oxidation (P), and Oxidation (W). PSM FDR was set at 1% for all 7 samples. The search also enabled the identification of dependent and secondary peptides. The remaining parameters were set to their default values. After the search, the deamidation rates of N and Q were calculated63 for each sample.
In addition, we determined the extra variable modification for Hexian H. erectus samples with PEAKS and pFind results. Cl(Y) and diCl (Y) were added for HX-S1 (both post-translational modifications were self-made in Configuration-Modifications), and Cl (Y) and Dehydrated (STY) were included for HX-S2.
Considering the high resolution of the instrumentation used, we also reduced the precursor ion (MS1) mass tolerance to 5 ppm and repeated all analyses in PEAKS and pFind. The results were similar, and the two main SAPs (AMBN 253 and AMBN 273) within the Homo genus were consistently identified (Supplementary Table 8). As the final results do not change significantly, we focus mainly on the results from the commonly used 10 ppm search to make our results more comparable to previous studies.
Construction of the consensus protein sequences and phylogenetic analysis
Consensus sequences of endogenous proteins were reconstructed for phylogenetic analysis (Supplementary Data 1). Peptides shorter than eight amino acids or with abnormal or artificial post-translational modifications were excluded for the consensus sequence reconstruction of each endogenous protein. Only alleles with a PSM count ≥2, a PSM ratio ≥10%, and an intensity ratio ≥10% were considered reliable and retained. For heterozygous sites, at least two peptides were required per allele. The heterozygous variant site, observed only in the Harbin specimen, was identified with 79 peptides for the V allele and 38 for the M allele at position 273 in AMBN (Table 1). We include both alleles of the Harbin consensus in Supplementary Data 2. We performed additional AMELY sequence correction as described in Supplementary Note 4.
The protein data in the phylogenetic tree include Denisova 3 (ref. 52), two Neanderthals (Altai, Vindija33.19)45,53, two modern humans from the HGDP project (San, HGDP0987; Bougainville, HGDP01027)30, a chimpanzee (Pan)29, and a Gorilla29. We used PartitionFinder2 (v2.1.1)64 to identify the best partitioning schemes and amino acid substitution models for our dataset. Then, we built a consensus Bayesian phylogenetic tree using the software MrBayes (v3.2.6)65, running 8 Markov Chain Monte Carlo (MCMC) chains for 1 million iterations in 2 independent runs. Sampling was done every 500 generations, and the first 200,000 iterations were discarded as burn-in. The tree was plotted with FigTree v1.4.4 (http://tree.bio.ed.ac.uk/software/figtree).
DNA analysis
Sliding window analysis
Variant sites were identified where either Denisova 3 or two Neanderthals (Vindija33.19 and Altai) differed from two modern African diploid genomes (S_Khomani_San-1.DG, S_Mandenka-2.DG) from the phased Simons Genome Diversity Project panel (SGDP) (https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/phased_data2021)66,67. For each of 129 overlapping 20 kb windows (with 2 kb steps), pairwise matching rates were calculated for variants between each African haploid and each archaic artificial haploid of either Denisova 3 or the Altai Neanderthal (where each variant is randomly assigned to a haploid to preserve all variants), covering in total 276 kb of DNA sequence surrounding the rs564905233 SNP. For each window, statistical significance of the differences between the African–Denisovan and African–Neanderthal matching rates were assessed using a Wilcoxon rank-sum test, and a permutation test (n = 1,000 permutations). The Wilcoxon rank-sum test was used to compare distributions of pairwise matching rates, with W representing the sum of ranks in the African–Denisovan group. The permutation test was used to evaluate differences in mean matching rates. No adjustments were made for multiple comparisons, since each window was analysed independently to identify localized regions of divergence. All statistical tests were performed as two-sided tests. Full results, including W, exact P values, and the number of comparisons per window, are provided in Supplementary Data 3.
Ethics statement
Permission to test for ancient proteins in the human specimens from this study was granted by the collection room of the IVPP, the Hexian Culture, Tourism, and Sports Bureau, and the Luanchuan County Culture, Radio, Television, and Tourism Bureau. The work was conducted in collaboration with local researchers, who are co-authors because of their contributions to assembling archaeological materials and/or discussions that informed the study.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Online content
Any methods, additional references, Nature Portfolio reporting summaries, source data, extended data, supplementary information, acknowledgements, peer review information; details of author contributions and competing interests; and statements of data and code availability are available at 10.1038/s41586-026-10478-8.
Supplementary information
Supplementary Notes 1–6, Supplementary Figs. 1–4, Supplementary Tables 1–8 and legends for Supplementary Data 1–3
Supplementary Data 1–3
Acknowledgements
We thank F. Bai, K. Li, L. Duan and Z. Zou for their initial visualization of parts of Figs. 1 and 3b; X. Yang for initial testing; S. Pääbo for comments; J. Kelso and S. Peyrégne for permission to use the data of one position of the Denisova 25 genome; the BSI team for making improvements to PEAKS Online; and Natural Earth for providing the map data used in this study. This work was supported by the National Natural Science Foundation of China (L2424324), the Chinese Academy of Sciences (CAS) (YSBR-019), the Archaeological Talent Promotion Program of China (2024-278) and the New Cornerstone Science Foundation.
Extended data figures and tables
Author contributions
Q.F. conceived and designed the research project. S.X., Q.J., Z.D., X.G., Y.D. and J.X. prepared the archaeological samples and materials. Q.F., H.R. and K.Z. performed or supervised wet laboratory work. Z.D., X.G., Y.D. and J.X. provided archaeological interpretations. Q.F., Z.W., and X.F. processed the data. Q.F., Z.W. and E.A.B. analysed data. Q.F. wrote the manuscript. Q.F. and E.A.B. revised the manuscript. Q.F., Z.W., E.A.B., H.R. and K.Z. wrote the methods. Q.F., S.X. and Z.W. wrote and prepared the supplementary information. All authors discussed, critically revised and approved the final version of the manuscript.
Peer review
Peer review information
Nature thanks Enrico Cappellini who co-reviewed with Palesa Madupe; Daniel Green and Ryan Paterson for their contribution to the peer review of this work. Peer reviewer reports are available.
Data availability
All the proteomic mass spectrometry data and searching database have been deposited in the ProteomeXchange Consortium via the PRIDE68 partner repository with the dataset identifier PXD068897. Protein consensus sequences for six specimens and the Harbin individual are available in Supplementary Data 2. Reference palaeoproteomic data is available through the PRIDE partner repository (https://www.ebi.ac.uk/pride/archive/) with the dataset identifier PXD054412 (the Penghu 1 specimen), PXD040221 (4 Paranthropus robustus specimens: SK14132, SK830, SK835 and SK850), PXD014342 (Dmanisi H. erectus specimen D4163, and Atapuerca H. antecessor specimen ATD6-92), PXD018721 (specimen Laos_TNH2, and modern human specimens 1692 and 1693), PXD009781 (17 modern human specimens) and PXD012587 (3 modern human specimens). The ‘Hominid Palaeoproteomic Reference Dataset’ is available through the Zenodo repository (10.5281/zenodo.7728060 (ref. 56)). All other reference enamel protein sequences are available in the UniProt Knowledgebase (https://www.uniprot.org/). Human reference genome hg19 is available through the National Center for Biotechnology Information (https://www.ncbi.nlm.nih.gov/) under accession number PRJNA31257, and the corresponding aligned chimpanzee genome is available at http://hgdownload.cse.ucsc.edu/goldenPath/hg19/vsPanTro6. The previously reported ancient DNA datasets used in this study are available through the Allen Ancient DNA Resource v62.0 (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW). All modern human genomes from the HGDP project are available at ftp://ngs.sanger.ac.uk/production/hgdp/hgdp_wgs.2019051. The worldwide base maps of land, ocean, and rivers are available in Natural Earth (https://www.naturalearthdata.com/).
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data
is available for this paper at 10.1038/s41586-026-10478-8.
Supplementary information
The online version contains supplementary material available at 10.1038/s41586-026-10478-8.
References
- 1.Antón, S. C. Natural history of Homo erectus. Am. J. Biol. Anthropol.122, S126–S170 (2003). [DOI] [PubMed] [Google Scholar]
- 2.Herries, A. I. R. et al. Contemporaneity of Australopithecus, Paranthropus, and early Homo erectus in South Africa. Science368, eaaw7293 (2020). [DOI] [PubMed] [Google Scholar]
- 3.Mussi, M. et al. Early Homo erectus lived at high altitudes and produced both Oldowan and Acheulean tools. Science382, 713–718 (2023). [DOI] [PubMed] [Google Scholar]
- 4.Ferring, R. et al. Earliest human occupations at Dmanisi (Georgian Caucasus) dated to 1.85–1.78 Ma. Proc. Natl Acad. Sci. USA108, 10432–10436 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Lordkipanidze, D. et al. A complete skull from Dmanisi, Georgia, and the evolutionary biology of early Homo. Science342, 326–331 (2013). [DOI] [PubMed] [Google Scholar]
- 6.Huguet, R. et al. The earliest human face of Western Europe. Nature640, 707–713 (2025). [DOI] [PubMed] [Google Scholar]
- 7.Kamberov, Y. G. et al. Modeling recent human evolution in mice by expression of a selected EDAR variant. Cell152, 691–702 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zhu, Z. Y. et al. New dating of the Homo erectus cranium from Lantian (Gongwangling), China. J. Hum. Evol.78, 144–157 (2015). [DOI] [PubMed] [Google Scholar]
- 9.Zhu, Z. et al. Hominin occupation of the Chinese Loess Plateau since about 2.1 million years ago. Nature559, 608–612 (2018). [DOI] [PubMed] [Google Scholar]
- 10.Liu, W., Wu, X. & Xing, S. The morphological evidence for the regional continuity and diversity of Middle Pleistocene human evolution in China. Acta Archaeol. Sin.38, 473–490 (2019). [Google Scholar]
- 11.Welker, F. et al. The dental proteome of Homo antecessor. Nature580, 235–238 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Shen, G., Gao, X., Gao, B. & Granger, D. E. Age of Zhoukoudian Homo erectus determined with 26Al/10Be burial dating. Nature458, 198–200 (2009). [DOI] [PubMed] [Google Scholar]
- 13.Grün, R. et al. ESR analysis of teeth from the palaeoanthropological site of Zhoukoudian, China. J. Hum. Evol.32, 83–91 (1997). [DOI] [PubMed] [Google Scholar]
- 14.Antón, S. C. Evolutionary significance of cranial variation in Asian Homo erectus. Am. J. Phys. Anthropol.118, 301–323 (2002). [DOI] [PubMed] [Google Scholar]
- 15.Liu, W. & Zhang, Y. The cranial metric diversity of Chinese Homo erectus. Acta Anthropologica Sinica24, 121–136 (2005). [Google Scholar]
- 16.Grun, R. et al. ESR and U-series analyses of teeth from the palaeoanthropological site of Hexian, Anhui Province, China. J. Hum. Evol.34, 555–564 (1998). [DOI] [PubMed] [Google Scholar]
- 17.Zhao, J.-X., Hu, K., Collerson, K. D. & Xu, H.-K. Thermal ionization mass spectrometry U-series dating of a hominid site near Nanjing, China. Geology29, 27–30 (2001). [Google Scholar]
- 18.Zhao, L. et al. Middle Pleistocene hominins from the Sunjiadong site, Luanchuan county, Henan Province, Central China. Acta Anthropol. Sin.37, 192–205 (2018). [Google Scholar]
- 19.Zhang, Y., Zhang, S., Gu, X. & Li, X. A study on the perforated bone from Sunjiadong site in Henan Province. Acta Anthropol. Sin.36, 457–464 (2017). [Google Scholar]
- 20.Liu, W. et al. A mandible from the Middle Pleistocene Hexian site and its significance in relation to the variability of Asian Homo erectus. Am. J. Phys. Anthropol.162, 715–731 (2017). [DOI] [PubMed] [Google Scholar]
- 21.Kaifu, Y. Archaic hominin populations in Asia before the arrival of modern humans: their phylogeny and implications for the “Southern Denisovans”. Curr. Anthropol58, S418–S433 (2017). [Google Scholar]
- 22.Jia, L. P. Report on excavations at the Sinanthropus site in 1958. Palaeovertebrata Paleoanthropoligica1, 21–26 (1959). [Google Scholar]
- 23.Woo, J.-K. & Chia, L. New discoveries of Sinanthropus pekinensis in Choukoutien. Acta Palaeontol. Sin.2, 267–288 (1954). [Google Scholar]
- 24.Huang, P.-H. et al. Study of ESR dating for burying age of the first skull of Pekinig Man and chronological scale of the cave deposit in Zhoukoudian site Loc. 1. Acta Anthropol. Sin.10, 107–115 (1991). [Google Scholar]
- 25.Pei, J. in Multi-Disciplinary Study of the Peking Man Site at Zhoukoudian (eds Wu, R. K. et al.) 258–260 (Science Press, 1985).
- 26.Huang, W., Fang, D. & Ye, Y. Preliminary studyon the fossil hominid skull and fauna of Hexian, Anhui. Vertabrata Palasiatica20, 248–256 (1982). [Google Scholar]
- 27.Ji, Q., Wu, W., Ji, Y., Li, Q. & Ni, X. Late Middle Pleistocene Harbin cranium represents a new Homo species. Innovation2, 100132 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Ni, X. et al. Massive cranium from Harbin in northeastern China establishes a new Middle Pleistocene human lineage. Innovation2, 100130 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Patramanis, I., Ramos-Madrigal, J., Cappellini, E. & Racimo, F. PaleoProPhyler: a reproducible pipeline for phylogenetic inference using ancient proteins. Peer Commun. J.3, e112 (2023). [Google Scholar]
- 30.Bergström, A. et al. Insights into human genetic variation and population history from 929 diverse genomes. Science367, eaay5012 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Madupe, P. P. et al. Enamel proteins reveal biological sex and genetic variability in southern African Paranthropus. Science388, 969–973 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Cox, J. & Mann, M. MaxQuant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat. Biotechnol.26, 1367–1372 (2008). [DOI] [PubMed] [Google Scholar]
- 33.Xin, L. et al. A streamlined platform for analyzing tera-scale DDA and DIA mass spectrometry data enables highly sensitive immunopeptidomics. Nat. Commun.13, 3108 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Chi, H. et al. Comprehensive identification of peptides in tandem mass spectra using an efficient open search engine. Nat. Biotechnol.36, 1059–1061 (2018). [DOI] [PubMed] [Google Scholar]
- 35.Tsutaya, T. et al. A male Denisovan mandible from Pleistocene Taiwan. Science388, 176–180 (2025). [DOI] [PubMed] [Google Scholar]
- 36.Fu, Q. et al. The proteome of the late Middle Pleistocene Harbin individual. Science389, eadu9677 (2025). [DOI] [PubMed] [Google Scholar]
- 37.Stewart, N. A., Gerlach, R. F., Gowland, R. L., Gron, K. J. & Montgomery, J. Sex determination of human remains from peptides in tooth enamel. Proc. Natl Acad. Sci. USA114, 13649–13654 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Parker, G. J. et al. Sex estimation using sexually dimorphic amelogenin protein fragments in human enamel. J. Archaeol. Sci.101, 169–180 (2019). [Google Scholar]
- 39.Lugli, F. et al. Enamel peptides reveal the sex of the Late Antique ‘Lovers of Modena’. Sci. Rep.9, 13130 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Welker, F. Elucidation of cross-species proteomic effects in human and hominin bone proteome identification through a bioinformatics experiment. BMC Evol. Biol.18, 23 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Presslee, S. et al. Palaeoproteomics resolves sloth relationships. Nat. Ecol. Evol.3, 1121–1130 (2019). [DOI] [PubMed] [Google Scholar]
- 42.Bai, F., Wu, Z., Xing, S. & Fu, Q. Rapid and robust sex determination from ancient enamel proteomes using protSexInferer. J. Genet. Genomics10.1016/j.jgg.2026.04.012 (2026). [DOI] [PubMed]
- 43.Demeter, F. et al. A Middle Pleistocene Denisovan molar from the Annamite Chain of northern Laos. Nat. Commun.13, 2557 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Wall, J. D. et al. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature576, 106–111 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Prüfer, K. et al. The complete genome sequence of a Neanderthal from the Altai Mountains. Nature505, 43–49 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Hubisz, M. J., Williams, A. L. & Siepel, A. Mapping gene flow between ancient hominins through demography-aware inference of the ancestral recombination graph. PLoS Genet.16, e1008895 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Peyrégne, S. et al. A high-coverage genome from a 200,000-year-old Denisovan. Preprint at bioRxiv10.1101/2025.10.20.683404 (2025).
- 48.Feng, X. et al. The phylogenetic position of the Yunxian cranium elucidates the origin of Homo longi and the Denisovans. Science389, 1320–1324 (2025). [DOI] [PubMed] [Google Scholar]
- 49.Xia, H. et al. Middle and Late Pleistocene Denisovan subsistence at Baishiya Karst Cave. Nature632, 108–113 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Chen, F. et al. A late Middle Pleistocene Denisovan mandible from the Tibetan Plateau. Nature569, 409–412 (2019). [DOI] [PubMed] [Google Scholar]
- 51.Sawafuji, R., Tsutaya, T., Takahata, N., Pedersen, M. W. & Ishida, H. East and Southeast Asian hominin dispersal and evolution: a review. Q. Sci. Rev.333, 108669 (2024). [Google Scholar]
- 52.Meyer, M. et al. A high-coverage genome sequence from an archaic Denisovan individual. Science338, 222–226 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Prüfer, K. et al. A high-coverage Neandertal genome from Vindija Cave in Croatia. Science358, 655–658 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Fu, Q. et al. Genome sequence of a 45,000-year-old modern human from western Siberia. Nature514, 445–449 (2014). [DOI] [PMC free article] [PubMed]
- 55.Chen, S. et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature625, 92–100 (2024). [DOI] [PMC free article] [PubMed]
- 56.Patramanis, I., Ramos Madrigal, J., Cappellini, E. & Racimo, F. Hominid Palaeoproteomic Reference Dataset (1.0.1) [Data set]. Zenodo10.5281/zenodo.7728060 (2022).
- 57.Rao, H. et al. Palaeoproteomic analysis of Pleistocene cave hyenas from east Asia. Sci. Rep.10, 16674 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Buckley, M., Collins, M., Thomas-Oates, J. & Wilson, J. C. Species identification by analysis of bone collagen using matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry. Rapid Commun. Mass Spectrom.23, 3843–3854 (2009). [DOI] [PubMed] [Google Scholar]
- 59.Strohalm, M., Kavan, D., Novák, P., Volný, M. & Havlíček, V. mMass 3: a cross-platform software environment for precise analysis of mass spectrometric data. Anal. Chem.82, 4648–4651 (2010). [DOI] [PubMed] [Google Scholar]
- 60.Jágr, M. et al. Proteomic analysis of dentin–enamel junction and adjacent protein-containing enamel matrix layer of healthy human molar teeth. Eur. J. Oral Sci.127, 112–121 (2019). [DOI] [PubMed] [Google Scholar]
- 61.Gil-Bona, A. & Bidlack, F. B. Tooth enamel and its dynamic protein matrix. Int. J. Mol. Sci.21, 4458 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Hendy, J. et al. A guide to ancient protein studies. Nat. Ecol. Evol.2, 791–799 (2018). [DOI] [PubMed] [Google Scholar]
- 63.Mackie, M. et al. Palaeoproteomic profiling of conservation layers on a 14th century Italian wall painting. Angew. Chem. Int. Ed. Engl.57, 7369–7374 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Lanfear, R., Frandsen, P. B., Wright, A. M., Senfeld, T. & Calcott, B. PartitionFinder 2: new methods for selecting partitioned models of evolution for molecular and morphological phylogenetic analyses. Mol. Biol. Evol.34, 772–773 (2017). [DOI] [PubMed] [Google Scholar]
- 65.Ronquist, F. & Huelsenbeck, J. P. MrBayes 3: Bayesian phylogenetic inference under mixed models. Bioinformatics19, 1572–1574 (2003). [DOI] [PubMed] [Google Scholar]
- 66.Mallick, S. et al. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature538, 201–206 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Rubinacci, S., Ribeiro, D. M., Hofmeister, R. J. & Delaneau, O. Publisher Correction: Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat. Genet.53, 412 (2021). [DOI] [PubMed] [Google Scholar]
- 68.Y, P.-R. et al. The PRIDE database at 20 years: 2025 update. Nucleic Acids Res.53, D543–D553 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Notes 1–6, Supplementary Figs. 1–4, Supplementary Tables 1–8 and legends for Supplementary Data 1–3
Supplementary Data 1–3
Data Availability Statement
All the proteomic mass spectrometry data and searching database have been deposited in the ProteomeXchange Consortium via the PRIDE68 partner repository with the dataset identifier PXD068897. Protein consensus sequences for six specimens and the Harbin individual are available in Supplementary Data 2. Reference palaeoproteomic data is available through the PRIDE partner repository (https://www.ebi.ac.uk/pride/archive/) with the dataset identifier PXD054412 (the Penghu 1 specimen), PXD040221 (4 Paranthropus robustus specimens: SK14132, SK830, SK835 and SK850), PXD014342 (Dmanisi H. erectus specimen D4163, and Atapuerca H. antecessor specimen ATD6-92), PXD018721 (specimen Laos_TNH2, and modern human specimens 1692 and 1693), PXD009781 (17 modern human specimens) and PXD012587 (3 modern human specimens). The ‘Hominid Palaeoproteomic Reference Dataset’ is available through the Zenodo repository (10.5281/zenodo.7728060 (ref. 56)). All other reference enamel protein sequences are available in the UniProt Knowledgebase (https://www.uniprot.org/). Human reference genome hg19 is available through the National Center for Biotechnology Information (https://www.ncbi.nlm.nih.gov/) under accession number PRJNA31257, and the corresponding aligned chimpanzee genome is available at http://hgdownload.cse.ucsc.edu/goldenPath/hg19/vsPanTro6. The previously reported ancient DNA datasets used in this study are available through the Allen Ancient DNA Resource v62.0 (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW). All modern human genomes from the HGDP project are available at ftp://ngs.sanger.ac.uk/production/hgdp/hgdp_wgs.2019051. The worldwide base maps of land, ocean, and rivers are available in Natural Earth (https://www.naturalearthdata.com/).







