Abstract
The origins and movements of early East Asians remain unclear due to limited ancient genomes and Y chromosome data. Here, we report a large-scale Y chromosome resource, including 1045 newly sequenced genomes from previously underrepresented groups, revealing high-resolution phylogenetic patterns in F and N1b lineages. We constructed a time-calibrated phylogeny that traces early paternal roots to diverged C/D/F lineages and indicates Paleolithic migrations from the southern Himalayas. These early movements, together with later Neolithic expansions, shaped the paternal landscape observed today. We find long-lasting bottlenecks in early F lineages and reveal inland and coastal migration routes linking South China and Southeast Asia, followed by rapid diversification during the Neolithic period associated with millet and rice farming. Distinct inland and coastal southward expansions of N1a/b hunting-gathering and farming groups, including those linked to Tibeto-Burman speakers, and demographic shifts among northern coastal populations further shaped genomic diversity. These findings provide insights into how ancient divergence and agriculture-driven expansions influenced East Asian paternal history.
Ancient Y chromosome lineages trace Paleolithic migrations and Neolithic expansions that shaped East Asia’s paternal history
INTRODUCTION
Archaeological and genetic findings have traced the migration routes of southern and northern populations around the Himalayas, extending from Central Asia through the Tongtiandong site in the Altai region to northern East Asia, as well as from coastal Southeast Asia to southern East Asia (1, 2). These migrations played a key role in shaping early East Asian development, a region characterized by its genetic, phenotypic, cultural, and biological diversity. Genomes from pre-Last Glacial Maximum (LGM)—including the 40,000-year-old Tianyuan individual, the Ancient North Eurasian-related Mal’ta boy in the north, and Hòabìnhian hunter-gatherer lineages from the south—indicate limited direct genetic influence on post-LGM populations (3, 4). The survival of early Asian ancestry in marginal regions of eastern Eurasia is evidenced by Y chromosome lineages, such as various D sublineages found among the Andamanese Onge, highland Tibetans, and Japanese, as well as autosomal and mitochondrial genetic traits seen in tropical indigenous groups (5–7). An in-depth analysis of these deeply divergent Y chromosome lineages is essential for understanding early population dynamics in East Asia and the Qinghai-Xizang Plateau. However, low-resolution Y chromosome genotyping using single nucleotide polymorphisms (SNPs) and short tandem repeats has provided only a basic understanding of paternal genetic structure in China (2, 8), with limited capacity for precise timing and comparison with other East Asian groups, especially in South China, where ancient DNA data are lacking.
During the post-LGM period, global warming prompted shifts in human subsistence strategies, transitioning from merely surviving in extreme environments to actively modifying those environments. This cultural change, especially evident during the Neolithic era in regions of agricultural innovation such as the Yangtze and Yellow River basins, coincided with large-scale population movements, mixing, and replacement among spatiotemporally diverse millet and rice-farming groups (4, 5, 9–11). Ancient genomes from early Neolithic sites in Fujian (Liangdao and Qihe), Guangxi (Longlin, Dushan, and Baojianshan), Yunnan (Xingyi), Shandong (Xiaogao, Xiaojingshan, Boshan, and Bianbian), Inner Mongolia (Yumin), and the Amur River Basin in China, as well as Vietnam, reveal early population divergence between northern and southern East Asians (4, 5, 9–11). Genetic differences among groups, including Neolithic Tibetans, millet farmers, and Siberian hunter-gatherers, reflect a complex history of expansion and mixture, reshaping the genetic landscape of both ancient and present-day East Asian populations (12). Ancient DNA reveals coastal connections and gene flow between northern and southern groups in China, the far east in Russia, and coastal Vietnam, as determined through autosomal admixture analyses (11). A notable genetic link exists between ancient Tibetans and Neolithic millet farmers from northern China (13). While these insights shed light on the formation and genetic history of East Asians, much ancient DNA from southern Chinese archeological sites remains fragmentary due to the degradation of remains in hot, humid, and acidic environments.
The permanent settlement of East Asia has been systematically examined through archeological and genetic evidence (14, 15). Early research, which primarily focused on widespread founding lineages O, D, Q, and R (16), frequently overlooked early-diverged rare lineages. South China and Southeast Asia are likely to have served as the initial entry points for anatomically modern humans into East Asia ~50 thousand years ago (kya) (10, 11). Archeological findings suggest that the eastern and northeastern Qinghai-Xizang Plateau served as essential routes for human settlement in the highlands, containing early deep population genetic legacies that predate interactions with later migrants, including F lineages among Lahu people and D lineages among Tibetans (16–18). These regions—connected to Southeast Asia, lowland East Asia, and the Tibetan highlands—exhibit the highest levels of linguistic and genetic diversity and are recognized as biodiversity hot spots. Their roles as centers of evolution and primary migration corridors have spurred multidisciplinary research focused on reconstructing the complex histories of ethnolinguistically distinct populations (19). The region’s dynamic history, characterized by diverse origins, migrations, admixture, and cultural exchange, has produced the intricate linguistic and genetic patterns observed in present-day populations in southwestern China and Southeast Asia. Coastal genetic connections have recently been evidenced via ancient DNA from Japan, Fujian, and Vietnam (11). An autosomal ancient genomic resource from the Mengzi Ren in inland Southwest China, representing an early diversified human lineage, shows genetic links among early Asian populations (20). The ancient DNA from the Xingyi site has revealed a direct connection between early Yunnan people and ancient highland Tibetans (21). In addition, the Middle Holocene people of central Yunnan exhibited a close relationship with modern Austroasiatic peoples (21). Considering human population movements around the southern and eastern regions of the Qinghai-Xizang Plateau, along the parallel rivers in the Tibetan-Yi corridor, from northern East Asia to Southeast Asia, especially along the Lancang/Mekong Rivers (1, 20, 21), we hypothesize that the ancient inland route also contributed to the initial formation of East Asian populations and to subsequent complex migrations, in conjunction with active population movements along coastal routes (11). However, the scarcity of large-scale whole-genome sequencing (WGS) data from ethnolinguistically diverse groups limits a comprehensive understanding of their genomic diversity and evolutionary processes (22). To address this, it is essential to include genetically underrepresented populations in broader genomic studies across southwestern China and the Qinghai-Xizang Plateau to reconstruct deep population histories and investigate the impacts of population expansion and migration.
Nonrecombining genomic regions, such as the Y chromosome and mitochondrial DNA, have been beneficial for reconstructing human evolution (23, 24). Large-scale Y chromosome sequences from modern human populations provide detailed insights into both deep and recent evolutionary histories, enabling precise estimates of genomic diversity and divergence times (24–28). The reconstruction of European patrilineages from 3.7 million Y chromosome variants in 334 individuals across Europe and the Middle East revealed recent large-scale expansions of three major lineages related to I1, R1a, and R1b (29). Other Y chromosome resources from Sardinia, Africa, Oceania, and the Americas have shed light on patterns of paternal divergence and Neolithic demographic processes, as well as their influence on the genetic diversity of patrilocal and matrilocal modern populations (26, 27, 30). Globally, the analysis of 1244 Y chromosome sequences identified out-of-Africa and Neolithic bottlenecks, revealing a lower effective population size in males compared to females, which affected Y chromosome diversity (31). Further research demonstrated that patrilineal segmentary systems, along with changes in social and subsistence structures, led to the loss of paternal lineages with low reproductive success, serving as a key mechanism for demographic dynamics (25). Uniparental genetic lines also provided deeper insights into the human evolutionary past, sex-biased mixing, and patterns such as patrilocality and matrilocality in ancient and modern East Asians. Population genetic analyses indicate that southern Han populations primarily formed through male-driven migration from northern China, with local maternal lineages enduring and mixing (8, 32). Mitochondrial data also suggest multiple gene flow events across the Yellow and Yangtze River basins (7). Recent large-scale sequencing of the Y chromosome and analyses of diverse ancestral groups have revealed multiple migration waves and cultural exchanges that have influenced the paternal lineage of Chinese minority groups and the Han Chinese (24, 32). Contributions from western Eurasian groups (R/J haplogroups) and Siberian hunter-gatherers (C/Q haplogroups), alongside indigenous O lineages, have been identified in Chinese paternal gene pools (24). The paternal genetic diversity across Han Chinese populations reflects both the effect of mixing and isolation (32). Overall, the combination of ancient and modern autosomal, Y chromosome, and mitochondrial genomes demonstrates intricate patterns of genetic diversity, resulting from early population differentiation and subsequent large-scale admixture events (11, 24, 33).
Although more than a million whole-genome sequences from autosomal genomic resources have substantially advanced our understanding of human genetic history and its medical implications (22, 34), uniparental marker-based resources remain limited, especially in East Asia. Here, we present a large-scale, high-quality Y chromosome genomic resource from 17,740 unrelated modern individuals and 977 ancient individuals from China and neighboring regions, aiming to fill gaps in Chinese Y chromosome data and explore the deep history of ancient DNA that was previously unavailable (Fig. 1, A and B). We identified multiple early Asian lineages, including the underexplored F and N lineages, which meaningfully contribute to the genetic heritage of eastern Eurasia. F-related and other early Asian lineages migrated into East Asia via coastal and inland routes, remaining prominent in modern southwestern Chinese and other East Asian populations. Unlike early-diverged lineages D, F, and C, which may be linked to hunter-gatherer groups, Neolithic farmer-related lineages, especially O and N, experienced multiple expansions driven by farming and herding innovations. In addition, we reconstructed the complex demographic history of N lineages, which were shaped by long-range migrations, contributing to the genetic legacy of both Sino-Tibetan and Siberian populations through different inland and coastal peopling and expansion routes.
Fig. 1. Phylogeny and geographical distribution of Y chromosome sequences identified in this study.
(A) Maximum parsimony Y chromosome tree illustrating phylogenetic relationships among modern Y chromosome sequences, rooted with four haplogroup B samples. Branches label key mutations and haplogroups, and colors distinguish major haplogroups. (B) Spatial distribution of modern and ancient samples. Different symbols and colors represent distinct data resources: red stars, newly sampled individuals; blue squares, reference samples; green stars, genotyped samples; purple circles, ancient DNA samples. The purple shading indicates the temporal depth of ancient samples. (C) Ethnic and geographic distributions of Y chromosome sequences from ethnolinguistically diverse Chinese individuals. (D) Ethnic and geographic distributions of Y chromosome genotyping data from 15,027 ethnolinguistically distinct individuals. (E) π and Tajima’s D across ethnically and geographically distinct Chinese populations.
RESULTS
Variation discovery, branch heterozygosity, and genomic diversity patterns
We generated 1045 Y chromosome sequences from 26 ethnic groups covering 25 provincial administrative regions of China and merged them with worldwide reference populations from the Human Genome Diversity Project (HGDP) (18), the expanded 1000 Genomes Project (e1KGP) (35), the Simons Genome Diversity Project (SGDP) (36), and other Y chromosome resources from neighboring regions, adding a total of 1668 samples to our analysis (Fig. 1, B and C, and table S1) (37–39). In addition, we genotyped 15,027 unrelated individuals from 38 ethnic groups sampled across all provinces of China (Fig. 1, B and D, and tables S1 and S2). We reanalyzed 977 ancient genomes from raw sequencing data and reclassified haplogroups based on advanced algorithms and our high-resolution phylogenetic framework (fig. S1 and table S3). This strategy aimed to reconstruct an Asian-specific Y chromosome phylogeny and to explore complex patterns of paternal diversification and their impact on the genomic diversity of modern and ancient geographically distinct East Asians. Following alignment, variant calling, and strict quality control, we identified 74,565 SNPs and 8866 insertion and deletion polymorphisms (InDels) within the high-quality 8-Mb regions (table S4), primarily comprising X-degenerated and ampliconic regions. In addition, the average number of Y chromosome variant differences observed between 46 father-son pairs in our final variant calling files was 1.13, indicating the reliability and robustness of our dataset (table S5).
We classified all sequenced and genotyped Chinese individuals, identifying 506 terminal haplogroups in the sequencing dataset and 48 N-related terminal haplogroups in the microarray-based genotyping resources (fig. S2). We then estimated diversity indices, including π and Tajima’s D, across ethnically and geographically diverse populations in the sequencing dataset (Fig. 1E). Across ethnic groups, π was highest in Xibe and Uyghur and remained elevated in Tujia and Mongolian, consistent with larger long-term effective size and/or admixture, although the small sample sizes for Xibe and Uyghur make these rankings less stable. Zhuang showed the lowest π, suggesting stronger drift or past bottlenecks, while its strongly negative Tajima’s D indicates an excess of rare variants compatible with recent growth and/or pervasive purifying selection, supporting multistage demography. Tajima’s D was negative across all major ethnicities, with the most extreme values in well-sampled groups (Han, Hui, Manchu, and Zhuang). Regionally, π peaked in Hubei, Xinjiang, and Gansu but was lowest in Guangdong and Sichuan; Tajima’s D was uniformly negative, especially in Guangxi and Yunnan, supporting broad recent expansion and/or purifying selection. Our sampling enriches underrepresented ethnic groups at the intersection of East Asia and Southeast Asia, but broader sampling from other regions is needed to capture the full scope of ethnolinguistic diversity. Because mutation rates per generation vary across genomic regions, we compared derived-mutation burdens within major Y chromosome lineages and observed lineage-specific differences consistent with distinct histories of bottlenecks or recent expansions. Our estimates of branch heterozygosity further revealed uneven accumulation of mutations along specific branches (fig. S3).
Paternal genomic insights into Asia’s peopling and diversification episodes
We reconstructed a high-resolution Y chromosome phylogeny using deeply rooted African lineages together with modern and ancient Asian genomes from ethnolinguistically and spatiotemporally diverse groups to explore paternal evolution, migration, and dynamics (fig. S1). The fitted maximum-likelihood (ML) model and Bayesian evolutionary analysis by sampling trees (BEAST)-based clustering pattern effectively recapitulated major evolutionary trajectories of Asian Y chromosome lineages. O-related individuals made up 45.4% (1231) of modern samples, followed by N (440, 16.2%), R (260, 9.6%), C (232, 8.6%), and F (132, 4.9%) lineages (Fig. 2A). These clustering patterns thoroughly characterize the phylogenetic relationships among spatiotemporally distinct ancient populations and linguistically diverse modern populations, aligning with prior Y chromosome diversity surveys (24, 26, 40). Notably, this study presents the largest-scale F-related resource to date (n = 132), uncovering rare lineages that highlight the importance of the East-Southeast Asian crossroads, a region historically underrepresented in human genetic research. F lineages were primarily confined to Yunnan and Guangdong provinces and extended into adjacent Southeast Asia, consistent with earlier findings involving the Lahu people from the HGDP (18). Conversely, N1a was found widely across Siberia, Finland, and northern China, whereas N1b was less common and contributed more prominently to the genetic heritage of southwestern China.
Fig. 2. Overview of the Y chromosome phylogeny and geographic distribution of key lineages.
(A) Major founding lineages in 2713 modern individuals at different haplogroup levels. (B) Time-calibrated phylogenetic tree. Branch lengths are proportional to time. Major haplogroups are represented by collapsed triangles, with base width proportional to sample size. Colors within each triangle represent the geographic distribution of samples. (C) Geographic distribution of ancient DNA samples older than 30,000 years. (D) Timeline for ancient samples older than 30,000 years. The detailed information is listed in table S3. (E to H) Geographic distributions of C and D lineages in modern [(E) and (G)] and ancient [(F) and (H)] datasets. Base map source: https://github.com/GuanglinHe1/SA-Neolithic-farmer-migration/ or https://zenodo.org/records/18573902.
The reconstructed phylogeny places the most recent common ancestor of African B and non-African lineages at around 95 kya and reveals deep separation among three ancestral lineages across geographically diverse populations (Fig. 2B). We infer two major divergence waves in out-of-Africa lineages at approximately 70 and 50 kya. The first divergence, which produced the DE-154, C-M130, and FT-M89 lineages, occurred between 70.34 kya [5% highest posterior density (HPD): 65.90 to 74.50 kya] and 69.62 kya (95% HPD: 65.00 to 73.54 kya) (Fig. 2B). The second divergence was followed by a 17,590-year bottleneck, with DE, FT, and C sublineages contributing to the genetic makeup of eastern Eurasians (Fig. 2, B to H). Most of the C sublineages have an important influence on East Asian ancestry, whereas C1-M8 sublineages affected South Asian and Oceanian populations (Fig. 2E). We observed that many Paleolithic Europeans carried C lineages, but their absence in post-LGM and modern European populations suggests potential Y chromosome lineage turnover in pre-Neolithic Europe (Fig. 2F). Ancestral populations with D-related lineages likely formed some of the earliest settlers in East Asia along the southern migration route, and these lineages persist in both modern and ancient groups from the Qinghai-Xizang Plateau and Japan (Fig. 2, G and H).
The FT-M89 lineage played a crucial role in the ancestry of modern non-African populations, with a second wave of divergence and expansion forming distinct continental distribution patterns (Fig. 2B). The root lineages of FT-M89 were confined to East Asia, Southeast Asia, and adjacent regions, whereas later diverging lineages contributed to the ancestry of South Asians (G, H, IJ, and LT), East Asians, Oceanians, and Europeans (Fig. 2B). The large number of F lineages identified here allows for a deeper reconstruction of Asian ancestry, particularly regarding rare Y chromosome lineages. The geographic distribution of modern samples belonging to C, D, and FT indicates a Southeast/East Asia origin of non-African lineages (40), inconsistent with evidence from archeology, linguistics, ancient DNA, and modern autosomal data (41, 42). To investigate prehistoric migration patterns, we systematically reviewed ancient samples dating back more than 30 kya and identified a Paleolithic European individual (F6-620) belonging to the FT-M89 lineage (Fig. 2, C and D). F6-620 individual carries only six derived variants on the branch of F-F24261 (in total, 14 variants). It does not belong to any of the three known sublineages of F-F24261, indicating an extinct lineage in modern humans. Taking ancient DNA and our phylogeny into account, we have argued that the divergence and expansion of the FT-M89 branch, likely near the Middle East/West Asia, followed by stable migrations, contributed to the current Y chromosome diversity in non-African lineages. We estimated the first divergence between East and South Asian lineages at 52.03 kya (95% HPD, 49.48 to 54.77 kya), with consistent divergence times observed among the C and D sublineages. While we cannot exclude the possibility of multiple recurrent divergence events between South Asian and East/Southeast Asian lineages, a deep divergence event was estimated at 66.91 kya (95% HPD, 62.73 to 71.09 kya). Phylogeographical analysis suggested that ancient C lineages were widely distributed in Eurasia (Fig. 2F). However, modern C lineages were only concentrated in eastern Eurasia and Oceania (Fig. 2D), implying a complex pattern of paternal turnover and replacement in West Eurasia, as supported by the genetic replacement of ancient DNA in West Eurasia until ~8.00 kya (43).
Coastal and inland migrations inferred from the deep-rooted F super-haplogroup
The F lineage is a rare and deeply rooted Y chromosome lineage that provides valuable insights into the ancient paternal history of eastern Eurasians, particularly where ancient autosomal DNA is lacking. Despite its importance, research on this lineage remains scarce, with only six F-related individuals previously reported among the Lahu people in the HGDP (18). Here, we assembled the largest-scale F-related dataset to date, comprising 132 individuals from China, and constructed a time-calibrated phylogeny (Fig. 3, A and B; fig. S5; and tables S6 and S7). We estimated that ancestors of the F haplogroup diverged from related lineages about 52.03 kya (95% HPD, 49.48 to 54.79 kya). Divergence within the F lineage occurred around 48.59 kya (95% HPD, 45.73 to 51.49 kya), splitting into three branches: F1-F14885, F2-M427, and F4-Z40734. After a bottleneck lasting ~21,840 years, the F1 lineage diverged into two branches at around 26.75 kya (95% HPD, 24.17 to 29.21 kya). Major lineages then experienced another bottleneck followed by expansion around 5.62 kya (95% HPD, 4.32 to 6.90 kya). The F2 lineage underwent prolonged bottlenecks and later gave rise to several younger lineages around the Neolithic transition (13.76 kya). F2 exhibited more recent splitting events than F3, which was only observed in a Kinh individual (Fig. 3B).
Fig. 3. Phylogeography of haplogroup F in East and Southeast Asia.
(A) Schematic phylogenetic tree of haplogroup F, including East Asian samples. The outgroup was collapse. Branch-defining variants are listed in tables S6 and S7. (B) Time-scaled schematic phylogeny of haplogroup F. Branch lengths are proportional to time, and clades are collapsed as triangles with base widths proportional to sample size. Bars denote the 95% highest posterior density intervals for divergence times. (C) Phylogenetic tree of haplogroup F. Node symbols indicate haplogroup affiliations (top), and the geographical distribution of samples in the phylogenetic tree is shown (bottom). The base map was downloaded from a published dataset (80), with dark gray indicating the inferred coastline ~50 kya. (D) Bayesian skyline plot for haplogroup F based on 132 high-coverage Y chromosome sequences. (E to G) Spatial autocorrelation of haplogroups F, F2-M427, and F1-F14885. The base map could be downloaded from https://github.com/GuanglinHe1/SA-Neolithic-farmer-migration/ or https://zenodo.org/records/18573902.
Rare lineages preserve early population history, making them informative for reconstructing modern paternal structure. To explore migration routes and early peopling models of East Asia, we analyzed the phylogeographical distribution and demographic dynamics of F lineages. We integrated modern Southeast Asian samples belonging to F with relatively low genome coverage into the high-resolution phylogenetic tree (Fig. 3C and fig. S5). The resulting patterns showed that F1 lineages clustered predominantly in Guangdong, while F2 concentrated in Yunnan and nearby regions, indicating a clear separation between inland and coastal areas. These findings suggest that both inland and coastal migration routes contributed to the gene pool of populations at the intersection of East and Southeast Asia. Effective population size estimates indicated that all F-related groups remained small until the middle Neolithic expansion, consistent with the phylogenetic evidence (Fig. 3D). The inference of divergence centers identified two hot spots: one for F2 on the southeastern Qinghai-Xizang Plateau and another in coastal Vietnam and South China (Fig. 3, E and F). These results suggest that deep-rooted F lineages maintained restricted ranges, settling at the crossroads of East and Southeast Asia through both inland and coastal routes. Only one F4-Z40734 sample was identified, located in Vietnam. Overall, the phylogeographic distribution of F points to two routes for modern humans entering East Asia, with F1-F14885 tracing a coastal route and F2-M427 an inland route. Our fine-scale reconstruction of paternal genetic structure supports a gradual expansion model influenced by large-scale population movements and replacements. Notably, F lineages showed close affinity to deep-rooted F lineages in European hunter-gatherers, emphasizing the need for additional Paleolithic genomes from Central and South Asia to clarify male-driven population migration and replacement.
Subsistence-related population divergence events and the Neolithic revolution in East Asia
ML phylogenetic analysis, divergence time estimates, and haplogroup classification assigned 704 individuals to O2-M122 and 527 to O1-F265 (Fig. 4, A and B, and figs. S6 and S7). BEAST-based dating indicated that O1 lineages diverged from O2 lineages before the LGM at 34.16 kya (95% HPD, 31.58 to 36.69 kya) and subsequently underwent distinct demographic trajectories shaped by bottlenecks, expansions, and admixture. Within O1, O1a and O1b diverged at ~33.01 kya (95% HPD, 30.62 to 35.58 kya). O1a experienced a bottleneck at ~16.43 kya, followed by rapid diversification into four lineages within a 680-year window. The dominant O1a1 lineage expanded during the Neolithic transition (~9.79 kya; 95% HPD, 8.74 to 10.84 kya). O1b split into O1b2 and O1b1 at ~30.88 kya (95% HPD, 28.44 to 33.31 ka ago). O1b2 remained in a long-term standstill for ~23,130 years before expanding at ~7.75 kya (95% HPD, 6.62 to 8.79 kya) but produced limited descendant lineages. O1b1 further diversified into O1b1a, O1b1a2, and O1b1a1 at ~23.76 kya (95% HPD, 21.85 to 25.92 kya), with O1b1a1 becoming dominant at ~13.94 kya (95% HPD, 12.65 to 15.29 kya). Phylogeographic analysis suggested a South China origin for O1-related lineages, likely associated with early rice-farming populations, as inferred from phylogeographic distributions, expansion patterns, and multidisciplinary evidence (Fig. 4C and figs. S5 and S8). For O2, the time-calibrated phylogeny placed the initial divergence at ~31.02 kya (95% HPD, 28.61 to 33.35 kya), resulting in a minor O2b lineage (1.12%, comprising eight individuals) and a dominant O2a lineage (98.88%, comprising 709 individuals). Within O2a, O2a2 (524 individuals) and O2a1 (185 individuals) diverged at ~25.87 kya (95% HPD, 23.91 to 27.90 kya), followed by a 5150-year standstill. Four founding lineages emerged: O2a2b (420 individuals) and O2a2a (104 individuals) diverged at ~24.96 kya (95% HPD, 23.02 to 26.86 kya), whereas O2a1b (29 individuals) and O2a1a (156 individuals) diverged at ~18.57 kya (95% HPD, 16.95 to 20.25 kya; Fig. 4B). These sublineages showed prominent post-LGM diversification and Neolithic expansions. Phylogeographic analysis further revealed high frequencies of O2 sublineages in northern China, particularly among ancient northern East Asians and millet-farming populations, linking O2 lineages to the Paleolithic-Neolithic transition of millet agriculture (Fig. 4D and figs. S7 and S9).
Fig. 4. Divergence tempo of the Y chromosome phylogeny for O1-F265, O2-M122, N-M231, and C-M130.
(A) Time-calibrated phylogenetic tree of O1-F265. (B) Time-calibrated phylogenetic tree of O2-M122. (C and D) Geographic distribution of ancient O1 and O2 lineages. The detailed information is listed in table S3. (E) Saturation curves relating sample size to internal node counts across O1-F265, O2-M122, N-M231, and C2-M217. The x axis represents the sample percentage, and the y axis represents the node percentage. Curve colors denote time periods. (F) Interdisciplinary evidence aligned with divergence tempo. Top: Divergence tempo for O1-F265, O2-M122, N-M231, and C2-M217. Middle: Counts of archeological sites in North and South China from (81). Bottom: Holocene temperature variability in China from (82). Base map source: https://github.com/GuanglinHe1/SA-Neolithic-farmer-migration/ or https://zenodo.org/records/18573902.
The tempo of divergence in a time-calibrated Y chromosome phylogeny serves as a proxy for historical population expansions and bottlenecks. We quantified node counts and tree files for the major lineages O1, O2, C2, and N. Node accumulation curves for O1, O2, N, and C2 approached saturation before the Early Neolithic (Fig. 4E), indicating that our sampling captures most pre-Neolithic paternal diversity in ancient China. By integrating genetic, archeological, and paleoclimatic evidence, we placed these demographic patterns in a broader sociocultural context. Divergence tempo increased sharply at ~10 kya and reached a first expansion peak at ~5 to 6 kya (Fig. 4F). The timing and magnitude of this Neolithic expansion varied among O subclades. Previous studies have shown that millet farmers in the Yellow River basin predominantly carried O2. In contrast, rice farmers from the Yangtze River Basin carried O1 (24). The earlier expansion of O2 lineages aligns with the earlier emergence and demographic success of millet agriculture, whereas the later expansion of O1 lineages coincides with the subsequent spread of rice agriculture. Archeological site counts rose after ~9 kya and peaked at ~6 kya, with a stronger and more sustained increase in northern than southern China. Paleoclimatic records indicate a sharp temperature rise at the onset of the Holocene, followed by the Holocene Climatic Optimum (9 to 4 kya) and the subsequent 4.2-ka cooling event. The temporal concordance of genetic, archeological, and climatic evidence supports the hypothesis that differences in subsistence strategy shaped the heterogeneous trajectories of Neolithic paternal expansion.
Millet agriculture-catalyzed Neolithic expansions and their inference around trans-Eurasia
A systematic analysis of spatiotemporally diverse sites from northern China reveals notable regional and temporal dynamics in the paternal haplotype composition of ancient populations, dating back to the Neolithic era. At the Jiangjialing site and other Shandong Houli culture sites (9 to 6 kya), haplogroup N1 predominated among hunter-gatherer populations (11). During the middle Neolithic, N1 remained prevalent among early agricultural communities in the West Liao River and Amur River basins. At the same time, haplogroups O2 and O1 emerged as the dominant lineages in the Yangshao agricultural communities of the Yellow River basin (8). These findings, along with those of previous studies, indicate that O-related ancestral paternal lineages were foundational to early Chinese farmers in the pre-Neolithic and Neolithic periods (24). However, the genetic legacy of N lineages limits our understanding of their role in the Neolithic transition in ancient China. We explored phylogenetic relationships and phylogeographic patterns on these ancient samples and 429 modern N lineages, comprising 173 N1a and 256 N1b Y chromosome genomes (Fig. 5A and figs. S10 and S11). BEAST-based coalescent time estimation revealed that N lineages diverged into N1 and N2 at 25.45 kya (95% HPD, 23.31 to 27.76 kya) (Fig. 5A). The N2 lineage was extremely rare, with only one sample assigned to this haplogroup. The N1 lineages underwent a bottleneck of ~5354 years and diverged into N1a and N1b during the LGM of ~20.08 kya (95% HPD, 18.43 to 21.67 kya). N1a, which is widely distributed in Northern Eurasia and Siberia, split into three founding lineages: N1a2, N1a3, and N1a1, between 16.91 and 17.98 kya. N1a2 and N1a3 are prevalent among the northern Han Chinese, Mongolian, and Manchu populations, whereas N1a1 is dominant in the Yakut, Finnish, and other Western Eurasian populations. N1b remained stable for 3000 years before separating into N1b1 and N1b2 at ~17.08 kya (95% HPD, 15.75 to 18.56 kya). N1b1 made a limited contribution to modern Chinese populations, with only 51 sequences classified into this lineage. In contrast, 205 N1b2 lineages were identified in southern Han Chinese and Tibeto-Burman-speaking populations.
Fig. 5. The phylogeny of haplogroup N.
(A) Calibrated phylogeny of haplogroup N inferred with BEAST, with branch lengths proportional to time. Key mutations and haplogroups are listed on the corresponding branches. Tip-circle colors represent geographic distribution: North China (gray), Southwest China (light blue), South China (yellow), Northwest China (pink), Europe (green), Siberia (blue), Central Asia (orange), and others (purple). Colored blocks on the right denote the language family of samples within each branch. Blue numbers on specific branches indicate the elapsed time from the lineage’s origin to the emergence of new sublineages. Branch colors indicate N-derived sublineages. Branch-specific variants are listed in tables S8 and S9. (B) Geographic distribution of 21,562 N samples across Eurasia. The schematic tree summarizes relationships among N sublineages. Bar height represents the normalized frequency. (C) Geographic distribution of ancient N lineages. (D) Spatial and temporal distributions of ancient DNA belonging to haplogroup N. The color of each dot represents geographic distribution. Detailed information is available in table S3. (E) Fine-scale phylogeographic distribution of N1b-F2603 among ancient and modern samples. The blue bar below the phylogeny denotes modern samples, while red and green bars indicate ancient samples belonging to N1b2 and N1b1, respectively. Line colors denote lineages, and the color gradient for ancient samples reflects archeologically defined time periods.
Since some Y chromosome data were derived from phylogenetic studies rather than large-scale population-based analyses, which can introduce sampling bias, to address this, we examined the geographic distribution of N-M231 and its sublineages among 15,027 Chinese individuals and 6521 Siberian samples across Eurasia (Fig. 5B and fig. S12) (37). A high proportion of N1a was found in northern Europe, Siberia, and Mongolia, and both N1b and N1a showed distinct geographic distributions across northern, southern, northwestern, and southwestern China (Fig. 5B). N1a2-P43 was primarily present in Siberia, followed by Mongolia and Europe, while its sister lineage, N1a2-F1101, was predominantly distributed in northwest China. N1a3-F949 and N1a3-F4063 were predominantly found in North China, whereas their downstream sublineages, N1a1-B211, N1a1-M2019, and N1a1-CTS6967, were mainly distributed in Siberia, Central Asia, and Europe. Considering the phylogeographic distribution of N1a-F1206 and its sublineages, the spread from southern Siberia into Eastern Europe likely occurred between 9.84 and 11.91 kya. N1b-F2603 is predominantly found in China, with N1b2-M1819 occurring primarily in Southwest China, followed by Northwest China, South China, and North China.
We analyzed ancient DNA from eastern Eurasia and identified 167 individuals with N lineages, including 92 N1a lineages, primarily found across Siberia. These included six Eneolithic Shamanka individuals belonging to N1a2a1a and 62 N1b lineages from northern China and the Qinghai-Xizang Plateau (Fig. 5, C and D). The oldest ancient DNA assigned to N2-B482 was mainly recovered from southern Siberia (Fig. 5, C and D), indicating this region as the likely origin of N-M231 based on its phylogenetic features. Ancient European samples are primarily linked to N1a1-M2019 and N1a1-CTS6967, with the earliest dating back to around 2.61 kya (Estonia_IA) (Fig. 5, C and D). N1b1 was geographically limited to middle Neolithic Shandong, with additional occurrences at 3.75-kya Erdaojingzi and 1.65-kya LongLongRak (N1b) (Fig. 5E). However, no evidence of this lineage has been found in Shandong during the historical period, suggesting a possible paternal lineage replacement or discontinuity between ancient and modern populations of the Shandong Peninsula (table S3). Spatial autocorrelation analysis shows that modern N1b1 lineages are concentrated in Anhui, Zhejiang, and Fujian, indicating a likely southward coastal migration after the Neolithic (fig. S13). Ancient DNA places early occurrences of N1b2 in northeastern Qinghai-Xizang Plateau, particularly at the Zongri site (Fig. 5E). By ~4.00 kya, its distribution shifted to southwestern China and the core region of Qinghai-Tibet Plateau, aligning with that observed in present-day populations (fig. S13). The high prevalence of this lineage among Tibeto-Burman speakers (Fig. 2A) suggests that southward movements were associated with the spread of Tibeto-Burman groups. These findings reveal distinct geographic centers of origin for N sublineages (Fig. 6). N1a was centered in Siberia, while N1b sublineages expanded along different inland and coastal routes (figs. S12 and S13). N1b1 was concentrated in northern coast of eastern China, whereas N1b2 had a high-frequency center in southwestern China (figs. S12 and S13).
Fig. 6. Summary landscape of deep population history across eastern Eurasia inferred from Y chromosome sequences.
Three major migration events shaped the genetic landscape of East and Southeast Asia. The black and gray arrows represents the initial entry of early Asians carrying deeply diverged F lineages, which moved along both inland and coastal routes. Blue and yellow arrows denote a second wave of population expansion driven by East Asian indigenous O lineages. The light orange arrows highlights three subsequent migration patterns: westward inland, southward coastal, and southward inland movements. The map labels key ancient samples that anchor and contextualize these migration waves.
DISCUSSION
Rapid diversification of early Asians and coastal and inland migration routes for Paleolithic peopling of East Asia
East Asia, serving as a vital corridor between Southeast Asia and Siberia, played a pivotal role in the Paleolithic migration of eastern Eurasia and the spread of Neolithic cultures. This region exhibits remarkable genetic, linguistic, cultural, biological, and phenotypic diversity (44). To reconstruct Chinese paternal lineages at high resolution, we compiled Y chromosome sequences from 2713 ethnolinguistically diverse individuals and 977 ancient people, along with high-density genotyping data from 15,027 individuals. A fully resolved, time-calibrated Y chromosome phylogeny was built to capture early Asian diversification and subsequent population dynamics. The findings revealed that multiple out-of-Africa lineages established early footholds in Asia, underwent intricate and rapid Paleolithic diversifications, followed by widespread Neolithic expansions. High-resolution phylogenetic analysis concentrated on early Asian lineages, including O, C, D, and F haplogroups, as well as Neolithic-expanded lineages such as coastal and inland F, farmer-related O1 and O2, and hunter-gatherer and farmer-sharing N lineages. The reconstructed phylogeny confirmed the rapid diversification of multiple C-related sublineages among Aboriginal Australians, South Asians, and East Asians, highlighting their deep connections with D lineages from China and Japan (1, 2, 16, 26). A previously unreported F-related phylogeny was also identified based on the largest-scale sequencing effort to date, revealing the divergence of Chinese aboriginal F genomes from Southeast Asia before the LGM (40).
These F genomes formed three deeply rooted lineages that contributed to both Southeast Asian and East Asian populations through coastal and inland migration routes. Most F genomes in our dataset come from Guangdong and Yunnan provinces, whereas prior studies mainly reported F lineages in ethnolinguistically diverse Southeast Asian populations (45, 46). The phylogenetic patterns and the widespread distribution of rare F lineages suggest that the settlement of East Asia occurred via both inland and coastal routes. During pre-LGM low sea levels, geographic connectivity across Southeast Asia may have strengthened coastal connections between Southeast Asia and East Asia during the Paleolithic (11). Archeological evidence also supports the existence of inland migration routes across Thailand, Myanmar, and southwestern China (47). Ancient DNA from the Xingyi site reveals direct genetic connections between early Neolithic populations of Yunnan, ancient Tibetans in the highlands, and Southeast Asian Hòabìnhian hunter-gatherers, shedding light on the genetic origins of early Asians (4, 21). Furthermore, Holocene population movements involving modern Austroasiatic groups from central Yunnan Neolithic populations indicate that migration routes along the eastern and southeastern Qinghai-Xizang Plateau were crucial (21). Historically, major trade and cultural routes such as the Tea-Horse Ancient Road, the Southern Silk Road (Shu-Shen-Du Dao), and the Lancang-Mekong River Corridor connected southwestern China with Laos, Myanmar, Thailand, and Cambodia, facilitating trade and transportation. Geographic corridors from southwestern China and Southeast Asia also highlight the importance of inland routes in shaping populations across Southeast and East Asia (5, 17, 48). F lineages, although rare, have substantial potential for tracing the deep paternal genetic legacy of East Asian populations. The phylogenetic evidence suggests that both Paleolithic coastal connections and inland interactions contributed to the region’s early paternal genetic makeup (Fig. 6). Our work acknowledges the complexity of migration dynamics, including southward migrations from core areas to Southeast Asia, as demonstrated by populations such as those in the Malay Archipelago. As complex cultural or geographical corridors exist in this region, further sampling collections may provide deep insights into the complex routes for past human migrations.
Neolithic expansion events and their associations with subsistence strategies
Ancient DNA studies have uncovered complex population histories among early East Asians (4, 5, 11, 12, 21, 49), including repeated movements and replacements near major Chinese agricultural centers and surrounding areas such as the Amur River basin and the Qinghai-Xizang Plateau (49). While the paternal genetic legacy carried by the Y chromosome remains elusive, it offers valuable insights into evolutionary timelines and paternal lineages. Here, we assembled the largest modern and ancient Y chromosome dataset focused on O1, O2, and N, which diverged from other non-African lineages or from each other, during the Middle Paleolithic. We then mapped their Paleolithic divergence and Neolithic spread, aligning with archeological records of early settlement in East Asia (Fig. 6). Our ML-based phylogenetic reconstruction supported established models and revealed lineage-specific histories that include early splits, prolonged bottlenecks, and rapid expansions. We detailed the demographic histories of ancient non-African lineages C, D, and F, which have few modern descendants in farming regions, as well as O lineages associated with rice and millet agriculture, and N lineages shared among farmers, hunter-gatherers, and herders. C and D lineages remained stable through the pre-Neolithic period, contributing to early settlements in Australia, the Qinghai-Xizang Plateau, and Japan (16, 24–26, 31, 45, 46). FT-M89 and GHIJK-CTS2254 lineages showed patterns consistent with strong sensitivity to resource constraints and survival advantages. Coastal lineages diverged earlier but formed smaller populations than inland lineages that diverged later. O1-F265 and O2-M122 exhibited distinct expansion histories in timing, scale, and geographic reach, with O1 splitting earlier and producing more descendant lineages than O2. Meanwhile, N lineages, once widespread among ancient Siberian and Shandong hunter-gatherers, are now common in northern Europe, Siberia, and both coastal and inland East Asia. These lineages display different sublineages with varied expansion routes. Our detailed reconstruction of their origins and migration paths, based on paternal DNA from archeological remains and high-resolution modern samples, suggests that these lineages originated from the Mongolian Plateau and southern Siberia during the Paleolithic and Neolithic periods. Variations in subsistence strategies during the late Paleolithic led to diverse migration pathways: N1a lineages spread westward, influencing the origins of Altaic and Indo-European groups in western Eurasia and northern Siberia; farming-related N1b lineages were dispersed inland and southward along coastal routes, shaping the genetic diversity of coastal regions such as ancient Shandong, modern Jiangzhe, the highland Qinghai-Xizang Plateau, and southwestern China.
Ancestral East Asians likely adopted various subsistence strategies, survival advantages, descent rules, and reproductive success rates, contributing to different patterns of paternal genetic expansion. Distinct paternal lineages linked to specific subsistence strategies may have emerged from diverse social structures, inbreeding dynamics, and historical developments. This can be inferred from the genetic legacy of the Y chromosome. Hunter-gatherer–associated F lineages showed prolonged stasis and limited contributions to modern populations, and they remain geographically localized. In contrast, farmer-associated lineages diversified rapidly and spread widely, consistent with demographic growth linked to cultural innovation. N lineages may reflect shared ancestral connections between ancient hunter-gatherers and early farmers, with distinct sublineages showing varied anthropological traits in terms of origin, distribution, and migration history. Archeological evidence in Siberia suggests the prevalence of hunter-gatherer cultures, whose subsistence strategies were shaped by the region’s harsh climatic conditions. The LGM constrained human populations to resource-scarce environments, favoring the mobile hunting of large game species such as mammoths and reindeer, which supported the possibility of the northern route of human migration and followed complex admixture. Moreover, early farming cultures in North China emerged during the Neolithic period, with the cultivation of millet marking a noteworthy shift in subsistence (50). These populations benefited from a temperate climate that facilitated the development of sedentary agricultural societies. South China followed a distinct subsistence trajectory, with rice cultivation becoming the dominant form of agriculture in its subtropical environment, characterized by warm temperatures and high rainfall (51, 52). This climate enabled the domestication of wetland rice, which, in turn, led to the development of complex social hierarchies and settlement patterns (4, 5, 11). The tropical climate in Southeast Asia supported diverse subsistence strategies, from foraging to early farming (4, 45, 46). Hunter-gatherer groups in this region exploited the rich biodiversity of tropical forests, whereas early agriculturalists adopted rice and tuber cultivation as part of a mixed subsistence strategy (4). These subsistence practices reflect adaptations to varying environmental conditions and resource availability across the region. Despite differences in climate and resources, both foraging and farming cultures display noteworthy innovation and flexibility, underscoring the importance of subsistence strategies in shaping human adaptation to diverse ecological landscapes. By combining the different patterns of Y chromosome paternal genetic history and variations in sociocultural factors, we suggest that subsistence strategies are the most noteworthy influencing factors of prehistoric migrations in eastern Eurasia.
In summary, we present a high-resolution Y chromosome resource integrating high-coverage sequencing and dense genotyping across ancient and modern samples. This resource allowed us to build the most comprehensive time-calibrated Y chromosome phylogeny for Asia, revealing both common and previously undocumented paternal lineages. Our analyses showed that many modern Asian men trace ancestry to O, C, N, Q, and R, while rare F and N1b lineages, mainly found in southwestern China, represent deeply diverged roots of early Asian lineages. Phylogeographic modeling suggests that early Asian settlement involved both inland and coastal migration routes, with Paleolithic expansions around the Himalayas contributing to eastern Eurasian diversity. Neolithic transitions, driven by millet- and rice-farming populations, led to the rapid diversification of O and N1b sublineages. N1b lineages are closely linked to the emergence of Tibeto-Burman groups, while N1a lineages persisted among pre-Neolithic hunter-gatherers and herders. We also find evidence of contrasting demographic outcomes: continuity among N1b2-related inland millet farmers and genetic turnover among coastal East Asians, especially in N1b1 founders. These findings demonstrate how deep Paleolithic divergences, prolonged bottlenecks, and multiple Neolithic expansions collectively shaped the complex paternal genetic landscape of East Asia and provide a framework for integrating Y chromosome data into broader studies of human migration and adaptation.
MATERIALS AND METHODS
Ethical approval and sample collection
This study received approval from the Medical Ethics Committee at West China Hospital, Sichuan University (approval no. 2021-190, 2024-2010, and 2024-1870). All participants provided written informed consent before sample collection. The procedures were strictly followed in accordance with the guidelines of the Human Genetic Resources Administration of China and complied with the ethical principles outlined in the 2013 revision of the Declaration of Helsinki (53). The project involved collaboration with scientists from the YanHuang Cohort Consortium and the GSRD-100KWCH working groups, bringing together experts in anthropology, genetics, statistics, and medicine. Fieldwork and genomic research were thoroughly conducted on ethnolinguistically diverse Chinese populations, with linguistically skilled researchers ensuring effective communication. Contributions from specialists in ancient DNA, population genetics, linguistics, ethnic studies, and medical science provided careful ethical oversight. Engagement with local minority communities was maintained through transparent sharing of research goals and results. A total of 15,961 samples from the YanHuang cohort and 111 samples from GSRD-100KWCH were collected with informed consent.
Genotyping and sequencing
A total of 15,027 saliva samples were genotyped using high-density Illumina microarrays, while the remaining 934 saliva samples were sequenced with a Y chromosome capture panel targeting 8.4 Mb of high-quality regions on the Salus Pro sequencer (Shenzhen Salus BioMed Co. Ltd.). We also generated WGS data for 111 individuals on the BGI DNBSEQ-T7 platform at an average depth of 30×. Within the sequenced set, we prioritized previously underrepresented N (n = 269) and F (n = 122) lineages sampled across different Chinese populations. To investigate paternal genetic history in Southwest China, we analyzed 654 high-coverage genomes from the crossroads of East and Southeast Asia, specifically Yunnan and Guangxi provinces, spanning Sino-Tibetan, Tai-Kadai, and Hmong-Mien speakers. To fully understand the region’s demographic history, we combined our data with 918 samples from previously published studies for phylogenetic analysis. Furthermore, we obtained FASTQ files from the HGDP (18), e1KGP (35), the SGDP (36), and other Y chromosome resources from neighboring regions (37, 38), adding 1664 samples to our dataset. An additional four samples attributed to the B lineage from SGDP were used to root the phylogenetic tree. The final Y chromosome sequencing dataset comprised 2713 samples, including 1045 recently generated samples (table S1). To systematically investigate the phylogeographic distribution of N lineages in East Asia and Siberia, we compiled our dataset with the paternal genetic resource of 6521 male individuals from Eurasian populations drawn from a previous study (table S2) (37).
Mapping, variant calling, filtering, and validation
Trimmomatic was used to perform quality control and remove adapters from the raw FASTQ files. The clean reads were then mapped to the GRCh37 human reference assembly using Burrows-Wheeler Aligner v0.7.17 (mem -M) (54). Sambamba v0.5.1 was used to remove duplicates and sort the Binary Alignment Map (BAM) files. Variant calling was performed in haploid mode for each BAM file using Genome Analysis Toolkit (GATK) HaplotypeCaller with a Phred-scaled quality score ≥ 20 (55). The high-confidence overlapping region of the Y chromosome between WGS and targeted sequencing was approximately 8.4 Mb (table S4). Consequently, variant calling was conducted via GATK GenotypeGVCFs (including nonvariant sites), with a focus on the 8.4-Mb region. To ensure the accuracy of downstream analysis, we applied several filters via BCFtools and Perl scripts: (i) InDels and multiple alleles were removed; (ii) sites with genotype quality (QUAL) > 30 were retained; (iii) heterozygous calls with the fraction of reads supporting the called allele ≤0.85 were set to the missing data; (iv) sites with at least two supporting reads were retained; and (v) sites with more than 5% missing genotypes across all samples were removed. After filtering, our dataset was reduced to 7,099,789 base pairs (bp) and 74,565 high-confidence SNPs from 2713 samples, which were used for phylogenetic reconstruction and coalescent time estimates. To ensure the robustness and accuracy of our filters, we also compared the average number of variant differences in our final VCF file between 46 father-son pairs. Theoretically, the average number of variant differences between father-son pairs is less than 1.2 in the 8.4-Mb region (table S5).
Tree reconstruction and coalescence analysis
We reconstructed the Y chromosome phylogeny via ML and Bayesian Markov chain Monte Carlo (MCMC) approaches. ML tree reconstruction was performed using RAxML v8.2.12 with 200 rapid bootstrap replicates and the GTRCAT substitution model (56). The haplogroup label of each sample was assigned using HaploGrouper based on the International Society of Genetic Genealogy v15.7 nomenclature (57). Haplogroup B was used to precisely root the tree. Phynder was used to assign all identified variants to the branch of the ML tree, which was subsequently curated manually (https://github.com/richarddurbin/phynder). The tree file was visualized and embellished in Interactive Tree of Life v6 (58). The annotated files were generated using the R package itol.toolkit (59). Coalescence time estimates were generated using Bayesian MCMC approaches implemented in BEAST v1.10.4 (60). To reduce computational costs, we conducted four lineage-based analyses for F-M89, O1-F265, O2-M122, and N-M214. JModelTest v2.1.10 was used to determine the most suitable model (61). Four parallel analyses were conducted for each haplogroup using different random numbers of seeds but identical parameters. The analyses used the GTR + I + G substitution model with a strict clock, a Bayesian skyline coalescent tree prior, and a piecewise-linear coalescent model with the number of groups set to 10. A mutation rate of 0.76 × 10−9 mutations/nucleotide/year [95% confidence interval (CI): 0.67 × 10−9 to 0.86 × 10−9] with a normal distribution was used in the analysis. An extended sample set was selected to increase the representativeness of the phylogeny, including four haplogroup B samples used as outgroups. For both BEAST analyses, node GT-F1329, with an estimated age of 51,557 years (95% CI: 50,206 to 52,917), served as the calibration point. For haplogroup F-M89, each run consisted of 100 million chains, with logging occurring every 10,000 steps. For haplogroups O1-F265, O2-M122, and N-M214, each run consisted of 60 million chains, with logging every 3000 steps. Each log file was inspected using Tracer v1.7 to ensure an effective sample size of at least 200. The four parallel chains were combined via LogCombiner v1.10.4. The maximum clade credibility (MCC) tree was generated via TreeAnnotator, with the initial 25% discarded as burn-in. The MCC tree was visualized via FigTree v1.4.4 (http://tree.bio.ed.ac.uk/software/figtree/). In addition, we used the rho statistics to evaluate the coalescent ages of each node in haplogroups O1-F265, O2-M122, F-M89, and N-M214 for unbiased and overlapping time estimates. A mutation rate of 0.76 × 10−9 (95% CI: 0.67 × 10−9 to 0.86 × 10−9) mutations per nucleotide per year was applied. With this mutation rate, a mutation in the 7.1-Mbp region corresponds to 185.32 (163.78 to 210.22) years (table S10).
Bayesian skyline plots
We further examined the demographic history of haplogroup F-M89 using 132 high-coverage Y chromosome sequences within a Bayesian skyline analysis framework. The parameters were similar to those described earlier, except that each MCMC simulation ran for 500 million generations, sampling every 5000 generations. The Bayesian skyline plot was reconstructed and visualized using Tracer v1.7. The dynamics of the other three lineages related to O and N were also estimated using the same parameters mentioned earlier.
Placing ancient DNA in the phylogenetic tree
To elucidate the evolutionary history of haplogroups N-M231 and F-M89, we analyzed these haplogroups in 15,461 ancient samples from our in-house database and website (https://indo-european.eu/ancient-dna/), which contains the most comprehensive dataset reported to date. A total of 168 ancient samples belonging to haplogroups N and F were collected for further analysis. In addition, we used ancient DNA from Eurasia to explore paternal evolutionary history comprehensively (62–76). We used phynder to assign informative SNPs related to haplogroups N-M231 and F-M89 to the branch of the phylogenetic tree. Detailed information related to the position and annotation of the variants is listed in tables S6 to S9. PathPhynder (77) was used to place these ancient DNA in the phylogenetic tree using the best path method. In this study, we integrated ancient DNA into the phylogenetic tree of modern humans and positioned modern individuals from Southeast Asia within the reconstructed F-related phylogenetic tree (46, 78, 79). This approach was used to investigate the phylogeographic distribution of the rare haplogroup F at the intersection of human migration. The sample information allocated to haplogroup F from Southeast Asia is listed in table S11.
Spatial geographic distribution
Surfer v8.0 (Golden Software, Inc., Golden, CO, USA) was used to generate contour maps of spatial frequencies related to the N-M231 and F-M89 lineages and their sublineages via the kriging procedure. The maps were generated via RStudio software and a published dataset (80). Spatial autocorrelation and haplogroup frequency spectra were characterized via ArcMap.
Estimates of the divergence tempo of the Y chromosome phylogeny
We estimated the divergence tempo for O1-F265, O2-M122, N-M231, and C2-M217 by counting the number of internal nodes. We performed a sliding window approach with a length of 200 years and a shift step of 200 years. The tempo curve was smoothly implemented as the geom_stream() function in R (ggplot2 package), with the smoothing parameter type set to “ridge” to apply ridge line smoothing.
Divergence tempo saturation in phylogenetic trees
We initially randomly selected 10 samples from the phylogenetic tree to establish the baseline tree topology. Subsequently, an iterative process was implemented where an additional 10 samples were incrementally added in each iteration. For each generated phylogenetic tree, the number of internal nodes was recorded as the sample size increased. A sliding window analysis was performed, starting 25,000 years ago, using a window size of 1000 years with a step size of 1000 years. At each time point, the number of internal nodes within the window was calculated. This process was repeated across 10 iterations to account for sampling variability. The internal node counts were averaged over these iterations and compared against the internal node count of the full phylogenetic tree, yielding a saturation curve that reflects the relationship between sample size and the number of internal nodes.
Integrating interdisciplinary data and evidence
To enhance our understanding of the tempo of divergence in the Y chromosome phylogeny, we integrated archeological and paleoclimate data within the same temporal framework. For the archeological data, we calculated the number of archeological sites in North China and South China according to Hosner et al. (81). The trend curve was smoothed via the same approach as described above. For paleoclimatic data, we used Holocene temperature variability in China, inferred from branched glycerol dialkyl glycerol tetraethers (brGDGTs) extracted from sediments of Lake Yangzonghai (82).
Acknowledgments
We thank the China National Supercomputing Center in Chengdu for providing sequencing data storage and computational analysis for this study. We express our gratitude to all the volunteers who contributed to this study.
Funding:
This study was supported by the National Natural Science Foundation of China (82572153 for M.W. and 82402203 for G.H.). G.H. also acknowledges support from the Major Project of the National Social Science Foundation of China (23&ZD203), the Center for Archaeological Science of Sichuan University (23SASA01), and the Sichuan Science and Technology Program (2024NSFSC1518). M.W. also acknowledges support from the Open Project of the Key Laboratory of Forensic Genetics of the Ministry of Public Security (2024FGKFKT02). Huijun Yuan acknowledges support from the 1.3.5 Project for Disciplines of Excellence, West China Hospital, Sichuan University (ZYJC20002).
Author contributions:
Conceptualizations: M.W., Z.W., G.H., and Huijun Yuan. Methodology: M.W., Z.W., G.H., and Y.J. Formal analysis: M.W., Z.W., Haibing Yuan, and L.L. Visualization: M.W., Z.W., Y.F., and J.C. Validation: Z.W. and J.C. Data curation: Huijun Yuan, Y.F., and R.T. Writing—original draft: M.W., G.H., Z.W., and Huijun Yuan. Writing—review and editing: M.W., G.H., Z.W., and Huijun Yuan. Investigation: Huijun Yuan, Y.J., L.L., Y.F., Y.L., and F.B. Resources: Huijun Yuan, R.T., and C.L. Funding acquisition: M.W., G.H., and Huijun Yuan. Project administration: M.W., G.H., and Huijun Yuan. Supervision: G.H. and Huijun Yuan.
Competing interests:
The authors declare that they have no competing interests.
Data, code, and materials availability:
All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials. There are no new materials generated in the study. Genome-wide variation data for ancient individuals were obtained from the Allen Ancient DNA Resource (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW&ve1984). The results of the analyses have been included in the Supplementary Materials and deposited in the OMIX database (https://ngdc.cncb.ac.cn/omix/) under accession number OMIX007713. The variation data reported in this paper have been deposited in the Genome Variation Map at the National Genomics Data Center, Beijing Institute of Genomics, Chinese Academy of Sciences and China National Center for Bioinformation under accession number GVM000878 [Genome Variation Map (cncb.ac.cn)]. Base maps are available in the GitHub repository (https://github.com/GuanglinHe1/SA-Neolithic-farmer-migration/) and Zenodo (https://zenodo.org/records/18573902). The analysis code is available via Zenodo (https://zenodo.org/records/18753608). Access to and use of data must comply with the regulations of the People’s Republic of China on the administration of human genetic resources. To comply with the informed-consent agreements, the data are available only upon written request for population history research; they may not be used for medical or disease-related research, natural-selection analyses, commercial purposes, or individual identification. Requesters may not share the data with third parties or deposit it in other public databases without prior written approval. The raw data can be provided by K. Wu pending scientific review and a completed material transfer agreement through Sichuan University. Requests for the raw data should be submitted to K. Wu (wuke2026@wchscu.cn).
Supplementary Materials
The PDF file includes:
Figs. S1 to S13
Legends for tables S1 to S11
Legend for data S1
Other Supplementary Material for this manuscript includes the following:
Tables S1 to S11
Data S1
REFERENCES
- 1.Jin L., Su B., Natives or immigrants: Modern human origin in east Asia. Nat. Rev. Genet. 1, 126–133 (2000). [DOI] [PubMed] [Google Scholar]
- 2.Su B., Xiao J., Underhill P., Deka R., Zhang W., Akey J., Huang W., Shen D., Lu D., Luo J., Chu J., Tan J., Shen P., Davis R., Cavalli-Sforza L., Chakraborty R., Xiong M., Du R., Oefner P., Chen Z., Jin L., Y-Chromosome evidence for a northward migration of modern humans into Eastern Asia during the last Ice Age. Am. J. Hum. Genet. 65, 1718–1724 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Yang M. A., Gao X., Theunert C., Tong H., Aximu-Petri A., Nickel B., Slatkin M., Meyer M., Paabo S., Kelso J., Fu Q., 40,000-year-old individual from Asia provides insight into early population structure in Eurasia. Curr. Biol. 27, 3202–3208.e9 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.McColl H., Racimo F., Vinner L., Demeter F., Gakuhari T., Moreno-Mayar J. V., van Driem G., Gram Wilken U., Seguin-Orlando A., de la Fuente Castro C., Wasef S., Shoocongdej R., Souksavatdy V., Sayavongkhamdy T., Saidin M. M., Allentoft M. E., Sato T., Malaspinas A. S., Aghakhanian F. A., Korneliussen T., Prohaska A., Margaryan A., de Barros Damgaard P., Kaewsutthi S., Lertrit P., Nguyen T. M. H., Hung H. C., Minh Tran T., Nghia Truong H., Nguyen G. H., Shahidan S., Wiradnyana K., Matsumae H., Shigehara N., Yoneda M., Ishida H., Masuyama T., Yamada Y., Tajima A., Shibata H., Toyoda A., Hanihara T., Nakagome S., Deviese T., Bacon A. M., Duringer P., Ponche J. L., Shackelford L., Patole-Edoumba E., Nguyen A. T., Bellina-Pryce B., Galipaud J. C., Kinaston R., Buckley H., Pottier C., Rasmussen S., Higham T., Foley R. A., Lahr M. M., Orlando L., Sikora M., Phipps M. E., Oota H., Higham C., Lambert D. M., Willerslev E., The prehistoric peopling of Southeast Asia. Science 361, 88–92 (2018). [DOI] [PubMed] [Google Scholar]
- 5.Wang C. C., Yeh H. Y., Popov A. N., Zhang H. Q., Matsumura H., Sirak K., Cheronet O., Kovalev A., Rohland N., Kim A. M., Mallick S., Bernardos R., Tumen D., Zhao J., Liu Y. C., Liu J. Y., Mah M., Wang K., Zhang Z., Adamski N., Broomandkhoshbacht N., Callan K., Candilio F., Carlson K. S. D., Culleton B. J., Eccles L., Freilich S., Keating D., Lawson A. M., Mandl K., Michel M., Oppenheimer J., Ozdogan K. T., Stewardson K., Wen S., Yan S., Zalzala F., Chuang R., Huang C. J., Looh H., Shiung C. C., Nikitin Y. G., Tabarev A. V., Tishkin A. A., Lin S., Sun Z. Y., Wu X. M., Yang T. L., Hu X., Chen L., Du H., Bayarsaikhan J., Mijiddorj E., Erdenebaatar D., Iderkhangai T. O., Myagmar E., Kanzawa-Kiriyama H., Nishino M., Shinoda K. I., Shubina O. A., Guo J., Cai W., Deng Q., Kang L., Li D., Li D., Lin R., Nini, Shrestha R., Wang L. X., Wei L., Xie G., Yao H., Zhang M., He G., Yang X., Hu R., Robbeets M., Schiffels S., Kennett D. J., Jin L., Li H., Krause J., Pinhasi R., Reich D., Genomic insights into the formation of human populations in East Asia. Nature 591, 413–419 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Deng L., Pan Y., Wang Y., Chen H., Yuan K., Chen S., Lu D., Lu Y., Mokhtar S. S., Rahman T. A., Hoh B. P., Xu S., Genetic connections and convergent evolution of tropical indigenous peoples in Asia. Mol. Biol. Evol. 39, msab361 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Luo L., Wang M., Liu Y., Li J., Bu F., Yuan H., Tang R., Liu C., He G., Sequencing and characterizing human mitochondrial genomes in the biobank-based genomic research paradigm. Sci. China Life Sci. 68, 1610–1625 (2025). [DOI] [PubMed] [Google Scholar]
- 8.Wen B., Li H., Lu D., Song X., Zhang F., He Y., Li F., Gao Y., Mao X., Zhang L., Qian J., Tan J., Jin J., Huang W., Deka R., Su B., Chakraborty R., Jin L., Genetic evidence supports demic diffusion of Han culture. Nature 431, 302–305 (2004). [DOI] [PubMed] [Google Scholar]
- 9.Mao X., Zhang H., Qiao S., Liu Y., Chang F., Xie P., Zhang M., Wang T., Li M., Cao P., Yang R., Liu F., Dai Q., Feng X., Ping W., Lei C., Olsen J. W., Bennett E. A., Fu Q., The deep population history of northern East Asia from the Late Pleistocene to the Holocene. Cell 184, 3256–3266.e13 (2021). [DOI] [PubMed] [Google Scholar]
- 10.Wang T., Wang W., Xie G., Li Z., Fan X., Yang Q., Wu X., Cao P., Liu Y., Yang R., Liu F., Dai Q., Feng X., Wu X., Qin L., Li F., Ping W., Zhang L., Zhang M., Liu Y., Chen X., Zhang D., Zhou Z., Wu Y., Shafiey H., Gao X., Curnoe D., Mao X., Bennett E. A., Ji X., Yang M. A., Fu Q., Human population history at the crossroads of East and Southeast Asia since 11,000 years ago. Cell 184, 3829–3841.e21 (2021). [DOI] [PubMed] [Google Scholar]
- 11.Yang M. A., Fan X., Sun B., Chen C., Lang J., Ko Y. C., Tsang C. H., Chiu H., Wang T., Bao Q., Wu X., Hajdinjak M., Ko A. M., Ding M., Cao P., Yang R., Liu F., Nickel B., Dai Q., Feng X., Zhang L., Sun C., Ning C., Zeng W., Zhao Y., Zhang M., Gao X., Cui Y., Reich D., Stoneking M., Fu Q., Ancient DNA indicates human population shifts and admixture in northern and southern China. Science 369, 282–288 (2020). [DOI] [PubMed] [Google Scholar]
- 12.Liu Y., Mao X., Krause J., Fu Q., Insights into human history from the first decade of ancient human genomics. Science 373, 1479–1484 (2021). [DOI] [PubMed] [Google Scholar]
- 13.Wang H., Yang M. A., Wangdue S., Lu H., Chen H., Li L., Dong G., Tsring T., Yuan H., He W., Ding M., Wu X., Li S., Tashi N., Yang T., Yang F., Tong Y., Chen Z., He Y., Cao P., Dai Q., Liu F., Feng X., Wang T., Yang R., Ping W., Zhang Z., Gao Y., Zhang M., Wang X., Zhang C., Yuan K., Ko A. M., Aldenderfer M., Gao X., Xu S., Fu Q., Human genetic history on the Tibetan Plateau in the past 5100 years. Sci. Adv. 9, eadd5582 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.He G., Wang M., Zou X., Chen P., Wang Z., Liu Y., Yao H., Wei L. H., Tang R., Wang C. C., Yeh H. Y., Peopling history of the Tibetan Plateau and multiple waves of admixture of Tibetans inferred from both Ancient and Modern genome-wide data. Front. Genet. 12, 725243 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Chen F. H., Dong G. H., Zhang D. J., Liu X. Y., Jia X., An C. B., Ma M. M., Xie Y. W., Barton L., Ren X. Y., Zhao Z. J., Wu X. H., Jones M. K., Agriculture facilitated permanent human occupation of the Tibetan Plateau after 3600 B.P. Science 347, 248–250 (2015). [DOI] [PubMed] [Google Scholar]
- 16.Qi X., Cui C., Peng Y., Zhang X., Yang Z., Zhong H., Zhang H., Xiang K., Cao X., Wang Y., Ouzhuluobu, Basang, Ciwangsangbu, Bianba, Gonggalanzi, Wu T., Chen H., Shi H., Su B., Genetic evidence of paleolithic colonization and neolithic expansion of modern humans on the tibetan plateau. Mol. Biol. Evol. 30, 1761–1778 (2013). [DOI] [PubMed] [Google Scholar]
- 17.Sun Y., Wang M., Sun Q., Liu Y., Duan S., Wang Z., Zhou Y., Zhong J., Huang Y., Huang X., Yang Q., Li X., Su H., Cai Y., Jiang X., Chen J., Yan J., Nie S., Hu L., Yang J., Tang R., Wang C. C., Liu C., Deng X., Yun L., He G., Distinguished biological adaptation architecture aggravated population differentiation of Tibeto-Burman-speaking people. J. Genet. Genomics 51, 517–530 (2024). [DOI] [PubMed] [Google Scholar]
- 18.Bergstrom A., McCarthy S. A., Hui R., Almarri M. A., Ayub Q., Danecek P., Chen Y., Felkel S., Hallast P., Kamm J., Blanche H., Deleuze J. F., Cann H., Mallick S., Reich D., Sandhu M. S., Skoglund P., Scally A., Xue Y., Durbin R., Tyler-Smith C., Insights into human genetic variation and population history from 929 diverse genomes. Science 367, eaay5012 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhang M., Yan S., Pan W., Jin L., Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic. Nature 569, 112–115 (2019). [DOI] [PubMed] [Google Scholar]
- 20.Zhang X., Ji X., Li C., Yang T., Huang J., Zhao Y., Wu Y., Ma S., Pang Y., Huang Y., He Y., Su B., A Late Pleistocene human genome from Southwest China. Curr. Biol. 32, 3095–3109.e5 (2022). [DOI] [PubMed] [Google Scholar]
- 21.Wang T., Yang M. A., Zhu Z., Ma M., Shi H., Speidel L., Min R., Yuan H., Jiang Z., Hu C., Li X., Zhao D., Bai F., Cao P., Liu F., Dai Q., Feng X., Yang R., Wu X., Liu X., Zhang M., Ping W., Liu Y., Wan Y., Yang F., Zhou R., Kang L., Dong G., Stoneking M., Fu Q., Prehistoric genomes from Yunnan reveal ancestry related to Tibetans and Austroasiatic speakers. Science 388, eadq9792 (2025). [DOI] [PubMed] [Google Scholar]
- 22.He G., Yao H., Duan S., Luo L., Sun Q., Tang R., Chen J., Wang Z., Sun Y., Li X., Hu L., Yun L., Yang J., Yan J., Nie S., Zhu Y., 10K_CPGDP Consortium, Wang C. C., Liu B., Hu L., Liu C., Wang M., Pilot work of the 10K Chinese People Genomic Diversity Project along the Silk Road suggests a complex east-west admixture landscape and biological adaptations. Sci. China Life Sci. 68, 914–933 (2025). [DOI] [PubMed] [Google Scholar]
- 23.Wang M., Chen H., Luo L., Huang Y., Duan S., Yuan H., Tang R., Liu C., He G., Forensic investigative genetic genealogy: Expanding pedigree tracing and genetic inquiry in the genomic era. J. Genet. Genomics 52, 460–472 (2025). [DOI] [PubMed] [Google Scholar]
- 24.Wang M., Huang Y., Liu K., Wang Z., Zhang M., Yuan H., Duan S., Wei L., Yao H., Sun Q., Zhong J., Tang R., Chen J., Sun Y., Li X., Su H., Yang Q., Hu L., Yun L., Yang J., Nie S., Cai Y., Yan J., Zhou K., Wang C., 10K_CPGDP Consortium, Zhu B., Liu C., He G., Multiple human population movements and cultural dispersal events shaped the landscape of chinese paternal heritage. Mol. Biol. Evol. 41, msae122 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Guyon L., Guez J., Toupance B., Heyer E., Chaix R., Patrilineal segmentary systems provide a peaceful explanation for the post-Neolithic Y-chromosome bottleneck. Nat. Commun. 15, 3243 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Karmin M., Flores R., Saag L., Hudjashov G., Brucato N., Crenna-Darusallam C., Larena M., Endicott P. L., Jakobsson M., Lansing J. S., Sudoyo H., Leavesley M., Metspalu M., Ricaut F. X., Cox M. P., Episodes of diversification and isolation in Island Southeast Asian and near oceanian male lineages. Mol. Biol. Evol. 39, msac045 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Pinotti T., Bergstrom A., Geppert M., Bawn M., Ohasi D., Shi W., Lacerda D. R., Solli A., Norstedt J., Reed K., Dawtry K., Gonzalez-Andrade F., Paz Y. M. C., Revollo S., Cuellar C., Jota M. S., Santos J. E. Jr., Ayub Q., Kivisild T., Sandoval J. R., Fujita R., Xue Y., Roewer L., Santos F. R., Tyler-Smith C., Y Chromosome sequences reveal a short beringian standstill, rapid expansion, and early population structure of native American founders. Curr. Biol. 29, 149–157.e3 (2019). [DOI] [PubMed] [Google Scholar]
- 28.Zeng T. C., Aw A. J., Feldman M. W., Cultural hitchhiking and competition between patrilineal kin groups explain the post-Neolithic Y-chromosome bottleneck. Nat. Commun. 9, 2077 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Batini C., Hallast P., Zadik D., Delser P. M., Benazzo A., Ghirotto S., Arroyo-Pardo E., Cavalleri G. L., de Knijff P., Dupuy B. M., Eriksen H. A., King T. E., de Munain A. L., Lopez-Parra A. M., Loutradis A., Milasin J., Novelletto A., Pamjav H., Sajantila A., Tolun A., Winney B., Jobling M. A., Large-scale recent expansion of European patrilineages shown by population resequencing. Nat. Commun. 6, 7152 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.D’Atanasio E., Trombetta B., Bonito M., Finocchio A., Di Vito G., Seghizzi M., Romano R., Russo G., Paganotti G. M., Watson E., Coppa A., Anagnostou P., Dugoujon J. M., Moral P., Sellitto D., Novelletto A., Cruciani F., The peopling of the last Green Sahara revealed by high-coverage resequencing of trans-Saharan patrilineages. Genome Biol. 19, 20 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Poznik G. D., Xue Y., Mendez F. L., Willems T. F., Massaia A., Sayres M. A. W., Ayub Q., McCarthy S. A., Narechania A., Kashin S., Chen Y., Banerjee R., Rodriguez-Flores J. L., Cerezo M., Shao H., Gymrek M., Malhotra A., Louzada S., Desalle R., Ritchie G. R., Cerveira E., Fitzgerald T. W., Garrison E., Marcketta A., Mittelman D., Romanovitch M., Zhang C., Zheng-Bradley X., Abecasis G. R., McCarroll S. A., Flicek P., Underhill P. A., Coin L., Zerbino D. R., Yang F., Lee C., Clarke L., Auton A., Erlich Y., Handsaker R. E., 1000 Genomes Project Consortium, Bustamante C. D., Tyler-Smith C., Punctuated bursts in human male demography inferred from 1,244 worldwide Y-chromosome sequences. Nat. Genet. 48, 593–599 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Wang Z., Liu K., Yuan H., Duan S., Liu Y., Luo L., Jiang X., Chen S., Wei L., Tang R., Hu L., Chen J., Li X., Yang Q., Sun Y., Sun Q., Huang Y., Su H., Zhong J., Yao H., Yun L., Li J., Yang J., Cai Y., Deng H., Yan J., Zhu B., 10K_CPGDP Consortium, Zhou K., Nie S., Liu C., Wang M., He G., YanHuang paternal genomic resource suggested a weakly-differentiated multi-source admixture model for the formation of Han’s founding ancestral lineages. Genomics Proteomics Bioinformatics, qzaf049 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Li Y. C., Ye W. J., Jiang C. G., Zeng Z., Tian J. Y., Yang L. Q., Liu K. J., Kong Q. P., River valleys shaped the maternal genetic landscape of Han Chinese. Mol. Biol. Evol. 36, 1643–1652 (2019). [DOI] [PubMed] [Google Scholar]
- 34.All of Us Research Program Genomics Investigators , Genomic data in the all of us research program. Nature 627, 340–346 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Byrska-Bishop M., Evani U. S., Zhao X., Basile A. O., Abel H. J., Regier A. A., Corvelo A., Clarke W. E., Musunuri R., Nagulapalli K., Fairley S., Runnels A., Winterkorn L., Lowy E., Human Genome Structural Variation Consortium, Flicek P., Germer S., Brand H., Hall I. M., Talkowski M. E., Narzisi G., Zody M. C., High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440.e19 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Mallick S., Li H., Lipson M., Mathieson I., Gymrek M., Racimo F., Zhao M., Chennagiri N., Nordenfelt S., Tandon A., Skoglund P., Lazaridis I., Sankararaman S., Fu Q., Rohland N., Renaud G., Erlich Y., Willems T., Gallo C., Spence J. P., Song Y. S., Poletti G., Balloux F., van Driem G., de Knijff P., Romero I. G., Jha A. R., Behar D. M., Bravi C. M., Capelli C., Hervig T., Moreno-Estrada A., Posukh O. L., Balanovska E., Balanovsky O., Karachanak-Yankova S., Sahakyan H., Toncheva D., Yepiskoposyan L., Tyler-Smith C., Xue Y., Abdullah M. S., Ruiz-Linares A., Beall C. M., Di Rienzo A., Jeong C., Starikovskaya E. B., Metspalu E., Parik J., Villems R., Henn B. M., Hodoglugil U., Mahley R., Sajantila A., Stamatoyannopoulos G., Wee J. T., Khusainova R., Khusnutdinova E., Litvinov S., Ayodo G., Comas D., Hammer M. F., Kivisild T., Klitz W., Winkler C. A., Labuda D., Bamshad M., Jorde L. B., Tishkoff S. A., Watkins W. S., Metspalu M., Dryomov S., Sukernik R., Singh L., Thangaraj K., Paabo S., Kelso J., Patterson N., Reich D., The simons genome diversity project: 300 genomes from 142 diverse populations. Nature 538, 201–206 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Ilumae A. M., Reidla M., Chukhryaeva M., Jarve M., Post H., Karmin M., Saag L., Agdzhoyan A., Kushniarevich A., Litvinov S., Ekomasova N., Tambets K., Metspalu E., Khusainova R., Yunusbayev B., Khusnutdinova E. K., Osipova L. P., Fedorova S., Utevska O., Koshel S., Balanovska E., Behar D. M., Balanovsky O., Kivisild T., Underhill P. A., Villems R., Rootsi S., Human Y chromosome haplogroup N: A non-trivial time-resolved phylogeography that cuts across language families. Am. J. Hum. Genet. 99, 163–173 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Wong L. P., Ong R. T., Poh W. T., Liu X., Chen P., Li R., Lam K. K., Pillai N. E., Sim K. S., Xu H., Sim N. L., Teo S. M., Foo J. N., Tan L. W., Lim Y., Koo S. H., Gan L. S., Cheng C. Y., Wee S., Yap E. P., Ng P. C., Lim W. Y., Soong R., Wenk M. R., Aung T., Wong T. Y., Khor C. C., Little P., Chia K. S., Teo Y. Y., Deep whole-genome sequencing of 100 southeast Asian Malays. Am. J. Hum. Genet. 92, 52–66 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.He G., Wang M., Luo L., Sun Q., Yuan H., Lv H., Feng Y., Liu X., Cheng J., Bu F., Population genomics of Central Asian peoples unveil ancient Trans-Eurasian genetic admixture and cultural exchanges. hLife 2, 554–562 (2024). [Google Scholar]
- 40.Hallast P., Agdzhoyan A., Balanovsky O., Xue Y., Tyler-Smith C., A Southeast Asian origin for present-day non-African human Y chromosomes. Hum. Genet. 140, 299–307 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Bennett E. A., Fu Q., Ancient genomes and the evolutionary path of modern humans. Cell 187, 1042–1046 (2024). [DOI] [PubMed] [Google Scholar]
- 42.Zeberg H., Jakobsson M., Paabo S., The genetic changes that shaped Neandertals, Denisovans, and modern humans. Cell 187, 1047–1058 (2024). [DOI] [PubMed] [Google Scholar]
- 43.Mathieson I., Alpaslan-Roodenberg S., Posth C., Szecsenyi-Nagy A., Rohland N., Mallick S., Olalde I., Broomandkhoshbacht N., Candilio F., Cheronet O., Fernandes D., Ferry M., Gamarra B., Fortes G. G., Haak W., Harney E., Jones E., Keating D., Krause-Kyora B., Kucukkalipci I., Michel M., Mittnik A., Nagele K., Novak M., Oppenheimer J., Patterson N., Pfrengle S., Sirak K., Stewardson K., Vai S., Alexandrov S., Alt K. W., Andreescu R., Antonovic D., Ash A., Atanassova N., Bacvarov K., Gusztav M. B., Bocherens H., Bolus M., Boroneant A., Boyadzhiev Y., Budnik A., Burmaz J., Chohadzhiev S., Conard N. J., Cottiaux R., Cuka M., Cupillard C., Drucker D. G., Elenski N., Francken M., Galabova B., Ganetsovski G., Gely B., Hajdu T., Handzhyiska V., Harvati K., Higham T., Iliev S., Jankovic I., Karavanic I., Kennett D. J., Komso D., Kozak A., Labuda D., Lari M., Lazar C., Leppek M., Leshtakov K., Vetro D. L., Los D., Lozanov I., Malina M., Martini F., McSweeney K., Meller H., Mendusic M., Mirea P., Moiseyev V., Petrova V., Price T. D., Simalcsik A., Sineo L., Slaus M., Slavchev V., Stanev P., Starovic A., Szeniczey T., Talamo S., Teschler-Nicola M., Thevenet C., Valchev I., Valentin F., Vasilyev S., Veljanovska F., Venelinova S., Veselovskaya E., Viola B., Virag C., Zaninovic J., Zauner S., Stockhammer P. W., Catalano G., Krauss R., Caramelli D., Zarina G., Gaydarska B., Lillie M., Nikitin A. G., Potekhina I., Papathanasiou A., Boric D., Bonsall C., Krause J., Pinhasi R., Reich D., The genomic history of southeastern Europe. Nature 555, 197–203 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Cavalli-Sforza L. L., The Chinese human genome diversity project. Proc. Natl. Acad. Sci. U.S.A. 95, 11501–11503 (1998). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Liu D., Duong N. T., Ton N. D., Van Phong N., Pakendorf B., Van Hai N., Stoneking M., Extensive ethnolinguistic diversity in Vietnam reflects multiple sources of genetic diversity. Mol. Biol. Evol. 37, 2503–2519 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Kutanan W., Kampuansai J., Srikummool M., Brunelli A., Ghirotto S., Arias L., Macholdt E., Hubner A., Schroder R., Stoneking M., Contrasting paternal and maternal genetic histories of Thai and Lao populations. Mol. Biol. Evol. 36, 1490–1506 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Schepartz L. A., Miller-Antonio S., Bakken D. A., Upland resources and the early Palaeolithic occupation of Southern China, Vietnam, Laos, Thailand and Burma. World Archaeol. 32, 1–13 (2000). [Google Scholar]
- 48.Zhang Z., Zhang Y., Wang Y., Zhao Z., Yang M., Zhang L., Zhou B., Xu B., Zhang H., Chen T., Dai W., Zhou Y., Shi S., Nielsen R., Li S. C., Li S., The Tibetan-Yi region is both a corridor and a barrier for human gene flow. Cell Rep. 39, 110720 (2022). [DOI] [PubMed] [Google Scholar]
- 49.He G., Sun Y., Duan S., Luo L., Sun Q., Li B., Yun L., Liu C., Wang M., Ancient genomes give insight into 160,000 years of East Asian population dynamics and biological adaptation. Genome Biol. 26, 420 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Lu H., Zhang J., Liu K. B., Wu N., Li Y., Zhou K., Ye M., Zhang T., Zhang H., Yang X., Shen L., Xu D., Li Q., Earliest domestication of common millet (Panicum miliaceum) in East Asia extended to 10,000 years ago. Proc. Natl. Acad. Sci. U.S.A. 106, 7367–7372 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Ma T., Rolett B. V., Zheng Z., Zong Y., Holocene coastal evolution preceded the expansion of paddy field rice farming. Proc. Natl. Acad. Sci. U.S.A. 117, 24138–24143 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Zuo X., Lu H., Jiang L., Zhang J., Yang X., Huan X., He K., Wang C., Wu N., Dating rice remains through phytolith carbon-14 study reveals domestication at the beginning of the Holocene. Proc. Natl. Acad. Sci. U.S.A. 114, 6486–6491 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.World Medical Association , World Medical Association Declaration of Helsinki: Ethical principles for medical research involving human subjects. JAMA 310, 2191–2194 (2013). [DOI] [PubMed] [Google Scholar]
- 54.H. Li, Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv:1303.3997 [q-bio.GN] (2013).
- 55.McKenna A., Hanna M., Banks E., Sivachenko A., Cibulskis K., Kernytsky A., Garimella K., Altshuler D., Gabriel S., Daly M., DePristo M. A., The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 20, 1297–1303 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Stamatakis A., RAxML version 8: A tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30, 1312–1313 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Jagadeesan A., Ebenesersdottir S. S., Guethmundsdottir V. B., Thordardottir E. L., Moore K. H. S., Helgason A., HaploGrouper: A generalized approach to haplogroup classification. Bioinformatics 37, 570–572 (2021). [DOI] [PubMed] [Google Scholar]
- 58.Letunic I., Bork P., Interactive Tree of Life (iTOL) v6: Recent updates to the phylogenetic tree display and annotation tool. Nucleic Acids Res. 52, W78–W82 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Zhou T., Xu K., Zhao F., Liu W., Li L., Hua Z., X., itol.toolkit accelerates working with iTOL (Interactive Tree of Life) by an automated generation of annotation files. Bioinformatics 39, btad339 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Suchard M. A., Lemey P., Baele G., Ayres D. L., Drummond A. J., Rambaut A., Bayesian phylogenetic and phylodynamic data integration using BEAST 1.10. Virus Evol. 4, vey016 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Darriba D., Taboada G. L., Doallo R., Posada D., jModelTest 2: More models, new heuristics and parallel computing. Nat. Methods 9, 772 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Mallick S., Micco A., Mah M., Ringbauer H., Lazaridis I., Olalde I., Patterson N., Reich D., The Allen Ancient DNA Resource (AADR) a curated compendium of ancient human genomes. Sci. Data 11, 182 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Bai F., Liu Y., Wangdue S., Wang T., He W., Xi L., Tsho Y., Tsering T., Cao P., Dai Q., Liu F., Feng X., Zhang M., Ran J., Ping W., Payon D., Mao X., Tong Y., Tsring T., Chen Z., Fu Q., Ancient genomes revealed the complex human interactions of the ancient western Tibetans. Curr. Biol. 34, 2594–2605.e7 (2024). [DOI] [PubMed] [Google Scholar]
- 64.Guo J. X., He H. F., Xie G. M., Tao L., Mai X., Zhu K. Y., Qin Q. S., Yang X. M., Xie Y., Wang R., Ma H., Zhao J., Li D. W., Gong S. Y., Wang C. C., Genetic affinity of cave burial and Hmong-Mien populations in Guangxi inferred from ancient genomes. Archaeol. Anthrop. Sci. 16, 121 (2024). [Google Scholar]
- 65.Li S., Wang R., Ma H., Tu Z., Qiu L., Chen H., Jiang L., Geng Y., Liu H., Wang J., Shen Q., Jin L., Li C., Wang C. C., Wei X., Ancient genomic time transect unravels the population dynamics of Neolithic middle Yellow River farmers. Sci. Bull. 69, 3365–3370 (2024). [DOI] [PubMed] [Google Scholar]
- 66.Zhu K., Hu C., Yang M., Zhang X., Guo J., Xie M., Yang X., Ma H., Wang R., Zhao J., Tao L., He H., Wan W., Zhang Q., Jin L., Zuo Y., Zhou B., Huang J., Wang C. C., The demic diffusion of Han culture into the Yunnan-Guizhou plateau inferred from ancient genomes. Natl. Sci. Rev. 11, nwae387 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Xiong J., Wang R., Chen G., Yang Y., Du P., Meng H., Ma M., Allen E., Tao L., Wang H., Jin L., Wang C. C., Wen S., Inferring the demographic history of Hexi Corridor over the past two millennia from ancient genomes. Sci. Bull. 69, 606–611 (2024). [DOI] [PubMed] [Google Scholar]
- 68.Yang X., Gao Y., Wangdue S., Ran J., Wang Q., Chen S., Yang J., Wang T., Gu Z., Zhang Y., Cao P., Dai Q., Chen S., Tong Y., Jia N., Sun Q., Huang Y., Perry L., d’Alpoim Guedes J., Han X., Liu F., Feng X., Yang Q., Wang Y., Hu S., Tian Y., Guo J., Liang X., You T., Li Y., Zhang Y., Deng Z., Qin L., Wu X., Zhuang Y., Liu Y., Fu Q., Chen F., Lake-centred sedentary lifestyle of early Tibetan Plateau Indigenous populations at high elevation 4,400 years ago. Nat. Ecol. Evol. 8, 2297–2308 (2024). [DOI] [PubMed] [Google Scholar]
- 69.Wang F., Wang R., Ma H., Zeng W., Zhao Y., Wu H., Tang Z., He H., Fang H., Wang C. C., Neolithization of Dawenkou culture in the lower Yellow River involved the demic diffusion from the Central Plain. Sci. Bull. 69, 3677–3681 (2024). [DOI] [PubMed] [Google Scholar]
- 70.Du P., Zhu K., Wang M., Sun Z., Tan J., Sun B., Sun B., Wang P., He G., Xiong J., Huang Z., Meng H., Sun C., Xie S., Wang B., Ge D., Ma Y., Sheng P., Ren X., Tao Y., Xu Y., Qin X., Allen E., Zhang B., Chang X., Wang K., Bao H., Yu Y., Wang L., Ma X., Du Z., Guo J., Yang X., Wang R., Ma H., Li D., Pan Y., Li B., Zhang Y., Zheng X., Han S., Jin L., Chen G., Li H., Wang C. C., Wen S., Genomic dynamics of the Lower Yellow River Valley since the Early Neolithic. Curr. Biol. 34, 3996–4006.e11 (2024). [DOI] [PubMed] [Google Scholar]
- 71.Yu Y., Yang X., Liu D., Du P., Meng H., Huang Z., Xiong J., Ding Y., Ren X., Allen E., Wang H., Han S., Jin L., Wang C. C., Wen S., Ancient genomic analysis of a Chinese hereditary elite from the Northern and Southern Dynasties. J. Genet. Genomics 52, 473–482 (2025). [DOI] [PubMed] [Google Scholar]
- 72.Ma H., Zhou Y., Wang R., Yan F., Chen H., Qiu L., Zhao J., Jin L., Wang C. C., Ancient genomes shed light on the long-term genetic stability in the Central Plain of China. Sci. Bull. 70, 333–337 (2025). [DOI] [PubMed] [Google Scholar]
- 73.Zhang F., Zhang X., Bai B., Hu C., Duan C., Yuan H., Zhang R., Ma P., Zhou B., Ning C., Ancient genomes provide insights into the genetic history in the historical era of southwest China. Archaeol. Anthrop. Sci. 16, 120 (2024). [Google Scholar]
- 74.Shen Q., Wu Z., Zan J., Yang X., Guo J., Ji Z., Wang B., Liu Y., Mao X., Wang X., Zou X., Zhou H., Peng Y., Ma H., He H., Bai T., Xu M., Wen S., Jin L., Zhang Q., Wang C. C., Ancient genomes illuminate the demographic history of Shandong over the past two millennia. J. Genet. Genomics 52, 494–501 (2025). [DOI] [PubMed] [Google Scholar]
- 75.Fang H., Liang F., Ma H., Wang R., He H., Qiu L., Tao L., Zhu K., Wu W., Ma L., Zhang H., Chen S., Zhu C., Chen H., Xu Y., Zhao Y., Liu H., Wang C. C., Dynamic history of the Central Plain and Haidai region inferred from Late Neolithic to Iron Age ancient human genomes. Cell Rep. 44, 115262 (2025). [DOI] [PubMed] [Google Scholar]
- 76.Liu J., Liu Y., Zhao Y., Zhu C., Wang T., Zeng W., Sun B., Wang F., Han H., Li Z., Feng X., Cao P., Luan F., Liu F., Dai Q., Guo J., Wang Z., Wei C., Wei Q., Yang R., Hou W., Ping W., Bai F., Miao B., Wang W., Yang M. A., Fu Q., East Asian Gene flow bridged by northern coastal populations over past 6000 years. Nat. Commun. 16, 1322 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Martiniano R., De Sanctis B., Hallast P., Durbin R., Placing ancient DNA sequences into reference phylogenies. Mol. Biol. Evol. 39, msac017 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Kutanan W., Shoocongdej R., Srikummool M., Hubner A., Suttipai T., Srithawong S., Kampuansai J., Stoneking M., Cultural variation impacts paternal and maternal genetic lineages of the Hmong-Mien and Sino-Tibetan groups from Thailand. Eur. J. Hum. Genet. 28, 1563–1579 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Macholdt E., Arias L., Duong N. T., Ton N. D., Van Phong N., Schroder R., Pakendorf B., Van Hai N., Stoneking M., The paternal and maternal genetic history of Vietnamese populations. Eur. J. Hum. Genet. 28, 636–645 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.C. Amante, B. W. Eakins, ETOPO1 arc-minute global relief model: Procedures, data sources and analysis. (NOAA, 2009).
- 81.Hosner D., Wagner M., Tarasov P. E., Chen X. C., Leipe C., Spatiotemporal distribution patterns of archaeological sites in China during the Neolithic and Bronze Age: An overview. Holocene 26, 1576–1593 (2016). [Google Scholar]
- 82.Wu J., Shen C. M., Yang H., Qian S., Xie S. C., Holocene temperature variability in China. Quat. Sci. Rev. 312, 108184 (2023). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Figs. S1 to S13
Legends for tables S1 to S11
Legend for data S1
Tables S1 to S11
Data S1
Data Availability Statement
All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials. There are no new materials generated in the study. Genome-wide variation data for ancient individuals were obtained from the Allen Ancient DNA Resource (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW&ve1984). The results of the analyses have been included in the Supplementary Materials and deposited in the OMIX database (https://ngdc.cncb.ac.cn/omix/) under accession number OMIX007713. The variation data reported in this paper have been deposited in the Genome Variation Map at the National Genomics Data Center, Beijing Institute of Genomics, Chinese Academy of Sciences and China National Center for Bioinformation under accession number GVM000878 [Genome Variation Map (cncb.ac.cn)]. Base maps are available in the GitHub repository (https://github.com/GuanglinHe1/SA-Neolithic-farmer-migration/) and Zenodo (https://zenodo.org/records/18573902). The analysis code is available via Zenodo (https://zenodo.org/records/18753608). Access to and use of data must comply with the regulations of the People’s Republic of China on the administration of human genetic resources. To comply with the informed-consent agreements, the data are available only upon written request for population history research; they may not be used for medical or disease-related research, natural-selection analyses, commercial purposes, or individual identification. Requesters may not share the data with third parties or deposit it in other public databases without prior written approval. The raw data can be provided by K. Wu pending scientific review and a completed material transfer agreement through Sichuan University. Requests for the raw data should be submitted to K. Wu (wuke2026@wchscu.cn).






