Skip to main content
Fundamental Research logoLink to Fundamental Research
. 2026 Apr 7;6(3):1332–1342. doi: 10.1016/j.fmre.2026.04.001

Genetic and linguistic landscapes of Gansu-Qinghai populations

Haodong Chen a,1, Huangzhen Huang a,1, Boxuan Zhao b,1, Yiling Jiang c,1, Mengting Xu c, Limin Qiu c, Jin Sun d, Yourong Ye e, Hongbing Yao f, Xinyi Jia f, Kai Sun g, Jiangtao Shen h, Hao Li i, Dan Xu i,j,, Chuan-Chao Wang a,, Shaoqing Wen b,
PMCID: PMC13247448  PMID: 42272488

Abstract

The Gansu-Qinghai (GQ) region is a critical Eurasian contact zone with exceptional linguistic and genetic diversity. However, genomic studies integrating genetic and linguistic evidence in areas with intensive language mixing remain limited. We conducted fine-scale genomic analyses on 153 new samples from four linguistically mixing or mixed regions: Linxia and Xiahe in Gansu, and Gan’gou and Wutun in Qinghai. Our findings reveal three primary ancestral components: Yellow River, western Eurasian steppe, and Tibetan Plateau (TP)-related lineages. While non-Han populations show pronounced genetic heterogeneity and varying patterns of language maintenance, Han populations exhibit asymmetric convergence. Specifically, the Wutun Han show concordant genetic and linguistic shifts, whereas other Han groups exhibit linguistic structural convergence without corresponding gene flow. Widespread east-west admixture was detected in both Han Chinese and non-Han populations from the GQ region, primarily dating to the Tang/Song and Yuan Dynasties. Together, our interdisciplinary genetic-linguistic framework reveals complex and asymmetric processes underlying population interaction and language evolution in the GQ contact zone.

Keywords: Population genetics, Gansu-Qinghai region, Gene flow, Language admixture, Co-evolution

Graphical abstract

Image, graphical abstract

1. Introduction

The co-evolution of genetic lineages and human languages has become a foundational framework for reconstructing population histories. This interdisciplinary approach traces its roots to the pioneering work of Luigi Luca Cavalli-Sforza (1922–2018), who introduced the concept of “cultural genes”, drawing parallels between the transmission of languages and technologies and genetic inheritance [[1], [2], [3]]. For example, the understanding of Indo-European language spread has been refined; recent evidence suggests a “Proto-Indo-Anatolian” language, spoken by peoples of the Caucasus-Lower Volga (CLV) cline, was ancestral to both the Anatolian and other Indo-European branches [[4], [5], [6], [7], [8]]. Similarly, phylogenetic studies have shown that Sino-Tibetan languages emerged around 5900 years ago in northern China’s Yellow River (YR) basin, in close association with the Yangshao culture and the spread of agriculture [9]. Building on this foundation, large-scale lexical and genetic analyses of Han Chinese reveal a hybrid model of demic and cultural diffusion: northern Mandarin dialects correlate with north-to-south genetic clines shaped by historical migrations, while southern dialects retain distinct linguistic features influenced by localized cultural transmission and assimilation [10]. These findings underscore the value of combining genetic and linguistic analyses to explore the complex dynamics of population movements and cultural interactions in the evolution of both genes and languages.

The Gansu-Qinghai (GQ) region in Northwest China, located along the eastern section of the Silk Road, provides an excellent setting to study the co-evolution of genetics and languages throughout human history [11]. This area is notable for its remarkable ethnolinguistic diversity [12], where speakers from various language families have coexisted and interacted for centuries. In addition to Sinitic varieties such as Zhongyuan Mandarin (including Qinlong, Longzhong, and Hezhou subgroups), the region is home to multiple non-Han language families, including Mongolic (e.g., Dongxiang, Bonan, Tu, and Eastern Yugur), Turkic (e.g., Salar and Western Yugur), and Tibetic languages (e.g., Amdo varieties) [[13], [14], [15], [16]].

Previous genetic studies indicate that populations in this region harbor both East Asian-related and West/Central Eurasian-related ancestries, reflecting long-term population movements and interactions along the Silk Road [[17], [18], [19], [20]]. Although broad correspondences between genetic structure and linguistic affiliation have been observed, uniparental marker studies have revealed notable exceptions in several non-Han groups, such as the Dongxiang, Bonan, and Yugur populations [[21], [22], [23]]. However, genome-wide autosomal studies that systematically compare genetic patterns across different ethnic and linguistic groups in the GQ region remain limited, constraining a comprehensive understanding of genetic-linguistic relationships in this complex contact zone.

Additionally, the GQ region showcases particularly fascinating cases of language mixing [24]. Mixing or mixed languages extensively borrow morphosyntactic traits from other languages, rather than vocabulary. Previous studies that focused solely on the enumeration and differentiation of lexical traits have not demonstrated significant efficacy. Languages such as Wutun [25] (a recognized mixed language reshaped by Amdo Tibetan and Mongolic influences) and Tangwang [26] (a Sinitic dialect with Mongolic case marking and Dongxiang substrate effects) exemplify typologically rare outcomes of intensive language contact. To date, only one Y-chromosome uniparental study has been conducted on the Tangwang people [27], revealing diverse paternal lineages (e.g., C3*-Mongolic, O3a1c-Han) that align with historical records of clan migrations [28]. However, significant gaps remain in genetic studies of mixed-language populations (such as Wutun) and other emerging linguistic varieties (Gan’gou, Linxia, and Xiahe), limiting our understanding of genetic-linguistic correlations in this contact zone.

Our study addresses these gaps by conducting the first integrated analysis of genome-wide data alongside linguistic findings. To achieve this, we generated new genome-wide data from 153 individuals and combined it with comprehensive previously published datasets encompassing 67 modern and 103 ancient populations from Eurasia [[29], [30], [31]]. We collected 153 saliva samples from four major ethnic groups (Hui, Tu, Tibetan, and Han) in regions characterized by language mixing, specifically Linxia/Xiahe in Gansu province and Wutun/Gan’gou in Qinghai province (Fig. 1). By combining previously published autosomal single nucleotide polymorphism (SNP) data of multiple ethnic groups in this region with existing linguistic typologies, we aim to: (1) evaluate the correspondence between genetic structure and language diversity across the GQ region, and (2) identify the underlying factors shaping the complex coevolutionary dynamics in mixing- and mixed-language Han Chinese.

Fig. 1.

Fig 1 dummy alt text

Geographic location and present-day samples of the GQ region. The format of the points corresponds to that in the PCA plot (Fig. 2A). The geographic base map was based on the Standard Map Service System of the Ministry of Natural Resources of China (http://bzdt.ch.mnr.gov.cn) (GS(2020)4619).

2. Materials and methods

2.1. Sample collection, genotyping, quality control and data merging

The studies involving human participants were reviewed and approved by the Medical Ethics Committee of Xiamen University (XDYX202412K88). The participants provided their written informed consent to participate in this study.

We collected 153 saliva samples from four autonomous prefectures and counties in Gansu and Qinghai Provinces, encompassing Han, Tibetan, Hui, and Tu groups, specifically from Gannan Tibetan Autonomous Prefecture, Linxia Hui Autonomous Prefecture, Huangnan Tibetan Autonomous Prefecture, and Minhe Hui and Tu Autonomous County. Genome-wide genotyping was performed on the Illumina Infinium Chinese Genotyping Array BeadChip: Illumina WeGene V3 Arrays covering about 700,000 SNPs at the WeGene Clinical Laboratory, Shenzhen. We used PLINK1.9 to clean the original binary files (bim, bed, fam) under the following cleaning conditions: –geno 0.05, –mind 0.05, –hwe 1e‐6 [32]. We used KING to filter the quality control data further, thereby avoiding the influence of potential kinship within three generations on the analysis results [33]. Ultimately, 504,909 SNPs and 135 individuals remained for subsequent population genetic analysis.

We merged our data with published modern and ancient East and Southeast Asian populations from the Human Origins (HO) data set and the 1240k data set included in the Allen Ancient DNA Resource (AADR) [29] using mergeit from EIGENSOFT. We also merged our data with recently published modern and ancient samples [30,31]. The “merged HO” dataset contains 67,567 overlapped SNPs, and the “merged 1240k” dataset contains 179,742 overlapped SNPs.

2.2. Y chromosomal and mtDNA haplogroup assignment

The WeGene V3 Arrays were designed to identify all known Y-chromosome and mtDNA lineages with 24,048 Y-chromosome and 3746 mtDNA phylogenetically relevant SNPs. We first corrected Y-chromosome and mitochondrial SNP data against the hg19 reference genome using bcftools [34]. Then, we utilized Haplogrep2 [34] and Y-LineageTracker [35] for haplogroup assignment.

2.3. Allele frequency-based analysis

2.3.1. Principal component analysis

We performed principal component analysis (PCA) based on the “merged HO” dataset using smartpca in the EIGENSOFT package [36]. The PCA analysis was performed at the individual level to characterize the genetic structure of all of our samples in Gansu province and the reference populations. The option “lsqproject: YES” was used to project ancient individuals to the calculated principal components based on modern populations. The ggplot2 package in the R software (http://www.r-project.org/) was used for visualization. Principal component analyses were performed at two spatial scales: a broad East-West Eurasian contrast and a fine-scale examination within East Asia. In the PCA within East Asian populations, the studied populations were visualized separately according to their ethnolinguistic affiliations to facilitate clearer inspection of their genetic clustering patterns.

2.3.2. ADMIXTURE

We retained 59,362 SNPs on the “merged HO” dataset after pruning strong linkage disequilibrium with each other using PLINK1.9 [37] by parameters “–indep-pairwise 200 25 0.4” and then ran unsupervised ADMIXTURE [38] with parameters from K = 2 to 10 in 100 bootstraps with different random seeds. We chose the best run according to the lowest cross-validation (CV) error value at K = 5.

2.3.3. Pairwise Fst genetic distance

We calculated Wright’s fixation index pairwise Fst genetic distance among 46 modern populations using smartpca as implemented in EIGENSOFT [36] with default parameters and “inbreed: YES” and “fsonly: YES”. In the following, the neighbor-joining (N-J) phylogenetic relationship was performed via MEGA11 [39].

2.3.4. TreeMix

We utilized a two-step procedure to minimize the influence of migration specification on the underlying tree topology and ensure that inferred migration events are robust and not driven by stochastic variation across runs (reference method: https://github.com/carolindahms/TreeMix/tree/main):

  • Step 1: Construction of the backbone tree without migration. We first inferred the population tree without specifying any migration events (m = 0) using TreeMix [40]. We used blocks of 500 SNPs for estimating the covariance matrix with the option “-k 500'', and performed 100 bootstrap replicates using the “-bootstrap” option to assess the robustness of the inferred topology. The Mbuti population was used as an outgroup to generate a rooted maximum-likelihood (ML) tree. The resulting bootstrapped trees were summarized using the PHYLIP v3.697 consense program [41] to construct a consensus tree, which served as a stable backbone topology for subsequent analyses.

  • Step 2: Inference of migration edges based on the consensus topology. Using the consensus tree obtained in Step 1 as a fixed tree topology (read tree option, -tf), we conducted TreeMix analyses allowing between 0 and 8 migration events. For each migration scenario, 100 bootstrap replicates were performed to evaluate the statistical support of inferred migration edges, and PHYLIP v3.697 consense was again used to generate consensus trees from the bootstrapped results. To further ensure the robustness and reproducibility of the inferred models, we performed 20 independent TreeMix runs for each migration number. For presentation, we selected the unique ML tree with the highest likelihood under each migration scenario.

2.3.5. f statistics

We computed f3 and f4 statistics via the qp3Pop and qpDstat programs in the ADMIXTOOLS package [42]. We conducted an outgroup-f3 analysis in the form of f3 (Target, X; Mbuti) to measure the genetic affinity between our studied populations and other modern and ancient populations in East Eurasia. We further computed f4 statistics to explore shared alleles between the studied populations and other modern and ancient populations in Eurasia, with the “f4mode: YES” parameter.

2.3.6. qpWave and qpAdm

We performed pairwise qpWave as implemented in ADMIXTOOLS [42] on the “merged 1240k” dataset with the default parameters and “allsnps: YES” and “inbreed: YES” to test whether our studied populations were genetically homogeneous or not, in relation to a set of outgroups. A rank = 0 with P‐value > 0.05 indicated that the pairwise populations were genetically homogeneous. We then used the qpAdm program with default parameters, “allsnps: YES” and “inbreed: YES” to estimate admixture proportions for studied populations as a combination of specified source populations by exploiting the shared genetic drift with a set of outgroups. Ancient genomes were included as temporal references to contextualize the ancestry components of modern populations from the GQ region.

2.3.7. Alder

We applied LD-based ALDER [43] software to infer admixture time using the default parameters and “checkmap: NO”. We used 29 years per generation to scale time [44].

2.3.8. MultiWaver

We used MultiWaver v2.0 [45] to infer population admixture history in northwestern China, which implements a more flexible framework to automatically select a best-fit admixture model among discrete models and continuous models.

2.4. Haplotype sharing-based analysis

2.4.1. IBD segment inference

We employed SHAPEIT v2.r90 [46] for haplotype phasing with additional parameters “–burn 10 –prune 10 –main 30” to enhance the accuracy. IBD fragments between each pair of individuals were inferred from the refined IBD [47] and merged by merge-ibd-segments.17Jan20.102.jar [48]. For the pairwise population level, we normalized the total shared IBD blocks by dividing the product of the sample size.

2.4.2. fineSTRUCTURE

The coancestry matrix results from ChromoPainter were subsequently used to identify the fine-scale population structure in fineSTRUCTURE through Bayesian clustering [49]. To ensure a well-balanced donor population set, we randomly selected 10 individuals from groups with more than 10 members. After phasing these modern samples via SHAPEIT, we employed fineSTRUCTURE using the following parameters: “-s3iters 100,000 -s4iters 50,000 -s1minsnps 1000 -s1indfrac 0.1'”.

3. Results

3.1. Genetic profiles of populations around the GQ region

We first assigned uniparental haplogroups to the newly sampled individuals (Table S1). Analysis of paternal Y chromosome lineages revealed that haplogroup O was predominant, followed by N, D1, C2, and Q, among others. Haplogroups O and N represent the most widespread paternal lineages across East Asia. Notably, haplogroup D1—a lineage characteristically prevalent in Tibeto-Burman-speaking populations—was primarily identified in the Han and Tu individuals from Wutun village. Haplogroup C2, typically associated with Northern Eurasian and Turkic-speaking groups, was also detected in our GQ individuals. Furthermore, we identified several lineages centered in Central and Western Eurasia, including Q, J, and R, reflecting past genetic interactions across the Eurasian steppe. Regarding maternal mtDNA lineages, the majority of individuals belonged to East Asian-dominant haplogroups, such as A, C, D, F, and M. Specifically, multiple individuals carried haplogroup M9 and its subclades, which are frequently observed in populations of the Tibetan Plateau (TP). In addition, we identified maternal lineages with potential Western or Central Eurasian affinities, including X2 and H.

We initially conducted Principal Component Analysis (PCA) within the context of ancient/modern Eurasian populations to examine the genetic profiles (Fig. 2A). Our observations revealed that populations sharing linguistic or geographical origins tend to cluster together, forming two distinct genetic clines: one along PC1, comprising Turkic-speaking populations and ancient populations from the Western Region (present-day Xinjiang) and the Eurasian steppe; the other cline includes Sinitic, Tibetan groups, and Southern Chinese (such as Tai-Kadai, Hmong-Mien, Austronesian, etc.). East Asian populations are compressed into a genetic cline.

Fig. 2.

Fig 2 dummy alt text

(A) The principal component analysis (PCA) with populations in Eurasia. Ancient samples were projected onto dimensions computed by present-day populations. (B) pairwise qpWave indicated the genetic relationship among different populations (outgroup: Mbuti.DG, Russia_Shamanka_EBA.SG, Chokhopani, Boisman_MN, DevilsCave_N.SG, Shandong_EN, Tanshishan, Tarim_EMBA1, Russia_MLBA_Sintashta, GaoHuaHua) (C) Results of unsupervised ADMIXTURE clustering analysis. It presents a subset of representative populations, whereas the complete ADMIXTURE results including all reference populations are shown in Fig. S1.

Given this compression and the extensive shared ancestry among East Asian populations, fine-scale differentiation among the studied groups was not fully resolved in the broad Eurasian PCA. To improve resolution, we therefore conducted an additional PCA restricted to modern East Asian populations and visualized the studied groups separately according to their ethnic affiliations (Fig. S1).

In parallel, to further explore fine-grained population structure beyond visual clustering patterns, we performed pairwise qpWave analyses among the studied populations. The qpWave results revealed significant genetic heterogeneity among different ethnic groups (P-value > 0.05), supporting the presence of substructure suggested by the PCA (Fig. 2B).

In the PCA plot, our target populations clustered with others from the GQ region (e.g., Tu, Dongxiang, Eastern Yugur, Bonan, Tibetan_Gannan, Tibetan_Gangcha, etc.). The mixed-language Han Chinese in Linxia, Xiahe, and Gan’gou (designated as Han_Linxia, Han_Xiahe, and Han_Gangou) exhibited a slight genetic shift from East Asians towards Central and Western Eurasian populations, occupying an intermediate genetic position between modern Han Chinese and Tibetan highlanders. They showed closer genetic affinity to ancient populations associated with the YR region. However, the mixed-language Han Chinese in Wutun, specifically Han_Wutun2, clustered closer to Tibetan highlanders compared to Han_Wutun1 and other mixed-language Han Chinese. They also show closer proximity to ancient populations from the highland TP region in the PCA plot. Additionally, Hui individuals in Linxia (designated as Hui_Linxia1, Hui_Linxia2) and Tu individuals in Gan’gou and Wutun (designated as Tu_Gangou and Tu_Wutun) also showed a genetic shift towards Central and Western Eurasian populations. Tibetans in Gan’gou and Wutun (designated as Tibetan_Gangou and Tibetan_Wutun), consistent with other Tibetans in the GQ region (e.g., Tibetan_Gannan, Tibetan_Gangcha), deviated from the genetic cluster of core Tibetans (e.g., Tibetan_Chamdo, Tibetan_Lhasa, Tibetan_Nagqu, Tibetan_Shannan, and Tibetan_Shigatse). Notably, Tibetan_Wutun2 clustered closer to Tibetan highlanders compared to Tibetan_Wutun1 and Tibetan_Gangou. Furthermore, we identified several distinctive genetic outliers based on PCA: a Han individual from Wutun (designated as Han_Wutun_o) clustered with lowland Han populations (e.g., Han_Henan, Han_Shandong, Han_Jiangsu, Han_Guangdong and Han_Sichuan), a Hui individual from Gansu showed significant genetic contributions from Central and Western Eurasian populations (Hui_Linxia_o), and two Tu individuals from Wutun demonstrated closer genetic proximity to core Tibetan population clusters (Tu_Wutun_o). Overall, a complex genetic landscape was observed in our studied populations. Among the populations in Wutun, subgroups within different ethnic groups exhibited a genetic trend towards Tibetan highlanders (e.g., Han_Wutun2, Tibetan_Wutun2, Tu_Wutun_o).

We performed an unsupervised ADMIXTURE clustering analysis and observed the lowest cross-validation error at K = 5. At this value of K, multiple ancestral components were identified across the studied populations (Fig. 2C). Among these, four major ancestral components were observed within our newly sampled populations. These included a northern East Asian-related component (represented by orange), a southern East Asian-related component (represented by pink), an East Asian component enriched in ancient and present-day TP populations (represented by green), and a Central and Western Eurasian-related component (represented primarily by the blue and yellow components). Notably, the green component was detected in both lowland and highland populations, but occurred at substantially higher proportions in highland TP populations. This pattern suggests that the green component represents an East Asian ancestral lineage that is widespread across the region but has been differentially enriched in TP populations, rather than a genetic component exclusively restricted to highland groups. Compared with other lowland East Asian populations (e.g., lowland Han Chinese), populations from northwestern China—including our newly sampled groups (Han_Linxia, Han_Xiahe, and Han_Gangou), Mongolic-speaking populations (Bonan, Dongxiang, and Eastern Yugur), and Turkic-speaking populations (Kazakh, Turkmen, and Uyghur)—exhibited a higher proportion of Central and Western Eurasian-related ancestry, as reflected by the increased blue and yellow components (Fig. 2C presents a subset of representative populations, whereas the complete ADMIXTURE results including all reference populations are shown in Fig. S2).

3.2. Genetic relationships between target populations and east Asian reference populations

We note that PCA primarily captures major axes of genetic variation and is not specifically designed to reflect linguistic boundaries, especially in regions characterized by long-term gene flow and admixture. To explore the genetic relationships between populations from the GQ region and other East Asian regions in greater detail, we employed several analytical methods, including those based on allele frequencies, shared identical-by-descent (IBD) segments, and a shared haplotype model.

In the Neighbor-Joining phylogenetic tree constructed using Fst distances (Fig. S3), our target populations, with a few exceptions, clustered closely with other Sino-Tibetan language speakers. Specifically, some populations, such as Han_Linxia, Han_Xiahe, Han_Gangou, Tu_Gangou, and Tibetan_Gangou, were more closely aligned with lowland populations. Meanwhile, populations in Wutun, such as Han_Wutun2, Tibetan_Wutun2, and Tu_Wutun_o, showed closer clustering with the Tibetans. This clustering pattern was further supported by the maximum likelihood phylogeny inferred by TreeMix. The TreeMix results with migration edge set to 0 (Fig. S4A) revealed three major clades: one comprising West Eurasians and Turkic-speaking populations, another consisting of lowland Han Chinese, and a third comprising core Tibetic-speaking populations.

These observations were further validated through analyses based on shared IBD segments (Fig. S4B), which demonstrated a tendency for populations within the same language family to cluster together. Notably, different populations in Wutun shared the most IBD segments among themselves, indicating more frequent genetic exchange within the local Wutun people.

Using the haplotype-sharing model, we conducted a fine-scale genetic structure analysis of our target samples and surrounding populations via fineSTRUCTURE analysis (Fig. S4C). The results showed that GQ populations tend to cluster more closely with each other compared to Tibetan highlanders and lowland Han Chinese populations.

To explore patterns of gene flow that might have occurred, we ran TreeMix with migration edges ranging from 0 to 8 (Fig. S5). We noticed that the overall log-likelihood increased monotonically with additional migration edges, but consistent patterns emerged across models. In all tested values of m, strong migration edges were inferred from West Eurasian populations into reference northwestern Chinese groups, particularly Uyghur, Kazakh, and Kyrgyz, indicating a robust and dominant signal of West Eurasian ancestry in these populations. As model complexity increased, additional migration edges from West Eurasian-related sources into populations from the GQ region became apparent. These edges were generally associated with smaller weights and appeared only at higher values of m, suggesting that they likely reflect the spatial diffusion and long-term accumulation of West Eurasian ancestry rather than discrete, independent migration events.

Based on these recurrent patterns, we subsequently applied f-statistics to formally test and quantify West Eurasian gene flow into western and northwestern Chinese populations. We then calculated outgroup-f3 statistics in the form of f3 (Target, X; Mbuti) to assess the shared genetic drift between our studied populations and other East Asian reference populations (Fig. S6; Table S2). Focusing on mixing/mixed-language Han Chinese, we found they shared the most genetic drift with other Han Chinese populations (e.g., Han_Shanghai, Han_Hubei, Han_Chongqing, Han_Henan), followed by some Sino-Tibetan linguistic groups (e.g., Qiang_Danba, Tibetan_Xinlong, Tibetan_Yunnan). We also noted significant negative values (Z-scores < −3) in the form of f4 (Mbuti, West Eurasians; mixing/mixed-language studied Han, lowland Han Chinese) (Table S3), except in f4 (Mbuti, West Eurasians; Han_Wutun_o, lowland Han Chinese), suggesting that West Eurasians share more alleles with most of our target Han Chinese compared to other lowland Han Chinese. Furthermore, for Han_Wutun2, which exhibits a closer genetic relationship with Tibetan highlanders as indicated by previous analyses, many negative results of f4 (Mbuti, Tibetan highlanders; Han_Wutun2, other mixing/mixed-language studied Han) (−3 < Z-scores < −2) (Table S4) suggested that Han_Wutun2 may possess a higher proportion of highland TP-related genetic components compared to other mixing/mixed-language Han Chinese.

3.3. Reconstruction of the ancestral origin and composition of the mixing/mixed-language Han Chinese

Ancient genomes were included as temporal references to contextualize the ancestry components of modern populations. The outgroup-f3 analysis (mixed-language Han, ancient East Asian; Mbuti) revealed that the mixed-language Han Chinese in the GQ region share the most genetic drift with YR-related ancestries, such as the YR-related ancestries during the Late Bronze and Iron Ages (YR_LBIA), the YR-related ancestries during the Late Neolithic (YR_LN), the Upper YR-related ancestries during the Late Neolithic (Upper_YR_LN), and the Upper YR-related ancestries during the Iron Age (Upper_YR_IA) (Fig. S6). Our analyses indicated an influence from the West Eurasian steppe. Considering previous observations of potential genetic influences from the highland TP on some populations, we conducted f4 statistics in the form of f4 (Mbuti, target; highland TP-related ancestries, YR-related ancestries) and found numerous statistically significant positive values (Z-scores > 3) (Table S11). Due to long-term genetic stability in the Central Plains, genetic continuity related to ancient YR populations has persisted from the Neolithic period to the present day [50]. When using YR_LBIA—a YR-related ancestry from the Longshan culture period—as a representative of YR ancestral components, and Kyang—a Nepal ancient population from highland TP—as a representative of highland TP-related ancestral components, the f4 (Mbuti, Target; Kyang, YR_LBIA) indicated significant positive results (Z-scores > 3), with Han_Wutun2 emerging as a unique exception (Z-scores < 3). We then conducted an outgroup-dropping pairwise qpWave test to investigate the cause of heterogeneity within the paired group. When Kyang was removed from the outgroups, we observed significant genetic homogeneity (P-value > 0.05 for rank = 0) between Han_Wutun2 and Tibetan_Gangou/Tibetan_Wutun2. This finding suggests that highland TP-related ancestral components might have a unique gene flow with Han_Wutun2 (Fig. S7).

We then used plausible YR-related ancestry (represented by YR_LBIA), highland TP-related ancestry (represented by Kyang), and West Eurasian steppe-related ancestry (represented by Russia_Alan) as sources to construct admixture models using qpAdm (Fig. 3; Table S12). The results revealed that the mixed-language Han Chinese (Han_Linxia, Han_Xiahe, Han_Gangou, and Han_Wutun) in GQ harbored approximately 49.6%–92.9% genetic contributions from YR-related ancestries and 5.5%–10.7% genetic contributions from West Eurasian steppe-related ancestries. However, one Han Chinese subgroup from Wutun (Han_Wutun2) rejected the two-way admixture model comprising YR-related and West Eurasian steppe-related components, requiring an additional 44.9% of highland TP-related ancestry components. The results indicate the presence of at least three genetic influxes originating from the YR Basin, West/Central Eurasian steppe, and the Qinghai-Tibet Plateau, which collectively influenced the mixed-language populations in the GQ region. Notably, highland TP-related ancestry was detected only in certain Han Chinese from Wutun, where it constituted a substantial proportion (44.9%), suggesting significant genetic interactions between them and Tibetan highlanders in the past.

Fig. 3.

Fig 3 dummy alt text

The result of qpAdm indicated at least three genetic influxes (from the YR Basin, west/central Eurasia steppe, and the Qinghai-Tibet Plateau) influenced the mixing/mixed-language Han Chinese in the GQ region. The geographicbase map was based on the Standard Map Service System of the Ministry of Natural Resources of China (http://bzdt.ch.mnr.gov.cn) (GS(2020)4619).

We estimated the dates of East-West genetic admixture for diverse ethnic groups across the four sampling sites using ALDER. The analysis utilized modern East Asian populations (comprising lowland Han and Sino-Tibetan speakers) and West Eurasian groups, including GBR.SG (SG means shotgun sequencing genomes), as well as Estonian.DG, Georgian.DG, and Greek.DG (DG means high-coverage shotgun genomes with diploid genotype calls) (Fig. 4; Table S13). The admixture event occurred 29.8–40.3 generations ago (approximately 860–1160 years ago, assuming 29 years per generation), corresponding to the Tang and Song dynasties, during the flourishing period of the Silk Road. While the ALDER results suggest an admixture time around 29.8–40.3 generations ago, it is important to note that this method primarily captures the most recent or the strongest pulse of gene flow and does not preclude earlier episodes of gene flow. In the context of the GQ populations, the historical formation likely involved continuous or multiple waves of interactions. The population genetic studies have shown that east-west genetic interactions in the GQ region began as early as the Bronze Age and were further shaped by southward movements of nomadic populations during the early Iron Age. We therefore interpret our results as reflecting a renewed or intensified phase of admixture during the Tang-Song transition, likely associated with large-scale population movements, sociopolitical restructuring, and demographic reorganization in northwestern China, rather than the initial onset of east-west contact.

Fig. 4.

Fig 4 dummy alt text

The east-west genetic admixture date of different populations in the GQ region inferred by ALDER.

To further characterize the complexity of admixture beyond single-pulse models, we applied MultiWaver v2.0 [45] to infer the admixture history of populations from northwestern China. Based on prior results from qpAdm and ALDER, representative East Asian (Han Chinese) and West Eurasian (Great Britains) populations were selected as proxies. MultiWaver evaluates multiple admixture scenarios under both discrete and continuous models. The best-supported model was a multiple-wave admixture model, with a bootstrap support ratio of 89% (Table S14). Under this model, two distinct east-west admixture waves were detected, occurring approximately 21–26 and 29–37 generations ago. This means an additional east-west admixture event in the GQ region that dates to approximately the Yuan dynasty.

4. Discussion and conclusion

4.1. The genetic and linguistic landscape of non-Han populations in the GQ region

Our genome-wide analyses reveal pronounced genetic heterogeneity among non-Han populations in the GQ region, reflecting distinct demographic histories that are only partially aligned with current linguistic affiliations. Despite their shared residence within a single contact zone, Tibetic-, Mongolic-, and Turkic-speaking populations exhibit markedly different proportions of YR-related, highland TP-related, and West Eurasian Steppe-related ancestries, indicating mlultiple population sources and complex admixture processes.

Our findings reveal genetic differentiation between Tibetans from the GQ region and those from the TP. Outgroup-f3 statistics indicate that the studied Tibetan groups (Tibetan_Gangou, Tibetan_Wutun1, and Tibetan_Wutun2) share more genetic drift with lowland Han Chinese than with core Tibetan populations. This pattern is further supported by significantly positive f4-statistics of the form f4 (Mbuti, studied Tibetan; core Tibetan, lowland Han Chinese) (Z-scores > 3) (Table S5), with the exception of Tibetan_Wutun2. The lack of a significant signal in Tibetan_Wutun2 may suggest a relatively higher proportion of genetic ancestry related to core Tibetan populations in this group, a possibility that warrants further investigation. We also examined genetic disparities between GQ Tibetans and core Tibetans from the highland TP using f4 (Mbuti, X; studied Tibetan, core Tibetan) (Table S6). When X represented populations from southern China and Central and Western Eurasia, f4 (Mbuti, X; Tibetan_Gangou/Tibetan_Wutun1, core Tibetan) showed numerous statistically significant negative values (Z-scores < −3). In contrast, significant negative values (-3 < Z-scores < −2) appeared only when X comprised specific ancient Western populations. Some positive f4 values (Mbuti, Tibetan highlanders; Tibetan_Wutun1, Tibetan_Wutun2) (2 < Z-scores < 3) (Table S7) suggest that Tibetan_Wutun2 shares more alleles with Tibetan highlanders compared to Tibetan_Wutun1. Using qpAdm, we quantified ancestry, revealing that Tibetan_Gangou could be modeled as 91.5% YR-related and 8.5% West Eurasia Steppe-related. Similarly, Tibetan_Wutun1 was modeled as 91.6% YR-related and 8.4% West Eurasia Steppe-related. In contrast, a three-way admixture model comprising 60.5% YR-related, 32.5% highland TP-related, and 7% West Eurasia Steppe-related was necessary to model Tibetan_Wutun2.

Previous pairwise qpWave results showed genetic heterogeneity among Tu_Gangou, Tu_Wutun, and Tu_Wutun_o (Fig. 2B). Significant negative f4 values (Mbuti, YR; Tu_Gangou, other Tu) (Z-scores < −3) suggest that Tu_Gangou shares more alleles with YR-related populations. Comparatively, Tu_Wutun shares more alleles with Central and Western Eurasians, indicated by negative f4 values (Mbuti, West Eurasians; Tu_Wutun, other Tu) (-3 < Z-scores < −2) (Table S8). Positive f4 values (Mbuti, highland TP-related; Tu_Wutun, Tu_Wutun_o) (Z-scores > 3) indicate that Tu_Wutun_o shares more alleles with highland TP-related ancestries compared to Tu_Wutun. Tu_Gangou could be modeled as 92.1% YR-related and 7.9% West Eurasia Steppe-related ancestries, Tu_Wutun as 67.5% YR-related, 22.7% highland TP-related, and 9.8% West Eurasia Steppe-related, while Tu_Wutun_o as 52.6% YR-related, 38.9% highland TP-related, and 8.5% West Eurasia Steppe-related ancestries.

We observed a stronger Western Eurasian influence in the Hui population. Significant negative f4 values (Mbuti, West Eurasians; studied Hui, lowland Han Chinese & other target populations) (Z-scores < −3) were detected (Table S9). Additionally, negative f4 results (Mbuti, Central/Western Eurasians; Hui_Linxia_o, Hui_Linxia1/Hui_Linxia2) indicate that the outlier individual Hui_Linxia_o shares more alleles with Central/Western Eurasian populations compared to other Hui individuals (Table S10). The Hui in Linxia can be successfully modeled using a YR-related and West Eurasia Steppe-related two-way admixture model, with Hui_Linxia_o showing the highest proportion of West Steppe-related ancestry at 19.6%.

In summary, we identified widespread Western Eurasian-related ancestral components among GQ populations, not only in Sinitic-speaking populations (e.g., Han, Hui) but also in Mongolic-speaking (e.g., Bonan, Dongxiang, Eastern Yugur, Tu), Tibetic-speaking (e.g., Tibetan in Gan’gou, Wutun, Xunhua, Gangcha), and Turkic-speaking (e.g., Salar) populations (Fig. 5; Table S12). Most non-Sinitic speaking populations exhibited significantly greater West Eurasian Steppe-related ancestry compared to Han Chinese from Linxia, Xiahe, Gan’gou, and Wutun, with the Turkic-speaking Salar having the highest proportion. Interestingly, Ancient Northeast Asia (ANA)-related ancestry was only detected in Inner Mongolian Mongols, not in other Mongolic-speaking groups from the GQ region, highlighting the heterogeneous origins of these populations (Fig. 5; Table S12). These genetic-linguistic mismatches can be interpreted within the elite dominance model [51] and the cultural dominance model [23] of language replacement. Historical records and paternal genetic evidence indicate that Mongolic-speaking populations (Dongxiang, Bonan, and Eastern Yugur) in the GQ region originated from West/Central Eurasia [22,23,52]. However, these groups, whether voluntarily or coercively, abandoned their ancestral languages and adopted the Mongolic languages of their rulers. In contrast, the Turkic-speaking Salar people, who also shared West-Central Eurasian-related ancestry, have retained their Turkic language.

Fig. 5.

Fig 5 dummy alt text

(A) A simple Han and non-Han language diagram with reference to previous linguistic research [28]; (B) Ancestral components of different language families inferred by qpAdm. For populations exhibiting substructure (e.g., Han Chinese in Wutun), the mean ancestral proportions across genetic clusters were calculated to represent their composite ancestry profiles. Here, the ancestral component of Wutun represents the average of Han_Wutun1 and Han_Wutun2. Additionally, given the extensive geographical scope of the Amdo Tibetan region and the lack of a well-defined population that can adequately represent its genetic characteristics, no specific population was included in this study. This issue warrants further investigation in future research.

The elite dominance model [51] provides a useful framework for interpreting cases in which linguistic affiliation is decoupled from genetic ancestry. In the GQ region, our results suggest that at least two forms of elite-driven language replacement operated. The formation of Dongxiang and Bonan populations is more consistent with a ruling-group elite dominance scenario, in which language shift occurred under sustained political and social dominance. Historical and interdisciplinary evidence indicates that their ancestors originated from Central or Western Eurasia and were relocated into China during the westward expansions of the Mongol Empire in the 13th century [23,52,53]. Through prolonged interaction and admixture with Mongolic, Han, and Tibetan populations, their original languages (e.g., Turkic, Arabic, or Persian) were gradually replaced by Mongolic languages. This process plausibly explains the observed mismatch between linguistic classification and genetic ancestry: although Dongxiang and Bonan are classified within the Mongolic branch, our whole genome-based ancestry modeling reveals substantial West Eurasian Steppe-related ancestry (∼10.2%−16%), supporting a scenario of language replacement without complete population turnover.

In contrast, the linguistic formation of the Yugur populations appears more compatible with a leader-group elite dominance model. Despite the long-term coexistence of Eastern Yugur (Mongolic) and Western Yugur (Turkic) languages, no significant linguistic convergence is observed. Ethnohistorical and linguistic studies suggest that multiple tribes integrated under socially dominant leadership groups, adopting their languages in the process. Consistent with this interpretation, molecular anthropological evidence indicates that Eastern Yugur individuals carry low frequencies of Y-chromosome haplogroups typical of Mongolic populations and elevated frequencies of West Eurasian-associated haplogroups [23]. Our ancestry modeling based on published whole-genome data further identifies a West Eurasian Steppe-related component (∼6.3%) in Eastern Yugur, implying that their ancestors also likely derived from Central or Western Eurasia and underwent a language shift after migration into China.

Interestingly, language maintenance rather than replacement is also observed in the same region, as exemplified by the Turkic-speaking Salar population. Similar to the Dongxiang and Bonan, the ancestors of the Salar are generally considered to have originated from Central and Western Eurasia. However, unlike the former groups, the migration of the Salar ancestors appears to have been largely voluntary, motivated by the desire to escape political oppression by local rulers. Historical sources suggest that the Salar enjoyed a relatively high social status during the Yuan dynasty; Saguchi noted their elevated status [54], and Gui’e Yu and colleagues reported that the Salar participated in Mongol military campaigns and were consequently favored by Mongol rulers [55]. In this sociopolitical context, the persistence of the Salar language, rather than its replacement by Mongolic languages, is therefore understandable. Molecular anthropological studies further show that the Central Asian-associated Y chromosomal haplogroup R1a1a*-M17 occurs at notably higher frequencies in the Salar than in the Dongxiang and Bonan populations. Consistently, our genome-wide ancestry modeling indicates that the Salar harbor approximately 18.3% West Eurasian Steppe-related ancestry.

These patterns demonstrate a dual process in the GQ region: non-Han populations underwent genetic assimilation with Han Chinese while experiencing complex linguistic admixture and replacement.

4.2. The co-diffusion of genes and languages in the Han Chinese of the GQ region

In contrast to many non-Han populations, Han Chinese in the GQ region exhibit variable degrees of correspondence between genetic admixture and linguistic convergence. Genome-wide analyses reveal modest but detectable West Eurasian Steppe-related ancestry (approximately 5.5%−10.7%) in most mixing or mixed-language Han populations, while the Han_Wutun2 group shows substantial genetic contribution from highland TP-related ancestry (∼44.9%). These results indicate that genetic exchange accompanied language contact in some contexts, but not uniformly across all Han populations.

From a linguistic perspective, the GQ region represents an extreme case of long-term structural convergence driven by sustained multilingual contact. While language contact phenomena are common globally, the convergence observed in GQ Sinitic varieties is typologically significant because it involves structural changes at the core syntactic level. These changes include a shift toward object-verb (OV) word order, reduction or loss of tonal systems, extension of plural marking to inanimate nouns, and the incorporation of non-Sinitic morphological affixes [56]. Crucially, this convergence proceeds in a direction that runs counter to the typologically dominant profile of Sinitic languages, which are characteristically verb-object (VO), tonal, and morphologically isolating. The co-occurrence of multiple non-Sinitic structural features within individual Sinitic varieties therefore reflects sustained and intensive contact rather than sporadic borrowing.

We examined the phylogenetic relationships of different language families (Fig. 5A, a simple diagram of Han and non-Han languages, referencing previous studies) in relation to the ancestral composition of their speakers (Fig. 5B), revealing that populations across all language families primarily possess ancestral components from the YR Basin, with influences from West Eurasian Steppe-related ancestry. In Wutun village, we observe a clear signal of coupled genetic and linguistic diffusion, characterized by substantially highland TP–related ancestry (approximately 44.9% in Han_Wutun2) alongside well-documented Tibetan structural influence on the Wutun language [24,25,57]. Linguistic studies further indicate that the Wutun language was additionally influenced by Mongolic languages. Typologically, the Wutun language may be characterized as a Tibetanized form of Chinese with some Mongolic features [25]. From a genetic perspective, Wutun shows a pattern similar to other Mongolic-speaking populations in the GQ region (e.g., Bonan, Dongxiang, Eastern Yugur, and Tu), characterized by the absence of detectable ANA-related ancestry. Although the lack of ANA-related ancestry does not directly imply genetic input from Mongolic-speaking populations, this shared genetic pattern, when considered alongside linguistic evidence, reflects regionally specific processes of language contact and cultural transmission.

In contrast, other Han Chinese with emerging linguistic varieties in the region do not exhibit detectable highland Tibetan-related genetic components, despite clear evidence of language contact and structural convergence. Based on the findings, we propose that language mixing can also occur in the absence of substantial gene flow, likely mediated by sociopolitical dominance, cultural transmission, or prolonged bilingualism rather than direct population replacement. Nevertheless, all mixing- and mixed-language Han Chinese populations examined here share a common genetic feature with other populations in the GQ region—the West/Central Eurasian-related ancestry. The frontier nature of the GQ region, characterized by frequent population mobility, military deployment, and multi-ethnic coexistence across historical periods, created a social environment in which genetic admixture and language contact could proceed asynchronously. Fine-scale analyses (Fst, IBD, and fineSTRUCTURE) further indicate that geographic proximity and frequent population movements promote genetic clustering across linguistic boundaries in regions characterized by intensive language contact.

Temporal analyses provide additional insight into the demographic context of these processes. The ALDER-based estimates suggest that east-west genetic admixture in the GQ region occurred approximately 29.8–40.3 generations ago, broadly corresponding to the Tang-Song period. Previous ancient human DNA analysis showed that clear East-West genetic admixture signals are only detected in a small number of outlier individuals from Dunhuang, dated to the Cao-Wei and Tang periods, but not to Han and Tang dynasties, when the Silk Road was most prosperous. It may suggest a limited demographic impact of Silk Road-era contacts. However, earlier east-west interactions, although historically well documented, may have been temporally diffuse or subsequently diluted by later demographic processes, rendering them less detectable by linkage disequilibrium-based methods. In this context, the MultiWaver analysis provides complementary insights by supporting a multiple-wave admixture model, indicating that the genetic landscape of the GQ region was shaped by temporally layered east-west interactions rather than a single event, with another east-west admixture event may have occurred more recently, during the Yuan Dynasty. Such a temporally layered process is consistent with archaeological and historical evidence documenting long-term population mobility and intermittent contact between East Asian and West Eurasian groups across different historical periods.

Additionally, historical records suggest that the Yuan dynasty created sociopolitical conditions conducive to increased population contact in northwestern China, including large-scale military deployments, administrative resettlement of Central and Western Eurasian groups, and long-term garrisoning in frontier regions such as GQ. These structural conditions likely facilitated sustained local interactions and intermarriage, thereby enhancing genetic exchange. Accordingly, the admixture signals detected by ALDER and MultiWaver are best understood as reflecting a recent intensification of long-standing east-west interactions under changing political and demographic contexts, rather than the initial emergence of such contacts.

Together, we can identify the primary factors driving the divergent co-evolutionary dynamics in these Han populations. First, contact intensity and demographic history played decisive roles: the mixed-language Wutun population emerged from a coupled process where substantial genetic admixture with Tibetan highlanders (demographic integration) paralleled the formation of a heavily Tibetanized Sinitic language. In contrast, the mixing-language Han groups in Linxia and Xiahe exhibit an uncoupled pattern, where sociopolitical mechanisms, such as the military garrison system and trade dominance along the Silk Road, facilitated the maintenance of YR-related genetic ancestry despite profound linguistic structural convergence induced by prolonged multilingualism. Furthermore, geographic structure acted as a scaffold for these interactions, with the GQ corridor serving as a contact zone where distinct ecological zones (lowland vs. highland) intersected. Finally, the timing of these interactions, intensified during the Tang-Song transition and Yuan dynasty, suggests that these evolutionary trajectories were not gradual drifts but were punctuated by specific historical phases of intensified East-West interaction.

It is important to note that the current study has limitations in sample coverage. First, geographic coverage limitations, particularly low sampling density in areas where Chinese dialects are actively in contact, result in gaps in the phylogenetic completeness of linguistic variation. Second, the existing data have not fully captured the dynamic evolutionary trajectories of complex language contact phenomena.

Human language and genetic systems exhibit deep structural similarities: both are information systems that rely on human carriers for transmission and evolution, and they share analogous evolutionary mechanisms in population diffusion, intergenerational transmission, and adaptive evolution [58]. Our study establishes a gene-language coupling model, providing a novel interdisciplinary analytical framework for understanding the co-evolution of these systems. This integration of biology and linguistics advances the reconstruction of human civilization’s evolutionary history, particularly by elucidating the synchrony between early population migrations and language differentiation.

CRediT authorship contribution statement

Haodong Chen: Writing – original draft, Visualization, Investigation, Formal analysis, Data curation. Huangzhen Huang: Writing – original draft, Visualization, Investigation, Formal analysis, Data curation. Boxuan Zhao: Investigation, Formal analysis, Data curation. Yiling Jiang: Investigation, Formal analysis, Data curation. Mengting Xu: Investigation. Limin Qiu: Investigation. Jin Sun: Investigation. Yourong Ye: Investigation. Hongbing Yao: Investigation. Xinyi Jia: Investigation. Kai Sun: Investigation. Jiangtao Shen: Investigation. Hao Li: Investigation. Dan Xu: Writing – review & editing, Supervision, Project administration, Conceptualization. Chuan-Chao Wang: Writing – review & editing, Supervision, Project administration, Funding acquisition, Conceptualization. Shaoqing Wen: Writing – review & editing, Writing – original draft, Supervision, Project administration, Investigation, Funding acquisition, Conceptualization.

Declaration of competing interest

The authors declare that they have no conflicts of interest in this work.

Acknowledgments

This work was funded by the European Research Council (ERC-2019-ADG, 883700-TRAM), the Lantai Youth Scholar Program (2022LTON602), the National Natural Science Foundation of China (T2425014 and 32270667), the Natural Science Foundation of Fujian Province of China (2023J06013), the Major Project of the National Social Science Foundation of Chin (21&ZD285), Open Research Fund of State Key Laboratory of Genetic Engineering at Fudan University (SKLGE-2310), Open Research Fund of Forensic Genetics Key Laboratory of the Ministry of Public Security (2023FGKFKT07), and National Key Research and Development Program of China (2023YFC3303701-02, 2024YFC3306701). The authors extend our sincere gratitude to the linguists for sharing their valuable research findings. The authors want to thank Saiyinjiya Caidengduoerji (for Mongolian languages), Barbara Kozhevina (for Turkic languages), Li Ting (for Tibetan languages), Liu Keyou and Wang Cong (for Sinitic languages). Besides, the authors are grateful to Zhu Kongyang for sharing the qpAdm script used for ancestral population simulation.

Biographies

Dan Xu, the PI of the ERC-2019-AdG-883700-TRAM, obtained her PhD from Sorbonne University, Paris, in 1987. She worked at the Chinese Academy of Social Sciences and as a professor at INALCO, Paris, and at JGU, Mainz, Germany. She was elected Senior Member of the Institut Universitaire de France in 2009 and Member of Academia Europaea in 2021.

Chuan-Chao Wang obtained his B.S. degree (2010) from Ocean University of China and his Ph.D. degree (2015) from Fudan University. He received postdoctoral training at Harvard Medical School and the Max Planck Institute. He joined Xiamen University in 2017 as a professor and Fudan University in 2024 as a distinguished professor. His work primarily focuses on using a genetic approach to study the genetic history of human populations.

Shaoqing Wen obtained his Ph.D. degree from Fudan University in 2017. He joined the Institute of Archaeological Science at Fudan University in 2019 as a junior research fellow and was promoted to associate professor in 2022. His research interest focuses on molecular archaeology.

Footnotes

Peer review under the responsibility of Editorial Board of Fundamental Research.

Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.fmre.2026.04.001.

Contributor Information

Dan Xu, Email: dxusong@uni-mainz.de.

Chuan-Chao Wang, Email: chuanchaowang@fudan.edu.cn.

Shaoqing Wen, Email: wenshaoqing@fudan.edu.cn.

Appendix. Supplementary materials

mmc1.docx (13.5KB, docx)
mmc2.xlsx (1.8MB, xlsx)
mmc3.pdf (378.9KB, pdf)
mmc4.pdf (393.1KB, pdf)
mmc5.pdf (127KB, pdf)
mmc6.pdf (535KB, pdf)
mmc7.pdf (10MB, pdf)
mmc8.pdf (607.8KB, pdf)
mmc9.pdf (197.7KB, pdf)

References

  • 1.Cavalli-Sforza L.L., Feldman M.W. Cultural transmission and evolution: A quantitative approach. Monogr. Popul. Biol. 1981;16:1–388. [PubMed] [Google Scholar]
  • 2.Cavalli-Sforza L.L. Genes, peoples, and languages. Proc. Natl. Acad. Sci. 1997;94(15):7719–7724. doi: 10.1073/pnas.94.15.7719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Colucci M., Leonardi M., Hodgson J., et al. The legacy of Luca Cavalli-Sforza on human evolution. Hum. Popul. Genet. Genom. 2025;5(1):0001. [Google Scholar]
  • 4.Haak W., Lazaridis I., Patterson N., et al. Massive migration from the steppe was a source for Indo-European languages in Europe. Nature. 2015;522(7555):207–211. doi: 10.1038/nature14317. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Allentoft M.E., Sikora M., Sjögren K.-G., et al. Population genomics of Bronze Age Eurasia. Nature. 2015;522(7555):167–172. doi: 10.1038/nature14507. [DOI] [PubMed] [Google Scholar]
  • 6.Narasimhan V.M., Patterson N., Moorjani P., et al. The formation of human populations in South and Central Asia. Science. 2019;365(6457):eaat7487. doi: 10.1126/science.aat7487. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ning C., Wang C.C., Gao S., et al. Ancient genomes reveal Yamnaya-related ancestry and a potential source of Indo-European speakers in iron age Tianshan. Curr. Biol. 2019;29(15):2526–2532.e4. doi: 10.1016/j.cub.2019.06.044. [DOI] [PubMed] [Google Scholar]
  • 8.Lazaridis I., Patterson N., Anthony D., et al. The genetic origin of the Indo-Europeans. Nature. 2025;639(8053):132–142. doi: 10.1038/s41586-024-08531-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zhang M.H., Yan S., Pan W.Y., et al. Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic. Nature. 2019;569(7754):112–115. doi: 10.1038/s41586-019-1153-z. [DOI] [PubMed] [Google Scholar]
  • 10.Yang C.K., Zhang X.X., Yan S., et al. Large-scale lexical and genetic alignment supports a hybrid model of Han Chinese demic and cultural diffusions. Nat. Hum. Behav. 2024;8(6):1163–1176. doi: 10.1038/s41562-024-01886-9. [DOI] [PubMed] [Google Scholar]
  • 11.Campbell L. In: Linguistic Areas: Convergence in Historical and Typological Perspective. Matras Y., McMahon A., Vincent N., editors. Palgrave Macmillan; UK, London: 2006. Areal linguistics: A closer scrutiny; pp. 1–31. [Google Scholar]
  • 12.Zang X.W. Edward Elgar Publishing; 2016. Handbook on Ethnic Minorities in China. [DOI] [Google Scholar]
  • 13.Dede K. The Chinese language in Qinghai, Stud. Orient. Electron. 2003;95:321–346. [Google Scholar]
  • 14.Peyraube A. In: Languages and Genes in Northwestern China and Adjacent Regions. Xu D., Li H., editors. Springer; Singapore: 2017. The case system in three Sinitic languages of the Qinghai-Gansu linguistic area; pp. 121–139. [Google Scholar]
  • 15.Dwyer A.M. Altaic elements in the Línxià dialect: Contact-induced change on the YR Plateau. J. Chin. Linguist. 1992;20(1):160–179. [Google Scholar]
  • 16.Chen N.X. Inner Mongolia People’s Publishing House; 1986. Baoan and Mongolian Languages.books.google.com/books?id=vJiQYgEACAAJ [Google Scholar]
  • 17.Yao H.B., Wang M.G., Zou X., et al. New insights into the fine-scale history of western-eastern admixture of the northwestern Chinese population in the Hexi Corridor via genome-wide genetic legacy. Mol. Genet. Genom. 2021;296(3):631–651. doi: 10.1007/s00438-021-01767-0. [DOI] [PubMed] [Google Scholar]
  • 18.Ma B., Chen J.W., Yang X.M., et al. The genetic structure and east-west population admixture in Northwest China inferred from genome-wide array genotyping. Front. Genet. 2021;12 doi: 10.3389/fgene.2021.795570. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.He G.L., Yao H.B., Duan S.H., et al. Pilot work of the 10K Chinese people genomic diversity project along the Silk Road suggests a complex east-west admixture landscape and biological adaptations. Sci. China Life Sci. 2025;68(4):914–933. doi: 10.1007/s11427-024-2748-4. [DOI] [PubMed] [Google Scholar]
  • 20.Pan Y.W., Zhang C., Lu Y., et al. Genomic diversity and post-admixture adaptation in the Uyghurs. Natl. Sci. Rev. 2022;9(3):nwab124. doi: 10.1093/nsr/nwab124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Zhou R.X., An L.Z., Wang X.L., et al. Testing the hypothesis of an ancient Roman soldier origin of the Liqian people in northwest China: A Y-chromosome perspective. J. Hum. Genet. 2007;52(7):584–591. doi: 10.1007/s10038-007-0155-0. [DOI] [PubMed] [Google Scholar]
  • 22.Zhou R.X., Yang D.Q., Zhang H., et al. Origin and evolution of two Yugur sub-clans in Northwest China: A case study in paternal genetic landscape. Ann. Hum. Biol. 2008;35(2):198–211. doi: 10.1080/03014460801922927. [DOI] [PubMed] [Google Scholar]
  • 23.Xu D., Wen S.Q. In: Languages and Genes in Northwestern China and Adjacent Regions. Xu D., Li H., editors. Springer; Singapore: 2017. The Silk road: Language and population admixture and replacement; pp. 55–78. [Google Scholar]
  • 24.Xu D. Why are some languages mixed in the GQ area? -The role of vocabulary and syntax in language mixing (in Chinese) Yuyanxue Luncong. 2024;4:77–88. [Google Scholar]
  • 25.Erika S. In: New Perspectives on Mixed Languages. Maria M., Eeva S., editors. De Gruyter Mouton; Berlin, Boston: 2021. Wutun as a mixed language; pp. 325–360. [Google Scholar]
  • 26.Ibrahim A. A linguistic sketch of Tangwang speech in Gansu province (in Chinese) Minor. Lang. China. 1985;6:33–47. [Google Scholar]
  • 27.Xu D., Wen S.Q. Springer; Singapore: 2017. Formation of a “Mixed Language” in Northwest China—The Case of Tangwang, Languages and Genes in Northwestern China and Adjacent Regions; pp. 87–105. [Google Scholar]
  • 28.Xu D. Springer; 2017. The Tangwang Language. link.springer.com/ [DOI] [Google Scholar]
  • 29.Mallick S., Micco A., Mah M., et al. The Allen Ancient DNA Resource (AADR) a curated compendium of ancient human genomes. Sci. Data. 2024;11(1):182. doi: 10.1038/s41597-024-03031-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Wang H.R., Yang M.A., Wangdue S., et al. Human genetic history on the TP in the past 5100 years. Sci. Adv. 2023;9(11):eadd5582. doi: 10.1126/sciadv.add5582. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Xiong J.X., Wang R., Chen G.K., et al. Inferring the demographic history of Hexi Corridor over the past two millennia from ancient genomes. Sci. Bull. (Beijing) 2024;69(5):606–611. doi: 10.1016/j.scib.2023.12.031. [DOI] [PubMed] [Google Scholar]
  • 32.Purcell S., Neale B., Todd-Brown K., et al. PLINK: A tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 2007;81(3):559–575. doi: 10.1086/519795. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Manichaikul A., Mychaleckyj J.C., Rich S.S., et al. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26(22):2867–2873. doi: 10.1093/bioinformatics/btq559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Weissensteiner H., Pacher D., Kloss-Brandstatter A., et al. HaploGrep 2: Mitochondrial haplogroup classification in the era of high-throughput sequencing. Nucleic Acids Res. 2016;44(W1):W58–W63. doi: 10.1093/nar/gkw233. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Chen H., Lu Y., Lu D.S., et al. Y-LineageTracker: A high-throughput analysis framework for Y-chromosomal next-generation sequencing data. BMC Bioinform. 2021;22(1):114. doi: 10.1186/s12859-021-04057-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Patterson N., Price A.L., Reich D. Population structure and eigenanalysis. PLoS Genet. 2006;2(12):e190. doi: 10.1371/journal.pgen.0020190. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Chang C.C., Chow C.C., Tellier L.C., et al. Second-generation PLINK: Rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7. doi: 10.1186/s13742-015-0047-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Alexander D.H., Novembre J., Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 2009;19(9):1655–1664. doi: 10.1101/gr.094052.109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Tamura K., Stecher G., Kumar S. MEGA11: Molecular evolutionary genetics analysis Version 11. Mol. Biol. Evol. 2021;38(7):3022–3027. doi: 10.1093/molbev/msab120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Pickrell J.K., Pritchard J.K. Inference of population splits and mixtures from genome-wide allele frequency data. PLoS Genet. 2012;8(11) doi: 10.1371/journal.pgen.1002967. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Felsenstein J. University of Washington; Seattle: 2005. PHYLIP (Phylogeny Inference Package) Version 3.6, Distributed by Author Department of Genome Sciences. evolution.genetics.washington.edu/phylip.html. [Google Scholar]
  • 42.Patterson N., Moorjani P., Luo Y., et al. Ancient admixture in human history. Genetics. 2012;192(3):1065–1093. doi: 10.1534/genetics.112.145037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Loh P.R., Lipson M., Patterson N., et al. Inferring admixture histories of human populations using linkage disequilibrium. Genetics. 2013;193(4):1233–1254. doi: 10.1534/genetics.112.147330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Fenner J.N. Cross-cultural estimation of the human generation interval for use in genetics-based population divergence studies. Am. J. Phys. Anthr. 2005;128(2):415–423. doi: 10.1002/ajpa.20188. [DOI] [PubMed] [Google Scholar]
  • 45.Ni X.M., Yuan K., Liu C., et al. MultiWaver 2.0: Modeling discrete and continuous gene flow to reconstruct complex population admixtures. Eur. J. Hum. Genet. 2019;27(1):133–139. doi: 10.1038/s41431-018-0259-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Delaneau O., Zagury J.-F., Marchini J. Improved whole-chromosome phasing for disease and population genetic studies. Nat. Methods. 2013;10(1):5–6. doi: 10.1038/nmeth.2307. [DOI] [PubMed] [Google Scholar]
  • 47.Browning B.L., Browning S.R. Improving the accuracy and efficiency of identity-by-descent detection in population data. Genetics. 2013;194(2):459–471. doi: 10.1534/genetics.113.150029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Browning S.R., Browning B.L., Daviglus M.L., et al. Ancestry-specific recent effective population size in the Americas. PLoS Genet. 2018;14(5) doi: 10.1371/journal.pgen.1007385. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Lawson D.J., Hellenthal G., Myers S., et al. Inference of population structure using dense haplotype data. PLoS Genet. 2012;8(1) doi: 10.1371/journal.pgen.1002453. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Ma H., Zhou Y.W., Wang R., et al. Ancient genomes shed light on the long-term genetic stability in the Central Plain of China. Sci. Bull. (Beijing) 2025;70(3):333–337. doi: 10.1016/j.scib.2024.07.024. [DOI] [PubMed] [Google Scholar]
  • 51.Colin R. Cambridge University Press; Cambridge: 1987. Archaeology and Language. [Google Scholar]
  • 52.Shou W.H., Qiao E.F., Wei C.Y., et al. Y-chromosome distributions among populations in Northwest China identify significant contribution from Central Asian pastoralists and lesser influence of western Eurasians. J. Hum. Genet. 2010;55(5):314–322. doi: 10.1038/jhg.2010.30. [DOI] [PubMed] [Google Scholar]
  • 53.Yao H.B., Wang C.C., Tao X.L., et al. Genetic evidence for an East Asian origin of Chinese Muslim populations Dongxiang and Hui. Sci. Rep. 2016;6 doi: 10.1038/srep38656. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Dwyer, Arienne M. The Turkic stratigraphy of Salar: An Oghuz in Chagatay clothes? Turkic Languages. 1998;2(1):49–83. [Google Scholar]
  • 55.Gui’e Y. Xinjiang Meishusheying/Dianziyinxiang Publishing House; Urumqi: 2010. The Salar Population (in Chinese) [Google Scholar]
  • 56.Xu D., Wen S.Q., Wang C.C. Shanghai Scientific & Technical Publishers; Shanghai: 2024. Language, Genes and Archaeology (in Chinese) [Google Scholar]
  • 57.Xu D. Mixed languages and mechanisms of language mixing within China (in Chinese) Chin. J. Lang. Policy Plan. 2018;2 doi: 10.19689/j.cnki.cn10-1361/h.20180206. [DOI] [Google Scholar]
  • 58.Pagel M. Human language as a culturally transmitted replicator. Nat. Rev. Genet. 2009;10(6):405–415. doi: 10.1038/nrg2560. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

mmc1.docx (13.5KB, docx)
mmc2.xlsx (1.8MB, xlsx)
mmc3.pdf (378.9KB, pdf)
mmc4.pdf (393.1KB, pdf)
mmc5.pdf (127KB, pdf)
mmc6.pdf (535KB, pdf)
mmc7.pdf (10MB, pdf)
mmc8.pdf (607.8KB, pdf)
mmc9.pdf (197.7KB, pdf)

Articles from Fundamental Research are provided here courtesy of The Science Foundation of China Publication Department, The National Natural Science Foundation of China

RESOURCES