Abstract
The Qiang people are widely recognized as a basal layer of the Chinese gene pool, yet their genetic origins and evolutionary history remain unclear. We analyzed 20 deep-sequenced Qiang genomes together with genomic data from Tibetan highlanders and neighboring lowlanders. The Qiang genetic structure has been profoundly shaped by historical admixture with Han and Tibetans, giving rise to distinct subpopulations. Using ~450 ancient Asian genomes, we traced the most recent common ancestor of the Qiang to ancient Yellow River farmers ~5300 years ago, indicating shared ancestry with other Chinese populations. We identified several highly differentiated variants related to ethanol metabolism and pigmentation between Qiang subpopulations, which likely arose from joint effects of admixture and selection. Notably, a highly prevalent missense variant in a blood pressure regulation gene was detected, suggesting a potential role in altitude adaptation. Collectively, our findings illuminate the genetic history of the Qiang and highlight how admixture and selection shaped the diversity of Tibetan-Plateau fringe populations.
High-coverage Qiang genome study reveals ancient Yellow River roots, Han-Tibetan admixture, and adaptive evolution.
INTRODUCTION
The Qiang people, also known as Rrmea or “the people living among the clouds,” primarily inhabit the rugged highland region bridging the Sichuan Basin to the Tibetan Plateau in Southwest China. Recognized as one of the oldest ethnic groups in East Asia, the Qiang people are believed to be direct descendants of the ancient Di-Qiang tribes (1), whose existence was documented more than 3000 years ago in oracle bone inscriptions. These tribes played a pivotal role in the early ethnogenesis of the Chinese nation (1, 2), notably influencing the cultural and demographic trajectory of Tibeto-Burman populations (1–3). Archaeological discoveries in the upper and middle Yellow River basin, as well as in Mao County—the largest contemporary settlement of the Qiang—have unearthed relics such as pottery and millet seeds (4). These artifacts suggest cultural exchange with the Yangshao Neolithic culture, underscoring a deep historical connection between the Di-Qiang and early northern Asian civilizations. Such findings highlight their importance in regional cultural evolution and cross-cultural interactions (4). While archaeologists and anthropologists have proposed that inhabitants of the southeastern Tibetan Plateau are the descendants of the Di-Qiang, the extent of genetic continuity remains debated. The term “Di-Qiang” itself broadly encompasses multiple ancient tribes, complicating lineage tracing. Moreover, centuries of migration and intermixing have obscured the origins of the modern Qiang people, transforming their ancestral narrative into a complex web of historical, cultural, and genetic interplay.
Previous genetic investigations have primarily focused on the paternal and maternal genetic architectures of the Qiang people (5–8). Notably, the pronounced prevalence of northern Asia maternal haplogroups (A, C, D, and G) suggests a northern origin for the population and corroborates archaeological evidence (8). Analysis of Y-chromosome single nucleotide polymorphism (Y-SNP) and Y-chromosome short tandem repeats (Y-STR) markers further reveal high haplotype diversity among the Qiang, strongly indicating that they represent an ancient population with a prolonged demographic history (5–7). Y haplotype analyses identified O-M175 and D-M174 as preponderant lineages (7, 8), followed by C, N, and Q, reflecting a complex admixture history involving interactions with Han Chinese and Tibetans. However, the limited resolution of Y chromosome and mitochondrial DNA (mtDNA) poses notable challenges in reconstructing such intricate admixture dynamics. One recent genome-wide study genotyped 20 Qiang individuals using the Human Origin array, revealing their closest genetic affinity to ancient populations from the Yellow River Basin (5). The authors proposed that the Qiang emerged from a two-way admixture between northern and southern Chinese ancestries, suggesting an initial northern origin followed by subsequent admixture with southern Chinese forebears. Nevertheless, microarray-based approaches are constrained in resolving fine-scale population structures and often fail to detect population-specific variants critical for studying minority populations (9, 10). Given the Qiang’s unique geographic position—transitioning between plains and plateaus—also makes them a pivotal group for investigating genomic adaptations, particularly those linked to high-altitude adaptation. Their genetic profile positions as a crucial intermediary for understanding adaptive traits in East Asia. Given these gaps, a comprehensive, high-resolution genomic study is urgently needed to elucidate their demographic history, admixture dynamics, and adaptive evolution.
In this study, we present high-coverage whole-genome sequencing (WGS) data (32×) from 20 individuals representing the Qiang ethnic group, sampled from Mao County—the core settlement area of this population. Through comprehensive comparative analyses with both contemporary and ancient reference populations, we elucidate the genetic architecture, ancestry composition, and adaptive mechanisms to high-altitude environments. Our findings provide in-depth insights into the demographic history of the Qiang people and highlight the crucial role of historical admixture events and selective pressures in shaping the genetic diversity landscape among populations inhabiting high mountainous regions.
RESULTS
Population structure and genetic makeup of the Qiang people
To investigate the genetic relationships between the Qiang people and global populations, a sequence of principal components analyses (PCAs) was executed leveraging data panel B (see the “Sample collection and data processing” section). The Qiang consistently demonstrated strongest genetic affinity with East Asian populations across all analytical frameworks, showing marked differentiation from European, Central Asian, and South Asian groups (fig. S1), confirming their East Asian ancestry. Within East Asian contexts, PCA revealed bimodal distribution of Qiang samples along PC1, forming two distinct clusters: one aligned closely with Han Chinese and another associated Tibetans (Fig. 1A), indicating substructure within the population. Subsequent Han-Tibetan-Qiang–focused PCA (Fig. 1B) confirmed this dual clustering pattern. To systematically characterize this substructure, we conducted unsupervised hierarchical clustering using the first two principal components from Fig. 1B. This divided the 20 Qiang samples equally into two genetically distinct clusters (Fig. 1C). Spatial remapping demonstrated that one cluster is closely aligned with Han Chinese, and the other cluster shows a stronger affinity with Tibetans (Fig. 1D). Accordingly, we designated them as Qiang_H and Qiang_T, respectively. We constructed a maximum likelihood tree (Fig. 1E and text S1) to further investigate the phylogenetic relationships between the Qiang subgroups and representative East Asian populations. The phylogenetic topology showed that both Qiang_H and Qiang_T clustered within the Tibeto-Burman linguistic branch (Fig. 1E), consistent with their ethnolinguistic classification. Notably, Qiang_T exhibited tight clustering with Tibetans, while Qiang_H showed closer proximity to Han Chinese, corroborating PCA results.
Fig. 1. Population structure of the Qiang people.
(A) PCA of Qiang samples and East Asian populations. Tibeto-Burman populations and Han are colored on the plot. Numbers in parentheses indicates the variance explained by each PC when considering the top 10 PCs. (B) PCA under the context of Han-Tibetan-Qiang. (C) Unsupervised hierarchical clustering of Qiang samples using the genetic coordinates of the first two PCs from (B), sample IDs are shown in the plot. (D) Recolored Qiang samples after classification by unsupervised hierarchical clustering in the PCA, using the genetic coordinates of the first two PCs from (B). (E) Maximum likelihood phylogenetic tree of Qiang_H, Qiang_T, and other representative populations, the brown clad represents Tibeto-Burman populations, mainly located in north/south west China, the blue clad represents north east Asian populations, the green clad represents south east Asian populations. Numbers on the branch indicates support values from 100 times bootstrapping.
To characterize the ancestral composition of Qiang subgroups, we used ADMIXTURE analysis, presuming the number (K) of ancestral populations fluctuated from 2 to 20. When K was set at 2, the ancestral components of East Asia and West Eurasia were separated (fig. S2). In line with the PCA outcomes (fig. S1), the Qiang people had nearly identical ancestral makeups as those of the East Asian populations (fig. S2). As K augmented, a more elaborate population structure came to the fore. Starting from K = 5, when distinct southern East Asian ancestral components (predominantly present in Chinese Taiwan aboriginal Ami and Atayal) and northern Tibeto-Burman ancestral components (widely distributed among Tibeto-Burman populations and maximally represented in the typical Tibeto-Burman groups Drung and Nu) began to diverge, it was noted that Qiang_H harbored a greater proportion of southern East Asian ancestral components compared to Qiang_T. Conversely, in Qiang_T, there were more northern Tibeto-Burman ancestral components than in Qiang_H (fig. S2). This pattern persisted as K continued to increase (fig. S2). We attained the minimum cross-validation (CV) error when assuming 12 ancestral components (K = 12) (fig. S3). At this juncture, the principal ancestral makeup of Qiang_T was the deep blue-colored northern ancestry (69.33%), which constituted the majority ancestry and was most pronounced in Tibetan (69.61%) (Fig. 2). In Qiang_H, this ancestry accounted for 51%, and in the Han Chinese population, it was 42.51%. Moreover, the southern East Asian ancestry (with the highest prevalence in Atayal) was measured at 33.21% in Qiang_H, 48.91% in Han, and meager 3.84% in Qiang_T, being virtually nonexistent in Tibetan groups. In addition, a notable fraction of Drung/Lahu-specific ancestral components was detected in Qiang_H (3.99%/10.21%), Qiang_T (7.55%/16.89%), and the Tibetans (8.41%/15.45%), while these components were relatively scarce in the Han Chinese (total, <5%) (Fig. 2). Overall, we discerned that Qiang_T consistently exhibited a higher degree of shared ancestral components with the Tibetans, a tendency that became more conspicuous as K grew larger (fig. S4). On the contrary, Qiang_H demonstrated a more proximate ancestry makeup to the Han Chinese. These findings were congruous with the PCA and phylogenetic analysis.
Fig. 2. ADMIXTURE inference of the ancestral makeup of Qiang_H, Qiang_T and other Eurasian populations.
The results are based on the ADMIXTURE analysis assuming 12 ancestral populations (K = 12). The inner pie chart shows the detailed proportion of ancestral components for Qiang_H, Qiang_T, Han, and Tibetan, respectively.
Briefly, our analysis revealed previously unidentified insights into Qiang population substructure, identifying distinct genetic architectures between subgroups that suggest divergent demographic histories. Qiang_H shows southern East Asian/Han-influenced ancestry, while Qiang_T maintains stronger northern Tibeto-Burman/Tibetan connections, indicating complex admixture processes shaping this highland population.
Admixture-driven subpopulation divergence
Genetic differentiation measured by FST revealed distinct divergence patterns between Qiang subgroups (Fig. 3, A and B, and table S1). Precisely, Qiang_H showed minimal genetic distance from northern Han Chinese (Han_N, FST = 0.0015 ± 0.0004), comparable to north-south Han differentiation (FST = 0.001) (11). Qiang_H demonstrated significantly lower divergence (Wilcoxon signed-rank test, P < 2.2 × 10−16) from Han than from Qiang_T (Qiang_H-Qiang_T, FST = 0.0046) (Fig. 3, B and C), suggesting that Qiang_H might have experienced a substantial recent gene flow from Han after the divergence of the Qiang subpopulations. In addition, the FST between Qiang_T and Tibetan (FST = 0.0051 ± 0.0009) was significantly lower (Wilcoxon signed-rank test, P < 2.2 × 10−16) than that between Qiang_H and Tibetan (FST = 0.0080 ± 0.0008) (Fig. 3C), indicating more genetic connection between Qiang_T and Tibetan after the divergence of the Qiang subpopulations. The FST between Qiang_H and Qiang_T was only ranked 17th in the pairwise FST calculated for Qiang_H (table S1). However, when taking into account the FST in Qiang_T, the genetic difference between these two subgroups was found to be the smallest (Fig. 3B), indicating that Qiang_T might be a more isolated population compared to Qiang_H. The isolation of Qiang_T was further corroborated by its less genetic sharing with surrounding populations (measured by D-statistics and f3-statistics) (figs. S5 and S6), a slower rate of linkage disequilibrium (LD) decay (fig. S7A) (12), lower nuclear diversity (fig. S7B) (13), and longer runs of homozygosity (fig. S7C) (text S2) (14).
Fig. 3. Genetic differentiation between Qiang_H/Qiang_T and East Asian populations.
(A and B) Genetic differences (FST) between Qiang_H (A) or Qiang_T (B) and other East Asian populations. Each branch represents a comparison between Qiang_H/Qiang_T and one tested population, with branch lengths sorted and proportional to the FST values. The accompanying table shows the five smallest pairwise FST values. (C) Pairwise FST values between populations were estimated using 100 times bootstrapping with 10 randomly selected samples per run. Each point represents one bootstrap iteration.
The genetic differentiation between the Qiang_H and Qiang_T was even substantially greater than that witnessed between northern and southern Han, which implies the existence of distinct evolutionary histories within the Qiang population. To explore the driving force behind the formation of subpopulations of Qiang_H and Qiang_T, we adopted an FST-based statistic, namely, the relative difference (RD) (15), to analyze and interpret the comprehensive situation regarding the relative genetic connections of Qiang_H and Qiang_T with surrounding populations. We observed that Qiang_H exhibited closer genetic connections with nearly all populations across East Asia, spanning from northeastern China (e.g., Man) to northwestern China (e.g., Xibo) and southeastern China (e.g., Zhuang) (Fig. 4A and table S2). Geographically speaking, this pattern was unlikely to result from a direct gene flow from these populations into Qiang_H. Instead, it was more likely be attributed by the remarkable gene flow from the widely distributed and populous Han Chinese. To verify this hypothesis, we substituted Qiang_H with Han in the RD analysis. We obtained a high correlation (R = 0.94, P < 2.2 × 10−16) between RD (Qiang_H, Qiang_T) and RD (Han, Qiang_T) (Fig. 4B), which indicates a robust genetic connection between Han and Qiang_H. However, when we replaced Qiang_H with southern populations like the Zhuang, the correlation decreased substantially (R = 0.28, P = 0.035) (fig. S8), suggesting that the additional genetic affinity (Fig. 3C) of Qiang_H with Southeast Asian populations was more likely due to the gene flow from Han. The gene flow from Han to Qiang_H was also corroborated by Treemix analysis (fig. S9). As a relatively more isolated population, only a few populations exhibited a close genetic connection with Qiang_T (Fig. 4A), such as Tibetan and Pumi (Fig. 4A and table S2). Subsequently, we replaced Tibetan with Qiang_T and obtained a notably high correlation (R = 0.99, P < 2.2 × 10−16) between RD (Qiang_H and Qiang_T) and RD (Qiang_H and Tibetan) (Fig. 4C), which implies the recent gene flow from Tibetan to Qiang_T. Although we did not detect a direct gene flow from Tibetan to Qiang_T in the Treemix results, Qiang_T consistently clustered together with Tibetan, indicating their close genetic relationship (fig. S9). To further investigate the genetic connectivity patterns among Han, Tibetan, and Qiang subpopulations, we used two complementary approaches: identity-by-descent (IBD) sharing analysis using hap-IBD (16) and genome-wide local ancestry inference conducted with HAPMIX (17). Our IBD analysis revealed distinct sharing patterns: Qiang_H exhibited significantly greater IBD segment sharing with Han populations compared to Qiang_T (Wilcoxon rank sum test, P = 4.33 × 10−5), while Qiang_T demonstrated stronger IBD sharing affinity with Tibetan populations than Qiang_H (Wilcoxon rank sum test, P = 1.5 × 10−3), demonstrating asymmetric bidirectional gene flow from the ancestral Han-Tibetan populations into Qiang subgroups (fig. S10). Complementing these findings, the local ancestry segment distribution (fig. S11) showed differential admixture patterns: Qiang_H carried substantially longer Han-derived ancestry segments compared to Tibetan-derived segments, whereas Qiang_T displayed the reverse pattern with predominant Tibetan ancestry segments exceeding Han contributions. These contrasting ancestry architectures highlight differential temporal dynamics and intensity of recent gene flow events from source populations into the Qiang subpopulation.
Fig. 4. Gene flow from surrounding populations drives the subpopulation formation within the Qiang population.
(A) RD between tested populations and target populations, where Qiang_H and Qiang_T as the target populations and other East Asian populations as the tested populations. Red points indicate a stronger genetic connection between the tested population and Qiang_H compared to Qiang_T, while blue points indicate a stronger connection to Qiang_T relative to Qiang_H. (B and C) Correlation between RD values: RD (Qiang_H, Qiang_T) and RD (Han, Qiang_T) (B) or RD (Qiang_H, Tibetan) (C).
In summary, our integrative analysis demonstrated that asymmetric recent admixture—substantial Han introgression into Qiang_H versus predominant Tibetan influence in Qiang_T—underlies the emergence of distinct Qiang subpopulations. This differential gene flow pattern suggests separate demographic trajectories shaped by regional population interactions.
Genetic origin and demographic history of the Qiang people
To investigate the genetic origins of the Qiang people, we incorporated 463 ancient specimens spanning 40 ka (thousand years ago) to 400 B.P. (before the present), with chronological concentration between 7 and 2 ka (table S3). The contemporary Qiang population occupies the geographical epicenter of these ancient sampling locations (fig. S12), making them particularly suitable as ancestral proxies for reconstructing Qiang population history. PCA with ancient comparators positioned the Qiang population at the intersection of major genetic clines (fig. S13). Specifically, Qiang individuals showed closer genetic affinities to (i) high-altitude Neolithic groups from the Qinghai-Tibet Plateau (Zongri and Chamdo) and (ii) mid-Holocene agricultural populations from the Yellow River Basin (China_Upper_YR_LN and China_YR_MN). Notably, systematic differentiation emerged between Qiang subgroups: (i) Qiang_H clustered with ancient Southeast Asian coastal populations; (ii) Qiang_T aligned with ancient Tibetan Plateau groups. These differential clustering patterns mirror modern reference population relationships. Modern Tibetans exhibited substantial overlap with ancient Tibetan specimens (Shigatse and Chamdo), demonstrating strong genetic continuity over three millennia (fig. S13). Conversely, modern Han populations displayed dual genetic signatures: They show both genetic affinity with Yellow River Neolithic farmers (China_YR_LN and China_YR_LBIA) and southern Neolithic/Metal Age groups (Gongguan and Hanben_IA) (fig. S13). This genetic architecture supports the established model of Holocene admixture between northern millet cultivators and southern rice agriculturalists that shaped the modern East Asian diversity (18). Next, we used outgroup f3-statistics to quantify the genetic continuity between these ancient samples and modern populations (Fig. 5, A and B, and fig. S14). To minimize potential analytical bias, we conducted a series of outgroup f3-statistics analyses using geographically representative modern populations across East Asia. Our analysis revealed a pronounced isolation-by-distance pattern, where modern groups showed strongest genetic affinity with geographically proximal ancient populations: (i) Tibetans displayed maximal f3 values with Qinghai-Tibet Plateau specimens (f3 = 0.309 ± 0.005); (ii) Dai populations clustered preferentially with Southeast China Neolithic samples (f3 = 0.312 ± 0.009); (iii) Yakut individuals exhibited peak affinity to ancient Northeast Asian groups (f3 = 0.298 ± 0.019). These spatial genetic patterns corroborated PCA clustering observed in fig. S13. Notably, both Qiang subgroups (Qiang_H/Qiang_T) showed significantly higher shared genetic drift with middle-upper Yellow River Basin populations than with other regional groups (Welch’s t test, P < 2.2 × 10−16), strongly supporting their ancestral origins in this region—consistent with historical records of the settlement of the ancient Di-Qiang tribes (Fig. 5, A and B, and tables S4 and S5). Critical evidence emerged from six ancient sites from the Yellow River Basin exceeding intra-Qiang genetic sharing: (i) China_Miaozigou_MN (5.2 ka); (ii) China_Shimao_LN (4.1 to 4.6 ka); (iii) China_Upper_YR_LN (3.9 to 4.5 ka); (iv) China_YR_LN (3.9 to 4.5 ka); (v) China_Upper_YR_LBIA (2.2 to 3.3 ka); (vi) China_Upper_YR_IA (1.8 to 1.9 ka). The Miaozigou_MN specimen demonstrated exceptional affinity, achieving peak f3 values with both Qiang subgroups [f3(Qiang_H, X; Mbuti) = 0.3195 ± 0.0031 and f3(Qiang_T, X; Mbuti) = 0.3196 ± 0.0031] (tables S4 and S5). D-statistics confirmed near-identical allele sharing between Miaozigou_MN and Qiang subgroups [D (Mbuti, Miaozigou_MN; Qiang_H, Qiang_T) = 0.0001; Z = 0.064] (Fig. 5C), establishing this ~5300-year-old population as the most recent common ancestral source before Qiang population divergence into the present-day Qiang_H and Qiang_T.
Fig. 5. Genetic origin of the Qiang people.
(A and B) Out group f3-statistics were applied to estimate the shared genetic drift between ancient individuals and the target populations Qiang_H (A) or Qiang_T (B). The text-annotated dots represent the top 10 samples that shared the most genetic drift with the target population. Deep red text annotations highlight samples that share even more genetic drift with the target population than observed between Qiang_H and Qiang_T. The shaded area indicates the most likely originating region of the Qiang population. (C) D-statistics, D (Mbuti, X; Qiang_H, Qiang_T), were used for testing excess allele sharing of ancient individuals with Qiang_H versus Qiang_T, and |Z| ≥ 3 was used for threshold. Blue points represents individuals sharing excess alleles with Qiang_T relative to Qiang_H, whereas the red point indicates individuals sharing excess alleles with Qiang_H relative to Qiang_T. Deep red points correspond to the deep red–annotated samples from (A) and (B). (D) A possible demographic model of the Qiang population. Ancestry proportions were estimated using MultiWaveX and rescaled after incorporating ancient and private ancestry into the model.
To elucidate the population history of subgroups, we implemented a multimethod approach combining MSMC2 and MSMC-IM for coalescent-based analyses to estimate their historical effective population size (Ne) as well as the divergent time between populations (fig. S15). Qiang_T exhibited notable reduced Ne [median = 20,300; 95% highest posterior density (HPD): 13,100 to 44,100] compared to Qiang_H (median = 58,500; 95% HPD: 39,800 to 175,600), consistent with its isolation characteristics (fig. S15A). The estimated divergent time between Han and Tibetan was approximately 9.7 ka (95% confidence interval: 6.8 to 10.4 ka), aligning with previous estimates (table S6) (19). We estimated that the divergent times between Qiang_H and Han and between Qiang_H and Tibetan were 4.8 ka (4.1 to 5.7 ka) and 7.7 ka (5.9 to 8.4 ka), respectively. For Qiang_T, the divergent with Han and Tibetan was 8.4 ka (7.3 to 8.7 ka) and 3.8 ka (3.7 to 4.0 ka), respectively. The divergent time between Qiang_H and Qiang_T was estimated to be 4.1 ka (3.3 to 4.3 ka) (table S6). Accounting for asymmetric gene flow from Han and Tibetan into Qiang_H and Qiang_T, we reconstructed ancestral Qiang (aQiang) divergence windows: aQiang-Han: 4.8 to 8.4 ka and aQiang-Tibetan: 3.8 to 7.7 ka. The observed Qiang_H/Qiang_T divergence (4.1 ka) likely represents a maximum estimation due to differential introgression effects from the more divergent Han-Tibetan pair (FST = 0.0124 ± 0.0011; Fig. 3C).
We further investigated the demographic history of the Qiang population by using MultiWaverX (20) to model their admixture trajectories with Han and Tibetan groups. Both Qiang_H and Qiang_T fit well into a biphasic 2-2 admixture model (R2 ≥ 0.98; fig. S16, A, C, and D). Given the genetic links between the Qiang and ancient Yellow River populations, MultiWaverX might miss older genetic signals predating the defined sources, as the local ancestry inference process assigns all chromosomal segments to present-day reference groups. To address this, we quantified the ancient and private genetic component in modern Qiang subgroups (text S3), identifying ~1.47% in Qiang_H and ~1.74% in Qiang_T. Consistent with historical records, we interpret these components as ancient Di-Qiang–specific ancestry, potentially representing an early foundational layer of the Chinese gene pool. We then refined the initial model (fig. S16A) by incorporating these components from the first admixture event. The updated model (Fig. 5D and fig. S16B) revealed that the primary admixture occurred approximately 4.7 ka (4.2 to 5.0 ka), with contributions from ancient Han (36.3 to 56.1%), ancient Tibetans (38.4 to 56.6%), and ancient Di-Qiang–specific ancestries (5.5 to 7.1%), leading to the formation of the aQiang (Fig. 5D and fig. S16). Subsequent admixture events followed distinct patterns: one around 1.4 and 0.7 ka, with Tibetans (54.2 ± 1.6%) and Han Chinese (41.7 ± 1.0%) contributions, respectively, resulting in the emergence of Qiang_H (Fig. 5D); the another around 1.7 ka and 0.8 ka, with Han Chinese (54.8 ± 2.1%) and Tibetan (45.8 ± 0.8%) contributions, leading to the formation of Qiang_T (Fig. 5D). These trajectories explain (i) differential genetic affinity patterns in PCA, phylogenetic topology, FST, and RD analysis (Figs. 1, 3, and 4); (ii) residual ancestry proportions (Qiang_H: 56.7% Han derived; Qiang_T: 59.7% Tibetan derived) (fig. S16B); (iii) temporal concordance with historical migrations: Wei and Jin Dynasties Han migrations (3rd to 5th centuries CE); Tubo Empire expansions (7th to 9th centuries CE); Han-Tibetan interactions since the Yuan Dynasty (13th centuries CE); and (iv) the foundational layers of the Chinese gene pool.
Genetic differentiation between Qiang subpopulations
Our analyses revealed substantial genetic differentiation within the Qiang population. To characterize this divergence, we implemented locus-specific FST analysis (top 0.1% threshold), identifying 6664 highly differentiated variants between Qiang_H and Qiang_T (fig. S17A and table S7). Notably, 59% of these variants functioned as expression quantitative trait loci (eQTLs), which significantly exceeding genome-wide eQTL density (36%; χ2 test, P < 2.2 × 10−16). ANNOVAR (21) mapped these variants to 1243 genes showing significant enrichment in neuronal pathways: synaptic membrane composition [false discovery rate (FDR), P = 1 × 10−10], neuron recognition (FDR, P = 3.9 × 10−6), and axon guidance (FDR, P = 1 × 10−5) (fig. S17B and table S8). Moreover, 56 highly differentiated variants overlapped genome-wide association study catalog entries (table S9), including rs4976033-G, a type 2 diabetes risk allele (22) showing markedly lower frequency in Qiang_T (0.25) versus Qiang_H (0.8). Given the asymmetric bidirectional gene flow from Han and Tibetan, the Qiang represent an invaluable model for exploring the evolutionary trajectories of Han-Tibetan–differentiated genes. Accordingly, we focused on the well-documented genes associated with traits: high-altitude adaptation (23–26), skin pigmentation (8, 27, 28), and ethanol metabolism genes (29, 30). While no differentiation emerged in hypoxia-related genes (EPAS1/EGLN1/TMEM247), we identified 87 highly differentiated variants associated with reported pigmentation candidate genes (table S10) (31), including MED1 (32), GPR161 (33), and BMPR1B (table S10) (34). Notably, two variants, rs2114533 and rs6815969, which are eQTLs for BMPR1B with the highest expression in Nerve-Tibial tissue (fig. S18B), might contribute to skin color pigmentation. These derived alleles (frequency: Qiang_H = 0.6 versus Qiang_T = 0.1) up-regulate BMPR1B, which inhibits melanogenesis (34). Ancestry decomposition via RFMix (35) showed Tibetan ancestry predominance at these loci in Qiang_T versus Han ancestry in Qiang_H (fig. S18C), indicating admixture-driven differentiation. Population branch statistics (PBS) analysis revealed positive selection signals for pigmentation genes GPR161/TFAP2A/MED1 in Qiang_T (figs. S18 and S19).
We identified five highly differentiated variants associated with ethanol metabolism genes: four are eQTLs of ADH1C (rs1442493/rs7692081/rs1442492/rs17028973) (fig. S20A), and one is linked to ALDH1A3 (rs3809523). The derived alleles of the four ADH1C eQTLs, which up-regulated ADH1C expression, were found at higher frequencies in Qiang_H (0.8 to 0.85) compared to Qiang_T (0.25) (fig. S20, A and B). This suggests a potentially enhanced capacity for ethanol-to-acetaldehyde conversion in Qiang_H relative to Qiang_T (36). Notably, three of the four eQTLs exhibited extended haplotype homozygosity [normalized |integrated haplotype score (iHS)| ≥ 2] in Qiang_H but not in Qiang_T or the source populations, Han and Tibetan (fig. S20C). To further validate this signal, we performed cross-population extended haplotype homozygosity (XP-EHH) analysis using Qiang_T as the control, which confirmed recent positive selection at these loci in Qiang_H (fig. S20C and table S11). We propose that the differentiation of these variants may have been driven by postadmixture selection specific to Qiang_H.
In addition, a key variant (rs1229984) in the ethanol metabolism gene ADH1B (rs1229984-T, ADH1B*48His, alcoholism-protective variant) (37) also showed differential frequencies between Qiang subpopulations and their source populations (Qiang_H = 0.65, Han = 0.76 versus Qiang_T = 0.2, Tibetan = 0.17). Haplotype network analysis revealed that Qiang_H and Han predominantly carry the H7 haplotype (rs1229984-T), whereas Qiang_T and Tibetan mainly retain non-H7 haplotypes (fig. S21A) (30, 38). The rs1229984-T has been previously identified as under positive selection in southeastern Chinese populations but not in Tibetans (29), consistent with our observations of EHH patterns in Han (normalized |iHS| = 2.46) and Tibetan (normalized |iHS| = 0.70) (fig. S21B). A similar pattern was observed between Qiang_H (normalized |iHS| = 2.67) and Qiang_T (normalized |iHS| = 0.07) (fig. S21B). Furthermore, ADH1B haplotype diversity (Han = 0.904, Qiang_H = 0.932, Qiang_T = 0.974, Tibetan = 0.979) inversely correlated with Han ancestry proportion, suggesting that the strong EHH signal in Qiang_H may have been introduced through recent admixture with Han. Supporting this, PBS analysis indicated that rs1229984 in Qiang_H likely reflects admixture (negative PBS bottom 0.1%) from source populations (fig. S19A). In conjunction with the demographic model (Fig. 5D), we propose that the differentiation of rs1229984-T between Qiang subpopulation was driven by admixture.
Local adaptation
The Qiang population inhabits high mountain regions, a geographic setting that distinguishes them from both the lowland Han Chinese and highland Tibetan populations. To investigate potential local adaptation signals, we analyzed variants ranked in the top 0.1% from PBS tests, using either Qiang_H or Qiang_T as the target population with Han Chinese samples in the Human Genome Diversity Project (HGDP-Han) and Tibetan as reference populations (see the “Natural selection” section) (tables S12 and S13). We identified 95 overlapping putative variants under selection in the PBS tests (fig. S22 and table S14), which likely reflect common selective pressures acting on the Qiang following their divergence from Han and Tibetan. Notably, 81% (77 of 95) of these variants clustered within three genomic regions, 14 variants within a 17-kb segment on chromosome 8 (Chr8:18097747-18115076), 18 variants in a 13-kb region of chromosome 5 (Chr5:67199578-67213032), and 45 variants in a 740-kb region on chromosome 2 (Chr2:96910922-97655343) (table S14). The 14 variants on chromosome 8 showed associations with ASAH1-AS1 and NAT1. The derived allele of these variants exhibited substantially higher frequencies in the Qiang (0.325) compared to Han (0.045) and Tibetan (0.045) groups, as well as most global populations (fig. S23). Elevated frequencies were also observed in Oceania populations including the NAN_Melanesian (0.25) and Papuan (0.46 to 0.5), although the derived allele rs144282517 was absent in these groups (fig. S23). NAT1 encodes a key enzyme involved in xenobiotic and drug metabolism, playing crucial roles in detoxification processes and chemical carcinogen metabolism (39). ASAH1-AS1, an antisense long noncoding RNA gene, regulates ASAH1—an acid ceramidase participating in sphingolipid metabolism and the innate immune responses (40). We hypothesize that these noncoding variants may modulate NAT1 and ASAH1 expression through transcriptional regulatory mechanisms, although functional validation remains necessary. Among the 18 variants on chromosome 5, 12 were identified as eQTLs for CD180, with all alternative alleles associated with increased CD180 expression levels (fig. S24). CD180 encodes a cell surface protein predominantly expressed on antigen-presenting cells, where it collaborates with MD-1 and TLR4 to mediate the innate immune response to bacterial lipopolysaccharide in B cells (41). These eQTLs demonstrated markedly higher alternative allele frequencies in Qiang populations (0.275 to 0.325) compared to Han and Tibetan groups (0 to 0.045) (fig. S25). Notably, elevated frequencies were also observed in African populations (Yoruba and Bantu >0.5), Northeast Asian Hezhen (~0.3), and South American Karitiana (~0.4), contrasting with minimal frequencies in European and other Asian populations (figs. S25 and S26).
The chromosome 2 region contained nearly half (47%, 45 of 95) of candidate variants, including the sole missense variant rs143216880 (ANKRD36: c.3065C>T, p.Ser1022Phe; CADD_phred = 22.5) (table S11). The rs143216880-T allele showed low frequency in Han (0.045) and complete absence in Tibetans but substantial prevalence in Qiang populations (Qiang_H: 0.3; Qiang_T: 0.35) (Fig. 6A and table S15). Global analysis revealed rs143216880-T’s extreme rarity worldwide—absent in African, European, and American populations, with only minor representation in Central/East Asian groups (Fig. 6A and table S15). The Gin population in Guangxi province showed the next highest frequency (0.139) after Qiang groups (table S15). Eleven of the remaining 44 variants functioned as ANKRD36 eQTLs, displaying strong LD with rs143216880 (r2 ≥ 0.9) (Fig. 6B). All of the alternative alleles at these eQTLs down-regulate ANKRD36 expression (fig. S27) and showed Qiang-specific prevalence (0.3 to 0.35) compared to other East Asian populations (Fig. 6C). ANKRD36 involvement in blood pressure regulation is supported by its reduced expression in essential hypertension patients and elevated blood pressure in Ankrd36 knockout mice (42). The high-frequency eQTLs together with the missense variant may collectively reduce ANKRD36 expression and alter protein function. This pattern could plausibly reflect an adaptive response to chronic high-altitude hypoxia, where elevated blood pressure may enhance organ perfusion and oxygen delivery (43–45).
Fig. 6. Adaptation signals in ANKRD36.
(A) Allele frequency of rs143216880-T in worldwide populations. (B) PBS signal in ANKRD36; annotated points indicate the positions of the 11 eQTLs and the missense rs143216880. (C) Allele frequency of the 11 eQTLs and the missense rs143216880. (D) Haplotype network was constructed by PopART using 249 SNPs located near ANKRD36. Because of the large sample size of Han (HGDP+AAGC) and Yao, not all haplotypes of these two populations are plotted. For Han, 19 individuals carrying rs143216880-T and 17 randomly selected individuals without rs143216880-T are displayed, totaling 72 haplotypes. For Yao, 14 individuals carrying rs143216880-T and 20 randomly selected individuals without rs143216880-T are displayed, totaling 68 haplotypes.
Haplotype-based analyses of rs143216880-T revealed contrasting patterns: Qiang populations lacked extended haplotype homozygosity, while low-frequency groups like Han and Yao exhibited long-range homozygosity patterns (fig. S28). The rapid decay of EHH in the Qiang suggests ancestral haplotype preservation. Using Genealogical Estimation of Variant Age (GEVA) (46), we estimated rs143216880-T allele ages of ~1582 generations (~46 ka, 29-year generation time) in Qiang, closely matching KGP3 global estimates (1691 generations, ~49 ka) (46), whereas Han and Yao showed younger allele ages of ~803 and 701 generations (~20 to 23 ka), respectively. Median-joining network analysis of a 216-kb region (Chr2:96,990,429-97,205,943; 249 SNPs) encompassing all 11 eQTLs and rs143216880 revealed distinct clustering: rs143216880-T haplotypes formed a monophyletic branch, with 5 of 13 Qiang haplotypes positioned near ancestral nodes versus terminal placement of Han/Yao haplotypes (Fig. 6D), This phylogenetic pattern further supports the ancestral status of Qiang haplotypes compared to Han and Yao.
DISCUSSION
This study represents a high-coverage whole-genome analysis of the Qiang people, offering critical insights into their genomic history and their central role in the ethnogenesis of Tibeto-Burman groups in Southwest China (1, 8, 47). Our findings not only resolve long-standing debates about Qiang origins but also establish a foundational genomic references for regional population studies.
Ancient origins and neolithic migrations
Archaeological evidence suggests that hydrological decline and population expansion during the Neolithic period (6000 to 5000 years ago) triggered migrations from the middle and upper Yellow River into northwestern Sichuan (48). These migrants moved from the cold continent into the warmer Sichuan Basin via natural waterway corridors (known as the Tibetan-Yi corridor) along the southeastern margin of the Tibetan Plateau (48). Mao County, located at the entrance of this corridor, likely served as a focal point for these migrations. We propose that the Qiang people originated from this movement, as the timing broadly aligns with our admixture model for the initial formation of the Qiang ancestral population (~4700 B.P.) (Fig. 5D and fig. S16B) and coincides with the dating of the Yingpangshan archaeological site (5300 to 4600 B.P.).
Ancient DNA analyses further reveal that a ~5300 B.P. Yangshao culture–associated population (China_Miaozigou_MN) shares nearly identical genetic drift with modern Qiang subgroups (Fig. 5, A to C). This population may represent a preadmixture Han-like ancestral migrant group, given its closer genetic affinity to present-day Han Chinese (fig. S14B). Our two-way admixture model also suggests the contribution of a Tibetan-like ancestral migrant group, although this component remains unidentified in current ancient DNA databases. This gap underscores the need for broader archaeological sampling and genomic sequencing to recover this missing ancestral lineage.
However, the limitations of ancient DNA data must be acknowledged. DNA degradation can bias statistical estimates, and the geographical concentration of available ancient DNA in the Yellow River basin may overlook genetic connections with other historically relevant regions, such as Southwest China. In addition, the absence of ancient genomes from Yingpangshan—the site most directly linked to the Qiang—prevents a direct test of local genetic continuity. Addressing these limitations through expanded excavations and sequencing will be essential to refine our understanding of the Qiang migration history.
Deep ancestry
Beyond Neolithic migrations, we identified several ancient genomic components in the Qiang, including a putative adaptive haplotype within the gene ANKRD36. The key mutation rs143216880-T in the Qiang, estimated to have arisen ~46,000 years ago, coincides with the initial peopling of East Asia (49). This mutation, which is exclusive to East Asians and absent in African and European populations, may have persisted because of potential fitness benefits in high-altitude environments. We hypothesize that this haplotype was introduced into the Qiang through early admixture waves or derived from an unmodeled “ghost” ancestral lineage. More broadly, our haplotype analysis revealed the persistence of an Ancient Di-Qiang–specific ancestry in both Qiang_H and Qiang_T. Rather than being erased by later admixture, this ancestral component appears to represent one of the deepest layers of the Chinese gene pool, underscoring the importance of the Qiang as a genetic reservoir for reconstructing East Asian demographic history.
Genetic replacement and population structure
Population structure analyses revealed that the Qiang population is divided into two distinct subgroups. This pattern remains evident even after integrating previously published microarray data (fig. S29), indicating a core genetic architecture of the Qiang people. This division also aligns with earlier anthropological observations (50). Signals from RD (Fig. 4), IBD sharing (fig. S10), and local ancestry distribution (fig. S11) collectively suggest that this differentiation was shaped by recent asymmetric gene flow from Han and Tibetan populations. Haplotype-sharing analyses (51) further show that the Qiang harbor fewer private haplotypes, confirming them as recipients rather than sources (figs. S31 to S33 and text S4). Notably, Qiang_H consistently shares more haplotypes with Han (fig. S31), whereas Qiang_T shares more with Tibetans (fig. S32), highlighting differential admixture histories as a major driver of their genetic divergence. Contrary to earlier hypotheses based on Y chromosome and mtDNA data, our genome-wide analyses suggest that the Qiang genetic makeup has been substantially influenced by Han and Tibetan ancestry (Han derived: 38.59 to 56.68%; Tibetan-derived: 41.85 to 59.67%).
Historical evidence provides a compelling backdrop to these genetic patterns. Multiple migration and conquest events likely contributed to the genetic architecture observed in the Qiang. During the Wei and Jin Dynasties (220 to 420 CE), successive waves of Han Chinese migrated southward to escape political turmoil and warfare, promoting the mixture of Han and Qiang tribes (1). In the seventh to ninth centuries, the Tubo Empire expanded aggressively into the eastern fringe of the Plateau, incorporating local tribes into its political and military system, consolidating their control for over a century (47). The Yuan Dynasty (13th to 14th centuries) further reshaped the landscape by actively promoting Tibetan Buddhism across the Empire (52) and the following Ming Dynasty (beginning in the 14th century), large migrations of Han Chinese into Sichuan, thereby intensifying demographic and cultural exchange between Han, Tibetan, and local Qiang peoples (47). These events collectively transformed the genetic makeup of the corridor, embedding successive layers of Han- and Tibetan-related ancestry within the Qiang and driving their gradual Tibetanisation and Sinicization.
Genetic differentiation and adaptive signatures
The Qiang serve as an informative model for studying how admixture and selection shape human genetic diversity. We focused on functional loci associated with high-altitude adaptation, pigmentation, and alcohol metabolism. While genes involved in high-altitude adaptation, such as EPAS1 and EGLN1, remain relatively homogeneous across Qiang subgroups, we observed substantial differentiation in loci related to pigmentation and alcohol metabolism genes. Notably, Qiang_H may have enhanced alcohol metabolic capacity, while Qiang_T may exhibit darker pigmentation, with both admixture and selection jointly contributing these differences.
High-altitude adaptation has been a focal point of genomic research over past decades, yet populations inhabiting transitional altitudes, like the Qiang, have been largely overlooked. We identified a candidate adaptive haplotype in ANKRD36, which may contribute to physiological responses to chronic hypoxia at moderate altitudes. Nevertheless, these findings remain preliminary, and further research—including genotype-phenotype association studies and experimental validation—is needed to establish a direct role of this haplotype in altitude adaptation.
Together, our genomic reconstruction reveals that the Qiang people are shaped by Neolithic Yellow River ancestry and postagricultural admixture dynamics. However, the extensive genetic replacement observed underscores critical gaps in current reference frameworks. To recover the obscured heritage of historically pivotal groups like the Qiang, we advocate for (i) broader sampling and clinical investigation to improve genetic and medical equality of underrepresented ethnic groups; (ii) population-specific reference genome (11) to capture unrepresented variation; and (iii) pangenome references (53) to enable comprehensive ancestral haplotype reconstruction. These advances will be crucial for disentangling complex admixture histories across Southwest China’s ethnic tapestry and advancing our understanding of human genetic diversity in Asia.
MATERIALS AND METHODS
Sample collection and data processing
Peripheral blood samples were collected from 20 unrelated Qiang individuals from Mao County in Sichuan Province, China, the largest settlement of the Qiang people. Each individual was the offspring of the same ethnicity at least three generations. The family history and self-reported ethnicity were recorded through a questionnaire or an interview conducted in the local dialect. Written informed consent was obtained from all participants after providing them with detailed information about the study objectives and procedures, and consent discussions were conducted in the local dialect to ensure full understanding. All procedures were carried out in accordance with the ethical standards of the 1964 Helsinki Declaration, its later amendments (2000), or comparable ethical standards. This study was approved by the Biomedical Research Ethics Committee of Fudan University (FE23277I, FE21086, and FE23265I).
High coverage (32×) WGS was carried out on an Illumina Hiseq X10 platform (Wuxi NextCODE, Shanghai, China). Sequencing reads were aligned to the human reference genome GRCh38 using BWA-MEM v0.7.10 (54), and variants were called following the Genome Analysis Toolkit (GATK v4.1) Best Practices workflow. In the 20 deep-sequenced Qiang genomes, we identified 6,720,460 high-quality biallelic single-nucleotide variants (SNVs), including 210,620 novel SNVs not reported in single nucleotide polymorphism database (dbSNP) build 151. To achieve comprehensive analytical insights, we collected 1123 high-coverage WGS samples from 33 diverse East Asian populations including 131 Han and 33 Tibetans via the Chinese Pangenome Consortium (19, 53, 55) and 783 samples from 46 global populations through the Human Genome Diversity Project (HGDP) (56). Ultimately, we generated a high-quality jointly called WGS panel comprising 1926 samples (panel A), containing more than 100 million biallelic SNVs for downstream analyses. We also performed variant phasing on data panel A using SHAPEIT4 (57) for haplotype-based analysis.
On the basis of the data panel A, we generated two additional combined datasets for different analytical purposes. The first, for population structure analysis (panel B), was generated by intersecting panel A with the Human Origins dataset (58). After filtering out second-degree relatives using KING v2.2.8, we obtained 3128 unrelated individuals and 533,878 autosomal SNVs, confirming that all 20 Qiang individuals were unrelated. The second, for ancient DNA analysis (panel C), incorporated 75 ancient Tibetan Plateau samples from Wang et al. (59), 11 ancient southwest China samples from Tao et al. (4), and 376 ancient East Asia samples from the 1240K dataset (table S3). We used mergit program from EIGENSOFT v8.0.0 (60) to combine these datasets with the 20 Qiang individuals, 33 Tibetans, 33 Han samples from HGDP, and other representative modern individuals (table S3) extracted from the data panel A, yielding a final dataset containing 771,531 SNVs across 711 individuals.
Population structure
PCA of modern human samples was performed with EIGENSOFT v8.0.0 based on the data panel B (60). Variants with a missing rate ≥ 10% or minor allele frequency (MAF) ≤ 5% were excluded. To mitigate the influence of LD, we thinned the SNVs that were less than 50 kb apart from each other by using PLINK v1.90 (61). We performed a series of PCA under different contexts—Eurasia (using 42,273 SNVs), East Asia (41,356 SNVs), and Han-Qiang-Tibetan (41,239 SNVs) contexts—with subset samples extracted through bcftools v1.14 (62). Unsupervised hierarchical clustering of the Qiang samples was conducted by stats package hclust function based on R v4.2.3.
We used the maximum likelihood algorithm contml in PHYLIP to estimate the population phylogenetic tree (63). Populations of interest were subset from data panel B, and variants with a MAF ≤ 5% were filtered out, yielding 328,600 SNVs for inference. The consensus tree was generated on the basis of 100 bootstrap down-samplings.
Ancestry composition
We used ADMIXTURE (64) to infer the ancestral composition of both our target and reference populations based on data panel B. SNVs were filtered for MAF and LD using PLINK v1.9 (61), resulting in 181,448 SNVs for ADMIXTURE analysis. We ran ADMIXTURE by assuming different numbers of ancestral sources (K) ranging from 2 to 20. CV errors were calculated, with the minimum CV error (0.485) observed at K = 12 (fig. S3). To minimize the effects of local optima, each K was run 10 times with different random seeds. Last, we used pong (65) to identify a representative run for each K, and the ancestry components were visualized using AncestryPainter 2.0 (66, 67).
Genetic differentiation and recent gene flow
We estimated the overall population genetic distance using Weir and Cockerham’s FST. To minimize sample size bias, 10 individuals were randomly selected from each population; for populations with fewer than 10 samples, all available individuals were included. This random sampling was repeated 100 times for bootstrapping. SNVs with a missing rate greater than 10% were excluded, resulting in 378,535 SNVs for the final analysis.
To further dissect recent population admixture, we used an FST-based method as described by Gao et al. (15). Specifically, we calculated the RD of genomic affinity between two target populations (A and B) and other reference populations (X). The formula is as follows
A positive RD indicates greater genetic communication between population X and population B, while a negative RD suggests stronger genetic communication between population X and population A.
Ancient DNA analysis
PCA of ancient samples was performed by using EIGENSOFT v8.0.0 based on data panel C, with the parameter “lsproject: YES.” The ancient samples were classified according to the sample location as described by Tao et al. (4). Outgroup-f3 statistics between Qiang subpopulations and ancient samples were calculated using the qp3Pop program from ADMIXTOOLS (68) with the parameter “inbreed: YES.” We used the f3 statistic in the form f3(Qiang_H/Qiang_T, ancient; Mbuti) to assess shared genetic drift between Qiang subpopulations and ancient samples. In addition, D-statistics were computed using the qpDstat program from ADMIXTOOLS in the form D (Mbuti, ancient; Qiang_H, Qiang_T) to measure allele sharing between Qiang subpopulations and ancient samples. A significantly positive D-statistic (Z ≥ 3) suggests excess allele sharing between ancient samples and Qiang_T, while a significantly negative value (Z ≤ −3) indicates excess allele sharing between ancient samples and Qiang_H.
Demographic history inference
The historical effective population size (Ne) and divergence time were inferred by MSMC2 (69) and MSMC-IM (8) based on the phased data panel A. For Ne estimation, we sampled eight haplotypes from the same population, and the estimation was repeated five times using MSMC2. To estimate divergence time, we sampled four haplotypes from each of two populations and calculated coalescence rates with MSMC2. We then applied MSMC-IM to estimate time-dependent migration rates. The point where the cumulative migration probability dropped below 0.5 was considered the divergence time between populations. The median of the five repeated estimates was used as the time for population split.
We used MultiWaverX (20), an automated tool that selects and fits the best admixture model based on the distribution of ancestral segments, to estimate the admixture model for the Qiang subpopulations. We applied HAPMIX (17) to infer the local ancestry of the Qiang subpopulations. HGDP-Han and Tibetan were used as the ancestral populations. We carried out MultiWaverX analysis with the default parameters based on ancestral segments generated by HAPMIX. The Multi 2-2 model was determined as the best-fitting model as MultiWaverX described.
Natural selection
We applied FST-based PBS (26) for genome-wide natural selection signal detection based on the data panel A. The PBS was defined as
where T = −log(1 − FST), A is the target population, and B and C are the reference populations. We set Qiang_H or Qiang_T as the target population and HGDP-Han and Tibetan as the reference population, respectively. We subset these samples from the data panel A and filter the variants with missing rates larger than 10%, yielding 105,116,82 SNVs for natural selection signal scan. Locus-specific FST between populations were calculated by PLINK 2.0 (61). We focused on variants where PBS values fell within the top 0.1%, with the highest positive values indicating potential signals of positive selection and the lowest negative values (bottom 0.1%) indicating potential signals caused by admixture (13).
We also used selscan (70) for haplotype-based selection signal scan, including iHS and XP-EHH, based on phased data panel A. We used the same set of samples as those used in the PBS analysis. We applied a 0.05-cM window size and a 0.025-cM step size to split the genome. Within each window, the maximum absolute normalized iHS or XP-EHH value was used to represent the window. Window values ranked in the top 0.5% were considered as candidate regions under selection. SNVs within these windows with an absolute normalized iHS or XP-EHH value greater than 2 were considered core variants under selection.
Allele age estimation
Allele age was estimated with the GEVA based on the recombination and mutation joint clock model (46) using the phased panel A. To improve haplotype resolution for dating, we expanded the Han set to 164 individuals (rs143216880-T frequency = 0.058), resulting in 19 Han, 14 Yao, and 13 Qiang haplotypes for analysis. Genetic maps of each chromosome from HapMap were used to convert data from phased Variant Call Format (VCF) to the input files of GEVA. The initial effective population size and mutation rate were set according to the manual as 1 × 104 and 1.25 × 10−8, respectively.
Ethics statement
The use of genetic data in this study has been approved by the National Health Commission of the People’s Republic of China (no. 2025BAT01058).
Acknowledgments
We thank M. Stoneking for the helpful discussions on this manuscript. We thank L. Chen and X. Bai for sharing the blood pressure data of natural populations. We thank all the participants of this project. The computational work in this study was supported by the CFFF Computing Platform of Fudan University.
Funding:
This study was supported by the National Key Research and Development Program of China (no. 2023YFC2605400 to S.X.); the National Natural Science Foundation of China (NSFC) grants (32288101 to S.X., 32030020 to S.X., 32470649 to Y.L., 323B2013 to S.X., 32300499 to Y.G., and 32270665 to S.X.); the Shanghai Science and Technology Commission Program (25JS2810100 to S.X., QNKJ2024023 to Y.L., and 23JS1410100 to S.X.); the Office of Global Partnerships (Key Projects Development Fund to S.X.); and the Yunnan Leading Medical Scientist Training Program (L-2018003 to Z.Y.). The funders had no role in the study design, data collection, analysis, decision to publish, or preparation of the manuscript.
Author contributions:
Conceptualization: S.X. and J.C. Methodology: S.X., W.Z., and Z.Y. Validation: W.Z., S.X., and Z.Y. Formal analysis: W.Z., S.X., Y.G., C.L., Y.Y., and X.W. Investigation: W.Z., S.X., and Z.Y. Resources: S.X., J.C., Y.L., and Z.Y. Data curation: S.X., J.C., W.Z., Y.L., and Z.Y. Writing—original draft: S.X. and W.Z. Writing—review and editing: S.X., W.Z., Y.G., J.C., Y.L., and Z.Y. Visualization: W.Z. and J.C. Supervision: S.X., J.C., Y.L., and Z.Y. Project administration: S.X., J.C., and Z.Y. Funding acquisition: S.X. and Z.Y.
Competing interests:
The authors declare they have no competing interests.
Data and materials availability:
The variants of 20 Qiang samples have been deposited in the Genome Variation Map (GVM, https://ngdc.cncb.ac.cn/gvm) in the National Genomics Data Center under accession number GVM001004. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.
Supplementary Materials
The PDF file includes:
Supplementary Text
Figs. S1 to S33
Legends for tables S1 to S15
Other Supplementary Material for this manuscript includes the following:
Tables S1 to S15
REFERENCES
- 1.Fei X., The formation and development of the Chinese nation with multi-ethnic groups. Int. J. Anthropol. Ethnol. 1, 1 (2017). [Google Scholar]
- 2.Su B., Xiao C., Deka R., Seielstad M. T., Kangwanpong D., Xiao J., Lu D., Underhill P., Cavalli-Sforza L., Chakraborty R., Jin L., Y chromosome haplotypes reveal prehistorical migrations to the Himalayas. Hum. Genet. 107, 582–590 (2000). [DOI] [PubMed] [Google Scholar]
- 3.Yang Z., Chen H., Lu Y., Gao Y., Sun H., Wang J., Jin L., Chu J., Xu S., Genetic evidence of tri-genealogy hypothesis on the origin of ethnic minorities in Yunnan. BMC Biol. 20, 166 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Tao L., Yuan H., Zhu K., Liu X., Guo J., Min R., He H., Cao D., Yang X., Zhou Z., Wang R., Zhao D., Ma H., Chen J., Zhao J., Li Y., He Y., Suo D., Zhang R., Li S., Li L., Yang F., Li H., Zhang L., Jin L., Wang C.-C., Ancient genomes reveal millet farming-related demic diffusion from the Yellow River into southwest China. Curr. Biol. 33, 4995–5002.e7 (2023). [DOI] [PubMed] [Google Scholar]
- 5.He G., Adnan A., Al-Qahtani W. S., Safhi F. A., Yeh H. Y., Hadi S., Wang C. C., Wang M., Liu C., Yao J., Genetic admixture history and forensic characteristics of Tibeto-Burman-speaking Qiang people explored via the newly developed Y-STR panel and genome-wide SNP data. Front. Ecol. Evol. 10, 1–19 (2022). [Google Scholar]
- 6.Kang L., Lu Y., Wang C., Hu K., Chen F., Liu K., Li S., Jin L., Li H., The genographic consortium, Y-chromosome O3 haplogroup diversity in Sino-Tibetan populations reveals two migration routes into the Eastern Himalayas: Y-chromosomes of Eastern Himalayans. Ann. Hum. Genet. 76, 92–99 (2012). [DOI] [PubMed] [Google Scholar]
- 7.Song M., Wang Z., Lyu Q., Ying J., Wu Q., Jiang L., Wang F., Zhou Y., Song F., Luo H., Hou Y., Song X., Ying B., Paternal genetic structure of the Qiang ethnic group in China revealed by high-resolution Y-chromosome STRs and SNPs. Forensic Sci. Int.: Genet. 61, 102774 (2022). [DOI] [PubMed] [Google Scholar]
- 8.Wang C. C., Wang L. X., Shrestha R., Zhang M., Huang X. Y., Hu K., Jin L., Li H., Genetic structure of Qiangic populations residing in the Western Sichuan corridor. PLOS ONE 9, e103772 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bocher O., Willer C. J., Zeggini E., Unravelling the genetic architecture of human complex traits through whole genome sequencing. Nat. Commun. 14, 3520 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Dering C., Hemmelmann C., Pugh E., Ziegler A., Statistical analysis of rare sequence variants: An overview of collapsing methods. Genet. Epidemiol. 35, S12–S17 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Lou H., Gao Y., Xie B., Wang Y., Zhang H., Shi M., Ma S., Zhang X., Liu C., Xu S., Haplotype-resolved de novo assembly of a Tujia genome suggests the necessity for high-quality population-specific genome references. Cell Syst. 13, 321–333.e6 (2022). [DOI] [PubMed] [Google Scholar]
- 12.Zhang C., Dong S.-S., Xu J.-Y., He W.-M., Yang T.-L., PopLDdecay: A fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics 35, 1786–1788 (2019). [DOI] [PubMed] [Google Scholar]
- 13.Pan Y., Zhang C., Lu Y., Ning Z., Lu D., Gao Y., Zhao X., Yang Y., Guan Y., Mamatyusupu D., Xu S., Genomic diversity and post-admixture adaptation in the Uyghurs. Natl. Sci. Rev. 9, nwab124 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Pemberton T. J., Absher D., Feldman M. W., Myers R. M., Rosenberg N. A., Li J. Z., Genomic patterns of homozygosity in worldwide human populations. Am. J. Hum. Genet. 91, 275–292 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Gao Y., Zhang X., Chen H., Lu Y., Ma S., Yang Y., Zhang M., Xu S., Reconstructing the ancestral gene pool to uncover the origins and genetic links of Hmong–Mien speakers. BMC Biol. 22, 59 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Zhou Y., Browning S. R., Browning B. L., A fast and simple method for detecting identity-by-descent segments in large-scale data. Am. J. Hum. Genet. 106, 426–437 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Price A. L., Tandon A., Patterson N., Barnes K. C., Rafaels N., Ruczinski I., Beaty T. H., Mathias R., Reich D., Myers S., Sensitive detection of chromosomal segments of distinct ancestry in admixed populations. PLOS Genet. 5, e1000519 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Yang M. A., Fan X., Sun B., Chen C., Lang J., Ko Y.-C., Tsang C., Chiu H., Wang T., Bao Q., Wu X., Hajdinjak M., Ko A. M.-S., Ding M., Cao P., Yang R., Liu F., Nickel B., Dai Q., Feng X., Zhang L., Sun C., Ning C., Zeng W., Zhao Y., Zhang M., Gao X., Cui Y., Reich D., Stoneking M., Fu Q., Ancient DNA indicates human population shifts and admixture in northern and southern China. Science 369, 282–288 (2020). [DOI] [PubMed] [Google Scholar]
- 19.Lu D., Lou H., Yuan K., Wang X., Wang Y., Zhang C., Lu Y., Yang X., Deng L., Zhou Y., Feng Q., Hu Y., Ding Q., Yang Y., Li S., Jin L., Guan Y., Su B., Kang L., Xu S., Ancestral origins and genetic history of tibetan highlanders. Am. J. Hum. Genet. 99, 580–594 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Zhang R., Ni X., Yuan K., Pan Y., Xu S., MultiWaverX : Modeling latent sex-biased admixture history. Brief. Bioinform. 23, bbac179 (2022). [DOI] [PubMed] [Google Scholar]
- 21.Wang K., Li M., Hakonarson H., ANNOVAR: Functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Mahajan A., Taliun D., Thurner M., Robertson N. R., Torres J. M., Rayner N. W., Steinthorsdottir V., Scott R. A., Grarup N., Cook J. P., Schmidt E. M., Wuttke M., Sarnowski C., Mägi R., Nano J., Gieger C., Trompet S., Lecoeur C., Preuss M., Prins B. P., Guo X., Bielak L. F., Bennett A. J., Bork-Jensen J., Brummett C. M., Canouil M., Eckardt K.-U., Fischer K., Kardia S. L., Kronenberg F., Läll K., Liu C.-T., Locke A. E., Luan J., Ntalla I., Nylander V., Schönherr S., Schurmann C., Yengo L., Bottinger E. P., Brandslund I., Christensen C., Dedoussis G., Florez J. C., Ford I., Franco O. H., Frayling T. M., Giedraitis V., Hackinger S., Hattersley A. T., Herder C., Ikram M. A., Ingelsson M., Jørgensen M. E., Jørgensen T., Kriebel J., Kuusisto J., Ligthart S., Lindgren C. M., Linneberg A., Lyssenko V., Mamakou V., Meitinger T., Mohlke K. L., Morris A. D., Nadkarni G., Pankow J. S., Peters A., Sattar N., Stančáková A., Strauch K., Taylor K. D., Thorand B., Thorleifsson G., Thorsteinsdottir U., Tuomilehto J., Witte D. R., Dupuis J., Peyser P. A., Zeggini E., Loos R. J. F., Froguel P., Ingelsson E., Lind L., Groop L., Laakso M., Collins F. S., Jukema J. W., Palmer C. N. A., Grallert H., Metspalu A., Dehghan A., Köttgen A., Abecasis G., Meigs J. B., Rotter J. I., Marchini J., Pedersen O., Hansen T., Langenberg C., Wareham N. J., Stefansson K., Gloyn A. L., Morris A. P., Boehnke M., McCarthy M. I., Fine-mapping of an expanded set of type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps. Nat. Genet. 50, 1505–1513 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Lou H., Lu Y., Lu D., Fu R., Wang X., Feng Q., Wu S., Yang Y., Li S., Kang L., Guan Y., Hoh B.-P., Chung Y.-J., Jin L., Su B., Xu S., A 3.4-kb Copy-number deletion near EPAS1 is significantly enriched in high-altitude Tibetans but Absent from the denisovan sequence. Am. J. Hum. Genet. 97, 54–66 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Xu S., Li S., Yang Y., Tan J., Lou H., Jin W., Yang L., Pan X., Wang J., Shen Y., Wu B., Wang H., Jin L., A genome-wide search for signals of high-altitude adaptation in Tibetans. Mol. Biol. Evol. 28, 1003–1011 (2011). [DOI] [PubMed] [Google Scholar]
- 25.Deng L., Zhang C., Yuan K., Gao Y., Pan Y., Ge X., He Y., Yuan Y., Lu Y., Zhang X., Chen H., Lou H., Wang X., Lu D., Liu J., Tian L., Feng Q., Khan A., Yang Y., Jin Z.-B., Yang J., Lu F., Qu J., Kang L., Su B., Xu S., Prioritizing natural-selection signals from the deep-sequencing genomic data suggests multi-variant adaptation in Tibetan highlanders. Natl. Sci. Rev. 6, 1201–1222 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Yi X., Liang Y., Huerta-Sanchez E., Jin X., Cuo Z. X. P., Pool J. E., Xu X., Jiang H., Vinckenbosch N., Korneliussen T. S., Zheng H., Liu T., He W., Li K., Luo R., Nie X., Wu H., Zhao M., Cao H., Zou J., Shan Y., Li S., Yang Q., Asan, Ni P., Tian G., Xu J., Liu X., Jiang T., Wu R., Zhou G., Tang M., Qin J., Wang T., Feng S., Li G., Huasang, Luosang J., Wang W., Chen F., Wang Y., Zheng X., Li Z., Bianba Z., Yang G., Wang X., Tang S., Gao G., Chen Y., Luo Z., Gusang L., Cao Z., Zhang Q., Ouyang W., Ren X., Liang H., Zheng H., Huang Y., Li J., Bolund L., Kristiansen K., Li Y., Zhang Y., Zhang X., Li R., Li S., Yang H., Nielsen R., Wang J., Wang J., Sequencing of 50 human exomes reveals adaptation to high altitude. Science 329, 75–78 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Yang Z., Bai C., Pu Y., Kong Q., Guo Y., Ouzhuluobu, Gengdeng, Liu X., Zhao Q., Qiu Z., Zheng W., He Y., Lin Y., Deng L., Zhang C., Xu S., Peng Y., Xiang K., Zhang X., Baimayangji, Cirenyangji, Cui C., Baimakangzhuo, Gonggalanzi, Bianba, Pan Y., Xin J., Wang Y., Liu S., Wang L., Guo H., Feng Z., Wang S., Shi H., Jiang B., Wu T., Qi X., Su B., Genetic adaptation of skin pigmentation in highland Tibetans. Proc. Natl. Acad. Sci. U.S.A. 119, e2200421119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Deng L., Xu S., Adaptation of human skin color in various populations. Hereditas 155, 1 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Peng Y., Shi H., Qi X., Xiao C., Zhong H., Ma R. Z., Su B., The ADH1B Arg47His polymorphism in East Asian populations and expansion of rice domestication in history. BMC Evol. Biol. 10, 15 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Lu Y., Kang L., Hu K., Wang C., Sun X., Chen F., Kidd J. R., Kidd K. K., Li H., High diversity and no significant selection signal of human ADH1B gene in Tibet. Investig. Genet. 3, 23 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Baxter L. L., Watkins-Chow D. E., Pavan W. J., Loftus S. K., A curated gene list for expanding the horizons of pigmentation biology. Pigment Cell Melanoma Res. 32, 348–358 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Thaler R., Yoshizaki K., Nguyen T., Fukumoto S., Den Besten P., Bikle D. D., Oda Y., Mediator 1 ablation induces enamel-to-hair lineage conversion in mice through enhancer dynamics. Commun. Biol. 6, 766 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Mukhopadhyay S., Wen X., Ratti N., Loktev A., Rangell L., Scales S. J., Jackson P. K., The ciliary G-protein-coupled receptor Gpr161 negatively regulates the sonic hedgehog pathway via cAMP signaling. Cell 152, 210–223 (2013). [DOI] [PubMed] [Google Scholar]
- 34.Yaar M., Wu C., Park H.-Y., Panova I., Schutz G., Gilchrest B. A., Bone morphogenetic protein-4, a novel modulator of melanogenesis. J. Biol. Chem. 281, 25307–25314 (2006). [DOI] [PubMed] [Google Scholar]
- 35.Maples B. K., Gravel S., Kenny E. E., Bustamante C. D., RFMix: A discriminative modeling approach for rapid and robust local-ancestry inference. Am. J. Hum. Genet. 93, 278–288 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Hurley T. D., Edenberg H. J., Genes encoding enzymes involved in ethanol metabolism. Alcohol Res. 34, 339–344 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Edenberg H. J., Xuei X., Chen H.-J., Tian H., Wetherill L. F., Dick D. M., Almasy L., Bierut L., Bucholz K. K., Goate A., Hesselbrock V., Kuperman S., Nurnberger J., Porjesz B., Rice J., Schuckit M., Tischfield J., Begleiter H., Foroud T., Association of alcohol dehydrogenase genes with alcohol dependence: A comprehensive analysis. Hum. Mol. Genet. 15, 1539–1549 (2006). [DOI] [PubMed] [Google Scholar]
- 38.Li H., Gu S., Han Y., Xu Z., Pakstis A. J., Jin L., Kidd J. R., Kidd K. K., Diversification of the ADH1B gene during expansion of modern humans. Ann. Hum. Genet. 75, 497–507 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Choudhury C., Butcher N. J., Minchin R. F., Polymorphism in the human arylamine N-acetyltransferase 1 gene 3′-untranslated region determines polyadenylation signal usage. Biochem. Pharmacol. 200, 115020 (2022). [DOI] [PubMed] [Google Scholar]
- 40.Rother N., Yanginlar C., Prévot G., Jonkman I., Jacobs M., van Leent M. M. T., van Heck J., Matzaraki V., Azzun A., Morla-Folch J., Ranzenigo A., Wang W., van der Meel R., Fayad Z. A., Riksen N. P., Hilbrands L. B., Lindeboom R. G. H., Martens J. H. A., Vermeulen M., Joosten L. A. B., Netea M. G., Mulder W. J. M., van der Vlag J., Teunissen A. J. P., Duivenvoorden R., Acid ceramidase regulates innate immune memory. Cell Rep. 42, 113458 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Edwards K., Lydyard P. M., Kulikova N., Tsertsvadze T., Volpi E. V., Chiorazzi N., Porakishvili N., The role of CD180 in hematological malignancies and inflammatory disorders. Mol. Med. 29, 97 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Yan Y., Wang J., Yu L., Cui B., Wang H., Xiao X., Zhang Y., Zheng J., Wang J., Hui R., Wang Y., ANKRD36 is involved in hypertension by altering expression of ENaC genes. Circ. Res. 129, 1067–1081 (2021). [DOI] [PubMed] [Google Scholar]
- 43.Calbet J. A. L., Chronic hypoxia increases blood pressure and noradrenaline spillover in healthy humans. J. Physiol. 551, 379–386 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Naeije R., Physiological adaptation of the cardiovascular system to high altitude. Prog. Cardiovasc. Dis. 52, 456–466 (2010). [DOI] [PubMed] [Google Scholar]
- 45.Faulhaber M., Gatterer H., Haider T., Linser T., Netzer N., Burtscher M., Heart rate and blood pressure responses during hypoxic cycles of a 3-week intermittent hypoxia breathing program in patients at risk for or with mild COPD. Int. J. Chron. Obstruct. Pulmon. Dis. 10, 339–345 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Albers P. K., McVean G., Dating genomic variants and shared ancestry in population-scale sequencing data. PLoS Biol. 18, e3000586 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Shi S., Ethnic flows in the Tibetan-Yi corridor throughout history. Int. J. Anthropol. Ethnol. 2, 2 (2018). [Google Scholar]
- 48.Liu L., Chen J., Wang J., Zhao Y., Chen X., Archaeological evidence for initial migration of Neolithic Proto Sino-Tibetan speakers from Yellow River valley to Tibetan Plateau. Proc. Natl. Acad. Sci. U.S.A. 119, e2212006119 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Yang M. A., A genetic history of migration, diversification, and admixture in Asia. Hum. Popul. Genet. Genom. 2, 0001 (2022). [Google Scholar]
- 50.W. Ming-ke, “Making history while being in history: The histories of the Qiang and Rma” in Chasing Traces: History and Ethnography in the Uplands of Socialist Asia, P. Petit, J. Michaud, Eds. (University of Hawai’i Press, 2024), pp. 176–200. http://jstor.org/stable/jj.6505269.11.
- 51.Xu S., Jin W., Jin L., Haplotype-sharing analysis showing uyghurs are unlikely genetic donors. Mol. Biol. Evol. 26, 2197–2206 (2009). [DOI] [PubMed] [Google Scholar]
- 52.Sun W., Banbur D., Li Y., Diversified communication and harmony of chinese culture: A historical investigation into the religious-culture exchanges among the Hans, Tibetans, and Mongolians. Int. J. Anthropol. Ethnol. 8, 22 (2024). [Google Scholar]
- 53.Gao Y., Yang X., Chen H., Tan X., Yang Z., Deng L., Wang B., Kong S., Li S., Cui Y., Lei C., Wang Y., Pan Y., Ma S., Sun H., Zhao X., Shi Y., Yang Z., Wu D., Wu S., Zhao X., Shi B., Jin L., Hu Z., Lu Y., Chu J., Ye K., Xu S., A pangenome reference of 36 Chinese populations. Nature 619, 112–121 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Li H., Durbin R., Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25, 1754–1760 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Gao Y., Zhang C., Yuan L., Ling Y. C., Wang X., Liu C., Pan Y., Zhang X., Ma X., Wang Y., Lu Y., Yuan K., Ye W., Qian J., Chang H., Cao R., Yang X., Ma L., Ju Y., Dai L., Tang Y., The Han100K Initiative, Zhang G., Xu S., PGG.Han: The Han Chinese genome database and analysis platform. Nucleic Acids Res. 48, D971–D976 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Bergström A., McCarthy S. A., Hui R., Almarri M. A., Ayub Q., Danecek P., Chen Y., Felkel S., Hallast P., Kamm J., Blanché H., Deleuze J.-F., Cann H., Mallick S., Reich D., Sandhu M. S., Skoglund P., Scally A., Xue Y., Durbin R., Tyler-Smith C., Insights into human genetic variation and population history from 929 diverse genomes. Science 367, eaay5012 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Delaneau O., Zagury J.-F., Robinson M. R., Marchini J. L., Dermitzakis E. T., Accurate, scalable and integrative haplotype estimation. Nat. Commun. 10, 5436 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Lazaridis I., Patterson N., Mittnik A., Renaud G., Mallick S., Kirsanow K., Sudmant P. H., Schraiber J. G., Castellano S., Lipson M., Berger B., Economou C., Bollongino R., Fu Q., Bos K. I., Nordenfelt S., Li H., de Filippo C., Prüfer K., Sawyer S., Posth C., Haak W., Hallgren F., Fornander E., Rohland N., Delsate D., Francken M., Guinet J.-M., Wahl J., Ayodo G., Babiker H. A., Bailliet G., Balanovska E., Balanovsky O., Barrantes R., Bedoya G., Ben-Ami H., Bene J., Berrada F., Bravi C. M., Brisighelli F., Busby G. B. J., Cali F., Churnosov M., Cole D. E. C., Corach D., Damba L., van Driem G., Dryomov S., Dugoujon J.-M., Fedorova S. A., Romero I. G., Gubina M., Hammer M., Henn B. M., Hervig T., Hodoglugil U., Jha A. R., Karachanak-Yankova S., Khusainova R., Khusnutdinova E., Kittles R., Kivisild T., Klitz W., Kučinskas V., Kushniarevich A., Laredj L., Litvinov S., Loukidis T., Mahley R. W., Melegh B., Metspalu E., Molina J., Mountain J., Näkkäläjärvi K., Nesheva D., Nyambo T., Osipova L., Parik J., Platonov F., Posukh O., Romano V., Rothhammer F., Rudan I., Ruizbakiev R., Sahakyan H., Sajantila A., Salas A., Starikovskaya E. B., Tarekegn A., Toncheva D., Turdikulova S., Uktveryte I., Utevska O., Vasquez R., Villena M., Voevoda M., Winkler C. A., Yepiskoposyan L., Zalloua P., Zemunik T., Cooper A., Capelli C., Thomas M. G., Ruiz-Linares A., Tishkoff S. A., Singh L., Thangaraj K., Villems R., Comas D., Sukernik R., Metspalu M., Meyer M., Eichler E. E., Burger J., Slatkin M., Pääbo S., Kelso J., Reich D., Krause J., Ancient human genomes suggest three ancestral populations for present-day Europeans. Nature 513, 409–413 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Wang H., Yang M. A., Wangdue S., Lu H., Chen H., Li L., Dong G., Tsring T., Yuan H., He W., Ding M., Wu X., Li S., Tashi N., Yang T., Yang F., Tong Y., Chen Z., He Y., Cao P., Dai Q., Liu F., Feng X., Wang T., Yang R., Ping W., Zhang Z., Gao Y., Zhang M., Wang X., Zhang C., Yuan K., Ko A. M.-S., Aldenderfer M., Gao X., Xu S., Fu Q., Human genetic history on the Tibetan Plateau in the past 5100 years. Sci. Adv. 9, eadd5582 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Patterson N., Price A. L., Reich D., Population structure and eigenanalysis. PLOS Genet. 2, e190 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Chang C. C., Chow C. C., Tellier L. C., Vattikuti S., Purcell S. M., Lee J. J., Second-generation PLINK: Rising to the challenge of larger and richer datasets. GigaScience 4, 7 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Danecek P., Bonfield J. K., Liddle J., Marshall J., Ohan V., Pollard M. O., Whitwham A., Keane T., McCarthy S. A., Davies R. M., Li H., Twelve years of SAMtools and BCFtools. GigaScience 10, giab008 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Joseph F., PHYLIP—Phylogeny inference package (Version 3.2). Cladistics 5, 164–166 (1989). [Google Scholar]
- 64.Alexander D. H., Novembre J., Lange K., Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 19, 1655–1664 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Behr A. A., Liu K. Z., Liu-Fang G., Nakka P., Ramachandran S., pong: Fast analysis and visualization of latent clusters in population genetic data. Bioinformatics 32, 2817–2823 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Chen S., Lei C., Zhao X., Pan Y., Lu D., Xu S., AncestryPainter 2.0: Visualizing ancestry composition and admixture history graph. Genome Biol. Evol. 16, evae249 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Feng Q., Lu D., Xu S., AncestryPainter: A graphic program for displaying ancestry composition of populations and individuals. Genomics Proteomics Bioinformatics 16, 382–385 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Patterson N., Moorjani P., Luo Y., Mallick S., Rohland N., Zhan Y., Genschoreck T., Webster T., Reich D., Ancient admixture in human history. Genetics 192, 1065–1093 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.S. Schiffels, K. Wang, “MSMC and MSMC2: The Multiple Sequentially Markovian Coalescent” in Statistical Population Genomics, J. Y. Dutheil, Ed. (Springer US, 2020), pp. 147–166. 10.1007/978-1-0716-0199-0_7. [DOI] [PubMed]
- 70.Szpiech Z. A., Hernandez R. D., selscan: An efficient multithreaded program to perform EHH-based scans for positive selection. Mol. Biol. Evol. 31, 2824–2827 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Text
Figs. S1 to S33
Legends for tables S1 to S15
Tables S1 to S15
Data Availability Statement
The variants of 20 Qiang samples have been deposited in the Genome Variation Map (GVM, https://ngdc.cncb.ac.cn/gvm) in the National Genomics Data Center under accession number GVM001004. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.






