Abstract
Understanding fine-scale genetic and geographic ancestry in East and Southeast Asia is difficult due to complex population histories and limited high-resolution genomic data. Here, we introduce a comprehensive framework that combines ancestry-informative single nucleotide polymorphism (AISNP) panels with machine learning to jointly determine genetic ancestry and geographic origins in 1,703 individuals from 67 East and Southeast Asian groups. We developed seven nested AISNP panels, from 50 to 2,000 SNPs, and tested six classification algorithms: logistic regression, support vector machines, k-nearest neighbors, random forest, convolutional neural networks, and eXtreme Gradient Boosting (XGBoost). The best results came from the optimized XGBoost model, which achieved 95.6% accuracy and an AUC of 0.999 with 2,000 AISNPs. For geographic localization, we used the Locator model, a deep neural network that predicts latitude and longitude directly from unphased genotypes. Notably, Locator trained on just 2,000 AISNPs performed nearly as well as models built on high-density genomic data (597,569 SNPs). Overall, these findings show that carefully designed AISNP panels combined with suitable machine learning techniques can provide highly accurate and efficient ancestry inference, offering valuable insights for population genetics, forensic science, and biogeography in East and Southeast Asia.
Supplementary Information
The online version contains supplementary material available at 10.1186/s40246-025-00837-3.
Keywords: Ancestry inference, Ancestry-informative SNP panels, Machine learning, Population genetics, Geographic localization, East and Southeast Asia
Introduction
Asia, one of the world’s most populous regions, harbors a complex genetic landscape shaped by admixture among diverse source populations with distinct cultural and historical backgrounds [1–5]. Recent across-Asian continent projects, such as the GenomeAsia 100 K Project [6], Indian genomic resources [7], SEA3K [8], and regional population genomic cohorts, such as the 10K_CPGDP [9], YanHuang Cohort [10, 11], STROMICS [12], China Kadoorie Biobank [13], NyuWa Genome Resource [14], ChinaMAP [15], Born in Guangzhou Cohort Study [16], the Healthy Zhejiang One Million Cohort Project [17], and others [18–22], have identified population-specific variants, fine-scale population Asian structures and publicly available genomic datasets. These advances within East Asia and Southeast Asia, along with unique evolutionary paths and repeated trans-Eurasian migrations, have created complex geographic-based stratifications in both ancient and modern populations [23–27]. These patterns of genetic diversity among East Asian and Southeast Asian populations reflect intricate historical migration patterns [20, 27–29]. Conventional approaches often fail to capture subtle genetic distinctions, particularly in scenarios involving multi-ethnic admixture or complex migratory histories [30, 31]. While genome-wide datasets provide high resolution, their scale and cost limit their routine use in forensic science, archaeology, and population genetics [20]. This has driven the need for reduced yet highly informative marker sets that capture fine-scale population structure while minimizing genotyping effort.
Biogeographic ancestry inference from DNA has significant value for criminal investigations, legal proceedings, and population studies [32]. Ancestry-informative markers (AIMs), defined as variants with large inter-population frequency differences, form the cornerstone of biogeographic ancestry inference [33]. Among different types of AIMs, SNPs are particularly well suited, given their stability, dense distribution, and strong frequency differentiation across populations [30]. Although several AIM panels have been applied in East Asian and Southeast Asian populations [34–41], achieving fine-scale resolution remains difficult. Variation in target populations, marker selection, and analytical models further limit their broader application. Despite these challenges, algorithms such as Geographic Population Structure (GPS) have shown that ancestry profiles derived from ADMIXTURE could be translated into geographic coordinates with remarkable precision, sometimes down to the village level [42]. These algorithmic advances underscore the potential of computational tools for ancestry inference. At the same time, they highlight both the feasibility of computational frameworks for geographic inference and the need for systematic strategies that integrate optimized marker sets with modern machine learning.
Advances in high-throughput genomic sequencing have greatly expanded the resources available for ancestry analysis [22, 43–50]. However, the complexity and high dimensionality of these data often exceed the capacity of traditional statistical approaches [50–52]. Although current methodologies exhibit robust differentiation among continental populations, their performance declines significantly within genetically homogeneous populations or geographically proximate groups. In East Asia and Southeast Asia, where complex admixture has blurred genetic-geographic correlations, this limitation is especially pronounced [53]. Machine learning approaches are increasingly filling this gap, offering scalability, flexibility, and the ability to capture non-linear relationships in genomic data [54, 55]. The deep learning-based Locator framework, for example, predicts geographic origin directly from unphased genotypes without requiring explicit spatial models. Simulations show that the Locator can infer locations within roughly four generations of diffusion, and it runs substantially faster than traditional models [56]. Other studies have demonstrated that deep learning and ensemble methods outperform classical approaches in resolving fine-scale structures [40, 57–59]. For example, Eugenio et al.. utilized deep learning architectures to predict fine-scale population structure with higher resolution than traditional methods [60]. Chen et al.. demonstrated that random forests (RF) models improved the accuracy of population assignment in substructure datasets [41]. These studies illustrate that machine learning approaches have been reliably used to enhance ancestry inference, providing a strong foundation for our framework that integrates optimized AISNP panels with diverse machine learning models. We designed a framework that integrates nested AISNP panels with diverse machine learning models, including logistic regression (LR), support vector machines (SVM), k-nearest neighbors (KNN), RF, convolutional neural networks (CNN), and XGBoost, to optimize ancestry inference across East Asia and Southeast Asia. Additionally, we applied the Locator model, a previously published deep learning framework, to predict geographic origins (latitude and longitude) from reference genotype datasets and compared its performance via reduced AISNP panels versus genome-wide data. This strategy enhances ancestry inference resolution and has significant practical implications for criminal investigations by narrowing geographic search areas during forensic analyses.
Materials and methods
Sample preparation and quality control
We implemented a multi-step framework integrating sample preparation, SNP selection, population structure analysis, and machine learning for ancestry inference. The overall workflow is illustrated in Fig. 1. We curated genotype data for 1,703 individuals representing 67 population groups across East Asia and Southeast Asia [29, 61–63]. These data were derived from the publicly available Human Origins dataset generated by the Reich Laboratory (accessible via https://reich.hms.harvard.edu/datasets). Family relationships were assessed by estimating kinship coefficients via the KING algorithm [64]. Samples were excluded if their genotype call rates ≥ 10% (--mind 0.1), SNPs with missing call rates ≥ 10% (--geno 0.1), minor allele frequencies < 1% (--maf 0.01), or Hardy-Weinberg equilibrium (HWE) p-values < 0.001 (--hwe 0.001) using PLINK 1.9 [65]. After quality control, 597,569 high-quality SNPs across 1,703 samples were retained for downstream analyses.
Fig. 1.
Overview of the integrated framework for biogeographical ancestry inference. The framework integrates machine learning based genetic ancestry classification (light blue dashed boxes) with geographic origin prediction using the established Locator model (Battey et al. 2020; red dashed box). The combined approach (green dashed box) uses reduced AISNP panels to achieve accurate and efficient inference. Created in BioRender
Ancestry classification and population grouping
We used ADMIXTURE [66] to estimate ancestry components from genome-wide data, applying default tenfold cross-validation (--cv = 10) across K = 2–20 with 100 bootstraps. To evaluate the impact of linkage disequilibrium (LD), we pruned SNPs with PLINK (--indep-pairwise 200 25 0.2) and re-ran ADMIXTURE on both pruned and unpruned datasets. The cross-validation (CV) error and differences between consecutive K values (ΔCV) were recorded to identify the range of clustering solutions. Based on these analyses, K = 9 provided the most supported structure in our dataset.
To provide a high-quality training dataset with precise geographical or cultural labels in this pilot work, we relabeled each individual based on genetic affinity and geographical labels. We assigned individuals to an ancestry group if their maximum ancestry proportion exceeded 50% at K = 9. This threshold was adopted as a pragmatic and interpretable criterion to define dominant ancestry components in model-based clustering. Alternative thresholds of 25% and 75% were also tested to assess the distribution of group assignments, but were not used in downstream analyses. Using this approach, 1,028 individuals were classified as “Certain”, while the remaining 675 individuals with more evenly distributed ancestry proportions (≤ 50%) were labeled as “Uncertain”. To reclassify these admixed individuals, we trained a RF classifier using ancestry components from K = 2 to K = 9 as input features, thereby capturing hierarchical population structure. Model performance was optimized through grid search over three hyperparameters: mtry (5 to 30), number of trees (100 to 2,000), and min_n (2 to 10). Tuning was performed using 10-fold stratified cross-validation with ROC-AUC as the primary evaluation metric, while out-of-bag (OOB) error was monitored for stability. Model performance plateaued at approximately 100 trees, and the configuration yielding the highest mean ROC-AUC (mtry = 5, trees = 100, min_n = 2) was selected. This tuned RF model provided stable posterior probabilities and successfully reclassified all 675 admixed individuals into one of the nine ancestry groups, reducing the risk of misclassification associated with a single K threshold.
Through this two-step unsupervised-plus-supervised framework, we assigned all 1,703 individuals to nine ancestry groups defined by genetic components, geographic distribution, and linguistic affiliation: Central Asia_Turkic (CA_Turkic, n = 115), North China_Sino-Tibetan_Altaic (NC_Sino-Tibetan_Altaic, n = 254), Northeast Asia_Tungusic (NEA_Tungusic, n = 285), South China_Tai-Kadai_Sinitic (SC_Tai-Kadai_Sinitic, n = 302), Southeast Asia_Austroasiatic (SEA_Austroasiatic, n = 234), Southeast Asia_Tai-Kadai_Sino-Tibetan (SEA_Tai-Kadai_Sino-Tibetan, n = 159), Southeast Coast China_East Asian Islands_Austronesian (SECC_ISEA_Austronesian, n = 60), Southwest Asia_Hmong-Mien (SWA_Hmong-Mien, n = 54), and West Siberia_Turkic (WS_Turkic, n = 240) (Fig. S1).
Ancestry informative SNP selection
AISNPs were identified via the AIM generator [67], which incorporates genotype and allele frequency data from PLINK, performs LD pruning, and ranks SNPs via Rosenberg’s In statistic [68]. The selection criteria were as follows: (1) chromosomal distribution, requiring SNPs to be located on different chromosomes or at least 1 Mb apart; (2) exclusion of duplicate SNPs and those on sex chromosomes; (3) high differentiation (FST, In) among populations; and (4) compliance of loci with HWE. This yielded 2,007 candidate SNPs. We computed allele frequencies and LD (r²) with PLINK and then selected nested panels of 50, 100, 250, 500, 1,000, 1,500, and 2,000 AISNPs, retaining the top-ranked variants by In values. Fig. S2 shows the even distribution of selected SNPs across all chromosomes.
Evaluation of AISNP panels
PCA was conducted using the smartPCA program [69] from the EIGENSOFT package [70] to examine population stratification. Ancestry proportions were subsequently estimated through unsupervised clustering via ADMIXTURE software, with tenfold cross-validation employed to increase the robustness of the models. Non-linear dimensionality reduction and visualization of complex population structures were performed via t-distributed stochastic neighbor embedding (t-SNE) analysis [71] via the Rtsne package in R v4.3.2.
Machine learning classification and panel evaluation
Data preprocessing
We randomly sampled 80% of the participants (1,363 individuals) as our training set, whereas the remaining 20% (340 individuals) composed the testing set. The input features included the first ten principal components (PC1-PC10) from PCA and ancestry proportions from the optimal ADMIXTURE K.
Machine learning models and implementations
We employed six machine learning algorithms, including LR, KNN, RF, SVM, XGBoost, and CNN, to identify the optimal classifier for ancestry prediction. Each model was trained with nested AISNP panels (50 − 2,000 markers).
LR was implemented as the baseline model due to its widespread application and reliability in classification tasks. The KNN model provided a non-parametric comparison to the baseline LR model. To achieve the optimal performance, we trained the KNN across a range of neighbors (k = 1–20) via 5-fold cross-validation. For RF, the mtry and ntree are two important hyperparameters that should be fine-tuned prior to the training process. We employed a grid search over candidate values of mtry (2, 3, 5, 7, 10) and ntree (100, 300, 500, 700) to identify the optimal combination. Each configuration was evaluated via 5-fold cross-validation. For the SVM, the radial basis function (RBF) was chosen due to its superior performance in classification tasks compared to other kernels. A total of 1,000 random combinations of the regularization parameter (C, 2-⁵-2⁵) and kernel parameter (σ, 2-⁵-2³) were sampled to determine the optimal RBF configuration, with each combination evaluated via 5-fold cross-validation. XGBoost requires more extensive hyperparameter tuning due to its complex structure. Six hyperparameters with significant impacts on model performance were tuned via Bayesian optimization: tree depth (max_depth: 3–10), learning rate (eta: 0.01–0.3), column sampling ratio (colsample_bytree: 0.5-1), subsample ratio (subsample: 0.5-1), minimum child weight (min_child_weight: 1–10), and loss reduction threshold (gamma: 0–5). During the training process, a 5-fold cross-validation was employed to improve model robustness. Finally, a one-dimensional CNN (1D-CNN) was utilized for classification. The architecture included two convolutional layers with max-pooling, a fully connected dense layer with dropout, and a Softmax output layer. The CNN was implemented in TensorFlow (version 2.18.0) (https://www.tensorflow.org/) with Keras (version 3.10.0) (https://keras.io/), and hyperparameters were tuned using the Keras Tuner (https://keras.io/keras_tuner/) in Python. We conducted 200 trials of random search to tune model hyperparameters, exploring ranges for the first convolutional layer (16–64 filters, step 16), kernel size (3 or 5), filters in the second convolutional layer (32–128, step 32), dense layer units (64–256, step 64), dropout rate (0.2–0.5, step 0.1), and Adam optimizer learning rate (0.001, 0.0005, 0.0001). Each candidate was trained for up to 150 epochs with early stopping (20% validation split). For all six models, the optimal configuration was selected based on the one-vs-rest macro AUC criterion.
Model evaluations
We assessed classification using micro-averaged ROC curves, the area under the curve (AUC), accuracy, sensitivity, specificity, PPV, and NPV. Metrics were calculated as follows: sensitivity = TP/(TP + FN); specificity = TN/(TN + FP); PPV = TP/(TP + FP); NPV = TN/(TN + FN); and accuracy = (TP + TN)/(TP + TN + FP + FN), where TP represents true positives, TN represents true negatives, FN represents false negatives, and FP represents false positives. For visualization, we plotted ROC-AUC curves for the best-performing model, XGBoost. To benchmark informativeness, we compared tuned XGBoost models trained on AISNPs with models trained on randomly selected SNPs of equal size. We evaluated performance differences via Wilcoxon rank-sum tests.
Model training and geographic prediction via deep neural networks
We used Locator [56] to predict sample coordinates (latitude, longitude) from unphased genotypes. Locator is an open-source deep learning model designed for geographic inference from genetic data and was originally developed by Battey et al.. (2020, Elife). Input VCFs were converted to allele count vectors. We split the dataset into training (75%, n = 1,278) and validation (25%, n = 425) sets, masking the validation coordinates. The models were trained on nested AISNP panels and the full genome-wide dataset (597,569 SNPs). Accuracy was assessed via nine validation samples and visualized at 95%, 50%, and 10% confidence levels in R v4.3.2. Training employed the Adam optimizer (Kingma and Ba, 2014) with the Euclidean distance as the loss function, which is defined as follows:
![]() |
Results
Population stratification in East Asia and Southeast Asia based on genome-wide data
We performed ADMIXTURE analysis on genome-wide SNPs to investigate hierarchical stratification across East Asia and Southeast Asia. For visualization, we focused on K = 2–9 (Fig. S3), which captured major continental to regional components. Cross-validation error analysis for K = 2–20 revealed a steady decrease until K = 9–10, after which plateaus or fluctuations suggested overfitting (Table S1, Fig. S4). The minimum CV error was observed at K = for unpruned data (0.39232) and at K = 10 for LD-pruned data (0.31280). Although K = 10 introduced an additional component, it was confined primarily to 11 individuals (10 BoY and one Han_Fujian) and did not represent a stable ancestry group. The ΔCV values also showed only marginal improvement beyond K = 9, supporting nine clusters as the best-supported and biologically interpretable structure. LD pruning did not alter these results, indicating robustness to marker dependence
To assign individuals into ancestry groups, we evaluated maximum ancestry proportions at thresholds of 25%, 50%, and 75%. At ≥ 25%, 1,486 individuals (87.2%) were assigned to a dominant ancestry; at ≥ 50%, 1,028 (60.4%); and at ≥ 75%, 643 (37.7%) (Table S2, Fig. S5). These results suggest that the 50% threshold provides a reasonable balance between classification confidence and population coverage, supporting its use as an operational grouping strategy for ancestry inference. Individuals not exceeding 50% were labeled “Uncertain” and reclassified with a RF model incorporating multi-K ancestry features, enabling consistent assignment of all 1,703 samples into nine ancestry groups defined by geography and linguistic affiliation (Fig. S1). PCA and ADMIXTURE on the genome-wide dataset (597,569 SNPs) further confirmed the differences among these nine groups (Fig. 2): PC1 separated northern from southern populations, whereas ADMIXTURE highlighted a finer substructure. Populations sharing geographic proximity and linguistic ties exhibited greater genetic affinity, aligning with known migration and demographic patterns in East and Southeast Asia.
Fig. 2.
Principal component analysis (PCA) and ADMIXTURE analysis (K = 9) of populations after prediction via the Random Forest method based on genome-wide sequencing data
AISNP selection and assessment of genetic ancestry inference
We constructed seven AISNP panels via the AIM generator based on data from 1,703 individuals grouped by geographic region or linguistic affiliation (Table S3). PCA of these panels revealed clear genetic relationships among the nine subpopulations (Fig. S6). With 50 or 100 markers, the populations clustered distinctly but with some overlap. Discriminative power increased markedly with larger panels, underscoring the importance of marker density for fine-scale ancestry inference. ADMIXTURE analysis with the 2,000-marker panel further resolved nine primary ancestry components and detected an additional minor component (represented in yellow) enriched in Southeast Asian groups (Fig. S7), including the Tai-Kadai, Sino-Tibetan, Austroasiatic, and Hmong-Mien groups. These results demonstrate that carefully selected AISNP panels can capture meaningful substructures across East Asia and Southeast Asia.
t-SNE analysis confirmed the impact of marker density on population resolution (Fig. 3). For 50 markers, the clusters overlapped extensively; for 100, 250, 500, and 1,000 markers, the population boundaries became progressively clearer. For the 1,500-2,000 markers, the clusters clearly separated and had minimal overlap, closely mirroring patterns from the genome-wide data (597,569 SNPs). Visualization further revealed genetic relationships, highlighting the close clustering between the NC_Sino-Tibetan_Altaic and NEA_Tungusic groups, which is indicative of genetic affinity. Similarly, the SC_Tai-Kadai_Sinitic and SWA_Hmong-Mien groups exhibited tighter clustering with increased AISNP numbers, underscoring their genetic distinctiveness. Overall, our analysis demonstrated that panels of 1,500-2,000 strategically selected AISNPs can achieve ancestry resolution comparable to that of genome-wide methods while enabling robust and computationally efficient inference.
Fig. 3.
t-SNE analysis plots for the seven AISNP sets and genome-wide sequencing data across the nine populations
Ancestry prediction for diverse populations via six machine learning methods and nested AISNP panels
We evaluated the performance of six machine learning methods (CNN, KNN, LR, RF, SVM, and XGBoost) for ancestry prediction across nine populations via nested AISNP panels of varying sizes (Tables S4, 5, Fig. 4). The overall accuracy improved steadily with increasing marker density, but the model behaviors varied considerably. LR provided stable performance across all panels, reflecting its ability to model linear relationships but its limited ability to capture complex structures. SVM initially performed comparably to LR, but hyperparameter tuning unexpectedly reduced accuracy, likely owing to sensitivity to kernel selection and parameter settings in high-dimensional ancestry data. This behavior highlights the need for cautious tuning and validation when applying SVMs to population genetic inference. KNN achieved high accuracy in certain populations (e.g., Class 9, WS_Turkic) but showed inconsistent overall performance. Interestingly, KNN performed worse when trained on genome-wide SNPs than when the 2,000-AISNP panel was used, underscoring the importance of informative marker selection over raw density. CNN underperformed on smaller panels, particularly below 250 markers, but showed substantial improvement once trained on larger sets and optimized through hyperparameter tuning. These improvements likely reflect the CNN’s ability to exploit local dependencies and hierarchical patterns in genotype data, although its computational demands and sensitivity to input size remain challenges for practical implementation.
Fig. 4.
Curve showing the average classification accuracy of the six machine learning algorithms in relation to the number of SNPs
Among tree-based approaches, RF demonstrated strong performance on larger panels, nearly matching genome-wide accuracy with 2,000 markers. XGBoost consistently outperformed all other models, achieving the highest accuracy and AUC across panel sizes. With 2,000 SNPs, XGBoost reached 95.6% accuracy and an AUC of 0.999, nearly identical to the genome-wide results. Importantly, XGBoost delivered balanced sensitivity and specificity across most populations, including groups with close genetic affinities, making it the most robust and interpretable classifier for ancestry inference in this dataset. Nevertheless, some populations remain difficult to distinguish due to overlapping genetic backgrounds. For example, with a CNN at 2,000 SNPs, Class 3 (SC_Tai-Kadai_Sinitic) achieved 0.869 sensitivity and 0.975 specificity, whereas Class 5 (SEA_Tai-Kadai_Sino-Tibetan) had lower sensitivity (0.659) despite high specificity (0.977). These discrepancies underscore the persistent challenge of resolving closely related or admixed groups. XGBoost was selected as the primary model for ancestry group classification in subsequent analyses. To assess the effect of model optimization, we compared tuned versus untuned XGBoost models across marker panels. Tuning consistently improved the sensitivity and predictive balance, especially for intermediate panel sizes (500-1,500 SNPs). For example, at 1,500 SNPs, the Class 5 sensitivity increased from 78.05% to 85.37%, the Class 1 sensitivity and NPV improved, and the Class 6 PPV rose from 89.19% to 91.67%. Although tuning occasionally reduced the PPV in some classes, it generally produced more stable and balanced predictions across metrics (Table S5). We further evaluated overall model performance via AUC values from the best-performing XGBoost classifier. The AUC increased with marker density, from 0.949 with 50 SNPs to 0.998 with 1,000–1,500 SNPs, reaching 0.999 with 2,000 SNPs and 1.000 with genome-wide data (597,569 SNPs) (Fig. S8). These results show that 1,000–2,000 carefully selected AISNPs can achieve near-genome-wide accuracy, providing an efficient alternative for practical applications. To evaluate statistical significance, we benchmarked informative AISNP panels against randomly selected SNP panels of equal size. Informative panels consistently outperformed random sets across all sizes. The 1,000–2,000 AISNP panels nearly matched the genome-wide SNP performance, whereas the random sets plateaued below 87% accuracy (Table S6, Fig. S9). Notably, Wilcoxon rank-sum tests indicated that the AISNP panels significantly outperformed random SNP sets in both accuracy (p = 0.0031) and AUC (p = 0.003).
Collectively, these results demonstrate that both marker selection and model optimization critically shape ancestry inference. Although LR, SVM, KNN, and CNN each contributed valuable insights, RF, particularly XGBoost, proved most effective, with performance nearly equivalent to that of genome-wide data when 1,000–2,000 informative SNPs were used. The combination of hyperparameter-tuned models with carefully selected AISNP panels enables accurate, balanced, and computationally efficient ancestry classification across diverse East and Southeast Asian populations. These findings establish a practical framework for integrating reduced marker sets with advanced algorithms in forensic and population genomic applications.
Notably, genetic differentiation between certain closely related populations, particularly SC_Tai-Kadai_Sinitic (Class 3) and SEA_Tai-Kadai_Sino-Tibetan (Class 5), remains more difficult to resolve, as reflected by their lower sensitivity metrics. This challenge is attributable to deep historical connections and prolonged gene flow, which have produced a genetic continuum spanning Southeast Asia and South China. This continuum is supported by their close clustering in PCA and t-SNE space (Figs. 2 and 3, S6) and their shared ancestry components in ADMIXTURE analysis (Figs. 2, S7). Although this genetic overlap reduces classification precision, it reflects meaningful biogeographic relationships. Accurately identifying individuals within such admixed contexts still provides substantial forensic value, and future studies may improve resolution by incorporating higher-density or multi-modal markers.
Geographic location prediction via deep neural networks
We applied the Locator model [56], a previously published deep learning framework for geographic inference, to evaluate the ability of reduced AISNP panels to predict sample coordinates. The Locator was trained on nested AISNP panels and genome-wide data (597,569 SNPs), and masked validation samples were used to assess performance. The 2,000-AISNP panel consistently achieved geographic prediction accuracy comparable to that of the genome-wide dataset and, in some cases, achieved comparable performance in some cases (Fig. 5). Additionally, Fig. S10 compares geographic predictions for the same individual (NEA_Tungusic) via Locator models trained on various SNP panels. Our findings indicate that the selected set of 2,000 AISNPs consistently provided robust geographic origin predictions across all tested populations, with predictive performance improving as the number of AISNPs increased. Notably, the 2,000 AISNP model achieved superior accuracy in predicting the geographic origin of individual NEA_Tungusic compared with the model utilizing the full 597,569 SNP dataset. This highlights the effectiveness of our AISNP selection approach for inferring biogeographical ancestry.
Fig. 5.
The geographic coordinates of the nine samples were predicted via 2000 AISNPs as inputs. The blue circles represent the geographic coordinates of the training samples, the black dots represent the predicted geographic coordinates from five repeated predictions, and the red circles indicate the true geographic coordinates of the samples. The contours show the 95%, 50%, and 10% quantiles of a two-dimensional kernel density across windows
The positional errors, which are calculated as the distances between the predicted and actual geographic coordinates, are presented in Fig. 6. These error metrics demonstrate that the Locator models, which are based on selected SNP panels, effectively predict geographic locations. Notably, the model employing 2,000 AISNPs achieved predictive accuracy comparable to that of the genome-wide SNP model, confirming the ability of our AISNP screening method to capture essential biogeographic signals efficiently. This targeted AISNP panel not only improved population differentiation based on geographic region but also significantly reduced computational demands, thereby enhancing the practical application and efficiency of biogeographical prediction analyses.
Fig. 6.
Box plots and histogram of the kernel peak error and centroid error. The red box plot represents the kernel peak error, which is calculated as the Euclidean distance between the true location and the location with the highest kernel density estimated from the predicted coordinates (longitude and latitude). It reflects the accuracy of the most likely predicted location. The green box plot represents the centroid error, which is calculated as the Euclidean distance between the geometric centroid of the predicted locations and the true geographic location. It reflects the overall central tendency of the predictions. The histograms and box plots visualize the distribution of these errors, highlighting the model’s performance in predicting geographic locations
Discussion
Large-scale genomic resources from Eastern Eurasian populations have created new opportunities for advancing forensic applications and population genetics [18, 20]. Resolving population substructures and accurately assigning individuals are critical steps for ancestry inference. Model-based approaches such as STRUCTURE and ADMIXTURE have been widely used for detecting population structure [72], but they assume marker independence and do not account for LD or local ancestry. In this study, ADMIXTURE analyses across K = 2–20, followed by LD pruning, consistently support nine major clusters, with only marginal improvements at K = 10. This robustness was further validated by supervised reclassification, in which RF models reliably reassigned individuals with admixed profiles. This approach successfully classified all 1,703 samples into nine distinct groups, which is generally consistent with previously reported admixture models in the reconstruction of complex demographic histories in population genetic investigations [51]. To ensure robustness, the RF classifier was optimized via 10-fold stratified cross-validation with a grid search across three hyperparameters (mtry, trees, and min_n). The best-performing model (mtry = 5, trees = 100, min_n = 2) achieved the highest mean ROC-AUC and was finalized for reclassification. We selected RF because of its stability with correlated ancestry predictors, resistance to overfitting in moderate-sized datasets, and interpretability of feature importance [73–75]. Nevertheless, alternative classifiers such as SVM, XGBoost, or neural networks may offer complementary advantages. Future work could benchmark these algorithms in a systematic ensemble framework to further improve the classification of highly admixed individuals. Our evaluation of seven nested AISNP panels demonstrated that carefully selected markers can recover fine-scale substructures with high resolution. Panels of 1,500-2,000 SNPs achieved ancestry resolution comparable to that of genome-wide data, as confirmed by PCA, ADMIXTURE, and t-SNE. These findings highlight the efficiency of reduced, information-rich marker sets, which substantially lower genotyping costs while retaining robust inference power. Such panels are particularly valuable in forensic settings, where practical considerations often preclude whole-genome sequencing. We adopted the 50% threshold as a practical and interpretable criterion for initial group assignment. Although not a biologically defined standard, it enabled consistent labeling for model training and downstream classification. Our findings show that high-resolution ancestry and geographic inference can still be achieved even when initial group definitions are based on such simplified thresholds.
We applied six machine learning models (RF, LR, KNN, SVM, CNN, and XGBoost) to evaluate ancestry classification performance across panels containing 50, 250, 500, 1000, 1500, and 2000 AISNPs (Tables S4-5). LR provided stable performance but plateaued as SNP numbers increased, reflecting the inability of linear models to capture non-linear ancestry patterns. The SVM initially performed comparably to the LR, but its accuracy decreased after hyperparameter tuning, likely due to its sensitivity to kernel scaling and parameter selection in high-dimensional genotype data. KNN performed well for certain groups but was inconsistent overall, particularly with genome-wide data where uninformative markers diluted the ancestry signal. This behavior aligns with the algorithm’s reliance on local similarity metrics. The CNN underperformed on smaller panels (< 250 SNPs) but improved substantially with larger inputs and tuning, reflecting its ability to capture local dependencies and hierarchical patterns. However, its computational demands and dependence on large marker sets may limit its utility in forensic applications.
Among the tree-based models, XGBoost consistently delivered the highest overall performance across AISNP panel sizes, achieving 95.6% accuracy and an AUC of 0.999 with 2,000 SNPs. Its advantage was particularly evident at smaller panel sizes (50 − 1,500 SNPs), where it significantly outperformed RF and other classifiers, with accuracy improving from 0.657 (50 SNPs) to 0.956 (2,000 SNPs). RF showed comparable performance when applied to genome-wide data, achieving 97.4% with 597,569 SNPs, highlighting its robustness with high-dimensional inputs. However, XGBoost maintained more balanced sensitivity and specificity across most populations, and its scalability with reduced marker sets made it the most effective classifier in our framework. Both ensemble methods captured complex feature interactions and reduced overfitting, but XGBoost provided superior predictive stability across varying panel sizes. Benchmark analyses further confirmed that XGBoost trained on informative AISNP panels consistently outperformed randomly selected SNP panels in both accuracy and AUC (Wilcoxon rank-sum test, p < 0.01), reinforcing that performance gains were due to marker informativeness rather than chance.
For geographic prediction, we applied the previously published Locator model [56], which directly infers latitude and longitude from unphased genotypes. This study provides several methodological extensions to existing geographic inference frameworks. While we build on the Locator model, we expanded its application in several key ways (Fig. 1). First, we integrate geographic prediction with genetic ancestry classification into a unified pipeline, enabling complementary biogeographic inference. Second, we optimize and evaluate Locator within the complex genetic landscape of East and Southeast Asia, rather than on global datasets. Third, we demonstrate that Locator trained on just 2,000 AISNPs achieves geographic accuracy comparable to genome-wide SNP data, and in some cases, yields even lower positional error. These findings demonstrate the effectiveness of reduced, information-rich marker panels for accurate and cost-efficient geographic inference. However, Locator’s performance depends strongly on the quality and representativeness of the training data. Sampling bias and limited geographic coverage may constrain its broader applicability. Moreover, the model’s computational efficiency comes at the cost of reduced biological interpretability due to its black-box nature. Training and inference with large datasets also remain computationally intensive, and performance may vary across geographic regions [56]. Expanding geographic sampling, incorporating additional marker types (e.g., microhaplotypes, structural variants), and developing hybrid models that combine interpretability with accuracy will further improve geographic prediction [72].
A major challenge in forensic ancestry inference and human origin research is recent admixture signals [51]. Human migration and widespread admixture are key parts of our history [26, 76], making precise ancestry determination from AISNP panels inherently difficult. We highlighted how overlapping ancestry components, recent gene flow, and reference-panel biases limit advanced clustering and assignment techniques in distinguishing mixed populations in forensic biogeographical ancestry inference. Additional factors, such as model misspecification, LD structure, and allele-frequency convergence among ethnolinguistically diverse populations [50], further restrict the use of prune-AISNP panels for broad populations and increase uncertainty in accurate forensic ancestry predictions. One approach, like human genetic research, involves capturing the full spectrum of human genetic diversity and analyzing the variant features within each population [21]. Looking ahead, large-language genome models [77, 78], such as those integrating various deep learning methods, especially transformer-based models that perform local ancestry inference at the haplotype or block level with strong uncertainty estimates, show promise for revealing detailed, locus-specific signals of demographic history. High-density local ancestry inference techniques in population genetics, combined with these computational advances, can deepen our understanding of molecular anthropology and provide valuable insights for precision medicine.
In summary, this study demonstrates that 1,500-2,000 strategically selected AISNPs, when paired with optimized machine learning methods, enable ancestry and geographic inference at near-genome-wide resolution. This integrated framework bridges the gap between large-scale sequencing and practical applications, offering a powerful tool for forensic identification, population history reconstruction, and biogeographic research in East and Southeast Asia. However, our study has several limitations. Although the AISNP panels performed well in this dataset, their generalizability to other populations or broader Asian contexts remains to be tested. Moreover, the Locator model’s performance may degrade in under-sampled or geographically heterogeneous regions due to sparse training data. Future studies should address these limitations by expanding geographic sampling, incorporating diverse marker types, and exploring hybrid modeling strategies that balance interpretability and accuracy.
Conclusion
We demonstrate that 1,500-2,000 strategically selected AISNPs, combined with optimized machine learning models, enable ancestry and geographic inference at near-genome-wide resolution. By systematically evaluating six classifiers and benchmarking against random SNP panels, we show that model choice and marker informativeness are both critical for robust inference. XGBoost consistently achieved the best overall performance, whereas the previously published Locator framework demonstrated that reduced AISNP panels can accurately predict geographic origin. Together, these findings establish a practical and computationally efficient framework for forensic identification, population history reconstruction, and biogeographic research in East and Southeast Asia.
Supplementary Information
Supplementary material 1. Fig. S1 Geographical locations of the collected samples. Fig. S2 Density distribution of AISNPs across 22 chromosomes. Fig. S3 ADMIXTURE plot (K=2-9) of populations after prediction via the Random Forest method based on Genome-wide sequencing data. Fig. S4 Cross-validation (CV) error before and after LD pruning across different values of K. Fig. S5 Distribution of maximum ancestry proportions across 1,703 individuals. The histogram shows the number of individuals according to their maximum ancestry proportion. The dashed vertical lines indicate assignment thresholds of 25% (blue), 50% (red), and 75% (green). Fig. S6 PCA plots for the 7 AISNP sets and genome-wide sequencing data across the nine populations. Fig. S7 Results of PCA and model-based ADMIXTURE clustering analysis (K = 10) using 2000 AISNPs among the nine groups. Each point or bar represents an individual sample. Fig. S8 XGBoost model classification assessment via ROC analysis for 50, 100, 250, 500, 1000, 1500 and 2000 AISNPs. Fig. S9 Comparison of classification performance between AISNP panels and random SNP panels. (A) Accuracy and (C) AUC values of the AISNP panel and randomly selected SNP sets with increasing panel size. (B, D) Boxplots illustrate the distribution of accuracy and AUC values for informative and random SNP panels, respectively. Wilcoxon rank-sum tests were performed to assess statistical significance. Fig. S10 Geographical location prediction of NEA_Tungusic individual using the Locator deep neural network model, based on AISNP panels of 50, 100, 250, 500, 1,000, 1,500, and 2,000 SNPs, as well as the genome-wide dataset (597,569 SNPs).
Supplementary material 2. Table S1 Comparison of ADMIXTURE cross-validation (CV) error before and after LD pruning across K=2-20. ΔCV represents the change in the CV error relative to the previous K.
Supplementary material 3. Table S2 Distribution of maximum ancestry proportions across 1,703 individuals, with cumulative counts under different thresholds.
Supplementary material 4. Table S3 Chromosomal locations of the AISNP markers across different panels. Numbers 1 to 2000 indicate the SNPs with the top 1 to 2000 In values. The 50 AISNP panel consists of SNPs from number 1 to number 50. The 100 AISNP panel consists of SNPs from number 1 to number 100. The 250 AISNP panel consists of SNPs from number 1 to number 250. The 500 AISNP panel consists of SNPs from number 1 to number 500. The 1000 AISNP panel consists of SNPs from number 1 to number 1000. The 2000 AISNP panel consists of SNPs from number 1 to number 2000.
Supplementary material 5. Table S4 Prediction accuracy for populations using six machine learning methods based on 2000/1500/1000/500/250/100/50 AISNPs. Class 1: NC_Sino-Tibetan_Altaic; Class 2: NEA_Tungusic; Class 3: SC_Tai-Kadai_Sinitic; Class 4: SWA_Hmong-Mien; Class 5: SEA_Tai-Kadai_Sino-Tibetan; Class 6: SEA_Austroasiatic; Class 7: SECC_ISEA_Austronesian; Class 8: CA_Turkic; Class 9: WS_Turkic.
Supplementary material 6. Table S5 Sensitivity, specificity, positive predictive value and negative predictive value of 2000/1500/1000/500/250/100/50 AISNPs. Class 1: NC_Sino-Tibetan_Altaic; Class 2: NEA_Tungusic; Class 3: SC_Tai-Kadai_Sinitic; Class 4: SWA_Hmong-Mien; Class 5: SEA_Tai-Kadai_Sino-Tibetan; Class 6: SEA_Austroasiatic; Class 7: SECC_ISEA_Austronesian; Class 8: CA_Turkic; Class 9: WS_Turkic.
Supplementary material 7. Table S6 Performance comparison of AISNP panels and random SNP panels.
Acknowledgements
We would like to thank BioRender for providing the drawing materials.
Author contributions
J.Y. and G.H. conceived and supervised the project. G.H., M.W. and H.F. collected and integrated the genomic datasets. J.C. and Y.H. analyzed the data and wrote the draft of the manuscript. J.Y. and G.H. revised the manuscript. All the authors read and approved the final manuscript.
Funding
This work was supported by the National Natural Science Foundation of China (82402203, 82030058, and 82202078), the Major Project of the National Social Science Foundation of China (23&ZD203), the Open Project of the Key Laboratory of Forensic Genetics of the Ministry of Public Security (2022FGKFKT05), the Center for Archaeological Science of Sichuan University (23SASA01), and the Sichuan Science and Technology Program (2024NSFSC1518).
Data availability
No datasets were generated or analysed during the current study.
Declarations
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Haoliang Fan, Email: fanhaoliang198931@163.com.
Mengge Wang, Email: Menggewang2021@163.com.
Guanglin He, Email: guanglinhescu@163.com.
Jiangwei Yan, Email: yanjw@sxmu.edu.cn.
References
- 1.Jeong C, Wang K, Wilkin S, Taylor WTT, Miller BK, Bemmann JH, Stahl R, Chiovelli C, Knolle F, Ulziibayar S, et al. A dynamic 6,000-Year genetic history of eurasia’s Eastern steppe. Cell. 2020;183(4):890–e904829. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Flegontov P, Altinisik NE, Changmai P, Rohland N, Mallick S, Adamski N, Bolnick DA, Broomandkhoshbacht N, Candilio F, Culleton BJ, et al. Palaeo-Eskimo genetic ancestry and the peopling of Chukotka and North America. Nature. 2019;570(7760):236–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Mao X, Zhang H, Qiao S, Liu Y, Chang F, Xie P, Zhang M, Wang T, Li M, Cao P, et al. The deep population history of Northern East Asia from the late pleistocene to the holocene. Cell. 2021;184(12):3256–e32663213. [DOI] [PubMed] [Google Scholar]
- 4.Wang T, Wang W, Xie G, Li Z, Fan X, Yang Q, Wu X, Cao P, Liu Y, Yang R, et al. Human population history at the crossroads of East and Southeast Asia since 11,000 years ago. Cell. 2021;184(14):3829–41. e3821. [DOI] [PubMed] [Google Scholar]
- 5.Robbeets M, Bouckaert R, Conte M, Savelyev A, Li T, An DI, Shinoda KI, Cui Y, Kawashima T, Kim G, et al. Triangulation supports agricultural spread of the Transeurasian languages. Nature. 2021;599(7886):616–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.GenomeAsia KC. The genomeasia 100K project enables genetic discoveries across Asia. Nature. 2019;576(7785):106–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Kerdoncuff E, Skov L, Patterson N, Banerjee J, Khobragade P, Chakrabarti SS, Chakrawarty A, Chatterjee P, Dhar M, Gupta M, et al. 50,000 years of evolutionary history of india: impact on health and disease variation. Cell. 2025;188(13):3389–404. e3386. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.He Y, Zhang X, Peng MS, Li YC, Liu K, Zhang Y, Mao L, Guo Y, Ma Y, Zhou B, et al. Genome diversity and signatures of natural selection in Mainland Southeast Asia. Nature. 2025;643(8071):417–26. [DOI] [PubMed] [Google Scholar]
- 9.He G, Yao H, Duan S, Luo L, Sun Q, Tang R, Chen J, Wang Z, Sun Y, Li X, et al. Pilot work of the 10K Chinese people genomic diversity project along the silk road suggests a complex east-west admixture landscape and biological adaptations. Sci China Life Sci. 2025;68(4):914–33. [DOI] [PubMed] [Google Scholar]
- 10.Wang M, Huang Y, Liu K, Wang Z, Zhang M, Yuan H, Duan S, Wei L, Yao H, Sun Q et al. Multiple human population movements and cultural dispersal events shaped the landscape of Chinese paternal heritage. Mol Biol Evol. 2024;41(7):msae122. [DOI] [PMC free article] [PubMed]
- 11.Wang Z, Liu K, Yuan H, Duan S, Liu Y, Luo L, Jiang X, Chen S, Wei L, Tang R et al. YanHuang Paternal Genomic Resource Suggested A Weakly-Differentiated Multi-Source Admixture Model for the Formation of Han’s Founding Ancestral Lineages. Genomics, Proteomics & Bioinformatics. 2025;4:qzaf049. [DOI] [PMC free article] [PubMed]
- 12.Cheng S, Xu Z, Bian S, Chen X, Shi Y, Li Y, Duan Y, Liu Y, Lin J, Jiang Y, et al. The STROMICS genome study: deep whole-genome sequencing and analysis of 10K Chinese patients with ischemic stroke reveal complex genetic and phenotypic interplay. Cell Discovery. 2023;9(1):75. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Walters RG, Millwood IY, Lin K, Schmidt Valle D, McDonnell P, Hacker A, Avery D, Edris A, Fry H, Cai N, et al. Genotyping and population characteristics of the China kadoorie biobank. Cell Genom. 2023;3(8):100361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Zhang P, Luo H, Li Y, Wang Y, Wang J, Zheng Y, Niu Y, Shi Y, Zhou H, Song T, et al. NyuWa genome resource: A deep whole-genome sequencing-based variation profile and reference panel for the Chinese population. Cell Rep. 2021;37(7):110017. [DOI] [PubMed] [Google Scholar]
- 15.Cao Y, Li L, Xu M, Feng Z, Sun X, Lu J, Xu Y, Du P, Wang T, Hu R, et al. The ChinaMAP analytics of deep whole genome sequences in 10,588 individuals. Cell Res. 2020;30(9):717–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Huang S, Liu S, Huang M, He JR, Wang C, Wang T, Feng X, Kuang Y, Lu J, Gu Y, et al. The born in Guangzhou cohort study enables generational genetic discoveries. Nature. 2024;626(7999):565–73. [DOI] [PubMed] [Google Scholar]
- 17.Zhou D, Wu M, Tan Q, Sun L, Tu Y, Zheng W, Zhu Y, Yang M, Hu K, Hu F, et al. Non-coding genetic elements of lung cancer identified using whole genome sequencing in 13,722 Chinese. Nat Commun. 2025;16(1):7365. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Jiang T, Guo H, Liu Y, Li G, Cui Z, Cui X, Liu Y, Li Y, Zhang A, Cao S, et al. A comprehensive genetic variant reference for the Chinese population. Sci Bull (Beijing). 2024;69(24):3820–5. [DOI] [PubMed] [Google Scholar]
- 19.Yang MY, Zhong JD, Li X, Tian G, Bai WY, Fang YH, Qiu MC, Yuan CD, Yu CF, Li N, et al. SEAD reference panel with 22,134 haplotypes boosts rare variant imputation and genome-wide association analysis in Asian populations. Nat Commun. 2024;15(1):10839. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wang M, Duan S, Li X, Yang J, Yuan H, Liu C, He G. Genome-driven Chinese precision medicine: Biobank-scale genomic research as a new paradigm. Innov Life. 2025.
- 21.Wang M, Luo L, Yeh H-Y, Wang C-C, Yuan H, Liu C, Tang R, He G. Mighty oaks from little acorns: High-quality genomes of underrepresented populations enhance health equity in precision medicine. hLife. 2025. 10.1016/j.hlife.2025.05.014IF.
- 22.Yang Q, Duan S, Huang Y, Liu C, Wang M, He G. The large-scale whole-genome sequencing era expedited medical discovery and clinical translation. EngMedicine. 2025;2(1):100055. [Google Scholar]
- 23.He Y, Zhang X, Peng M-S, Li Y, Liu K, Zhang Y, Mao L, Guo Y, Ma Y, Zhang Y et al. Genome diversity and signatures of natural selection in mainland Southeast Asia. Nature. 2025;643:417–26. [DOI] [PubMed]
- 24.He G, Sun Y, Duan S, Luo L, Sun Q, Li B, Yun L, Liu C, Wang M. Ancient genomes give insight into 160,000 years of East Asian population dynamics and biological adaptation. Genome Biol 2025.
- 25.Yang MA, Fu Q. Insights into modern human prehistory using ancient genomes. Trends Genet. 2018;34(3):184–96. [DOI] [PubMed] [Google Scholar]
- 26.Liu Y, Mao X, Krause J, Fu Q. Insights into human history from the first decade of ancient human genomics. Science. 2021;373(6562):1479–84. [DOI] [PubMed] [Google Scholar]
- 27.Wang Z, Liu K, Yuan H, Duan S, Liu Y, Luo L, Jiang X, Chen S, Wei L-H, Tang R, et al. YanHuang paternal genomic resource suggested a Weakly-Differentiated Multi-Source Admixture model for the formation of Han’s founding ancestral lineages. Genomics, Proteomics & Bioinformatics. 2025;4:qzaf049. [DOI] [PMC free article] [PubMed]
- 28.McColl H, Racimo F, Vinner L, Demeter F, Gakuhari T, Moreno-Mayar JV, van Driem G, Gram Wilken U, Seguin-Orlando A, de la Fuente Castro C, et al. The prehistoric peopling of Southeast Asia. Science. 2018;361(6397):88–92. [DOI] [PubMed] [Google Scholar]
- 29.Lipson M, Cheronet O, Mallick S, Rohland N, Oxenham M, Pietrusewsky M, Pryce TO, Willis A, Matsumura H, Buckley H, et al. Ancient genomes document multiple waves of migration in Southeast Asian prehistory. Science. 2018;361(6397):92–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.He G, Wang Z, Wang M, Luo T, Liu J, Zhou Y, Gao B, Hou Y. Forensic ancestry analysis in two Chinese minority populations using massively parallel sequencing of 165 ancestry-informative SNPs. Electrophoresis. 2018;39(21):2732–42. [DOI] [PubMed] [Google Scholar]
- 31.Wang Z, He G, Luo T, Zhao X, Liu J, Wang M, Zhou D, Chen X, Li C, Hou Y. Massively parallel sequencing of 165 ancestry informative SNPs in two Chinese Tibetan-Burmese minority ethnicities. Forensic Sci Int Genet. 2018;34:141–7. [DOI] [PubMed] [Google Scholar]
- 32.Wang M, Chen H, Luo L, Huang Y, Duan S, Yuan H, Tang R, Liu C, He G. Forensic investigative genetic genealogy: expanding pedigree tracing and genetic inquiry in the genomic era. J Genet Genomics. 2025;52(4):460–72. [DOI] [PubMed] [Google Scholar]
- 33.He G, Liu C, Wang M. Perspectives and opportunities in forensic human, animal, and plant integrative genomics in the pangenome era. Forensic Sci Int. 2025;367:112370. [DOI] [PubMed] [Google Scholar]
- 34.Li CX, Pakstis AJ, Jiang L, Wei YL, Sun QF, Wu H, Bulbul O, Wang P, Kang LL, Kidd JR, et al. A panel of 74 aisnps: improved ancestry inference within Eastern Asia. Forensic Sci Int Genet. 2016;23:101–10. [DOI] [PubMed] [Google Scholar]
- 35.Shi CM, Liu Q, Zhao S, Chen H. Ancestry informative SNP panels for discriminating the major East Asian populations: Han Chinese, Japanese and Korean. Ann Hum Genet. 2019;83(5):348–54. [DOI] [PubMed] [Google Scholar]
- 36.Yahya P, Sulong S, Harun A, Wangkumhang P, Wilantho A, Ngamphiw C, Tongsima S, Zilfalil BA. Ancestry-informative marker (AIM) SNP panel for the Malay population. Int J Legal Med. 2020;134(1):123–34. [DOI] [PubMed] [Google Scholar]
- 37.Chen L, Zhou Z, Zhang Y, Xu H, Wang S. EASplex: A panel of 308 AISNPs for East Asian ancestry inference using next generation sequencing. Forensic Sci Int Genet. 2022;60:102739. [DOI] [PubMed] [Google Scholar]
- 38.Cao Y, Zhu Q, Huang Y, Li X, Wei Y, Wang H, Zhang J. An efficient ancestry informative SNPs panel for further discriminating East Asian populations. Electrophoresis. 2022;43(16–17):1774–83. [DOI] [PubMed] [Google Scholar]
- 39.Phillips C, de la Puente M, Ruiz-Ramirez J, Staniewska A, Ambroa-Conde A, Freire-Aradas A, Mosquera-Miguel A, Rodriguez A, Lareu MV. Eurasiaplex-2: shifting the focus to SNPs with high population specificity increases the power of forensic ancestry marker sets. Forensic Sci Int Genet. 2022;61:102780. [DOI] [PubMed] [Google Scholar]
- 40.Wang C, Wang S, Zhao Y, Liu J, Zhang D, Wang F, Fan H, Li C, Jiang L. A biogeographical ancestry inference pipeline using PCA-XGBoost model and its application in Asian populations. Forensic Sci Int Genet. 2025;77:103239. [DOI] [PubMed] [Google Scholar]
- 41.Chen J, Huang Y, Zhong J, Wang M, He G, Yan J. Bioinformatic insights into five Chinese population substructures inferred from the East Asian-specific AISNP panel. BMC Genomics. 2025;26(1):748. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Elhaik E, Tatarinova T, Chebotarev D, Piras IS, Maria Calò C, De Montis A, Atzori M, Marini M, Tofanelli S, Francalacci P, et al. Geographic population structure analysis of worldwide human populations infers their biogeographical origins. Nat Commun. 2014;5:3513. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Verma A, Huffman JE, Rodriguez A, Conery M, Liu M, Ho YL, Kim Y, Heise DA, Guare L, Panickan VA, et al. Diversity and scale: genetic architecture of 2068 traits in the VA million veteran program. Science. 2024;385(6706):eadj1182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.All of Us Research Program Genomics I. Genomic data in the all of Us research program. Nature. 2024;627(8003):340–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Rubinacci S, Hofmeister RJ, Sousa da Mota B, Delaneau O. Imputation of low-coverage sequencing data from 150,119 UK biobank genomes. Nat Genet. 2023;55(7):1088–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R, Nagulapalli K, et al. High-coverage whole-genome sequencing of the expanded 1000 genomes project cohort including 602 trios. Cell. 2022;185(18):3426–40. e3419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Taliun D, Harris DN, Kessler MD, Carlson J, Szpiech ZA, Torres R, Taliun SAG, Corvelo A, Gogarten SM, Kang HM, et al. Sequencing of 53,831 diverse genomes from the NHLBI topmed program. Nature. 2021;590(7845):290–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Bergstrom A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science. 2020;367(6484):1339. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.He G, Yao H, Duan S, Luo L, Sun Q, Tang R, Chen J, Wang Z, Sun Y, Li X et al. Pilot work of the 10K Chinese People Genomic Diversity Project along the Silk Road suggests a complex east-west admixture landscape and biological adaptations. Science China Life sciences. 2025;68(4):914-33. [DOI] [PubMed]
- 50.Yang Q, Sun Y, Duan S, Nie S, Liu C, Deng H, Wang M, He G. High-quality Population-specific Haplotype-resolved Reference Panel in the Genomic and Pangenomic Eras. Genomics, Proteomics & Bioinformatics. 2025;10:qzaf022. [DOI] [PMC free article] [PubMed]
- 51.He G, Wang M, Luo L, Sun Q, Yuan H, Lv H, Feng Y, Liu X, Cheng J, Bu F, et al. Population genomics of central Asian peoples unveil ancient Trans-Eurasian genetic admixture and cultural exchanges. hLife. 2024;2(11):554–62. [Google Scholar]
- 52.Luo L, Wang M, Liu Y, Li J, Bu F, Yuan H, Tang R, Liu C, He G. Sequencing and characterizing human mitochondrial genomes in the biobank-based genomic research paradigm. Sci China Life Sci. 2025;68(6):1610-25. [DOI] [PubMed]
- 53.Bose A, Platt DE, Parida L, Drineas P, Paschou P. Integrating Linguistics, social Structure, and geography to model genetic diversity within India. Mol Biol Evol. 2021;38(5):1809–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Eraslan G, Avsec Z, Gagneur J, Theis FJ. Deep learning: new computational modelling techniques for genomics. Nat Rev Genet. 2019;20(7):389–403. [DOI] [PubMed] [Google Scholar]
- 55.Huang X, Rymbekova A, Dolgova O, Lao O, Kuhlwilm M. Harnessing deep learning for population genetic inference. Nat Rev Genet. 2024;25(1):61–78. [DOI] [PubMed] [Google Scholar]
- 56.Battey CJ, Ralph PL, Kern AD. Predicting geographic location from genetic variation with deep neural networks. eLife. 2020;9:e54507. [DOI] [PMC free article] [PubMed]
- 57.Barash M, McNevin D, Fedorenko V, Giverts P. Machine learning applications in forensic DNA profiling: A critical review. Forensic Sci Int Genet. 2024;69:102994. [DOI] [PubMed] [Google Scholar]
- 58.Heinzel CS, Purucker L, Hutter F, Pfaffelhuber P. Advancing biogeographical ancestry predictions through machine learning. Forensic Sci Int Genet. 2025;79:103290. [DOI] [PubMed] [Google Scholar]
- 59.Kloska A, Giełczyk A, Grzybowski T, Płoski R, Kloska SM, Marciniak T, Pałczyński K, Rogalla-Ładniak U, Malyarchuk BA, Derenko MV et al. A Machine-Learning-Based approach to prediction of biogeographic ancestry within Europe. Int J Mol Sci. 2023;24(20):15095. [DOI] [PMC free article] [PubMed]
- 60.Alladio E, Poggiali B, Cosenza G, Pilli E. Multivariate statistical approach and machine learning for the evaluation of biogeographical ancestry inference in the forensic field. Sci Rep. 2022;12(1):8974. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Mallick S, Micco A, Mah M, Ringbauer H, Lazaridis I, Olalde I, Patterson N, Reich D. The Allen ancient DNA resource (AADR) a curated compendium of ancient human genomes. Sci Data. 2024;11(1):182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Liu D, Duong NT, Ton ND, Van Phong N, Pakendorf B, Van Hai N, Stoneking M. Extensive ethnolinguistic diversity in Vietnam reflects multiple sources of genetic diversity. Mol Biol Evol. 2020;37(9):2503–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Wang CC, Yeh HY, Popov AN, Zhang HQ, Matsumura H, Sirak K, Cheronet O, Kovalev A, Rohland N, Kim AM, et al. Genomic insights into the formation of human populations in East Asia. Nature. 2021;591(7850):413–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26(22):2867–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Alexander DH, Novembre J, Lange K. Fast model-based Estimation of ancestry in unrelated individuals. Genome Res. 2009;19(9):1655–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Daya M, van der Merwe L, Galal U, Möller M, Salie M, Chimusa ER, Galanter JM, van Helden PD, Henn BM, Gignoux CR, et al. A panel of ancestry informative markers for the complex five-way admixed South African coloured population. PLoS ONE. 2013;8(12):e82224. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Rosenberg NA, Li LM, Ward R, Pritchard JK. Informativeness of genetic markers for inference of ancestry. Am J Hum Genet. 2003;73(6):1402–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Patterson N, Price AL, Reich D. Population structure and eigenanalysis. PLoS Genet. 2006;2(12):e190. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Price AL, Patterson NJ, Plenge RM, Weinblatt ME, Shadick NA, Reich D. Principal components analysis corrects for stratification in genome-wide association studies. Nat Genet. 2006;38(8):904–9. [DOI] [PubMed] [Google Scholar]
- 71.Maaten, Lvd. Hinton gejjomlr: visualizing data using t-SNE. 2008, 9:2579–605.
- 72.Liu X, Koyama S, Tomizuka K, Takata S, Ishikawa Y, Ito S, Kosugi S, Suzuki K, Hikino K, Koido M, et al. Decoding triancestral origins, archaic introgression, and natural selection in the Japanese population by whole-genome sequencing. Sci Adv. 2024;10(16):eadi8419. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Hallett MJ, Fan JJ, Su XG, Levine RA, Nunn ME. Random forest and variable importance rankings for correlated survival data, with applications to tooth loss. STAT MODEL. 2014;14(6):523–47.
- 74.Strobl C, Boulesteix AL, Zeileis A, Hothorn T. Bias in random forest variable importance measures: illustrations, sources and a solution. BMC Bioinformatics. 2007;8:25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Halabaku E, Bytyçi E. Overfitting in machine learning: A comparative analysis of decision trees and random forests. 2024, 39(6):987–1006.
- 76.Avila-Arcos MC, Raghavan M, Schlebusch C. Going local with ancient DNA: A review of human histories from regional perspectives. Science. 2023;382(6666):53–8. [DOI] [PubMed] [Google Scholar]
- 77.Consens ME, Dufault C, Wainberg M, Forster D, Karimzadeh M, Goodarzi H, Theis FJ, Moses A, Wang B. Transformers and genome Language models. Nat Mach Intell. 2025;7(3):346–62. [Google Scholar]
- 78.Benegas G, Ye C, Albors C, Li JC, Song YS. Genomic Language models: opportunities and challenges. Trends Genet. 2025;41(4):286–302. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary material 1. Fig. S1 Geographical locations of the collected samples. Fig. S2 Density distribution of AISNPs across 22 chromosomes. Fig. S3 ADMIXTURE plot (K=2-9) of populations after prediction via the Random Forest method based on Genome-wide sequencing data. Fig. S4 Cross-validation (CV) error before and after LD pruning across different values of K. Fig. S5 Distribution of maximum ancestry proportions across 1,703 individuals. The histogram shows the number of individuals according to their maximum ancestry proportion. The dashed vertical lines indicate assignment thresholds of 25% (blue), 50% (red), and 75% (green). Fig. S6 PCA plots for the 7 AISNP sets and genome-wide sequencing data across the nine populations. Fig. S7 Results of PCA and model-based ADMIXTURE clustering analysis (K = 10) using 2000 AISNPs among the nine groups. Each point or bar represents an individual sample. Fig. S8 XGBoost model classification assessment via ROC analysis for 50, 100, 250, 500, 1000, 1500 and 2000 AISNPs. Fig. S9 Comparison of classification performance between AISNP panels and random SNP panels. (A) Accuracy and (C) AUC values of the AISNP panel and randomly selected SNP sets with increasing panel size. (B, D) Boxplots illustrate the distribution of accuracy and AUC values for informative and random SNP panels, respectively. Wilcoxon rank-sum tests were performed to assess statistical significance. Fig. S10 Geographical location prediction of NEA_Tungusic individual using the Locator deep neural network model, based on AISNP panels of 50, 100, 250, 500, 1,000, 1,500, and 2,000 SNPs, as well as the genome-wide dataset (597,569 SNPs).
Supplementary material 2. Table S1 Comparison of ADMIXTURE cross-validation (CV) error before and after LD pruning across K=2-20. ΔCV represents the change in the CV error relative to the previous K.
Supplementary material 3. Table S2 Distribution of maximum ancestry proportions across 1,703 individuals, with cumulative counts under different thresholds.
Supplementary material 4. Table S3 Chromosomal locations of the AISNP markers across different panels. Numbers 1 to 2000 indicate the SNPs with the top 1 to 2000 In values. The 50 AISNP panel consists of SNPs from number 1 to number 50. The 100 AISNP panel consists of SNPs from number 1 to number 100. The 250 AISNP panel consists of SNPs from number 1 to number 250. The 500 AISNP panel consists of SNPs from number 1 to number 500. The 1000 AISNP panel consists of SNPs from number 1 to number 1000. The 2000 AISNP panel consists of SNPs from number 1 to number 2000.
Supplementary material 5. Table S4 Prediction accuracy for populations using six machine learning methods based on 2000/1500/1000/500/250/100/50 AISNPs. Class 1: NC_Sino-Tibetan_Altaic; Class 2: NEA_Tungusic; Class 3: SC_Tai-Kadai_Sinitic; Class 4: SWA_Hmong-Mien; Class 5: SEA_Tai-Kadai_Sino-Tibetan; Class 6: SEA_Austroasiatic; Class 7: SECC_ISEA_Austronesian; Class 8: CA_Turkic; Class 9: WS_Turkic.
Supplementary material 6. Table S5 Sensitivity, specificity, positive predictive value and negative predictive value of 2000/1500/1000/500/250/100/50 AISNPs. Class 1: NC_Sino-Tibetan_Altaic; Class 2: NEA_Tungusic; Class 3: SC_Tai-Kadai_Sinitic; Class 4: SWA_Hmong-Mien; Class 5: SEA_Tai-Kadai_Sino-Tibetan; Class 6: SEA_Austroasiatic; Class 7: SECC_ISEA_Austronesian; Class 8: CA_Turkic; Class 9: WS_Turkic.
Supplementary material 7. Table S6 Performance comparison of AISNP panels and random SNP panels.
Data Availability Statement
No datasets were generated or analysed during the current study.







