Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Feb 6;6(2):e70321. doi: 10.1002/cpz1.70321

Strategies in Global Ancestry and Local Ancestry Inference

Bilcag Akgun 1, Farid Rajabli 1,2,
PMCID: PMC12880525  PMID: 41649483

Abstract

Genetic ancestry inference has become essential in population and medical genetics, especially for studies of admixed populations. Accurate determination of both global ancestry (GA) proportions and local ancestry (LA) segmental origins requires careful selection of computational methods and reference panels. Here, we present a practical, protocol‐oriented guide that (i) clarifies key concepts (GA vs. LA, reference panel selection, phasing requirements), (ii) organizes methods into model‐based clustering and dimensionality‐reduction approaches for GA and hidden Markov model–based, window‐based machine learning, and deep learning frameworks for LA and (iii) provides concise guidance on tool selection for GA and LA. Step‐by‐step protocols are provided for a typical ADMIXTURE‐based GA analysis and for a SHAPEIT5 + RFMix LA inference pipeline, with practical considerations for genotype array and whole‐genome sequencing data. We also discuss quality control, method validation, and downstream applications of ancestry inference. Finally, we address current challenges and highlight recent advances, including fast algorithms, deep learning models, improved phasing, and integrative tools. This guide aims to help researchers select and implement appropriate ancestry inference methods for diverse study designs and datasets. © 2026 The Author(s). Current Protocols published by Wiley Periodicals LLC.

Basic Protocol 1: Global ancestry analysis (ADMIXTURE pipeline)

Basic Protocol 2: Local ancestry analysis (phasing + RFMix pipeline)

Keywords: ancestry inference, global ancestry, local ancestry, population genetics, population structure

INTRODUCTION

Over the past two decades, advancements in high‐throughput genotyping and sequencing technology have expanded our understanding of human population genetics and demographic history (see Current Protocols article: Slatko et al., 2018). The study of genetic ancestry has progressed from basic population clustering methods to advanced approaches that can trace genomic contributions of ancestral populations at high resolution across the genome. These methodological advancements are particularly important for studying admixed populations, whose genomes are mosaics of segments from multiple ancestral sources due to historical migration and mixing events (Atkin et al., 2022). Ancestry inference now plays multiple roles in genomics research. In genetic association analyses, ancestry information is essential for controlling population stratification in genome‐wide association studies (GWASs) and for improving the accuracy of disease‐associated variant discovery (Tan & Atkinson, 2023). In precision medicine, ancestry‐specific allele frequency and effect size differences influence risk prediction and drug responses (Krainc & Fuentes, 2022). In evolutionary genetics, ancestral patterns elucidate natural selection, adaptation, and historical demographic events (Racimo et al., 2015).

Two complementary methods exist for inferring genetic ancestry in admixed individuals: global ancestry (GA) and local ancestry (LA) (Padhukasahasram, 2014). GA inference estimates the overall proportion of each ancestral population across an individual's genome. In contrast, LA inference assigns an ancestry label to genomic segments along each chromosome, identifying which chromosomal segments were inherited from which ancestral populations, thereby providing a high‐resolution view of the mosaic structure of admixed genomes. GA methods are widely used to capture broad population structure and are commonly included as covariates in genetic association studies. LA allow for ancestry‐aware analyses, including admixture mapping, fine‐mapping in admixed genomes, and tests for recent selection on particular lineages. LA inference is crucial for downstream analyses that depend on ancestry at specific genomic loci, as individuals with similar GA proportions may still have very different LA profiles (Tan & Atkinson, 2023).

Despite significant progress, important challenges remain. Biobank‐scale datasets including hundreds of thousands of genomes require approaches that are both accurate and computationally scalable (Peterson et al., 2019). Many populations (e.g., Amerindian, certain African, and South Asian groups) remain underrepresented in reference panels, limiting inference accuracy and equity (Bien et al., 2019; Fatumo et al., 2022). Methods must also account for complex demographic scenarios, such as multi‐wave or ancient admixture events, which can violate standard model assumptions. Here, we provide a comprehensive guide to selecting and implementing ancestry inference methods, with an emphasis on admixed populations and practical protocols for commonly used tools.

Global Ancestry Inference

Overview

GA inference (Basic Protocol 1) estimates the genome‐wide ancestry proportions of an individual or population (Alexander et al., 2009). Most methods model the genome as a mixture of contributions from K ancestral groups, using genotype data to infer these proportions (Pritchard et al., 2000). GA estimates are widely used to capture population structure and are routinely included as covariates in downstream analyses, for example, to correct for population stratification in GWASs.

Method categories

Model‐based clustering approaches

Model‐based methods assign individuals probabilistically to clusters (populations) under assumptions of Hardy‐Weinberg and linkage equilibrium within clusters (Falush et al., 2003). STRUCTURE uses a Bayesian framework with Markov chain Monte Carlo (MCMC) to jointly estimate ancestry proportions and allele frequencies for K populations (Pritchard et al., 2000). STRUCTURE provides a full posterior distribution of ancestry for each individual, offering measures of confidence. However, it is computationally intensive for large single nucleotide polymorphism (SNP) datasets, and the standard STRUCTURE implementation assumes loci are unlinked.

ADMIXTURE is a widely used faster alternative that employs maximum‐likelihood estimation via an expectation‐maximization (EM) algorithm (Alexander et al., 2009). It converges more quickly than STRUCTURE by sacrificing the full Bayesian posterior in favor of point estimates of ancestry proportions. ADMIXTURE can handle large genotype datasets efficiently and includes a cross‐validation (CV) procedure to help choose the optimal number of clusters K. It assumes individuals are unrelated and that populations are in Hardy‐Weinberg equilibrium, and like STRUCTURE, it can be run in supervised mode if reference population labels are provided.

fastSTRUCTURE further accelerates model‐based inference by using a variational Bayesian framework with a sparsity prior (Raj et al., 2014). It automates the choice of K by maximizing model evidence. fastSTRUCTURE is substantially faster and more memory efficient than STRUCTURE for large datasets, though it provides approximate rather than fully Bayesian inference and may be less flexible in modeling complex demographic histories.

Dimensionality reduction approaches

Unsupervised principal component (PC) analysis (PCA) is another fundamental approach for characterizing GA and population structure (Patterson et al., 2006). PCA decomposes the genomic relationship matrix (GRM) into PCs that capture the major axes of genetic variation within a dataset. The top PCs often correspond to geographic or ancestral gradients (Sankararaman et al., 2008), enabling intuitive visualization of population structure. PCA does not require explicit population model assumptions and is extremely scalable, making it a standard tool for identifying clusters and continuous clines of ancestry and for correcting population stratification in association studies. However, PCA is an indirect method for ancestry inference: individual PC coordinates can be converted to ancestry estimates by projection or by correlating with reference groups, but interpretation can be challenging when multiple admixture events or nonlinear population structure is present.

Performance

GA approaches are often highly accurate for distinguishing major continental‐level ancestries (Martin et al., 2017). However, their resolution is more limited for closely related groups, especially when reference populations are genetically similar or when the true source population is not represented in the reference panel, which can lead to misassigned ancestry proportions (Hellenthal et al., 2014). Additionally, GA assumes each individual can be summarized by a single set of proportions averaged across the genome, which may not capture recent recombination events or locus‐specific ancestry differences. Despite these limitations, GA analysis is still an important first step in studies of admixed populations, informing sample selection, identifying population structure, and providing covariates to avoid confounding in association studies.

Key global ancestry methods

Table 1 summarizes several widely used GA inference methods, with their strengths and limitations.

Table 1.

Global Ancestry Inference Methods

Method Algorithm type Strengths Limitations
STRUCTURE Bayesian MCMC Provides posterior probability distribution for admixture estimates and quantifies uncertainty Computationally intensive (slow) for large datasets; assumes unlinked markers (requires LD pruning)
ADMIXTURE Maximum‐likelihood (EM algorithm) Very fast convergence; can handle genome‐wide SNP data and includes cross‐validation to choose optimal number of clusters Provides point estimates only (no posterior probabilities); assumes independence between markers (sensitive to LD if SNPs not pruned)
fastSTRUCTURE Variational Bayes More scalable than STRUCTURE for large sample sizes and high SNP counts; can automatically suggest the best number of ancestral populations K Uses an approximation, which may be less accurate for complex admixture scenarios; fewer modeling options (less flexible) than full Bayesian approaches
PCA Eigen‐decomposition (unsupervised) Extremely fast even on large datasets; makes no explicit population assumptions; captures continuous gradients of variation Not a direct admixture percentage method; interpreting PCs as ancestry fractions is indirect; results can be influenced by outlier individuals or uneven sampling

Local Ancestry Inference

Overview

LA inference (Basic Protocol 2) aims to determine the ancestry of each segment of an admixed genome, essentially painting each locus with an ancestry label (e.g. European, African, or Amerindian) according to the source population from which it originated (Loh et al., 2013). This is a finer‐grained analysis than GA, and it requires methods that model linkage disequilibrium (LD) patterns and recombination. Phased haplotypes are typically required as input for LA methods because distinguishing ancestries along a chromosome is easier when the sequence of alleles on each chromosome copy is known. Compared to GA, LA analysis is more computationally intensive, as it often employs sliding‐window approaches or hidden Markov models (HMMs) that scan millions of markers and explicitly model transitions between ancestry states across the genome (Sankararaman et al., 2008).

Method categories

HMM‐based methods

HMM‐based approaches model the unknown ancestry sequence along the genome as a latent state sequence (Sankararaman et al., 2008; Tang et al., 2005). These methods explicitly model recombination events as transitions between ancestry states, often assuming a certain number of generations since admixture to calibrate transition probabilities (Price et al., 2009). Some early examples include LAMP (Local Ancestry in adMixed Populations), which used an HMM for two ancestral populations (Sankararaman et al., 2008), and HAPMIX, which used an HMM for two‐way admixture that modeled allele frequencies of each ancestor (Price et al., 2009). Later, LAMP‐LD extended LAMP to handle multiple (>2) ancestral populations and to incorporate LD by analyzing SNP windows within an HMM (Baran et al., 2012). These HMM‐based approaches are statistically robust and can provide confidence estimates, but they can be slow and may require specification of parameters such as admixture timing. Recent HMM‐based tools aim to improve speed and scalability. For example, FLARE (Fast Local Ancestry in the Rare Admixed Ensemble) uses an efficient haplotype compression strategy to reduce the state space, allowing for much faster biobank‐scale LA inference while maintaining accuracy (Browning et al., 2023). Recomb‐Mix is another advanced method that uses a graph‐based optimization within an HMM framework, which collapses similar ancestry paths to speed up computation and has shown high accuracy across diverse scenarios (Wei et al., 2025).

Graph/dynamic programming methods

These approaches aim to identify an optimal path of ancestry assignments across the genome, often without specifying a full probabilistic generative model. For instance, Loter formulates LA inference as a graph path–finding problem: it builds a graph in which nodes represent possible ancestry states at each position and edges represent switches between ancestries, and then it finds the path that best matches the observed genotypes to reference haplotypes (Dias‐Alves et al., 2018). Loter's algorithm is deterministic and does not require specification of a genetic map or admixture‐generation parameters, making it relatively robust and applicable even beyond human data. However, purely combinatorial solutions of this type may not capture uncertainty as well as HMM‐based approaches and may perform less well when ancestry patterns are complex and would benefit from an explicit statistical model. Additionally, some methods, such as Recomb‐Mix, blend HMM‐ and graph‐based ideas (Wei et al., 2025).

Window‐based machine learning methods

These approaches first divide the genome into small windows of contiguous SNPs, classify the ancestry of each window using a supervised classifier trained on reference haplotypes, and then apply a smoothing step to correct inconsistencies between adjacent windows (Maples et al., 2013). RFMix is a prime example: it trains a random forest on reference haplotypes to predict window ancestry and then uses a conditional random field (CRF) to smooth these predictions along the chromosome (Maples et al., 2013). This combination yields fast and robust ancestry predictions, and RFMix v2 further improves accuracy by incorporating genetic map information and tuning model parameters. Gnomix is another window‐based approach, which first assigns preliminary ancestry calls via logistic regression and then refines them with a gradient‐boosted trees model that learns recombination patterns (Hilmarsson et al., 2021). Such methods often perform particularly well for recently admixed individuals and can accommodate multiple ancestry groups with relatively lower computational cost than full HMM‐based methods (Hilmarsson et al., 2021).

Deep learning methods

A few recent LA tools use neural networks to directly model complex patterns in genomic data. SALAI‐Net uses a convolutional neural network (CNN) with an attention mechanism to infer LA in a species‐agnostic way; it is trained on simulated data to generalize across ancestry compositions and even across species, reducing the need to retrain for each new dataset (Oriol Sabat et al., 2022). SALAI‐Net matches query haplotypes to reference haplotypes within windows and uses multi‐head attention to integrate information across windows, outputting ancestry probabilities for each segment. Similarly, Orchestra is a deep learning model (with convolutional and attention layers) trained on >10,000 individuals from 35 populations; it achieves state‐of‐the‐art accuracy, particularly in complex multi‐way admixture scenarios, and substantially outperforms prior methods in a 35‐way admixture benchmark (Lerga‐Jaso et al., 2025). These methods demonstrate the potential for high accuracy even with many ancestry categories, but they often require large training datasets and specialized training pipelines and can pose challenges for interpretability and generalization.

Performance

No single LA method is uniformly best for all scenarios. For example, methods such as FLARE, Gnomix, and Loter tend to perform well for older admixture events (e.g., over 100 generations) or subtle ancestry differences (where ancestry segments are short and numerous) (Wei et al., 2025). By contrast, RFMix and Recomb‐Mix often show superior accuracy in recently admixed populations (Wei et al., 2025). Deep learning–based approaches (SALAI‐Net, Orchestra) extend these capabilities by accommodating many ancestral populations and complex multi‐way admixture histories within a single framework. In a comprehensive benchmark, Wei et al. (2025) reported that Orchestra and Recomb‐Mix achieved the most robust overall performance, whereas other approaches, such as FLARE and Gnomix, showed advantages in specific challenging settings (e.g., very ancient admixture or closely related source populations). Collectively, these results indicate that method selection should be guided by the study context, including the number of ancestries, time since admixture, availability of training data, and trade‐offs between accuracy and computational cost.

Once LA has been inferred along the genome, researchers can derive additional ancestry‐informed insights. The ancestry tract length distribution (the sizes of continuous segments of a given ancestry in the genome) is informative about admixture timing: on average, shorter ancestry tracts indicate older admixture events (Hellenthal et al., 2014). By analyzing tract lengths, one can estimate the number of generations since major admixture events and detect non‐random patterns of ancestry along the genome that may indicate natural selection. For example, admixture mapping leverages LA results to test whether differences in LA across individuals help explain variation in a trait or disease, under the expectation that causal variants differ in frequency between ancestral populations; admixture mapping can be more powerful than standard SNP‐based GWASs because it tests ancestry segments rather than individual variants, reducing the multiple‐testing burden and enabling detection of risk loci that might otherwise be missed (Atkinson et al., 2021). Tools such as TRACTOR (Atkinson et al., 2021) incorporate LA in association tests, allowing admixed individuals to be included in GWASs while accounting for their mosaic genomes. LA information can also improve imputation in admixed populations by using ancestry‐specific reference panels to impute missing genotypes more accurately within each ancestry segment. In summary, LA analysis not only provides a finer‐resolution picture of ancestry but also enables analyses of admixture timing and history and ancestry‐specific genetic effects that complement GA findings.

Key local ancestry methods

Table 2 highlights a selection of LA inference tools, illustrating the variety of algorithmic approaches. Note that many other methods exist; the ones listed are representative of the most‐used or recent tools.

Table 2.

Local Ancestry Inference Methods

Method Algorithm type Input requirements Strengths Limitations
RFMix v2 Random forest + CRF (two‐stage) Phased genotypes of admixed individuals; reference panel haplotypes labeled by ancestry High accuracy for recent admixture; incorporates recombination map; robust to small phasing errors by post hoc smoothing Requires well‐phased data; computationally intensive for very large datasets
FLARE Extended Li‐Stephens HMM (haplotype compression) Unphased or phased genotypes; reference panel haplotypes Fast and scalable to biobank‐scale data; maintains high accuracy by reducing model complexity A newer method with relatively limited independent validation so far; may require parameter tuning for optimal compression levels
Orchestra Deep learning (CNN + attention) Phased genotypes; large training reference (>10k individuals) with known ancestries Very high accuracy, especially as number of ancestry groups increases; pre‐trained model available for human populations Less interpretable; training is computationally heavy; performance depends on how closely the training panel matches the study populations
LAMP‐LD Window‐based HMM Unphased genotypes acceptable (internally considers haplotype pairs); reference panel genotypes Accounts for linkage disequilibrium within windows; relatively robust to phasing errors; was one of the first to handle more than two ancestries Requires tuning window size and other parameters; can be less accurate for very small ancestry segments

Reference Panels for Ancestry Inference

Accurate ancestry inference hinges on the quality of the reference panel. A reference panel consists of reference individuals from the ancestral source populations, with known or assumed ancestry, that provide the baseline allele frequencies or haplotype patterns for each population (Cann et al., 2002). The choice of reference panel affects both GA and LA inference (Martin et al., 2017). Important considerations include the following:

Population matching

The reference populations should closely match the true ancestral groups of the study samples (Popejoy & Fullerton, 2016). If an ancestry present in the admixed individuals is missing in the reference, methods may misattribute that ancestry to the closest available population, a pitfall known as reference‐panel bias. For example, an individual with ancestry from a population not represented in the reference might have those segments incorrectly labeled as another population (Hellenthal et al., 2014). To mitigate this, one should include reference groups that capture the known historical sources of admixture in the study population. Additionally, reference individuals are ideally chosen to be unrelated and unadmixed themselves (they should be clear representatives of a single ancestry). Admixed reference samples can confuse the algorithms because they carry mosaic genomes (Peterson et al., 2019). Some reference datasets (like the Human Genome Diversity Project, or HGDP) intentionally sampled isolated populations with minimal recent admixture to serve as clean reference points (Bergstrom et al., 2020).

Diversity and representativeness

The panel should capture the genetic diversity of each ancestral group. If the reference for a given ancestry is too narrow (e.g., individuals sampled only from one sub‐region), segments from a different subgroup might be misclassified. Studies have shown that including more diverse subpopulations can improve inference accuracy (Martin et al., 2017). However, increasing reference diversity can also introduce closely related populations that are harder to distinguish. A balance is required, and sometimes it may be necessary to merge closely related reference groups or use hierarchical inference strategies (first assign broad continental ancestry, then refine to sub‐ancestry levels).

Sample size

Generally, larger reference panels improve inference accuracy up to a point of diminishing returns (Browning & Browning, 2016). For GA, even on the order of tens of individuals per population can suffice to estimate allele frequency differences (e.g., the 1000 Genomes Project used ∼100 samples per population). In practice, around 50 individuals per population is often cited as a bare minimum for stable global proportion estimates, whereas >100 is recommended for more robust inference (Martin et al., 2017; Pritchard et al., 2000). LA benefits even more from large reference sample sizes because haplotype variation needs to be well represented; several hundred phased haplotypes per ancestry are typically desirable (Martin et al., 2017). When reference sample size is limited, some methods (especially those using HMMs or deep learning) can partially compensate by sharing information across individuals or using internal regularization and modeling features. Conversely, extremely large panels can pose computational challenges for methods that scale worse than linearly with reference size. Newer algorithms such as FLARE and Recomb‐Mix have been designed with scalability in mind and can handle thousands of reference haplotypes (Browning et al., 2023; Wei et al., 2025).

Resource examples

Several major publicly available reference panels are widely used in genetic studies. Table 3 lists some of the most commonly used resources and their key features. These panels differ in population coverage and sample size. For instance, the 1000 Genomes Project (phase 3) includes individuals from Africa, Europe, East and South Asia, and the Americas, providing broad global coverage with high‐quality phased genomes (Auton et al., 2015). The HGDP is smaller but includes many indigenous populations chosen for high genetic diversity and minimal recent admixture, making it particularly valuable for ancestry inference (Bergstrom et al., 2020; Cann et al., 2002). gnomAD is much larger in sample size (>140,000 genomes/exomes) but was primarily designed to provide allele frequency estimates rather than to serve as a dedicated ancestry reference panel (Karczewski et al., 2020). The All of Us Research Program (USA) (Ramirez et al., 2022) is a contemporary resource with genotypes on hundreds of thousands of participants from US, many of whom have admixed ancestries. Additionally, regional initiatives such as H3Africa (Rotimi et al., 2014) for African populations or Latin American initiatives (Borda et al., 2024; Bruxel et al., 2025; Sohail & Moreno‐Estrada, 2024) are helping fill gaps in global reference representation. Choosing a reference panel that is well matched to the study population is critical, as inaccurate or biased references will directly propagate into inaccurate ancestry inference.

Table 3.

Major Reference Panels for Ancestry Inference

Reference panel Populations covered Sample size Key features Access Compute environment
1000 Genomes Project (Phase 3) 26 populations worldwide (Africa, Europe, East Asia, South Asia, Americas) 2,504 Broad global coverage; high‐quality phased haplotypes; common baseline for many studies Publicly available Desktop feasible for typical use; HPC or cloud recommended for large cohorts
Human Genome Diversity Project (HGDP) 51 populations worldwide (Africa, Middle East, Europe, South/Central Asia, East Asia, Oceania, Americas) 1,043 Maximizes genetic diversity; includes several geographically isolated groups Publicly available Desktop feasible for typical use; HPC or cloud recommended for large cohorts
gnomAD (Genome Aggregation Database) Multiple (primarily aggregated data from various populations) >800,000 Very large sample size; great for allele frequency reference; however, contains admixed individuals and many exome sequences; mainly used for frequency‐based analyses Public aggregated summaries; individual‐level genotypes not distributed via gnomAD Desktop feasible for summary use; HPC or cloud recommended for large scale VCF processing
All of Us (Research Program) Diverse US population (many with mixed ancestry; e.g., large African American, Hispanic/Latino representation >245,000 (ongoing) Ultra‐large, real‐world cohort; can create ancestry‐specific reference subsets to improve representation of admixed groups Controlled access by approved application and agreement Typically analyzed within the All of Us Researcher Workbench cloud environment

Phasing for Ancestry Inference

Phasing refers to the determination of haplotype structure, that is, the assignment of alleles to maternal and paternal chromosomes (Browning & Browning, 2011). This is important for ancestry inference because ancestry is inherited in haplotype chunks: during admixture, long stretches of the chromosome are copied from an ancestor. If genotype data are unphased, LA methods either have to consider many possible phasings (dramatically increasing computational complexity) or ignore long‐range linkage, both of which reduce accuracy. Therefore, high‐quality phasing prior to LA inference is strongly recommended (Browning & Browning, 2011).

Phasing methods

Modern phasing algorithms leverage population‐level genotype datasets to statistically infer haplotypes. They assume that within a large sample, segments of chromosomes will be shared identical‐by‐descent between individuals, providing information about the underlying phase (Browning & Browning, 2016). Leading phasing tools include BEAGLE 5.4, which uses localized haplotype clustering (an HMM‐based approach) and can perform genotype imputation simultaneously (Browning & Browning, 2016). SHAPEIT5 is a newer method that introduces an efficient algorithm based on the Positional Burrows‐Wheeler Transform (PBWT) to phase millions of genomes with linear scalability (Delaneau et al., 2011; Delaneau et al., 2019; Hofmeister et al., 2023). It achieves high accuracy even for rare variants and is optimized for huge datasets (Hofmeister et al., 2023). Eagle is another widely used phasing tool that also uses a PBWT‐inspired approach and is known for very fast run times with modest memory usage (Loh et al., 2016). These population‐based phasing methods work best for unrelated or distantly related individuals. In cases where family data are available, family‐based phasing can be applied (the known relationships are used to assign haplotypes within pedigrees) (Browning & Browning, 2011). Family phasing can yield nearly error‐free phasing for long segments, which is particularly helpful for phasing rare variants that population‐based methods might struggle with. However, not all studies have family data, and population phasing methods are generally quite effective for common variants in large datasets.

Impact on local ancestry

Phasing errors can directly degrade LA accuracy. A switch error (where a segment of a haplotype is incorrectly flipped between chromosomes) will appear to an LA algorithm as an ancestry switch that never actually happened biologically (Avadhanam & Williams, 2025), creating spurious short ancestry tracts. Some LA methods attempt to be robust to phasing errors. For example, LAMP‐LD was designed to allow unphased input by effectively considering both phasing possibilities at once for each window (Baran et al., 2012). RFMix's smoothing CRF can sometimes correct small phasing mistakes post hoc by realigning segments (Maples et al., 2013). Nonetheless, the general best practice is to minimize phasing errors upfront. Using the latest phasing algorithms and performing rigorous quality control (QC) will lead to more reliable LA calls (Browning et al., 2023). Even for GA, phasing can help in certain analyses, such as detecting long continuous stretches of homozygosity or identifying segments of identity‐by‐descent (IBD) that relate to ancestry (Peterson et al., 2019). In summary, phasing is an essential step that underpins the resolution and reliability of LA inference.

Key phasing tools

Table 4 outlines a few commonly used phasing programs and their characteristics, as relevant to ancestry studies.

Table 4.

Phasing Methods

Method Algorithm type Input requirements Strengths Limitations
BEAGLE 5.4 Li‐Stephens local haplotype model (HMM‐based) Unphased genotypes, reference panel (optional); can also impute missing data Fast and accurate for phasing and imputation; handles very large reference panels; robust performance across a range of sample sizes Performance may drop if reference data are sparse; results can be less accurate in very small datasets without reference
SHAPEIT5 PBWT‐based iterative refinement Unphased genotypes; can incorporate a reference panel or use haplotype library Extremely fast and memory efficient; high accuracy, including for rare variants; scalable to biobank‐scale data Newer method; may require substantial computational resources for peak performance
Eagle PBWT + haplotype graph minimization Unphased genotypes Very rapid phasing for large datasets; produces long‐range phase information accurately Designed mainly for very large sample sizes; in small sample sets, might not outperform other algorithms; less accurate for rare variants

Method Selection Framework

Choosing the right ancestry inference strategy depends on the research question, the data characteristics, and available resources. However, these approaches are not mutually exclusive. In many studies, GA analysis is first performed to characterize population structure, identify outliers, and verify or refine sample labels. Then, LA analysis is run on a targeted set of samples or on a specific chromosome or genomic region of interest. Conversely, LA calls can be aggregated across the genome to obtain GA proportions, which may in some settings be more accurate than direct GA estimates when the LA method itself is highly accurate.

Basic Protocol 1. GLOBAL ANCESTRY ANALYSIS (ADMIXTURE PIPELINE)

This section provides a step‐by‐step protocol for performing a typical GA analysis using the ADMIXTURE software (Alexander et al., 2009). We assume genotype data are available for the study individuals in PLINK format (e.g., SNP array genotypes), along with a suitable reference panel or set of reference individuals from relevant populations. The goal is to estimate each individual's ancestry proportions. The steps include preparing the data, running ADMIXTURE, and interpreting the results.

Software and computational requirements

  • PLINK (v1.9 or later) and ADMIXTURE (v1.3 or later) are required. Input may come from SNP arrays or from whole‐genome sequencing (WGS) data converted to PLINK. The PLINK QC and LD pruning steps in this protocol are typically feasible on a desktop or laptop for standard SNP array datasets. For WGS scale marker sets, large cohorts, or repeated ADMIXTURE runs across many K values with CV, a high‐performance computing (HPC) or cloud environment is recommended, and analyses can be parallelized across K.

Data preparation and quality control

  • 1

    Start by creating a merged genotype dataset that includes the study samples together with any reference individuals (if reference samples will be included in the analysis). Perform standard genotype QC (Marees et al., 2018) using a tool like PLINK (Chang et al., 2015) (e.g., remove individuals with high missingness and markers with low call rate or very low minor allele frequency). Additionally, filter out obvious genotype errors or duplicates. LD pruning is critical: ADMIXTURE assumes markers are unlinked, so prune or thin the SNP set (e.g., remove one SNP from any pair with R2 > 0.1 within a 50‐kb window, sliding by 5 kb).

    This can be done in PLINK with commands like the following:
    
    
    plink --bfile raw_data \
    --geno 0.02\
    --maf 0.01\
    --indep‐pairwise 50 5 0.1\
    --make‐bed
    --out data_QC
    

    The example above is an ADMIXTURE‐oriented QC pipeline that filters SNPs [e.g., geno 2% missingness threshold, maf (minor allele frequency) 1% threshold] and then performs LD pruning, resulting in a pruned dataset (data_QC.bed) suitable for ADMIXTURE. Ensure that all individuals have appropriate population labels for downstream reference‐based interpretation.

Running ADMIXTURE for ancestry inference

  • 2
    Use ADMIXTURE to take genotype data in PLINK .bed format and estimate ancestry proportions for a specified number of clusters K. It is often useful to run ADMIXTURE for several values of K to find the best fit. ADMIXTURE's CV feature can be used to identify the optimal value of K. For each K of interest, run the following:
    
    
    admixture --cv=10 data_QC.bed K
    

    This command will output two main files: data_QC.K.Q (the Q matrix of ancestry proportions for each individual) and data_QC.K.P (the P matrix of allele frequencies for each ancestry cluster). It also prints the CV error for that K to the console. Run this for K = 2,3,4, … up to a reasonable maximum number (if up to five ancestral populations are expected, K = 5 could be tested). Then, examine the CV errors; the model with the lowest CV error is often considered the best‐supported K.

Results interpretation and visualization

  • 3

    After selecting an optimal K, inspect the Q matrix (which can be loaded in R or Python or viewed as text). Each individual will have K columns summing to 1, representing the fraction of their genome assigned to each ancestry cluster. If known reference individuals were included, each inferred cluster can be labeled according to which reference population has values near 1 in that cluster (e.g., if African reference individuals have Q ∼ (1,0,0) for K = 3, that first component corresponds to African ancestry). It is also helpful to plot the Q matrix as a bar plot, with individuals grouped by population or sorted by ancestry proportion. Each bar represents an individual, partitioned into K colored segments. This provides an intuitive visualization of admixture.

    By following this pipeline, you obtain robust GA estimates for each individual. This provides a foundation for more targeted analyses, such as LA or admixture mapping, and ensures any population structure in your data is acknowledged in subsequent statistical analyses (Peterson et al., 2019).

Basic Protocol 2. LOCAL ANCESTRY ANALYSIS (PHASING + RFMix PIPELINE)

This section outlines a typical pipeline for LA inference using RFMix v2 (Maples et al., 2013) as the LA inference tool. We assume you have a dataset of admixed individuals and a reference panel with haplotype data from the ancestral source populations. Given that RFMix requires phased input, we first phase the data (using SHAPEIT5 as an example phasing tool) (Hofmeister et al., 2023) and then run RFMix to obtain LA calls for each genomic segment.

Software and computational requirements

  • This protocol assumes SHAPEIT5 (v5.1.0) and RFMix v2 are installed, along with standard Variant Call Format (VCF) utilities (bcftools v1.10 or later and htslib tabix v1.10 or later). Small pilot datasets or single‐chromosome runs may be feasible on a desktop or laptop. For whole‐genome analyses, dense WGS‐scale data, or moderate to large cohorts, an HPC or cloud environment is recommended, and analyses can be parallelized by chromosome.

  • The following are prerequisites: (a) genotype data for admixed study individuals (in VCF or PLINK format), (b) genotype data for reference individuals from each ancestry (with the same SNP set or at least substantial overlap), and (c) a genetic map or recombination rate file for the SNP coordinates.

Pre‐phase the genotypes

  • 1

    Phase the study individuals using a phasing tool like SHAPEIT5, ensuring the algorithm leverages the reference panel haplotypes. Instead of merging the reference with the target dataset, provide the external reference panel explicitly to the phasing software. This way, all individuals are effectively phased together with the reference haplotypes guiding the phase, without physically combining the files.

    For example, using SHAPEIT5 (module phase_common for common variants) for each chromosome, you might run the following:
    
    
    phase_common --input study_samples_chr1.vcf.gz \
    --reference ref_panel_chr1.vcf.gz \
    --map genetic_map_chr1.txt --region 1 \
    --output phased_chr1.vcf.gz --thread 8
    

    The above command phases chromosome 1 of the study samples (‐‐input) using a reference panel of haplotypes for chr1 (‐‐reference), with a genetic map provided (‐‐map). We specify the chromosome via ‐‐region and use eight threads. In practice, SHAPEIT5 has specific modules (separate tools for common and rare variants) and many options for reference panels and scaffolds. You would repeat the phasing for all chromosomes. The result is a phased VCF where each individual's genotypes are split into phase 1 and phase 2 (maternal vs. paternal haplotypes).

Prepare input for RFMix

  • 2

    For RFMix v2 to take a phased VCF, provide the following: (a) a phased VCF containing haplotypes for the study samples, (b) a phased VCF of the reference panel haplotypes, (c) classes file (also called a reference population file) that maps each haplotype of each reference individual to an ancestry ID, and (d) a genetic map file (chromosome, position, genetic_map_position) for the markers when using recombination distances (recommended for LA inference). Ensure that the SNP set and allele coding are consistent between the study and reference data (e.g., filter out strand‐ambiguous SNPs, indels, and rare variants that may cause artifacts or mismatches).

Run RFMix for local ancestry inference

  • 3

    Execute RFMix v2 with the prepared inputs.

    The exact command depends on the software version; for RFMix v2, a general command might look like the following:
    
    
    rfmix \
    ‐f phased_chr1.vcf.gz \
    ‐r ref_panel_chr1.phased.vcf.gz \
    ‐m rfmix_ref_panel_classes.txt \
    ‐g chr1.b38.rfmix.gmap \
    ‐n 5 \
    ‐c 0.2 \
    ‐s 0.2 \
    ‐o LA_chr1 \
    --chromosome=chr1
    
    • In this example,
      • ‐f phased_chr1.vcf.gz specifies the phased VCF of the target samples for chromosome 1.
      • ‐r ref_panel_chr1.phased.vcf.gz specifies the phased VCF of the reference panel haplotypes.
      • ‐m rfmix_ref_panel_classes.txt is the classes file listing each reference haplotype's ancestry ID.
      • ‐g chr1.b38.rfmix.gmap is the genetic map for chromosome 1.
      • ‐n 5 sets the number of trees per random forest, and ‐c 0.2 ‐s 0.2 set the window and EM parameters.
      • ‐o LA_chr1 sets the output prefix. The - - chromosome = chr1 option specifies chromosome 1.

    RFMix builds a random forests across the windows along the chromosome and then uses a CRF to smooth the window‐level predictions. It outputs LA assignments for each haplotype at each SNP or window. Typically, output files include per‐site ancestry calls (two ancestry states per position per individual) and optionally posterior probabilities.

Downstream integration

  • 4

    Integrate LA calls into association, demographic, or functional analyses.

    For example, for admixture mapping, you test whether LA at a locus (e.g., African vs. European ancestry) is associated with a phenotype across individuals, leveraging differences in allele frequencies between ancestral populations (Atkinson et al., 2021). Tools like TRACTOR or custom scripts can combine the LA calls with association testing. If the focus is demographic history, you can use the ancestry tract length information to infer admixture timing (short average tract = older admixture) (Hellenthal et al., 2014). You can also visualize results by plotting each chromosome colored by ancestry or by plotting ancestry proportion along the genome to detect regions deviating from the genome‐wide average.

    By following this pipeline, you obtain high‐resolution LA maps for each individual genome. These results enable powerful analyses specific to admixed populations, complementing the GA overview obtained from methods like ADMIXTURE. As with GA, interpretation of LA results should be made in the context of the reference panel and modeling assumptions, and for critical applications, it can be informative to compare results from multiple LA inference methods (Wei et al., 2025).

COMMENTARY

Ancestry inference is a rapidly evolving field, with continuous improvements in algorithms, reference data, and computational efficiency. In recent years, new methods such as FLARE and Recomb‐Mix have been introduced, scaling to biobank‐sized datasets without sacrificing accuracy (Browning et al., 2023; Wei et al., 2025), and deep learning approaches (SALAI‐Net, Orchestra) have begun to handle complex multi‐ancestry scenarios that were previously difficult to model (Lerga‐Jaso et al., 2025; Oriol Sabat et al., 2022). Similarly, phasing methods such as SHAPEIT5 have dramatically increased the speed of haplotype assembly, enabling ancestry analysis on millions of genomes (Hofmeister et al., 2023). On the reference data front, initiatives like All of Us and H3Africa are broadening representation, which will help reduce the reference‐panel bias and improve accuracy for historically understudied populations (Popejoy & Fullerton, 2016).

Despite this progress, challenges remain. One major issue is reference panel mismatch, namely, when the available reference populations do not perfectly match the genetic ancestry of the study samples (Martin et al., 2017; Pearson & Durbin, 2023). Future methods are likely to incorporate robustness against such mismatches, perhaps by down‐weighting less relevant reference haplotypes or by simultaneously learning ancestry clusters from the dataset itself. There is also growing interest in integrating ancient DNA to calibrate and validate admixture timing inferences, as well as extending LA methods to non‐human species (Oriol Sabat et al., 2022).

Method selection should be guided by research objectives, computational resources, and data characteristics. GA methods are appropriate for population structure correction and demographic analysis (Qin & Zhu, 2012), whereas LA approaches enable fine‐scale mapping and studies of selection (Cuadros‐Espinoza et al., 2022; Qin et al., 2010). Proper QC, reference panel selection, and validation are critical for reliable results across all applications.

Another frontier is the application of ancestry inference in medical and functional genomics. As precision medicine studies include more diverse cohorts, controlling for ancestry and discovering ancestry‐specific genetic effects become crucial (Peterson et al., 2019). Methods that combine LA with association analysis (Atkinson et al., 2021) or incorporate functional genomic data (e.g., expression quantitative trait locus, or eQTL, data) are emerging (Gay et al., 2020). We anticipate further development of integrative tools that seamlessly connect ancestry inference with downstream analyses like GWASs, admixture mapping, polygenic risk scoring, and detection of natural selection.

In summary, researchers should choose GA vs. LA approaches based on their specific goals; often, a combination of both is ideal. GA gives a broad overview and is computationally lightweight, whereas LA provides granular insight and power for certain analyses at the cost of greater computational and modeling complexity. By following best practices as outlined above, one can obtain reliable ancestry inferences. These inferences not only illuminate the demographic history behind a dataset but also enable more equitable and informed genetic studies in our increasingly admixed world.

Author Contributions

Bilcag Akgun: Writing—original draft; conceptualization; methodology. Farid Rajabli: Writing—review and editing; supervision; conceptualization; resources; funding acquisition; methodology; project administration.

Conflict of Interest

The authors declare no conflicts of interest.

Acknowledgments

The authors acknowledge the participants in genetic studies whose contributions have enabled advances in population genetics and medical genetics. This work was supported by grants U19AG074865 and R01AG070864 from the National Institute on Aging (NIA) of the National Institutes of Health (NIH), in addition to a grant from the Coins for Alzheimer's Research Trust (CART).

Akgun, B. , & Rajabli, F. (2026). Strategies in global ancestry and local ancestry inference. Current Protocols, 6, e70321. doi: 10.1002/cpz1.70321

Published in the Human Genetics section

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analyzed during the current work.

Literature Cited

  1. Alexander, D. H. , Novembre, J. , & Lange, K. (2009). Fast model‐based estimation of ancestry in unrelated individuals. Genome Research, 19(9), 1655–1664. 10.1101/gr.094052.109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Atkin, A. L. , Christophe, N. K. , Stein, G. L. , Gabriel, A. K. , & Lee, R. M. (2022). Race terminology in the field of psychology: Acknowledging the growing multiracial population in the U.S. American Psychologist, 77(3), 381–393. 10.1037/amp0000975 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Atkinson, E. G. , Maihofer, A. X. , Kanai, M. , Martin, A. R. , Karczewski, K. J. , Santoro, M. L. , Ulirsch, J. C. , Kamatani, Y. , Okada, Y. , Finucane, H. K. , Koenen, K. C. , Nievergelt, C. M. , Daly, M. J. , & Neale, B. M. (2021). Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nature Genetics, 53(2), 195–204. 10.1038/s41588-020-00766-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Auton, A. , Brooks, L. D. , Durbin, R. M. , Garrison, E. P. , Kang, H. M. , Korbel, J. O. , Marchini, J. L. , McCarthy, S. , McVean, G. A. , Abecasis, G. R. , & 1000 Genomes Project Consortium . (2015). A global reference for human genetic variation. Nature, 526(7571), 68–74. 10.1038/nature15393 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Avadhanam, S. , & Williams, A. L. (2025). Phase‐free local ancestry inference mitigates the impact of switch errors on phase‐based methods. G3 (Bethesda), 15(8), jkaf122. 10.1093/g3journal/jkaf122 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Baran, Y. , Pasaniuc, B. , Sankararaman, S. , Torgerson, D. G. , Gignoux, C. , Eng, C. , Rodriguez‐Cintron, W. , Chapela, R. , Ford, J. G. , Avila, P. C. , Rodriguez‐Santana, J. , Burchard, E. G. , & Halperin, E. (2012). Fast and accurate inference of local ancestry in Latino populations. Bioinformatics, 28(10), 1359–1367. 10.1093/bioinformatics/bts144 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Bergstrom, A. , McCarthy, S. A. , Hui, R. , Almarri, M. A. , Ayub, Q. , Danecek, P. , Chen, Y. , Felkel, S. , Hallast, P. , Kamm, J. , Blanche, H. , Deleuze, J. F. , Cann, H. , Mallick, S. , Reich, D. , Sandhu, M. S. , Skoglund, P. , Scally, A. , Xue, Y. , … Tyler‐Smith, C. (2020). Insights into human genetic variation and population history from 929 diverse genomes. Science, 367(6484). 10.1126/science.aay5012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Bien, S. A. , Wojcik, G. L. , Hodonsky, C. J. , Gignoux, C. R. , Cheng, I. , Matise, T. C. , Peters, U. , Kenny, E. E. , & North, K. E. (2019). The future of genomic studies must be globally representative: Perspectives from PAGE. Annual Review of Genomics and Human Genetics, 20, 181–200. 10.1146/annurev-genom-091416-035517 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Borda, V. , Loesch, D. P. , Guo, B. , Laboulaye, R. , Veliz‐Otani, D. , French, J. N. , Leal, T. P. , Gogarten, S. M. , Ikpe, S. , Gouveia, M. H. , Mendes, M. , Abecasis, G. R. , Alvim, I. , Arboleda‐Bustos, C. E. , Arboleda, G. , Arboleda, H. , Barreto, M. L. , Barwick, L. , Bezzera, M. A. , … O'Connor, T. D. (2024). Genetics of Latin American Diversity Project: Insights into population genetics and association studies in admixed groups in the Americas. Cell Genomics, 4(11), 100692. 10.1016/j.xgen.2024.100692 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Browning, B. L. , & Browning, S. R. (2016). Genotype imputation with millions of reference samples. The American Journal of Human Genetics, 98(1), 116–126. 10.1016/j.ajhg.2015.11.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Browning, S. R. , & Browning, B. L. (2011). Haplotype phasing: existing methods and new developments. Nature Reviews Genetics, 12(10), 703–714. 10.1038/nrg3054 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Browning, S. R. , Waples, R. K. , & Browning, B. L. (2023). Fast, accurate local ancestry inference with FLARE. The American Journal of Human Genetics, 110(2), 326–335. 10.1016/j.ajhg.2022.12.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Bruxel, E. M. , Rovaris, D. L. , Belangero, S. I. , Chavarria‐Soley, G. , Cuellar‐Barboza, A. B. , Martinez‐Magana, J. J. , Nagamatsu, S. T. , Nievergelt, C. M. , Nunez‐Rios, D. L. , Ota, V. K. , Peterson, R. E. , Sloofman, L. G. , Adams, A. M. , Albino, E. , Alvarado, A. T. , Andrade‐Brito, D. , Arguello‐Pascualli, P. Y. , Bandeira, C. E. , Bau, C. H. D. , … Montalvo‐Ortiz, J. L. (2025). Psychiatric genetics in the diverse landscape of Latin American populations. Nature Genetics, 57(5), 1074–1088. 10.1038/s41588-025-02127-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Cann, H. M. , de Toma, C. , Cazes, L. , Legrand, M. F. , Morel, V. , Piouffre, L. , Bodmer, J. , Bodmer, W. F. , Bonne‐Tamir, B. , Cambon‐Thomsen, A. , Chen, Z. , Chu, J. , Carcassi, C. , Contu, L. , Du, R. , Excoffier, L. , Ferrara, G. B. , Friedlaender, J. S. , Groot, H. , … Cavalli‐Sforza, L. L. (2002). A human genome diversity cell line panel. Science, 296(5566), 261–262. 10.1126/science.296.5566.261b [DOI] [PubMed] [Google Scholar]
  15. Chang, C. C. , Chow, C. C. , Tellier, L. C. , Vattikuti, S. , Purcell, S. M. , & Lee, J. J. (2015). Second‐generation PLINK: rising to the challenge of larger and richer datasets. Gigascience, 4, 7. 10.1186/s13742-015-0047-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Cuadros‐Espinoza, S. , Laval, G. , Quintana‐Murci, L. , & Patin, E. (2022). The genomic signatures of natural selection in admixed human populations. The American Journal of Human Genetics, 109(4), 710–726. 10.1016/j.ajhg.2022.02.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Delaneau, O. , Marchini, J. , & Zagury, J. F. (2011). A linear complexity phasing method for thousands of genomes. Nature Methods, 9(2), 179–181. 10.1038/nmeth.1785 [DOI] [PubMed] [Google Scholar]
  18. Delaneau, O. , Zagury, J. F. , Robinson, M. R. , Marchini, J. L. , & Dermitzakis, E. T. (2019). Accurate, scalable and integrative haplotype estimation. Nature Communications, 10(1), 5436. 10.1038/s41467-019-13225-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Dias‐Alves, T. , Mairal, J. , & Blum, M. G. B. (2018). Loter: A software package to infer local ancestry for a wide range of species. Molecular Biology and Evolution, 35(9), 2318–2326. 10.1093/molbev/msy126 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Falush, D. , Stephens, M. , & Pritchard, J. K. (2003). Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies. Genetics, 164(4), 1567–1587. 10.1093/genetics/164.4.1567 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Fatumo, S. , Chikowore, T. , Choudhury, A. , Ayub, M. , Martin, A. R. , & Kuchenbaecker, K. (2022). A roadmap to increase diversity in genomic studies. Nature Medicine, 28(2), 243–250. 10.1038/s41591-021-01672-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Gay, N. R. , Gloudemans, M. , Antonio, M. L. , Abell, N. S. , Balliu, B. , Park, Y. , Martin, A. R. , Musharoff, S. , Rao, A. S. , Aguet, F. , Barbeira, A. N. , Bonazzola, R. , Hormozdiari, F. , Ardlie, K. G. , Brown, C. D. , Im, H. K. , Lappalainen, T. , Wen, X. , Montgomery, S. B. , & GTex Consortium . (2020). Impact of admixture and ancestry on eQTL analysis and GWAS colocalization in GTEx. Genome Biology, 21(1), 233. 10.1186/s13059-020-02113-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Hellenthal, G. , Busby, G. B. J. , Band, G. , Wilson, J. F. , Capelli, C. , Falush, D. , & Myers, S. (2014). A genetic atlas of human admixture history. Science, 343(6172), 747–751. 10.1126/science.1243518 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Hilmarsson, H. , Kumar, A. S. , Rastogi, R. , Bustamante, C. D. , Montserrat, D. M. , & Ioannidis, A. G. (2021). High resolution ancestry deconvolution for next generation genomic data. bioRxiv, 2021.2009.2019.460980. 10.1101/2021.09.19.460980 [DOI] [Google Scholar]
  25. Hofmeister, R. J. , Ribeiro, D. M. , Rubinacci, S. , & Delaneau, O. (2023). Accurate rare variant phasing of whole‐genome and whole‐exome sequencing data in the UK Biobank. Nature Genetics, 55(7), 1243–1249. 10.1038/s41588-023-01415-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Karczewski, K. J. , Francioli, L. C. , Tiao, G. , Cummings, B. B. , Alfoldi, J. , Wang, Q. , Collins, R. L. , Laricchia, K. M. , Ganna, A. , Birnbaum, D. P. , Gauthier, L. D. , Brand, H. , Solomonson, M. , Watts, N. A. , Rhodes, D. , Singer‐Berk, M. , England, E. M. , Seaby, E. G. , Kosmicki, J. A. , … MacArthur, D. G. (2020). The mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 581(7809), 434–443. 10.1038/s41586-020-2308-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Krainc, T. , & Fuentes, A. (2022). Genetic ancestry in precision medicine is reshaping the race debate. Proceedings of the National Academy of Sciences of the United States of America, 119(12), e2203033119. 10.1073/pnas.2203033119 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Lerga‐Jaso, J. , Novkovic, B. , Unnikrishnan, D. , Bamunusinghe, V. , Hatorangan, M. R. , Manson, C. , Pedersen, H. , Osama, A. , Terpolovsky, A. , Bohn, S. , de Marino, A. , Mahmoud, A. A. , Bircan, K. O. , Khan, U. , Grabherr, M. G. , & Yazdi, P. G. (2025). Tracing human genetic histories and natural selection with precise local ancestry inference. Nature Communications, 16(1), 4576. 10.1038/s41467-025-59936-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Loh, P. R. , Danecek, P. , Palamara, P. F. , Fuchsberger, C. , Y, A. R. , H, K. F. , Schoenherr, S. , Forer, L. , McCarthy, S. , Abecasis, G. R. , Durbin, R. , & A, L. P. (2016). Reference‐based phasing using the Haplotype Reference Consortium panel. Nature Genetics, 48(11), 1443–1448. 10.1038/ng.3679 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Loh, P. R. , Lipson, M. , Patterson, N. , Moorjani, P. , Pickrell, J. K. , Reich, D. , & Berger, B. (2013). Inferring admixture histories of human populations using linkage disequilibrium. Genetics, 193(4), 1233–1254. 10.1534/genetics.112.147330 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Maples, B. K. , Gravel, S. , Kenny, E. E. , & Bustamante, C. D. (2013). RFMix: a discriminative modeling approach for rapid and robust local‐ancestry inference. The American Journal of Human Genetics, 93(2), 278–288. 10.1016/j.ajhg.2013.06.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Marees, A. T. , de Kluiver, H. , Stringer, S. , Vorspan, F. , Curis, E. , Marie‐Claire, C. , & Derks, E. M. (2018). A tutorial on conducting genome‐wide association studies: Quality control and statistical analysis. International Journal of Methods in Psychiatric Research, 27(2), e1608. 10.1002/mpr.1608 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Martin, A. R. , Gignoux, C. R. , Walters, R. K. , Wojcik, G. L. , Neale, B. M. , Gravel, S. , Daly, M. J. , Bustamante, C. D. , & Kenny, E. E. (2017). Human demographic history impacts genetic risk prediction across diverse populations. The American Journal of Human Genetics, 100(4), 635–649. 10.1016/j.ajhg.2017.03.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Oriol Sabat, B. , Mas Montserrat, D. , Giro, I. N. X. , & Ioannidis, A. G. (2022). SALAI‐Net: species‐agnostic local ancestry inference network. Bioinformatics, 38(Suppl_2), ii27–ii33. 10.1093/bioinformatics/btac464 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Padhukasahasram, B. (2014). Inferring ancestry from population genomic data and its applications. Frontiers in Genetics, 5, 204. 10.3389/fgene.2014.00204 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Patterson, N. , Price, A. L. , & Reich, D. (2006). Population structure and eigenanalysis. PLOS Genetics, 2(12), e190. 10.1371/journal.pgen.0020190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Pearson, A. , & Durbin, R. (2023). Local ancestry inference for complex population histories. bioRxiv, 2023.2003.2006.529121. 10.1101/2023.03.06.529121 [DOI] [Google Scholar]
  38. Peterson, R. E. , Kuchenbaecker, K. , Walters, R. K. , Chen, C. Y. , Popejoy, A. B. , Periyasamy, S. , Lam, M. , Iyegbe, C. , Strawbridge, R. J. , Brick, L. , Carey, C. E. , Martin, A. R. , Meyers, J. L. , Su, J. , Chen, J. , Edwards, A. C. , Kalungi, A. , Koen, N. , Majara, L. , … Duncan, L. E. (2019). Genome‐wide association studies in ancestrally diverse populations: Opportunities, methods, pitfalls, and recommendations. Cell, 179(3), 589–603. 10.1016/j.cell.2019.08.051 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Popejoy, A. B. , & Fullerton, S. M. (2016). Genomics is failing on diversity. Nature, 538(7624), 161–164. 10.1038/538161a [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Price, A. L. , Tandon, A. , Patterson, N. , Barnes, K. C. , Rafaels, N. , Ruczinski, I. , Beaty, T. H. , Mathias, R. , Reich, D. , & Myers, S. (2009). Sensitive detection of chromosomal segments of distinct ancestry in admixed populations. PLOS Genetics, 5(6), e1000519. 10.1371/journal.pgen.1000519 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Pritchard, J. K. , Stephens, M. , & Donnelly, P. (2000). Inference of population structure using multilocus genotype data. Genetics, 155(2), 945–959. 10.1093/genetics/155.2.945 [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Qin, H. , Morris, N. , Kang, S. J. , Li, M. , Tayo, B. , Lyon, H. , Hirschhorn, J. , Cooper, R. S. , & Zhu, X. (2010). Interrogating local population structure for fine mapping in genome‐wide association studies. Bioinformatics, 26(23), 2961–2968. 10.1093/bioinformatics/btq560 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Qin, H. , & Zhu, X. (2012). Allowing for population stratification in association analysis. Methods in Molecular Biology, 850, 399–409. 10.1007/978-1-61779-555-8_21 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Racimo, F. , Sankararaman, S. , Nielsen, R. , & Huerta‐Sanchez, E. (2015). Evidence for archaic adaptive introgression in humans. Nature Reviews Genetics, 16(6), 359–371. 10.1038/nrg3936 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Raj, A. , Stephens, M. , & Pritchard, J. K. (2014). fastSTRUCTURE: variational inference of population structure in large SNP data sets. Genetics, 197(2), 573–589. 10.1534/genetics.114.164350 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Ramirez, A. H. , Sulieman, L. , Schlueter, D. J. , Halvorson, A. , Qian, J. , Ratsimbazafy, F. , Loperena, R. , Mayo, K. , Basford, M. , Deflaux, N. , Muthuraman, K. N. , Natarajan, K. , Kho, A. , Xu, H. , Wilkins, C. , Anton‐Culver, H. , Boerwinkle, E. , Cicek, M. , Clark, C. R. , Roden, D. M. , … All of Us Research Program . (2022). The All of Us Research Program: Data quality, utility, and diversity. Patterns (NY), 3(8), 100570. 10.1016/j.patter.2022.100570 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Rotimi, C. , Abayomi, A. , Abimiku, A. , Adabayeri, V. M. , Adebamowo, C. , Adebiyi, E. , Ademola, A. D. , Adeyemo, A. , Adu, D. , Affolabi, D. , Agongo, G. , Ajayi, S. , Akarolo‐Anthony, S. , Akinyemi, R. , Akpalu, A. , Alberts, M. , Alonso Betancourt, O. , Alzohairy, A. M. , Zar, H. , … H3Africa Consortium . (2014). Research capacity. Enabling the genomic revolution in Africa. Science, 344(6190), 1346–1348. 10.1126/science.1251546 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Sankararaman, S. , Sridhar, S. , Kimmel, G. , & Halperin, E. (2008). Estimating local ancestry in admixed populations. The American Journal of Human Genetics, 82(2), 290–303. 10.1016/j.ajhg.2007.09.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Slatko, B. E. , Gardner, A. F. , & Ausubel, F. M. (2018). Overview of next‐generation sequencing technologies. Current Protocols in Molecular Biology, 122(1), e59. 10.1002/cpmb.59 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Sohail, M. , & Moreno‐Estrada, A. (2024). The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity building. Disease Models & Mechanisms, 17(1), dmm050522. 10.1242/dmm.050522 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Tan, T. , & Atkinson, E. G. (2023). Strategies for the genomic analysis of admixed populations. Annual Review of Biomedical Data Science, 6, 105–127. 10.1146/annurev-biodatasci-020722-014310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Tang, H. , Peng, J. , Wang, P. , & Risch, N. J. (2005). Estimation of individual admixture: analytical and study design considerations. Genetic Epidemiology, 28(4), 289–301. 10.1002/gepi.20064 [DOI] [PubMed] [Google Scholar]
  53. Wei, Y. , Zhi, D. , & Zhang, S. (2025). Recomb‐Mix: Fast and accurate local ancestry inference. Bioinformatics, 41(Supplement_1), i180–i188. 10.1093/bioinformatics/btaf227 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analyzed during the current work.


Articles from Current Protocols are provided here courtesy of Wiley

RESOURCES