Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2025 Jan 2;112(2):224–234. doi: 10.1016/j.ajhg.2024.12.005

Characterizing features affecting local ancestry inference performance in admixed populations

Jessica Honorato-Mauer 1, Nirav N Shah 1, Adam X Maihofer 2, Clement C Zai 3, Sintia Belangero 4, Caroline M Nievergelt 2; Psychiatric Genomics Consortium for PTSD Ancestry Working Group, Marcos Santoro 5,7,, Elizabeth G Atkinson 1,6,7,∗∗
PMCID: PMC11866949  PMID: 39753130

Summary

In recent years, significant efforts have been made to improve methods for genomic studies of admixed populations using local ancestry inference (LAI). Accurate LAI is crucial to ensure that downstream analyses accurately reflect the genetic ancestry of research participants. Here, we test analytic strategies for LAI to provide guidelines for optimal accuracy, focusing on admixed populations reflective of Latin America’s primary continental ancestries—African (AFR), Amerindigenous (AMR), and European (EUR). Simulating linkage-disequilibrium-informed admixed haplotypes under a variety of 2- and 3-way admixture models, we implemented a standard LAI pipeline, testing the impact of reference panel composition, DNA data type, demography, and software parameters to quantify ancestry-specific LAI accuracy. We observe that across all models, AMR tracts have notably reduced LAI accuracy as compared to EUR and AFR tracts, with true positive rate means for AMR ranging from 88% to 94%, EUR from 96% to 99%, and AFR from 98% to 99%. When LAI miscalls occurred, they most frequently erroneously called EUR ancestry in true AMR sites. Concerning reference panel curation, we find that using a reference panel well matched to the target population, even with a smaller sample size, was accurate and the most computationally efficient. Imputation did not harm LAI performance in our tests; rather, we observed that higher variant density improved accuracy. While directly responsive to admixed Latin American cohort compositions, these trends are broadly useful for informing best practices for LAI across admixed populations. Our findings reinforce the need for the inclusion of more underrepresented populations in sequencing efforts to improve reference panels.

Keywords: ancestry, local ancestry inference, population genetics, genetic admixture, bioinformatics, reference panels


This work examines accuracy rates and error modes for local ancestry inference (LAI) using simulated admixed cohorts. Our findings inform best practices for high-quality LAI, which is vital for genomics applications in admixed populations. We further highlight the need for increased inclusion of underrepresented populations in reference panels.

Introduction

Admixed populations present a challenge in genome-wide analyses, as their genomes contain components from different continental ancestries, which vary from person to person along the genome, even if two people have the same overall ancestry proportions. This makes it statistically challenging to control for population structure, which can bias tests if left uncorrected.1,2 Despite recent advances in complex trait genetics, limitations remain in our understanding of the architecture of genetic disorders in diverse populations due to their exclusion from many genomic studies.3,4,5 For example, Latin American populations currently represent only 1.3% of all genome-wide association study (GWAS) samples despite accounting for 8.4% of the world population and contributing disproportionately to GWAS findings.6 As large-scale efforts begin to focus more heavily on admixed groups, there is an unmet need for the design of well-suited pipelines to appropriately study these underrepresented populations.7

Local ancestry inference (LAI) is a machine learning approach for assigning each genomic region to a specific ancestry group by comparing phased genotype data to a reference sample containing representative phased whole-genome sequence (WGS) data. This method allows the researcher to assign the ancestral origin for each ancestry component for subsequent analyses, providing a better framework for control over the population structure than only considering global admixture, as is accomplished when including principal components as covariates in statistical testing. There are several newly established analysis methods and tools that are tailored to admixed populations that implement LAI to deconvolute continental ancestry components in admixed samples. Previous work has shown that LAI can improve the discovery power in GWASs for identifying ancestry-specific hits,8 polygenic risk scoring,9 evolutionary research,10 and characterizing gene-gene interactions,11 as well as aiding in the development of more inclusive precision medicine.12,13 To ensure and maximize the success of improving accuracy and statistical power in analyses involving LAI for admixed samples, it is paramount that local ancestry is correctly called in the individual haplotypes. However, limitations in available reference panels for many understudied populations hinder analysis, and the reference panel characteristics and algorithm parameters that result in optimal LAI accuracy across populations are still not firmly established.

One of the main features affecting LAI analysis is the reference panel used to infer local ancestry on the target sample. Reference panels are broadly required for genomic pipelines, including LAI, yet are often sparse for admixed populations, particularly those who have some Amerindigenous (AMR) ancestry. Further, many reference samples are themselves admixed, which complicates assigning ancestral tracts, often resulting in admixed populations having diminished accuracy if reference panel homogeneity is assumed. As such, LAI may perform differently across ancestry components or populations, highlighting the need for establishing best-practice guidelines. This is especially important for diverse cohorts that may lack large, well-matched reference samples to train algorithms effectively.

Here, we comprehensively test strategies for conducting LAI using existing reference resources with simulated “truth” genomic datasets reflective of the demographic history of Latin America to identify features that result in the best true positive rates (TPRs). Latin American populations have a complex ancestry makeup resulting from past admixture events from multiple continental areas. Though the specific patterns vary between different geographic regions, historical admixture events generally involved substantial contributions from AMR, European (EUR), and/or African (AFR) populations.14,15,16 Thus, the genomes of Latin American individuals are complex mosaics of different ancestral tracts that vary in length depending on the historical timing when pulses of admixture occurred. We wish to clarify that in this manuscript, we use the term AMR to refer to the Amerindigenous ancestry present in modern-day Latin America rather than as a population label for admixed American samples, as has been occasionally done in prior efforts. We also wish to highlight that we only describe the inference of individuals’ genetic ancestry throughout this manuscript rather than any metric of self-identification.

In our tests, we modify key parameters affecting LAI performance, including (1) how well matched the reference panel is to the sample, (2) the absolute size of the panel, (3) the presence of admixture in the reference sample, (4) the genomic data type/number of variants (i.e., genotyping arrays versus WGS data), (5) demographic features of the cohort (global admixture proportions and timing of admixture events), and (6) parameter selection in LAI models (e.g., window size, number of expectation-maximization [EM] iterations). This informs best practice for researchers when conducting LAI on Latin American and other admixed populations to produce the highest-accuracy results.

Methods

Dataset generation and quality control

To generate both a simulated truth dataset and comparison reference panels, we used data from the jointly called dataset of the 1000 Genomes Project (1KG) and Human Genome Diversity Project (HGDP) on GRCh38.17,18,19 For our 3-way admixed analyses, we used data from American populations from the HGDP (Karitiana, Surui, Colombian, Maya, Pima) as well as Peruvians from Lima, Peru (PEL) and, in one test, all East Asian (EAS) populations from the 1KG to capture AMR ancestry. Each of the AMR populations from the HGDP was randomly split in half, with one half used for admixture simulations (n = 31) and the other used as a reference sample for LAI (n = 31). To keep sample sizes balanced between ancestries for simulations, we selected 30 Iberians in Spain (IBS) samples from the 1KG to capture southern EUR ancestry and 30 Yoruba in Ibadan, Nigeria (YRI) samples from the 1KG for western AFR ancestry. For the reference panel for LAI, we used the remaining samples from IBS and YRI populations (n = 77 each) and the other half of the AMR samples to represent a common analytic scenario. We appreciate the high amount of intra-continental AFR and EUR diversity present that is omitted with this strategy. However, in testing the inclusion of all AFR/EUR samples as a reference, our results were highly unreliable, with significantly different global ancestry proportions called compared to those modeled. This is likely explained by the highly unbalanced sample size (30 EUR:30 AFR:1 AMR). We filtered to keep only unrelated individuals and excluded multiallelic or duplicated variants, as well as those with a missingness rate > 10% and a minor-allele frequency (MAF) < 0.5%. Genomic phasing of the complete dataset was conducted using SHAPEIT4,20 and after phasing, we subset the populations of interest as described below for our various simulations, with some samples used to model truth individuals and some used as an LAI reference. By using distinct samples for our sample generation and reference panels, we have an unbiased estimate of accuracy, albeit at the cost of reducing the reference sample size.

Simulating truth admixed haplotypes

Because of the reference sample size limitations, we simulated 60 haplotypes for each admixed cohort. Sample sizes for the reference component ancestries were selected to be equivalent to avoid biases due to unbalanced representation, and the terminal node size flag (-n 5) was implemented in LAI runs to further account for any sample size differences.

Latin America is a highly diverse region, and cohort admixture proportions vary widely depending on the country and even within each country.14,15,16 Here, we simulated cohorts with six global ancestry patterns based on common ancestry proportions observed across Latin America. Briefly, in these simulations, one pulse of admixture is simulated at a designated point in time with specified global ancestry proportions contributed from the relevant source populations, after which haplotypes taken from the reference dataset are copied from the previous generation until the present, with tract switches informed by a recombination map. We used here the hg38 HapMap combined recombination map, which includes representatives from relevant global populations.21 This results in a simulated truth dataset that is highly similar to modern empirical Latin American cohorts but has known phase and local ancestry, which can be used for method benchmarking.

We tested four 2-way models of AMR/EUR admixture and two 3-way models of AMR/EUR/AFR admixture. In the 2-way models, we compared the TPR rates for the average ancestry proportions for a 2-way admixed Latin American individual (70% AMR/30% EUR, termed “average 2-way model”)15 and even AMR/EUR proportions (“even model”), as well as two models to analyze the effect of extreme ancestry proportions, each with 5% of one ancestry and 95% of the other (“extreme models”). This allows us to assess the performance that may be expected in a typical 2-way admixed empirical sample, as well as assess features of sample composition influencing accuracy performance.

For 3-way models, we tested a model of average proportions for a 3-way admixed Latin American individual (15% AMR/60% EUR/25% AFR—average proportions for a Brazilian individual, termed “average 3-way model”)22 and an even-proportioned model. Simulations were conducted using the admix-simu tool.23 We simulated the average 3-way model in three different admixture demographic scenarios, considering a single pulse of admixture at 9, 12, or 17 generations ago.14 This allowed us to evaluate the impact of varying tract lengths on true positive LAI rates. All other models were simulated considering a single pulse of admixture 9 generations ago for the sake of comparability. In the simulation of the admixture model that has 3-way average Latin American proportions and a pulse of admixture 12 generations ago, we used data for all autosomes to obtain the highest precision. This admixture model was landed upon because it is reflective of the intermediate admixture pulse in a population migration model for the Brazilian population according to Kehdy et al.14 For all other simulations, we simulated only chromosome 1 for the sake of computational efficiency.

For comparisons of DNA data generation type, we created a pseudo-genotype array dataset by selecting all SNVs present in the Global Screening Array (Illumina GSA) from our WGS-density simulation reference dataset, a genotyping array that has been regularly used for non-EUR datasets. To test the effect of imputation on LAI accuracy, given that imputation is a typical step in cohort data processing for genomic analyses such as GWASs, we imputed the simulated haplotypes with SNP array-density sites using the TOPMed panel24 imputation server and filtered imputed sites with imputation quality scores (INFO) >0.8 and MAFs >0.005.

LAI

Local ancestry was deconvoluted using RFMix v.1.5.4.25 We used the TrioPhased option, with a base window size of 0.2 cM, a terminal node size of 5, and 2 EM iterations, with reference panels reanalyzed in EM to account for any admixture present in the reference (flags -w 0.2, -n 5, -e 2, and --use-reference-panels-in-EM, respectively), and the number of generations since admixture was specified depending on the simulation model (9, 12, or 17). For reference panel testing, we used three different reference panel combinations from the HGDP and 1KG (AMR/EUR or AMR/EUR/AFR) that varied only in the AMR reference samples, given that this group has much less representation in reference panels relative to the other two ancestries. This was done to benchmark how variations in the reference for this ancestry impact LAI accuracy, with particular attention paid to improving AMR accuracy given the limitations of available reference resources. The three reference panels for LAI were constructed using, for EUR and AFR components, respectively, the remaining IBS and YRI samples from the 1KG not used in the simulations (IBS n = 77, YRI n = 77). For the AMR component, the three panels varied as follows: (1) well matched to the target but low sample size: using the other half of HGDP-AMR samples (n = 30); (2) medium sample size containing some admixture in AMR component: using the 1KG sample from PEL (n = 85); and (3) large sample size but AMR component poorly matched to the target: 1KG-PEL (n = 85) plus 1KG EAS populations (n = 505). We included the EAS population in panel 3 to capture highly diverged AMR ancestry considering the demographic history of human migrations, as the ancestors of modern AMR peoples of the Americas migrated from East Asia across the Bering Strait around 15,000 years ago.26 As such, AMR and EAS ancestries are less diverged than other ancestral components, but even so, this composition makes for a poorly matched reference panel to the target data. Panel 2 represents the procedure most commonly conducted in current studies. Additionally, this comparison aimed to observe the impact of homogeneous versus admixed samples in the reference since the HGDP AMR populations used in the first reference panel described above are more homogeneous compared to the 1KG PEL used in the second and third reference panel compositions. We used the LAI reference panel containing the AMR samples described in panel 1 for all comparisons that did not involve reference panel testing.

Statistical analysis

We quantified the LAI TPRs for each run to assess their respective performance. To calculate TPRs, we divided the sum of genomic positions where the ancestry was correctly identified by the sum of positions associated with that ancestry (“truth” local ancestry output by the simulation) for each simulated haplotype (Figure S1). We analyzed the best-guess ancestry calls output by RFMix, regardless of confidence (forward-backward) estimates. We computed ancestry-specific TPRs to assess if there was differential LAI performance depending on the background truth ancestry and tested for statistically significant differences using the Wilcoxon rank-sum test. Significance was considered when the Bonferroni-adjusted p value (p-adj) was <0.05.

Results

Impact of demography and ancestry proportions on LAI performance

We compared the effect of different demographic models on LAI performance. Specifically, we assessed the impact of varying component ancestry proportions in 2- and 3-way models as well as different generation times at which an admixture pulse occurred (considering a single pulse) by simulating 9, 12, and 17 generations since admixture.

In general, LAI accuracy for a given ancestry increased as that global ancestry proportion increased (Figure 1; Table 1). A low global ancestry percentage tended both to decrease the accuracy and result in larger standard deviations, as observed in the extreme-proportions simulations (95% EUR/5% AMR and 5% EUR/95% AMR; Figure 1). In all tested models, we additionally observed significantly lower TPRs for the AMR component (p-adj < 0.05). This result was consistent for all chromosomes and demographic models tested (Figures 1 and 2; Table S1).

Figure 1.

Figure 1

LAI TPRs in between-ancestry comparisons

(A) True positive rates (TPRs) for LAI in six simulated cohorts with varying proportions of 2- or 3-way admixture between AFR/EUR/AMR (displayed in order of decreasing mean TPRs). These simulated haplotypes consist of chromosome 1 and are considered a pulse of admixture 9 generations ago.

(B) TPRs for LAI in varying generations since admixture models for the simulated haplotype data. The haplotypes in this comparison had 15% AMR/60% EUR/25% AFR proportions of admixture in all autosomes. Significance level: p ≤ 0.05, ∗∗p ≤ 0.01, ∗∗∗p ≤ 0.001, and ∗∗∗∗p ≤ 0.0001.

Table 1.

LAI true positive rate estimates per ancestry per comparison

Simulation model name Analysis parameters Mean TPR (SD) per ancestry
AMR EUR AFR
Impact of demography: proportions

average 3-way 15%/60%/25%, 9gen, chr1 0.884 (0.224) 0.986 (0.018) 0.986 (0.026)
even 3-way 33%/33%/34%, 9gen, chr1 0.933 (0.082) 0.966 (0.130) 0.991 (0.008)
even 2-way 50%/50%/0%, 9gen, chr1 0.918 (0.120) 0.992 (0.008) N/A
average 2-way 70%/30%/0%, 9gen, chr1 0.935 (0.058) 0.987 (0.014) N/A
extreme proportions, high AMR 95%/5%/0%, 9gen, chr1 0.937 (0.054) 0.961 (0.134) N/A
extreme proportions, high EUR 5%/95%/0%, 9gen, chr1 0.845 (0.280) 0.998 (0.003) N/A

Impact of demography: generations since admixture

average 3-way 15%/60%/25%, 9gen, chr1 0.884 (0.224) 0.986 (0.018) 0.986 (0.026)
average 3-way 15%/60%/25%, 12gen, chr1 0.906 (0.117) 0.982 (0.016) 0.989 (0.013)
average 3-way 15%/60%/25%, 17gen, chr1 0.896 (0.162) 0.977 (0.017) 0.984 (0.011)

Features of data/analysis: reference panel

average 3-way; low-N and well-matched AMR reference 15%/60%/25%, 9gen, chr1; HGDP AMR reference 0.884 (0.224) 0.986 (0.018) 0.986 (0.026)
average 3-way; medium-N and admixed AMR reference 15%/60%/25%, 9gen, chr1; 1KG PEL as AMR reference 0.878 (0.224) 0.986 (0.018) 0.989 (0.013)
average 3-way; high-N and admixed + unmatched AMR reference 15%/60%/25%, 9gen, chr1; 1KG PEL + EAS as AMR reference 0.875 (0.223) 0.983 (0.019) 0.991 (0.009)

Features of data/analysis: data type

average 3-way 15%/60%/25%, 12gen, all autosomes; WGS density 0.935 (0.025) 0.983 (0.005) 0.989 (0.004)
average 3-way 15%/60%/25%, 12gen, all autosomes; SNP array density 0.938 (0.029) 0.977 (0.008) 0.982 (0.004)
average 3-way 15%/60%/25%, 12gen, all autosomes; imputed SNP array density 0.927 (0.026) 0.982 (0.004) 0.987 (0.003)

Features of data/analysis: window size

average 3-way 15%/60%/25%, 12gen, all autosomes, 0.1 cM 0.939 (0.023) 0.979 (0.005) 0.991 (0.003)
average 3-way 15%/60%/25%, 12gen, all autosomes, 0.2 cM (default) 0.935 (0.025) 0.983 (0.005) 0.989 (0.004)
average 3-way 15%/60%/25%, 12gen, all autosomes, 0.4 cM 0.923 (0.026) 0.984 (0.005) 0.980 (0.005)

Analysis parameters: admixture proportions of simulated cohort (AMR/EUR/AFR), number of generations since admixture, simulated chromosome, and other parameters. gen, generation; chr, chromosome; N/A, not available.

Figure 2.

Figure 2

LAI true positive rates in within-ancestry comparisons

(A) True positive rates for LAI in three reference panel comparisons that vary in the AMR component, separated by ancestry component. Benchmarking was run on the model reflecting a pulse of admixture 9 generations ago with 15% AMR/60% EUR/25% AFR proportions in chromosome 1.

(B) True positive rates for LAI in WGS versus SNP array (GSA) versus imputed data. Benchmarking was run on the model reflecting a pulse of admixture 12 generations ago with 15% AMR/60% EUR/25% AFR proportions in all autosomes.

(C) True positive rates for LAI runs varying the RFMix window size parameter in centimorgans (cM). Benchmarking was run on the model reflecting a pulse of admixture 12 generations ago with 15% AMR/60% EUR/25% AFR proportions in all autosomes. Significance level: p ≤ 0.05, ∗∗p ≤ 0.01, ∗∗∗p ≤ 0.001, and ∗∗∗∗p ≤ 0.0001.

The number of generations since admixture had a slight impact on TPRs, which modestly decreased in older admixture events (17 generations ago) compared to more recent ones (9 generations ago). We observed statistically significant differences in TPRs between the 9 and 12 generation models in the AFR and EUR components and between the 12 and 17 generation models in the AFR component (p-adj < 0.05). In all models, the 17 generations since admixture model had the lowest accuracy (Figure 1B; Tables 1 and S1). This is likely due to increased difficulty in painting short ancestry tracts; the further back in time a pulse occurred, the shorter the relative ancestral tracts will be in the current day, as recombination breaks ancestral stretches down over time.27 The general trends observed in relative TPRs per ancestry proportion were the same regardless of admixture pulse generation times.

Impact of reference panel compositions

To benchmark the performance of different reference panel compositions, we tested multiple AMR/EUR and AMR/EUR/AFR reference panel combinations comprising individuals from the HGDP and 1KG. We organized our test reference panels to reflect (1) a very well-matched panel but with low sample size, (2) a moderate-sized panel that includes admixed individuals in the reference versus restricting to only homogeneous individuals, or (3) a very large reference that is poorly matched. To illustrate the differences between panels, we ran an ADMIXTURE analysis (Figure S2).

The average TPRs were similar across the three tested reference panels and had no statistically significant differences (p-adj < 0.05; Figure 2A; Table S2), however, the small but well-matched reference panel (n = 184) resulted in a considerably faster running time compared to the admixed AMR reference panel (n = 239) and the large but unmatched reference panel (n = 659), which took three and sixteen times longer to complete an LAI run for chromosome 1, respectively. Figure 2A and Table 1 summarize the TPR results for each reference panel.

Additionally, we tested analytical scenarios where LAI would be performed on a 2-way EUR/AMR cohort with a 3-way AFR/EUR/AMR reference and on a 3-way AFR/EUR/AMR cohort with a 2-way AFR/EUR reference. In the first scenario, considering that all ancestral components present in the target haplotypes are present in the reference, TPRs did not change significantly compared with using a 2-way reference (Figure S3). In the second scenario, we observed that the TPRs for EUR and AFR were not impacted significantly compared to the 3-way reference panel analysis. However, because the AMR portion in the simulated admixed haplotypes would be completely ignored by the LAI software in this scenario (due to the lack of this component in the reference), this entire ancestral component was miscalled, which impacts the overall accuracy (Figure S4).

Impact of genetic data type

We assessed the impact of genetic data type on LAI accuracy with simulations generated from the WGS-density reference, a subset of SNP array variants, and a dataset of imputed variants (using the subset of SNP array variants as input) in the average 3-way (15% AMR/60% EUR/25% AFR) admixture model simulation considering 12 generations since admixture. We selected genomic variants targeted by the GSA chip for these tests, as this is a commonly used array for diverse datasets.

LAI run on the WGS-density simulated dataset achieved better TPRs for AFR and EUR ancestry components than the SNP array-density dataset (p-adj < 0.05), likely due to fuller haplotype coverage, but was roughly 6 times slower to complete for all autosomes. Imputation slightly improved LAI calls for these ancestry components compared to the SNP-array-only runs, indicating increased SNP density improved performance. These trends were different in the AMR component, however, in which we observed no statistically significant differences between either WGS, SNP array, or imputed datasets and observed a qualitative decrease in TPRs following imputation (the lowest TPR in this component), although it was not statistically significant (Tables 1 and S2; Figure 2B). Additionally, we performed a validation analysis of the imputation accuracy results by selecting only the original SNP array sites from the LAI results of the imputed dataset to observe if imputation changed LAI on these sites, which could lead to changes in accuracy performance. We observed no significant changes in TPRs compared to the full imputed dataset (Figure S5).

RFMix window size parameter changes do not improve LAI accuracy from the default value

As RFMix calls LA with a sliding window approach, we tested whether halving or doubling the default window size improved calls, which could change results especially at the borders of chromosomes that may only have anchoring haplotype information on one side of the window.

We found that halving the default window size to 0.1 cM did not significantly change the TPRs for the AFR and AMR components and that it did significantly lower the TPR in the EUR component compared to the default 0.2 cM (p-adj < 0.0001). Doubling the default window size to 0.4 cM significantly decreased the TPR for the AFR component (p-adj < 1e−5) but did not significantly change for the other components compared to the default (Figure 2C; Tables 1 and S2). As such, we recommend retaining the default 0.2 cM window size for RFMix runs.

We additionally examined the ForwardBackward probability estimates from RFMix to check how confidently the algorithm estimated the wrong calls, as such confidence estimates could be a readily implemented filter to remove poorly called loci. We observed, however, that miscalls had high-confidence estimates; therefore, setting a stringent filter for ForwardBackward probabilities in an attempt to reduce LAI miscalls would not be sufficient to improve the results.

LAI miscalls are more frequent in certain genomic locations but vary between cohorts

When miscalls in LAI occurred, we observed that although they may occur at any point in the genome, they were more frequent around telomere and centromere regions (Figures 3A, 3B, and S6–S11). Considering sites with over 10% miscalls in both 3-way model runs and a window of 1 kb upstream and downstream, we observed that these regions can span or flank genes (Tables S3 and S4), most of which have been previously associated in GWASs according to the GWAS Catalog (Tables S5 and S6). We compared these sites with low-complexity regions from the UCSC RepeatBrowser hg38 dataset and observed an overlap of >98% in both models. Filtering for low-complexity regions could therefore be a potential strategy for removing areas prone to miscalls. It is important to note, however, that the sites/regions with over 10% miscalls varied between the admixture models in which we ran this analysis; therefore, we do not supply a narrower list of regions that will be more frequently miscalled, as this may vary between cohorts.

Figure 3.

Figure 3

Trends in LAI wrong calls

(A) Percentage of wrong calls per site on chromosome 1, in total and separated by error mode for LAI run on the model reflecting a pulse of admixture 12 generations ago with 15% AMR/60% EUR/25% AFR proportions.

(B) Percentage of wrong calls per site on chromosome 1, in total and separated by error mode for LAI run on the model reflecting a pulse of admixture 12 generations ago with 33% AMR/33% EUR/34% AFR proportions.

(C) Miscall counts separated by error mode summing all autosomes for LAI run on the model reflecting a pulse of admixture 12 generations ago with 15% AMR/60% EUR/25% AFR proportions.

(D) Miscall counts separated by error mode summing all autosomes for LAI run on the model reflecting a pulse of admixture 12 generations ago with 33% AMR/33% EUR/34% AFR proportions.

LAI miscalls occur with a consistent error mode

We summed miscall counts and divided them between error modes, i.e., how many truth sites of one ancestry are being miscalled as each of the other ancestries. This allowed us to characterize trends in miscall directions and observe whether one ancestry was systematically being over- or under-called than another. We observed that the most common direction for miscalls to occur was for truth AMR sites to be incorrectly called EUR (Figures 3C, 3D, and S12–S15). The second most frequent error mode was EUR positions being miscalled AFR.

Discussion

In this study, we evaluated characteristics impacting the performance of LAI for a range of 2- and 3-way admixed demographic models reflective of many Latin American populations. Specifically, we assessed the impact of reference panel composition; demographic features such as the proportions of major ancestry groups and number of generations since admixture in the cohort; genetic data technology (genotyping arrays versus WGS); the impact of imputation; and LAI analytic thresholds on the performance of diverse cohorts. Given the high LAI accuracy observed in the literature for 2-way admixed AFR/EUR cohorts,8 we focused our analyses in this manuscript on determining the best practices for cohorts involving AMR, as the smaller divergence time between EUR and AMR tracts poses a challenge for deconvolution, as does the particularly limited availability of relevant reference individuals for AMR ancestry. Thus, we focused the construction of our reference panel tests in the service of optimizing AMR accuracy.

These benchmarks allow us to provide a set of recommendations for parameter and panel selection to achieve optimal LAI performance in Latin American populations. Specifically, comparing the performances of different reference panel compositions, we observed that there was no significant difference in accuracy across the three panels (well-matched but small sample size, medium size with some degree of admixture in the reference, and large but poorly matched to the target cohort), although we do observe a large difference in runtime, with the small but well-matched panel running substantially faster than the other panels. Given the high computational burden required by LAI, having a quicker runtime for analysis is an important point of consideration in practical use. As such, a curated reference panel reflective of the ancestries present in the target cohort appears to be the best option for LAI reference panel construction. Importantly, across all demographic and reference panel models tested, AMR ancestry tracts suffer from notably reduced accuracy as compared to EUR and AFR tracts. This is likely due to there being less representative (and less homogeneous) reference data for AMR ancestry in existing reference resources. Moving forward, it will be vital for efforts to focus on the ethical recruitment of more diverse and geographically distributed reference samples in large-scale data collection efforts to maximize the performance of LAI across all ancestry backgrounds.

Regarding ancestry proportions in the admixed simulations, we observed that, overall, having a higher proportion of an ancestry in the simulation improved the TPRs for that ancestry in LAI. When ancestries represent a very small (e.g., 5%) global proportion, that ancestry suffers reduced performance, with AMR suffering the most.

We investigated LAI miscalls in the realistic 3-way simulation models to evaluate the typical error mode when wrong calls are produced. Specifically, we examined the rates of miscalls for each ancestry component to characterize trends in the relative amount and direction of miscalls, see if a particular ancestry was being systematically over- or under-called, document if there were genomic regions where miscalls were most frequent, identify other factors that could be driving error modes, and assess if alterations to RFMix parameters could improve miscall rates. By investigating the typical error modes when LAI miscalls occur, we observed a much higher frequency of miscalls in the direction of calling simulated AMR regions as EUR compared to other miscall directions. This may be explained by the smaller genetic divergence between AMR and EUR haplotype tracts than either is to AFR, resulting in closer haplotype similarity. This finding implies that reference panels will need to be grown substantially to confidently assess within-continental ancestral components for many geographic regions. The specific direction of AMR/EUR miscalls being dominated in the direction of AMR to EUR, rather than vice versa, can be explained by the substantially (2×) larger sample size available for EUR compared to AMR. Another point of consideration is that the AMR reference samples themselves have some degree of admixture with EUR ancestry, which adds uncertainty to the model, though we did implement EM procedures to attempt to correct for this. These results are consistent with miscall trends observed in other studies of diverse populations.28 As this prior cited work was done with older LAI software than RFMix, we have confirmed that this error mode is consistent between different LAI algorithms and therefore likely to be driven by the genetic data rather than a feature specific to RFMix.

Beyond error modes, we observe that miscall regions do not appear randomly across the genome but are most likely to fall in areas that mark the edges of haplotypes, like centromere and telomeres (Figures 3A, 3B, and S3–S8). We note several areas that had elevated miscall rates (higher than 10% miscalls). ForwardBackward probabilities for the LAI algorithm were still confident in such areas, and tweaking RFMix parameters was insufficient to correct them. However, 98% of miscall-prone areas overlap with low-complexity regions, such as short and long interspersed nuclear elements (SINEs/LINEs), DNA repeats, and micro-satellites, offering a potential strategy for filtering problematic sites and hinting that technical artifacts in genotype calling may be driving LAI errors. This could be due to the fact that low-complexity regions are usually more challenging to map29,30 and/or are evolutionarily conserved,31,32 with little variation across ancestry groups, and therefore these regions may be more prone to error in LAI. The development of methods that incorporate repeat polymorphisms, multiallelic variants, and other complex forms of genetic variation in genome-wide analyses may help improve LAI accuracy. Additional care should be considered when analyzing these regions and/or genes in close proximity. Good LAI accuracy is vital for gene discovery and other statistical genomics efforts, as misclassification both soaks up power in LAI-informed GWASs and can lead to false positive associations due to technical ancestry miscalls.8 Importantly, miscall regions may contain genes of interest, so care should be taken to validate, for example, GWAS hits in border haplotype areas that show elevated miscall rates. The inflation of miscalls at particular regions could also impact the interpretation of other statistical genetics efforts, such as admixture mapping or evolutionary scans of selection that utilize local ancestry enrichment.

Examining how different DNA data types impact LAI performance, we observe that WGS and SNP array simulated data resulted in similar TPR estimates for genotyped sites, although having more variants in the dataset improved estimates. We also observe that LAI performs nearly as well on imputed data as directly genotyped data when a large and diverse reference panel is used. This suggests that, provided imputation can be performed with a representative reference panel, LAI calls on imputed data may be confidently utilized for downstream efforts. Out of an abundance of caution, we recommend setting a stringent INFO threshold (e.g., 0.8) for imputed sites to ensure high-confidence calls. We note, additionally, that non-significant differences in TPRs in the context of this work do not mean that the differences that we observe are not relevant, as small differences in LAI accuracy can impact statistical power in downstream applications8 and may represent a difference observed in a large number of sites in the genome. Expanding available reference samples to contain representative haplotypes from diverse and understudied populations would improve the quality of imputation as well as LAI.

Of course, this work has some important limitations that must be considered. As the focus of the present study is Latin American populations, we limited our demographic models to those involving 2-way admixture between AMR and EUR or 3-way admixture between AMR, EUR, and AFR, which represents the majority of Latinx populations. We note, however, that some Latin American populations have other patterns than those directly benchmarked here. Despite this, the broader trends in LAI performance identified in this work should hold across demographic models beyond the specific use cases simulated in this manuscript. We also note that while we appreciate that there is a high level of diversity within continental regions,14,33 only continental-level ancestry was able to be assessed here due to limitations in available reference panel geographic coverage. Similarly, having a small number of available reference AMR samples limited the number of individuals available for simulating and running LAI, which limits variability in the data for this component in comparison to EUR and AFR. Improved LAI call rates and finer-scale LAI resolution would be possible in the future if reference panels are expanded. Regarding software, we have benchmarked only RFMix v.1 in this work, as prior work has demonstrated that RFMix v.1 performed the best in comparison to other methods for multi-way admixed samples.34 We expect the trends observed here to be consistent across LAI software, though further benchmarking would be needed to confirm this.

In conclusion, in our reference panel benchmarking, the best cost-benefit in terms of LAI accuracy and speed is to use a well-matched reference even if it has a lower sample size. Examining the ancestry-specific performance of LAI across reference panels, we observed consistently lower performance for the AMR ancestry component across all simulation settings compared to EUR and AFR. Unfortunately, this inequity could not be overcome by any of the tested modifications to the reference panel, LAI software parameters, or features of genetic data. The best way to improve AMR performance would be to increase the well-matched reference panel’s sample size, underscoring the importance of furthering recruitment of larger and more representative reference samples for understudied populations. Given the high proportion of the global population that contains admixed ancestry and the fact that populations are getting increasingly admixed over time,35 it is timely to establish the optimal methods for well-calibrated genomic analyses in admixed populations.

Data and code availability

Code generated in this project for simulating admixed data and quantifying LAI TPRs is freely available on GitHub at https://github.com/Atkinson-Lab/LAI-sims-accuracy.

Acknowledgments

This study was supported by Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP; fellowship no. 2021/09584-1). Financial support for the PTSD-PGC Ancestry Working Group was provided by the National Institute of Mental Health (NIMH; R01MH106595). E.G.A. was supported by the National Institute of Mental Health (K01 MH121659 and R01 HG012869), the Caroline Wiess Law Fund for Research in Molecular Medicine, and the ARCO Foundation Young Teacher-Investigator Fund at Baylor College of Medicine.

Author contributions

J.-H.M. conducted the analysis and wrote the manuscript. A.X.M. and N.N.S. assisted with software. C.C.Z. reviewed the manuscript. C.M.N. and S.B. advised on the project. E.G.A. and M.S. conceptualized, supervised, and funded the project, as well as contributed to the manuscript. All authors reviewed and approved the final manuscript.

Declaration of interests

The authors declare no competing interests.

Published: January 2, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2024.12.005.

Contributor Information

Marcos Santoro, Email: santoro@unifesp.br.

Elizabeth G. Atkinson, Email: elizabeth.atkinson@bcm.edu.

Web resources

1000 Genomes Project, http://ftp.1000genomes.ebi.ac.uk/

GitHub, admix-simu, https://github.com/williamslab/admix-simu/

GitHub, ancestry pipeline, https://github.com/armartin/ancestry_pipeline/

GitHub, RFMix v.1, https://github.com/indraniel/rfmix

gnomAD, https://gnomad.broadinstitute.org/downloads#v3-hgdp-1kg

HapMap GRCh38 recombination map, http://bochet.gcc.biostat.washington.edu/beagle/genetic_maps/plink.GRCh38.map.zip

The Human Genome Diversity Project, ftp://ngs.sanger.ac.uk/production/hgdp/hgdp_wgs.20190516/statphase/

Plotgardener, https://phanstiellab.github.io/plotgardener/

Shapeit v.4, https://odelaneau.github.io/shapeit4/

TOPMed Imputation Server, https://imputation.biodatacatalyst.nhlbi.nih.gov/

UCSC RepeatBrowser, https://repeatbrowser.ucsc.edu/data/

Supplemental information

Document S1. Figures S1–S15
mmc1.pdf (20.6MB, pdf)
Data S1. Tables S1–S6
mmc2.xlsx (459.7KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (23.3MB, pdf)

References

  • 1.Sul J.H., Martin L.S., Eskin E. Population structure in genetic studies: Confounding factors and mixed models. PLoS Genet. 2018;14 doi: 10.1371/journal.pgen.1007309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Sohail M., Maier R.M., Ganna A., Bloemendal A., Martin A.R., Turchin M.C., Chiang C.W., Hirschhorn J., Daly M.J., Patterson N., et al. Polygenic adaptation on height is overestimated due to uncorrected stratification in genome-wide association studies. Elife. 2019;8 doi: 10.7554/eLife.39702. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Popejoy A.B., Fullerton S.M. Genomics is failing on diversity. Nature. 2016;538:161–164. doi: 10.1038/538161a. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Sirugo G., Williams S.M., Tishkoff S.A. The Missing Diversity in Human Genetic Studies. Cell. 2019;177:1080. doi: 10.1016/j.cell.2019.04.032. [DOI] [PubMed] [Google Scholar]
  • 5.Martin A.R., Gignoux C.R., Walters R.K., Wojcik G.L., Neale B.M., Gravel S., Daly M.J., Bustamante C.D., Kenny E.E. Human Demographic History Impacts Genetic Risk Prediction across Diverse Populations. Am. J. Hum. Genet. 2020;107:788–789. doi: 10.1016/j.ajhg.2020.08.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Gurdasani D., Barroso I., Zeggini E., Sandhu M.S. Genomics of disease risk in globally diverse populations. Nat. Rev. Genet. 2019;20:520–535. doi: 10.1038/s41576-019-0144-0. [DOI] [PubMed] [Google Scholar]
  • 7.Peterson R.E., Kuchenbaecker K., Walters R.K., Chen C.-Y., Popejoy A.B., Periyasamy S., Lam M., Iyegbe C., Strawbridge R.J., Brick L., et al. Genome-wide Association Studies in Ancestrally Diverse Populations: Opportunities, Methods, Pitfalls, and Recommendations. Cell. 2019;179:589–603. doi: 10.1016/j.cell.2019.08.051. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Atkinson E.G., Maihofer A.X., Kanai M., Martin A.R., Karczewski K.J., Santoro M.L., Ulirsch J.C., Kamatani Y., Okada Y., Finucane H.K., et al. Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nat. Genet. 2021;53:195–204. doi: 10.1038/s41588-020-00766-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Marnetto D., Pärna K., Läll K., Molinaro L., Montinaro F., Haller T., Metspalu M., Mägi R., Fischer K., Pagani L. Ancestry deconvolution and partial polygenic score can improve susceptibility predictions in recently admixed individuals. Nat. Commun. 2020;11:1628. doi: 10.1038/s41467-020-15464-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Padhukasahasram B. Inferring ancestry from population genomic data and its applications. Front. Genet. 2014;5:204. doi: 10.3389/fgene.2014.00204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Aschard H., Gusev A., Brown R., Pasaniuc B. Leveraging local ancestry to detect gene-gene interactions in genome-wide data. BMC Genet. 2015;16:124. doi: 10.1186/s12863-015-0283-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Duconge J., Ruaño G. The Emerging Role of Admixture in the Pharmacogenetics of Puerto Rican Hispanics. J. Pharmacogenomics Pharmacoproteomics. 2010;1 doi: 10.4172/2153-0645.1000101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Goetz L.H., Uribe-Bruce L., Quarless D., Libiger O., Schork N.J. Admixture and clinical phenotypic variation. Hum. Hered. 2014;77:73–86. doi: 10.1159/000362233. [DOI] [PubMed] [Google Scholar]
  • 14.Kehdy F.S.G., Gouveia M.H., Machado M., Magalhães W.C.S., Horimoto A.R., Horta B.L., Moreira R.G., Leal T.P., Scliar M.O., Soares-Souza G.B., et al. Origin and dynamics of admixture in Brazilians and its effect on the pattern of deleterious mutations. Proc. Natl. Acad. Sci. USA. 2015;112:8696–8701. doi: 10.1073/pnas.1504447112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Ruiz-Linares A., Adhikari K., Acuña-Alonzo V., Quinto-Sanchez M., Jaramillo C., Arias W., Fuentes M., Pizarro M., Everardo P., de Avila F., et al. Admixture in Latin America: geographic structure, phenotypic diversity and self-perception of ancestry based on 7,342 individuals. PLoS Genet. 2014;10 doi: 10.1371/journal.pgen.1004572. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Homburger J.R., Moreno-Estrada A., Gignoux C.R., Nelson D., Sanchez E., Ortiz-Tello P., Pons-Estel B.A., Acevedo-Vasquez E., Miranda P., Langefeld C.D., et al. Genomic insights into the ancestry and demographic history of South America. PLoS Genet. 2015;11 doi: 10.1371/journal.pgen.1005602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Byrska-Bishop M., Evani U.S., Zhao X., Basile A.O., Abel H.J., Regier A.A., Corvelo A., Clarke W.E., Musunuri R., Nagulapalli K., et al. High coverage whole genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022;185:3426–3440.e19. doi: 10.1016/j.cell.2022.08.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Bergström A., McCarthy S.A., Hui R., Almarri M.A., Ayub Q., Danecek P., Chen Y., Felkel S., Hallast P., Kamm J., et al. Insights into human genetic variation and population history from 929 diverse genomes. Science. 2020;367 doi: 10.1126/science.aay5012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Koenig Z., Yohannes M.T., Nkambule L.L., Zhao X., Goodrich J.K., Kim H.A., Wilson M.W., Tiao G., Hao S.P., Sahakian N., et al. A harmonized public resource of deeply sequenced diverse human genomes. Genome Res. 2024;34:796–809. doi: 10.1101/gr.278378.123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Delaneau O., Zagury J.-F., Robinson M.R., Marchini J.L., Dermitzakis E.T. Accurate, scalable and integrative haplotype estimation. Nat. Commun. 2019;10:5436. doi: 10.1038/s41467-019-13225-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.The International HapMap Consortium The International HapMap Project. Nature. 2003;426:789–796. doi: 10.1038/nature02168. [DOI] [PubMed] [Google Scholar]
  • 22.Souza A.M.d., Resende S.S., Sousa T.N.d., Brito C.F.A.d. A systematic scoping review of the genetic ancestry of the Brazilian population. Genet. Mol. Biol. 2019;42:495–508. doi: 10.1590/1678-4685-GMB-2018-0076. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Williams A. 2016. Admix-simu: Admix-simu: Program to simulate admixture between multiple populations (Zenodo) [DOI] [Google Scholar]
  • 24.Taliun D., Harris D.N., Kessler M.D., Carlson J., Szpiech Z.A., Torres R., Taliun S.A.G., Corvelo A., Gogarten S.M., Kang H.M., et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature. 2021;590:290–299. doi: 10.1038/s41586-021-03205-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Maples B.K., Gravel S., Kenny E.E., Bustamante C.D. RFMix: a discriminative modeling approach for rapid and robust local-ancestry inference. Am. J. Hum. Genet. 2013;93:278–288. doi: 10.1016/j.ajhg.2013.06.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Raghavan M., Steinrücken M., Harris K., Schiffels S., Rasmussen S., DeGiorgio M., Albrechtsen A., Valdiosera C., Ávila-Arcos M.C., Malaspinas A.-S., et al. POPULATION GENETICS. Genomic evidence for the Pleistocene and recent population history of Native Americans. Science. 2015;349 doi: 10.1126/science.aab3884. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Gravel S. Population genetics models of local ancestry. Genetics. 2012;191:607–619. doi: 10.1534/genetics.112.139808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Seldin M.F., Pasaniuc B., Price A.L. New approaches to disease mapping in admixed populations. Nat. Rev. Genet. 2011;12:523–528. doi: 10.1038/nrg3002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Tørresen O.K., Star B., Mier P., Andrade-Navarro M.A., Bateman A., Jarnot P., Gruca A., Grynberg M., Kajava A.V., Promponas V.J., et al. Tandem repeats lead to sequence assembly errors and impose multi-level challenges for genome and protein databases. Nucleic Acids Res. 2019;47:10994–11006. doi: 10.1093/nar/gkz841. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Treangen T.J., Salzberg S.L. Repetitive DNA and next-generation sequencing: computational challenges and solutions. Nat. Rev. Genet. 2011;13:36–46. doi: 10.1038/nrg3117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Enright J.M., Dickson Z.W., Golding G.B. Low Complexity Regions in Proteins and DNA are Poorly Correlated. Mol. Biol. Evol. 2023;40 doi: 10.1093/molbev/msad084. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Lenz C., Haerty W., Golding G.B. Increased substitution rates surrounding low-complexity regions within primate proteins. Genome Biol. Evol. 2014;6:655–665. doi: 10.1093/gbe/evu042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Atkinson E.G., Dalvie S., Pichkar Y., Kalungi A., Majara L., Stevenson A., Abebe T., Akena D., Alemayehu M., Ashaba F.K., et al. Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa. Am. J. Hum. Genet. 2022;109:1667–1679. doi: 10.1016/j.ajhg.2022.07.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Schubert R., Andaleon A., Wheeler H.E. Comparing local ancestry inference models in populations of two- and three-way admixture. PeerJ. 2020;8 doi: 10.7717/peerj.10090. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Institute of Medicine Committee on Future Directions for the National Healthcare Quality and Disparities Reports . National Academies Press; 2010. Future Directions for the National Healthcare Quality and Disparities Reports. [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S15
mmc1.pdf (20.6MB, pdf)
Data S1. Tables S1–S6
mmc2.xlsx (459.7KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (23.3MB, pdf)

Data Availability Statement

Code generated in this project for simulating admixed data and quantifying LAI TPRs is freely available on GitHub at https://github.com/Atkinson-Lab/LAI-sims-accuracy.


Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES