Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2023 Jan 16;110(2):336–348. doi: 10.1016/j.ajhg.2022.12.013

Pathogen exposure misclassification can bias association signals in GWAS of infectious diseases when using population-based common control subjects

Dylan Duchen 1, Candelaria Vergara 1, Chloe L Thio 2, Prosenjit Kundu 3, Nilanjan Chatterjee 3, David L Thomas 2, Genevieve L Wojcik 1,4, Priya Duggal 1,4,∗
PMCID: PMC9943744  PMID: 36649706

Summary

Genome-wide association studies (GWASs) have been performed to identify host genetic factors for a range of phenotypes, including for infectious diseases. The use of population-based common control subjects from biobanks and extensive consortia is a valuable resource to increase sample sizes in the identification of associated loci with minimal additional expense. Non-differential misclassification of the outcome has been reported when the control subjects are not well characterized, which often attenuates the true effect size. However, for infectious diseases the comparison of affected subjects to population-based common control subjects regardless of pathogen exposure can also result in selection bias. Through simulated comparisons of pathogen-exposed cases and population-based common control subjects, we demonstrate that not accounting for pathogen exposure can result in biased effect estimates and spurious genome-wide significant signals. Further, the observed association can be distorted depending upon strength of the association between a locus and pathogen exposure and the prevalence of pathogen exposure. We also used a real data example from the hepatitis C virus (HCV) genetic consortium comparing HCV spontaneous clearance to persistent infection with both well-characterized control subjects and population-based common control subjects from the UK Biobank. We find biased effect estimates for known HCV clearance-associated loci and potentially spurious HCV clearance associations. These findings suggest that the choice of control subjects is especially important for infectious diseases or outcomes that are conditional upon environmental exposures.

Keywords: genetic epidemiology, GWAS, population-based controls, common controls, misclassification bias, infectious disease


Pathogen exposure is a necessary but insufficient cause of infectious disease. Through simulation and empirical genome-wide association comparisons, Duchen et al. show that ignoring pathogen exposure can bias genetic associations. Control selection is important to accurately characterize the genetics underlying outcomes conditional upon environmental exposures, including infectious diseases.

Introduction

Genome-wide association studies (GWASs) are focused on identifying genetic associations with health outcomes. This has been successfully accomplished using well-characterized affected individuals and control subjects, often from epidemiologic cohort or case-control studies.1 More recently, GWASs have utilized biobanks and consortia to expand sample populations. This can include comparing well-characterized affected individuals to phenotypically uncharacterized or population-based common control subjects which has led to the identification of novel associations with low to modest effect sizes (see Pan-UKBB in web resources).2,3,4,5,6,7,8 Concerns related to the use of common control subjects include not adequately accounting for population substructure and the likelihood of non-differential misclassification across the outcome (i.e., some proportion of the control subjects have developed the outcome of interest) which can often attenuate true associations.2,5,7,8,9

When studying infectious disease, it is key that an individual is exposed to a pathogen (i.e., virus, bacteria, protozoa) before they can develop disease. Host genetics and immunity along with pathogen genetics, co-infections, comorbidities, age, and sex are known to explain some of the heterogeneity in disease outcomes. Thus, it is critical to account for pathogen exposure because unexposed individuals are never at risk of developing the outcome. This can be achieved through antibody or antigen testing or obtaining documented history of exposure (i.e., vaccination records, contact tracing). While it is assumed that the internal validity of case-control studies, including GWASs, is maintained by the characterization of both the genetic exposure and the phenotypic outcome for every study participant, it is also assumed that affected individuals and control subjects are at risk of developing the outcome. For infectious disease studies involving population-based common control subjects, this presumption is not always true and can result in differential misclassification of pathogen exposure between affected individuals and control subjects.10,11 Whether epidemiological factors that influence pathogen exposure (i.e., increased transmission in certain populations or occupations, routes of transmission, or comorbidities) can subsequently result in spurious genetic association signals due to this misclassification is not fully explored.

In this study, we performed simulations comparing affected individuals to well-characterized control subjects with known exposures and to population-based common control subjects with unknown exposures to identify whether external variables associated with pathogen exposure can induce spurious associations. We also used empirical data to compare the genetic association results from a GWAS performed to identify host loci associated with recovery from hepatitis C virus (HCV) infection using either known HCV-exposed and persistently infected control subjects or population-based common control subjects from the UK Biobank (UKB).

Subjects and methods

Simulations to characterize pathogen-exposure-associated selection bias

The conceptual framework depicted in Figure 1 provides a graphical representation of how comparing exposed affected individuals to population-based common control subjects differs from comparisons made to well-characterized (pathogen-exposed) control subjects. For loci associated with pathogen exposure (e.g., a causal locus for a risk factor of pathogen exposure) unrelated to the outcome of interest, if unequivocally pathogen-exposed cases are compared to population-based common control subjects regardless of exposure status, differential misclassification of pathogen exposure can induce selection bias, resulting in spurious associations between the outcome and the loci associated with pathogen exposure.

Figure 1.

Figure 1

Conceptual framework of pathogen exposure-linked selection bias

Arrows dictate the direction of hypothesized causal effects. Red arrows highlight paths involved with pathogen exposure-linked selection bias, which results in observed associations between risk factors of pathogen exposure (and their linked loci) and the outcome of interest when affected individuals are compared to population-based common control subjects with unknown pathogen exposure. Hollow arrows reflect conditional relationships between exposure and outcome/case selection. The dashed arrow reflects an association induced via the differential misclassification of pathogen exposure. The comparison of exposed affected individuals to control subjects with unknown pathogen exposure (host genetics, red line comparison) results in differential misclassification of pathogen exposure, introducing associations between pathogen exposure (and risk factors for pathogen exposure) and the outcome. No spurious association is expected when pathogen-exposed affected individuals are compared to pathogen-exposed control subjects (host genetics, black line comparison). SNP, single-nucleotide polymorphism.

How strongly pathogen exposure is associated with case status determines the strength of this bias (i.e., the dashed red line reflecting the degree of differential misclassification in pathogen exposure in Figure 1). Simulations were performed to assess whether the association between a non-outcome-associated SNP can become spuriously associated with the outcome due to a relationship with pathogen exposure. Whether the prevalence of a pathogen contributes to this bias was also determined by assessing the degree to which inflated effect estimates for the “outcome of interest”-SNP relationship were observed across the various simulation scenarios. Simulations were performed using the lavaan package (v.0.6-8) in R as it allows for the control of statistical associations between simulated variables.12,13 The phenotypic data and SNP genotypes were simulated for a large cohort (n = 1,000,000).

Terms and variables

For this study on infectious diseases, clinical outcomes may include asymptomatic infection (i.e., SARS-CoV-2 positive without symptoms), infection-associated disease (i.e., COVID-19), or post-infection sequelae (i.e., long covid). Exposure to the pathogen (i.e., SARS-CoV-2) is required for all of these outcomes and therefore pathogen exposure misclassification bias may occur. “Affected individuals” were defined as individuals with a known pathogen exposure who developed the clinical outcome. “Well-characterized control subjects” were defined as individuals with a known pathogen exposure event who did not develop the clinical outcome. To investigate the effects of differential misclassification of pathogen exposure in the absence of non-differential misclassification of the outcome, “population-based common control subjects” were defined as any individual who did not develop the clinical outcome with unknown pathogen exposure status. In each simulated cohort, a fixed number of affected individuals, well-characterized control subjects, and population-based common control subjects were randomly sampled from individuals meeting the inclusion criteria. In each simulated cohort, unless otherwise stated, the prevalence of pathogen exposure was set at 25%. Affected individuals were selected from the exposed population, of which 50% had the observed clinical outcome and the remainder were defined as well-characterized control subjects. As population-based common control subjects were selected from the entire population, excluding affected individuals, a proportion of the population-based control subjects would be expected to have a simulated pathogen exposure.

Additional simulated variables included a single-nucleotide polymorphism (SNP, coded as 0, 1, 2 minor alleles) and one “unknown” variable associated with pathogen exposure (U1). This U1 variable could represent pathogen-specific characteristics like route of transmission, modifiable lifestyle factors like occupation, or comorbidities associated with increased risk of pathogen exposure (Figure 1). To explore the effect of confounding within a locus completely unrelated to the outcome of interest, neither the SNP nor the U1 variable were simulated to be associated with the outcome. For all simulated cohorts, the SNP had a fixed minor allele frequency (MAF) of 15% and a fixed SNP-U1 association (β) of 0.1, whereby each additional SNP allele was associated with an increase of 10% of U1. Similar effect sizes have been observed for markers associated with other complex diseases.14,15,16,17,18 The SNP was simulated with a MAF of 15% and in Hardy-Weinberg equilibrium (HWE). The effect estimates for associations between simulated variables were confirmed via logistic and linear model-based regressions using the speedglm package in R.12,13

Simulation scenarios

For effects of a pathogen exposure-specific variable (simulation scenario 1), we simulated 500,000 replicates of n = 1,000,000 individuals for each of the following relationships between U1 and pathogen exposure: no association (β = 0, OR = 1), a moderate association (β = log(1.2), OR = 1.2), or a strong association (β = log(2), OR = 2). Simulations were performed assuming the prevalence of pathogen exposure was 25% and that 50% of exposed individuals had the outcome (Table 1).

Table 1.

Simulation scenario parameters

Simulation Fixed values
Sample size: Affected (n):control (n)
Outcome: Post-exposure Pathogen exposure prevalence β: U1-“pathogen exposure”
β: U1-“pathogen exposure”

Scenario 1a 50% 25% log(1), log(1.2), log(2) 20,000:20,000, 20,000:200,000
Scenario 1b 50% 25% log(1), log(1.2), log(2) 20:000:20:000

Pathogen-exposure prevalence

Scenario 2a 50% 5%, 25%, 50%, 75%, 100% log(1.2), log(2.0) 20,000:20,000, 20,000:200,000
Scenario 2b 50% 5%, 25%, 50%, 75%, 100% log(1.2), log(2.0) 20:000:20:000

For each scenario, the parameter of interest (U1-“pathogen exposure” or pathogen-exposure prevalence), the fixed values, and the sample sizes of each simulation is listed. Scenario 1a, affected individuals vs. population-based control subjects; scenario 1b, affected individuals vs. well-characterized control subjects; scenario 2a, affected individuals vs. population-based control subjects; scenario 2b, affected individuals vs. well-characterized control subjects.

For effects of the prevalence of pathogen exposure (simulation scenario 2), we simulated 500,000 replicates of n = 1,000,000 individuals for each of the following pathogen exposure prevalences: 5%, 25%, 50%, 75%, or 100%. While these pathogen exposure prevalence parameters were selected to provide generalizable results, they reflect biologically plausible prevalence estimates for multiple viral pathogens. For example, in certain populations the prevalence of cytomegalovirus or Epstein-Barr virus range from 40% to >90% by geography or age, respectively (see Pan-UKBB in web resources),4 while hepatitis B or C virus ranges from <5% to >30% of many populations.5,6,7 Simulations were performed assuming either a moderate (β = log(1.2), OR = 1.2) or a strong (β = log(2), OR = 2) association between U1 and pathogen exposure and 50% of exposed individuals had the outcome (Table 1).

In each simulated replicate, 20,000 affected individuals, 20,000 well-characterized control subjects, and 20,000 population-based control subjects were randomly selected and included in logistic regressions to quantify the “outcome of interest”-SNP association. Regressions compared affected individuals to population-based control subjects (scenarios 1a and 2a) or well-characterized control subjects (scenarios 1 and 2b). To simulate a realistic use-case of GWAS involving many common control subjects, we re-performed the above simulations and regressions for both scenarios comparing 20,000 affected individuals to 200,000 population-based control subjects. Regressions were performed in R using the Rfast package.19

To test whether the effect sizes of the “outcome of interest”-SNP associations (βSNP) were significantly different when affected individuals were compared to population-based or well-characterized control subjects, we used a modified Z statistic to compare whether the obtained regression coefficients were significantly different from one another, as a standard Z test can result in increased type 1 errors for comparisons of regression coefficients.20 Rather than use the estimated SE of the difference between groups as the denominator to estimate each scenario’s parameter-specific Z score, the square root of the summed average sample variance terms was used. With two logistic regressions for each of the parameter-specific 500,000 replicates, 13 million regressions were performed when comparing 20,000 affected individuals to 20,000 population-based control subjects.

As a null association for the “outcome of interest”-SNP relationship was simulated for each cohort, any non-null observed “outcome of interest”-SNP association reflects spurious signals induced via the scenario-specific selection of affected individuals and control subjects. We calculated the proportion of cohorts with “outcome of interest”-SNP associations that reached the conservative but widely accepted genome-wide significance threshold of p ≤ 5 × 10−8.21 To determine whether the prevalence of pathogen exposure was associated with the magnitude of these spurious signals, we performed linear regressions to obtain the magnitude and direction of the average beta estimates obtained from each set of pathogen exposure prevalence parameter-specific simulations. Measures of heterogeneity were estimated using the meta R package for each set of parameter-specific simulations to confirm independent yet equivalent cohorts were simulated (Table S1).22 Additional simulations were considered and are included in the supplemental methods.

HCV GWAS using affected individuals and well-characterized control subjects or population-based control subjects

We leveraged data from The HCV Extended Genetic Consortium (HCV Consortium) which includes 1,869 individuals of African ancestry, 1,739 individuals of European ancestry, and 486 individuals of Hispanic ancestry who passed GWAS-related quality control metrics,23 as previously described.24,25,26 Consent was obtained for genotyping as approved by the governing institutional review board and DNA provided to Johns Hopkins School of Medicine without identifiers, as described previously.24,26 We only used individuals of European ancestry in this study (tightly clustered with the 1000 Genomes Project [1000G] Northern European population from Utah [CEU] and with an admixture estimate of at least 90% European ancestry, determined via fastSTRUCTURE).27 Of the 1,739 European ancestry HCV Consortium individuals, 702 were individuals with serological evidence of spontaneous clearance from HCV infection (affected individuals) and 1,037 were defined as individuals with persistent HCV infection (“well-characterized control subjects”).24 Of these individuals, 16% were living with HIV and 31% were female.24,25,26

The UK Biobank (UKB) is a large population-level cohort including 500,000 volunteers recruited from across the UK.28 “Population-based control subjects” were 370,702 unrelated individuals from the UKB who were identified to be of similar genetic ancestry to HCV Consortium individuals of European ancestry. For these individuals, we had no information on HCV infection status or previous exposure to HCV.

Selection of samples for analysis

Linkage disequilibrium (LD)-based pruning was used to iteratively remove variants in LD (r2 > 0.2) via a sliding 500k base-pair-wide window prior to performing principal components analysis (PCA) in the combined HCV Consortium-UKB cohort. This LD-based pruning was performed on all genetic markers with missingness <1.5% and MAF >2.5%, after excluding regions of long-range linkage disequilibrium.29 Given the extreme sample size of the combined HCV-UKB cohorts, FastPCA was used.30,31 LD-based pruning and PCA was carried out using the R package SNPRelate.32

Out of 487,409 UKB individuals with genetic data available, 407,192 UKB participants who passed previously described QC metrics were included within our analyses.28,33,34 We limited UKB population-based control subjects to 370,702 ancestry-matched individuals genetically similar to the European individuals from the HCV Consortium (Figures S1 and S2).35 In brief, genetic ancestry-based matching of affected individuals and control subjects was accomplished through multiple iterations of PCA. Initially, the UKB participants whose top three PCs were within six standard deviations of the mean PC values estimated from HCV Consortium individuals were retained for analysis (n = 383,724). A second PCA filtering step, using the top twenty PCs, excluded any UKB individual with estimated eigenvalues 5% higher or lower than the range of PC-specific values from HCV Consortium individuals. A total of 370,702 ancestry-matched UKB control subjects were identified.

Selection of markers for analysis

From the HCV Consortium, out of an initial 661,397 directly genotyped markers mapped to the GRCh37/hg19 reference sequence, a total of 656,340 markers had MAF >1%, variant-level call rates >97%, were in Hardy-Weinberg equilibrium (HWE) based on HWE exact tests (p > 1 × 10−5), and passed the “McCarthy Group Tools” Haplotype Reference Consortium (HRC) imputation preparation pipeline (https://www.well.ox.ac.uk/∼wrayner/tools/).36 No allele frequency difference threshold between this database and the HCV Consortium cohort individuals was used to remove genotyped markers. Imputation was performed using the HRC (v.r1.1, 2016) European reference panel via the Michigan Imputation Server and phased using Eagle (v.2.4).37,38 Markers with imputation quality R2 values >0.3 were retained and markers which were not bi-allelic or had a genotyping rate <97% or MAF <1% were removed. Additionally, imputed markers with variant-level INFO scores <90% or HWE p < 5 × 10−7 were removed. A total of 6,468,618 imputed markers passed these QC metrics.

From UK Biobank, a total of 670,739 directly genotyped autosomal QC-passed markers were used for imputation, as previously described.28 Phasing was performed with SHAPEIT3 and imputation carried out using IMPUTE4.28,39,40 UKB imputation involved a combination of the Haplotype Reference Consortium (HRC), the UK10K haplotype reference, and the 1000 Genomes phase 3 reference panels.28 To be considered a potential match for any QC-passed HCV Consortium-derived marker, the imputed UKB variant was required to have INFO score >90%, MAF>1%, and HWE p > 1 × 10−7. A total of 93,095,623 imputed UKB variants met this quality control threshold and were considered for analysis.

Imputed markers in the HCV Consortium were matched to UKB QC-passed imputed markers using position and allele information. A total of 6,025,969 markers were shared across the ancestry-matched UKB and HCV Consortium databases. A total of 6,009,835 markers were shared across the combined HCV Consortium and a set of 364,308 UKB control subjects who passed a more stringent set of sample-level genetic heterozygosity-focused QC metrics than those performed by the UKB Consortium, as described below and in the supplemental methods.

Additional exploratory GWASs using different subsets of UKB control subjects were performed to determine whether the observed association between certain genetic loci and HCV clearance was driven by mismatched fine-scale ancestry or by excess sample-level genetic heterozygosity, or was unaccounted for population structure due to the extreme number of population-based UKB control subjects (Figures S3–S8). Other exploratory GWASs assessed differential enrichment for certain epidemiological factors relative to the population-based common control subjects. This included a GWAS comparing the set of HCV clearance cases from non-hemophilia-focused cohorts (n = 513) to a more precisely matched case-control cohort of UKB control subjects (n = 4,908) (Figure S9). A GWAS comparing all the individuals from the HCV Consortium (n = 1,739) to the European ancestry-matched population-based controls from the UKB (n = 370,702) was also performed (Figure S10). Each exploratory GWAS occurred after performing all sample-level and imputed marker-level QC described above.

Statistical analysis

Association of affected individuals and well-characterized control subjects

Genetic associations between the 702 affected individuals and 1,037 well-characterized control subjects of the HCV Consortium were performed via logistic regression under an additive genetic model using PLINK,31,41 as the lack of any significant sample size imbalance between affected individuals and control subjects or intensive computational resources needed to perform these genetic associations did not warrant a more complex association approach. Covariates for this GWAS included sex and the first twenty principal components, estimated from the HCV Consortium dataset after linkage disequilibrium-based pruning.

Association of affected individuals and population-based control subjects

To account for the severe sample size imbalance between affected individuals and control subjects and to optimize computational performance at extreme differences in sample sizes, associations between the 702 affected individuals and 370,702 ancestry-matched population-based control subjects were performed under an additive genetic model using REGENIE with the saddle point approximation (SPA) test.3,42,43 While REGENIE partially accounts for population stratification, the top twenty PCs estimated from the combined HCV-UKB cohort were included as fixed effect covariates to ensure differences due to genetic ancestry or population structure among UKB individuals of European ancestry were accounted for (Figure S2).44,45 Sex was included as a covariate.

We utilized a genome-wide significance threshold of p ≤ 5 × 10−8.21 After performing the association testing, to minimize the risk of bias due to differences across genotyping arrays or imputation panels between our affected individuals and control subjects among putative HCV clearance-associated loci, a MAF-based filter was used to exclude markers suggestively associated with HCV clearance (p ≤ 5 × 10−5) using genome-wide allele frequency data for the gnomAD non-Finnish European population downloaded in September 2021, using the gnomAD v.2.1.1 release.9 Any suggestively associated marker with a MAF > 10% among gnomAD non-Finnish Europeans was excluded if the MAF among UKB control subjects and gnomAD non-Finnish Europeans differed in absolute terms by at least 5%. Similarly, any suggestively associated marker with a MAF < 10% in gnomAD non-Finnish Europeans was excluded if the relative difference between the gnomAD-derived MAF and UKB MAF was greater than 25%.

Manhattan and quantile-quantile plots for each GWAS were generated using the ggfastman or qqman packages in R (Figure S14).46 A table of all suggestively associated markers (p < 5 × 10−5) for the ancestry-matched UKB control subjects GWAS can be found in Table S2.

We performed Cochran’s Q test of heterogeneity using METAL to determine whether the results obtained from the GWAS of affected individuals vs. population-based control subjects from the UKB differed significantly from the GWAS of affected individuals vs. well-characterized control subjects from the HCV Consortium (Figure S11).47,48

Results

Simulations to characterize pathogen exposure-associated selection bias

Simulations were performed using a simplified version of the framework presented in Figure 1. In brief, the association between an SNP not associated with outcome but associated with pathogen exposure via the unknown variable U1, was compared between pathogen-exposed affected individuals and pathogen-exposed (well-characterized) or population-based common control subjects.

First, we simulated whether the presence and magnitude of the U1-“pathogen exposure” relationship affects the association between an SNP linked to U1 and the outcome of interest (Table 1). To provide a baseline for further explorations, the first model was simulated to have no association between U1 and pathogen exposure (OR = 1). As expected, we found no association between the SNP and outcome, regardless of the choice of well-characterized or population-based control subjects or whether affected individuals were compared to 20,000 population-based common control subjects or 200,000 population-based common control subjects (p = 1) (Figure 2). For all simulation scenarios, no replicates comparing affected individuals and well-characterized pathogen-exposed control subjects resulted in spurious “outcome of interest”-SNP associations (Table 2).

Figure 2.

Figure 2

Distribution of observed odds ratios for a non-outcome-associated locus across varying U1-“pathogen exposure” associations

Comparisons involved 20,000 affected individuals and equal numbers of well-characterized or population-based control subjects (left) or 20,000 affected individuals and equal numbers of well-characterized and 200,000 population-based control subjects (right). Boxplots reflect distribution of odds ratios obtained comparing affected individuals to population-based (red) or well-characterized (yellow) control subjects. Reported p values derived from a Z score based on the difference between averaged beta estimates when using population-based control subjects vs. well-characterized control subjects.

Table 2.

Proportion of spurious associations and odds ratios across scenario-specific simulated cohorts

Parameter value U1-“pathogen exposure” Controls 20,000 affected vs. 20,000 common controls
20,000 affected vs. 200,000 common controls
OR (mean) Spurious associations n (%) OR (mean) Spurious associations n (%)
Scenario 1

U1-“pathogen exposure” OR = 1 pathogen exposed 1 0 1 0
OR = 1.2 pathogen exposed 1 0 1 0
OR = 2 pathogen exposed 1 0 1 0
U1-“pathogen exposure” OR = 1 population-based 1 0 1 0
OR = 1.2 population-based 1.04 159 (0.03) 1.04 1,792 (0.36)
OR = 2 population-based 1.13 410,503 (82.1) 1.13 499,666 (99.93)

Scenario 2

Pathogen-exposure prevalence: 5% OR = 1.2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.06 2,830 (0.57) 1.06 36,815 (7.36)
OR = 2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.2 499,962 (99.99) 1.197 500,000 (100)
Pathogen-exposure prevalence: 25% OR = 1.2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.04 153 (0.03) 1.04 1,749 (0.35)
OR = 2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.13 407,864 (81.57) 1.132 499,647 (99.93)
Pathogen-exposure prevalence: 50% OR = 1.2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.03 13 (<0.01) 1.03 136 (0.03)
OR = 2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.095 101,203 (20.24) 1.095 399,122 (79.82)
Pathogen-exposure prevalence: 75% OR = 1.2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.019 3 (<0.01) 1.019 10 (<0.01)
OR = 2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1.06 2,725 (0.55) 1.059 35,680 (7.14)
Pathogen-exposure prevalence: 100% OR = 1.2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1 1 (<0.01) 1 0 (0)
OR = 2 pathogen exposed 1 0 (0) 1 0 (0)
population-based 1 1 (<0.01) 1 0 (0)

For each scenario, the parameter of interest, the fixed values, and the sample sizes of each simulation is listed. OR, odds ratio.

For cohorts with a moderate association between U1-“pathogen exposure” (OR = 1.2), there was an observed but non-significant inflation for comparisons involving 20,000 population-based common control subjects compared to 20,000 well-characterized control subjects (p = 0.15), based on the Z score estimate derived using the average effect estimates obtained when affected individuals were compared to population-based or well-characterized control subjects. Additionally, a small number of replicates (n = 159/500,000) had spurious genome-wide significant associations when affected individuals were compared to population-based common controls. When the association between U1-“pathogen exposure” was simulated to be stronger (OR = 2), significantly inflated odds ratios were observed for comparisons involving 20,000 population-based common control subjects (p = 5.3 × 10−6) (Figure 2). For these simulations, 82.1% of replicates (n = 410,503/500,000) had spurious associations (p < 5 × 10−8) when affected individuals were compared to 20,000 population-based common control subjects (Table 2).

To investigate a scenario closer to current practice, we simulated unbalanced datasets with 20,000 affected individuals and 200,000 population-based control subjects. In these simulations the reduced variance due to increased statistical power across “outcome of interest”-SNP estimates resulted in similarly inflated but higher proportions of spurious genome-wide significant replicates compared to replicates involving 20,000 population-based common control subjects for both moderate (n = 1,792/500,000, p = 0.11) and strong (n = 499,666/500,000, p = 1.9 × 10−7) U1-“pathogen exposure” association scenarios (Figure 2, Table 2).

In the second set of simulations we varied the pathogen exposure prevalence from 5% to 100% to consider different infectious disease prevalence scenarios. All prevalence values except for universal exposure (100%) resulted in inflated effect estimates for comparisons involving population-based common controls when assuming either moderate (OR = 1.2) or strong (OR = 2) U1-“pathogen exposure” relationships, although only the strong exposure was significant (Figures 3 and S12). Using 20,000 affected individuals vs. 200,000 population-based control subjects, we observed increased spurious associations for rare pathogen exposure prevalence (5%, n = 36,815/500,000) which decreased to 0 replicates with spurious associations as the pathogen exposure approached 100% (or all exposed). This was present for both moderate and strong U1-“pathogen exposure” associations (Table 2), suggesting that any observed bias was related to the differential misclassification of pathogen exposure when affected individuals were compared to population-based control subjects. For comparisons involving 20,000 affected individuals and 20,000 population-based control subjects, a single spurious replicate was observed for the scenario involving 100% pathogen exposure prevalence (Table 2).

Figure 3.

Figure 3

Distribution of observed odds ratios for a non-outcome strongly associated locus across varying pathogen-exposure prevalence

Comparisons involved 20,000 affected individuals and equal numbers of well-characterized or population-based control subjects (left) or 20,000 affected individuals and equal numbers of well-characterized and 200,000 population-based control subjects (right) assuming a strong U1-“pathogen exposure” relationship. Boxplots reflect distribution of odds ratios comparing affected individuals to population-based (red) or well-characterized (yellow) control subjects. Reported p values derived from a Z score based on the difference between averaged beta estimates when using population-based control subjects vs. well-characterized control subjects.

We also evaluated whether the number of affected individuals altered the magnitude of any inflated effect estimates since infectious disease studies are often much smaller than other complex disease studies due to availability and collection of samples. When assuming a strong U1-“pathogen exposure” relationship and fixing other parameters to those utilized in the first scenario simulations, no observable difference in the mean inflation of the “outcome of interest”-SNP association (OR ≈ 1.13) was observed for comparisons between varying numbers of affected individuals (500, 1,000, 5,000, 10,000, and 20,000) and 200,000 population-based controls (Figure S13), suggesting that the bias exists regardless of case sample size.

HCV GWAS using affected individuals and well-characterized or population-based control subjects

To assess the real-world consequences of differential pathogen exposure misclassification in GWASs of infectious disease, we explored the effects of control definitions using a previously published GWAS of HCV spontaneous clearance versus persistence.24 All affected individuals and control subjects included within this previous GWAS were unequivocally exposed to HCV and the outcome of interest focused on a specific clinical sequela of infection, HCV clearance. We compared the results of this HCV-clearance GWAS with the well-characterized (persistently HCV-infected) control subjects to the same affected individuals compared to population-based common control subjects from the UKB. The GWAS comparing HCV clearance affected individuals to persistently infected individuals replicated previously published loci.26,49,50,51

Replication of known HCV clearance-associated loci

In the GWAS using population-based UKB control subjects (702 affected individuals vs. 370,702 population-based control subjects), the two primary HCV clearance-associated loci within the human leukocyte antigen (HLA) region (HLA-DQB1) and IFNL3 locus were replicated (p < 5 × 10−8; Figure 4). However, the effect sizes of these markers were biased toward the null compared to the GWAS using well-characterized control subjects (Table S3). The genome-wide significant IFNL3 locus markers had odds ratios in the range of 0.37–0.45 when affected individuals were compared to well-characterized control subjects, while there were odds ratios of 0.56–0.63 for the population-based UKB controls comparison. The difference in effect size for each of these markers was significant (Cochran’s Q test, p < 0.05, Figure S11), including the chromosome 19 marker in Table 3, and consistent with bias toward the null due to non-differential outcome misclassification from using population-based control subjects.

Figure 4.

Figure 4

Manhattan plot of the GWASs comparing affected individuals with HCV clearance to persistently infected well-characterized control subjects and ancestry-matched population-based UKB control subjects

The top panel reflects results from a GWAS performed comparing HCV-clearance individuals (n = 702) to individuals persistently infected with HCV from the HCV Consortium (n = 1,037). The bottom panel reflects the results of a GWAS performed comparing the same HCV-clearance individuals to ancestry-matched population-based common control subjects from the UKB (n = 370,702). The same set of genome-wide markers were used in each GWAS (6,025,969 markers). Points reflects the p values for each marker across the genome, ordered along the X axis by chromosomal position. The Y axis reflects the −log (p value), with the most significantly associated markers farthest away from 0 and outside the genome-wide significance threshold (p < 5 × 10−8), indicated by the red lines. Loci of interest are highlighted and annotated with their nearest/overlapping gene, with novel HCV clearance-associated loci highlighted in the bottom panel.

Table 3.

HCV clearance-associated loci

Analyzed groups Chromosome 6: HLA-DQB1, rs9275241
Chromosome 19: IFNL3, rs11881222
Chromosome 4: STX18, rs58612183
OR 95% CI p value Affected (EAF, G) Controls (EAF, G) OR 95% CI p value Affected (EAF, G) Controls (EAF, G) OR 95% CI p value Affected (EAF, C) Controls (EAF, C)
Affected vs. well-characterized controls (HCV persistence) 0.61 0.53–0.71 1.90 × 10−11 39.46% 51.49% 0.42 0.36–0.50 4.06 × 10−23 20.09% 36.31% 1.22 0.97–1.54 8.40 × 10−2 10.97% 9.50%
Affected vs. ancestry- matched population-based controls (UKB) 0.69 0.62–0.77 4.35 × 10−11 39.46% 52.12% 0.61 0.54–0.69 4.04 × 10−16 20.09% 28.15% 1.93 1.56–2.41 3.01 × 10−9 10.97% 6.33%
No hemophiliacs: affected vs. ancestry-matched population-based controls (UKB, Matched 1:10) 0.70 0.61–0.80 1.95 × 10−7 39.08% 47.85% 0.60 0.51–0.70 5.42 × 10−10 20.27% 29.49% 1.47 1.17–1.84 8.55 × 10−4 9.55% 6.65%

Results of the association for the top HCV clearance-associated markers within the HLA-DQB1, IFNL3, and STX18 loci from the GWAS comparing affected individuals with well-characterized control subjects from the HCV Consortium, ancestry matched population-based common UKB control subjects, and the analysis of all affected individuals vs. all well-characterized control subjects GWAS and a matched case-control cohort GWAS excluding HCV-clearance affected individuals with hemophilia and their matched population-based control subjects. Measures of association and frequency of the effect allele for each locus are provided for each analysis. OR, odds ratio; 95% CI, 95% confidence interval of the OR; EAF, effect allele frequency.

Use of population-based control subjects identifies previously unreported HCV clearance-associated loci

In addition to the replication of known loci, two HCV clearance-associated loci were observed when affected individuals were compared to population-based common control subjects from the UKB. The first locus was on chromosome 4 within an intron of syntaxin 18 (STX18) (rs58612183, MAFUKB = 6.3%, OR = 1.93, p = 3.01 × 10−9, Figure 4) which had a reduced OR and was not significant in the GWAS involving well-characterized control subjects (MAFPersistence = 9.5%, OR = 1.22, p = 0.084). Interestingly, the minor allele frequency in the control groups differed but was within the European range of 6%–10%.52 In the GWAS comparing all HCV European ancestry individuals from the HCV Consortium (persistence and clearance, n = 1,739) to population-based UKB control subjects, a genome-wide significant association was identified at this locus (p = 3.9 × 10−10), suggesting that this locus is not HCV clearance specific but reflects any HCV infection (or membership in the HCV Consortium) compared to population control subjects with unknown viral exposure. The inflated signal in STX18 remained associated with HCV clearance in additional exploratory GWASs except for the associations excluding affected individuals known to have hemophilia and their ancestry-matched UKB control subjects (p = 8.55 × 10−4, OR = 1.47) (Table 3).

The second locus was on chromosome 2 within the long noncoding RNA (lncRNA) MIR3681HG (top marker rs10803744, MAFUKB = 12.5%, OR = 1.57, p = 7.99 × 10−8). Interestingly, the allele frequency differs between HCV-clearance (MAF = 17%) and HCV-persistent (MAF = 13%) individuals (OR = 1.42, p = 4.06 × 10−4) with European general populations matching the HCV-persistent (ensemble MAF 11%–15%).52 While some of the exploratory GWAS efforts resulted in less extreme effects for the top MIR3681HG marker, no set of exploratory GWAS consistently eliminated this signal (supplemental methods, Table S4).

To determine whether associations were driven by the imbalance in proportion of affected individuals to control subjects, we performed 1:1 and 1:10 matching for each affected individual with population-based common control subjects via two different commonly used matching methods (Mahalanobis distance and propensity score-based matching). While the Mahalanobis matched cohorts resulted in deflated estimates for the chromosome 2 MIR3681HG signal (OR ≈ 1.44, 1.46, respectively) closer to the HCV clearance vs. HCV persistence GWAS than all ancestry-matched UKB control subjects, 1:10 matching resulted in the locus remaining suggestively associated (p = 6.76 × 10−7). The use of propensity score-based matching resulted in a further inflation in the effect estimate for the locus (OR ≈ 1.76, 1.63, respectively; Table S4). It is possible that local ancestry/population structure is driving this signal.

Discussion

Understanding the circumstances when using common control subjects can put the internal validity of a GWAS at risk is necessary, especially with the growing availability of common control subjects in resources like biobanks and national cohorts.53 While previous work using population-based control subjects has suggested that non-differential misclassification of outcome can be partially compensated by increased sample sizes and statistical power,2,5,6,7,9 this current work demonstrates that for studies of infectious diseases, the comparison of affected individuals to control subjects of unknown pathogen exposure can result in spurious associations.

The simulations show that ignoring pathogen exposure among control subjects can result in a more inflated effect estimate for loci associated with exposure when the prevalence of the pathogen is rare compared to when it is common. Thus, GWASs involving common control subjects investigating sequelae associated with endemic viruses like cytomegalovirus or Epstein-Barr virus may be less susceptible to exposure-linked selection bias since it is likely all adults have been exposed.54,55 However, infectious outcomes specific to HIV, tuberculosis, or malaria need to consider whether the exposure profiles of the selected control subjects approximate that of the affected individuals. Ideally, all control subjects would be evaluated and tested, but if that is not feasible on a large scale then sampling from high endemic areas or from subpopulations with increased burdens of disease could reduce the risk of spurious associations. Similarly, simulations indicate that epidemiological factors even moderately associated with pathogen exposure in a population where the pathogen of interest is rare can result in spurious genetic associations. This bias may be especially problematic for emerging pathogens like SARS-CoV-2 where certain comorbidities and demographic characteristics may be associated with SARS-CoV-2 exposure (i.e., occupation, living conditions) and the prevalence of the viral exposure depends upon changing political, social, economic, immune, and viral factors.10,56

For example, in studies of disease severity and death due to COVID-19 involving population-based common control subjects, several associations have been observed within loci associated with risk factors for SARS-CoV-2 infection like blood type (e.g., ABO),5,57,58 obesity,59 alcohol use,60 and non-European genetic ancestry.58,59 Most notable is the ABO signal in chromosome 9, which is the most significantly associated locus for SARS-CoV-2 infection and significantly associated with both hospitalization and critical illness due to COVID-19 when using population-based control subjects, according to the COVID-19 Host Genetics Initiative’s meta-analysis results (r6 release).5 However, these markers fail to reach nominal statistical significance (p ≤ 0.05) when control subjects are limited to non-hospitalized individuals with COVID-19 infection,5 suggesting the ABO locus may not be a valid association. Loci associated with risk factors for SARS-CoV-2 could also result in spurious associations for other infectious disease-related outcomes when using common control subjects, as blood type is similarly associated with susceptibility to various bacterial infections and SARS,61,62 obesity with influenza and pneumonia,63,64 and alcohol use with contracting tuberculosis,65 HIV,66 and pneumonia.67 However, multiple COVID-19 severity-associated loci identified using population-based common control subjects may reflect real associations, as they were also successfully identified in GWASs limited to SARS-CoV-2-exposed control subjects (e.g., LZTFL1, IFNAR2).

Using empirical GWAS data we show that population-based common control subjects replicated HCV-clearance-associated loci, but markers within the known gene association locus of IFNL3 were significantly deflated (OR ≈ 0.6) as compared to the associations with pathogen-exposed control subjects (OR ≈ 0.4), likely due to non-differential misclassification. In contrast, the novel HCV clearance association within STX18 is likely spurious as it is attenuated after the removal of hemophilia-affected individuals and their matched control subjects. Relatively rare in the general population, hemophilia is a major risk factor for HCV exposure,68,69 and this association may have been driven by enrichment of hemophiliacs, or some other risk factor potentially related to the use of blood products, among affected individuals compared to UKB control subjects (U1-“pathogen exposure”). STX18, a member of the soluble N-ethylmaleimide-sensitive factor attachment protein receptors (SNARE) protein family and involved in vesicular transport, is not known to be associated with HCV clearance or infection and no association was observed for this locus in GWAS involving pathogen-exposed control subjects. Similar to our simulation results, the magnitude of the inflated effect for this locus remained largely unchanged when affected individuals were compared to fewer common control subjects (Table S4), suggesting that simply reducing the number of population-based control subjectsto approximate the number of affected individuals does not avoid the consequences of pathogen exposure misclassification.

The HCV association with lncRNA MIR3681HG remains unclear. The majority of previously published GWAS associations for this locus involve UKB control subjects,70,71,72,73 with significantly associated phenotypes including height,74 BMI,75 and educational attainment.76 Systemic differences between the UKB and the UK population have been described, with UKB participants being healthier than the general population (i.e., a “healthy volunteer” bias) and genetic loci have been identified which are linked participation within the UKB.73,77,78,79,80 These participation-associated loci have also been linked to participation and educational attainment within the UKB.80 Interestingly, the MIR3681HG signal in our GWAS involving UKB participants is also associated with BMI and educational attainment, further suggesting this may be an inflated effect. A reported association with COVID-19 susceptibility is also noted.81 Identified in a series of GWASs using cross-sectional snapshots of UKB participant data,81 MIR3681HG reached genome-wide significance once but failed to remain significant despite increasing cases of COVID-19.81

Genetic studies aim to limit spurious findings by setting clear quality control measures and stringent significance thresholds and by requiring replication of findings to reduce the costly consequences of implementing translational analysis of spurious signals. We should note that predicting the consequences of pathogen exposure misclassification is complicated by pathogen-specific host interactions, co-evolution between pathogen and human genomes, and heterogeneity between populations of both allele frequencies across the genome in addition to pathogen prevalence. These findings do not detract from the utility of GWASs involving population-based common control subjects, which remains an important and valuable approach to interrogate the genetic determinants of human health and disease. Rather, we recommend that findings from infectious disease-focused association studies involving common control subjects be interpreted with caution, considering the environmental and pathogen-specific context of the desired study outcome.

In sum, the use of population-based common control subjects may be more problematic for infectious disease-focused GWASs than previously described depending on the probability of exposure to the infectious pathogen.6,7 While true disease-associated loci can be identified, concerns related to selection bias, confounding, and misclassification are exacerbated by the inability to account for pathogen exposure among common control subjects.10 Efforts to increase the selection of control subjects from areas with high prevalence or endemicity of disease or selection of older age participants if there is known childhood exposure may attenuate the risk of false findings. Otherwise, control subjects should be carefully selected and screened for pathogen exposure.

Data and code availability

Access to individual-level phenotypic and genetic data from HCV extended genetics consortium individuals can be requested via dbGaP (phs000248.v1.p1). Access to individual-level data from the UKB can be requested at https://www.ukbiobank.ac.uk. Summary statistics for each GWAS performed can be viewed and downloaded at https://my.locuszoom.org/ (Study name: HCV Consortium vs. UKB). Code used to perform simulation experiments, bipartite PCA-based matching, and gnomAD MAF-based variant-level filtering can be accessed at https://github.com/dduchen/Population_Based_Controls_GWAS_ID_Manuscript.

Acknowledgments

Funding from the National Institute of Allergy and Infecitous Diseases, project number 2R01AI148049, along with a COVID-19 supplement under the same grant number (D.L.T., P.D., G.L.W.) and Burroughs Wellcome Fund, MD-GEM training grant (D.D.), supported this study. Access and use of the UK Biobank was approved using application number 17712. G.L.W. was additionally supported by the National Human Genome Research Institute (NHGRI) grant R35HG011944.

Author contributions

The study was designed by D.D., C.C., G.L.W., and P.D. Data collection was led by (D.L.T., C.L.T, N.C., P.K., P.D.). Simulation was performed by D.D. Statistical analyses were designed and performed by D.D., C.C., G.L.W., and P.D. Manuscript was first drafted by D.D., C.C., G.L.W., and P.D. All authors contributed to the final manuscript.

Declaration of interests

All authors declare no competing interests.

Published: January 16, 2023

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2022.12.013.

Web resources

Supplemental information

Document S1. Figures S1–S17 and Tables S1–S5
mmc1.pdf (3.3MB, pdf)
Document S2. Article plus supplemental information
mmc2.pdf (4.6MB, pdf)

References

  • 1.Visscher P.M., Wray N.R., Zhang Q., Sklar P., McCarthy M.I., Brown M.A., Yang J. 10 Years of GWAS Discovery: Biology, Function, and Translation. Am. J. Hum. Genet. 2017;101:5–22. doi: 10.1016/j.ajhg.2017.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Wellcome Trust Case Control Consortium. Clayton D.G., Cardon L.R., Craddock N., Deloukas P., Duncanson A., Kwiatkowski D.P., McCarthy M.I., Ouwehand W.H., Samani N.J., et al. Genome-wide association study of 14, 000 cases of seven common diseases and 3, 000 shared controls. Nature. 2007;447:661–678. doi: 10.1038/nature05911. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zhou W., Nielsen J.B., Fritsche L.G., Dey R., Gabrielsen M.E., Wolford B.N., LeFaive J., VandeHaar P., Gagliano S.A., Gifford A., et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nat. Genet. 2018;50:1335–1341. doi: 10.1038/s41588-018-0184-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Canela-Xandri O., Rawlik K., Tenesa A. An atlas of genetic associations in UK Biobank. Nat. Genet. 2018;50:1593–1599. doi: 10.1038/s41588-018-0248-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Callaway E. Mapping the human genetic architecture of COVID-19. Nature. 2021;596:472–473. doi: 10.1038/s41586-021-03767-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Mozzi A., Pontremoli C., Sironi M. Genetic susceptibility to infectious diseases: Current status and future perspectives from genome-wide approaches. Infect. Genet. Evol. 2018;66:286–307. doi: 10.1016/j.meegid.2017.09.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Mitchell B.D., Fornage M., McArdle P.F., Cheng Y.-C., Pulit S.L., Wong Q., Dave T., Williams S.R., Corriveau R., Gwinn K., et al. Using previously genotyped controls in genome-wide association studies (GWAS): application to the Stroke Genetics Network (SiGN) Front. Genet. 2014;5:1–7. doi: 10.3389/fgene.2014.00095. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wojcik G.L., Murphy J., Edelson J.L., Gignoux C.R., Ioannidis A.G., Manning A., Rivas M.A., Buyske S., Hendricks A.E. Opportunities and challenges for the use of common controls in sequencing studies. Nat. Rev. Genet. 2022;23:665–679. doi: 10.1038/s41576-022-00487-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Pairo-Castineira E., Clohisey S., Klaric L., Bretherick A.D., Rawlik K., Pasko D., Walker S., Parkinson N., Fourman M.H., Russell C.D., et al. Genetic mechanisms of critical illness in COVID-19. Nature. 2021;591:92–98. doi: 10.1038/s41586-020-03065-y. [DOI] [PubMed] [Google Scholar]
  • 10.Griffith G.J., Morris T.T., Tudball M.J., Herbert A., Mancano G., Pike L., Sharp G.C., Sterne J., Palmer T.M., Davey Smith G., et al. Collider bias undermines our understanding of COVID-19 disease risk and severity. Nat. Commun. 2020;11:5749. doi: 10.1038/s41467-020-19478-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Munafò M.R., Tilling K., Taylor A.E., Evans D.M., Davey Smith G. Collider scope: when selection bias can substantially influence observed associations. Int. J. Epidemiol. 2018;47:226–235. doi: 10.1093/ije/dyx206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Rosseel Y. lavaan : An R Package for Structural Equation Modeling. J. Stat. Soft. 2012;48 [Google Scholar]
  • 13.R Core Development Team . 2020. R: A Language and Environment for Statistical Computing. [Google Scholar]
  • 14.Yengo L., Sidorenko J., Kemper K.E., Zheng Z., Wood A.R., Weedon M.N., Frayling T.M., Hirschhorn J., Yang J., Visscher P.M., GIANT Consortium Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry. Hum. Mol. Genet. 2018;27:3641–3649. doi: 10.1093/hmg/ddy271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Fesinmeyer M.D., North K.E., Ritchie M.D., Lim U., Franceschini N., Wilkens L.R., Gross M.D., Bůžková P., Glenn K., Quibrera P.M., et al. Genetic risk factors for BMI and obesity in an ethnically diverse population: results from the population architecture using genomics and epidemiology (PAGE) study. Obesity. 2012;21:835–846. doi: 10.1002/oby.20268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Pulit S.L., Stoneman C., Morris A.P., Wood A.R., Glastonbury C.A., Tyrrell J., Yengo L., Ferreira T., Marouli E., Ji Y., et al. Meta-analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry. Hum. Mol. Genet. 2019;28:166–174. doi: 10.1093/hmg/ddy327. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Shrine N., Guyatt A.L., Erzurumluoglu A.M., Jackson V.E., Hobbs B.D., Melbourne C.A., Batini C., Fawcett K.A., Song K., Sakornsakolpat P., et al. New genetic signals for lung function highlight pathways and chronic obstructive pulmonary disease associations across multiple ancestries. Nat. Genet. 2019;51:481–493. doi: 10.1038/s41588-018-0321-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Astle W.J., Elding H., Jiang T., Allen D., Ruklisa D., Mann A.L., Mead D., Bouman H., Riveros-Mckay F., Kostadima M.A., et al. The allelic landscape of human blood cell trait variation and links to common complex disease. Cell. 2016;167:1415–1429.e19. doi: 10.1016/j.cell.2016.10.042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tsagris M., Papadakis M. Taking R to its limits: 70+ tips. PeerJ. 2018;6:1–15. [Google Scholar]
  • 20.Clogg C.C., Petkova E., Haritou A. Statistical methods for comparing regression coefficients between models. Am. J. Sociol. 1995;100:1261–1293. [Google Scholar]
  • 21.Panagiotou O.A., Ioannidis J.P.A., Genome-Wide Significance Project What should the genome-wide significance threshold be? Empirical replication of borderline genetic associations. Int. J. Epidemiol. 2012;41:273–286. doi: 10.1093/ije/dyr178. [DOI] [PubMed] [Google Scholar]
  • 22.Balduzzi S., Rücker G., Schwarzer G. How to perform a meta-analysis with R: a practical tutorial. Evid. Based. Ment. Health. 2019;22:153–160. doi: 10.1136/ebmental-2019-300117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Mitchell R., Hemani G., Dudding T., Corbin L., Harrison S., Paternoster L. 2019. UK Biobank Genetic Data: MRC-IEU Quality Control, Version 2. [DOI] [Google Scholar]
  • 24.Vergara C., Thio C.L., Johnson E., Kral A.H., O’Brien T.R., Goedert J.J., Mangia A., Piazzolla V., Mehta S.H., Kirk G.D., et al. Multi-ancestry genome-wide association study of spontaneous clearance of hepatitis C virus. Gastroenterology. 2019;156:1496–1507.e7. doi: 10.1053/j.gastro.2018.12.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Wojcik G.L., Thio C.L., Kao W.H.L., Latanich R., Goedert J.J., Mehta S.H., Kirk G.D., Peters M.G., Cox A.L., Kim A.Y., et al. Admixture analysis of spontaneous hepatitis C virus clearance in individuals of African descent. Genes Immun. 2014;15:241–246. doi: 10.1038/gene.2014.11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Duggal P., Thio C.L., Wojcik G.L., Goedert J.J., Mangia A., Latanich R., Kim A.Y., Lauer G.M., Chung R.T., Peters M.G., et al. Genome-wide association study of spontaneous resolution of hepatitis C virus infection: data from multiple cohorts. Ann. Intern. Med. 2013;158:235–245. doi: 10.7326/0003-4819-158-4-201302190-00003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Raj A., Stephens M., Pritchard J.K. fastSTRUCTURE: variational inference of population structure in large SNP data sets. Genetics. 2014;197:573–589. doi: 10.1534/genetics.114.164350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K., Motyer A., Vukcevic D., Delaneau O., O’Connell J., et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562:203–209. doi: 10.1038/s41586-018-0579-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Price A.L., Weale M.E., Patterson N., Myers S.R., Need A.C., Shianna K.V., Ge D., Rotter J.I., Torres E., Taylor K.D., et al. Long-range LD can confound genome scans in admixed populations. Am. J. Hum. Genet. 2008;83:132–135. doi: 10.1016/j.ajhg.2008.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Galinsky K.J., Bhatia G., Loh P.-R., Georgiev S., Mukherjee S., Patterson N.J., Price A.L. Fast principal-component analysis reveals convergent evolution of ADH1B in Europe and East Asia. Am. J. Hum. Genet. 2016;98:456–472. doi: 10.1016/j.ajhg.2015.12.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Chang C.C., Chow C.C., Tellier L.C., Vattikuti S., Purcell S.M., Lee J.J. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience. 2015;4:7. doi: 10.1186/s13742-015-0047-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Zheng X., Levine D., Shen J., Gogarten S.M., Laurie C., Weir B.S. A high-performance computing toolset for relatedness and principal component analysis of SNP data. Bioinformatics. 2012;28:3326–3328. doi: 10.1093/bioinformatics/bts606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K., Motyer A., Vukcevic D., Delaneau O., O’Connell J., et al. Genome-wide genetic data on ∼500,000 UK Biobank participants. bioRxiv. 2017 doi: 10.1101/166298. Preprint at. [DOI] [Google Scholar]
  • 34.UK Biobank . UK Biobank; 2015. Genotyping and Quality Control of UK Biobank, a Large-Scale, Extensively Phenotyped Prospective Resource: Information for Researchers; pp. 1–27. Interim Data Release, 2015. [Google Scholar]
  • 35.Manichaikul A., Mychaleckyj J.C., Rich S.S., Daly K., Sale M., Chen W.-M. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26:2867–2873. doi: 10.1093/bioinformatics/btq559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.McCarthy S., Das S., Kretzschmar W., Delaneau O., Wood A.R., Teumer A., Kang H.M., Fuchsberger C., Danecek P., Sharp K., et al. A reference panel of 64, 976 haplotypes for genotype imputation. Nat. Genet. 2016;48:1279–1283. doi: 10.1038/ng.3643. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Loh P.-R., Danecek P., Palamara P.F., Fuchsberger C., A Reshef Y., K Finucane H., Schoenherr S., Forer L., McCarthy S., Abecasis G.R., et al. Reference-based phasing using the haplotype reference consortium panel. Nat. Genet. 2016;48:1443–1448. doi: 10.1038/ng.3679. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Das S., Forer L., Schönherr S., Sidore C., Locke A.E., Kwong A., Vrieze S.I., Chew E.Y., Levy S., McGue M., et al. Next-generation genotype imputation service and methods. Nat. Genet. 2016;48:1284–1287. doi: 10.1038/ng.3656. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Howie B., Fuchsberger C., Stephens M., Marchini J., Abecasis G.R. Fast and accurate genotype imputation in genome-wide association studies through pre-phasing. Nat. Genet. 2012;44:955–959. doi: 10.1038/ng.2354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.O’Connell J., Sharp K., Shrine N., Wain L., Hall I., Tobin M., Zagury J.-F., Delaneau O., Marchini J. Haplotype estimation for biobank-scale data sets. Nat. Genet. 2016;48:817–820. doi: 10.1038/ng.3583. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Purcell S., Neale B., Todd-Brown K., Thomas L., Ferreira M.A.R., Bender D., Maller J., Sklar P., De Bakker P.I.W., Daly M.J., Sham P.C. PLINK: A tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 2007;81:559–575. doi: 10.1086/519795. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Butler R.W. Cambridge University Press; 2007. Saddlepoint Approximations with Applications. [Google Scholar]
  • 43.Mbatchou J., Barnard L., Backman J., Marcketta A., Kosmicki J.A., Ziyatdinov A., Benner C., O’Dushlaine C., Barber M., Boutkov B., et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet. 2021;53:1097–1103. doi: 10.1038/s41588-021-00870-7. [DOI] [PubMed] [Google Scholar]
  • 44.Galinsky K.J., Loh P.-R., Mallick S., Patterson N.J., Price A.L. Population structure of UK Biobank and ancient eurasians reveals adaptation at genes influencing blood pressure. Am. J. Hum. Genet. 2016;99:1130–1139. doi: 10.1016/j.ajhg.2016.09.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Spence J.P., Song Y.S. Inference and analysis of population-specific fine-scale recombination maps across 26 diverse human populations. Sci. Adv. 2019;5:eaaw9206. doi: 10.1126/sciadv.aaw9206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.D Turner S. qqman: an R package for visualizing GWAS results using Q-Q and manhattan plots. J. Open Source Softw. 2018;3:731. [Google Scholar]
  • 47.Willer C.J., Li Y., Abecasis G.R. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26:2190–2191. doi: 10.1093/bioinformatics/btq340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Cochran W.G. The combination of estimates from different experiments. Biometrics. 1954;10:101. [Google Scholar]
  • 49.Rauch A., Kutalik Z., Descombes P., Cai T., Di Iulio J., Mueller T., Bochud M., Battegay M., Bernasconi E., Borovicka J., et al. Genetic variation in IL28B is associated with chronic hepatitis C and treatment failure: a genome-wide association study. Gastroenterology. 2010;138:1338–1345. doi: 10.1053/j.gastro.2009.12.056. [DOI] [PubMed] [Google Scholar]
  • 50.Thomas D.L., Thio C.L., Martin M.P., Qi Y., Ge D., O’hUigin C., Kidd J., Kidd K., Khakoo S.I., Alexander G., et al. Genetic variation in IL28B and spontaneous clearance of hepatitis C virus. Nature. 2009;461:798–801. doi: 10.1038/nature08463. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Ge D., Fellay J., Thompson A.J., Simon J.S., Shianna K.V., Urban T.J., Heinzen E.L., Qiu P., Bertelsen A.H., Muir A.J., et al. Genetic variation in IL28B predicts hepatitis C treatment-induced viral clearance. Nature. 2009;461:399–401. doi: 10.1038/nature08309. [DOI] [PubMed] [Google Scholar]
  • 52.Cunningham F., Allen J.E., Allen J., Alvarez-Jarreta J., Amode M.R., Armean I.M., Austine-Orimoloye O., Azov A.G., Barnes I., Bennett R., et al. Ensembl 2022. Nucleic Acids Res. 2022;50:D988–D995. doi: 10.1093/nar/gkab1049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Karlsen T.H. Understanding COVID-19 through genome-wide association studies. Nat. Genet. 2022;54:368–369. doi: 10.1038/s41588-021-00985-x. [DOI] [PubMed] [Google Scholar]
  • 54.Tzellos S., Farrell P.J. Epstein-barr virus sequence variation-biology and disease. Pathogens. 2012;1:156–174. doi: 10.3390/pathogens1020156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Cannon M.J., Schmid D.S., Hyde T.B. Review of cytomegalovirus seroprevalence and demographic characteristics associated with infection. Rev. Med. Virol. 2010;20:202–213. doi: 10.1002/rmv.655. [DOI] [PubMed] [Google Scholar]
  • 56.Wolff D., Nee S., Hickey N.S., Marschollek M. Risk factors for Covid-19 severity and fatality: a structured literature review. Infection. 2021;49:15–28. doi: 10.1007/s15010-020-01509-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zhao J., Yang Y., Huang H., Li D., Gu D., Lu X., Zhang Z., Liu L., Liu T., Liu Y., et al. Relationship Between the ABO Blood Group and the Coronavirus Disease 2019 (COVID-19) Susceptibility. Clin. Infect. Dis. 2021;73:328–331. doi: 10.1093/cid/ciaa1150. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Shelton J.F., Shastri A.J., Ye C., Weldon C.H., Filshtein-Sonmez T., Coker D., Symons A., Esparza-Gordillo J., 23andMe COVID-19 Team. Aslibekyan S., Auton A. Trans-ancestry analysis reveals genetic and nongenetic associations with COVID-19 susceptibility and severity. Nat. Genet. 2021;53:801–808. doi: 10.1038/s41588-021-00854-7. [DOI] [PubMed] [Google Scholar]
  • 59.Rozenfeld Y., Beam J., Maier H., Haggerson W., Boudreau K., Carlson J., Medows R. A model of disparities: risk factors associated with COVID-19 infection. Int. J. Equity Health. 2020;19:126. doi: 10.1186/s12939-020-01242-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Kianersi S., Ludema C., Macy J.T., Chen C., Rosenberg M. Relationship between high-risk alcohol consumption and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) seroconversion: a prospective sero-epidemiological cohort study among American college students. Addiction. 2022;117:1908–1919. doi: 10.1111/add.15835. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Cooling L. Blood groups in infection and host susceptibility. Clin. Microbiol. Rev. 2015;28:801–870. doi: 10.1128/CMR.00109-14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Guillon P., Clément M., Sébille V., Rivain J.-G., Chou C.-F., Ruvoën-Clouet N., Le Pendu J. Inhibition of the interaction between the SARS-CoV Spike protein and its cellular receptor by anti-histo-blood group antibodies. Glycobiology. 2008;18:1085–1093. doi: 10.1093/glycob/cwn093. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Falagas M.E., Koletsi P.K., Baskouta E., Rafailidis P.I., Dimopoulos G., Karageorgopoulos D.E. Pandemic A(H1N1) 2009 influenza: review of the Southern Hemisphere experience. Epidemiol. Infect. 2011;139:27–40. doi: 10.1017/S0950268810002037. [DOI] [PubMed] [Google Scholar]
  • 64.Nie W., Zhang Y., Jee S.H., Jung K.J., Li B., Xiu Q. Obesity survival paradox in pneumonia: a meta-analysis. BMC Med. 2014;12:61. doi: 10.1186/1741-7015-12-61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Simou E., Britton J., Leonardi-Bee J. Alcohol consumption and risk of tuberculosis: a systematic review and meta-analysis. Int. J. Tuberc. Lung Dis. 2018;22:1277–1285. doi: 10.5588/ijtld.18.0092. [DOI] [PubMed] [Google Scholar]
  • 66.Rumbwere Dube B.N., Marshall T.P., Ryan R.P., Omonijo M. Predictors of human immunodeficiency virus (HIV) infection in primary care among adults living in developed countries: a systematic review. Syst. Rev. 2018;7:82. doi: 10.1186/s13643-018-0744-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Simou E., Britton J., Leonardi-Bee J. Alcohol and the risk of pneumonia: a systematic review and meta-analysis. BMJ Open. 2018;8:e022344. doi: 10.1136/bmjopen-2018-022344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Goedert J.J., Chen B.E., Preiss L., Aledort L.M., Rosenberg P.S. Reconstruction of the hepatitis C virus epidemic in the US hemophilia population, 1940-1990. Am. J. Epidemiol. 2007;165:1443–1453. doi: 10.1093/aje/kwm030. [DOI] [PubMed] [Google Scholar]
  • 69.Berntorp E., Fischer K., Hart D.P., Mancuso M.E., Stephensen D., Shapiro A.D., Blanchette V. Haemophilia. Nat. Rev. Dis. Primers. 2021;7:45. doi: 10.1038/s41572-021-00278-x. [DOI] [PubMed] [Google Scholar]
  • 70.Ruth K.S., Day F.R., Tyrrell J., Thompson D.J., Wood A.R., Mahajan A., Beaumont R.N., Wittemans L., Martin S., Busch A.S., et al. Using human genetics to understand the disease impacts of testosterone in men and women. Nat. Med. 2020;26:252–258. doi: 10.1038/s41591-020-0751-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Saevarsdottir S., Olafsdottir T.A., Ivarsdottir E.V., Halldorsson G.H., Gunnarsdottir K., Sigurdsson A., Johannesson A., Sigurdsson J.K., Juliusdottir T., Lund S.H., et al. FLT3 stop mutation increases FLT3 ligand level and risk of autoimmune thyroid disease. Nature. 2020;584:619–623. doi: 10.1038/s41586-020-2436-0. [DOI] [PubMed] [Google Scholar]
  • 72.Wu Y., Byrne E.M., Zheng Z., Kemper K.E., Yengo L., Mallett A.J., Yang J., Visscher P.M., Wray N.R. Genome-wide association study of medication-use and associated disease in the UK Biobank. Nat. Commun. 2019;10:1891. doi: 10.1038/s41467-019-09572-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Buniello A., MacArthur J.A.L., Cerezo M., Harris L.W., Hayhurst J., Malangone C., McMahon A., Morales J., Mountjoy E., Sollis E., et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 2019;47:D1005–D1012. doi: 10.1093/nar/gky1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Kichaev G., Bhatia G., Loh P.-R., Gazal S., Burch K., Freund M.K., Schoech A., Pasaniuc B., Price A.L. Leveraging polygenic functional enrichment to improve GWAS power. Am. J. Hum. Genet. 2019;104:65–75. doi: 10.1016/j.ajhg.2018.11.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Zhu Z., Guo Y., Shi H., Liu C.-L., Panganiban R.A., Chung W., O’Connor L.J., Himes B.E., Gazal S., Hasegawa K., et al. Shared genetic and experimental links between obesity-related traits and asthma subtypes in UK Biobank. J. Allergy Clin. Immunol. 2020;145:537–549. doi: 10.1016/j.jaci.2019.09.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Lee J.J., Wedow R., Okbay A., Kong E., Maghzian O., Zacher M., Nguyen-Viet T.A., Bowers P., Sidorenko J., Karlsson Linnér R., et al. Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Nat. Genet. 2018;50:1112–1121. doi: 10.1038/s41588-018-0147-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Fry A., Littlejohns T.J., Sudlow C., Doherty N., Adamska L., Sprosen T., Collins R., Allen N.E. Comparison of sociodemographic and health-related characteristics of UK Biobank participants with those of the general population. Am. J. Epidemiol. 2017;186:1026–1034. doi: 10.1093/aje/kwx246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Pirastu N., Cordioli M., Nandakumar P., Mignogna G., Abdellaoui A., Hollis B., Kanai M., Rajagopal V.M., Parolo P.D.B., Baya N., et al. Genetic analyses identify widespread sex-differential participation bias. Nat. Genet. 2021;53:663–671. doi: 10.1038/s41588-021-00846-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Alten S.V., Domingue B.W., Galama T., Marees A.T. Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering. medRxiv. 2022 doi: 10.1101/2022.05.16.22275048. Preprint at. [DOI] [Google Scholar]
  • 80.Tyrrell J., Zheng J., Beaumont R., Hinton K., Richardson T.G., Wood A.R., Davey Smith G., Frayling T.M., Tilling K. Genetic predictors of participation in optional components of UK Biobank. Nat. Commun. 2021;12:886. doi: 10.1038/s41467-021-21073-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Thibord F., Melissa V. Chan, Ming-Huei Chen, Andrew D.Johnson. A year of Covid-19 GWAS results from the GRASP portal reveals potential SARS-CoV-2 modifiers v2. medRxiv. 2021 doi: 10.1101/2021.06.08.21258507. Preprint at. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S17 and Tables S1–S5
mmc1.pdf (3.3MB, pdf)
Document S2. Article plus supplemental information
mmc2.pdf (4.6MB, pdf)

Data Availability Statement

Access to individual-level phenotypic and genetic data from HCV extended genetics consortium individuals can be requested via dbGaP (phs000248.v1.p1). Access to individual-level data from the UKB can be requested at https://www.ukbiobank.ac.uk. Summary statistics for each GWAS performed can be viewed and downloaded at https://my.locuszoom.org/ (Study name: HCV Consortium vs. UKB). Code used to perform simulation experiments, bipartite PCA-based matching, and gnomAD MAF-based variant-level filtering can be accessed at https://github.com/dduchen/Population_Based_Controls_GWAS_ID_Manuscript.


Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES