Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2026 Mar 28;27(2):bbag142. doi: 10.1093/bib/bbag142

Estimating population structure using epigenome-wide methylation data

Ziqing Wang 1,2, Kent D Taylor 3, Jerome I Rotter 4, Stephen S Rich 5, Yinan Zheng 6, Lifang Hou 7, Xiuqing Guo 8, Jan Bressler 9, Laura M Raffield 10, Yongmei Liu 11, Robert Kaplan 12,13, Donald M Lloyd-Jones 14, Alanna C Morrison 15, Myriam Fornage 16,17, Bruce M Psaty 18,19, Jennifer A Brody 20, Tamar Sofer 21,22,23,24,; the TOPMed Epigenetics working group
PMCID: PMC13032028  PMID: 41902502

Abstract

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (= 929), CARDIA (= 1123), JHS (= 1365), ARIC (= 2338), and HCHS/SOL (= 1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R² ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Keywords: population stratification, DNA methylation, epigenome-wide association study (EWAS)

Introduction

Epigenome-wide association studies (EWASs) and genome wide association studies (GWASs) are powerful approaches to identify epigenetic and genetic associations, respectively, with complex traits and diseases. In both types of analyses, it is essential to address population stratification, a form of bias caused by linkage disequilibrium and allele frequency divergence between subpopulations, which arise from populations’ evolutionary and demographic histories [1, 2]. Failure to account for population structure can lead to test statistic inflation and increased false positive findings, thus complicating the identification of true biological signals. The most common method to address population stratification is to include genetic principal components (GPCs), derived from principal components analysis (PCA) of genome-wide genetic data, as covariates in the association analysis. Consequently, the absence or incompleteness of genetic data often limits the sample size and thereby reduces the statistical power of EWAS association analyses.

DNA methylation (DNAm) is an epigenetic modification that can influence phenotype and disease outcomes by regulating gene expression through transcription and chromosomal inactivation [3, 4]. Notably, DNAm is shaped by both genetic factors and environmental exposures [5, 6], ranging from lifestyle and socioeconomic stressors [7] to infectious agents [8], and is characterized by its long-term stability, with potential transgenerational impact [9]. That said, previous studies have attempted to compute methylation-based PCs by PCA on genome-wide DNAm data or using CpG sites highly correlated with single nucleotide polymorphisms (SNPs) to capture population structure [8, 10, 11]. For instance, Rahmani and colleagues proposed to construct methylation PC (MPC) by applying PCA on a set of genetically informative CpGs in close vicinity to SNPs selected by linear model [10]. These findings demonstrate that genetic ancestry information is also embedded in one’s DNAm profiles, offering a promising opportunity for addressing population structure in epigenome-wide association studies when genetic data are unavailable.

One potential challenge with using PCs obtained by PCA of genome-wide methylation rather than genetic data, is that they may inadvertently capture technical or demographic variation, such as differences in cell type composition or age, since DNAm is influenced by a wide range of factors beyond genetic ancestry [11, 12]. In this study, we develop methylation-based measures of GPCs, which we term “Methylation population scores (MPSs)”. The MPSs are methylation scores that predict GPCs. We develop them using a feature selection approach informed by GPCs, while accounting for additional confounding variables, including age, sex, estimated cell type proportions, and environmental factors, including geographical locations and key lifestyles. We use methylation and genetic data from five multi-ethnic cohorts participating in the Trans-Omics in Precision Medicine (TOPMed) program, encompassing individuals self-reporting of White, Black, Chinese, and Hispanic backgrounds. We further compared our MPSs to MPC computed through PCA using reference list provided by Rahmani et al [10]. The significance of this study is two-fold, (i) develop MPSs based on available, continuous, genetic ancestry information (in the form of genetic PCs), meanwhile controlling for environmental and technical confounders through covariate adjustment; and (ii) incorporate a broader range of population groups and a larger sample size compared to previous studies, which primarily focused on individuals of White and Black backgrounds.

Methods

Figure 1 provides an overview of the study. We used methylation data (Illumina 850K/EPIC array in primary analysis, limiting to CpG sites available on the Illumina HumanMethylation450 BeadChip in secondary analysis, from genetically unrelated, individuals from five studies representing multiple race/ethnicity backgrounds from the TOPMed program: Multi-Ethnic Study of Atherosclerosis (MESA), Coronary Artery Risk Development in Young Adults (CARDIA), Jackson Heart Study (JHS), Atherosclerosis Risk in Communities (ARIC) study, and the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). Within each cohort, we randomly assigned 85% of participants to a training dataset and the remaining to a test dataset. We developed MPSs using CpG sites identified by feature selection methods trained on GPCs, and performed multiple analyses to evaluate their performance.

Figure 1.

Flowchart illustrating the study analysis plan across five cohorts: ARIC, CARDIA, JHS, MESA, and HCHS/SOL, showing the sequence of analytical steps from data input to final results.

Analysis plan of the study. ARIC, Atherosclerosis risk in communities study; CARDIA, Coronary artery risk development in young adults; JHS, Jackson heart study; MESA, Multi-ethnic study of atherosclerosis; HCHS/SOL, Hispanic community health study/study of Latinos.

The trans-omics for precision medicine program

The TOPMed program aims to improve diagnosis, treatment, and prevention of heart, lung, blood, and sleep disorders by elucidating genetic and phenotypic data over 130 000 participants from more than 80 participating studies [13], available through dbGaP (Database of Genotypes and Phenotypes) [14]. To minimize batch bias across cohorts, sequencing centers and experiment years, TOPMed standardized laboratory methods to a single pipeline and performed variant and genotype calling jointly on all samples in a given TOPMed freeze [13]. For this study, we obtained TOPMed whole genome sequencing (WGS) data from the Freeze 10 release.

Whole genome sequencing

DNA samples were extracted from whole blood samples collected from each participating cohort, before being processed for pair-ended 150 bp sequencing with a mean depth of 30x using Illumina HiSeq X Ten instruments at TOPMed sequencing centers [14]. As aforementioned, generated reads were first mapped to the human-genome build GRCh38 following a common pipeline across all sequencing centers. Variant discovery and genotyping were performed jointly across all available samples in freeze 10 using GotCloud pipeline [15]. Quality control included variant inclusion by minor allele frequency (MAF > 0.01), genotypes with a minimal depth of more than 10x, and sample-level checks for pedigree errors, self-reported and genetic sex discrepancies, and discordance with previous genotyping array data [14]. GPCs and genetic relatedness were also centrally computed using the PC-AiR and PC-Relate algorithms [16]. Among our study participants, a set of unrelated individuals was defined as a maximal set of individuals where twice the kinship coefficient was lower than 0.0625*2 for all pairs of individuals, limiting relatedness to third degree.

Epigenome-wide methylation quantification

Methylation profiling for ARIC, CARDIA, CHS, and HCHS/SOL was performed at the University of Washington Northwest Genomics Center, where DNA samples underwent normalization and bisulfite conversion using the EZ-96 DNA Methylation kit (#D5003; ZYMO Research). Methylation data for “NHLBI TOPMed: Multi-Ethnic Study of Atherosclerosis (MESA)” was performed at Keck Molecular Genomics Core Facility. DNAm profiles were generated using the MethylationEPIC BeadChip and quality checked using Illumina GenomeStudio software (v2.0.3) to include samples with a genotyping call rate of 0.98 or higher. Samples passing quality control were normalized by normal-exponential deconvolution using out-of-band probes (Noob) background subtraction [17] using the R package Minfi [18] to obtain beta-values. We estimated cell type subpopulations from the processed methylation data for each cohort using reference-based Houseman’s method [19].

HCHS/SOL DNAm samples were also profiled using the Infinium MethylationEPIC array at the HCHS/SOL Data Coordinating Center (DCC) at the University of North Carolina at Chapel Hill, and examined for sex and SNPs mismatches, control and blind duplicate, followed by normalization using the R package SeSAMe [20] to mask 105 454 probes, dye bias correction, and Noob normalization [17, 18]. The obtained beta-values were further corrected for type-2 probe bias using BMIQ method from the WateRmelon package [21], and underwent ComBat batch correction [22]. Subsequently, cell type proportions were also estimated using Houseman’s method [19].

Development of methylation population scores: methylation predictors of genetic principal components

We implemented a two-step feature selection method to identify the most representative CpG for each GPC in MPSs construction using training data. As shown in Fig. 1, the first step involved a preliminary selection of CpG sites for each GPC based on their associations. As meta-analysis is shown to be equivalent to aggregated data analysis [23], we estimated CpG site associations with GPCs by linear regression within each cohort’s training dataset in order to reduce computational burden. Each linear model is adjusted for age, sex, smoking and alcohol use status (if applicable), BMI, cell type proportions, study center, and race/ethnicity (if applicable; as identification with race/ethnicity group is often associated with lifestyle and environmental exposure patterns that may impact methylation). Specifically, smoking and alcohol use were modeled as categorical variable (0 = never, 1 = former or current). We then meta-analyzed these associations across cohorts by inverse variance, fixed effects meta-analysis. We applied False-Discovery Rate (FDR) correction on the resulting P-values using the Benjamini–Hochberg procedure [24], and selected CpG sites with an FDR-adjusted q-value<0.05. We then aggregated the training datasets across all cohorts and applied weighted-Lasso regression to further refine the selection CpG sites and compute weights.

Weighted-Lasso was implemented as follows. We first constructed an initial MP (MPSi) for each GPC as weighted sums of the CpG sites identified by a first Lasso regression, adjusting for the same set of covariates (except for alcohol use due to partial availability) as were used in the initial regression analysis. Next, we added a weighting step because the variance of GPCs varies by genetic ancestry due to patterns of allele frequency and linkage disequilibrium. Thus, each GPC was regressed on its corresponding initial MPSi without covariate adjustment to compute individual-specific weights, calculated as normalized squared residuals such that their sum equaled the sample size. According to the distribution of these constructed weights (Fig. S1), we truncated the extreme values above 90th percentile to the 90th percentile for stability, by limiting the contributions from extreme values to model results [25]. These individual weights were then implemented as observation weights in a second Lasso regression to improve the accuracy of MPSs by accounting for differences in prediction variation related to genetic diversity (where genetic structure impacts genetic variance [26]). In addition, we selected and provided CpG sites based on those present in the Infinium 450K array (MPS_450K_Lasso), after excluding Epic-specific CpG sites, using the same method.

Methylation population scores evaluation

We evaluated the performance of the MPSs, computed as weighted sums of CpG sites identified in the second Lasso regression, by computing their correlations with GPCs, and comparing them via visualization to (i) GPCs; (ii) MPCs constructed via PCA based on SNP adjacent CpG sites; and (iii) MPCs constructed as weighted sums based oof the CpGs obtained via EPISTRUCTURE by Rahmani and colleagues (MPC_Rah) [10] in the test dataset. Missing values in the methylation data were imputed using the mean across all samples before performing PCA. In particular, we performed visualization of population structure highlighting self-reported race/ethnicity groups, because race/ethnicity groups have shared patterns of genetic ancestry due to historical geographic migration patterns. We computed the variance explained by each MPSs for its corresponding GPC using linear regression adjusting for the same covariates aforementioned, and between−/within-group ratio (BW ratio), defined as the ratio of the race of between-group to within-group sum-of-squares matrices using Multivariate Analysis of Variance. The performance of MPSs were further evaluated in an independent cohort Cardiovascular Health Study (CHS) (Supplementary information). Moreover, we constructed an additional set of MPS in CHS (MPS_deconf), in which CpG sites were residualized for the same confounding variables, age, sex, smoking status, BMI, cell type proportions, study center, and race, prior to MPS construction.

For sensitivity analyses, we constructed three alternative MPSs: (i) MPS_450K, restricted to CpGs measured by the 450K array; (ii) MPS_EN, derived from CpGs selected by Elastic Net; and (iii) MPS_res, based on CpGs residualized for confounding variables. Specifically, the optimal regularization strength parameter (α) for Elastic Net model was selected across a set of values (0.3, 0.5, 0.7, 0.9, 1) using 5-fold cross validation.

Comparing alternative approaches to population stratification adjustment in epigenome-wide associations study of diabetes mellitus in Hispanic community health study/study of Latinos

We also conducted EWAS using MethParquet package [27] in HCHS/SOL participants with diabetes as the exposure, which was defined according to medical history and lab criteria defined by American Diabetes Association, as previously described [28]. These analyses compared models that adjusted for MPSs in the subset of individuals who have genetic data and therefore have GPCs (n = 1475), MPSs in a larger dataset of individuals who have methylation data (and therefore MPSs) but not necessarily GPCs (n = 2695), models adjusted for GPCs alone (n = 1475), and unadjusted models. In addition, we identified significant associations (adjusted P-value using Bonferroni correction less than 0.05) from the EWAS results in HCHS/SOL and compared them to previously identified diabetes associated CpG sites to evaluate biological relevance [29]. To further evaluate generalizability, we constructed MPSs in HCHS/SOL participants without genetic data and assessed their ability to differentiate individuals of different Hispanic/Latino backgrounds.

Results

Descriptive statistics of demographic characteristics for the TOPMed cohorts, stratified by self-reported race/ethnicity, are summarized in Table 1. Across all cohorts, women comprised a higher proportion of participants. The average age ranges from 40 years in CARDIA to 60 years in MESA.

Table 1.

Demographic characteristics of study cohorts.

Race/
ethnicity
N Age
(Mean (SD))
Sex
(N, %)—Female
BMI
(Mean (SD))
ARIC
 White 1461 52.34 (5.49) 869 (59.5%) 26.10 (4.19)
 Black 596 53.03 (5.62) 366 (61.4%) 29.80 (6.03)
CARDIA
 White 613 40.61 (3.40) 366 (59.7%) 27.13 (6.05)
 Black 510 39.68 (3.84) 355 (69.6%) 30.98 (7.70)
JHS
 Black 1365 55.89 (12.01) 853 (62.5%) 31.94 (7.36)
MESA
 White 396 60.91 (9.77) 199 (50.3%) 28.02 (4.81)
 Black 183 60.96 (9.57) 106 (57.9%) 30.64 (5.62)
 Chinese 69 60.97 (10.19) 32 (46.4%) 24.72 (3.05)
 Hispanic/Latino 281 58.65 (9.35) 154 (54.8%) 29.66 (5.03)
HCHS/SOL
 Central American 88 56.99 (7.87) 58 (65.9%) 30.57 (6.11)
 Cuban 359 57.39 (7.32) 197 (54.9%) 29.69 (5.31)
 Dominican 222 56.36 (7.87) 164 (73.9%) 29.92 (5.13)
 Mexican 298 55.88 (7.52) 211 (70.8%) 30.77 (5.92)
 Puerto Rican 412 57.75 (7.73) 261 (63.3%) 30.88 (6.25)
 South American 59 55.97 (6.42) 43 (72.9%) 29.54 (5.42)
 Other/more than 1 36 57.44 (8.45) 18 (50.0%) 30.74 (5.69)

MESA: Multi-Ethnic Study of Atherosclerosis; CARDIA: Coronary Artery Risk Development in Young Adults; JHS: Jackson Heart Study; ARIC: Atherosclerosis Risk in Communities study; HCHS/SOL: Hispanic Community Health Study/Study of Latinos

Methylation population scores development

We developed MPSs for the first 10 GPCs. The number of CpG sites selected for each GPC in step 1 (FDR <0.05) and by the weighted Lasso regression in step 2 are reported in Table S1, with GPC1 having the largest number of associated CpGs (n = 32 172), and GPC7 the fewest (n = 44). Detailed lists of CpGs selected for MPSs construction, along with weights are provided in the Supplementary Data 1. The top four PCs accounted for the majority of selected CpGs, whereas PC7 was associated with only 25 CpGs (Table S1).

Methylation population scores evaluation

We constructed the MPSs in the aggregated TOPMed validation dataset (n = 1090). Figure 2 visualizes the Pearson correlations between the 10 MPSs and GPCs. The MPSs were moderately to highly correlated with their corresponding GPCs. The strongest correlation was between the first MPSs and the first GPC (R2 = 0.98), with the lowest correlation between MPS7 and GPC7 (R2 = 0.28). Similar to the correlation pattern among GPCs, MPS2 through MPS4 showed relatively higher correlations compared to other components. Interestingly, MPS5 and MPS8 demonstrated moderate correlations with GPC2 through GPC4 as well as their corresponding MPS—a pattern not observed for GPC5 and GPC8.

Figure 2.

Correlation matrix displaying pairwise correlations between 10 methylation population scores and genetic principal components, indicating the degree of overlap between epigenetic and genetic ancestry measures.

Correlation analysis between the 10 MPSs and GPC.

Out of 4913 CpGs previously reported by Rahmani et al. [10] for construction of population stratification indicators, 1369 were available in our methylation data and used to construct the methylation principal components (MPC_Rah) in the aggregated test dataset across five TOPMed cohorts.

We visualized the top three principal components of our MPSs, MPC_Rah, MPC based on SNP adjacent CpG sites, and GPCs using scatterplots to assess their performance in differentiating the four self-reported race/ethnicity groups in our test dataset (Fig. 3), which show group-level separation in GPCs space due to correlation of these groupings with genetic ancestry at the population level. We anticipated that methylation derived from MPSs and MPCs would show similar patterns to GPCs. As shown in Fig. 3A and D, MPSs derived from feature selection in this study exhibited clear separation among the four race/ethnicity groups, closely mirroring the clustering observed with GPCs, though the within-group dispersion appeared greater for MPSs. Additionally, the separation by MPSs between Chinese and Hispanic/Latino groups was less pronounced (Fig. 3A). On the other hand, the top 3 MPC_Rahmani effectively distinguished Black, White, and Hispanic/Latino participants, but not Chinese (Fig. 3B). The clustering of the three race/ethnicity groups was more dispersed and less compact compared to MPSs (Fig. 3A and B). Such performance is reflected by the BW ratio, with the MPSs showing higher and slightly smaller BW ratio than MPS_Rahs and GPCs, respectively (Fig. 3). MPCs constructed using PCA on SNP adjacent CpG sites distinguishes Hispanic/Latino participants fairly well from other race/ethnicity groups, but Black and White participants are mixed together (Fig. 3C). Moreover, a group of participants from all four race/ethnicity groups were separated out by MPCs (Fig. 3C). Interestingly, MPCs showed the highest BW ratio among the four population structure measures, probably driven by large intergroup distances that nevertheless failed to capture coherent clustering of the race/ethnicity groups (Fig. 3C).

Figure 3.

Four-panel scatter plots comparing methylation and genetic population structure colored by race/ethnicity. Panel A shows study-derived methylation population scores; Panel B shows methylation principal components constructed based on CpG sites provided by Rahmani et al.; Panel C shows SNP-adjacent CpG-derived methylation principal components; Panel D shows genetic principal components. Between-to-within group variance ratios are reported for each panel.

Scatter plots of GPCs, MPSs, and MPC_Rahmani, colored by race/ethnicity groups present in the data. (A) MPSs constructed in the current study; (B) MPCs constructed via PCA using CpG sites from previously published paper by Rahmani et al. [10]; (C) MPCs constructed in test dataset through PCA on SNP adjacent CpGs; (D) GPCs. BW ratio: Between−/within-group variance ratio.

Figure S2 illustrates the separation of self-reported Hispanic backgrounds in HCHS/SOL participants by MPSs, combining those in test data and without genetic data (n = 689) (left), compared with the ones in the test dataset by GPCs (n = 221) (right). While the self-reported Hispanic/Latino groups appear less distinct from each other in the MPSs space (left), it broadly replicates the structure observed with GPCs. In both cases, individuals of Mexican, Central American, and South American backgrounds are more mixed together, while individuals of Dominican and Puerto Rican backgrounds were grouped closer to each other (Fig. S2). According to the parallel coordinate plot, these backgrounds diverge more clearly at GPC3 and GPC8, as well as at MPS3 and MPS8 (Fig. S3). However, Hispanic/Latino backgrounds exhibit minimal differentiation at MPS5 to MPS7 (Fig. S3B).

Evaluation in cardiovascular health study

We further constructed MPSs in an independent cohort CHS (n = 2974), which is predominantly composed of White and Black participants. Similar to the observed correlations in test data, MPS1 and GPC1 also showed the highest correlation (R2 = 0.97) in CHS (Fig. S4). However, the correlations for the second and third components decreased to 0.62 and 0.5, respectively (Fig. S4), likely due to reduced available CpG sites for MPS2 and MPS3 in CHS. Still, top MPSs could separate Black and White following similar pattern as GPCs, despite smaller BW ratio and larger within-group dispersion (Fig. 4). Interestingly, top components of MPS_deconf showed much smaller correlations with GPCs than the later ones, with an R2 of 0.19 for MPS_deconf1 compared with values exceeding 0.8 for the final two (Fig. S5).

Figure 4.

Four-panel scatter plot evaluating transferability of methylation population structure scores in the Cardiovascular Health Study, colored by race and ethnicity. Panel A shows study-derived methylation population scores; Panel B shows methylation principal components constructed based on CpG sites provided by Rahmani et al.; Panel C shows SNP-adjacent CpG-derived methylation principal components; Panel D shows genetic principal components. Between-to-within group variance ratios are reported for each panel.

Scatter plots of GPCs, MPSs, colored by race/ethnicity groups present in CHS. (A) MPSs constructed in the current study; (B) MPCs constructed via PCA using CpG sites from previously published paper by Rahmani et al. [10]; (C) MPCs constructed in test dataset through PCA on SNP adjacent CpGs; (D) GPCs. BW ratio: Between−/within-group variance ratio.

Sensitivity analyses

Table S2 list the variance explained by all the constructed population structure measures. Notably, the top three MPSs explained more than 80% of variance in their corresponding GPC, whereas the variance explained by MPS7 and MPS8 dropped below 10% (Table S2). The MPS_450K_Lasso follow a similar pattern, albeit with slightly less variance explained (Table S2). MPS_EN, obtained via another feature selection method Elastic Net, explained a level of variance highly comparable to that of the LASSO-based MPSs, while MPS_RES and MPS_450K showed lower proportions of variance explained for the GPCs (Table S2). Residualizing CpGs for confounding variables (MPS_res) demonstrated clearer separation for the Chinese group, and reduced ability to distinguish White and Black groups, as shown by the smaller distance between the individuals (Fig. S6).

Comparing population stratification adjustment in epigenome-wide association study of diabetes in Hispanic community health study/study of Latinos

When using participants with genetic data (n = 1475), i.e. using exactly the same sample size to compare population stratification adjustment approaches, all three models appear inflated in both QQ plots and Manhattan plots (Fig. 5A, Fig. S7). The genomic inflation factors were 1.57 for no adjustment, 1.51 and 1.50 for adjusting for 5 GPCs and 5 MPSs, respectively, indicating moderate inflation across all three analyses, and some attenuation of inflation when adjusted for GPCs or MPSs. The number of statistically significant sites differed. Adjusting for 5 GPCs or 5 MPSs resulted in 153 and 157 significant CpGs, respectively, while no adjustment for population structure yielded 211 CpGs associated with diabetes (Fig. S7). Figure 5B illustrates the results expanding the analysis to all available samples for no adjustment and adjusting for MPSs (n = 2695), with 2672 and 1749 significant CpGs, respectively. We also benchmarked the identified CpG sites against a previously validated set of 56 CpG sites associated with type 2 diabetes [29]. Of these sites, 15 have FDR P-value<.05 (computed over the 56 sites) in the analysis that did not adjust for population structure and MPS-adjusted EWAS, and 14 for GPC-adjusted EWAS (Table S3). In the analysis using all available samples, 26 and 24 were identified for no-adjustment and MPS-adjusted EWAS, respectively (Table S4).

Figure 5.

Two-panel quantile–quantile plots comparing epigenome-wide association study results under three covariate adjustment strategies: five genetic principal components, five methylation population scores, and no adjustment. Panel A uses a matched sample size of 1475 across all models; Panel B uses optimized sample sizes with 2695 participants for the methylation population scores and unadjusted models and 1475 for the genetic principal component model.

QQplot of results from EWAS adjusting for 5GPCs, 5MPSs and neither (none). (A) Using same sample size (n = 1475); (B) using optimized sample size (n = 2695 for MPSs adjusted model (5MPSs) and unadjusted model (none); n = 1475 for GPC adjusted model).

Discussion

We developed MPSs using DNAm data from multiple multi-ethnic cohorts via a two-step feature selection approach and evaluated their ability to capture population structure. The top three MPSs exhibited strong correlations with corresponding GPC and effectively separated self-reported race/ethnicity groups in the independent test dataset and CHS. On the other hand, MPCs derived from PCA using SNP adjacent CpG sites show inferior performance to MPSs in distinguishing across different race/ethnicity groups. The grouping of several individuals from distinct population backgrounds may be attributed to a certain degree of environmental effect captured by MPCs or the noise generated during the imputation step. The distinction of self-reported Hispanic backgrounds by MPSs among HCHS/SOL participants without genetic data further supports their consistency in capturing meaningful population structure. The closer distance between Mexican, Central American, and South American Hispanic/Latino groups may be caused by more substantial proportions of Amerindian genetic ancestry compared to other Hispanic/Latino groups [30]. In contrast, the genetic ancestry of Dominican and Puerto Rican individuals is less well captured by the available CpG sites in the current dataset. This limitation could result from the preprocessing/quality control step of the methylation data, where some CpG sites were removed/masked in order to reduce bias.

While GPCs remain the gold standard to account for population structure in EWAS, our results indicate that MPSs offer comparable performance to GPCs and are superior to unadjusted analyses, particularly when genetic data are unavailable. Adjusting for population structure reduces confounding arising from ancestry as causal factor, via ancestry-specific genetic variants and frequencies, of both methylation patterns and health outcomes. However, adjustment for measures of population structure does not account for other common causes of methylation and outcomes of interest, nor to other potential artifacts. The moderate inflation factor observed in our diabetes EWAS analysis suggests residual confounding that remains unaccounted for, likely attributable to unmeasured confounders, including environmental, socioeconomic, and lifestyle factors which are common causes of both methylation levels and diabetes, as well as technical artifacts [31, 32]. Residual inflation resulted from such biases could be further addressed using approaches that correct, or account for, the test statistics distribution [32]. This residual inflation underscores the importance of rigorous study design and analytical approaches to minimize confounding in EWAS [33, 34].

Similar to the previously proposed methods [8, 10, 11, 35], our approach also relies on the assumption that some DNAm sites are associated with genetics [6, 36]. In absence of genetic data, EpiAnceR+ [35] and EPISTRUCTURE [10] can be used to infer population structure by applying PCA over CpGs near SNPs. While these two methods work in different ways, in principle, they both remove variation, i.e. unrelated to population structure by regressing out covariate effects from methylation beta values, identifying methylation sites that are near SNPs (EpiAnceR+ further uses the methylation probes that specifically target SNPs), and applying PCA [35]. EPISTRUCTURE further selected CpGs that had association with SNPs, rather than were only near SNPs [10]. In contrast, the approach that we present here (i) filtered CpG sites based on association with precomputed GPCs, not using information about proximity to SNPs in the CpG filtering process; (ii) used penalized regression to generate scores “predicting” the reference GPCs, instead of generating scores using PCA. While it is biologically reasonable to model CpG sites as outcomes, we retained genetic PCs as outcomes for consistency (Fig. 1) due to single-outcome requirement for flexible feature selection method. Note that it is appropriate to use GPCs as outcomes and CpG as exposures because covariates effects are “regressed out” of the CpGs based on the Frisch–Waugh–Lovell theorem, just like covariates would be “regressed out” of the GPCs if they are used as exposures while CpGs are outcomes [37]. This method is also less computational demanding than traditional PCA when applied to large cohort data. Such efficiency was gained by first filtering out DNAm methylation sites not associated with genetic PCs, which can also be achieved by preselecting those associated with genetic variants. To further account for environmental confounders and tissue heterogeneity, factors previously observed to influence the selection of methylation sites in SNP-based methods [11], we incorporated these variables as covariates in our models.

In primary analysis we did not regress out covariates first, because the regression model associated with CpGs with GPCs accounts for the covariates in the regression. In a sensitivity analysis residualizing the beta values for cell proportions and confounding covariates before penalized regression, the resulting MPSs perform slightly less well than those from analysis that included the covariates in the penalized regression. We also assessed whether, once CpG and weights are already selected for the MPSs, it is useful to residualize CpGs for important determinants of methylation values (age, sex, cell proportions, and other covariates) prior to construction of the MPSs in a new cohort. As a result, residualizing CpGs before MPS construction led to diminished performance of the resulting MPS in CHS. However, this does not inform of PCA-based MPSs (such as in EpiAnceR+ and EPISTRUCTURE): as PCA cannot account for covariates, it is important to regress-out covariates first in PC-based analyses. A limitation of our work is that we did not compare to EpiAnceR+, because it requires raw red/green intensities data rather than beta values, and, while we computed MPS_Rahmani (EPISTRUCTURE), these are based on list of CpGs compiled for the 450K array, as the list has not been updated for EPIC array.

We also applied Elastic Net regression (MPS_EN). Again, the resulting MPSs are similar to those from the main analysis. While this manuscript focused on linear models for MPSs development, non-linear, machine learning (ML) models have increasingly and successfully implemented for development of omics-based prediction models of complex traits [38–40]. It is a topic of future work to develop non-linear ML models to capture population structure using methylation data, in the forms of MPSs or other non-linear latent representations [41–43].

Nevertheless, the greater dispersion patterns of points within each group in the scatterplots (Fig. 3) and lower correlations between non-top MPSs and GPCs imply that these MPSs do not entirely capture the genetic structure in the data. The decreased correlations between later MPSs and GPCs reflect the smaller proportions of population structure captured by these later components, which may render the development of these MPSs more susceptible to noise. The observed lower correlations suggest diminishing contribution of later GPCs to methylation variation—at least as captured by the available CpG sites. It is possible that heterogeneity in methylation data is attributable to cohort-specific differences in quality control and data processing limited inference and performance of MPSs. Another limitation of this study is the small number of Chinese participants (n = 69, Table 1), who primarily have East Asian genetic ancestry, whereas TOPMed individuals self-reporting other race/ethnicities typically have low levels of East Asian genetic ancestry (see Supplementary Fig. S3 available online at http://bib.oxfordjournals.org/ in Kurniansyah et al. [44]). This could partly explain their closer grouping towards Hispanics/Latinos participants, compared to when using GPCs (Fig. 3). This impacts the generalizability of our MPSs in distinguishing the Chinese population in other independent datasets.

Both MPC_Rahmani derived from ancestry-informative SNPs and MPSs trained on GPCs effectively capture the prominent population structure. One limitation for the comparison with MPC created based on Rahmani et al. [10] is the limited overlap between their selected CpG and those available in our integrated dataset (1469 out of 4913 CpGs). Moreover, their reference CpGs were selected using 450K DNAm data from individuals of European ancestry, which may restrict generalizability to other populations and to DNAm data generated using different arrays. Given the strong correlation and high reproducibility between Infinium 450K/EPIC array data and whole-genome bisulfite sequencing (WGBS) [45, 46], WGBS can also be used to construct MPSs based on the CpG sites identified in the current study. Furthermore, because of its comprehensive coverage of the human genome [47], WGBS is well suited for identifying CpG sites related to genetic PCs using our proposed feature-selection approach. Future research could therefore focus on enhancing the reproducibility and transferability of DNAm-based population structure prediction across diverse platforms and ancestries, e.g. by creating imputation models for missing CpG sites. To facilitate the use of our computed MPSs for future studies, we have made the lists of CpGs selected for each GPC, along with the corresponding R code, publicly available at Supplementary Data 1 and on Zenodo repository (DOI: 10.5281/zenodo.17074835).

Conclusion

Methylation-based scores predicting GPC, developed while integrating feature selection and adjustment for relevant confounders, provide a reliable estimate of population structure, showing strong concordance with GPCs, effective differentiation of racial and ethnic groups, as well as robust control of inflation in EWAS when genetic data are absent. With appropriate application, MPSs can complement GPCs to account for population structure in large cohorts when genetic data are unavailable for all or some individuals.

Key Points

  • Methylation population scores (MPSs) provide a reliable estimate of population structure in the data and can complement genetic principal components (GPCs) when genetic data are absent.

  • Unlike previous methods based on unsupervised principal component analysis, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations.

  • Like genetic principal components, MPSs also effectively differentiate racial and ethnic groups, and can reduce some of the inflation in epigenome-wide association analysis.

Supplementary Material

20260309_Supplementary_Tables_and_figures_bbag142
Supplementary_Data_1_bbag142

Acknowledgements

Molecular data for the Trans-Omics in Precision Medicine (TOPMed) program was supported by the National Heart, Lung and Blood Institute (NHLBI). See the TOPMed Omics Support Table (Supplementary Note 1) for study specific omics support information. Core support including centralized genomic read mapping and genotype calling, along with variant quality metrics and filtering were provided by the TOPMed Informatics Research Center (3R01HL-117626-02S1; contract HHSN268201800002I). Core support including phenotype harmonization, data management, sample-identity QC, and general program coordination were provided by the TOPMed Data Coordinating Center (R01HL-120393; U01HL-120393; contract HHSN268201800001I). We gratefully acknowledge the studies and participants who provided biological samples and data for TOPMed. Study specific acknowledgements will be provided in the supplementary information.

Contributor Information

Ziqing Wang, CardioVascular Institute, Beth Israel Deaconess Medical Center, 330 Brookline Ave, Boston, MA 02215, United States; Department of Medicine, Harvard Medical School, 25 Shattuck Street, Boston, MA 02215, United States.

Kent D Taylor, The Institute for Translational Genomics and Population Sciences, Department of Pediatrics, The Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center, 1124 W Carson Street, Torrance, CA 90502, United States.

Jerome I Rotter, The Institute for Translational Genomics and Population Sciences, Department of Pediatrics, The Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center, 1124 W Carson Street, Torrance, CA 90502, United States.

Stephen S Rich, Department of Public Health Genomics, University of Virginia School of Medicine, 1415 Jefferson Park Avenue, Charlottesville, VA 22903, United States.

Yinan Zheng, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, 420 East Superior Street, Chicago, IL 60611, United States.

Lifang Hou, Department of Preventive Medicine, Northwestern University Feinberg School of Medicine, 420 East Superior Street, Chicago, IL 60611, United States.

Xiuqing Guo, The Institute for Translational Genomics and Population Sciences, Department of Pediatrics, The Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center, 1124 W Carson Street, Torrance, CA 90502, United States.

Jan Bressler, Human Genetics Center, Department of Epidemiology, School of Public Health, The University of Texas Health Science Center at Houston, 1200 Pressler Street, Houston, TX 77030, United States.

Laura M Raffield, Department of Genetics, University of North Carolina at Chapel Hill, 250 E. Franklin Street, Chapel Hill, NC 27514, United States.

Yongmei Liu, Department of Medicine, Divisions of Cardiology and Neurology, Duke University Medical Center, 10 Duke Medicine, Durham, NC 27710, United States.

Robert Kaplan, Department of Epidemiology and Population Health, Albert Einstein College of Medicine, 1300 Morris Park Avenue, Bronx, NY 10461, United States; Division of Public Health Sciences, Fred Hutchinson Cancer Center, 1100 Fairview Ave N, Seattle, WA 98109, United States.

Donald M Lloyd-Jones, Department of Preventive Medicine, Boston University Chobanian & Avedisian School of Medicine, 72 E. Concord St., Boston, MA 02118, United States.

Alanna C Morrison, Human Genetics Center, Department of Epidemiology, School of Public Health, The University of Texas Health Science Center at Houston, 1200 Pressler Street, Houston, TX 77030, United States.

Myriam Fornage, Human Genetics Center, Department of Epidemiology, School of Public Health, The University of Texas Health Science Center at Houston, 1200 Pressler Street, Houston, TX 77030, United States; Brown Foundation Institute of Molecular Medicine, McGovern Medical School, University of Texas Health Science Center at Houston, 7000 Fannin St, Houston, TX 77030, United States.

Bruce M Psaty, Cardiovascular Health Research Unit, Department of Medicine, University of Washington School of Public Health, 3980 15th Ave NE, Seattle, WA 98195, United States; Department of Epidemiology, University of Washington, 1400 NE Campus Parkway Seattle, WA 98195, United States.

Jennifer A Brody, Cardiovascular Health Research Unit, Department of Medicine, University of Washington School of Public Health, 3980 15th Ave NE, Seattle, WA 98195, United States.

Tamar Sofer, CardioVascular Institute, Beth Israel Deaconess Medical Center, 330 Brookline Ave, Boston, MA 02215, United States; Department of Medicine, Harvard Medical School, 25 Shattuck Street, Boston, MA 02215, United States; Division of Sleep Medicine and Circadian Disorders, Department of Medicine, Brigham and Women’s Hospital, 75 Francis St, Boston, MA 02115, United States; Department of Biostatistics, Harvard T.H Chan School of Public Health, 677 Huntington Avenue, Boston, MA 02115, United States.

Conflict of interest

None declared.

Funding

This work was supported by the National Heart, Lung, and Blood Institute (grant number R01HL161012).

Data availability

TOPMed freeze 10 WGS, methylation, and phenotype data are available by application to dbGaP according to the study specific accession: ARIC: “phs001211”, CARDIA: “phs001612”, CHS: “phs001368”, JHS: “phs000964”, HCHS/SOL: ‘phs001395”. JHS methylation used in this manuscript are available via application to dbGaP, via accession “phs000286”. JHS methylation data can also be accessed through data use agreement to coordinating center (https://www.jacksonheartstudy.org/). HCHS/SOL methylation data used in this manuscript are available through application to the database of Genotypes and Phenotypes (dbGaP) accession “phs000810”, or via data use agreement with the HCHS/SOL Data Coordinating Center (DCC) at the University of North Carolina at Chapel Hill, see collaborators website: https://sites.cscc.unc.edu/hchs/. MPSs CpGs and weights will be provided at the Zenodo repository.

References

  • 1. Peterson  RE, Kuchenbaecker  K, Walters  RK  et al. Genome-wide association studies in ancestrally diverse populations: opportunities, methods, pitfalls, and recommendations. Cell  2019;179:589–603. 10.1016/j.cell.2019.08.051 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Jones  SC, Cardone  KM, Bradford  Y  et al. The impact of ancestry on genome-wide association studies. Pac Symp Biocomput  2025;30:251–67. 10.1142/9789819807024_0019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Jin  B, Li  Y, Robertson  KD. DNA methylation: superior or subordinate in the epigenetic hierarchy?  Genes Cancer  2011;2:607–17. 10.1177/1947601910393957 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Yong  W-S, Hsu  F-M, Chen  P-Y. Profiling genome-wide DNA methylation. Epigenetics Chromatin  2016;9:26. 10.1186/s13072-016-0075-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Jaenisch  R, Bird  A. Epigenetic regulation of gene expression: how the genome integrates intrinsic and environmental signals. Nat Genet  2003;33:245–54. 10.1038/ng1089 [DOI] [PubMed] [Google Scholar]
  • 6. Min  JL, Hemani  G, Hannon  E  et al. Genomic and phenotypic insights from an atlas of genetic effects on DNA methylation. Nat Genet  2021;53:1311–21. 10.1038/s41588-021-00923-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Champagne  FA. Epigenetic influence of social experiences across the lifespan. Dev Psychobiol  2010;52:299–311. 10.1002/dev.20436 [DOI] [PubMed] [Google Scholar]
  • 8. Husquin  LT, Rotival  M, Fagny  M  et al. Exploring the genetic basis of human population differences in DNA methylation and their causal impact on immune gene regulation. Genome Biol  2018;19:222. 10.1186/s13059-018-1601-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Byun  H-M, Nordio  F, Coull  BA  et al. Temporal stability of epigenetic markers: sequence characteristics and predictors of short-term DNA methylation variations. PLoS One  2012;7:e39220. 10.1371/journal.pone.0039220 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Rahmani  E, Shenhav  L, Schweiger  R  et al. Genome-wide methylation data mirror ancestry information. Epigenetics Chromatin  2017;10:1. 10.1186/s13072-016-0108-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Barfield  RT, Almli  LM, Kilaru  V  et al. Accounting for population stratification in DNA methylation studies. Genet Epidemiol  2014;38:231–41. 10.1002/gepi.21789 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Koestler  DC, Christensen  B, Karagas  MR  et al. Blood-based profiles of DNA methylation predict the underlying distribution of cell types: a validation analysis. Epigenetics  2013;8:816–26. 10.4161/epi.25430 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Taliun  D, Harris  DN, Kessler  MD  et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed program. Nature  2021;590:290–9. 10.1038/s41586-021-03205-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Mailman  MD, Feolo  M, Jin  Y  et al. The NCBI dbGaP database of genotypes and phenotypes. Nat Genet  2007;39:1181–6. 10.1038/ng1007-1181 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Jun  G, Wing  MK, Abecasis  GR  et al. An efficient and scalable analysis framework for variant extraction and refinement from population-scale DNA sequence data. Genome Res  2015;25:918–25. 10.1101/gr.176552.114 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Conomos  MP, Reiner  AP, Weir  BS  et al. Model-free estimation of recent genetic relatedness. Am J Hum Genet  2016;98:127–48. 10.1016/j.ajhg.2015.11.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Triche  TJ, Weisenberger  DJ, Van Den Berg  D  et al. Low-level processing of Illumina Infinium DNA methylation BeadArrays. Nucleic Acids Res  2013;41:e90–0. 10.1093/nar/gkt090 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Aryee  MJ, Jaffe  AE, Corrada-Bravo  H  et al. Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays. Bioinformatics  2014;30:1363–9. 10.1093/bioinformatics/btu049 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Houseman  EA, Accomando  WP, Koestler  DC  et al. DNA methylation arrays as surrogate measures of cell mixture distribution. BMC Bioinformatics  2012;13:86. 10.1186/1471-2105-13-86 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Ding  W, Kaur  D, Horvath  S  et al. Comparative epigenome analysis using Infinium DNA methylation BeadChips. Brief Bioinform  2023;24:bbac617. 10.1093/bib/bbac617 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Pidsley  R, CC  YW, Volta  M  et al. A data-driven approach to preprocessing Illumina 450K methylation array data. BMC Genomics  2013;14:293. 10.1186/1471-2164-14-293 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Leek  JT, Johnson  WE, Parker  HS  et al. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics  2012;28:882–3. 10.1093/bioinformatics/bts034 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Lin  DY, Zeng  D. Meta-analysis of genome-wide association studies: no efficiency gain in using individual participant data. Genet Epidemiol  2010;34:60–6. 10.1002/gepi.20435 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Benjamini  Y, Hochberg  Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Source. J R Stat Soc B Methodol  1995;57:289–300. 10.1111/j.2517-6161.1995.tb02031.x [DOI] [Google Scholar]
  • 25. Cole  SR, Hernán  MA. Constructing inverse probability weights for marginal structural models. Am J Epidemiol  2008;168:656–64. 10.1093/aje/kwn164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Sofer  T, Zheng  X, Laurie  CA  et al. Variant-specific inflation factors for assessing population stratification at the phenotypic variance level. Nat Commun  2021;12:3506. 10.1038/s41467-021-23655-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Wang  Z, Cassidy  M, Wallace  DA  et al. MethParquet: an R package for rapid and efficient DNA methylation association analysis adopting apache parquet. Bioinformatics  2024;40:btae410. 10.1093/bioinformatics/btae410 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Wang  Z, Wallace  DA, Spitzer  BW  et al. Methylation risk score of C-reactive protein associates sleep health with related health outcomes. Commun Biol  2025;8:821. 10.1038/s42003-025-08226-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Fraszczyk  E, Spijkerman  AMW, Zhang  Y  et al. Epigenome-wide association study of incident type 2 diabetes: a meta-analysis of five prospective European cohorts. Diabetologia  2022;65:763–76. 10.1007/s00125-022-05652-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Rao  H, Weiss  MC, Moon  JY  et al. Advancements in genetic research by the Hispanic community health study/study of Latinos: a 10-year retrospective review. HGG Adv  2025;6:100376. 10.1016/j.xhgg.2024.100376 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Guintivano  J, Shabalin  AA, Chan  RF  et al. Test-statistic inflation in methylome-wide association studies. Epigenetics  2020;15:1163–6. 10.1080/15592294.2020.1758382 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. van  Iterson  M, van  Zwet  EW, BIOS Consortium  et al. Controlling bias and inflation in epigenome- and transcriptome-wide association studies using the empirical null distribution. Genome Biol  2017;18:19. 10.1186/s13059-016-1131-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Campagna  MP, Xavier  A, Lechner-Scott  J  et al. Epigenome-wide association studies: current knowledge, strategies and recommendations. Clin Epigenetics  2021;13:214. 10.1186/s13148-021-01200-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Zheng  Y, Chen  Z, Pearson  T  et al. Design and methodology challenges of environment-wide association studies: a systematic review. Environ Res  2020;183:109275. 10.1016/j.envres.2020.109275 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Höffler  KD, Katrinli  S, Halvorsen  MW  et al. Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches. Epigenetics Chromatin  2025;18:69. 10.1186/s13072-025-00627-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Hannon  E, Gorrie-Stone  TJ, Smart  MC  et al. Leveraging DNA-methylation quantitative-trait loci to characterize the relationship between Methylomic variation, gene expression, and complex traits. Am J Hum Genet  2018;103:654–65. 10.1016/j.ajhg.2018.09.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Lovell  MC. A simple proof of the FWL theorem. J Econ Educ  2008;39:88–91. 10.3200/JECE.39.1.88-91 [DOI] [Google Scholar]
  • 38. Tanvir  RB, Islam  MM, Sobhan  M  et al. MOGAT: a multi-omics integration framework using graph attention networks for cancer subtype prediction. Int J Mol Sci  2024;25:2788. 10.3390/ijms25052788 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Benkirane  H, Pradat  Y, Michiels  S  et al. CustOmics: a versatile deep-learning based strategy for multi-omics integration. PLoS Comput Biol  2023;19:e1010921. 10.1371/journal.pcbi.1010921 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Kaczmarek  E, Jamzad  A, Imtiaz  T  et al. Multi-Omic graph transformers for cancer classification and interpretation. Pac Symp Biocomput  2022;27:373–84. [PubMed] [Google Scholar]
  • 41. Uzel  K, Grossen  C, Çilingir  FG. lcUMAPtSNE: use of non-linear dimensionality reduction techniques with genotype likelihoods. 2024. 10.1101/2024.04.01.587545 [DOI]
  • 42. Diaz-Papkovich  A, Anderson-Trocmé  L, Ben-Eghan  C  et al. UMAP reveals cryptic population structure and phenotype heterogeneity in large genomic cohorts. PLoS Genet  2019;15:e1008432. 10.1371/journal.pgen.1008432 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Zhang  M, Liu  Y, Zhou  H  et al. A novel nonlinear dimension reduction approach to infer population structure for low-coverage sequencing data. BMC Bioinformatics  2021;22:348. 10.1186/s12859-021-04265-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Kurniansyah  N, Goodman  MO, Khan  AT  et al. Evaluating the use of blood pressure polygenic risk scores across race/ethnic background groups. Nat Commun  2023;14:3202. 10.1038/s41467-023-38990-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. de  Abreu  AR, Ibrahim  J, Lemonidis  V  et al. Comparison of current methods for genome-wide DNA methylation profiling. Epigenetics Chromatin  2025;18:57. 10.1186/s13072-025-00616-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Pidsley  R, Zotenko  E, Peters  TJ  et al. Critical evaluation of the Illumina MethylationEPIC BeadChip microarray for whole-genome DNA methylation profiling. Genome Biol  2016;17:208. 10.1186/s13059-016-1066-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Lister  R, Pelizzola  M, Dowen  RH  et al. Human DNA methylomes at base resolution show widespread epigenomic differences. Nature  2009;462:315–22. 10.1038/nature08514 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

20260309_Supplementary_Tables_and_figures_bbag142
Supplementary_Data_1_bbag142

Data Availability Statement

TOPMed freeze 10 WGS, methylation, and phenotype data are available by application to dbGaP according to the study specific accession: ARIC: “phs001211”, CARDIA: “phs001612”, CHS: “phs001368”, JHS: “phs000964”, HCHS/SOL: ‘phs001395”. JHS methylation used in this manuscript are available via application to dbGaP, via accession “phs000286”. JHS methylation data can also be accessed through data use agreement to coordinating center (https://www.jacksonheartstudy.org/). HCHS/SOL methylation data used in this manuscript are available through application to the database of Genotypes and Phenotypes (dbGaP) accession “phs000810”, or via data use agreement with the HCHS/SOL Data Coordinating Center (DCC) at the University of North Carolina at Chapel Hill, see collaborators website: https://sites.cscc.unc.edu/hchs/. MPSs CpGs and weights will be provided at the Zenodo repository.


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES