Skip to main content
Briefings in Bioinformatics logoLink to Briefings in Bioinformatics
. 2025 Sep 12;26(5):bbaf460. doi: 10.1093/bib/bbaf460

Transfer learning reveals the mediating mechanisms of cross-ethnic lipid metabolic pathways in the association between APOE gene and Alzheimer’s disease

Lulu Pan 1,2, Yahang Liu 2,2, Chen Huang 3,2, Ruilang Lin 4, Yongfu Yu 5,6,✉, Guoyou Qin 7,8,✉
PMCID: PMC12445873  PMID: 40966649

Abstract

Lipid-mediated effects play a crucial role in elucidating the pathological mechanisms linking the ε4 allele of the apolipoprotein E gene (APOE ε4) to Alzheimer’s disease (AD). However, traditional mediation analysis methods often suffer from insufficient statistical power in studies involving minority populations due to limited sample sizes. This study innovatively develops a high-dimensional mediation analysis model (TransHDM) based on a transfer learning framework. By leveraging information from source data with large-scale samples, it significantly enhances the ability to identify potential mediators in small sample target data. The method first constructs a high-dimensional regression model using aggregated data from the source data and target data, then applies transfer regularization to adjust for heterogeneity between the source and target domains, correcting for estimation bias in high-dimensional Lasso. Ultimately, it achieves parameter transfer across domains, addressing statistical bias and inferential uncertainty caused by small sample sizes. Simulation results demonstrate that, compared to traditional methods, this approach significantly improves the power in identifying true mediator variables while effectively controlling the family-wise error rate in multiple testing. When applied to the Alzheimer’s Disease Neuroimaging Initiative cohort, TransHDM transferred large-scale data from white and other ethnic groups, identifying additional lipid metabolic pathways mediating the influence of the APOE ε4 allele on AD pathological progression in African American populations compared to pre-transfer analysis. These pathways include glycerophospholipid metabolism, glycerolipid metabolism, sphingolipid metabolism, and ether lipid metabolism (false discovery rate < 0.05). The TransHDM framework not only provides a powerful methodological tool for small sample population research but also offers valuable insights for future research in exploring disease mechanisms and developing biomarkers for disease prediction.

Keywords: transfer learning, mediation analysis, Alzheimer’s disease, lipidomics, African Americans, external dataset

Introduction

Alzheimer’s disease (AD), a complex neurodegenerative disorder with multifactorial etiologies, has become a significant public health challenge amid the global aging population [1]. Epidemiological studies reveal substantial ethnic disparities in AD incidence, with African Americans exhibiting a markedly higher prevalence compared to non-Hispanic Whites and Asian Americans [2]. This population not only bears a greater disease burden but also faces limited access to healthcare resources, highlighting the urgent need to enhance AD awareness and research within the African American community [3]. The racial disparities in AD are driven by a combination of genetic and environmental factors. Among these, the ε4 allele of the apolipoprotein E gene (APOE ε4) was a strong risk factor for AD [4]. Research indicates that APOE ε4 is associated with an increased prevalence of metabolic syndrome, with these metabolic disturbances often preceding the onset of AD pathological changes [5, 6]. These findings suggest that alterations in lipid metabolism may serve as a key molecular mechanism linking APOE ε4 to AD pathogenesis, potentially offering novel biomarkers and therapeutic targets [7]. However, the role of lipid metabolism in AD across different racial groups remains poorly understood. While associations between lipid metabolism and AD have been established in non-Hispanic Whites, their applicability to African Americans remains unclear [8]. Therefore, comprehensive lipidomic analyses in African Americans are essential to explore both shared and unique lipid metabolic features across racial groups, thereby validating the generalizability of findings from White populations and informing the development of universal biomarkers and therapeutic strategies.

Previous studies into lipid metabolism as a mediator of the APOE ε4–AD association have focused on specific lipid molecules or metabolic pathways, failing to capture the complexity of systemic lipid metabolic changes [9, 10]. Although lipidomics—an innovative technology enabling simultaneous analysis of diverse lipid molecules—offers a more comprehensive approach [8], its application in AD research has predominantly focused on non-Hispanic Whites, with African Americans significantly underrepresented [11]. This disparity is largely attributable to the small sample sizes of African American cohorts in major studies. For instance, the Alzheimer’s Disease Neuroimaging Initiative (ADNI), despite encompassing 59 sites across North America and collecting data from thousands of individuals, includes African American participants in ~10% of its total sample [12]. This underrepresentation, a common issue in large datasets and AD-related clinical trials [13], severely limits statistical power, complicating mediation analysis in high-dimensional data settings and hindering a comprehensive understanding of lipid metabolism’s role in AD pathogenesis [14]. Consequently, the development of effective prevention and intervention strategies for minority racial groups remains challenging.

In recent years, mediation analysis has incorporated dimensionality reduction techniques such as variable selection and regularization to address the high-dimensionality and correlations inherent in multi-omics data. For example, the HIMA method proposed by Zhang et al. uses the minimax concave penalty to first select mediators and then test for joint mediation effects, which may improve false positives [15]. Building on this, Perera et al. introduced the HIMA2 method, which utilizes a debiased LASSO technique for simultaneous mediator selection and testing, better controlling for false positives and improving statistical power [16]. However, these methods still face limitations in small sample scenario, often resulting in reduced statistical power. Transfer learning has emerged as a promising solution to address these limitations. This approach enhances model performance on target data by transferring information from external source data, demonstrating significant potential in practice [17–20]. Advances in data acquisition technologies have made it increasingly feasible to obtain related datasets from source domains, enabling researchers to utilize large-sample data (i.e. non-Hispanic White) to improve model performance for African Americans. Recent developments in high-dimensional penalized regression-based transfer learning frameworks have further advanced the field. For example, Li et al. proposed a high-dimensional two-step transfer learning framework based on linear regression [21]. Building on this, Tian et al. and Li et al. introduced algorithms for constructing confidence intervals for transfer learning estimates using a debiased LASSO approach, providing theoretical proof for reliable statistical inference [22, 23]. Despite these advancements, the application of transfer learning in mediation analysis remains underexplored, presenting a significant opportunity for innovation.

To address these gaps, this study integrates transfer learning methods within the HIMA2 framework and proposes a transfer learning–based high-dimensional mediation analysis (TransHDM) framework, aiming to overcome the challenges of sparse target data in mediator selection and mediation effect inference. Figure 1 provides a graphical overview of our research. We assess the performance of our method through extensive numerical evaluations and apply it to the ADNI study to investigate the mediating role of lipid metabolism in the APOE ε4–AD association. Our analysis focuses on lipid metabolic features in African Americans, aiming to explore both shared mechanisms and potential heterogeneity compared to non-Hispanic Whites. This study not only advances our understanding of cross-racial lipid metabolic characteristics in APOE ε4–associated AD but also introduces an innovative framework for high-dimensional mediation analysis in data-scarce settings.

Figure 1.

This figure illustrates the overall workflow of the study. Panel a highlights the challenge of low statistical power in high-dimensional mediation analysis with small sample sizes. Panel b presents the proposed TransHDM framework, which includes key steps such as assessing source–target similarity, building a LASSO model with combined data, and applying transfer regularization to address heterogeneity and bias. Panel c provides a schematic diagram of the simulation settings and analysis steps used in the ANDI cohort application.

Overview of this study. (a) Challenges in high-dimensional mediation analysis AA. Current high-dimensional mediation analysis methods lack sufficient statistical power when the target sample size is small. (b) Transfer learning–based high-dimensional mediation analysis (TransHDM) framework and solutions to technical difficulty. First, the method evaluates whether external data meets transfer learning criteria by calculating the loss function in cross-validation and comparing the similarity in variable associations between source and target data. Next, a high-dimensional LASSO model is built using data from both the source and target domains to obtain summary estimates. Transfer regularization then adjusts for heterogeneity between the source and target domains while correcting for biases in high-dimensional LASSO estimates. (c) A schematic diagram of the simulation process under different settings and the analyzing steps in the ANDI cohort application. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity.

Materials and methods

TransHDM framework

In the analytical workflow, we first identify transferable data sources, then perform transfer learning–based mediation analysis on the complete dataset. The TransHDM framework is an extension of the HIMA2 method proposed by Perera et al. [16], which enhances the estimation of mediation effects by incorporating covariate adjustment and integrating transfer learning algorithms to leverage information from external data sources. Since transfer learning with debiased lasso regression is the core algorithm throughout this process, we introduce this algorithm first in Section 2.1.1 to provide readers with the necessary technical foundation for better understanding and application of the subsequent methods.

Transfer learning with debiased lasso regression

Suppose there is a target dataset Inline graphic with sample size Inline graphic and Inline graphic independent source datasets Inline graphic, where Inline graphic is the index set of transferable sources, which is assumed to be known. The sample size of the Inline graphic-th source is Inline graphic. The regression model corresponding to the target dataset or the Inline graphic-th source datasets is

graphic file with name DmEquation1.gif (1)

where Inline graphic is the coefficient of the target dataset and Inline graphic is the coefficient of the Inline graphic-th source for Inline graphic.

Our goal is to transfer useful information from the sources to the target, improving the estimation accuracy of Inline graphic. Algorithm 1 presents a transfer learning framework for high-dimensional regression using debiased lasso, which involves the following two key steps:

Algorithm 1.

Transfer learning with debiased lasso regression.

Input: Target data Inline graphic, source data Inline graphic, and known transferable set Inline graphic.
   1. Initial lasso estimation:
    Fit the regression model in all dataset, compute
              Inline graphic
   2. Debiased estimate (two-step correction)
    2.1 Transfer learning bias correction:
                      Inline graphic
    where Inline graphic
    2.2 Lasso bias correction
                     Inline graphic
                         Inline graphic
Output:  Inline graphic and Inline graphic.

Step 1. Initial lasso estimation:

We first consider that index set Inline graphic for the transferable source is known. The regression model is first fitted using both target and sources to obtain the initial lasso estimate:

graphic file with name DmEquation2.gif (2)

where Inline graphic denotes the loss function based on the data, Inline graphic is the regularization parameter, and Inline graphic enforces sparsity in the solution.

Step 2. Debiased estimate (two-step correction).

Step 2.1. First debiasing step (transfer learning bias correction):

Notably, the initial estimate Inline graphic obtained from the combined source and target data may be influenced by discrepancies between the two domains. To mitigate the potential bias introduced by the source data, we introduce a correction vector Inline graphic is introduced, which is a sparse vector that adjusts the estimate to better fit the target distribution:

graphic file with name DmEquation3.gif (3)

Inline graphic is the regularization parameter, and Inline graphic denotes the ℓ1-norm, which encourages sparsity in the correction vector Inline graphic. Then, the debiased estimator is given by

graphic file with name DmEquation4.gif (4)

Step 2.2. Second debiasing step (lasso bias correction):

While the first step addresses the bias introduced by transfer learning, the estimator Inline graphic may still be biased due to the nature of lasso regression in high-dimensional settings [24]. To correct this, a second debiasing step is performed using nodewise regression, which estimates the inverse covariance matrix Inline graphic of the features to further refine the coefficients [25]. Motivated by Tian et al. [22], the final debiased estimator is given by

graphic file with name DmEquation5.gif (5)

For each feature Inline graphic, nodewise regression estimates the conditional dependencies among features by regressing Inline graphic on the remaining features Inline graphic. The regression coefficients are estimated using

graphic file with name DmEquation6.gif (6)

where

graphic file with name DmEquation7.gif (7)

and

graphic file with name DmEquation8.gif (8)

Then, we define the matrix Inline graphic and diagonal matrix Inline graphic as

graphic file with name DmEquation11.gif (9)

and

graphic file with name DmEquation12.gif (10)

where Inline graphic.

Finally, the inverse covariance matrix is computed as Inline graphic, and the variance estimator of Inline graphic is defined as

graphic file with name DmEquation13.gif (11)

where Inline graphic.

Detecting transferable datasets

While the initial model fitting utilizes both the target and source datasets, not all source datasets may contribute effectively to improving the performance of target model. Thus, it is necessary to evaluate the transferability of each source dataset to ensure that only relevant information is incorporated.

To determine which source datasets can contribute effectively to the target model, the target dataset Inline graphicis first randomly split into two disjoint equal-size groups: a training set Inline graphic and a validation set Inline graphic. Using the training set only and run the Algorithm 1 to estimate the coefficients Inline graphic. For each source dataset Inline graphic, we then combine it with the target training set and run the Algorithm 1 to estimate the coefficients Inline graphic. These coefficients are evaluated on the target validation set Inline graphic, and the transferability of each source dataset is measured using the statistic

graphic file with name DmEquation14.gif (12)

The set of transferable source datasets is denoted as

graphic file with name DmEquation15.gif (13)

where Inline graphic can be chosen by cross-validation. The threshold 0.01 in the above formula is chosen based on prior literature [26]. Results corresponding to thresholds of 0.02 and 0.03 are provided in Table S1, indicating that the estimates remain stable within a small range.

This ensures that only relevant source datasets are retained for subsequent analyses, improving the robustness of the transfer learning framework.

Transfer learning for high-dimensional mediation analysis

This work is the first to integrate transfer learning into the mediation analysis framework, allowing for more accurate mediation effect estimation and mediator selection when dealing with high-dimensional mediators. In the case of high-dimensional mediators, we denote the exposure as Inline graphic, baseline covariates to be adjusted for as Inline graphic, where Inline graphic is a Inline graphic-dimensional vector, including all exposure–outcome, mediator–outcome, and exposure–mediator confounders. The mediators are denoted as Inline graphic, where Inline graphic is a Inline graphic-dimensional matrix and Inline graphic exceeds the sample size Inline graphic. The outcome model for Inline graphicand mediator model for Inline graphic can be expressed as the following mediator models and outcome model to estimate and test mediation effects:

graphic file with name DmEquation16.gif (14)

where Inline graphic and Inline graphic are the intercept terms for the mediator and outcome models, respectively. Inline graphic represents the direct effect of Inline graphic on Inline graphic adjusting for all mediators, and the mediation effect of Inline graphic on Inline graphic through Inline graphic can be defined as Inline graphic. Specifically, Inline graphic represent the coefficients for the effect of exposure Inline graphic on each of the mediators Inline graphic, adjusted for baseline covariates Inline graphic, and Inline graphic represent the coefficients for the effect of all mediators Inline graphic on the outcome Inline graphic, adjusting for exposure Inline graphic and covariates Inline graphic. Inline graphic and Inline graphic are the error terms, assumed to be independent and identically distributed normal variables with a mean of zero and constant variance.

In real-world data, mediation effects often interact with each other. For example, in biological pathways, changes in the expression or activity of one protein may trigger alterations in the expression or function of multiple downstream proteins, forming a complex network of interactions. Therefore, it is more reasonable for outcome model to treat all potential mediators as a whole. This approach allows for capturing the correlations among mediators without the need to explicitly define the causal relationships or mechanisms between them. Such an overall modeling strategy avoids biases caused by the omission of critical causal pathways and is particularly suitable for high-dimensional data, as it can reflect the cooperative effects and modular functional characteristics of complex biological systems to a certain extent.

The detailed estimation and testing procedure for the proposed high-dimensional mediation testing framework is given as follows:

Step 1. Sure independence screening.

When the number of mediators Inline graphic is much larger than the sample size Inline graphic, we can first employ the sure independence screening (SIS) to reduce the dimensionality [27]. We integrate SIS with transfer learning to effectively identify relevant mediators while leveraging information from both the target and source datasets. Notably, the application of transfer learning assumes that the target and source datasets share the same outcome, exposure, and mediator variables—an assumption commonly adopted in existing transfer learning frameworks [21, 22]. To reduce the number of mediators, we first consider the following marginal models for each mediator Inline graphic, Inline graphic:

graphic file with name DmEquation17.gif (15)

Using Algorithm 1, we estimate the coefficients Inline graphic and Inline graphic for mediator Inline graphic in the marginal models with target data Inline graphic and source data Inline graphic. Next, we compute the product term Inline graphic to roughly quantify the mediation effect of Inline graphic on Inline graphic through Inline graphic. To balance dimensionality reduction with the retention of potentially relevant mediators, mediators are ranked by the magnitude of Inline graphic. Following a reference criterion commonly adopted in high-dimensional settings, threshold Inline graphic can be utilized to guide the inclusion of mediators into set Inline graphic for the following analyses [15, 16]. This step effectively excludes irrelevant variables, reducing the dimensionality of the mediator set, improving model stability and computational efficiency in high-dimensional scenarios, and reducing the likelihood of false discoveries in mediator selection. Without using SIS, the model may face issues such as convergence difficulties, unstable parameter estimates, and a significant increase in computational costs. It is worth noting that although the Inline graphic is fixed at 100 in our simulations for simplicity, it can be manually adjusted based on sample size or specific domain knowledge to assess the robustness and stability of the method [15, 16].

Step 2. Estimation and testing of mediation effects.

For the mediators in the selected set Inline graphic, first, we estimate the exposure–mediator effect Inline graphic using Algorithm 1 applied to the exposure–mediator model, for each mediator Inline graphic. The corresponding Inline graphic-value for the test of significance of Inline graphic is calculated as

graphic file with name DmEquation18.gif (16)

where Inline graphic is the cumulative distribution function of Inline graphic. Similarly, the mediator–outcome effect Inline graphic is estimated using Algorithm 1 applied to the mediator–outcome model, and the significance of Inline graphic is tested using

graphic file with name DmEquation19.gif (17)

The variance estimates Inline graphic and Inline graphic refer to the variance of the coefficients Inline graphic and Inline graphic, respectively, and are calculated using Equation (11) in Section 2.1.1.

Step 3. Joint significance testing.

To determine the overall significance of each mediator, we perform a joint significance test by combining the Inline graphic-value for Inline graphic and Inline graphic:

graphic file with name DmEquation20.gif (18)

The P-values Inline graphic are adjusted to Inline graphic using the “JS-mixture” procedure, which controls the family-wise error rate (FWER) across multiple hypothesis tests [28]. Mediators with Inline graphic are considered significant, and the set of significant mediators Inline graphic is denoted as

graphic file with name DmEquation21.gif (19)

The details of the algorithm are shown in Algorithm 2.

Algorithm 2.

Transfer learning–based estimation and testing framework for high-dimensional mediation analysis (TransHDM).

Input target data Inline graphic, source data Inline graphic
 1. Sure independence screening
           Inline graphic
 2. Estimate and test exposure–mediator effects and mediator–outcome effects
   Compute Inline graphic and Inline graphic via Algorithm 1 based on outcome model (1) and (2)
   Inline graphic, for Inline graphic
   Inline graphic, for Inline graphic
 3. Joint significance test
   Inline graphic, for Inline graphic
   Inline graphic, where Inline graphic
Output  Inline graphic and Inline graphic

Simulation data generation

We conduct extensive simulations to evaluate the performance of the proposed method, TransHDM. The HDM method can be viewed as a simplified version of TransHDM, as it does not incorporate the transfer learning strategy. Both methods are extensions of the HIMA2 framework, with further adjustment for confounding variables in the model. However, while HDM uses debiased Lasso to estimate the model coefficients for the target data, TransHDM further integrates transfer learning, leveraging external source data to improve the mediation effect estimation. By comparing TransHDM with HDM, we demonstrate the advantages of the transfer learning strategy in enhancing the performance of mediation effect estimation and mediator selection.

For the target with Inline graphic and Inline graphic sources with Inline graphic, we consider two covariate distribution settings: (1) homogeneous design: the covariates Inline graphic are independently generated from Inline graphic for all source and target samples, where Inline graphic; (2) heterogeneous design: the covariates Inline graphic in each source Inline graphic are generated from Inline graphic, where Inline graphic. Here, Inline graphic is a random matrix with each entry equal to 0.3 with probability 0.3 and 0 otherwise. The target domain follows the homogeneous design with covariance Inline graphic. This setting induces covariate shift across source domains. Since both the exposure Inline graphic and mediators Inline graphic are generated conditional on the covariates Inline graphic, the heterogeneity in Inline graphic further propagates to Inline graphic and Inline graphic, resulting in induced distributional shifts in these variables as well.

The exposure Inline graphic is generated as

graphic file with name DmEquation22.gif

where Inline graphic and Inline graphic. The mediator Inline graphic for Inline graphic is generated as

graphic file with name DmEquation23.gif

where Inline graphic and Inline graphic, and the error terms Inline graphic are drawn from the multivariate normal distribution with mean zero and covariance matrix Inline graphic, where Inline graphic is the first-order autoregressive correlation structure with Inline graphic The outcome is generated as

graphic file with name DmEquation24.gif

where Inline graphic, Inline graphic  Inline graphic, and Inline graphic.

For the target, we set Inline graphic for the first eight mediators, and Inline graphic otherwise. We set the coefficient Inline graphic for the first eight mediators and Inline graphic otherwise. Therefore, we have (i) Inline graphic forInline graphic; (ii) Inline graphic but Inline graphic forInline graphic; (iii) Inline graphic but Inline graphic for Inline graphic; and (iv) Inline graphic and Inline graphic for Inline graphic.

We randomly select Inline graphic as the number of transferable sources. The coefficient Inline graphic is defined as

graphic file with name DmEquation25.gif

Similarly, for Inline graphic, we define

graphic file with name DmEquation26.gif

where Inline graphic is a random subset of Inline graphic with Inline graphic for Inline graphic, reflecting the heterogeneity between target and sources.

In the above setting, we set Inline graphic. We define Inline graphic and Inline graphic. All the simulations are based on 200 replications with 6 different settings: Inline graphicand Inline graphic.

Study sample

Participants in this study were from the ADNI cohort, a longitudinal study designed to identify clinical, genetic, imaging, and biological markers of early AD progression. The initial phase, ADNI-1, was launched in 2003 and recruited participants aged 55–90 from 63 sites across the USA and Canada. Subsequent phases (ADNI-GO, ADNI-2, and ADNI-3) tracked existing participants and enrolled additional cohorts. Detailed descriptions of ADNI’s design are available elsewhere [29, 30]. Demographic information, APOE genotype, questionnaire data, lipid metabolism data, neuroimaging data, and cerebrospinal fluid (CSF) biomarker data were obtained from the ADNI data repository (adni.loni.usc.edu).

For this analysis, 1524 eligible participants with complete plasma lipid metabolism data, exposure and covariate information, and participation in ADNI-1, ADNI-GO, or ADNI-2 were included. The exposure of interest was the number of APOE-ε4 alleles carried by participants, and 781 serum lipid metabolites measured at baseline were considered as mediators (detailed information can be found in Section 2.4). Covariates, selected based on prior literature and database availability, included age, gender, and years of education. The target outcomes were AD severity, assessed using three dimensions: (1) the 13-item version of the Alzheimer’s Disease Assessment Scale—Cognitive Subscale (ADAS-Cog 13) score, a widely accepted cognitive measure; (2) CSF concentrations of t-tau and amyloid-beta (1–42) (ABeta42), which reflect AD-related pathological changes; and (3) white matter hyperintensity (WMH) volume, an imaging marker of structural brain changes associated with AD. All outcomes were measured after baseline to ensure appropriate temporal relationships between exposure (APOE-ε4), mediators (lipid metabolism), and outcomes (AD severity). The detailed measurement processes for outcomes are provided in Sections 2.5–2.7.

To retain more samples for subsequent analysis, participants with missing values for each outcome (ADAS-Cog 13, t-tau, ABeta42, or WMH) were excluded separately, resulting in four analytic datasets comprising 1524, 1130, 1133, and 915 participants, respectively. The sample selection and exclusion process are detailed in Supplementary Fig. S1. All participants provided written informed consent, and the study protocol was approved by the Institutional Review Board at each participating site.

Lipid metabolomic

Lipid analysis data utilized in this study were obtained directly from the ADNI database. These lipid measurements had previously been conducted by ADNI using plasma samples collected from a subset of participants. The samples were analyzed at the Metabolomics Laboratory of the Baker Heart and Diabetes Institute, where lipid profiling of 781 plasma lipids was performed according to methods described previously [31]. See Table S2 for specific classes and species. To account for batch effects and differences between the ADNI-1 longitudinal cohort and ADNI-2/GO cohorts, the median concentration of each analyte was adjusted. Additionally, we normalized all data beforehand to ensure comparability across lipid molecules on the same scale.

Neuroimaging analysis

The WMH data were obtained from the ADNI dataset. As previously described, all scans were preprocessed using a standardized pipeline [32]. Detailed information on the method for WMH detection has been published [33]. Specifically, WMH detection was performed using a Bayesian Markov Random Field approach, which leverages a vector of three image intensities (proton density [PD], T1, and T2) associated with image pixels.

CSF biomarker analysis

CSF biomarker data were obtained from the ADNI dataset. As previously described, AD-related biomarkers, including t-tau and ABeta42, in CSF samples from ADNI-1/GO/2 were analyzed using the validated and highly automated Roche Elecsys electrochemiluminescence immunoassay [34]. This method significantly improved both within-laboratory and between-laboratory precision and accuracy, while also enhancing lot-to-lot consistency of the immunoassay kits.

Cognitive assessment

The cognitive performance was assessed using the modified Alzheimer’s Disease Assessment Scale—Cognitive Subscale. This scale comprises 13 items evaluating memory, language, praxis, and orientation abilities. Missing scores for any single item resulted in exclusion of the total ADAS-Cog 13 score. The scale ranges from 0 to 85, with higher scores indicating more severe cognitive impairment.

Statistical analysis

The datasets corresponding to the four AD severity indicators (ADAS-Cog 13, t-tau, ABeta42, WMH) were divided into target sample for African Americans (n = 62, 37, 37, 40, respectively) and source sample including non-Hispanic Whites, Asian, and other race (n = 1415, 1054, 1062, 841, respectively). The transferability of source datasets was pre-identified using a transferability recognition algorithm. Specifically, datasets from Non-Hispanic White, Asian, and Other racial groups were evaluated based on metrics including loss of source, loss of validation, and threshold values, as detailed in Table S3. Based on this analysis, datasets meeting the transferability threshold criteria were selected as transferable source datasets for subsequent analyses. High-dimensional mediation analysis was first conducted on African American individuals using HDM without transfer learning. To leverage external information, TransHDM was then applied to incorporate transferable knowledge from other racial groups. For comparison, HDM was also performed separately on non-Hispanic White individuals. Although the source dataset included participants from multiple racial backgrounds, we restricted the source population to non-Hispanic White individuals due to the insufficient sample sizes of other groups, which limited their suitability for reliable estimation. Lipid molecules with a selection frequency exceeding 80% from 100 repeated samplings are chosen as potential mediator variables to enhance the robustness of selection. Multiple testing adjustments for mediator variables were conducted using the “JS-mixture” method to control the FWER, with statistical significance defined as Inline graphic <.05.

Results

Simulation study

To assess the effectiveness of TransHDM in identifying true mediators, we simulated multiple experiments with varying source data sample sizes (Inline graphic200, 400, 600), mediator dimensions (Inline graphic1000 and 2000), different degrees of mediator correlation (low: Inline graphic, moderate: Inline graphic, high: Inline graphic), and covariate covariance structures (homogeneous and heterogeneous designs). These settings simulate common correlation patterns of high-dimensional omics data in real world (for data generation details, see “Simulation data generation” in the Methods section). Transferable detecting algorithm demonstrated robust performance in accurately identifying transferable sources, achieving perfect identification across all tested scenarios (Table S4). Importantly, TransHDM consistently outperformed traditional high-dimensional mediation (HDM) methods in accuracy and precision in mediator selection as well as mediation effect estimation (Figs. 2–3).

Figure 2.

This line graph illustrates how the accuracy of mediation selection and mediation effect estimates change with the increasing number of transferable sources, under different mediator dimensions and mediator correlations. It shows that TransHDM achieves higher power and lower false discovery rates (FWER) compared to HDM, while effectively reducing bias and variability in mediation effect estimates as the number of transferable sources increases.

Results of simulation experiments under homogeneous design. (a) Performance comparison of TransHDM and HDM across 18 simulation settings, focusing on the accuracy (Power) of selecting true mediators and the false discovery rate (FWER) of selecting irrelevant variables. (b–d) Performance comparison of TransHDM and HDM across 18 simulation settings in estimating the root mean squared error, relative bias and standard deviation of direct effects, mediated effects, and mediation proportions. Inline graphic was the correlation coefficient among mediators, Inline graphic was the dimension of mediators and Inline graphic was the number of transferable sources.

Figure 3.

This line graph illustrates how the accuracy of mediation selection and mediation effect estimates change with the increasing number of transferable sources, under different mediator dimensions and mediator correlations. The result shows that, under a heterogeneous design, TransHDM still outperforms HDM. Although the FWER is slightly higher compared to the homogeneous design, as the number of transferable sources increases, TransHDM effectively controls the FWER and maintains robust mediation effect estimates.

Results of simulation experiments under heterogeneous design. (a) Performance comparison of TransHDM and HDM across 18 simulation settings, focusing on the accuracy (Power) of selecting true mediators and the false discovery rate (FWER) of selecting irrelevant variables. (b–d) Performance comparison of TransHDM and HDM across 18 simulation settings in estimating the root mean squared error, relative bias, and SD of direct effects, mediated effects, and mediation proportions. Inline graphic was the correlation coefficient among mediators, Inline graphic was the dimension of mediators, and Inline graphic was the number of transferable sources.

Figure 2(a) and Fig. 3(a) report the FWER and statistical power of mediator selection under homogeneous and heterogeneous designs, respectively. As the number of sources (Inline graphic) increased, power improved, exceeding 0.95. Even with high correlation (Inline graphic), the method still achieved high power with sufficiently transferable datasets. FWER decreased significantly with increasing Inline graphic, demonstrating the ability of TransHDM to control false positive. Although FWER slightly increased at Inline graphic, it remained at a low level (FWER < 0.1), confirming robust error rate control. Figure 2(b–d) and Fig. 3(b–d) present the root mean squared error (rMSE), relative bias (rBias), and standard deviation (SD) of average direct effects (DE), indirect effects (IDE), and mediator proportions (MP) under different settings. Compared to HDM, TransHDM significantly reduced MSE, rBias, and SD for DE, IDE, and MP. Errors further decreased and approached zero with increasing Inline graphic, demonstrating TransHDM’s robustness across varying mediator dimensions (Inline graphic) and correlations (Inline graphic).

In further simulation studies, we introduced a heterogeneous design of the covariate variance–covariance matrix to mimic cross-domain covariate distribution shifts commonly encountered in practical analyses (Fig. 3). The results showed that under heterogeneous design, mediation analysis exhibited higher FWER compared to the homogeneous design, suggesting that differences in covariate distributions pose challenges to model accuracy. Nevertheless, as previously noted, with increasing Inline graphic, TransHDM demonstrated strong robustness across multiple evaluation indicators for effect estimation and mediator identification. Specifically, the FWER was effectively controlled. These findings highlight that TransHDM retains strong adaptability and robustness despite the challenges introduced by covariate heterogeneity.

In addition, we conducted simulations to evaluate the impact of retaining different numbers of mediators in the SIS step (Inline graphic 250, 500, 750, and 1000, whereInline graphic1000 represents no SIS) on TransHDM and HDM under the settings of Inline graphic and Inline graphic (Table S5). By comparing the results across different Inline graphic thresholds, we observed that as Inline graphic increases, power gradually improves, reflecting the inclusion of more potential mediators, thereby increasing detection efficiency. However, this also leads to a significant increase in FWER, possibly because more noise is retained, which elevates false positives in multiple testing. Nevertheless, across all s thresholds, transfer learning consistently enhances performance on the target data, and transfer from sufficient sources can effectively reduce FWER to acceptable levels. In terms of computational efficiency, the simulation time without SIS increases substantially. For Inline graphic and Inline graphic, the runtime of the algorithm without SIS is 10 to 15 times longer than that with SIS (Table S6).

Application for Alzheimer’s disease

Characteristics of participants

First, we conducted a comprehensive descriptive analysis of the similarities in baseline demographic and lipidomic characteristics between the target population (African American) and the source population (non-Hispanic White) to assess the rationale and feasibility of applying transfer learning in the context of this study.

Supplementary Table S7 summarizes baseline characteristics and AD severity–related outcomes across APOE ε4 carrier groups. The African American population (median age: 73.2 years; 37.7% male) and non-Hispanic White population (median age: 74.0 years; 56.2% male) both exhibited distinct AD pathological features in APOE ε4 carriers compared to non-carriers, with higher ADAS-Cog 13 scores, elevated t-tau levels, larger WMH volumes, and lower ABeta42 levels. Notably, non-APOE ε4 African Americans displayed milder AD-related pathology than their non-Hispanic White counterparts, with lower ADAS-Cog 13 scores, reduced t-tau levels, and higher ABeta42 levels. Correlation analysis revealed moderate associations of cognitive function (ADAS-Cog 13) with AD biomarkers (t-tau and ABeta42) (Pearson correlation: −0.40 to 0.44), while WMH showed weaker correlations with the other three indicators (Pearson correlation: −0.13 to 0.03) (Fig. S2).

In addition, lipid correlation patterns were highly consistent between African Americans and non-Hispanic Whites (Fig. S3). Significant positive correlations were observed within and between lipid classes, such as lysophosphatidylcholine (LPC) and lysophosphatidylethanolamine (LPE). Principal component analysis revealed substantial overlap in lipidomic profiles between African American and non-Hispanic White, indicating high similarity (Fig. S4).

Therefore, these findings provide direct evidence that the target and source populations are highly similar in terms of baseline demographic and lipidomic features, supporting both the plausibility of shared mediating factors between the two populations and the rationale for applying transfer learning from the source population to target population.

Identification of significant mediators

Figure 4(a) illustrates the number of lipid classes identified as mediators of the APOE ε4–AD association in both African American and non-Hispanic White. Transfer learning enhanced mediator identification in African Americans by incorporating useful information from external races, revealing seven additional lipid classes consistently identified across all outcomes: dihydroceramide (dhCer), alkenyl-phosphatidylcholine (PC (P)), phosphatidylethanolamine (PE), lysophosphatidylethanolamine (LPE (P)), diacylglycerol (DG), acylcarnitine (AC), and hydroxylated acylcarnitine (AC-OH). Results of lipid identification remain generally consistent between analyses of different outcomes. In non-Hispanic Whites, ~40 lipid classes were identified per outcome (ADAS-Cog 13: 38; t-tau: 41; ABeta42: 39; WMH: 38), with 32 lipids consistently significant across all outcomes. This consistency validates the robustness of lipid mediation effects across diverse AD severity measures.

Figure 4.

The figure shows the number of significant lipid mediators identified using HDM and TransHDM in African Americans, as well as the number identified using HDM in White individuals across different outcomes. It demonstrates that transfer learning in African Americans revealed seven additional lipid classes that were consistently significant across all outcomes, with results generally consistent across different outcomes.

Lipid mediator identification results. (a) The upset plot shows the number of lipid categories significantly mediating the association between APOE ε4 and four AD severity measures in both AA and NHW before and after transfer learning. The list on the right shows the lipid mediators that were significant in all four outcome analyses in AA, AA-T, and NHW. (b) Distribution profiles of significant lipid molecules. All lipid molecules are ordered sequentially, and the figure highlights those exhibiting significant mediation effects. All lipid molecules are arranged according to the order in Supplementary Tables S8–S15, and different colors indicate distinct lipid categories. Molecules with a significant mediation effect are filled with their category’s color, whereas nonsignificant molecules are shown in gray. The number of colored blocks within each category corresponds to the number of significant mediators in that category. Statistical significance was determined by the Joint Significance Test (Step 3 of Algorithm 2, TransHDM). The SIS thresholds were set at Inline graphic = 100 for HDM on AA population, Inline graphic = 150 for TransHDM on AA population, and Inline graphic = 400 for HDM on NHW population. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity.

We performed correlation analysis on identified lipids in African American and non-Hispanic Whites. Strong inter-lipid relationships, observed in African American result after transfer learning, such as PE with phosphatidylcholine (PC), PE with phosphatidylinositol (PI), PE with DG, and DG with triacylglycerol (TG), were also found in non-Hispanic Whites, reflecting the close biological or metabolic interrelationships between these lipids (Figs. S5–S6). These high correlations support the biological characteristic of lipids operating in a modular fashion within metabolic networks. When analyzing the African American population alone, the smaller sample size may have limited the identification of all lipids with synergistic functions. Therefore, the application of transfer learning could help identify lipid groups that are functionally related, allowing for the inclusion of lipid mediators that focus on key functional modules.

Figure 4(b) presents the lipid profiles significantly mediating the relationship between APOE ε4 and the four AD severity indicators. Analysis of lipid profiles corresponding to each evaluation metric revealed that non-Hispanic Whites exhibited a greater number of significant lipid molecules within each lipid class compared to African Americans. In the latter group, target population analysis identified fewer lipid categories and fewer significant lipid molecules per class. However, the application of transfer learning not only enabled the identification of additional lipid classes but also found more significant lipid molecules within previously identified classes, such as Cer(d), PC, and TG.

Comparative analysis of lipid profiles across the four outcome measures demonstrated consistency in both the overall lipid classes and the proportion of significant lipid molecules within each class. Notably, certain lipids exhibited population-specific mediation effects. For example, alkyl-phosphatidylethanolamine (PE (O)) was uniquely identified in non-Hispanic Whites, whereas sphingosine-1-phosphate (S1P) was specific to African Americans. Additionally, the significance of some lipids differed in various outcomes, potentially reflecting the distinct pathological aspects of AD captured by different evaluation metrics. These variations may also be influenced by factors such as sample size limitations and the inherent complexity of lipid–disease interactions. The estimated results of the mediation effects of specific lipids are shown in Tables S8–S15.

Lipid metabolomic pathway analysis

To explore the cross-ethnic mediation mechanisms of lipid metabolic pathways in the association between APOE4 ε4 and AD pathology, we performed pathway enrichment analysis on lipid classes demonstrating significant mediation effects in both African American and non-Hispanic White populations (Fig. 5, Table S16). In non-Hispanic Whites, lipid species were significantly enriched in four metabolic pathways: glycerolipid metabolism, glycerophospholipid metabolism, sphingolipid metabolism, and ether lipid metabolism. In African Americans, independent analysis revealed three distinct mediation patterns: (1) glycerophospholipid, sphingolipid, and ether lipid metabolism pathways significantly mediated the APOE–ADAS-Cog 13 association; (2) glycerophospholipid metabolism specifically mediated the association of APOE with t-tau and ABeta42 biomarkers; and (3) glycerophospholipid and sphingolipid metabolism mediated the APOE–WMH relationship (all false discovery rate (FDR) < 0.05). Notably, after integrating cross-ethnic data using transfer learning algorithms, we observed that glycerolipid, glycerophospholipid, and sphingolipid metabolism pathways concurrently mediated APOE associations with t-tau, ABeta42, and WMH in African Americans. Furthermore, all four metabolic pathways identified in non-Hispanic Whites consistently mediated APOE–ADAS-Cog 13 associations in African Americans.

Figure 5.

The Figure shows that after applying transfer learning, more significantly enriched lipid metabolic pathways were identified in African Americans, which were also found to be the same as those identified in the White population.

KEGG pathway enrichment analysis results. (a) The bar charts display the top 10 KEGG enrichment results for the lipids identified in AA, AA-T, and NHW across four outcomes. The bar charts marked with “***” indicate significant results with corrected FDR P-values. (b) The heatmap illustrates the significance of lipid species enriched in the four pathways. Blocks indicate lipids with significant mediating effects in their respective analysis groups, whereas empty cells indicate that the lipid category does not show a significant effect in that group. (c) The action network of lipids enriched in different pathways. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity; FDR, false discovery rate.

Figure 5(b) shows the identification of lipids enriched in the four aforementioned significant metabolic pathways across different outcomes and populations. In African American, certain lipids were not consistently identified across all outcome pathways. However, transfer learning significantly improved the identification of lipids with potential mediating effects. For example, DG, LPC, and PI in glycerophospholipid metabolism, Sphingomyelin (SM) in sphingolipid metabolism, DG and free fatty acid in glycerolipid metabolism, and lyso-alkyl-phosphatidylcholine (LPC(O)) in ether lipid metabolism were identified across all four outcomes, substantially improving the detection of metabolic pathways.

Although the significant metabolic pathways were largely similar between populations, subtle differences in specific enriched lipid categories were observed. For instance, sphingosine (Sph) in the sphingolipid metabolism pathway was uniquely identified in African Americans, while alkyl-phosphatidylethanolamine (PE(O)) in the ether lipid metabolism pathway was specific to non-Hispanic Whites, both before and after transfer learning. These results suggest that transfer learning, by incorporating external data and enhancing sample diversity, improves the recognition of complex and population-specific variations in lipid metabolism.

Sensitivity analysis

To further evaluate the robustness of the transfer learning approach in lipid identification, we conducted the following analyses. First, we constructed predictive models to assess whether the lipid mediators identified by TransHDM could improve the prediction of AD severity in the AA population. As shown in Table S17, the model incorporating lipid mediators identified by TransHDM (Model 2) achieved the best predictive performance. In contrast, Model 1, which included lipid mediators identified by HDM, showed inferior performance. Model 3, based on lipid mediators from the NHW population, showed similarly poor performance to the Model 4 with no mediators included, highlighting that TransHDM effectively captures population-specific characteristics of the target group. Overall, the lipid mediators identified through transfer learning captured substantial variation in AD severity, indirectly supporting the mediation analysis conclusion that these lipids may play important roles in AD pathology.

Second, to assess the stability of the TransHDM, we removed the lipid molecules initially identified by TransHDM from the AA population. No newly identified lipids were identified with a frequency >80%, suggesting that no new stable lipid mediators were found after removing the previously identified ones. However, when excluding the lipid mediators identified from HDM and reapplying TransHDM, new lipid molecules were rediscovered, and the lipids identified were largely consistent with those from the original analysis (Tables S18–S21). These results further confirm the stability and consistency of the transfer learning approach, indicating that the exclusion of previously identified lipids did not lead to divergent or inconsistent findings.

Additionally, to examine the impact of transferable source data size on mediator identification, we compared the number of lipid mediators detected under four settings: AA (without transfer learning), AAT1 (using one-third of random transferable samples), AAT2 (two-thirds), and AAT3 (the full set). As shown in Fig. S7 and Tables S22–S25, the number of identified mediators increases from AA to AAT3, suggesting that more transferable data enhance mediator detection. Moreover, most lipid mediators identified under different settings show substantial overlap, demonstrating the robustness of mediator selection.

We examined the impact of varying Inline graphic retained by SIS on the results (Fig. S8). When Inline graphic is low, stringent screening that missed many potential mediators may result in elevated Type II error rates. As s increases from 50 to 100 or 150, including more true mediators boosts detection ability, leading to more significant lipids identified. The lipids detected at lower Inline graphic exhibit overlap with those identified at higher Inline graphic, underscoring the robustness of the results across varying Inline graphic. However, when s further increases to 200, the inclusion of noise lipids may increase model complexity, reducing the identification of lipid mediators. Consequently, retaining 100 to 150 variables emerges as a relatively optimal balance, as adopted in our primary analysis.

Discussion

This study proposes a novel transfer learning–based high-dimensional mediation analysis framework designed to address the challenges of mediator selection and effect estimation in data-scarce settings. By integrating external transferable datasets, this approach significantly enhances the accuracy, robustness, and false discovery rate control in mediation analysis. Compared to traditional methods, TransHDM effectively leverages external data to overcome for small sample sizes limitations, thereby improving statistical power in high-dimensional data analysis. Applied to the ADNI cohort, this framework identified four significant lipid metabolism pathways—glycerolipid metabolism, glycerophospholipid metabolism, sphingolipid metabolism, and ether lipid metabolism—that significantly mediate the APOE ε4–AD association in both non-Hispanic White and African American populations. These pathways, critical for lipid homeostasis and neuroinflammation regulation, reveal shared metabolic mechanisms underlying AD across racial groups.

TransHDM integrates transfer learning with high-dimensional mediation analysis, providing a robust solution for small sample scenarios. The successful application not only advances AD research in African American populations but also offers a promising approach for studying other data-scarce groups, such as rare disease cohorts, specific subpopulations, and early-stage disease patients [35]. Traditional methods often struggle to identify potential biomarkers or elucidate pathological mechanisms in these contexts due to insufficient sample sizes. By integrating external data, transfer learning can mitigate these limitations and enhance the exploration of complex disease mechanisms, ultimately contributing to the development of targeted interventions tailored to diverse populations.

Given these advantages, the proposed framework demonstrates both methodological innovation and excellent practical applicability in disease mechanism research and personalized treatment strategy development. Our findings reveal consistent lipid metabolism pathways mediating the APOE ε4–AD relationship in both African American and non-Hispanic White populations. The four pathways play crucial roles in lipid homeostasis, cell membrane integrity, and neuroinflammation regulation, suggesting that despite differences in genetic background, environmental exposure, and lifestyle, the role of lipid metabolism may universally contribute to AD pathology. Our findings indicate that lipid metabolism-based AD biomarkers and therapeutic strategies developed in White populations may also be applicable to Black populations. This cross-racial applicability not only accelerates the translational application of AD research but also offers new approaches to address health disparities arising from insufficient research and delayed development of effective biomarkers and treatments.

Our findings are supported by previous studies demonstrating minimal differences in serum lipid levels between African American (AA) and non-Hispanic White (NHW) populations [36]. This consistency further validates the shared lipid metabolism pathways across racial groups. Numerous lipidomics studies in NHW populations have identified significant differential expression of glycerophospholipids, glycerolipids, and sphingolipids in AD patients [37–40]. Similarly, in African Americans, glycerophospholipids and sphingolipids have shown prospective associations with mild cognitive impairment and dementia [41, 42]. Collectively, these studies underscore the consistent role of lipid metabolism in AD across racial groups, strongly supporting our findings.

In our analysis of the African American population, ether lipid metabolism significantly mediated the association between APOE ε4 and ADAS-Cog 13. However, no significant mediation effects were observed for t-tau, ABeta42, or WMH outcomes, even after applying transfer learning. This discrepancy may be due to the smaller sample sizes for these three outcomes, limiting statistical power and the ability to fully capture ether lipid metabolism’s mediating effects. Growing evidence suggests that inflammation plays an important role in AD pathogenesis, with ether lipid metabolism closely associated with inflammatory processes [43]. Recent studies indicated a strong association between ether lipids and ferroptosis, an iron-dependent form of cell death characterized by lipid peroxidation and oxidative stress accumulation, which is implicated in neurodegenerative diseases such as AD [44–46]. Ether lipid metabolism may regulate AD progression through inflammatory and immune modulation mechanisms [47]. Future research should expand sample size to further explore the critical role of ether lipids in mediating these biological processes.

In addition, the outcomes used in our study capture distinct dimensions of AD severity. WMH volume primarily reflects cerebrovascular lesions or small vessel diseases, while ADAS-Cog 13 score, t-tau levels, and ABeta42 levels more directly assess AD-specific pathological features, including cognitive function, tau protein pathology, and amyloid deposition [48]. Despite the limited correlation among these four indicators, the lipid molecular types identified through mediation analysis were largely consistent. This suggests that these lipid molecules may play an important role in multiple pathophysiological processes and different stages of AD. Specifically, APOE ε4 may influence not only amyloid metabolism but also cerebrovascular health and neuroinflammation through these lipid-mediated mechanisms.

A notable limitation of this study is its focus on serum lipids rather than CSF lipids. The relationship between peripheral metabolites and APOE ε4–related pathological processes in the central nervous system remains poorly understood. While changes in peripheral metabolites may serve as potential biomarkers for AD risk, their clinical utility and accuracy require further validation. Additionally, in this study, the SIS threshold was manually determined in advance, which may limit the flexibility of variable selection. Future research could potentially optimize the selection of threshold in SIS using methods such as cross-validation, thereby improving the flexibility and accuracy of variable selection in high-dimensional data. Furthermore, to validate the universality of the key lipid metabolism pathways identified here, future studies should expand the sample size of African American populations and integrate multi-omics data, such as genomics and proteomics. These pathways should also be explored further for their potential applications in early AD diagnosis, risk prediction, and targeted therapy.

Conclusion

In conclusion, this study leverages transfer learning to reveal consistent lipid metabolism pathways in both African American and non-Hispanic White populations, providing critical scientific evidence for cross-racial AD pathology research and therapeutic strategy development. The proposed TransHDM framework offers a robust tool for addressing the challenges of high-dimensional data analysis in small sample populations and specific heterogeneous groups, such as rare disease cohorts. These findings not only advance our understanding of AD mechanisms but also pave the way for more inclusive and precise biomedical research.

Key Points

  • TransHDM is the first transfer learning–based high-dimensional mediation framework specifically designed to overcome the limitations of traditional methods in small-sample studies. By leveraging large-scale external data and employing transfer regularization to mitigate population heterogeneity, it significantly enhances the detection power for true mediators while rigorously controlling family-wise error rates (demonstrated in simulation studies).

  • Applied to the ADNI cohort, TransHDM identified four APOE ε4–mediated lipid metabolism pathways (glycerophospholipid, glycerolipid, sphingolipid, and ether lipid metabolism) in African American populations. These pathways—previously undetectable with conventional methods due to limited sample sizes—provide novel insights into Alzheimer’s disease mechanisms and facilitate biomarker discovery for underrepresented groups.

  • These lipid pathways mediating APOE ε4 effects in African Americans were also validated in White populations, suggesting evolutionarily conserved mechanisms of lipid homeostasis and neuroinflammation across ethnicities. This cross-racial consistency indicates that lipid-targeted therapies or biomarkers developed in majority populations could be directly applicable to African American individuals, potentially reducing health disparities.

Supplementary Material

Appendix_Figure_20250723_bbaf460
Appendix_Table_20250723_bbaf460

Contributor Information

Lulu Pan, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.

Yahang Liu, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.

Chen Huang, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.

Ruilang Lin, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.

Yongfu Yu, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China; Shanghai Key Laboratory of Gene Editing and Cell Therapy for Rare Diseases, Fudan University, 83 Fen Yang Road, Xuhui District, Shanghai 200031, China.

Guoyou Qin, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China; Shanghai Institute of Infectious Disease and Biosecurity, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.

Author contributions

L.L.P., Y.H.L., and C.H. conceived the idea and contributed to statistical analysis, interpretation of data, and the draft of the manuscript. R.L.L. contributed to the analysis of the data and revised the manuscript. G.Y.Q. and Y.F.Y. contributed to the conception of the study, overall supervision, and final editing of the manuscript. All authors read and approved the final manuscript.

Conflict of interest

None declared.

Funding

This work was supported by the National Natural Science Foundation of China (No. 82273730 to Y.F.Y., 82173612 to G.Y.Q.), Shanghai Rising-Star Program (21QA1401300 to Y.F.Y.), Shanghai Municipal Natural Science Foundation (22ZR1414900 to Y.F.Y.), the Three-Year Public Health Action Plan of Shanghai (GWVI-11.2-XD10 and GWVI-11.1-01 to Y.F.Y.), Shanghai Talent Programs (BJKJ2024050 to Y.F.Y.), and Shanghai Municipal Science and Technology Major Project (ZD2021CY001 to G.Y.Q.).

Data availability

All data used in the analyses reported here are available in the ADNI data repository (adni.loni.usc.edu).

Standard protocol approvals, registrations, and patient consents

Written informed consent was obtained at the time of enrollment for imaging and sample collection, and protocols of consent forms were approved by the Institutional Review Board at each participating site.

Code availability

The source code supporting this work is publicly available at https://github.com/PanLululu/TransHDM.

References

  • 1. Scheltens  P, De Strooper  B, Kivipelto  M. et al.  Alzheimer’s disease. Lancet  2021;397:1577–90. 10.1016/S0140-6736(20)32205-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Matthews  KA, Xu  W, Gaglioti  AH. et al.  Racial and ethnic estimates of Alzheimer’s disease and related dementias in the United States (2015-2060) in adults aged >/=65 years. Alzheimers Dement  2019;15:17–24. 10.1016/j.jalz.2018.06.3063 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. 2021 Alzheimer’s disease facts and figures. Alzheimers Dement  2021;17:327–406. 10.1002/alz.12328 [DOI] [PubMed] [Google Scholar]
  • 4. Farrer  LA, Cupples  LA, Haines  JL. et al.  Effects of age, sex, and ethnicity on the association between apolipoprotein E genotype and Alzheimer disease. A meta-analysis. APOE and Alzheimer Disease Meta Analysis Consortium. JAMA  1997;278:1349–56. 10.1001/jama.1997.03550160069041 [DOI] [PubMed] [Google Scholar]
  • 5. Torres-Perez  E, Ledesma  M, Garcia-Sobreviela  MP. et al.  Apolipoprotein E4 association with metabolic syndrome depends on body fatness. Atherosclerosis  2016;245:35–42. 10.1016/j.atherosclerosis.2015.11.029 [DOI] [PubMed] [Google Scholar]
  • 6. Yang  LG, March  ZM, Stephenson  RA. et al.  Apolipoprotein E in lipid metabolism and neurodegenerative disease. Trends Endocrinol Metab  2023;34:430–45. 10.1016/j.tem.2023.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Polis  B, Samson  AO. A new perspective on Alzheimer’s disease as a brain expression of a complex metabolic disorder. In: Wisniewski  T (ed.), Alzheimer’s Disease. Brisbane (AU): Codon Publications, 2019. [PubMed] [Google Scholar]
  • 8. Stepler  KE, Robinson  RAS. The potential of ‘omics to link lipid metabolism and genetic and comorbidity risk factors of Alzheimer’s disease in African Americans. Adv Exp Med Biol  2019;1118:1–28. 10.1007/978-3-030-05542-4_1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Power  MC, Rawlings  A, Sharrett  AR. et al.  Association of midlife lipids with 20-year cognitive change: a cohort study. Alzheimers Dement  2018;14:167–77. 10.1016/j.jalz.2017.07.757 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Bernath  MM, Bhattacharyya  S, Nho  K. et al.  Serum triglycerides in Alzheimer disease: relation to neuroimaging and CSF biomarkers. Neurology  2020;94:e2088–98. 10.1212/WNL.0000000000009436 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Wang  T, Huynh  K, Giles  C. et al.  APOE epsilon2 resilience for Alzheimer’s disease is mediated by plasma lipid species: analysis of three independent cohort studies. Alzheimers Dement  2022;18:2151–66. 10.1002/alz.12538 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Gianattasio  KZ, Bennett  EE, Wei  J. et al.  Generalizability of findings from a clinical sample to a community-based sample: a comparison of ADNI and ARIC. Alzheimers Dement  2021;17:1265–76. 10.1002/alz.12293 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Franzen  S, Smith  JE, van den  Berg  E. et al.  Diversity in Alzheimer’s disease drug trials: the importance of eligibility criteria. Alzheimers Dement  2022;18:810–23. 10.1002/alz.12433 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Babulal  GM, Quiroz  YT, Albensi  BC. et al.  Perspectives on ethnic and racial disparities in Alzheimer’s disease and related dementias: update and areas of immediate need. Alzheimers Dement  2019;15:292–312. 10.1016/j.jalz.2018.09.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Zhang  H, Zheng  Y, Zhang  Z. et al.  Estimating and testing high-dimensional mediation effects in epigenetic studies. Bioinformatics  2016;32:3150–4. 10.1093/bioinformatics/btw351 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Perera  C, Zhang  H, Zheng  Y. et al.  HIMA2: high-dimensional mediation analysis and its application in epigenome-wide DNA methylation data. BMC Bioinformatics  2022;23:296. 10.1186/s12859-022-04748-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Torrey  L, Shavlik  J. Transfer Learning. In: Olivas ES, Guerrero JD, Martinez-Sober M, Magdalena-Benedito JR, Serrano López AJ (eds.), Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques. Hershey (PA): IGI global, 2010. 242–64. [Google Scholar]
  • 18. Davila  A, Colan  J, Hasegawa  Y. Comparison of fine-tuning strategies for transfer learning in medical image classification. Image Vis Comput  2024;146:105012. 10.1016/j.imavis.2024.105012 [DOI] [Google Scholar]
  • 19. Dieckhaus  H, Brocidiacono  M, Randolph  NZ. et al.  Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc Natl Acad Sci  2024;121:e2314853121. 10.1073/pnas.2314853121 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Mahmud  T, Barua  K, Habiba  SU. et al.  An explainable ai paradigm for Alzheimer’s diagnosis using deep transfer learning. Diagnostics  2024;14:345. 10.3390/diagnostics14030345 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Li  S, Cai  TT, Li  H. Transfer learning for high-dimensional linear regression: prediction, estimation and minimax optimality. J R Stat Soc Ser B Stat Methodol  2022;84:149–73. 10.1111/rssb.12479 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Tian  Y, Feng  Y. Transfer learning under high-dimensional generalized linear models. J Am Stat Assoc  2023;118:2684–97. 10.1080/01621459.2022.2071278 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Li  S, Zhang  L, Cai  TT. et al.  Estimation and inference for high-dimensional generalized linear models with knowledge transfer. J Am Stat Assoc  2024;119:1274–85. 10.1080/01621459.2023.2184373 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Van de Geer  S, Bühlmann  P, Ya  R. et al.  On asymptotically optimal confidence regions and tests for high-dimensional models. Ann Statist  2014;42:1166–202. 10.1214/14-AOS1221 [DOI] [Google Scholar]
  • 25. Meinshausen  N, Bühlmann  P. High-dimensional graphs and variable selection with the lasso. Ann Statist  2006;34:1436–62. 10.1214/009053606000000281 [DOI] [Google Scholar]
  • 26. Zhang  Y, Zhu  Z. Transfer learning for high-dimensional quantile regression via convolution smoothing. Stat Sin  2025;35:939–58. 10.5705/ss.202022.0396 [DOI] [Google Scholar]
  • 27. Fan  J, Lv  J. Sure independence screening for ultrahigh dimensional feature space. J R Stat Soc Ser B Stat Methodol  2008;70:849–911. 10.1111/j.1467-9868.2008.00674.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Dai  JY, Stanford  JL, LeBlanc  M. A multiple-testing procedure for high-dimensional mediation hypotheses. J Am Stat Assoc  2022;117:198–213. 10.1080/01621459.2020.1765785 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Weiner  MW, Veitch  DP, Aisen  PS. et al.  Recent publications from the Alzheimer’s Disease Neuroimaging Initiative: reviewing progress toward improved AD clinical trials. Alzheimers Dement  2017;13:e1–85. 10.1016/j.jalz.2016.11.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Saykin  AJ, Shen  L, Yao  X. et al.  Genetic studies of quantitative MCI and AD phenotypes in ADNI: progress, opportunities, and plans. Alzheimers Dement  2015;11:792–814. 10.1016/j.jalz.2015.05.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Huynh  K, Barlow  CK, Jayawardana  KS. et al.  High-throughput plasma lipidomics: detailed mapping of the associations with cardiometabolic risk factors. Cell Chem Biol  2019;26:71–84.e4. 10.1016/j.chembiol.2018.10.008 [DOI] [PubMed] [Google Scholar]
  • 32. DeCarli  C, Fletcher  E, Ramey  V. et al.  Anatomical mapping of white matter hyperintensities (WMH): exploring the relationships between periventricular WMH, deep WMH, and total WMH burden. Stroke  2005;36:50–5. 10.1161/01.STR.0000150668.58689.f2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Schwarz  C, Fletcher  E, DeCarli  C. et al.  Fully-automated white matter hyperintensity detection with anatomical prior knowledge and without FLAIR. Inf Process Med Imaging  2009;5636:239–51. 10.1007/978-3-642-02498-6_20 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Bittner  T, Zetterberg  H, Teunissen  CE. et al.  Technical performance of a novel, fully automated electrochemiluminescence immunoassay for the quantitation of beta-amyloid (1-42) in human cerebrospinal fluid. Alzheimers Dement  2016;12:517–26. 10.1016/j.jalz.2015.09.009 [DOI] [PubMed] [Google Scholar]
  • 35. Johansson  A, Andreassen  OA, Brunak  S. et al.  Precision medicine in complex diseases-molecular subgrouping for improved prediction and treatment stratification. J Intern Med  2023;294:378–96. 10.1111/joim.13640 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Carnethon  MR, Pu  J, Howard  G. et al.  Cardiovascular health in African Americans: a scientific statement from the American Heart Association. Circulation  2017;136:e393–423. 10.1161/CIR.0000000000000534 [DOI] [PubMed] [Google Scholar]
  • 37. Whiley  L, Sen  A, Heaton  J. et al.  Evidence of altered phosphatidylcholine metabolism in Alzheimer’s disease. Neurobiol Aging  2014;35:271–8. 10.1016/j.neurobiolaging.2013.08.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Klavins  K, Koal  T, Dallmann  G. et al.  The ratio of phosphatidylcholines to lysophosphatidylcholines in plasma differentiates healthy controls from patients with Alzheimer’s disease and mild cognitive impairment. Alzheimers Dement (Amst)  2015;1:295–302. 10.1016/j.dadm.2015.05.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Li  D, Misialek  JR, Boerwinkle  E. et al.  Plasma phospholipids and prevalence of mild cognitive impairment and/or dementia in the ARIC Neurocognitive Study (ARIC-NCS). Alzheimers Dement (Amst)  2016;3:73–82. 10.1016/j.dadm.2016.02.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Mielke  MM, Bandaru  VV, Haughey  NJ. et al.  Serum sphingomyelins and ceramides are early predictors of memory impairment. Neurobiol Aging  2010;31:17–24. 10.1016/j.neurobiolaging.2008.03.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Li  D, Misialek  JR, Boerwinkle  E. et al.  Prospective associations of plasma phospholipids and mild cognitive impairment/dementia among African Americans in the ARIC Neurocognitive Study. Alzheimers Dement (Amst)  2017;6:1–10. 10.1016/j.dadm.2016.09.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Khan  MJ, Chung  NA, Hansen  S. et al.  Targeted lipidomics to measure phospholipids and sphingomyelins in plasma: a pilot study to understand the impact of race/ethnicity in Alzheimer’s disease. Anal Chem  2022;94:4165–74. 10.1021/acs.analchem.1c03821 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Jofre-Monseny  L, Minihane  AM, Rimbach  G. Impact of apoE genotype on oxidative stress, inflammation and disease risk. Mol Nutr Food Res  2008;52:131–45. 10.1002/mnfr.200700322 [DOI] [PubMed] [Google Scholar]
  • 44. Zou  Y, Henry  WS, Ricq  EL. et al.  Plasticity of ether lipids promotes ferroptosis susceptibility and evasion. Nature  2020;585:603–8. 10.1038/s41586-020-2732-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Kim  JW, Lee  JY, Oh  M. et al.  An integrated view of lipid metabolism in ferroptosis revisited via lipidomic analysis. Exp Mol Med  2023;55:1620–31. 10.1038/s12276-023-01077-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Bao  WD, Pang  P, Zhou  XT. et al.  Loss of ferroportin induces memory impairment by promoting ferroptosis in Alzheimer’s disease. Cell Death Differ  2021;28:1548–62. 10.1038/s41418-020-00685-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Chen  X, Kang  R, Kroemer  G. et al.  Ferroptosis in infection, inflammation, and immunity. J Exp Med  2021;218:e20210518. 10.1084/jem.20210518 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Jack  CR  Jr, Bennett  DA, Blennow  K. et al.  NIA-AA Research Framework: toward a biological definition of Alzheimer’s disease. Alzheimers Dement  2018;14:535–62. 10.1016/j.jalz.2018.02.018 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix_Figure_20250723_bbaf460
Appendix_Table_20250723_bbaf460

Data Availability Statement

All data used in the analyses reported here are available in the ADNI data repository (adni.loni.usc.edu).


Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press

RESOURCES