Abstract
Lipid-mediated effects play a crucial role in elucidating the pathological mechanisms linking the ε4 allele of the apolipoprotein E gene (APOE ε4) to Alzheimer’s disease (AD). However, traditional mediation analysis methods often suffer from insufficient statistical power in studies involving minority populations due to limited sample sizes. This study innovatively develops a high-dimensional mediation analysis model (TransHDM) based on a transfer learning framework. By leveraging information from source data with large-scale samples, it significantly enhances the ability to identify potential mediators in small sample target data. The method first constructs a high-dimensional regression model using aggregated data from the source data and target data, then applies transfer regularization to adjust for heterogeneity between the source and target domains, correcting for estimation bias in high-dimensional Lasso. Ultimately, it achieves parameter transfer across domains, addressing statistical bias and inferential uncertainty caused by small sample sizes. Simulation results demonstrate that, compared to traditional methods, this approach significantly improves the power in identifying true mediator variables while effectively controlling the family-wise error rate in multiple testing. When applied to the Alzheimer’s Disease Neuroimaging Initiative cohort, TransHDM transferred large-scale data from white and other ethnic groups, identifying additional lipid metabolic pathways mediating the influence of the APOE ε4 allele on AD pathological progression in African American populations compared to pre-transfer analysis. These pathways include glycerophospholipid metabolism, glycerolipid metabolism, sphingolipid metabolism, and ether lipid metabolism (false discovery rate < 0.05). The TransHDM framework not only provides a powerful methodological tool for small sample population research but also offers valuable insights for future research in exploring disease mechanisms and developing biomarkers for disease prediction.
Keywords: transfer learning, mediation analysis, Alzheimer’s disease, lipidomics, African Americans, external dataset
Introduction
Alzheimer’s disease (AD), a complex neurodegenerative disorder with multifactorial etiologies, has become a significant public health challenge amid the global aging population [1]. Epidemiological studies reveal substantial ethnic disparities in AD incidence, with African Americans exhibiting a markedly higher prevalence compared to non-Hispanic Whites and Asian Americans [2]. This population not only bears a greater disease burden but also faces limited access to healthcare resources, highlighting the urgent need to enhance AD awareness and research within the African American community [3]. The racial disparities in AD are driven by a combination of genetic and environmental factors. Among these, the ε4 allele of the apolipoprotein E gene (APOE ε4) was a strong risk factor for AD [4]. Research indicates that APOE ε4 is associated with an increased prevalence of metabolic syndrome, with these metabolic disturbances often preceding the onset of AD pathological changes [5, 6]. These findings suggest that alterations in lipid metabolism may serve as a key molecular mechanism linking APOE ε4 to AD pathogenesis, potentially offering novel biomarkers and therapeutic targets [7]. However, the role of lipid metabolism in AD across different racial groups remains poorly understood. While associations between lipid metabolism and AD have been established in non-Hispanic Whites, their applicability to African Americans remains unclear [8]. Therefore, comprehensive lipidomic analyses in African Americans are essential to explore both shared and unique lipid metabolic features across racial groups, thereby validating the generalizability of findings from White populations and informing the development of universal biomarkers and therapeutic strategies.
Previous studies into lipid metabolism as a mediator of the APOE ε4–AD association have focused on specific lipid molecules or metabolic pathways, failing to capture the complexity of systemic lipid metabolic changes [9, 10]. Although lipidomics—an innovative technology enabling simultaneous analysis of diverse lipid molecules—offers a more comprehensive approach [8], its application in AD research has predominantly focused on non-Hispanic Whites, with African Americans significantly underrepresented [11]. This disparity is largely attributable to the small sample sizes of African American cohorts in major studies. For instance, the Alzheimer’s Disease Neuroimaging Initiative (ADNI), despite encompassing 59 sites across North America and collecting data from thousands of individuals, includes African American participants in ~10% of its total sample [12]. This underrepresentation, a common issue in large datasets and AD-related clinical trials [13], severely limits statistical power, complicating mediation analysis in high-dimensional data settings and hindering a comprehensive understanding of lipid metabolism’s role in AD pathogenesis [14]. Consequently, the development of effective prevention and intervention strategies for minority racial groups remains challenging.
In recent years, mediation analysis has incorporated dimensionality reduction techniques such as variable selection and regularization to address the high-dimensionality and correlations inherent in multi-omics data. For example, the HIMA method proposed by Zhang et al. uses the minimax concave penalty to first select mediators and then test for joint mediation effects, which may improve false positives [15]. Building on this, Perera et al. introduced the HIMA2 method, which utilizes a debiased LASSO technique for simultaneous mediator selection and testing, better controlling for false positives and improving statistical power [16]. However, these methods still face limitations in small sample scenario, often resulting in reduced statistical power. Transfer learning has emerged as a promising solution to address these limitations. This approach enhances model performance on target data by transferring information from external source data, demonstrating significant potential in practice [17–20]. Advances in data acquisition technologies have made it increasingly feasible to obtain related datasets from source domains, enabling researchers to utilize large-sample data (i.e. non-Hispanic White) to improve model performance for African Americans. Recent developments in high-dimensional penalized regression-based transfer learning frameworks have further advanced the field. For example, Li et al. proposed a high-dimensional two-step transfer learning framework based on linear regression [21]. Building on this, Tian et al. and Li et al. introduced algorithms for constructing confidence intervals for transfer learning estimates using a debiased LASSO approach, providing theoretical proof for reliable statistical inference [22, 23]. Despite these advancements, the application of transfer learning in mediation analysis remains underexplored, presenting a significant opportunity for innovation.
To address these gaps, this study integrates transfer learning methods within the HIMA2 framework and proposes a transfer learning–based high-dimensional mediation analysis (TransHDM) framework, aiming to overcome the challenges of sparse target data in mediator selection and mediation effect inference. Figure 1 provides a graphical overview of our research. We assess the performance of our method through extensive numerical evaluations and apply it to the ADNI study to investigate the mediating role of lipid metabolism in the APOE ε4–AD association. Our analysis focuses on lipid metabolic features in African Americans, aiming to explore both shared mechanisms and potential heterogeneity compared to non-Hispanic Whites. This study not only advances our understanding of cross-racial lipid metabolic characteristics in APOE ε4–associated AD but also introduces an innovative framework for high-dimensional mediation analysis in data-scarce settings.
Figure 1.
Overview of this study. (a) Challenges in high-dimensional mediation analysis AA. Current high-dimensional mediation analysis methods lack sufficient statistical power when the target sample size is small. (b) Transfer learning–based high-dimensional mediation analysis (TransHDM) framework and solutions to technical difficulty. First, the method evaluates whether external data meets transfer learning criteria by calculating the loss function in cross-validation and comparing the similarity in variable associations between source and target data. Next, a high-dimensional LASSO model is built using data from both the source and target domains to obtain summary estimates. Transfer regularization then adjusts for heterogeneity between the source and target domains while correcting for biases in high-dimensional LASSO estimates. (c) A schematic diagram of the simulation process under different settings and the analyzing steps in the ANDI cohort application. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity.
Materials and methods
TransHDM framework
In the analytical workflow, we first identify transferable data sources, then perform transfer learning–based mediation analysis on the complete dataset. The TransHDM framework is an extension of the HIMA2 method proposed by Perera et al. [16], which enhances the estimation of mediation effects by incorporating covariate adjustment and integrating transfer learning algorithms to leverage information from external data sources. Since transfer learning with debiased lasso regression is the core algorithm throughout this process, we introduce this algorithm first in Section 2.1.1 to provide readers with the necessary technical foundation for better understanding and application of the subsequent methods.
Transfer learning with debiased lasso regression
Suppose there is a target dataset
with sample size
and
independent source datasets
, where
is the index set of transferable sources, which is assumed to be known. The sample size of the
-th source is
. The regression model corresponding to the target dataset or the
-th source datasets is
![]() |
(1) |
where
is the coefficient of the target dataset and
is the coefficient of the
-th source for
.
Our goal is to transfer useful information from the sources to the target, improving the estimation accuracy of
. Algorithm 1 presents a transfer learning framework for high-dimensional regression using debiased lasso, which involves the following two key steps:
Algorithm 1.
Transfer learning with debiased lasso regression.
Input: Target data , source data , and known transferable set . |
| 1. Initial lasso estimation: |
| Fit the regression model in all dataset, compute |
|
| 2. Debiased estimate (two-step correction) |
| 2.1 Transfer learning bias correction: |
|
where
|
| 2.2 Lasso bias correction |
|
|
Output: and . |
Step 1. Initial lasso estimation:
We first consider that index set
for the transferable source is known. The regression model is first fitted using both target and sources to obtain the initial lasso estimate:
![]() |
(2) |
where
denotes the loss function based on the data,
is the regularization parameter, and
enforces sparsity in the solution.
Step 2. Debiased estimate (two-step correction).
Step 2.1. First debiasing step (transfer learning bias correction):
Notably, the initial estimate
obtained from the combined source and target data may be influenced by discrepancies between the two domains. To mitigate the potential bias introduced by the source data, we introduce a correction vector
is introduced, which is a sparse vector that adjusts the estimate to better fit the target distribution:
![]() |
(3) |
is the regularization parameter, and
denotes the ℓ1-norm, which encourages sparsity in the correction vector
. Then, the debiased estimator is given by
![]() |
(4) |
Step 2.2. Second debiasing step (lasso bias correction):
While the first step addresses the bias introduced by transfer learning, the estimator
may still be biased due to the nature of lasso regression in high-dimensional settings [24]. To correct this, a second debiasing step is performed using nodewise regression, which estimates the inverse covariance matrix
of the features to further refine the coefficients [25]. Motivated by Tian et al. [22], the final debiased estimator is given by
![]() |
(5) |
For each feature
, nodewise regression estimates the conditional dependencies among features by regressing
on the remaining features
. The regression coefficients are estimated using
![]() |
(6) |
where
![]() |
(7) |
and
![]() |
(8) |
Then, we define the matrix
and diagonal matrix
as
![]() |
(9) |
and
![]() |
(10) |
where
.
Finally, the inverse covariance matrix is computed as
, and the variance estimator of
is defined as
![]() |
(11) |
where
.
Detecting transferable datasets
While the initial model fitting utilizes both the target and source datasets, not all source datasets may contribute effectively to improving the performance of target model. Thus, it is necessary to evaluate the transferability of each source dataset to ensure that only relevant information is incorporated.
To determine which source datasets can contribute effectively to the target model, the target dataset
is first randomly split into two disjoint equal-size groups: a training set
and a validation set
. Using the training set only and run the Algorithm 1 to estimate the coefficients
. For each source dataset
, we then combine it with the target training set and run the Algorithm 1 to estimate the coefficients
. These coefficients are evaluated on the target validation set
, and the transferability of each source dataset is measured using the statistic
![]() |
(12) |
The set of transferable source datasets is denoted as
![]() |
(13) |
where
can be chosen by cross-validation. The threshold 0.01 in the above formula is chosen based on prior literature [26]. Results corresponding to thresholds of 0.02 and 0.03 are provided in Table S1, indicating that the estimates remain stable within a small range.
This ensures that only relevant source datasets are retained for subsequent analyses, improving the robustness of the transfer learning framework.
Transfer learning for high-dimensional mediation analysis
This work is the first to integrate transfer learning into the mediation analysis framework, allowing for more accurate mediation effect estimation and mediator selection when dealing with high-dimensional mediators. In the case of high-dimensional mediators, we denote the exposure as
, baseline covariates to be adjusted for as
, where
is a
-dimensional vector, including all exposure–outcome, mediator–outcome, and exposure–mediator confounders. The mediators are denoted as
, where
is a
-dimensional matrix and
exceeds the sample size
. The outcome model for
and mediator model for
can be expressed as the following mediator models and outcome model to estimate and test mediation effects:
![]() |
(14) |
where
and
are the intercept terms for the mediator and outcome models, respectively.
represents the direct effect of
on
adjusting for all mediators, and the mediation effect of
on
through
can be defined as
. Specifically,
represent the coefficients for the effect of exposure
on each of the mediators
, adjusted for baseline covariates
, and
represent the coefficients for the effect of all mediators
on the outcome
, adjusting for exposure
and covariates
.
and
are the error terms, assumed to be independent and identically distributed normal variables with a mean of zero and constant variance.
In real-world data, mediation effects often interact with each other. For example, in biological pathways, changes in the expression or activity of one protein may trigger alterations in the expression or function of multiple downstream proteins, forming a complex network of interactions. Therefore, it is more reasonable for outcome model to treat all potential mediators as a whole. This approach allows for capturing the correlations among mediators without the need to explicitly define the causal relationships or mechanisms between them. Such an overall modeling strategy avoids biases caused by the omission of critical causal pathways and is particularly suitable for high-dimensional data, as it can reflect the cooperative effects and modular functional characteristics of complex biological systems to a certain extent.
The detailed estimation and testing procedure for the proposed high-dimensional mediation testing framework is given as follows:
Step 1. Sure independence screening.
When the number of mediators
is much larger than the sample size
, we can first employ the sure independence screening (SIS) to reduce the dimensionality [27]. We integrate SIS with transfer learning to effectively identify relevant mediators while leveraging information from both the target and source datasets. Notably, the application of transfer learning assumes that the target and source datasets share the same outcome, exposure, and mediator variables—an assumption commonly adopted in existing transfer learning frameworks [21, 22]. To reduce the number of mediators, we first consider the following marginal models for each mediator
,
:
![]() |
(15) |
Using Algorithm 1, we estimate the coefficients
and
for mediator
in the marginal models with target data
and source data
. Next, we compute the product term
to roughly quantify the mediation effect of
on
through
. To balance dimensionality reduction with the retention of potentially relevant mediators, mediators are ranked by the magnitude of
. Following a reference criterion commonly adopted in high-dimensional settings, threshold
can be utilized to guide the inclusion of mediators into set
for the following analyses [15, 16]. This step effectively excludes irrelevant variables, reducing the dimensionality of the mediator set, improving model stability and computational efficiency in high-dimensional scenarios, and reducing the likelihood of false discoveries in mediator selection. Without using SIS, the model may face issues such as convergence difficulties, unstable parameter estimates, and a significant increase in computational costs. It is worth noting that although the
is fixed at 100 in our simulations for simplicity, it can be manually adjusted based on sample size or specific domain knowledge to assess the robustness and stability of the method [15, 16].
Step 2. Estimation and testing of mediation effects.
For the mediators in the selected set
, first, we estimate the exposure–mediator effect
using Algorithm 1 applied to the exposure–mediator model, for each mediator
. The corresponding
-value for the test of significance of
is calculated as
![]() |
(16) |
where
is the cumulative distribution function of
. Similarly, the mediator–outcome effect
is estimated using Algorithm 1 applied to the mediator–outcome model, and the significance of
is tested using
![]() |
(17) |
The variance estimates
and
refer to the variance of the coefficients
and
, respectively, and are calculated using Equation (11) in Section 2.1.1.
Step 3. Joint significance testing.
To determine the overall significance of each mediator, we perform a joint significance test by combining the
-value for
and
:
![]() |
(18) |
The P-values
are adjusted to
using the “JS-mixture” procedure, which controls the family-wise error rate (FWER) across multiple hypothesis tests [28]. Mediators with
are considered significant, and the set of significant mediators
is denoted as
![]() |
(19) |
The details of the algorithm are shown in Algorithm 2.
Algorithm 2.
Transfer learning–based estimation and testing framework for high-dimensional mediation analysis (TransHDM).
Input target data , source data
|
| 1. Sure independence screening |
|
| 2. Estimate and test exposure–mediator effects and mediator–outcome effects |
Compute and via Algorithm 1 based on outcome model (1) and (2) |
, for
|
, for
|
| 3. Joint significance test |
, for
|
, where
|
Output and
|
Simulation data generation
We conduct extensive simulations to evaluate the performance of the proposed method, TransHDM. The HDM method can be viewed as a simplified version of TransHDM, as it does not incorporate the transfer learning strategy. Both methods are extensions of the HIMA2 framework, with further adjustment for confounding variables in the model. However, while HDM uses debiased Lasso to estimate the model coefficients for the target data, TransHDM further integrates transfer learning, leveraging external source data to improve the mediation effect estimation. By comparing TransHDM with HDM, we demonstrate the advantages of the transfer learning strategy in enhancing the performance of mediation effect estimation and mediator selection.
For the target with
and
sources with
, we consider two covariate distribution settings: (1) homogeneous design: the covariates
are independently generated from
for all source and target samples, where
; (2) heterogeneous design: the covariates
in each source
are generated from
, where
. Here,
is a random matrix with each entry equal to 0.3 with probability 0.3 and 0 otherwise. The target domain follows the homogeneous design with covariance
. This setting induces covariate shift across source domains. Since both the exposure
and mediators
are generated conditional on the covariates
, the heterogeneity in
further propagates to
and
, resulting in induced distributional shifts in these variables as well.
The exposure
is generated as
![]() |
where
and
. The mediator
for
is generated as
![]() |
where
and
, and the error terms
are drawn from the multivariate normal distribution with mean zero and covariance matrix
, where
is the first-order autoregressive correlation structure with
The outcome is generated as
![]() |
where
,
, and
.
For the target, we set
for the first eight mediators, and
otherwise. We set the coefficient
for the first eight mediators and
otherwise. Therefore, we have (i)
for
; (ii)
but
for
; (iii)
but
for
; and (iv)
and
for
.
We randomly select
as the number of transferable sources. The coefficient
is defined as
![]() |
Similarly, for
, we define
![]() |
where
is a random subset of
with
for
, reflecting the heterogeneity between target and sources.
In the above setting, we set
. We define
and
. All the simulations are based on 200 replications with 6 different settings:
and
.
Study sample
Participants in this study were from the ADNI cohort, a longitudinal study designed to identify clinical, genetic, imaging, and biological markers of early AD progression. The initial phase, ADNI-1, was launched in 2003 and recruited participants aged 55–90 from 63 sites across the USA and Canada. Subsequent phases (ADNI-GO, ADNI-2, and ADNI-3) tracked existing participants and enrolled additional cohorts. Detailed descriptions of ADNI’s design are available elsewhere [29, 30]. Demographic information, APOE genotype, questionnaire data, lipid metabolism data, neuroimaging data, and cerebrospinal fluid (CSF) biomarker data were obtained from the ADNI data repository (adni.loni.usc.edu).
For this analysis, 1524 eligible participants with complete plasma lipid metabolism data, exposure and covariate information, and participation in ADNI-1, ADNI-GO, or ADNI-2 were included. The exposure of interest was the number of APOE-ε4 alleles carried by participants, and 781 serum lipid metabolites measured at baseline were considered as mediators (detailed information can be found in Section 2.4). Covariates, selected based on prior literature and database availability, included age, gender, and years of education. The target outcomes were AD severity, assessed using three dimensions: (1) the 13-item version of the Alzheimer’s Disease Assessment Scale—Cognitive Subscale (ADAS-Cog 13) score, a widely accepted cognitive measure; (2) CSF concentrations of t-tau and amyloid-beta (1–42) (ABeta42), which reflect AD-related pathological changes; and (3) white matter hyperintensity (WMH) volume, an imaging marker of structural brain changes associated with AD. All outcomes were measured after baseline to ensure appropriate temporal relationships between exposure (APOE-ε4), mediators (lipid metabolism), and outcomes (AD severity). The detailed measurement processes for outcomes are provided in Sections 2.5–2.7.
To retain more samples for subsequent analysis, participants with missing values for each outcome (ADAS-Cog 13, t-tau, ABeta42, or WMH) were excluded separately, resulting in four analytic datasets comprising 1524, 1130, 1133, and 915 participants, respectively. The sample selection and exclusion process are detailed in Supplementary Fig. S1. All participants provided written informed consent, and the study protocol was approved by the Institutional Review Board at each participating site.
Lipid metabolomic
Lipid analysis data utilized in this study were obtained directly from the ADNI database. These lipid measurements had previously been conducted by ADNI using plasma samples collected from a subset of participants. The samples were analyzed at the Metabolomics Laboratory of the Baker Heart and Diabetes Institute, where lipid profiling of 781 plasma lipids was performed according to methods described previously [31]. See Table S2 for specific classes and species. To account for batch effects and differences between the ADNI-1 longitudinal cohort and ADNI-2/GO cohorts, the median concentration of each analyte was adjusted. Additionally, we normalized all data beforehand to ensure comparability across lipid molecules on the same scale.
Neuroimaging analysis
The WMH data were obtained from the ADNI dataset. As previously described, all scans were preprocessed using a standardized pipeline [32]. Detailed information on the method for WMH detection has been published [33]. Specifically, WMH detection was performed using a Bayesian Markov Random Field approach, which leverages a vector of three image intensities (proton density [PD], T1, and T2) associated with image pixels.
CSF biomarker analysis
CSF biomarker data were obtained from the ADNI dataset. As previously described, AD-related biomarkers, including t-tau and ABeta42, in CSF samples from ADNI-1/GO/2 were analyzed using the validated and highly automated Roche Elecsys electrochemiluminescence immunoassay [34]. This method significantly improved both within-laboratory and between-laboratory precision and accuracy, while also enhancing lot-to-lot consistency of the immunoassay kits.
Cognitive assessment
The cognitive performance was assessed using the modified Alzheimer’s Disease Assessment Scale—Cognitive Subscale. This scale comprises 13 items evaluating memory, language, praxis, and orientation abilities. Missing scores for any single item resulted in exclusion of the total ADAS-Cog 13 score. The scale ranges from 0 to 85, with higher scores indicating more severe cognitive impairment.
Statistical analysis
The datasets corresponding to the four AD severity indicators (ADAS-Cog 13, t-tau, ABeta42, WMH) were divided into target sample for African Americans (n = 62, 37, 37, 40, respectively) and source sample including non-Hispanic Whites, Asian, and other race (n = 1415, 1054, 1062, 841, respectively). The transferability of source datasets was pre-identified using a transferability recognition algorithm. Specifically, datasets from Non-Hispanic White, Asian, and Other racial groups were evaluated based on metrics including loss of source, loss of validation, and threshold values, as detailed in Table S3. Based on this analysis, datasets meeting the transferability threshold criteria were selected as transferable source datasets for subsequent analyses. High-dimensional mediation analysis was first conducted on African American individuals using HDM without transfer learning. To leverage external information, TransHDM was then applied to incorporate transferable knowledge from other racial groups. For comparison, HDM was also performed separately on non-Hispanic White individuals. Although the source dataset included participants from multiple racial backgrounds, we restricted the source population to non-Hispanic White individuals due to the insufficient sample sizes of other groups, which limited their suitability for reliable estimation. Lipid molecules with a selection frequency exceeding 80% from 100 repeated samplings are chosen as potential mediator variables to enhance the robustness of selection. Multiple testing adjustments for mediator variables were conducted using the “JS-mixture” method to control the FWER, with statistical significance defined as
<.05.
Results
Simulation study
To assess the effectiveness of TransHDM in identifying true mediators, we simulated multiple experiments with varying source data sample sizes (
200, 400, 600), mediator dimensions (
1000 and 2000), different degrees of mediator correlation (low:
, moderate:
, high:
), and covariate covariance structures (homogeneous and heterogeneous designs). These settings simulate common correlation patterns of high-dimensional omics data in real world (for data generation details, see “Simulation data generation” in the Methods section). Transferable detecting algorithm demonstrated robust performance in accurately identifying transferable sources, achieving perfect identification across all tested scenarios (Table S4). Importantly, TransHDM consistently outperformed traditional high-dimensional mediation (HDM) methods in accuracy and precision in mediator selection as well as mediation effect estimation (Figs. 2–3).
Figure 2.
Results of simulation experiments under homogeneous design. (a) Performance comparison of TransHDM and HDM across 18 simulation settings, focusing on the accuracy (Power) of selecting true mediators and the false discovery rate (FWER) of selecting irrelevant variables. (b–d) Performance comparison of TransHDM and HDM across 18 simulation settings in estimating the root mean squared error, relative bias and standard deviation of direct effects, mediated effects, and mediation proportions.
was the correlation coefficient among mediators,
was the dimension of mediators and
was the number of transferable sources.
Figure 3.
Results of simulation experiments under heterogeneous design. (a) Performance comparison of TransHDM and HDM across 18 simulation settings, focusing on the accuracy (Power) of selecting true mediators and the false discovery rate (FWER) of selecting irrelevant variables. (b–d) Performance comparison of TransHDM and HDM across 18 simulation settings in estimating the root mean squared error, relative bias, and SD of direct effects, mediated effects, and mediation proportions.
was the correlation coefficient among mediators,
was the dimension of mediators, and
was the number of transferable sources.
Figure 2(a) and Fig. 3(a) report the FWER and statistical power of mediator selection under homogeneous and heterogeneous designs, respectively. As the number of sources (
) increased, power improved, exceeding 0.95. Even with high correlation (
), the method still achieved high power with sufficiently transferable datasets. FWER decreased significantly with increasing
, demonstrating the ability of TransHDM to control false positive. Although FWER slightly increased at
, it remained at a low level (FWER < 0.1), confirming robust error rate control. Figure 2(b–d) and Fig. 3(b–d) present the root mean squared error (rMSE), relative bias (rBias), and standard deviation (SD) of average direct effects (DE), indirect effects (IDE), and mediator proportions (MP) under different settings. Compared to HDM, TransHDM significantly reduced MSE, rBias, and SD for DE, IDE, and MP. Errors further decreased and approached zero with increasing
, demonstrating TransHDM’s robustness across varying mediator dimensions (
) and correlations (
).
In further simulation studies, we introduced a heterogeneous design of the covariate variance–covariance matrix to mimic cross-domain covariate distribution shifts commonly encountered in practical analyses (Fig. 3). The results showed that under heterogeneous design, mediation analysis exhibited higher FWER compared to the homogeneous design, suggesting that differences in covariate distributions pose challenges to model accuracy. Nevertheless, as previously noted, with increasing
, TransHDM demonstrated strong robustness across multiple evaluation indicators for effect estimation and mediator identification. Specifically, the FWER was effectively controlled. These findings highlight that TransHDM retains strong adaptability and robustness despite the challenges introduced by covariate heterogeneity.
In addition, we conducted simulations to evaluate the impact of retaining different numbers of mediators in the SIS step (
250, 500, 750, and 1000, where
1000 represents no SIS) on TransHDM and HDM under the settings of
and
(Table S5). By comparing the results across different
thresholds, we observed that as
increases, power gradually improves, reflecting the inclusion of more potential mediators, thereby increasing detection efficiency. However, this also leads to a significant increase in FWER, possibly because more noise is retained, which elevates false positives in multiple testing. Nevertheless, across all s thresholds, transfer learning consistently enhances performance on the target data, and transfer from sufficient sources can effectively reduce FWER to acceptable levels. In terms of computational efficiency, the simulation time without SIS increases substantially. For
and
, the runtime of the algorithm without SIS is 10 to 15 times longer than that with SIS (Table S6).
Application for Alzheimer’s disease
Characteristics of participants
First, we conducted a comprehensive descriptive analysis of the similarities in baseline demographic and lipidomic characteristics between the target population (African American) and the source population (non-Hispanic White) to assess the rationale and feasibility of applying transfer learning in the context of this study.
Supplementary Table S7 summarizes baseline characteristics and AD severity–related outcomes across APOE ε4 carrier groups. The African American population (median age: 73.2 years; 37.7% male) and non-Hispanic White population (median age: 74.0 years; 56.2% male) both exhibited distinct AD pathological features in APOE ε4 carriers compared to non-carriers, with higher ADAS-Cog 13 scores, elevated t-tau levels, larger WMH volumes, and lower ABeta42 levels. Notably, non-APOE ε4 African Americans displayed milder AD-related pathology than their non-Hispanic White counterparts, with lower ADAS-Cog 13 scores, reduced t-tau levels, and higher ABeta42 levels. Correlation analysis revealed moderate associations of cognitive function (ADAS-Cog 13) with AD biomarkers (t-tau and ABeta42) (Pearson correlation: −0.40 to 0.44), while WMH showed weaker correlations with the other three indicators (Pearson correlation: −0.13 to 0.03) (Fig. S2).
In addition, lipid correlation patterns were highly consistent between African Americans and non-Hispanic Whites (Fig. S3). Significant positive correlations were observed within and between lipid classes, such as lysophosphatidylcholine (LPC) and lysophosphatidylethanolamine (LPE). Principal component analysis revealed substantial overlap in lipidomic profiles between African American and non-Hispanic White, indicating high similarity (Fig. S4).
Therefore, these findings provide direct evidence that the target and source populations are highly similar in terms of baseline demographic and lipidomic features, supporting both the plausibility of shared mediating factors between the two populations and the rationale for applying transfer learning from the source population to target population.
Identification of significant mediators
Figure 4(a) illustrates the number of lipid classes identified as mediators of the APOE ε4–AD association in both African American and non-Hispanic White. Transfer learning enhanced mediator identification in African Americans by incorporating useful information from external races, revealing seven additional lipid classes consistently identified across all outcomes: dihydroceramide (dhCer), alkenyl-phosphatidylcholine (PC (P)), phosphatidylethanolamine (PE), lysophosphatidylethanolamine (LPE (P)), diacylglycerol (DG), acylcarnitine (AC), and hydroxylated acylcarnitine (AC-OH). Results of lipid identification remain generally consistent between analyses of different outcomes. In non-Hispanic Whites, ~40 lipid classes were identified per outcome (ADAS-Cog 13: 38; t-tau: 41; ABeta42: 39; WMH: 38), with 32 lipids consistently significant across all outcomes. This consistency validates the robustness of lipid mediation effects across diverse AD severity measures.
Figure 4.
Lipid mediator identification results. (a) The upset plot shows the number of lipid categories significantly mediating the association between APOE ε4 and four AD severity measures in both AA and NHW before and after transfer learning. The list on the right shows the lipid mediators that were significant in all four outcome analyses in AA, AA-T, and NHW. (b) Distribution profiles of significant lipid molecules. All lipid molecules are ordered sequentially, and the figure highlights those exhibiting significant mediation effects. All lipid molecules are arranged according to the order in Supplementary Tables S8–S15, and different colors indicate distinct lipid categories. Molecules with a significant mediation effect are filled with their category’s color, whereas nonsignificant molecules are shown in gray. The number of colored blocks within each category corresponds to the number of significant mediators in that category. Statistical significance was determined by the Joint Significance Test (Step 3 of Algorithm 2, TransHDM). The SIS thresholds were set at
= 100 for HDM on AA population,
= 150 for TransHDM on AA population, and
= 400 for HDM on NHW population. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity.
We performed correlation analysis on identified lipids in African American and non-Hispanic Whites. Strong inter-lipid relationships, observed in African American result after transfer learning, such as PE with phosphatidylcholine (PC), PE with phosphatidylinositol (PI), PE with DG, and DG with triacylglycerol (TG), were also found in non-Hispanic Whites, reflecting the close biological or metabolic interrelationships between these lipids (Figs. S5–S6). These high correlations support the biological characteristic of lipids operating in a modular fashion within metabolic networks. When analyzing the African American population alone, the smaller sample size may have limited the identification of all lipids with synergistic functions. Therefore, the application of transfer learning could help identify lipid groups that are functionally related, allowing for the inclusion of lipid mediators that focus on key functional modules.
Figure 4(b) presents the lipid profiles significantly mediating the relationship between APOE ε4 and the four AD severity indicators. Analysis of lipid profiles corresponding to each evaluation metric revealed that non-Hispanic Whites exhibited a greater number of significant lipid molecules within each lipid class compared to African Americans. In the latter group, target population analysis identified fewer lipid categories and fewer significant lipid molecules per class. However, the application of transfer learning not only enabled the identification of additional lipid classes but also found more significant lipid molecules within previously identified classes, such as Cer(d), PC, and TG.
Comparative analysis of lipid profiles across the four outcome measures demonstrated consistency in both the overall lipid classes and the proportion of significant lipid molecules within each class. Notably, certain lipids exhibited population-specific mediation effects. For example, alkyl-phosphatidylethanolamine (PE (O)) was uniquely identified in non-Hispanic Whites, whereas sphingosine-1-phosphate (S1P) was specific to African Americans. Additionally, the significance of some lipids differed in various outcomes, potentially reflecting the distinct pathological aspects of AD captured by different evaluation metrics. These variations may also be influenced by factors such as sample size limitations and the inherent complexity of lipid–disease interactions. The estimated results of the mediation effects of specific lipids are shown in Tables S8–S15.
Lipid metabolomic pathway analysis
To explore the cross-ethnic mediation mechanisms of lipid metabolic pathways in the association between APOE4 ε4 and AD pathology, we performed pathway enrichment analysis on lipid classes demonstrating significant mediation effects in both African American and non-Hispanic White populations (Fig. 5, Table S16). In non-Hispanic Whites, lipid species were significantly enriched in four metabolic pathways: glycerolipid metabolism, glycerophospholipid metabolism, sphingolipid metabolism, and ether lipid metabolism. In African Americans, independent analysis revealed three distinct mediation patterns: (1) glycerophospholipid, sphingolipid, and ether lipid metabolism pathways significantly mediated the APOE–ADAS-Cog 13 association; (2) glycerophospholipid metabolism specifically mediated the association of APOE with t-tau and ABeta42 biomarkers; and (3) glycerophospholipid and sphingolipid metabolism mediated the APOE–WMH relationship (all false discovery rate (FDR) < 0.05). Notably, after integrating cross-ethnic data using transfer learning algorithms, we observed that glycerolipid, glycerophospholipid, and sphingolipid metabolism pathways concurrently mediated APOE associations with t-tau, ABeta42, and WMH in African Americans. Furthermore, all four metabolic pathways identified in non-Hispanic Whites consistently mediated APOE–ADAS-Cog 13 associations in African Americans.
Figure 5.
KEGG pathway enrichment analysis results. (a) The bar charts display the top 10 KEGG enrichment results for the lipids identified in AA, AA-T, and NHW across four outcomes. The bar charts marked with “***” indicate significant results with corrected FDR P-values. (b) The heatmap illustrates the significance of lipid species enriched in the four pathways. Blocks indicate lipids with significant mediating effects in their respective analysis groups, whereas empty cells indicate that the lipid category does not show a significant effect in that group. (c) The action network of lipids enriched in different pathways. Abbreviations: AA, African American; AA-T, African American with transfer learning; NHW, non-Hispanic White; ADAS-Cog 13, the 13-item cognitive subscale of the Alzheimer’s Disease Assessment Scale; Abeta42, amyloid-beta (1–42); t-tau, total tau; WMH, white matter hyperintensity; FDR, false discovery rate.
Figure 5(b) shows the identification of lipids enriched in the four aforementioned significant metabolic pathways across different outcomes and populations. In African American, certain lipids were not consistently identified across all outcome pathways. However, transfer learning significantly improved the identification of lipids with potential mediating effects. For example, DG, LPC, and PI in glycerophospholipid metabolism, Sphingomyelin (SM) in sphingolipid metabolism, DG and free fatty acid in glycerolipid metabolism, and lyso-alkyl-phosphatidylcholine (LPC(O)) in ether lipid metabolism were identified across all four outcomes, substantially improving the detection of metabolic pathways.
Although the significant metabolic pathways were largely similar between populations, subtle differences in specific enriched lipid categories were observed. For instance, sphingosine (Sph) in the sphingolipid metabolism pathway was uniquely identified in African Americans, while alkyl-phosphatidylethanolamine (PE(O)) in the ether lipid metabolism pathway was specific to non-Hispanic Whites, both before and after transfer learning. These results suggest that transfer learning, by incorporating external data and enhancing sample diversity, improves the recognition of complex and population-specific variations in lipid metabolism.
Sensitivity analysis
To further evaluate the robustness of the transfer learning approach in lipid identification, we conducted the following analyses. First, we constructed predictive models to assess whether the lipid mediators identified by TransHDM could improve the prediction of AD severity in the AA population. As shown in Table S17, the model incorporating lipid mediators identified by TransHDM (Model 2) achieved the best predictive performance. In contrast, Model 1, which included lipid mediators identified by HDM, showed inferior performance. Model 3, based on lipid mediators from the NHW population, showed similarly poor performance to the Model 4 with no mediators included, highlighting that TransHDM effectively captures population-specific characteristics of the target group. Overall, the lipid mediators identified through transfer learning captured substantial variation in AD severity, indirectly supporting the mediation analysis conclusion that these lipids may play important roles in AD pathology.
Second, to assess the stability of the TransHDM, we removed the lipid molecules initially identified by TransHDM from the AA population. No newly identified lipids were identified with a frequency >80%, suggesting that no new stable lipid mediators were found after removing the previously identified ones. However, when excluding the lipid mediators identified from HDM and reapplying TransHDM, new lipid molecules were rediscovered, and the lipids identified were largely consistent with those from the original analysis (Tables S18–S21). These results further confirm the stability and consistency of the transfer learning approach, indicating that the exclusion of previously identified lipids did not lead to divergent or inconsistent findings.
Additionally, to examine the impact of transferable source data size on mediator identification, we compared the number of lipid mediators detected under four settings: AA (without transfer learning), AAT1 (using one-third of random transferable samples), AAT2 (two-thirds), and AAT3 (the full set). As shown in Fig. S7 and Tables S22–S25, the number of identified mediators increases from AA to AAT3, suggesting that more transferable data enhance mediator detection. Moreover, most lipid mediators identified under different settings show substantial overlap, demonstrating the robustness of mediator selection.
We examined the impact of varying
retained by SIS on the results (Fig. S8). When
is low, stringent screening that missed many potential mediators may result in elevated Type II error rates. As s increases from 50 to 100 or 150, including more true mediators boosts detection ability, leading to more significant lipids identified. The lipids detected at lower
exhibit overlap with those identified at higher
, underscoring the robustness of the results across varying
. However, when s further increases to 200, the inclusion of noise lipids may increase model complexity, reducing the identification of lipid mediators. Consequently, retaining 100 to 150 variables emerges as a relatively optimal balance, as adopted in our primary analysis.
Discussion
This study proposes a novel transfer learning–based high-dimensional mediation analysis framework designed to address the challenges of mediator selection and effect estimation in data-scarce settings. By integrating external transferable datasets, this approach significantly enhances the accuracy, robustness, and false discovery rate control in mediation analysis. Compared to traditional methods, TransHDM effectively leverages external data to overcome for small sample sizes limitations, thereby improving statistical power in high-dimensional data analysis. Applied to the ADNI cohort, this framework identified four significant lipid metabolism pathways—glycerolipid metabolism, glycerophospholipid metabolism, sphingolipid metabolism, and ether lipid metabolism—that significantly mediate the APOE ε4–AD association in both non-Hispanic White and African American populations. These pathways, critical for lipid homeostasis and neuroinflammation regulation, reveal shared metabolic mechanisms underlying AD across racial groups.
TransHDM integrates transfer learning with high-dimensional mediation analysis, providing a robust solution for small sample scenarios. The successful application not only advances AD research in African American populations but also offers a promising approach for studying other data-scarce groups, such as rare disease cohorts, specific subpopulations, and early-stage disease patients [35]. Traditional methods often struggle to identify potential biomarkers or elucidate pathological mechanisms in these contexts due to insufficient sample sizes. By integrating external data, transfer learning can mitigate these limitations and enhance the exploration of complex disease mechanisms, ultimately contributing to the development of targeted interventions tailored to diverse populations.
Given these advantages, the proposed framework demonstrates both methodological innovation and excellent practical applicability in disease mechanism research and personalized treatment strategy development. Our findings reveal consistent lipid metabolism pathways mediating the APOE ε4–AD relationship in both African American and non-Hispanic White populations. The four pathways play crucial roles in lipid homeostasis, cell membrane integrity, and neuroinflammation regulation, suggesting that despite differences in genetic background, environmental exposure, and lifestyle, the role of lipid metabolism may universally contribute to AD pathology. Our findings indicate that lipid metabolism-based AD biomarkers and therapeutic strategies developed in White populations may also be applicable to Black populations. This cross-racial applicability not only accelerates the translational application of AD research but also offers new approaches to address health disparities arising from insufficient research and delayed development of effective biomarkers and treatments.
Our findings are supported by previous studies demonstrating minimal differences in serum lipid levels between African American (AA) and non-Hispanic White (NHW) populations [36]. This consistency further validates the shared lipid metabolism pathways across racial groups. Numerous lipidomics studies in NHW populations have identified significant differential expression of glycerophospholipids, glycerolipids, and sphingolipids in AD patients [37–40]. Similarly, in African Americans, glycerophospholipids and sphingolipids have shown prospective associations with mild cognitive impairment and dementia [41, 42]. Collectively, these studies underscore the consistent role of lipid metabolism in AD across racial groups, strongly supporting our findings.
In our analysis of the African American population, ether lipid metabolism significantly mediated the association between APOE ε4 and ADAS-Cog 13. However, no significant mediation effects were observed for t-tau, ABeta42, or WMH outcomes, even after applying transfer learning. This discrepancy may be due to the smaller sample sizes for these three outcomes, limiting statistical power and the ability to fully capture ether lipid metabolism’s mediating effects. Growing evidence suggests that inflammation plays an important role in AD pathogenesis, with ether lipid metabolism closely associated with inflammatory processes [43]. Recent studies indicated a strong association between ether lipids and ferroptosis, an iron-dependent form of cell death characterized by lipid peroxidation and oxidative stress accumulation, which is implicated in neurodegenerative diseases such as AD [44–46]. Ether lipid metabolism may regulate AD progression through inflammatory and immune modulation mechanisms [47]. Future research should expand sample size to further explore the critical role of ether lipids in mediating these biological processes.
In addition, the outcomes used in our study capture distinct dimensions of AD severity. WMH volume primarily reflects cerebrovascular lesions or small vessel diseases, while ADAS-Cog 13 score, t-tau levels, and ABeta42 levels more directly assess AD-specific pathological features, including cognitive function, tau protein pathology, and amyloid deposition [48]. Despite the limited correlation among these four indicators, the lipid molecular types identified through mediation analysis were largely consistent. This suggests that these lipid molecules may play an important role in multiple pathophysiological processes and different stages of AD. Specifically, APOE ε4 may influence not only amyloid metabolism but also cerebrovascular health and neuroinflammation through these lipid-mediated mechanisms.
A notable limitation of this study is its focus on serum lipids rather than CSF lipids. The relationship between peripheral metabolites and APOE ε4–related pathological processes in the central nervous system remains poorly understood. While changes in peripheral metabolites may serve as potential biomarkers for AD risk, their clinical utility and accuracy require further validation. Additionally, in this study, the SIS threshold was manually determined in advance, which may limit the flexibility of variable selection. Future research could potentially optimize the selection of threshold in SIS using methods such as cross-validation, thereby improving the flexibility and accuracy of variable selection in high-dimensional data. Furthermore, to validate the universality of the key lipid metabolism pathways identified here, future studies should expand the sample size of African American populations and integrate multi-omics data, such as genomics and proteomics. These pathways should also be explored further for their potential applications in early AD diagnosis, risk prediction, and targeted therapy.
Conclusion
In conclusion, this study leverages transfer learning to reveal consistent lipid metabolism pathways in both African American and non-Hispanic White populations, providing critical scientific evidence for cross-racial AD pathology research and therapeutic strategy development. The proposed TransHDM framework offers a robust tool for addressing the challenges of high-dimensional data analysis in small sample populations and specific heterogeneous groups, such as rare disease cohorts. These findings not only advance our understanding of AD mechanisms but also pave the way for more inclusive and precise biomedical research.
Key Points
TransHDM is the first transfer learning–based high-dimensional mediation framework specifically designed to overcome the limitations of traditional methods in small-sample studies. By leveraging large-scale external data and employing transfer regularization to mitigate population heterogeneity, it significantly enhances the detection power for true mediators while rigorously controlling family-wise error rates (demonstrated in simulation studies).
Applied to the ADNI cohort, TransHDM identified four APOE ε4–mediated lipid metabolism pathways (glycerophospholipid, glycerolipid, sphingolipid, and ether lipid metabolism) in African American populations. These pathways—previously undetectable with conventional methods due to limited sample sizes—provide novel insights into Alzheimer’s disease mechanisms and facilitate biomarker discovery for underrepresented groups.
These lipid pathways mediating APOE ε4 effects in African Americans were also validated in White populations, suggesting evolutionarily conserved mechanisms of lipid homeostasis and neuroinflammation across ethnicities. This cross-racial consistency indicates that lipid-targeted therapies or biomarkers developed in majority populations could be directly applicable to African American individuals, potentially reducing health disparities.
Supplementary Material
Contributor Information
Lulu Pan, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.
Yahang Liu, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.
Chen Huang, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.
Ruilang Lin, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.
Yongfu Yu, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China; Shanghai Key Laboratory of Gene Editing and Cell Therapy for Rare Diseases, Fudan University, 83 Fen Yang Road, Xuhui District, Shanghai 200031, China.
Guoyou Qin, Department of Biostatistics, Key Laboratory of Public Health Safety of Ministry of Education, NHC Key Laboratory for Health Technology Assessment, School of Public Health, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China; Shanghai Institute of Infectious Disease and Biosecurity, Fudan University, 130 Dong’an Road, Xuhui District, Shanghai 200032, China.
Author contributions
L.L.P., Y.H.L., and C.H. conceived the idea and contributed to statistical analysis, interpretation of data, and the draft of the manuscript. R.L.L. contributed to the analysis of the data and revised the manuscript. G.Y.Q. and Y.F.Y. contributed to the conception of the study, overall supervision, and final editing of the manuscript. All authors read and approved the final manuscript.
Conflict of interest
None declared.
Funding
This work was supported by the National Natural Science Foundation of China (No. 82273730 to Y.F.Y., 82173612 to G.Y.Q.), Shanghai Rising-Star Program (21QA1401300 to Y.F.Y.), Shanghai Municipal Natural Science Foundation (22ZR1414900 to Y.F.Y.), the Three-Year Public Health Action Plan of Shanghai (GWVI-11.2-XD10 and GWVI-11.1-01 to Y.F.Y.), Shanghai Talent Programs (BJKJ2024050 to Y.F.Y.), and Shanghai Municipal Science and Technology Major Project (ZD2021CY001 to G.Y.Q.).
Data availability
All data used in the analyses reported here are available in the ADNI data repository (adni.loni.usc.edu).
Standard protocol approvals, registrations, and patient consents
Written informed consent was obtained at the time of enrollment for imaging and sample collection, and protocols of consent forms were approved by the Institutional Review Board at each participating site.
Code availability
The source code supporting this work is publicly available at https://github.com/PanLululu/TransHDM.
References
- 1. Scheltens P, De Strooper B, Kivipelto M. et al. Alzheimer’s disease. Lancet 2021;397:1577–90. 10.1016/S0140-6736(20)32205-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Matthews KA, Xu W, Gaglioti AH. et al. Racial and ethnic estimates of Alzheimer’s disease and related dementias in the United States (2015-2060) in adults aged >/=65 years. Alzheimers Dement 2019;15:17–24. 10.1016/j.jalz.2018.06.3063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. 2021 Alzheimer’s disease facts and figures. Alzheimers Dement 2021;17:327–406. 10.1002/alz.12328 [DOI] [PubMed] [Google Scholar]
- 4. Farrer LA, Cupples LA, Haines JL. et al. Effects of age, sex, and ethnicity on the association between apolipoprotein E genotype and Alzheimer disease. A meta-analysis. APOE and Alzheimer Disease Meta Analysis Consortium. JAMA 1997;278:1349–56. 10.1001/jama.1997.03550160069041 [DOI] [PubMed] [Google Scholar]
- 5. Torres-Perez E, Ledesma M, Garcia-Sobreviela MP. et al. Apolipoprotein E4 association with metabolic syndrome depends on body fatness. Atherosclerosis 2016;245:35–42. 10.1016/j.atherosclerosis.2015.11.029 [DOI] [PubMed] [Google Scholar]
- 6. Yang LG, March ZM, Stephenson RA. et al. Apolipoprotein E in lipid metabolism and neurodegenerative disease. Trends Endocrinol Metab 2023;34:430–45. 10.1016/j.tem.2023.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Polis B, Samson AO. A new perspective on Alzheimer’s disease as a brain expression of a complex metabolic disorder. In: Wisniewski T (ed.), Alzheimer’s Disease. Brisbane (AU): Codon Publications, 2019. [PubMed] [Google Scholar]
- 8. Stepler KE, Robinson RAS. The potential of ‘omics to link lipid metabolism and genetic and comorbidity risk factors of Alzheimer’s disease in African Americans. Adv Exp Med Biol 2019;1118:1–28. 10.1007/978-3-030-05542-4_1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Power MC, Rawlings A, Sharrett AR. et al. Association of midlife lipids with 20-year cognitive change: a cohort study. Alzheimers Dement 2018;14:167–77. 10.1016/j.jalz.2017.07.757 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Bernath MM, Bhattacharyya S, Nho K. et al. Serum triglycerides in Alzheimer disease: relation to neuroimaging and CSF biomarkers. Neurology 2020;94:e2088–98. 10.1212/WNL.0000000000009436 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Wang T, Huynh K, Giles C. et al. APOE epsilon2 resilience for Alzheimer’s disease is mediated by plasma lipid species: analysis of three independent cohort studies. Alzheimers Dement 2022;18:2151–66. 10.1002/alz.12538 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Gianattasio KZ, Bennett EE, Wei J. et al. Generalizability of findings from a clinical sample to a community-based sample: a comparison of ADNI and ARIC. Alzheimers Dement 2021;17:1265–76. 10.1002/alz.12293 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Franzen S, Smith JE, van den Berg E. et al. Diversity in Alzheimer’s disease drug trials: the importance of eligibility criteria. Alzheimers Dement 2022;18:810–23. 10.1002/alz.12433 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Babulal GM, Quiroz YT, Albensi BC. et al. Perspectives on ethnic and racial disparities in Alzheimer’s disease and related dementias: update and areas of immediate need. Alzheimers Dement 2019;15:292–312. 10.1016/j.jalz.2018.09.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Zhang H, Zheng Y, Zhang Z. et al. Estimating and testing high-dimensional mediation effects in epigenetic studies. Bioinformatics 2016;32:3150–4. 10.1093/bioinformatics/btw351 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Perera C, Zhang H, Zheng Y. et al. HIMA2: high-dimensional mediation analysis and its application in epigenome-wide DNA methylation data. BMC Bioinformatics 2022;23:296. 10.1186/s12859-022-04748-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Torrey L, Shavlik J. Transfer Learning. In: Olivas ES, Guerrero JD, Martinez-Sober M, Magdalena-Benedito JR, Serrano López AJ (eds.), Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques. Hershey (PA): IGI global, 2010. 242–64. [Google Scholar]
- 18. Davila A, Colan J, Hasegawa Y. Comparison of fine-tuning strategies for transfer learning in medical image classification. Image Vis Comput 2024;146:105012. 10.1016/j.imavis.2024.105012 [DOI] [Google Scholar]
- 19. Dieckhaus H, Brocidiacono M, Randolph NZ. et al. Transfer learning to leverage larger datasets for improved prediction of protein stability changes. Proc Natl Acad Sci 2024;121:e2314853121. 10.1073/pnas.2314853121 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Mahmud T, Barua K, Habiba SU. et al. An explainable ai paradigm for Alzheimer’s diagnosis using deep transfer learning. Diagnostics 2024;14:345. 10.3390/diagnostics14030345 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Li S, Cai TT, Li H. Transfer learning for high-dimensional linear regression: prediction, estimation and minimax optimality. J R Stat Soc Ser B Stat Methodol 2022;84:149–73. 10.1111/rssb.12479 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Tian Y, Feng Y. Transfer learning under high-dimensional generalized linear models. J Am Stat Assoc 2023;118:2684–97. 10.1080/01621459.2022.2071278 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Li S, Zhang L, Cai TT. et al. Estimation and inference for high-dimensional generalized linear models with knowledge transfer. J Am Stat Assoc 2024;119:1274–85. 10.1080/01621459.2023.2184373 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Van de Geer S, Bühlmann P, Ya R. et al. On asymptotically optimal confidence regions and tests for high-dimensional models. Ann Statist 2014;42:1166–202. 10.1214/14-AOS1221 [DOI] [Google Scholar]
- 25. Meinshausen N, Bühlmann P. High-dimensional graphs and variable selection with the lasso. Ann Statist 2006;34:1436–62. 10.1214/009053606000000281 [DOI] [Google Scholar]
- 26. Zhang Y, Zhu Z. Transfer learning for high-dimensional quantile regression via convolution smoothing. Stat Sin 2025;35:939–58. 10.5705/ss.202022.0396 [DOI] [Google Scholar]
- 27. Fan J, Lv J. Sure independence screening for ultrahigh dimensional feature space. J R Stat Soc Ser B Stat Methodol 2008;70:849–911. 10.1111/j.1467-9868.2008.00674.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Dai JY, Stanford JL, LeBlanc M. A multiple-testing procedure for high-dimensional mediation hypotheses. J Am Stat Assoc 2022;117:198–213. 10.1080/01621459.2020.1765785 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Weiner MW, Veitch DP, Aisen PS. et al. Recent publications from the Alzheimer’s Disease Neuroimaging Initiative: reviewing progress toward improved AD clinical trials. Alzheimers Dement 2017;13:e1–85. 10.1016/j.jalz.2016.11.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Saykin AJ, Shen L, Yao X. et al. Genetic studies of quantitative MCI and AD phenotypes in ADNI: progress, opportunities, and plans. Alzheimers Dement 2015;11:792–814. 10.1016/j.jalz.2015.05.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Huynh K, Barlow CK, Jayawardana KS. et al. High-throughput plasma lipidomics: detailed mapping of the associations with cardiometabolic risk factors. Cell Chem Biol 2019;26:71–84.e4. 10.1016/j.chembiol.2018.10.008 [DOI] [PubMed] [Google Scholar]
- 32. DeCarli C, Fletcher E, Ramey V. et al. Anatomical mapping of white matter hyperintensities (WMH): exploring the relationships between periventricular WMH, deep WMH, and total WMH burden. Stroke 2005;36:50–5. 10.1161/01.STR.0000150668.58689.f2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Schwarz C, Fletcher E, DeCarli C. et al. Fully-automated white matter hyperintensity detection with anatomical prior knowledge and without FLAIR. Inf Process Med Imaging 2009;5636:239–51. 10.1007/978-3-642-02498-6_20 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Bittner T, Zetterberg H, Teunissen CE. et al. Technical performance of a novel, fully automated electrochemiluminescence immunoassay for the quantitation of beta-amyloid (1-42) in human cerebrospinal fluid. Alzheimers Dement 2016;12:517–26. 10.1016/j.jalz.2015.09.009 [DOI] [PubMed] [Google Scholar]
- 35. Johansson A, Andreassen OA, Brunak S. et al. Precision medicine in complex diseases-molecular subgrouping for improved prediction and treatment stratification. J Intern Med 2023;294:378–96. 10.1111/joim.13640 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Carnethon MR, Pu J, Howard G. et al. Cardiovascular health in African Americans: a scientific statement from the American Heart Association. Circulation 2017;136:e393–423. 10.1161/CIR.0000000000000534 [DOI] [PubMed] [Google Scholar]
- 37. Whiley L, Sen A, Heaton J. et al. Evidence of altered phosphatidylcholine metabolism in Alzheimer’s disease. Neurobiol Aging 2014;35:271–8. 10.1016/j.neurobiolaging.2013.08.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Klavins K, Koal T, Dallmann G. et al. The ratio of phosphatidylcholines to lysophosphatidylcholines in plasma differentiates healthy controls from patients with Alzheimer’s disease and mild cognitive impairment. Alzheimers Dement (Amst) 2015;1:295–302. 10.1016/j.dadm.2015.05.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Li D, Misialek JR, Boerwinkle E. et al. Plasma phospholipids and prevalence of mild cognitive impairment and/or dementia in the ARIC Neurocognitive Study (ARIC-NCS). Alzheimers Dement (Amst) 2016;3:73–82. 10.1016/j.dadm.2016.02.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Mielke MM, Bandaru VV, Haughey NJ. et al. Serum sphingomyelins and ceramides are early predictors of memory impairment. Neurobiol Aging 2010;31:17–24. 10.1016/j.neurobiolaging.2008.03.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Li D, Misialek JR, Boerwinkle E. et al. Prospective associations of plasma phospholipids and mild cognitive impairment/dementia among African Americans in the ARIC Neurocognitive Study. Alzheimers Dement (Amst) 2017;6:1–10. 10.1016/j.dadm.2016.09.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Khan MJ, Chung NA, Hansen S. et al. Targeted lipidomics to measure phospholipids and sphingomyelins in plasma: a pilot study to understand the impact of race/ethnicity in Alzheimer’s disease. Anal Chem 2022;94:4165–74. 10.1021/acs.analchem.1c03821 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Jofre-Monseny L, Minihane AM, Rimbach G. Impact of apoE genotype on oxidative stress, inflammation and disease risk. Mol Nutr Food Res 2008;52:131–45. 10.1002/mnfr.200700322 [DOI] [PubMed] [Google Scholar]
- 44. Zou Y, Henry WS, Ricq EL. et al. Plasticity of ether lipids promotes ferroptosis susceptibility and evasion. Nature 2020;585:603–8. 10.1038/s41586-020-2732-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Kim JW, Lee JY, Oh M. et al. An integrated view of lipid metabolism in ferroptosis revisited via lipidomic analysis. Exp Mol Med 2023;55:1620–31. 10.1038/s12276-023-01077-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46. Bao WD, Pang P, Zhou XT. et al. Loss of ferroportin induces memory impairment by promoting ferroptosis in Alzheimer’s disease. Cell Death Differ 2021;28:1548–62. 10.1038/s41418-020-00685-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Chen X, Kang R, Kroemer G. et al. Ferroptosis in infection, inflammation, and immunity. J Exp Med 2021;218:e20210518. 10.1084/jem.20210518 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Jack CR Jr, Bennett DA, Blennow K. et al. NIA-AA Research Framework: toward a biological definition of Alzheimer’s disease. Alzheimers Dement 2018;14:535–62. 10.1016/j.jalz.2018.02.018 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data used in the analyses reported here are available in the ADNI data repository (adni.loni.usc.edu).






















































