Skip to main content
BMC Medicine logoLink to BMC Medicine
. 2025 Dec 3;24:12. doi: 10.1186/s12916-025-04564-3

KoMethylNet: a novel epigenetic clock based on neural network analysis of DNA methylation data and epigenetic age acceleration in a Korean population

Dabin Yun 1, Kwangyeon Oh 2, Yujin Kim 3, Yong Min Ahn 4,5, Hemang M Parikh 6, Xiaoxi Meng 2, Zhaoming Wang 6,7, Nan Song 1,✉
PMCID: PMC12781313  PMID: 41339892

Abstract

Background

Epigenetic clocks have been extensively investigated in individuals of European ancestry and may be suboptimal in East Asians. We developed a novel epigenetic clock (KoMethylNet) using neural network analysis of DNA methylation (DNAm) data from the Korean population to predict chronological ages.

Methods

DNAm data (367,785 CpG sites) from 2,747 participants (Infinium Human Methylation 450 K BeadChip: N = 397; Infinium MethylationEPIC BeadChip: N = 2,350) in the Korean Genome and Epidemiology Study (KoGES) were used to train the neural network on chronological ages. SHapley Additive exPlanation analysis was used to select the optimal number of CpG sites. KoMethylNet-epigenetic age acceleration (EAA)-phenotype analysis was conducted with linear regression, to identify aging-related phenotypes in the Korean population.

Results

KoMethylNet, which uses 300 CpG sites, achieved a mean absolute error (MAE) of 2.82 years, a mean squared error (MSE) of 12.68 years, and a Pearson’s correlation coefficient (R) of 0.90 with chronological age. In the external validation using healthy Korean individuals, KoMethylNet achieved the highest performance (MAE = 2.74, MSE = 12.29, R = 0.94). Seven phenotypes, including diabetes-related traits (diabetes, HbA1c, and urine glucose), were positively associated with KoMethylNet-EAA.

Conclusions

We developed a neural network-based DNAm aging clock using Korean population data that enables precise age prediction and offers potential opportunities for advancing aging-related research in Korea.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12916-025-04564-3.

Keywords: Deep learning, DNA methylation aging clock, Epigenetic age acceleration

Background

Aging in humans is a physiological process that leads to various changes over time [1]. These changes increase the risk of age-related diseases, including Alzheimer’s disease [2], inflammatory bowel disease [3], lung disease [4], cardiovascular disease [5], and cancer [6]. One promising approach to understanding aging and age-related diseases is epigenetic epidemiology research [7–9], specifically through DNA methylation (DNAm) on CpG sites [10–12]. Consequently, epigenetic clocks have been developed to study aging and age-related diseases, focusing on biological rather than chronological age [13–15].

The first epigenetic clock was developed by Horvath using DNAm data from multiple tissues to estimate the DNAm age of tissues and cells [16]. Hannum et al. proposed another epigenetic clock using elastic net regression on DNAm sites from the blood DNA methylation profile and was validated with excellent accuracy [17]. Both are first-generation epigenetic clocks designed to predict chronological age based on DNA methylation levels. Second-generation epigenetic clocks have been developed to predict both chronological age and phenotypic measures, including various biomarkers. For instance, Levine et al. developed ‘PhenoAge’ using 513 CpG sites in a two-step process [18]. First, the composite phenotypic age was estimated using nine clinical markers, along with the chronological age [18]. Second, the CpG sites were modeled and selected to predict the phenotypic age [18]. DNAm-PhenoAge outperformed earlier epigenetic clocks in predicting various age-related health outcomes, including all-cause mortality, health span, and physical function [18]. Likewise, Lu et al. introduced ‘GrimAge’ using a similar two-step approach [19]. Initially, they developed DNAm-based surrogates for physiological risk factors, encompassing plasma proteins and stress indicators [19]. Additionally, they incorporated a DNAm estimator of smoking pack-years to capture smoking history [19]. These biomarkers, along with selected CpG sites, were then integrated to create DNAm GrimAge, a composite biomarker that quantifies biological age in years [19]. In addition, Zhang et al. developed a DNAm aging clock using 13,661 samples of blood and saliva, which includes 600 Chinese motor neuron disease patients [20]. Their DNAm aging clock outperformed Horvath aging clock and Hannum aging clock in chronological age prediction [20].

While the first- and second-generation epigenetic clocks were developed based on elastic net regression, novel epigenetic clocks have emerged in line with the ongoing trend of applying deep learning techniques to the field of biomedical research for better prediction accuracy [21, 22]. Levy et al. presented ‘MethylNet’, a modular deep learning framework including variational auto-encoders, that estimates biological age from DNAm data [23]. They used 34 datasets from 9,500 samples, and the main dataset for predicting chronological age was GSE87571 [24], which consisted of blood DNAm data from healthy participants aged 15 to 95. It achieved a mean absolute error (MAE) of 3.00 years, demonstrating improved accuracy compared to Hannum’s aging clock (MAE = 5.6 years) and Horvath’s aging clock (MAE = 3.9 years) [23]. They also applied Shapley Feature Attribution methods to quantify the contributions of individual CpG sites in age prediction [23]. Additionally, Galkin et al. developed ‘DeepMAge’, a neural network-based epigenetic clock using 6,411 samples from 32 studies obtained from blood samples on Infinium Human Methylation 450 K and 27 K BeadChip platforms [25]. It outperformed previous clocks with an MAE of 2.77 years and predicts older age for individuals with diseases such as inflammatory bowel disease, ovarian cancer, and multiple sclerosis [25]. Similarly, Camillo et al. introduced 'AltumAge’, a neural network-based pan-tissue epigenetic clock trained on over 14,000 individuals [26]. Using 20,318 CpG sites and adversarial regularization, 'AltumAge' achieved an MAE of 2.15 years, surpassing Horvath's clock [26]. Their model encompassed CpG sites linked to aging-related pathways such as mTOR and AMPK [26].

Epigenetic age acceleration (EAA) refers to the difference between an individual’s chronological and biological ages as estimated using DNAm aging clocks. Positive EAA values indicate accelerated aging, whereas negative values indicate slower aging. Numerous studies have explored the relationships between EAA and clinical, environmental, and socioeconomical factors [27–30]. Factors such as body mass index, cardiovascular disease, and alcohol consumption have been identified as correlates of EAA in these studies.

A major limitation of existing epigenetic clocks is that ethnic and sociodemographic discrepancies between the training and test datasets may contribute to inaccurate prediction results [31, 32]. Particularly in East Asian populations, producing accurate predictions with existing epigenetic clocks is challenging [33, 34]. To overcome this limitation and provide accurate age predictions for the East Asian population, we developed a novel neural network-based DNAm aging clock (KoMethylNet), trained specifically on Korean DNAm data. Model performance was assessed by comparing our model with other epigenetic clocks. Furthermore, we calculated EAA using our model and investigated the relationship between EAA and various clinical, demographic, and socioeconomic factors.

Methods

Study population and data collection

Our study used the Korean Genome and Epidemiology Study (KoGES) database, a large prospective cohort study that investigated genetic and environmental factors in South Korea [35]. Two study groups, the Ansan and Ansung (AA) and Health Examinee (HEXA) studies, were selected from the KoGES population. Next, we utilized 397 (Infinium Human Methylation 450 K BeadChip) and 1,528 participants (Infinium MethylationEPIC BeadChip) who were subgrouped by two types of DNAm arrays from the Ansan and Ansung study (N = 10,030), and 822 participants from the HEXA study (N= 173,357) with available DNAm or genetic data. General characteristics of the KoGES were collected through a combination of questionnaires, physical examinations, and clinical investigations [35]. The questionnaires covered sociodemographic status, lifestyle (diet, smoking, alcohol consumption, and physical activity), reproductive history, psychological stress, social relationships, disease history, and clinical measurements [35]. For the analysis, age at blood sample collection was included as a continuous variable; body mass index (BMI) as a categorical variable; and sex, alcohol consumption status, and smoking status as binary variables. Regarding the number of disease histories, the participants were stratified into three groups based on the number of diagnoses across five diseases: hypertension, dyslipidemia, diabetes mellitus, myocardial infarction, and cardiovascular disease. These groups were defined as having no diagnosis, single diagnosis, or multiple (≥ 2) diagnoses. In addition, education, monthly income, physical activity, and stress were included as categorical variables. For education, participants were categorized into three groups: ‘High school or less’ for those with a high school diploma or less, ‘College, no four-year degree’ for those who attended university without graduating or earned a two-year degree, and ‘College, four-year degree or higher’ for those who graduated from a four-year university or graduate school. Monthly income was categorized into eight levels. Level ‘1’ represented an income of less than 500,000 Korean won (KRW), while level ‘8’ represented 6,000,000 KRW or more. Intermediate levels were defined by intervals of increasing size; 500,000 KRW increments from level ‘2’ to ‘4’, 1,000,000 KRW increments for levels ‘5’ and ‘6’, and 2,000,000 KRW increments for level ‘7’. Additionally, for physical activity, participants were categorized into four groups: ‘Infrequent’ for those exercising 1–2 times per week, ‘Moderate’ for 3–4 times per week, ‘Frequent’ for 5–6 times per week, and ‘Daily’ for 7 times per week. For stress, participants were categorized into three levels: ‘None’ for those reporting no stress, ‘Moderate’ for those reporting occasional stress, and ‘High’ for those reporting frequent stress. Additional clinical (e.g., albumin) and demographic (e.g., daily calories) measurements were included as continuous variables. A total of 63 variables were selected for this study (Additional File 1: Table S6). The study protocol was approved by the Institutional Review Board of the Korean National Institutes of Health (IRB: CBNU-2025-E-0001, CBNU-2025-E-0003).

DNA methylation array

DNA methylation data were generated using the Infinium Human Methylation 450 K BeadChip (N = 397; AA study) and the Infinium MethylationEPIC BeadChip (N = 1,528; AA study and N = 822; HEXA study). The Infinium Human Methylation 450 K BeadChip covered 413,745 CpG sites and the Infinium MethylationEPIC BeadChip covered 853,307 CpG sites. We applied the 'Chip Analysis Methylation Pipeline' (ChAMP) R package [36] to the raw data to produce beta values for the CpG sites. To reduce batch effects, we utilized the champ.SVD function to detect principal components and the champ.runcombat function for correction. These functions are integrated into the ChAMP R package. For quality control, we applied the following filtering criteria: detection P-value > 0.01; bead count < 3 in at least 5% of samples; location on a sex chromosome; location near East Asian-specific SNPs identified by Nordlund et al. [37]; aligned to multiple locations. Additionally, DNA methylation levels were presented as beta values and normalized using beta-mixture quantiles [38]. Furthermore, the beta values for each CpG site were standardized within their respective groups (training, validation, and testing) using a robust scale. This approach involved removing the median and scaling the data based on the quantile range, thus mitigating the influence of outliers on subsequent analyses. We initially used 732,046 CpG sites for the HEXA study, 724,792 CpG sites for the AA study using the Infinium MethylationEPIC BeadChip, and 412,970 CpG sites for the AA study using the Infinium Human Methylation 450 k BeadChip. After quality control and normalization, 367,785 CpG sites that were common across the datasets were selected for further analysis.

Training and test set distribution

We adopted a stratified random sampling method for model development and evaluation to avoid any bias related to the characteristics of study population (e.g., sex). We applied a hold-out approach within each subgroup, randomly dividing the AA and HEXA study participants into an 80% training set and 20% testing set. The training set was further divided into 90% training set and 10% validation set. We aggregated the training, validation, and test datasets within each group (Fig. 1A). For 10-fold cross-validation (CV), we split participants within each group into a 90% training set and a 10% test set for each fold. The datasets were then combined across groups within each fold (Fig. 1B).

Fig. 1.

Fig. 1

Flowchart of dataset split process. A Flowchart of dataset split in hold-out process. B Flowchart of dataset split in 10-fold cross-validation

Model development

For our DNAm aging clock, we employed a neural network comprising an initial input layer, multiple hidden layers, and a final output layer to predict the chronological age. Within the hidden layers, individual neurons sequentially incorporate linear transformations, activation functions, layer normalization, and dropout. Additionally, each beta value of CpG sites was processed with a robust scaler to remove the median and scales according to the quantile range.

We used Optuna [39] to select the best hyperparameter settings for our model. The hyperparameter search space included the following options: the number of hidden units (256, 512, 1,024, 2,048), the number of hidden layers (2, 3, 4, 5, 6), dropout rate (0.1, 0.2, 0.3, 0.4), activation function (CELU, SELU, ReLU), optimizer (AdamW), learning rate (1e-6, 1e-5, 1e-4) and weight decay (1e-6, 1e-5, 1e-4). During the Optuna hyperparameter search, we utilized a subdivided training set (90% for training and 10% for validation) and employed the Mean Squared Error (MSE) as the loss function. Using the optimal hyperparameters, we trained our model for a maximum of 1,000 epochs, with an early stopping patience of 10 epochs. Next, we performed model interpretation using SHapley Additive exPlanations (SHAP) [40] analysis, ranking each CpG site based on its importance. Whereas ‘DeepMAGE’ [25] selected the 1,000 most important CpG sites for its final model, we aimed to build a more efficient model by selecting a smaller number of CpG sites. We evaluated model performance with a validation set, using CpG sites ranging from 100 to 900, in increments of 100. The optimal CpG sites were selected based on the following criteria: lowest MAE and MSE values, and highest Pearson correlation coefficient (R).All subsequent training, validation, and testing processes used the best hyperparameter settings and optimal CpG site selections from the previous step. For 10-fold CV, the results of each fold were averaged across all folds to derive the model performance metrics. Additionally, we trained linear regression and elastic net regression models with covariates of sex and BMI, and then evaluated their performance on the test set. In this step, for 9 participants (0.5%) in the AA study with missing BMI values, we applied multiple imputations based on the research conducted by Mishra et al. [41]. Finally, we use pyaging [42] to assess the performance of our model against existing DNAm aging clocks using a test dataset.To compare with a neural network-based model, we developed and evaluated both a linear regression and an elastic net regression model using the same optimal CpG sites selected above. We included BMI and sex as covariates. Both models were trained on the training set and evaluated using the same performance metrics as the neural network-based model, including MAE, MSE, and R.

External validation

To verify our model’s generalization, we used three unseen datasets from East Asian populations: GSE92767 [43], GSE214901 [44], and Korean bipolar disorder cohort. GSE92767 [43] contains saliva DNA methylation data (Infinium Human Methylation 450 k BeadChip) from 54 healthy male Korean individuals aged 18 to 73 years. Additionally, GSE214901 [44] contains DNA methylation data (Infinium MethylationEPIC BeadChip) from multiple tissues (brain, blood, saliva, and buccal epithelial tissues) from 19 patients, aged 13 to 73 years, who underwent neurosurgery. We used only the blood-derived DNA methylation data from this dataset. Lastly, the Korean bipolar disorder cohort consists of 129 patients with blood-derived DNA methylation data, which was profiled using the Illumina EPIC v1 and v2 arrays. This cohort was composed of 56 males and 73 females, with an age range of 19–59 years (mean age: 36.8 years). Samples were collected from three clinical centers: Samsung Medical Center (N = 73), Seoul National University Bundang Hospital (N = 25), and Seoul National University Hospital (N = 31). From these datasets, we selected 100 individuals aged 40 or older (≥ 40 years) because our model’s training data was restricted to participants in this age group. The final validation dataset was composed as follows: GSE92767 [43] (32 individuals), GSE214901 [44] (17 individuals), and the Korean bipolar disorder cohort (51 individuals).

Statistics

Statistical and bioinformatic analysis

We conducted a comparative analysis of a neural-network-based model, linear regression, and elastic net regression. The performance was evaluated using MAE, MSE, and R. Additionally, we implemented the ‘train_test_split’ function and ‘KFold’ function for the data splitting step.

The CpG sites selected for the final model development were annotated using ‘minfi’ R package [45] with the Illumina platform manifest (‘IlluminaHumanMethylation450kanno.ilmn12.hg19’). To identify enriched Gene Ontology (GO) [46] biological processes, we conducted a pathway enrichment analysis on the annotated genes using ‘g:Profiler’ [47]. Statistical significance was defined as P-value < 0.05.

We calculated EAAs as residuals from a linear regression model of epigenetic age against the age at blood sample collection. In this step, we used epigenetic ages from 10-fold cross-validation. Linear regression analyses were performed to assess the association between EAA and phenotypes after adjusting for age at blood sample collection and sex. EAA was the dependent variable, and each phenotype was analyzed individually as the independent variable. The threshold for statistical significance was set at P-value < 0.05.

All statistical analyses were performed using the Python packages Scikit-learn (version 1.7.0) [48], SciPy (version 1.16.1) [49], Python (version 3.11.4) and R software (version 4.3.3) [50].

Results

Study population characteristics

Characteristics of the participants enrolled in the HEXA and AA studies are shown in Table 1. The HEXA study comprised 822 individuals with a mean age of 50.6 years (standard deviation (SD) = 5.83 years), while the AA study included 1,925 participants with a mean age of 52.2 years (SD = 8.73 years). The HEXA study included a higher proportion of male (N = 622, 75.7%) than female participants (N = 200, 24.3%). In contrast, the AA study had a more balanced sex distribution, with male (N = 1,006, 52.3%) and female participants (N = 919, 47.7%). Regarding BMI (kg/m2), 0.1% (N = 1) of the HEXA participants were underweight (BMI < 18.5), while 47.3% (N = 389) were normal weight (18.5 ≤ BMI < 23.0). Overweight participants (23.0 ≤ BMI < 25.0) represented 14.7% (N = 121), and obese participants (BMI ≥ 25.0) represented 37.8% (N = 311). In the AA study, 1.5% (N = 28) of participants were underweight (BMI < 18.5), while 29.1% (N = 562) were normal weight (18.5 ≤ BMI < 23.0). Overweight participants (23.0 ≤ BMI < 25.0) represented 25.5% (N = 494), and obese participants (BMI ≥ 25.0) represented 43.5% (N = 841). Regarding the number of disease histories, 75.5% (N = 621) of the HEXA participants had no history of disease, 17.8% (N = 146) had one, and 6.7% (N = 55) had two or more. In the AA study, 72.7% (N = 1,399) of participants had no history of disease, 22.0% (N = 423) had one, and 5.4% (N = 103) had two or more. Details of the remaining variables are depicted in Table 1.

Table 1.

Characteristics of participants

Characteristics HEXA study
(N = 822)
AA study
(N = 1,925)
Age at blood sample collection (year)
 Mean (SD) 50.6 (5.83) 52.2 (8.73)
Sex
 Male 622 (75.7%) 1,006 (52.3%)
 Female 200 (24.3%) 919 (47.7%)
BMI (kg/m2)
 Underweight (BMI < 18.5) 1 (0.1%) 28 (1.5%)
 Normal (18.5 ≤ BMI < 23.0) 389 (47.3%) 562 (29.1%)
 Overweight (23.0 ≤ BMI < 25.0) 121 (14.7%) 494 (25.5%)
 Obese (25.0 ≤ BMI) 311 (37.8%) 841 (43.5%)
 Unknown 0 (0%) 9 (0.5%)
The number of disease histories
 0 621 (75.5%) 1,399 (72.7%)
 1 146 (17.8%) 423 (22.0%)
 ≥ 2 55 (6.7%) 103 (5.4%)
Smoking status
 Never 0 (0.0%) 1,060 (55.1%)
 Ever 103 (12.5%) 833 (43.3%)
 Unknown 719 (87.5%) 32 (1.7%)
Alcohol consumption status
 Never 178 (21.7%) 816 (42.4%)
 Ever 644 (78.3%) 1,086 (56.4%)
 Unknown 0 (0.0%) 23 (1.2%)
Stress
 None 471 (57.3%) 1,029 (53.5%)
 Moderate 291 (35.4%) 735 (38.2%)
 High 56 (6.8%) 142 (7.4%)
 Unknown 4 (0.5%) 19 (1.0%)
Education
 High school or less 500 (60.8%) 1,625 (84.4%)
 College, no four-year degree 67 (8.2%) 285 (14.8%)
 College, four-year degree or higher 248 (30.2%) 0 (0.0%)
 Unknown 7 (0.9%) 15 (0.8%)
Physical activity
 Infrequent 63 (7.7%) 99 (5.1%)
 Moderate 69 (8.4%) 181 (9.4%)
 Frequent 40 (4.9%) 120 (6.2%)
 Daily 36 (4.4%) 157 (8.2%)
 Unknown 614 (74.7%) 1,368 (71.1%)
Monthly income
 1 15 (1.8%) 346 (18.0%)
 2 23 (2.8%) 315 (16.4%)
 3 46 (5.6%) 316 (16.4%)
 4 95 (11.6%) 238 (12.4%)
 5 185 (22.5%) 342 (17.8%)
 6 198 (24.1%) 199 (10.3%)
 7 143 (17.4%) 96 (5.0%)
 8 93 (11.3%) 42 (2.2%)
 Unknown 24 (2.9%) 31 (1.6%)

The entire variables were collected through a combination of questionnaires, physical examinations, and clinical investigations. Regarding the number of disease histories, it was stratified based on the number of diagnoses across five diseases (hypertension, dyslipidemia, diabetes mellitus, myocardial infarction, and cardiovascular disease). For education, participants were categorized into three groups: ‘High school or less’ for those with a high school diploma or less, ‘College, no four-year degree’ for those who attended university without graduating or earned a two-year degree, and ‘College, four-year degree or higher’ for those who graduated from a four-year university or graduate school. Monthly income was categorized into eight levels. Level ‘1’ represented an income of less than 500,000 Korean won (KRW), while level ‘8’ represented 6,000,000 KRW or more. Intermediate levels were defined by intervals of increasing size; 500,000 KRW increments from level ‘2’ to ‘4’, 1,000,000 KRW increments for level ‘5’ and ‘6’, and 2,000,000 KRW increments for level ‘7’. Additionally, for physical activity, participants were categorized into four groups: ‘Infrequent’ for those exercising 1–2 times per week, ‘Moderate’ for 3–4 times per week, ‘Frequent’ for 5–6 times per week, and ‘Daily’ for 7 times per week. For stress, participants were categorized into three levels: ‘None’ for those reporting no stress, ‘Moderate’ for those reporting occasional stress, and ‘High’ for those reporting frequent stress

BMI Body mass index, SD Standard deviation, HEXA Health examinee, AA Ansan and Ansung

Hyperparameter optimization

Hyperparameter optimization was performed exclusively on the training and validation datasets in a hold-out trial. As a result, the following optimal hyperparameter settings were identified: 2,048 hidden units, three hidden layers, a dropout rate of 0.4, CELU activation function, AdamW optimizer, a learning rate of 1e-4, and weight decay of 1e-5. Subsequently, we trained and evaluated the models with the optimal hyperparameter settings to determine the optimal number of CpG sites using MAE, MSE, and R. As shown in Additional File 1: Table S1 and Fig. 2A, the model trained on 300 CpG sites showed the lowest MAE (2.82) and MSE (12.68), and the highest R (0.90). Based on these findings, a model trained on 300 CpG sites was selected for further analysis.

Fig. 2.

Fig. 2

Model performance metrics and SHAP value for each CpG site. A Stem plot of performance metrics among models trained with different numbers of CpG sites. Abbreviations: MAE, mean absolute error; MSE, mean squared error; R, Pearson correlation coefficient. B Scatterplot of age predictions on the hold-out test set, colored by absolute error. a) The color intensity of each dot visually encodes the absolute error between the predicted age and chronological age in the test set, with lighter colors indicating larger errors in prediction. b) In the hold-out test dataset, Pearson's correlation coefficient (R) was 0.90. C Beta and SHAP values summary plot for 300 CpG sites. a) The x-axis presents 300 CpG sites, which are selected for the final model development, and the y-axis presents the SHAP value for each CpG site. b) The color intensity of each dot shows the mean beta value of its corresponding CpG site across the entire dataset. Abbreviations: SHAP, SHapley Additive exPlanations

Model performance

We evaluated the model performance using the MAE, MSE, and R on the test set for both the hold-out and 10-fold CV (Table 2). Our model demonstrated the highest performance compared to the linear regression and elastic net regression models. In the holdout evaluation, our model (KoMethylNet) showed an MAE of 2.82, lower than 2.85 for the linear regression and 3.10 for elastic net regression. Additionally, our model showed an MSE of 12.68, lower than 13.22 for linear regression and 14.31 for elastic net regression. Furthermore, R from our model was 0.90, the same as for linear regression and elastic net (Table 2 and Fig. 2B). Furthermore, our model achieved the best performance in 10-fold CV (MAE = 2.26, MSE = 9.11, and R = 0.93) compared to linear (MAE = 2.38, MSE = 9.31, and R = 0.92) and elastic net regressions (MAE = 2.96, MSE = 13.37, and R = 0.90) (Table 2). Details of the test results for the 10-fold CV are presented in Additional File 1: Table S2.

Table 2.

Model performance in hold-out and 10-fold cross-validation

Models Hold-out 10-fold cross-validation
MAE MSE R MAE MSE R
KoMethylNet 2.82 12.68 0.90 2.26 9.11 0.93
Linear regression 2.85 13.22 0.90 2.38 9.31 0.92
Elastic net regression 3.10 14.31 0.90 2.96 13.37 0.90

MAE Mean absolute error, MSE Mean squared error, R (Pearson’s correlation coefficient

Next, we evaluated the model performance on three unseen datasets using a hold-out validation (Table 3). In the first dataset, GSE92767 [43] (healthy Korean individuals), our model achieved the lowest MAE (2.74) and MSE (12.29). However, its R value (0.94) was lower than those of Zhang’s aging clock [20] (0.96) and AltumAge [26] (0.95). Additionally, our model’s performance in GSE214901 [44] (Japanese neurosurgery patients) showed the second-lowest MAE (5.77) and MSE (54.18), following Zhang’s aging clock (MAE = 4.52 and MSE = 52.68). Additionally, KoMethylNet’s R value (0.50) was lower than those of four other DNAm aging clocks (GrimAge [19]: 0.71; Zhang’s aging clock: 0.70; DNAm-PhenoAge [18]: 0.62; Hannum’s aging clock [17]: 0.59). Lastly, in Korean bipolar disorder cohort, our model achieved a higher MAE (5.65) and MSE (45.71) compared to three other DNAm aging clocks (Hannum’s aging clock: MAE = 2.97, MSE = 13.59; Horvath’s aging clock [16]: MAE = 3.85, MSE = 21.35; AltumAge [26]: MAE = 4.02, MSE = 25.83). The R value (0.63) was lower than those of all the other DNAm aging clocks except for GrimAge (0.03).

Table 3.

Model performance for external validation

Models GSE92767
(N = 32)
GSE214901
(N = 17)
Korean bipolar disorder cohort
(N = 51)
MAE MSE R MAE MSE R MAE MSE R
KoMethylNet 2.74 12.29 0.94 5.77 54.18 0.50 5.65 45.71 0.63
Horvath 3.94 20.95 0.86 7.17 105.71 0.30 3.85 21.35 0.74
Hannum 10.47 127.48 0.88 10.88 142.61 0.59 2.97 13.59 0.73
DNAm-PhenoAge 4.40 30.84 0.85 9.34 111.49 0.62 23.58 576.81 0.69
GrimAge 7.57 69.33 0.92 5.99 58.51 0.71 10.07 137.41 0.03
Zhang 3.26 17.03 0.96 4.52 56.08 0.70 13.07 185.91 0.86
AltumAge 4.10 28.83 0.95 5.92 62.81 0.47 4.02 25.83 0.78

Individuals aged 40 or older (≥ 40 years) were used because our model’s training data was restricted to participants in this age group

DNAm DNA methylation, MAE Mean absolute error, MSE Mean squared error, R Pearson’s correlation coefficient, Horvath Horvath’s aging clock, Hannum Hannum’s aging clock, Zhang Zhang’s aging clock

Characteristics and pathway analysis of 300 CpG sites

The 300 CpG sites used for KoMethylNet overlap with CpG sites selected by other DNAm aging clocks. The highest overlap is Zhang’s aging clock, which shares 50 of the 300 CpG sites. Following Zhang’s aging clock, AltumAge (18 CpG sites), Hannum’s aging clock (17 CpG sites), DNAm-PhenoAge (10 CpG sites), Horvath’s aging clock (5 CpG sites), and GrimAge (5 CpG sites) share their CpG sites with KoMethylNet. Furthermore, Zhang’s aging clock has an exclusive overlap of 34 CpG sites with KoMethylNet, followed by AltumAge (8 CpG sites), Hannum’s aging clock (3 CpG sites), and GrimAge (3 CpG sites), while Horvath’s aging clock and DNAm-PhenoAge do not have any exclusive overlap with KoMethylNet. Details of all shared CpG sites are described in Additional File 1: Table S3.

Overall SHAP values and mean beta values of 300 CpG sites are presented in Fig. 2C. Among these, genomic annotations and beta value metrics of the top 10 CpG sites ranked by SHAP value are described in Table 4. SHAP values in these CpG sites ranged from 0.00066 to 0.00118 and mean beta values ranged from 0.27 to 0.70. The CpG sites were distributed across eight chromosomes (1, 2, 3, 6, 7, 10, 12, and 15), with five mapping to annotated genes and five located in intergenic regions. Three CpG sites were positioned within a gene body: cg01620164 in FIGN (chr2, body, north shelf), cg04875128 in OTUD7A (chr15, body, CpG island), and cg01820962 in NT5DC1 (chr6, body, no island). Promoter-proximal regions were represented by cg07553761 in TRIM59 (chr3, TSS1500, CpG island) and cg14209784 in AGAP11 (chr10, TSS1500, north shore). Other CpG sites were located in island-related regions without gene annotation, including cg22943590 (chr2, north shelf), cg11807280 (chr2, south shore), and cg19663246 (chr7, south shore). Two CpG sites lacked gene and island annotation altogether: cg08993878 (chr12) and cg10501210 (chr1). Notably, three CpG sites (cg10501210, cg01820962, and cg14209784) showed an association with enhancer regions, highlighting their potential regulatory importance. The entire list of 300 CpG sites with their corresponding annotations and beta value metrics are described in Additional File 1: Table S4.

Table 4.

Characteristics of the top 10 CpG sites with the highest SHAP values

CpG site SHAP value Beta value (mean [min–max]) Chr Gene Genic region CpG island region Enhancer
cg01620164 0.00118 0.46 [0.11–0.99] 2 FIGN Body N Shelf -
cg08993878 0.00097 0.41 [0.06–0.87] 12 - - - -
cg10501210 0.00085 0.70 [0.38–0.95] 1 - - - Yes
cg04875128 0.00084 0.27 [0.07–0.57] 15 OTUD7A Body Island -
cg22943590 0.00075 0.51 [0.13–0.93] 2 - - N Shelf -
cg11807280 0.00072 0.31 [0.07–0.67] 2 - - S Shore -
cg19663246 0.00069 0.41 [0.09–0.74] 7 - - S Shore -
cg01820962 0.00069 0.66 [0.28–0.95] 6 NT5DC1 Body - Yes
cg07553761 0.00068 0.47 [0.18–0.74] 3 TRIM59 TSS1500 Island -
cg14209784 0.00066 0.61 [0.32–0.87] 10 AGAP11 TSS1500 N Shore Yes

All genomic annotations are based on the GRCh37/hg19 genome build. Beta values are calculated throughout the entire study population

Body Between the ATG and stop codon; irrespective of the presence of introns, exons, TSS, or promoters, Chr Chromosome, CpG island The location of the CpG site relative to the CpG island, TSS Transcription start site, TSS1500 200–1500 bases upstream of the TSS, Shore 0–2 kb from island, Shelf 2–4 kb from island, N Upstream (5’) of CpG island, S Downstream (3’) of CpG island, SHAP SHapley Additive exPlanations

We used GO biological process analysis on the genes annotated with 300 CpG sites. The analysis identified an enrichment of 20 biological processes (Additional File 1: Table S5), which were categorized into three key areas: organismal development, cellular signaling, and regulatory processes. The most significant biological process was ‘multicellular organism development’ (P-value = 1.91 × 10–4), followed by ‘cell communication’ (P-value = 1.95 × 10–4), ‘anterograde trans-synaptic signaling’ (P-value = 3.20 × 10–4).

The association between KoMethylNet-EAA and phenotypes

We found seven phenotypes associated with KoMethylNet-EAA that had P-values lower than 0.05 (Table 5). First, we identified three phenotypes related to diabetes: diabetes (β = 0.81, SE = 0.19, P-value = 1.29 × 10–5), urine glucose level (Negative: Reference; 2 + : β = 0.67, SE = 0.28, P-value = 1.93 × 10–2; 3 + : β = 1.11, SE = 0.34, P-value = 1.00 × 10–3) and HbA1c (β = 0.14, SE = 0.04, P-value = 5.54 × 10–4). Second, we identified that triglycerides and uric acid were associated with KoMethylNet-EAA: triglycerides (β = 0.001, SE = 0.001, P-value = 7.13 × 10–3), uric acid (β = 0.01, SE = 0.003, P-value = 1.17 × 10–3). Lastly, we found that two socioeconomic phenotypes were negatively associated with KoMethylNet-EAA: education (High school or less: Reference; College, four-year degree or higher: β = −0.41, SE = 0.20, P-value = 4.46 × 10–2) and monthly income (1: Reference; 8: β =−0.69, SE = 0.31, P-value = 2.61 × 10–2). The entire analysis results are presented in Additional File 1:Table S6.

Table 5.

Significant phenotypes associated with KoMethylNet-EAA

Phenotype β SE P-value
Diabetes mellitus 0.81 0.19 1.29 × 10–5
Education
 High school or less Ref
 College, four-year degree or higher −0.41 0.20 4.46 × 10–2
HbA1c 0.14 0.04 5.54 × 10–4
Monthly income
 1 Ref
 8 −0.69 0.31 2.61 × 10–2
Triglycerides 0.001 0.001 7.13 × 10–3
Uric acid 0.01 0.003 1.17 × 10–3
Urine glucose
 Negative Ref
 2 +  0.67 0.28 1.93 × 10–2
 3 +  1.11 0.34 1.00 × 10–3

The entire variables were collected through a combination of questionnaires, physical examinations, and clinical investigations. For education, participants were categorized into three groups: ‘High school or less’ for those with a high school diploma or less and ‘College, four-year degree or higher’ for those who graduated from a four-year university or graduate school. Monthly income was categorized into eight levels. Level ‘1’ represented an income of less than 500,000 Korean won (KRW), while level ‘8’ represented 6,000,000 KRW or more. Intermediate levels were defined by intervals of increasing size; 500,000 KRW increments from level ‘2’ to ‘4’, 1,000,000 KRW increments for level ‘5’ and ‘6’, and 2,000,000 KRW increments for level ‘7’

EAA Epigenetic age acceleration, SE Standard error, HbA1c Hemoglobin A1c, Ref Reference

Discussion

This is the first study to develop a neural network-based DNAm aging clock (KoMethylNet) in a Korean population. We used SHAP analysis to identify the optimal number of CpG sites (300 CpGs) for our model to achieve strong performance metrics. In the hold-out method, the results were MAE = 2.82, MSE = 12.68, and R = 0.90, whereas in the 10-fold CV method, the results were MAE = 2.26, MSE = 9.11, and R = 0.93. These results were superior compared to linear and elastic net regression models. In the external validation, our model achieved the highest performance metrics compared to other DNAm aging clocks in healthy South Korean individuals, while it was not the highest in Japanese neurosurgery patients and Korean bipolar disorder cohort. Lastly, we identified seven phenotypes with statistically significant associations with KoMethylNet-EAA: diabetes mellitus, education, HbA1c, monthly income, triglycerides, uric acid and urine glucose.

Overall, our model achieved higher accuracy than the DNA methylation aging clocks developed using linear and elastic net regression. This result is primarily owing to the capacity of the neural network model to capture and utilize nonlinear relationships among input features compared with the linear and elastic net regression model. Furthermore, in the external validation, our model achieved the lowest MAE and MSE in healthy Korean individuals (GSE92767 [43]). The closest finding was from Zhang’s aging clock [20], which shares the largest number of CpG sites with KoMethylNet. In contrast, for Japanese neurosurgery patients (GSE214901 [44]), Zhang’s aging clock recorded the lowest MAE and MSE, while KoMethylNet was the second-best performer. However, this trend did not hold for Korean bipolar disorder cohort, as KoMethylNet was inferior to three other DNAm aging clocks (Horvath [16], Hannum [17], and AltumAge [26]). Interestingly, these DNAm aging clocks showed larger MAE and MSE in healthy Korean individuals (GSE92767) than in Korean bipolar disorder cohort, a finding that is inconsistent with previous studies reporting epigenetic age acceleration in bipolar patients [51–54]. Collectively, these results indicate that KoMethylNet has potential utility for East Asian populations and may reflect epigenetic age acceleration in specific diseases. This finding is in line with previous studies on the importance of ethnicity in the study population [34, 55].

Among the top 10 CpG sites selected for our model, 20% (cg04875128 and cg07553761) were shared with Hannum’s and Zhang’s aging clock, 40% (cg01620164, cg11807280, cg19663246, and cg14209784) were shared with Zhang’s aging clocks, and the remaining 40% (cg08993878, cg10501210, cg22943590, and cg01820962) were novel CpG sites not used in the development of previous DNAm aging clocks (Horvath, Hannum, DNAm-PhenoAge, GrimAge, Zhang, and AltumAge). It is notable that a large portion of these top 10 CpG sites were shared with Zhang’s aging clock, which demonstrated generalizability in Korean and Japanese populations as our external validation results showed. Furthermore, cg01620164 exhibited the strongest SHAP value. Yusipov et al. found that the methylation level of cg01620164 is associated with not only age, but also sex [56]. Interestingly, cg01620164 is only shared with Zhang’s aging clock (Additional File 1:Table S3). Considering our external validation results, cg01620164 may be relevant to the mechanisms of aging in Asian populations. Additionally, cg01620164 is annotated to FIGN, which functions in DNA synthesis, mitosis, and cellular migration [57]. Furthermore, the overexpression of FIGN promotes hepatocyte invasion, which may contribute to tumor progression [57].

Diabetes-related phenotypes (diabetes, HbA1c, and urine glucose) exhibited positive associations with KoMethylNet-EAA. Chang et al. identified that diabetes-related phenotypes including fasting glucose and HbA1c, were proportionally associated with EAA in the Taiwan Biobank population, using second-generation DNAm aging clocks (GrimEAA, DunedinPACE, and PhenoEAA) [58]. Additionally, Miao et al. found that the presence of type 2diabetes or glycemic traits (fasting plasma glucose, HbA1c, and triglyceride-glucose index) were positively associated with EAA in the Chinese National Twin Registry population, using three DNAm aging clocks (PCGrimAge, PCPhenoAge, and Dunedin PACE) [59]. The results of these studies in Asian populations are consistent with our findings in the Korean population. Additionally, triglycerides were positively associated with KoMethylNet-EAA. Kawamura et al. reported that triglycerides were positively associated with EAA in the Japanese population, consistent with the findings of our study [60]. Furthermore, uric acid was positively associated with KoMethylNet-EAA. This is consistent with a study by Wu et al., which reported the positive association between uric acid and EAA in the Taiwanese population [61]. To our knowledge, our finding is the first to identify this association in the Korean population. Lastly, two socioeconomic phenotypes (education and monthly income) were negatively associated with KoMethylNet-EAA. These findings are consistent with the study by Kim et al., which reported that higher socioeconomic status including education and income was associated with lower EAA [62]. This consistency highlights the potential of KoMethylNet to capture the biological impact of socioeconomic status on aging in the Korean population.

Our study has two strengths. First, we propose the first Korean-based DNAm aging clock developed using a neural network model, KoMethylNet. Using KoMethylNet, we could provide better predictions of Korean-specific biological aging. Second, we identified the associations between KoMethylNet-EAA and phenotypes such as uric acid and socioeconomic status, which were novel or consistent in the Korean population. Our study has three limitations. First, due to insufficient DNAm data from the Korean population, our model’s training sample size was relatively small compared to existing DNAm aging clocks (e.g., Zhang’s aging clock: 8,050 samples). Second, we could not use individuals younger than 40 years (< 40 years) in the external validation, because our training data was restricted to individuals aged 40 or older (≥ 40 years). For further research, a larger Korean DNAm dataset is required for model development, and it will give us the opportunity to expand KoMethylNet’s generalization ability on the Korean population. Lastly, while we integrated 450 K and EPIC array data, the potential platform-related biases cannot be ruled out entirely. Such biases could influence downstream analyses and may limit generalizability to DNAm data derived from a single array type.

Conclusions

To summarize, we developed a neural-network-based DNAm aging clock (KoMethylNet) using a Korean population. KoMethylNet achieved the highest performance in the external validation using healthy Korean individuals. Furthermore, we observed associations between KoMethylNet-EAA and seven phenotypes including diabetes and urine glucose. Although further research with a larger dataset is necessary, our study provides a novel DNAm aging clock that can enable precise age prediction and offer potential opportunities for aging-related research in the Korean population.

Supplementary Information

12916_2025_4564_MOESM1_ESM.pdf (1.6MB, pdf)

Supplementary Material 1. Table S1. Model test results for different numbers of CpG sites. Table S2. Training and evaluation results of 10-fold cross-validation. Table S3. 72 CpG sites shared with previous DNA methylation aging clocks. Table S4. Characteristics of 300 CpG sites ranked by SHAP values. Table S5. Pathway enrichment analyses of genes annotated from 300 CpG sites. Table S6. Entire association between phenotypes and KoMethylNet-EAA.

Acknowledgements

The authors thank all individuals who participated in this study.

Abbreviations

AA

Ansan and Ansung

BMI

Body Mass Index

ChAMP

Chip Analysis Methylation Pipeline

CV

Cross-Validation

DNAm

DNA Methylation

EAA

Epigenetic Age Acceleration

HbA1c

Hemoglobin A1c

HEXA

Health Examinee

MAE

Mean Absolute Error

MSE

Mean Squared Error

SHAP

SHapley Additive exPlanations

SE

Standard Error

Authors’ contributions

D.Y. and N.S. conceptualized the study. D.Y. and N.S. designed and performed the analysis. Y.K. and Y.M.A. performed the validation analysis on the Korean bipolar disorder cohort data. K.Y.O. contributed to the preprocessing of the genetic data. Z.W. and H.M.P. contributed the interpretation of the results. D.Y. and N.S. drafted the initial version of the manuscript. Z.W., X.M., and K.Y.O. reviewed and edited the manuscript. All authors read and approved the final manuscript.

Funding

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2024-00358322, RS-2025-02273102). This research was supported by the Bio & Medical Technology Development Program of the National Research Foundation (NRF) funded by the Korean government (MSIT) (No. RS-2024-00440787).

Data availability

The demographic and DNA methylome data for the Ansan and Ansung, Health Examinee studies were provided by the National Biobank of Korea and the Center for Disease Control and Prevention, Republic of Korea. These data are not publicly available and can be requested via the website (https://biobank.nih.go.kr) following the requisite approval process. The demographic and DNA methylome data for the Korean bipolar disorder cohort are not publicly available. Therefore, the analysis was performed by internal researchers at Samsung Medical Center, Republic of Korea. Additionally, GSE92767 and GSE214901 are publicly available through GEO Datasets (https://www.ncbi.nlm.nih.gov/gds).

Declarations

Ethics approval and consent to participate

Data in this study were from the Korean Genome and Epidemiology Study (KoGES; 2023-043), National Institute of Health, Korea Disease Control and Prevention Agency, Republic of Korea. The study protocol was approved by IRB of Korean National Institutes of Health (IRB: CBNU-2025-E-0001, CBNU-2025-E-0003), Seoul National University Hospital (IRB No. 1905–150-1035), Samsung Medical Center (IRB No. 2019–02-038), and Seoul National University Bundang Hospital (IRB No. B-1908–559-404) and adhered to the principles outlined in the Declaration of Helsinki. Written informed consent was obtained from all participants before the interviews.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Dziechciaz M, Filip R. Biological psychological and social determinants of old age: bio-psycho-social aspects of human aging. Ann Agric Environ Med. 2014;21(4):835–8. [DOI] [PubMed] [Google Scholar]
  • 2.Trevisan K, Cristina-Pereira R, Silva-Amaral D, Aversi-Ferreira TA. Theories of aging and the prevalence of Alzheimer’s disease. BioMed Res Int. 2019;2019:9171424. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Faye AS, Colombel JF. Aging and IBD: a new challenge for clinicians and researchers. Inflamm Bowel Dis. 2022;28(1):126–32. [DOI] [PubMed] [Google Scholar]
  • 4.Cho SJ, Stout-Delgado HW. Aging and lung disease. Annu Rev Physiol. 2020;82:433–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Liberale L, Badimon L, Montecucco F, Luscher TF, Libby P, Camici GG. Inflammation, aging, and cardiovascular disease: JACC review topic of the week. J Am Coll Cardiol. 2022;79(8):837–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Hoeijmakers JH. DNA damage, aging, and cancer. N Engl J Med. 2009;361(15):1475–85. [DOI] [PubMed] [Google Scholar]
  • 7.Morris BJ, Willcox BJ, Donlon TA. Genetic and epigenetic regulation of human aging and longevity. Biochim Biophys Acta Mol Basis Dis. 2019;1865(7):1718–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Sen P, Shah PP, Nativio R, Berger SL. Epigenetic mechanisms of longevity and aging. Cell. 2016;166(4):822–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.la Torre A, Lo Vecchio F, Greco A. Epigenetic Mechanisms of Aging and Aging-Associated Diseases. Cells. 2023;12(8):1163. [DOI] [PMC free article] [PubMed]
  • 10.Wang K, Liu H, Hu Q, Wang L, Liu J, Zheng Z, et al. Epigenetic regulation of aging: implications for interventions of aging and diseases. Signal Transduct Target Ther. 2022;7(1):374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Johnson AA, Akman K, Calimport SR, Wuttke D, Stolzing A, de Magalhaes JP. The role of DNA methylation in aging, rejuvenation, and age-related disease. Rejuvenation Res. 2012;15(5):483–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Salameh Y, Bejaoui Y, El Hajj N. DNA methylation biomarkers in aging and age-related diseases. Front Genet. 2020;11:171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Noroozi R, Ghafouri-Fard S, Pisarek A, Rudnicka J, Spolnicka M, Branicki W, et al. DNA methylation-based age clocks: from age prediction to age reversion. Ageing Res Rev. 2021;68:101314. [DOI] [PubMed] [Google Scholar]
  • 14.Horvath S, Raj K. DNA methylation-based biomarkers and the epigenetic clock theory of ageing. Nat Rev Genet. 2018;19(6):371–84. [DOI] [PubMed] [Google Scholar]
  • 15.Bergsma T, Rogaeva E. DNA methylation clocks and their predictive capacity for aging phenotypes and healthspan. Neurosci Insights. 2020;15:2633105520942221. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Horvath S. DNA methylation age of human tissues and cell types. Genome Biol. 2013;14(10):R115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Hannum G, Guinney J, Zhao L, Zhang L, Hughes G, Sadda S, et al. Genome-wide methylation profiles reveal quantitative views of human aging rates. Mol Cell. 2013;49(2):359–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Levine ME, Lu AT, Quach A, Chen BH, Assimes TL, Bandinelli S, et al. An epigenetic biomarker of aging for lifespan and healthspan. Aging (Albany NY). 2018;10(4):573–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lu AT, Quach A, Wilson JG, Reiner AP, Aviv A, Raj K, et al. DNA methylation GrimAge strongly predicts lifespan and healthspan. Aging (Albany NY). 2019;11(2):303–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhang Q, Vallerga CL, Walker RM, Lin T, Henders AK, Montgomery GW, et al. Improved precision of epigenetic clock estimates across tissues and its implication for biological ageing. Genome Med. 2019;11(1):54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ching T, Himmelstein DS, Beaulieu-Jones BK, Kalinin AA, Do BT, Way GP, et al. Opportunities and obstacles for deep learning in biology and medicine. J R Soc Interface. 2018. 10.1098/rsif.2017.0387. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Cao C, Liu F, Tan H, Song D, Shu W, Li W, et al. Deep Learning and Its Applications in Biomedicine. Genomics Proteomics Bioinformatics. 2018;16(1):17–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Levy JJ, Titus AJ, Petersen CL, Chen Y, Salas LA, Christensen BC. Methylnet: an automated and modular deep learning approach for DNA methylation analysis. BMC Bioinformatics. 2020;21(1):108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Johansson A, Enroth S, Gyllensten U. Continuous aging of the human DNA methylome throughout the human lifespan. PLoS ONE. 2013;8(6):e67378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Galkin F, Mamoshina P, Kochetov K, Sidorenko D, Zhavoronkov A. DeepMAge: a methylation aging clock developed with deep learning. Aging Dis. 2021;12(5):1252–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.de Lima Camillo LP, Lapierre LR, Singh R. A pan-tissue DNA-methylation epigenetic clock based on deep learning. NPJ Aging. 2022;8(1):4. [Google Scholar]
  • 27.Yusipov I, Kalyakulina A, Trukhanov A, Franceschi C, Ivanchenko M. Map of epigenetic age acceleration: a worldwide analysis. Ageing Res Rev. 2024;100:102418. [DOI] [PubMed] [Google Scholar]
  • 28.Oblak L, van der Zaag J, Higgins-Chen AT, Levine ME, Boks MP. A systematic review of biological, social and environmental factors associated with epigenetic clock acceleration. Ageing Res Rev. 2021;69:101348. [DOI] [PubMed] [Google Scholar]
  • 29.Chervova O, Panteleeva K, Chernysheva E, Widayati TA, Baronik ZF, Hrbkova N, et al. Breaking new ground on human health and well-being with epigenetic clocks: A systematic review and meta-analysis of epigenetic age acceleration associations. Ageing Res Rev. 2024;102:102552. [DOI] [PubMed] [Google Scholar]
  • 30.Qin N, Li Z, Song N, Wilson CL, Easton J, Mulder H, et al. Epigenetic age acceleration and chronic health conditions among adult survivors of childhood cancer. J Natl Cancer Inst. 2021;113(5):597–605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Crimmins EM, Thyagarajan B, Levine ME, Weir DR, Faul J. Associations of age, sex, race/ethnicity, and education with 13 epigenetic clocks in a nationally representative U.S. sample: the health and retirement study. J Gerontol A Biol Sci Med Sci. 2021;76(6):1117–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Watkins SH, Testa C, Chen JT, De Vivo I, Simpkin AJ, Tilling K, et al. Epigenetic clocks and research implications of the lack of data on whom they have been developed: a review of reported and missing sociodemographic characteristics. Environ Epigenet. 2023;9(1):dvad005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Park J, Won CW, Saligan LN, Kim YJ, Kim Y, Lukkahatai N. Accelerated epigenetic age in normal cognitive aging of Korean community-dwelling older adults. Biol Res Nurs. 2021;23(3):464–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zheng Z, Li J, Liu T, Fan Y, Zhai QC, Xiong M, et al. DNA methylation clocks for estimating biological age in Chinese cohorts. Protein Cell. 2024;15(8):575–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kim Y, Han BG, Ko GESg. Cohort Profile: The Korean Genome and Epidemiology Study (KoGES) Consortium. Int J Epidemiol. 2017;46(2):e20. [DOI] [PMC free article] [PubMed]
  • 36.Morris TJ, Butcher LM, Feber A, Teschendorff AE, Chakravarthy AR, Wojdacz TK, et al. ChAMP: 450k chip analysis methylation pipeline. Bioinformatics. 2014;30(3):428–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Zhou W, Laird PW, Shen H. Comprehensive characterization, annotation and innovative use of Infinium DNA methylation BeadChip probes. Nucleic Acids Res. 2017;45(4):e22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Teschendorff AE, Marabita F, Lechner M, Bartlett T, Tegner J, Gomez-Cabrero D, et al. A beta-mixture quantile normalization method for correcting probe design bias in Illumina Infinium 450 k DNA methylation data. Bioinformatics. 2013;29(2):189–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Akiba T, Sano S, Yanase T, Ohta T, Koyama M. Optuna: A Next-generation Hyperparameter Optimization Framework. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2019;2623–31.
  • 40.Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. Neural Information Processing Systems. 2017;30.
  • 41.Mishra GD, Dobson AJ. Multiple imputation for body mass index: lessons from the Australian Longitudinal Study on Women’s Health. Stat Med. 2004;23(19):3077–87. [DOI] [PubMed] [Google Scholar]
  • 42.de Lima Camillo LP. Pyaging: a Python-based compendium of GPU-optimized aging clocks. Bioinformatics. 2024. 10.1093/bioinformatics/btae200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Hong SR, Jung SE, Lee EH, Shin KJ, Yang WI, Lee HY. DNA methylation-based age prediction from saliva: High age predictability by combination of 7 CpG markers. Forensic Sci Int Genet. 2017;29:118–25. [DOI] [PubMed] [Google Scholar]
  • 44.Nishitani S, Isozaki M, Yao A, Higashino Y, Yamauchi T, Kidoguchi M, et al. Cross-tissue correlations of genome-wide DNA methylation in Japanese live human brain and blood, saliva, and buccal epithelial tissues. Transl Psychiatry. 2023;13(1):72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Aryee MJ, Jaffe AE, Corrada-Bravo H, Ladd-Acosta C, Feinberg AP, Hansen KD, et al. Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays. Bioinformatics. 2014;30(10):1363–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, et al. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium Nat Genet. 2000;25(1):25–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Kolberg L, Raudvere U, Kuzmin I, Adler P, Vilo J, Peterson H. G:Profiler-interoperable web service for functional enrichment analysis and gene identifier mapping (2023 update). Nucleic Acids Res. 2023;51(W1):W207–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. the Journal of machine Learning research. 2011;12:2825–30.
  • 49.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.R Core Team. (2024). R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing. Retrieved from https://www.R-project.org/
  • 51.Fries GR, Bauer IE, Scaini G, Wu MJ, Kazimi IF, Valvassori SS, et al. Accelerated epigenetic aging and mitochondrial DNA copy number in bipolar disorder. Transl Psychiatry. 2017;7(12):1283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Fries GR, Bauer IE, Scaini G, Valvassori SS, Walss-Bass C, Soares JC, et al. Accelerated hippocampal biological aging in bipolar disorder. Bipolar Disord. 2020;22(5):498–507. [DOI] [PubMed] [Google Scholar]
  • 53.Lima CNC, Suchting R, Scaini G, Cuellar VA, Favero-Campbell AD, Walss-Bass C, et al. Epigenetic GrimAge acceleration and cognitive impairment in bipolar disorder. Eur Neuropsychopharmacol. 2022;62:10–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Jeremian R, Malinowski A, Chaudhary Z, Srivastava A, Qian J, Zai C, et al. Epigenetic age dysregulation in individuals with bipolar disorder and schizophrenia. Psychiatry Res. 2022;315:114689. [DOI] [PubMed] [Google Scholar]
  • 55.Horvath S, Gurven M, Levine ME, Trumble BC, Kaplan H, Allayee H, et al. An epigenetic clock analysis of race/ethnicity, sex, and coronary heart disease. Genome Biol. 2016;17(1):171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Yusipov I, Bacalini MG, Kalyakulina A, Krivonosov M, Pirazzini C, Gensous N, et al. Age-related DNA methylation changes are sex-specific: a comprehensive assessment. Aging (Albany NY). 2020;12(23):24057–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zhou B, Wang J, Gao J, Xie J, Chen Y. Fidgetin as a potential prognostic biomarker for hepatocellular carcinoma. Int J Med Sci. 2020;17(17):2888–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Chang XY, Lin WY. Epigenetic age acceleration mediates the association between smoking and diabetes-related outcomes. Clin Epigenetics. 2023;15(1):94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Miao K, Hong X, Cao W, Lv J, Yu C, Huang T, et al. Association between epigenetic age and type 2 diabetes mellitus or glycemic traits: a longitudinal twin study. Aging Cell. 2024;23(7):e14175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Kawamura T, Radak Z, Tabata H, Akiyama H, Nakamura N, Kawakami R, et al. Associations between cardiorespiratory fitness and lifestyle-related factors with DNA methylation-based ageing clocks in older men: WASEDA’s health study. Aging Cell. 2024;23(1):e13960. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Wu YR, Lin WY. Associations between lifestyle factors, physiological conditions, and epigenetic age acceleration in an Asian population. Biogerontology. 2025;26(2):51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Kim DJ, Kang JH, Kim JW, Kim SB, Lee YK, Cheon MJ, et al. Assessing the utility of epigenetic clocks for health prediction in South Korean. Front Aging. 2024;5:1493406. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12916_2025_4564_MOESM1_ESM.pdf (1.6MB, pdf)

Supplementary Material 1. Table S1. Model test results for different numbers of CpG sites. Table S2. Training and evaluation results of 10-fold cross-validation. Table S3. 72 CpG sites shared with previous DNA methylation aging clocks. Table S4. Characteristics of 300 CpG sites ranked by SHAP values. Table S5. Pathway enrichment analyses of genes annotated from 300 CpG sites. Table S6. Entire association between phenotypes and KoMethylNet-EAA.

Data Availability Statement

The demographic and DNA methylome data for the Ansan and Ansung, Health Examinee studies were provided by the National Biobank of Korea and the Center for Disease Control and Prevention, Republic of Korea. These data are not publicly available and can be requested via the website (https://biobank.nih.go.kr) following the requisite approval process. The demographic and DNA methylome data for the Korean bipolar disorder cohort are not publicly available. Therefore, the analysis was performed by internal researchers at Samsung Medical Center, Republic of Korea. Additionally, GSE92767 and GSE214901 are publicly available through GEO Datasets (https://www.ncbi.nlm.nih.gov/gds).


Articles from BMC Medicine are provided here courtesy of BMC

RESOURCES