Skip to main content
Aging Cell logoLink to Aging Cell
. 2025 Jun 19;24(9):e70149. doi: 10.1111/acel.70149

MskAge—An Epigenetic Biomarker of Musculoskeletal Age Derived From a Genetic Algorithm Islands Model

Daniel C Green 1,2, Louise N Reynard 2,3, James R Henstock 1,2, Sjur Reppe 4,5,6, Kaare Gautvik 5, Mandy J Peffers 1,2, Daryl P Shanley 3, Peter D Clegg 1,2, Elizabeth G Canty‐Laird 1,2,✉
PMCID: PMC12419849  PMID: 40538142

ABSTRACT

Age is a significant risk factor for functional decline and disease of the musculoskeletal system, yet few biomarkers exist to facilitate ageing research in musculoskeletal tissues. Multivariate models based on DNA methylation, termed epigenetic clocks, have shown promise as markers of biological age. However, the accuracy of existing epigenetic clocks in musculoskeletal tissues are no more, and often less accurate than a randomly sampled baseline model. We developed a highly accurate epigenetic clock, MskAge, that is specific to tissues and cells of the musculoskeletal system. MskAge was built using a penalised genetic algorithm islands model that addresses multi‐tissue clock bias. The final model was trained on the transformed principal components of CpGs selected by the genetic algorithm. We show that MskAge tracks epigenetic ageing ex vivo and in vitro. Epigenetic age estimates are rejuvenated with cellular reprogramming and are accelerated at a rate of 0.45 years per population doubling. MskAge explains more variance associated with in vitro ageing of fibroblasts than the purpose‐developed skin and blood clock. The precision of MskAge and its ability to capture perturbations in biological ageing make it a promising research tool for musculoskeletal and ageing biologists.

Keywords: ageing, biomarker, epigenetics, genetic algorithm, musculoskeletal system


Created in BioRender. Laird, E. (2025) https://BioRender.com/bqfajrn.

graphic file with name ACEL-24-e70149-g003.jpg

1. Introduction

Ageing is a complex and multifaceted process manifesting in a progressive decline of cellular function that is historically linked to the passing of time (Kirkwood 2005). While societal norms associate an individual's age with the number of years since birth, the rate of organismal development and decline is non‐linear across a lifespan and varies between tissues and organisms. This non‐linearity and heterogeneity differentiate biological ageing trajectories from those typically associated with chronological ageing. Owing to the complexity of the ageing process, accurate biomarkers of biological age are becoming increasingly valuable in research and clinical settings (Lu et al. 2019). Most existing biomarkers primarily focus on accessible tissues, such as blood, skin and saliva (Horvath et al. 2018; Levine et al. 2018; Lu et al. 2019). Although these biomarkers are useful for investigating ageing rates at the population level, they may not be as applicable to all body tissues, such as those in the musculoskeletal system. To maximise the likelihood of successfully increasing health span and lifespan, interventions will likely need to rejuvenate cells throughout the entire organism. Consequently, identifying biomarkers that facilitate ex vivo research on musculoskeletal tissues has the potential to elucidate anti‐ageing interventions currently beyond the scope of existing biomarkers.

In the past decade, DNA methylation (DNAm) clocks have emerged as some of the most promising and accurate biomarkers of age (Hannum et al. 2013; Horvath 2013a; Levine et al. 2018; Lu et al. 2019; Voisin et al. 2020). DNAm clocks are multivariate statistical or machine learning models that predict chronological age or age‐related outcomes based on the proportion of methyl groups (CH3) present on DNA at an optimised selection of CpGs. Notably, CpGs selected by clocks are not necessarily those within ageing genes but rather those with associations to age in thousands of samples (Horvath 2013b). The relative stability of DNA methylation allows DNAm clocks to achieve higher accuracy compared to biomarkers developed on other layers of molecular data (Xia et al. 2021). Moreover, the addition of a methyl group to cytosine by DNA methyl transferases (DNMTs) is reversible via its subsequent hydroxylation by the Ten Eleven Translocase (TET) enzymes, providing a dynamic mechanism that can be exploited to track and potentially modify the rate of biological ageing through the methylome (Wu and Zhang 2017). The divergence between predicted epigenetic age and chronological age is referred to as epigenetic age acceleration (Quach et al. 2017). It has been well established that DNAm clocks can predict age acceleration rates independently of chronological age, age‐related disease, and age‐related functional outcomes (e.g., hand grip strength), making them promising candidates as biomarkers of ageing (Horvath 2013b; Horvath et al. 2018; Levine et al. 2018; Lu et al. 2019; Quach et al. 2017).

Ageing profoundly impacts musculoskeletal function, leading to the deterioration of bone, cartilage and muscle, which can cause diseases such as osteoporosis and osteoarthritis, as well as the ageing syndrome of frailty (Roberts et al. 2016). Although attempts have been made to define biomarkers of ageing in musculoskeletal tissues, these primarily focus on biomarkers of bone turnover, skeletal muscle mass, or assessments of physical activity (Kemp et al. 2018). While these biomarkers are useful for in vivo monitoring of physical function, they do not directly assess the age of musculoskeletal cells and cannot be applied ex vivo. To facilitate ex vivo research on musculoskeletal tissues, highly accurate and specific epigenetic biomarkers or other molecular clocks are likely required.

Some existing epigenetic clocks, such as Horvath's original multi‐tissue clock, are applicable to a wide range of tissue and cell types (Horvath 2013b). However, the representation of musculoskeletal tissues in the training of these epigenetic clocks is limited compared to other tissues, making them less reliable for musculoskeletal ageing research due to the tissue‐specific nature of DNA methylation. The primary drawback of non‐specific epigenetic clocks is their poor accuracy and high variance in musculoskeletal tissues, which has implications for designing experiments with appropriate power. Obtaining musculoskeletal samples for research often involves highly invasive procedures, such as surgical operations or post‐mortem extractions, highlighting the need for more accurate and specific epigenetic clocks to design appropriately powered studies. Recently, a highly accurate epigenetic clock (error +/− 4.6 years) was developed to track the age of human skeletal muscle using ~19,000 CpGs common between the Infinium 27 k, Infinium 450 k and Infinium EPIC arrays (Voisin et al. 2020). The skeletal muscle clock exemplifies an instance where the limitations of Horvath's original multi‐tissue clock were addressed, particularly its poor calibration for underrepresented samples in the training set of the model.

In the present study, we demonstrate that existing epigenetic clocks are no more and, in some instances, less accurate than a randomly sampled baseline model across multiple musculoskeletal tissues. We employed a genetic algorithm‐based meta‐heuristic feature selection framework to construct a musculoskeletal tissue‐specific epigenetic clock, optimising CpG selection to minimise error within each tissue. Recently, it was shown that the algebraic transformation of CpG methylation into principal components for use as features in training an epigenetic clock addresses several technical issues arising from CpG feature models (Higgins‐Chen et al. 2022). Our final model, MskAge, uses a modified version of this framework, where principal component analysis is calculated after feature selection with the genetic algorithm. We demonstrate that this modified approach, applied to our dataset, improves accuracy compared to calculating principal components of the entire feature space, resulting in a highly accurate epigenetic clock (error +/− 3.51 years) across all musculoskeletal tissues and blood samples in the dataset. Although MskAge is designed to predict chronological age, it resets across multiple independent cellular rejuvenation datasets and strongly correlates with population doubling in vitro. A tutorial for calculating MskAge can be found on Github: https://github.com/laird‐lab/Msk‐Age.

2. Results

2.1. Acquired DNA Methylation Datasets

Infinium 450 k and EPIC methylation array data were acquired from searches of the public domain repository Gene Expression Omnibus (GEO) and novel data generated within the UK Center for Integrated Musculoskeletal Ageing (CIMA). A total of 1048 samples were identified across both platforms. Figure 1 shows an overview of the analysis pipeline for the generation of MskAge.

FIGURE 1.

FIGURE 1

Schematic of analysis workflow for the development of Musculoskeletal Age (MskAge). The schematic depicts the collated samples undergoing quality control and filtering, being merged into a matrix and finally overviews the process of training the genetic algorithm and principal component model.

2.2. Methods of Quantile Normalisation Skew Epigenetic Age Estimates

With a plethora of normalisation methods available for methylation array technologies, the impact of preprocessing on epigenetic age estimates is an important consideration (Aryee et al. 2014; Fortin et al. 2017; Teschendorff et al. 2013). To explore this issue and guide the choice of normalisation for MskAge development, we computed the epigenetic age of 16 cartilage samples with a relatively narrow age range (50–60 years) using different methods. Our analysis revealed that for PhenoAge, GrimAge and Hannum clocks, the Quantile normalisation employed by both Quantile and Funnorm methods significantly affected epigenetic age estimates (Figure S1). This effect is further highlighted by a correlation matrix computed for each respective normalisation method, demonstrating that Quantile normalisation methods produce distinct methylation profiles that cluster independently from non‐quantile methods (Figure S2). These methods lead to significantly exacerbated errors in epigenetic clocks not trained on quantile‐normalised data (FDR < 0.001) (Figure S1). Based on these results, we select single‐sample normal exponential using out‐of‐band probes (ssNoob) as our chosen method of normalisation, but reasoned that with the exception of Quantile and Funnorm, other normalisation methods would be applicable to MskAge.

2.3. Existing Epigenetic Clocks Are Inaccurate in Musculoskeletal Tissues Relative to a Bootstrapped Randomly Samples CpG Model

We assessed the feasibility of using existing epigenetic clocks as age biomarkers in cartilage, bone, mesenchymal stromal cells (MSCs), skeletal muscle and tendon tissue. To draw a relative comparison, we compared errors of existing clocks to a ‘randomly sampled’ model, which ran 2500 elastic net regression models each on 350 randomly sampled CpGs from the training dataset and evaluated the predictions they made on independent test data. The distribution of the errors can be seen in Figure 2A, with a median absolute error of +/− 9.83 years for all tissues indicated by the dashed line. We compared the randomly sampled model errors and errors from existing epigenetic clocks averaged across all musculoskeletal tissues (Figure 2B) and split by each musculoskeletal tissue (Figure 2C). No existing epigenetic clock performs significantly better in terms of error than the randomly sampled model (Table 1). The Horvath clock had comparable performance with no statistically significant difference between errors (p = 0.31), all other epigenetic clocks had significantly greater errors than the randomly sampled model (p < 0.05) (Table 1).

FIGURE 2.

FIGURE 2

A random CpG model outperforms existing epigenetic clocks in musculoskeletal tissues. (A) Histogram displaying the distribution of median absolute errors (MAE) from 2500 iterations of an elastic net regression model containing 350 randomly sampled CpGs evaluated on an independent test dataset; the dashed vertical line represents the median MAE (+/− 9.83 years). (B) Boxplots of absolute age acceleration derived from existing epigenetic clocks and the errors from the randomly sampled models. Boxplots display the combined errors across all musculoskeletal tissues (B) and the specific errors for each musculoskeletal tissue (C). Boxplots display median errors and the interquartile range. Units are expressed in years. (D) Power calculations derived from respective errors of each existing epigenetic clock for a range of given effect sizes 1–2.5 years; the y axis refers to the sample size required to achieve 80% power for each respective effect size.

TABLE 1.

Confidence intervals (CI) and false discovery rate (FDR) from pairwise Student's t‐tests comparing errors from existing epigenetic clocks to a model containing randomly sampled CpGs.

Epigenetic clock 95% CI False discovery rate Significance
Horvath −1.06 to 0.34 0.31 NS
SkinBlood 4.52–6.35 7.7e−28 ****
PhenoAge 27.12–30.37 5.8e−135 ****
GrimAge 4.34–5.84 2.7e−34 ****
Hannum 16.42–19.55 8.6e−78 ****

Note: ****p < 0.0001.

2.4. Errors Associated With Existing Epigenetic Clocks Reduce the Capacity to Detect Perturbations in Epigenetic Ageing in Musculoskeletal Tissues

To further quantify the impact of errors from existing epigenetic clocks, we simulated power calculations for theoretical experiments. Power calculations were simulated based on detecting an effect size difference in that experiment of 1–2.5 years at 80% power and a 5% false positive rate. Using existing epigenetic clocks, we demonstrate that the consequences of high errors diminish the feasibility of designing appropriately powered in vitro experiments across most musculoskeletal tissues tested (Figure 2D). One example is that of MSCs, a commonly used cell type in tissue engineering that has attracted many efforts that attempt to characterise and modulate aspects of biological age (Al‐Azab et al. 2022). Reliably detecting an effect size of 2.5 years in MSCs would require > 200 samples using any of the existing epigenetic clocks.

2.5. Development of a Highly Accurate Principal Component‐Based Clock in Musculoskeletal Tissues

To build MskAge, we collected data from the public domain and within CIMA of 1048 Infinium Methylation 450 k or Infinium Methylation EPIC array samples that comprised multiple musculoskeletal tissues and blood. To eliminate bias, model development was performed on a training set (70%) and evaluation of error on an independent test set (30%). A penalised genetic algorithm islands model was designed to perform feature selection. MskAge was the result of regressing the principal components on the CpGs selected by the genetic algorithm against a transformed version of chronological age. The median absolute error across all tissues in the independent test data was +/− 3.51 years. The R 2 of MskAge regressed against chronological age in the test set was 0.92 (Figure 3A), higher than that of the other clocks tested (Figure S3). With the exception of tendon, which had a median absolute error of 5.52 years, all other tissues displayed a median absolute error of less than 4 years (Figure 3B). Moreover, we demonstrate that MskAge errors are consistent across tissue types when expressed as a percentage of chronological age (Figure S4). The errors for each tissue are displayed in Table S1. Sample size calculations demonstrate that an effect size of 2.5 years can be achieved with less than 60 samples for any of the tissues (Figure S5), a substantial improvement over the comparator clocks (Figure 2D).

FIGURE 3.

FIGURE 3

MskAge is a highly accurate predictor of age in musculoskeletal tissues and blood. (A) Linear regression of predicted epigenetic age generated with MskAge (y axis) vs. chronological age (x axis) on independent test data (n = 307). Points are colored by tissue origin. (B) Boxplot of Median Absolute Errors (MAE) of epigenetic age acceleration and chronological age from predictions generated with MskAge on independent test data (n = 307) split by tissue origin. Boxplots display median errors and the interquartile range. Units are expressed in years.

To assess whether the PCA‐based approach reduced technical noise inherent in methylation data (Sugden et al. 2020), technical reproducibility was assessed for MSKage and Horvath's original multi‐tissue clock. When comparing 450 k and EPIC v1 array technologies, MskAge had a median difference in age estimates of −1.28 years, whereas Horvath had a difference of 3.66 years (Figure 4A). Comparing within array‐technology reproducibility for EPIC v1, MskAge had a median difference in age estimates of −1.32 years, whereas Horvath had a change of 2.43 years (Figure 4B). The tighter age predictions for MSKage demonstrate the increased reliability of MSKage on technical replicates.

FIGURE 4.

FIGURE 4

Evaluation of the performance of MSKage (3365 CpGs) and Horvath's original clock (353 CpGs) on technical replicates as a difference in age estimates. (A) Comparison of age predictions for gDNA from preserved control hip cartilage (n = 6) run on the Infinium 450 k array and EPIC array v1. (B) Comparison of age predictions for gDNA from MSCs (n = 4) run on the same EPICv1 methylation array chip.

To assess whether MSKage is affected by musculoskeletal cell‐type specific methylation, the epigenetic age of MSCs differentiated to osteoblasts, chondrocytes, and tenocytes (Peffers et al. 2016) was assessed (Figure S6). MskAge was not significantly affected by cell‐type shifts (p = 0.872), suggesting that, at least within the musculoskeletal lineage in vitro, MskAge is robust to variations in cellular composition.

2.6. CpGs Used to Construct MskAge Are Significantly Enriched for Terms Related to Skeletal and Mesenchyme Development

To further understand the biological context of methylation changes being captured by MskAge, we performed enrichment analysis for the genes that the CpGs in the model were located within. Using a network‐based approach which connects nodes (significant enrichment terms, FDR < 0.05) via edges (gene‐set similarity), we identify 4 clusters of Gene Ontology based enrichment terms (Figure S7). Notably, the largest cluster contained terms specifically related to mesenchyme and muscle development. We employed the same approach using the Kyoto Encylopedia of Genes and Genomes (KEGG). Significant KEGG terms (FDR < 0.05) included those in cAMP, Hippo, Wnt and calcium signalling pathways and pluripotency of stem cells (Figure S8).

2.7. The Musculoskeletal Clock Is Reset With Cellular Reprogramming of MSCs and Fibroblasts

Given that our outcome variable was a transformed version of chronological age, a question that remains is whether or not MskAge is a highly accurate predictor of chronological age, or whether the biomarker has the capacity to track biological age. We utilised three in vitro models of cellular reprogramming to address this (Frobel et al. 2014; Gill et al. 2022; Ohnuki et al. 2014). We demonstrate that reprogramming of MSCs to iPSC‐MSCs resets MskAge to approximately 0, consistent with the age observed in Embryonic Stem Cells (ESCs) (Figure 5A). Moreover, we observe the same age reduction over a longitudinal time course of fibroblasts being reprogrammed to iPSCs (Figure 5B). Importantly, MskAge was not trained on fibroblasts, but still has the capacity to detect their rejuvenation as they are programmed towards iPSCs. Likewise, for ESCs, MskAge has not observed the specific methylation profile of an ESC but detects methylation in ESCs to have an epigenetic age of 0. Finally, we computed MskAge predictions on fibroblasts that were not fully reprogrammed to iPSCs but instead only reprogrammed until the maturation phase, a point at which the cell still retains its cellular identity and can re‐differentiate (Gill et al. 2022) (Figure 5C). We observe that when fibroblasts are only partially reprogrammed, MskAge does not reset to 0 but instead decreases by approximately 30 years, which is in agreement with effect sizes observed by the Skin and Blood clock (Gill et al. 2022). It is widely accepted that cellular reprogramming is one of the most potent methods of cellular rejuvenation, thus highlighting the capacity for MskAge to track underlying changes in biological ageing.

FIGURE 5.

FIGURE 5

MskAge tracks cellular ageing and rejuvenation in vitro. (A) Boxplot of predicted epigenetic ages generated with MskAge of MSC donors and the same cells transformed to iPSC‐MSCs (Frobel et al. 2014). Epigenetic age predictions were also computed on ESCs from the same dataset as a positive control. (B) Boxplots of predicted epigenetic ages generated with MskAge on human dermal fibroblasts over a longitudinal experiment of reprogramming to iPSCs (Ohnuki et al. 2014). Epigenetic age predictions were also computed on ESCs from the same dataset as a positive control. (C) Boxplots of predicted epigenetic ages generated with MskAge on human dermal fibroblasts following a longitudinal experiment of iPSC reprogramming only to the maturation phase (Gill et al. 2022). (D) Linear regression of predicted epigenetic age generated with MskAge (y axis) versus population doubling of fibroblasts from 6 healthy donors (x axis) (Sturm et al. 2022).

2.8. Musculoskeletal Age Tracks In Vitro Cellular Ageing

We envision that one of the prominent use cases for MskAge will be to facilitate in vitro research on musculoskeletal cells. Thus, a key attribute of such a tool would be that it tracks age‐related methylated changes in vitro as well as it does on ex vivo tissues. We sought to test whether MskAge tracks epigenetic ageing in vitro over multiple cell divisions across a longitudinal in vitro time course. We reprocessed and analysed the Cellular Lifespan multi‐omics longitudinal fibroblast ageing study (Sturm et al. 2022) that includes six healthy control fibroblast cell donors cultured for up to 80 population doublings. We calculated MskAge and extracted epigenetic age predictions of existing clocks on the same data. MskAge acceleration exhibits the strongest linear relationship (R 2 = 0.73, FDR < 2.37e‐34) with population doublings relative to all clocks tested (Table 2), including the fibroblast‐specific SkinBlood clock (R2 = 0.51, FDR < 1.99e‐15) (Figure 5D). According to MskAge in this study, in vitro ageing occurs at an average rate of 0.45 years per population doubling (Table 2). The ability of MskAge to detect in vitro age acceleration as a product of population doubling is advantageous for its use in facilitating the identification of perturbations that can attenuate or accelerate the ageing process in musculoskeletal cells.

TABLE 2.

Output of linear regression models computed from epigenetic age predictions in MskAge and epigenetic clocks in the longitudinal cellular lifespan study.

Clock R 2 p FDR Intercept Slope
MskAge 0.73 2.15e‐35 2.37E‐34 11.64 0.45
Horvath 0.05 0.031 0.124 18.84 0.17
SkinBlood 0.51 1.99e−15 1.79e−14 −7.27 0.33
PhenoAge 0.05 0.033 0.124 3.09 0.2
GrimAge 0.03 0.079 0.158 26.82 −0.04
Hannum 0.66 6.90e−23 6.9e−22 −36.4 0.41
PCHorvath 0.18 2.58e−05 2.0e−4 23.08 0.14
PCSkinBlood 0.11 0.001 0.005 11.37 0.16
PCPhenoAge 0.4 1.25e−11 1.0e−10 59.21 0.27
PCGrimAge 0.02 0.165 0.166 52.75 0.03
PCHannum 0.39 3.77e−11 2.64e−10 36.77 0.22

Abbreviation: PC, principal component version of the clocks.

2.9. Musculoskeletal Age Is Accelerated in Lesioned Osteoarthritic Cartilage

MskAge was trained on non‐lesioned cartilage samples. To test whether MskAge is accelerated in the context of osteoarthiritis (OA) we reprocessed 298 hip and knee non‐OA and OA cartilage samples. Lesioned cartilage exhibits significant Musculoskeletal Age acceleration in the hip (4.029 years, FDR = 0.003) and borderline significant Musculoskeletal Age acceleration in the knee (3.16 years, FDR = 0.067) relative to preserved non‐OA controls (Figure 6A,B; Table S2). The correlation between MSKage and chronological age was also assessed (Figure S9) and was strong in both hip and knee, though it is not optimal given both control and OA samples are present. Interestingly, cartilage extracted from non‐lesioned hip OA patients still exhibits a trend towards being epigenetically older than preserved control hip cartilage (2.056 years, p = 0.094) (Figure 6A; Table S2). While the direction of change is also the same for preserved OA knee cartilage relative to preserved control knee cartilage (1.02 years) the difference is non‐significant (p = 0.231). In contrast, Horvath's clock detected no differences in age acceleration between cartilage samples (Figure S10).

FIGURE 6.

FIGURE 6

MskAge is accelerated in lesioned OA cartilage. (A) Predicted MskAge of Preserved Control (n = 73), Preserved OA (n = 39) and Lesioned OA (n = 26) Hip cartilage samples. (B) Predicted MskAge of Preserved Control (n = 58), Preserved OA (n = 85) and Lesioned OA (n = 17) Knee cartilage samples.

It is noteworthy that no significant differences in MskAge were observed when we compared healthy to osteopenic and osteoporotic bone samples (Figure S11) or with exercise intervention in muscle (Figure S12).

3. Discussion

The musculoskeletal system is crucial to the ageing process, and maintaining its proper function is key for extending both health span and lifespan in humans (Roberts et al. 2016). Although there has been extensive research on the ageing of the musculoskeletal system, few reliable biomarkers are available for musculoskeletal tissues (Kemp et al. 2018). In this study, we introduce MskAge, an accurate principal component‐based epigenetic clock that can be applied ex vivo and in vivo to investigate the epigenetic age of musculoskeletal tissues and cells. MskAge tracks biological age acceleration and reversal, making it a promising tool for studying pro or anti‐ageing interventions in musculoskeletal tissues and cells.

Recent work has emphasised the importance of preprocessing pipelines for epigenetic age prediction (Ori et al. 2022). In light of this, we evaluated important technical considerations during the construction of MskAge by assessing the impact of various array‐based normalisation methods on epigenetic age predictions. A previous study demonstrated that Horvath's age prediction is robust to the choice of data preprocessing methods (McEwen et al. 2018). While our results align with McEwen and colleagues, the same does not hold true for other clocks, particularly when employing quantile‐based normalisation methods for PhenoAge, GrimAge and Hannum. Quantile normalisation methods force data distributions to be equal, resulting in a stringent method of statistical normalisation. Notably, PhenoAge, GrimAge and Hannum clocks did not have quantile normalised data in their training sets, which may account for the exacerbated prediction errors. As a result, we recommend normalising data using ssNoob prior to making predictions with MskAge.

A common debate in the ageing biomarker literature concerns the absolute importance of accuracy when the outcome variable is chronological age (Bell et al. 2019; Zhang et al. 2019). It is indeed valid that outside of forensic settings, establishing a model that provides a perfect readout of chronological age is seldom useful. Zhang and colleagues show experimentally that as a blood‐based epigenetic clock's predictions approach near‐perfect chronological age accuracy, the clock loses its association with mortality risk, which is a proxy of biological age (Zhang et al. 2019). These findings create some subjectivity and uncertainty around the ideal level of accuracy for a biomarker to be useful. To address this question objectively, we used a two‐fold data‐driven approach to assess the applicability of existing epigenetic clocks were applicable to musculoskeletal tissues. First, we created a ‘Random CpG’ model, involving 2500 iterations of randomly selecting 350 CpGs from the 338,185 CpGs in the dataset, training a model on 70% of the data and evaluating its error on the 30% held‐out test set. We selected 350 CpGs because it is approximately equal to the mean number of CpGs in the existing epigenetic clocks tested. Remarkably, the errors from the “Random CpG” model were as accurate or, in most cases, more accurate than epigenetic age predictions using existing clocks on the same samples (Figure 2B,C; Table 1). This outcome underlines the extent to which existing epigenetic clock CpGs lack specificity for musculoskeletal tissues. A plausible explanation for the accuracy of the Random CpG model is that it reflects the degree to which ageing impacts methylation patterns across a substantial portion of the epigenome, which has been reported to be up to 20% (Porter et al. 2021). Indeed, it has been shown that accumulating stochastic variation can predict chronological and biological age (Meyer and Schumacher 2024). Accurate models of chronological age prediction have also been built using only a small number of CpGs, such as those in the genes ELOVL2, FHL2, KLF14 and TRIM59 (Woźniak et al. 2021). Secondly, we simulated power calculations based on the errors of existing epigenetic clocks across a range of theoretical effect sizes for future experiments. We propose that one of the primary applications for an epigenetic biomarker of age in musculoskeletal tissues would be to identify compounds that could perturb the ageing process in vitro. According to the power calculations, designing adequately powered studies using existing epigenetic biomarkers would be largely impractical and unfeasible both logistically and financially. The combined evaluation of both above methods objectively underscores that existing epigenetic clocks are not suitable for quantifying epigenetic ageing in musculoskeletal tissues. As a result, we developed MskAge with the goal of providing substantial improvements in predictive accuracy and musculoskeletal‐tissue specificity compared to other epigenetic clocks. Furthermore, we benchmarked MskAge to show that its enhanced age prediction accuracy did not compromise its ability to detect changes in biological age.

Developing a multi‐tissue epigenetic clock can be viewed as a multi‐objective optimisation problem. DNA methylation serves as a critical factor in determining cellular identity, and as a result, differences in methylation between tissues tend to be more pronounced than the more subtle changes occurring with age (Kim and Costello 2017). The creation of a multi‐tissue epigenetic clock adds a hierarchical structure to a dataset. When age distributions across various tissue types are unbalanced, a naïve model that overlooks tissue origin might inadvertently capture tissue‐specific methylation patterns rather than age‐related changes. MskAge was developed to address this issue by employing a multi‐level genetic algorithm islands model that minimises within‐tissue age prediction errors. Genetic algorithms are meta‐heuristic frameworks that optimise outcomes based on evolutionary principles (Scrucca 2013). We utilised a penalised ridge regression fitness function within the genetic algorithm framework to efficiently search and reduce the feature space. Ridge regression models have been successfully employed as learners in fitness functions for genetic algorithms across various applications (Ahn et al. 2012; Zhang and Horvath 2005). The L2 penalty of the ridge regression cost function causes the coefficient weights of correlated features to shrink towards each other without completely removing them (Friedman et al. 2010). This coefficient shrinkage is desirable in the context of CpG methylation, which exhibits high collinearity. Furthermore, since the coefficient weights are shrunken but not removed, as in lasso or elastic net, ridge regression ensures that the inclusion of CpGs is assigned at some weight, enabling them to be penalised for their inclusion in the model. To penalise the inclusion of CpGs, we assigned each CpG a penalty of 0.01 years in the genetic algorithm, promoting a fitness benefit that encourages convergence towards a model with fewer features, provided that the removal of a specific CpG does not inflate the model error by more than the specified penalty. Rather than a single genetic algorithm, we distributed the evolutionary optimisation task across four independent islands that exchanged information on their best solutions every ten iterations. Utilising multiple islands for the optimisation task enables a more efficient search of the feature space and diminishes the likelihood of converging on local optima (Scrucca 2017; Whitley et al. 2015). After 200 iterations, our ‘fittest’ model, which includes the selection of CpGs optimised by the genetic algorithm to achieve the lowest cross‐validated accuracy, contained 3365 CpGs. Rather than further reducing this selection, we transformed all 3365 CpGs into linear principal components using singular value decomposition (SVD). The use of principal components as a dimensionality reduction method in constructing multivariate models with omics data is not new (Mishra et al. 2011). However, the first attempt using this approach in the development of epigenetic clocks was published recently (Higgins‐Chen et al. 2022). What distinguishes our method from Higgins‐Chen and colleagues is we adapted this approach by performing SVD on M values after extensive variable selection through the genetic algorithm. This approach resulted in an improved test set error of +/− 3.51 years, compared to an error of +/− 5.21 years when creating a principal component model on the full CpG matrix before variable selection.

The advantages of using principal components as features in epigenetic clocks have been discussed previously (Higgins‐Chen et al. 2022). An epigenetic clock constructed on principal components of a larger number of age‐related CpGs improves both inter‐ and intra‐dataset predictions, addressing the significant challenge of array‐related ‘batch to batch’ variability (Higgins‐Chen et al. 2022). Taking such an approach means that MskAge is optimised for data generated from Infinium 450 or Infinium EPIC arrays due to the large amount of CpGs that need to be quantified. Although it is feasible to develop a subset of MskAge coefficients that explain maximum variance with minimal redundancy, this likely would reduce the overall accuracy of the model.

A potential limitation of the application of epigenetic clocks to bulk tissue samples is the heterogeneity in DNA methylation between cell types and the potential for cell‐type composition to change with age (Teschendorff and Horvath 2025). We were not able to apply a cell‐type deconvolution algorithm for MSKage, as comprehensive methylation‐specific cellular reference sets for musculoskeletal tissues are rare. Instead, we demonstrated that MSKage acceleration in MSCs differentiated to multiple musculoskeletal cell types in vitro was similar. However, it remains plausible that age‐related shifts in immune cell populations or in the vascular system could influence readouts from ex vivo tissue samples.

A further improvement in future interactions of epigenetic clocks could include accounting for the non‐linearity of the ageing process (Shen et al. 2024) as reflected in methylation levels (Vershinina et al. 2021). We included an age transformation for samples less than 20 years old, though this term may have a negligible effect as few samples tested here were in this age range. For adult samples, quadratic terms were found to refine clock performance with CpG interaction terms, higher order polynomials, and spline‐based models cited as future directions to enhance clock performance (Bernabeu et al. 2023).

To date, little mechanistic evidence has emerged from epigenetic clock coefficients, and it has often been challenging to infer functional relevance to the ageing process from CpG‐based models containing a few hundred CpGs (Bell et al. 2019). The highly co‐linear nature of the epigenome may inflate the chances of clocks being built on methylation changes that are correlated but not causal to the ageing process. The method employed in this study does not eliminate the inclusion of correlated, non‐causal features. We found that with a large number of correlated CpGs, there was significant enrichment for Gene Ontology processes associated with development. The enriched KEGG pathways identified key signalling pathways, such as Hippo, Calcium and Wnt, which are commonly implicated in both musculoskeletal ageing and disease. Wnt signalling regulates numerous cellular functions and importantly contributes to both bone and cartilage regeneration, and its dysregulation exacerbates osteoarthritis (Houschyar et al. 2018). Additionally, Hippo signalling has been implicated in regulating skeletal muscle mass (Watt et al. 2015). Interestingly, both the significant Gene Ontology and KEGG terms contained modules related to hormonal regulation and Cushing syndrome, respectively. Cushing syndrome is a condition related to the overproduction of cortisol, and hypercortisolism is proposed to induce a premature ageing phenotype (Aulinas et al. 2013). While further research is required to delineate the functional relevance of methylation changes in these pathways, the enrichment of MskAge coefficients in musculoskeletal and ageing‐specific pathways is promising.

We assessed MskAge's capacity to quantify biological ageing methylation changes using cellular reprogramming and longitudinal in vitro ageing. Cellular reprogramming, the gold standard of experimental cellular rejuvenation, reverts somatic cells into an ESC‐like state, reversing many ageing hallmarks and rejuvenating the transcriptome and methylome (Frobel et al. 2014; Gill et al. 2022; Horvath 2013b; Lapasset et al. 2011; Ohnuki et al. 2014). We reanalysed three independent cellular reprogramming datasets, demonstrating that MskAge consistently reports the age of iPSCs and ESCs as zero and tracks the biological reversal of fibroblasts and MSCs during reprogramming, even though some of these cell types were not included in MskAge's training (Frobel et al. 2014; Gill et al. 2022; Ohnuki et al. 2014). It is plausible that by using the principal components of a large number of CpGs (3365), MskAge has the scope to capture cell‐type specific and cell‐type independent features of biological ageing. In contrast to cellular reprogramming, serially passaged primary cells without genetic perturbations undergo systemic changes recapitulating ageing and eventually become senescent (Hayflick and Moorhead 1961). MskAge exhibited a strong linear relationship with population doubling in serially passaged fibroblasts, outperforming other clocks tested, including the SkinBlood clock (Horvath et al. 2018). MskAge's ability to track in vitro ageing is advantageous for musculoskeletal ageing research, which often relies on the extraction and culture of ex vivo cells. Furthermore, MskAge ticks at an approximate rate of 0.45 years per population doubling in primary fibroblasts, enabling researchers to objectively investigate perturbations that may modulate this rate. MskAge constitutes a valuable new tool for musculoskeletal biologists and ageing researchers that is a highly accurate predictor of epigenetic age in musculoskeletal cells and tissues ex vivo and in vitro. Its superior accuracy facilitates the design of experiments with appropriate statistical power compared to existing epigenetic clocks in musculoskeletal tissues. In addition to its precision, MskAge effectively tracks the biological age of cells through well‐known ageing perturbations and provides a means by which serial passaging of cells can be used as a model of epigenetic ageing in vitro. MskAge also facilitates the monitoring of pharmacological, genetic, and nutritional interventions on musculoskeletal cells.

4. Methods

4.1. Data Acquisition

Datasets used for the development of the Musculoskeletal Clock were acquired internally from the Center for Integrated Musculoskeletal Ageing (CIMA) or the public repository Gene Expression Omnibus (GEO). GEO search terms were restricted to data generated on the Infinium Methylation 450 k or Infinium Methylation EPIC array platforms. Sources and descriptions of the acquired data are given in Table S3, and additional information is in Supporting Information methods.

4.2. Data Processing

All data processing was performed in R (4.1.2). Raw IDAT files and raw methylated and unmethylated intensity matrices available for 16/18 datasets were reprocessed using the functionality within the minfi package (Aryee et al. 2014). CpG probes were removed from IDAT files if their methylated and unmethylated signal intensities were not significantly above that of control probes (detection p‐value > 0.01). Probes were further removed if they mapped within 2 base pairs of a single nucleotide polymorphism (SNP), to the X or Y chromosome, or were found to cross‐hybridise to multiple genomic locations (Pidsley et al. 2016). The remaining 2/18 datasets were processed as above, with the exception that signal intensity relative to control probes could not be assessed.

4.3. Epigenetic Age Calculations for Existing Epigenetic Clocks

All epigenetic age calculations for existing epigenetic clocks presented in this manuscript were calculated using the online DNA Methylation Age calculator with the ‘Normalise Data’ option selected (Horvath 2013a). Age calculations for the in vitro fibroblast ageing data were derived directly from the Cellular Lifespan Study shiny application Cellular Lifespan Study (Sturm et al. 2022). It is valuable to note that whilst two iterations of a skeletal muscle clock have been published, they could not be evaluated herein because all of the skeletal muscle samples used in this study were used in the training sets of the two iterations of the skeletal muscle clock (Voisin et al. 2020, 2021).

4.4. Data Normalisation

Following the investigation of various normalisation methods, datasets used in the development of the musculoskeletal clock for which raw IDAT files were available were normalised using single sample normal‐exponential out‐of‐band (ssNoob) normalisation (Fortin et al. 2017). Datasets for which only unnormalised beta values were available were normalised using Beta Mixture Quantile Normalisation (BMIQ) (Teschendorff et al. 2013).

4.5. Data Imputation

Probe‐wise and sample‐wise missingness were assessed by counting row‐wise and column‐wise missing values (NA's). The data was overall complete, with > 99.5% of samples having missing values for < 1% of their CpGs. Values for the CpGs that were missing were imputed using the impute.knn function (Troyanskaya et al. 2001).

4.6. Defining a Random CpG Model of DNAmAge in Musculoskeletal Tissues

To establish a random CpG model, the dataset was split into training (70%) and test data (30%) with respect to tissue origin for the samples, such that 70% of each tissue was used for training within each dataset. A 10‐fold cross‐validated elastic net regression model was trained using 350 CpGs sampled at random without replacement from the 338,185 CpGs that remained after filtering and integrating the data. This process was repeated 2500 times to provide an unbiased distribution of errors (Zou and Hastie 2005). The error was calculated as the absolute Age Acceleration (DNAmAge—Age).

4.7. Determining Power Calculations for Given Effect Sizes

For each respective error of existing clocks in musculoskeletal tissues, power calculations were simulated based on a fixed power of 80% and significance at 5% (p = 0.05). Cohen's d was calculated as the difference in means divided by the pooled standard deviation of DNAm Age Acceleration. Power calculations were simulated for effect sizes of 1–2.5 years in increments of 0.25 years.

4.8. Univariate Statistical Analysis

Univariate statistical tests were computed with either pairwise Student's t‐tests or linear models as described in the table and figure legends. Resulting p values were adjusted using the Benjamini Hochberg FDR method (Benjamini and Hochberg 1995).

4.9. Genetic Algorithm

4.9.1. Training and Test Set Splits

Data was split 70/30 into training and test sets with respect to tissue origin. Descriptions of samples within each set can be seen in Table S5.

4.9.2. Feature Reduction and Selection

To reduce dimensionality by removing potentially redundant features (CpGs) we calculated the Pearson correlation coefficient of each CpG with age across each tissue in the training data and filtered out any CpGs with an absolute Pearson correlation < 0.3. After filtering, 7677 CpGs remained. We employed a penalised genetic algorithm islands model to select CpGs that were most predictive across each musculoskeletal tissue. A description of the algorithm can be found in sections below.

4.9.3. Problem Description

Using the reduced set of CpGs, the aim was to develop a model that optimised the selection of CpGs for the prediction of age across multiple musculoskeletal tissues and blood that also accounted for the hierarchical nature and imbalance between tissues in the training dataset. To do this, we designed a genetic algorithm islands model as follows.

Given a set of training samples Z (Z1, Z2…Zi) with each Zi having a corresponding chronological age Ai, tissue original T (T1, T2…Tk) and an associated vector of M values for CpGs S (S1, S2…Sj). For convenience, we add a superscript index to each sample that denotes the tissue type that the sample belongs to. For example, Zt denotes all samples in Z that belong to tissue T. The aim is to select a combination of S that minimises the error derived from the fitness function of predicting age AT in samples ZT.

4.9.4. Encoding the Fitness Function

To select a subset of S predictive of age across all ZT, each CpG Si is binary encoded to a value of S1 or S0. Here we denote a chromosome C as the binary encoded selection of S for a particular iteration of the genetic algorithm, specifically (S1 in S). To define the fitness (accuracy) of each chromosome C, we build a ridge regression function with lambda parameter defined by 10‐fold cross‐validation in the training set. The ridge regression fitness function can be seen below.

∑i=1Zyi−β0−∑j=1Cβjxij2+λ∑j=1Cβj2

Using the ridge regression coefficients, 10‐fold cross‐validation on the training set Z is used to compute a mean absolute error (MAE) for each musculoskeletal tissue ZT. The MAE for each tissue is then averaged to calculate a single objective fitness value, which is equal to the mean MAE equally weighted for every tissue T in the training set. We penalise the inclusion of CpGs in the model by adding a weight of 0.01 years to the MAE, thus forcing the genetic algorithm to discard individual CpGs whose inclusion in the model does not reduce error by at least this amount.

∑kϵT∑iϵZky^i−yi2ZkT

4.9.5. Evolution of Chromosomes on Independent Islands

The task of minimising the fitness value by the selection of the fittest chromosomes is distributed across four different islands within the genetic algorithm. Initially, a population that contains 200 randomly sampled chromosomes is generated and split equally across eight island nodes. Each island evaluates its population of chromosomes as described above, resulting in 50 fitness values. The top 5% of chromosomes that yield the fittest individuals (i.e., those with the lowest fitness values) in each island are retained from that population. Single‐point crossover is applied to occur at a rate of 70% to the top 5% of fitness individuals in each island. Crossover produces child chromosomes that are the product of two parent chromosomes. Random mutations are permitted to occur in each of the child chromosomes probabilistically at a rate of 10%. The aforementioned process is repeated for 400 iterations, with the migration of the fittest chromosomes between islands every ten iterations. Algorithm parameters are given in Table S6.

4.10. Development of the Final Musculoskeletal Clock

To build the final model, we used the 3365 CpGs selected by the genetic algorithm to compute principal components. We trained an elastic net regression model on these principal components to predict a transformed version of age (Horvath 2013b). Age transformation was performed as follows.

If age < 20:

F(age) = log (age+1)—log(20 + 1).

Else:

F(age) = (age‐20)/(20 + 1).

To assess the accuracy of the model on hold‐out test data and new independent datasets, M values were first projected into the same principal component space as the training data. The resulting principal components were then used to make predictions with the model.

4.10.1. Functional Enrichment

Functional enrichment of the 3365 CpGs was conducted using the clusterProfiler R package (Wu et al. 2021) with all identified genes from our filtered dataset as the background. CpGs were mapped to genes based on the Illumina manifest file. Enrichment was performed using the enrichGO function that performs enrichment of a vector of genes within Gene Ontology terms. Redundant terms were removed using the simplify function with a similarity cutoff of 0.7. Significant gene ontology pathway networks were clustered and visualised using an Enrichment Map. Edges connecting the nodes denote the overlap in genes between the pathways.

Author Contributions

Conceptualisation: Peter D. Clegg, Elizabeth G. Canty‐Laird. Data curation: Daniel C. Green. Formal Analysis: Daniel C. Green. Funding acquisition: Peter D. Clegg, Elizabeth G. Canty‐Laird. Investigation: Daniel C. Green. Methodology: Daniel C. Green. Project administration: Elizabeth G. Canty‐Laird. Resources (data sets): Louise N. Reynard, Sjur Reppe, Kaare Gutvik, Mandy J. Peffers. Software: Daniel C. Green. Supervision: James R. Henstock, Daryl P. Shanley, Peter D. Clegg, Elizabeth G. Canty‐Laird. Visualisation: Daniel C. Green. Writing – original draft: Daniel C. Green. Writing – review and editing: Daniel C. Green, Louise N. Reynard, James R. Henstock, Sjur Reppe, Kaare Gutvik, Mandy J. Peffers, Daryl P. Shanley, Peter D. Clegg, Elizabeth G. Canty‐Laird.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Data S1.

Acknowledgements

We would like the acknowledge Dr. Arturas Grauslys and the Computational Biology Facility at the University of Liverpool for expert statistical discussion during model development stages. Moreover, we would like to acknowledge Dr. Mohammad Saniee Abadeh for advice conceptualising the genetic algorithm approach.

Funding: This work was funded by the Medical Research Council Versus Arthritis Centre for Integrated Research into Musculoskeletal Ageing (CIMA) [MR/R502182/1].

Data Availability Statement

The data are available from the corresponding author upon reasonable request.

References

  1. Ahn, J. J. , Byun H. W., Oh K. J., and Kim T. Y.. 2012. “Using Ridge Regression With Genetic Algorithm to Enhance Real Estate Appraisal Forecasting.” Expert Systems With Applications 39, no. 9: 8369–8379. 10.1016/j.eswa.2012.01.183. [DOI] [Google Scholar]
  2. Al‐Azab, M. , Safi M., Idiiatullina E., Al‐Shaebi F., and Zaky M. Y.. 2022. “Aging of Mesenchymal Stem Cell: Machinery, Markers, and Strategies of Fighting.” Cellular & Molecular Biology Letters 27, no. 1: 69. 10.1186/s11658-022-00366-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Aryee, M. J. , Jaffe A. E., Corrada‐Bravo H., et al. 2014. “Minfi: A Flexible and Comprehensive Bioconductor Package for the Analysis of Infinium DNA Methylation Microarrays.” Bioinformatics 30, no. 10: 1363–1369. 10.1093/bioinformatics/btu049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Aulinas, A. , Santos A., Valassi E., et al. 2013. “Telomeres, Aging and Cushing's Syndrome: Are They Related?” Endocrinología y Nutrición (English Edition) 60, no. 6: 329–335. 10.1016/j.endoen.2012.10.006. [DOI] [PubMed] [Google Scholar]
  5. Bell, C. G. , Lowe R., Adams P. D., et al. 2019. “DNA Methylation Aging Clocks: Challenges and Recommendations.” Genome Biology 20, no. 1: 249. 10.1186/s13059-019-1824-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Benjamini, Y. , and Hochberg Y.. 1995. “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing.” Journal of the Royal Statistical Society: Series B: Methodological 57, no. 1: 289–300. 10.1111/j.2517-6161.1995.tb02031.x. [DOI] [Google Scholar]
  7. Bernabeu, E. , McCartney D. L., Gadd D. A., et al. 2023. “Refining Epigenetic Prediction of Chronological and Biological Age.” Genome Medicine 15, no. 1: 12. 10.1186/s13073-023-01161-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Fortin, J. P. , Triche T. J. Jr., and Hansen K. D.. 2017. “Preprocessing, Normalization and Integration of the Illumina HumanMethylationEPIC Array With Minfi.” Bioinformatics 33, no. 4: 558–560. 10.1093/bioinformatics/btw691. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Friedman, J. , Hastie T., and Tibshirani R.. 2010. “Regularization Paths for Generalized Linear Models via Coordinate Descent.” Journal of Statistical Software 33, no. 1: 1–22. 10.18637/jss.v033.i01. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Frobel, J. , Hemeda H., Lenz M., et al. 2014. “Epigenetic Rejuvenation of Mesenchymal Stromal Cells Derived From Induced Pluripotent Stem Cells.” Stem Cell Reports 3, no. 3: 414–422. 10.1016/j.stemcr.2014.07.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Gill, D. , Parry A., Santos F., et al. 2022. “Multi‐Omic Rejuvenation of Human Cells by Maturation Phase Transient Reprogramming.” eLife 11: e71624. 10.7554/eLife.71624. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Hannum, G. , Guinney J., Zhao L., et al. 2013. “Genome‐Wide Methylation Profiles Reveal Quantitative Views of Human Aging Rates.” Molecular Cell 49, no. 2: 359–367. 10.1016/j.molcel.2012.10.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Hayflick, L. , and Moorhead P. S.. 1961. “The Serial Cultivation of Human Diploid Cell Strains.” Experimental Cell Research 25: 585–621. 10.1016/0014-4827(61)90192-6. [DOI] [PubMed] [Google Scholar]
  14. Higgins‐Chen, A. T. , Thrush K. L., Wang Y., et al. 2022. “A Computational Solution for Bolstering Reliability of Epigenetic Clocks: Implications for Clinical Trials and Longitudinal Tracking.” Nature Aging 2, no. 7: 644–661. 10.1038/s43587-022-00248-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Horvath, S. 2013a. “DNA Methylation Age Calculator.” https://dnamage.genetics.ucla.edu/.
  16. Horvath, S. 2013b. “DNA Methylation Age of Human Tissues and Cell Types.” Genome Biology 14, no. 10: R115. 10.1186/gb-2013-14-10-r115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Horvath, S. , Oshima J., Martin G. M., et al. 2018. “Epigenetic Clock for Skin and Blood Cells Applied to Hutchinson Gilford Progeria Syndrome and Ex Vivo Studies.” Aging (Albany NY) 10, no. 7: 1758–1775. 10.18632/aging.101508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Houschyar, K. S. , Tapking C., Borrelli M. R., et al. 2018. “Wnt Pathway in Bone Repair and Regeneration ‐ What Do we Know So Far.” Frontiers in Cell and Developmental Biology 6: 170. 10.3389/fcell.2018.00170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Kemp, G. J. , Birrell F., Clegg P. D., et al. 2018. “Developing a Toolkit for the Assessment and Monitoring of Musculoskeletal Ageing.” Age and Ageing 47, no. suppl_4: iv1–iv19. 10.1093/ageing/afy143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Kim, M. , and Costello J.. 2017. “DNA Methylation: An Epigenetic Mark of Cellular Memory.” Experimental & Molecular Medicine 49, no. 4: e322. 10.1038/emm.2017.10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Kirkwood, T. B. 2005. “Understanding the Odd Science of Aging.” Cell 120, no. 4: 437–447. 10.1016/j.cell.2005.01.027. [DOI] [PubMed] [Google Scholar]
  22. Lapasset, L. , Milhavet O., Prieur A., et al. 2011. “Rejuvenating Senescent and Centenarian Human Cells by Reprogramming Through the Pluripotent State.” Genes & Development 25, no. 21: 2248–2253. 10.1101/gad.173922.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Levine, M. E. , Lu A. T., Quach A., et al. 2018. “An Epigenetic Biomarker of Aging for Lifespan and Healthspan.” Aging (Albany NY) 10, no. 4: 573–591. 10.18632/aging.101414. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Lu, A. T. , Quach A., Wilson J. G., et al. 2019. “DNA Methylation GrimAge Strongly Predicts Lifespan and Healthspan.” Aging (Albany NY) 11, no. 2: 303–327. 10.18632/aging.101684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. McEwen, L. M. , Jones M. J., Lin D. T. S., et al. 2018. “Systematic Evaluation of DNA Methylation Age Estimation With Common Preprocessing Methods and the Infinium MethylationEPIC BeadChip Array.” Clinical Epigenetics 10, no. 1: 123. 10.1186/s13148-018-0556-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Meyer, D. H. , and Schumacher B.. 2024. “Aging Clocks Based on Accumulating Stochastic Variation.” Nature Aging 4, no. 6: 871–885. 10.1038/s43587-024-00619-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Mishra, D. , Dash R., Rath A. K., and Acharya M.. 2011. “Feature Selection in Gene Expression Data Using Principal Component Analysis and Rough Set Theory.” Advances in Experimental Medicine and Biology 696: 91–100. 10.1007/978-1-4419-7046-6_10. [DOI] [PubMed] [Google Scholar]
  28. Ohnuki, M. , Tanabe K., Sutou K., et al. 2014. “Dynamic Regulation of Human Endogenous Retroviruses Mediates Factor‐Induced Reprogramming and Differentiation Potential.” Proceedings of the National Academy of Sciences of the United States of America 111, no. 34: 12426–12431. 10.1073/pnas.1413299111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Ori, A. P. S. , Lu A. T., Horvath S., and Ophoff R. A.. 2022. “Significant Variation in the Performance of DNA Methylation Predictors Across Data Preprocessing and Normalization Strategies.” Genome Biology 23, no. 1: 225. 10.1186/s13059-022-02793-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Peffers, M. J. , Goljanek‐Whysall K., Collins J., et al. 2016. “Decoding the Regulatory Landscape of Ageing in Musculoskeletal Engineered Tissues Using Genome‐Wide DNA Methylation and RNASeq.” PLoS One 11, no. 8: e0160517. 10.1371/journal.pone.0160517. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Pidsley, R. , Zotenko E., Peters T. J., et al. 2016. “Critical Evaluation of the Illumina MethylationEPIC BeadChip Microarray for Whole‐Genome DNA Methylation Profiling.” Genome Biology 17, no. 1: 208. 10.1186/s13059-016-1066-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Porter, H. L. , Brown C. A., Roopnarinesingh X., et al. 2021. “Many Chronological Aging Clocks Can Be Found Throughout the Epigenome: Implications for Quantifying Biological Aging.” Aging Cell 20, no. 11: e13492. 10.1111/acel.13492. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Quach, A. , Levine M. E., Tanaka T., et al. 2017. “Epigenetic Clock Analysis of Diet, Exercise, Education, and Lifestyle Factors.” Aging (Albany NY) 9, no. 2: 419–446. 10.18632/aging.101168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Roberts, S. , Colombier P., Sowman A., et al. 2016. “Ageing in the Musculoskeletal System.” Acta Orthopaedica 87, no. sup363: 15–25. 10.1080/17453674.2016.1244750. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Scrucca, L. 2013. “GA: A Package for Genetic Algorithms in R.” Journal of Statistical Software 53, no. 4: 1–37. 10.18637/jss.v053.i04. [DOI] [Google Scholar]
  36. Scrucca, L. 2017. “On Some Extensions to GA Package: Hybrid Optimisation, Parallelisation and Islands Evolution.” R Journal 9: 187. 10.32614/RJ-2017-008. [DOI] [Google Scholar]
  37. Shen, X. , Wang C., Zhou X., et al. 2024. “Nonlinear Dynamics of Multi‐Omics Profiles During Human Aging.” Nature Aging 4, no. 11: 1619–1634. 10.1038/s43587-024-00692-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Sturm, G. , Monzel A. S., Karan K. R., et al. 2022. “A Multi‐Omics Longitudinal Aging Dataset in Primary Human Fibroblasts With Mitochondrial Perturbations.” Scientific Data 9, no. 1: 751. 10.1038/s41597-022-01852-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Sugden, K. , Hannon E. J., Arseneault L., et al. 2020. “Patterns of Reliability: Assessing the Reproducibility and Integrity of DNA Methylation Measurement.” Patterns (New York, N.Y.) 1, no. 2: 100014. 10.1016/j.patter.2020.100014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Teschendorff, A. E. , and Horvath S.. 2025. “Epigenetic Ageing Clocks: Statistical Methods and Emerging Computational Challenges.” Nature Reviews. Genetics 26: 350–368. 10.1038/s41576-024-00807-w. [DOI] [PubMed] [Google Scholar]
  41. Teschendorff, A. E. , Marabita F., Lechner M., et al. 2013. “A Beta‐Mixture Quantile Normalization Method for Correcting Probe Design Bias in Illumina Infinium 450 k DNA Methylation Data.” Bioinformatics 29, no. 2: 189–196. 10.1093/bioinformatics/bts680. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Troyanskaya, O. , Cantor M., Sherlock G., et al. 2001. “Missing Value Estimation Methods for DNA Microarrays.” Bioinformatics 17, no. 6: 520–525. 10.1093/bioinformatics/17.6.520. [DOI] [PubMed] [Google Scholar]
  43. Vershinina, O. , Bacalini M. G., Zaikin A., Franceschi C., and Ivanchenko M.. 2021. “Disentangling Age‐Dependent DNA Methylation: Deterministic, Stochastic, and Nonlinear.” Scientific Reports 11, no. 1: 9201. 10.1038/s41598-021-88504-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Voisin, S. , Harvey N. R., Haupt L. M., et al. 2020. “An Epigenetic Clock for Human Skeletal Muscle.” Journal of Cachexia, Sarcopenia and Muscle 11, no. 4: 887–898. 10.1002/jcsm.12556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Voisin, S. , Jacques M., Landen S., et al. 2021. “Meta‐Analysis of Genome‐Wide DNA Methylation and Integrative Omics of Age in Human Skeletal Muscle.” Journal of Cachexia, Sarcopenia and Muscle 12, no. 4: 1064–1078. 10.1002/jcsm.12741. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Watt, K. I. , Turner B. J., Hagg A., et al. 2015. “The Hippo Pathway Effector YAP Is a Critical Regulator of Skeletal Muscle Fibre Size.” Nature Communications 6: 7048. 10.1038/ncomms7048. [DOI] [PubMed] [Google Scholar]
  47. Whitley, L. D. , Rana S. B., and Heckendorn R. B.. 2015. “The Island Model Genetic Algorithm: On Separability, Population Size and Convergence.” Paper Presented at the International Conference on Computer and Information Technology.
  48. Woźniak, A. , Heidegger A., Piniewska‐Róg D., et al. 2021. “Development of the VISAGE Enhanced Tool and Statistical Models for Epigenetic Age Estimation in Blood, Buccal Cells and Bones.” Aging (Albany NY) 13, no. 5: 6459–6484. 10.18632/aging.202783. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Wu, T. , Hu E., Xu S., et al. 2021. “clusterProfiler 4.0: A Universal Enrichment Tool for Interpreting Omics Data.” Innovation (Cambridge, Mass.) 2, no. 3: 100141. 10.1016/j.xinn.2021.100141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Wu, X. , and Zhang Y.. 2017. “TET‐Mediated Active DNA Demethylation: Mechanism, Function and Beyond.” Nature Reviews. Genetics 18, no. 9: 517–534. 10.1038/nrg.2017.33. [DOI] [PubMed] [Google Scholar]
  51. Xia, X. , Wang Y., Yu Z., Chen J., and Han J. J.. 2021. “Assessing the Rate of Aging to Monitor Aging Itself.” Ageing Research Reviews 69: 101350. 10.1016/j.arr.2021.101350. [DOI] [PubMed] [Google Scholar]
  52. Zhang, B. , and Horvath S.. 2005. “Ridge Regression Based Hybrid Genetic Algorithms for Multi‐Locus Quantitative Trait Mapping.” International Journal of Bioinformatics Research and Applications 1, no. 3: 261–272. 10.1504/ijbra.2005.007905. [DOI] [PubMed] [Google Scholar]
  53. Zhang, Q. , Vallerga C. L., Walker R. M., et al. 2019. “Improved Precision of Epigenetic Clock Estimates Across Tissues and Its Implication for Biological Ageing.” Genome Medicine 11, no. 1: 54. 10.1186/s13073-019-0667-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Zou, H. , and Hastie T.. 2005. “Regularization and Variable Selection via the Elastic Net.” Journal of the Royal Statistical Society, Series B: Statistical Methodology 67, no. 2: 301–320. 10.1111/j.1467-9868.2005.00503.x. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1.

Data Availability Statement

The data are available from the corresponding author upon reasonable request.


Articles from Aging Cell are provided here courtesy of Wiley

RESOURCES