Abstract
Background
Frailty is an important factor in human aging associated with a broad range of adverse outcomes. Frailty metrics are time intensive to collect making them difficult for larger scale application.
Methods
We apply machine learning to predict these frailty metrics, associated risk factors, and adverse outcomes from activity data. We use activity data collected using Actigraphy wearable accelerometer sensors, which are devices that measure acceleration along three axes of movement. Models were evaluated using Area Under the receiver operator Curve (AUC), Area Under Precision Recall Curve (AUPRC), Spearman rank test, Mann-Whitney U test, or Kruskal-Wallis test on repeated subsampling of train and test sets. All statistical tests are reported using -log10(P-value).
Results
Machine learning models show strong predictive performance even with small amounts of accelerometry data available. They are also able to better determine adverse outcomes such as hospitalization and mortality than frailty metrics themselves in our geriatric population.
Conclusions
This approach of wearable activity data-based prediction of frailty offers a surrogate (proxy or estimate) for determining frailty metrics in a scalable manner. It can also be used to determine adverse outcomes such as hospitalizations and mortality, allowing frailty to be used as a metric in other studies or medical practices.
Subject terms: Predictive markers, Computational biology and bioinformatics
Plain language summary
Frailty occurs during human aging and is associated with a broad range of unfavourable outcomes. Frailty is measured using various scores but these often rely on subjective information, are labor intensive to measure, and are not assessed over time. This work presents objective measures indicative of frailty, based on wearable sensors that measure movement. This was tested in a group of people with an average age of 75. Application of a computational model using this data enabled long-term outcomes, including hospitalization and death, to be more accurately predicted than using existing frailty measures. This work demonstrates that 48 hours of data collection per patient is sufficient. This type of system could be used on a larger number of people, enabling those at risk of unfavourable outcomes to be targeted with medical interventions or support.
Culos, Manas et al. construct machine learning surrogate frailty metrics from activity data collected in a nominally intrusive manner. They show that a limited amount of activity data is necessary to model frailty metrics allowing for an increased proliferation of their application along with a robust source of data for other age-related outcomes.
Introduction
According to World Population Prospects 20191, there are 703 million people aged 65 and over in the world, while the number of older people is projected to double to 1.5 billion in 2050. The most problematic manifestation of aging societies is the clinical condition of frailty2,3. Campbell and Buchner defined the term of frailty in 19974 as ‘a condition or syndrome which results from a multi-system reduction in reserve capacity to the extent that a number of physiological systems are close to, or past, the threshold of symptomatic clinical failure. As a consequence, the frail person is at increased risk of disability and death from minor external stresses’. Frailty has since arisen as a priority in daily clinical practice and in the care of older adults5, thus numerous frailty measures or scales have been proposed. Although there is no accepted gold standard frailty metric, there are two widespread approaches, the frailty phenotype6 model and the frailty index of cumulative deficits7. The former is used in daily clinical practice and is based on a physical function model, while the latter extensively evaluates possible health deficits to quantify the cumulative effect7. Despite widespread use, these two scales have been criticized for their definition, methodology, risk classification, and ability to predict adverse events8,9. Overcoming some of the limitations of these frailty measures a new operative scale named Frailty Trait Scale (FTS), based on characteristics of the biological trait of frailty syndrome has been developed10. The FTS incorporates new and relevant frailty components according to recent findings on frailty pathophysiology and has been suggested as a more sensitive instrument for detecting changes in the individual’s biological status than the previously validated scales10. Authors also designed a short form derived from the FTS, that can be used in clinical settings in daily practice and detect changes in frailty status after an intervention (FTS5)11. Despite the development and use of all these scales, the implementation of frailty assessment in clinical settings remains a challenge5,12.
Regular physical activity has been shown to protect against the condition of frailty and reflect an individual’s frailty13,14. Wearable device derived measures15,16 of basal activity levels can then inform us about physical characteristics predictive of frailty or other outcomes of interest. Previous studies have associated activity with individual health characteristics17, classes of frailty18–20, or biological age21,22. However, models directly estimating multiple frailty metrics, to our knowledge, have not yet been developed.
This study presents Extreme Gradient Boosted (XGBoost) Machine Learning (ML) models that leverage wearable accelerography technologies to predict frailty metrics, risk factors, and long term outcomes in a geriatric population (Supplemental Fig. 1) from the Toledo Study for Healthy Aging (TSHA). We show models accurately reproduce multiple frailty metrics using a limited window of data collection. Furthermore, activity based models for identifying adverse outcomes, like hospitalization and mortality, often outperform traditional frailty metrics. These findings show that activity data can serve as a robust, objective, and scalable surrogate for assessing frailty in clinical and research settings while also empowering other clinically relevant predictive models.
Methods
Measurement of frailty
Participants were classified using 4 different frailty scales: (1) Fried Frailty Phenotype (FP) scale; (2) Rockwood Frailty Index (FI), (3) Frailty Trait Scale (FTS), and (4) the Frailty Trait Scale-short form (FTS5). In the FP scale, weakness and slowness were measured and low energy, low physical activity, and loss of weight were reported. Each item was assigned a point6. According to this, participants were classified as robust (0 points), prefrail (1‒2 points), and frail (>3 points). The FI is a count of 40 clinical deficits, including the presence and severity of current disease, ability in activities of daily living, and physical and neurologic signs of clinical examinations23. The presence of each item is scored as a 1. The FTS includes 7 dimensions (balance of energy nutrition, physical activity, nervous system, vascular system, strength, endurance, and walking speed). Each item is scored from 0 (best) to 4 (worst), except for the “chair test,” which scores from 0 to 5 points10. The FTS5 is a short scale that is easy to administer and has a similar performance to the FTS, and it can be used in clinical settings for frailty diagnosis and evolution11. The 5-item (FTS5) (range 0-50) provides an opportunity for tracking trajectories inside each category (frail or non-frail) and between categories, supporting the fact that frailty is not a “categorical” issue but a continuous, discrete one.
Study design and participants
Data were drawn from the Toledo Study for Healthy Aging. The full methodology of the TSHA has been previously described24. Briefly, the TSHA is a population prospective cohort study originally conceived to examine the determinants and consequences of aging and frailty in individuals older than 65 years from Toledo, Spain (Table 1). In the first stage a team of trained psychologists conducted computer assisted face to face interviews with subjects. In the second stage, a physical examination followed by clinical and performance tests at the subject’s home was performed by experienced nurses. In the third stage, participants went to their health center to provide a blood sample in a fasted state. In the fourth stage, anthropometry data and body composition were obtained with dual energy X-ray absorptiometry (DXA). At this later stage, participants were asked to wear an accelerometer for a week (Supplemental Fig. 2). The current study used data from participants with valid accelerometer records from the second (2012 to 2014) and third (2015 to 2017) TSHA waves25. Signed informed consent was obtained from all participants. The study and subsequent analysis of collected data was approved by the Clinical Research Ethics Committee of the Toledo Hospital (approval code: 2010/93) and was conducted in accordance with the Declaration of Helsinki for human studies. As this approval includes the analysis of the deidentified TSHA data no additional IRB approval was necessary.
Table 1.
Summary Statistics of TSHA Cohort: mean and standard deviation of TSHA cohort (n = 437) separated by gender and admitted wave. Only wave two and three included actigraphy measurements
| Cohort Wave | n | Age | Weight (Kg) | Height (cm) | BMI | Waist Circumference (cm) | Hip Circumference (cm) | |
|---|---|---|---|---|---|---|---|---|
| Male | 2 | 85 | 77.48 (5.06) | 77.24 (10.91) | 163.19 (7.55) | 28.98 (3.36) | 98.93 (9.62) | 103.39 (7.77) |
| 3 | 121 | 74.50 (4.90) | 77.50 (12.42) | 163.86 (6.26) | 28.86 (4.40) | 103.09 (12.08) | 103.62 (8.57) | |
| Female | 2 | 104 | 76.38 (4.10) | 69.79 (11.78) | 150.11 (6.53) | 30.97 (4.88) | 91.09 (11.87) | 107.84 (10.44) |
| 3 | 127 | 75.35 (5.41) | 69.80 (12.08) | 151.02 (5.09) | 30.59 (5.00) | 95.49 (11.21) | 109.57 (10.06) |
Accelerometry Data Processing
Physical activity measured with ActiGraph smartwatch (ActiGraph, LLC, FL, USA) was processed to extract features representing various aspects of daily physical activity (e.g., step count, vigorous bouts of activity, and sedentary bouts.)26. ActiGraph’s actigraphy platform ActiLife v6.13.3 was used to extract features which were later used to build predictive models of frailty. These devices and data types have proven to be efficacious when compared against laboratory tests of physical activity but have shown a tendency to underestimate total wake time and sleep onset but is generally considered a good choice without any clear superior option27–30.
Extraction of sleep and activity features is detailed in the associated ActiLife v6.13.3 documentation. Exact features including names and values can be located in the included github repository (see Code & Data Availability section in supplemental) “data/actigraph_export/*” including extracted data and the variables used in their generation.
Predictive Models
The ML algorithm used for all supervised learning tasks were gradient boosted decision trees via the XGboost package V1.5.2.131 in R version 4.1.2. This tree based method produces multiple classification or regression trees which are fit and then added to an ensemble. Specifically, individual trees are added to an ensemble as weak learners to improve the shortcomings of the previous weak learners, which is based on the gradient boosting framework by32,33. As the primary objective of this study was to demonstrate the potential of ML predictors for frailty XGboost was selected for its history of good performance, ease of application for subject matter experts, and ability to handle mixed and missing data.
Individual trees are represented as functions, that exist in the space of all possible classification and regression trees . All trees are then summed to form the final predictions on some observation .
This differs from traditional random forest models through the additive training procedure used as opposed to the trees being independently generated. Broken up into distinct steps the process of building an XGboost model can be seen as the following:
Such models with 50 classification or regression trees were built for all predictions made throughout this work; all other parameters were set to the default of the R package.
Model Robustness and Evaluation
A repeated subsampling holdout method, without replacement and stratification, was used to evaluate the model34. The data was randomly divided into two equal sized subsets for training and testing using a uniform random sampler. Models were trained on one half and tested on the other to ensure that the predictions are made on patients previously blinded to the algorithm. This process was then repeated 50 times to test the robustness of the models. This stringent process of predicting on previously unseen data shows if results are generalizable to future populations through variation in train and test splits. Results generated in this manner are similar to bootstrap estimates (difference being sampling with or without replacement) of model performance. For the exact process of model training, prediction, and evaluation see the script “actigraphy-frailty/scripts/analysis.R” (Code & Data Availability section in supplemental).
These test set predictions were then averaged for final evaluation and presented using a Spearman rank, Mann-Whitney U, or Kruskal-Wallis test: Application of these statistical tests were dependent on the target used for ML modeling being continuous, binary classes, or multiple classes respectively. Supplementary data files 1-8 contain model evaluation of risks, outcomes, frailty metrics, and components of the frailty trait scales as the prediction target; Including evaluation for each of the 50 repeated model predictions (Supplementary data files 1,2,4 and 8) and the average model predictions across repeats (Supplementary data files 3, 5, 6, and 7).
Statistics and reproducibility
Statistical analyses were performed on a total cohort of 437 participants using R version 4.1.2. A repeated subsampling holdout method, without replacement and stratification, was used to evaluate the model34. That is, the dataset was randomly divided into two equal sized subsets for training and testing (50/50 split) to ensure that the predictions are made on patients previously blinded to the algorithm. This procedure was repeated for 50 independent replicates to test the robustness of the models. This stringent process of repeated predicting on previously unseen data shows if results are generalizable to future populations through variation in train and test splits. As models were specified a priori using default parameters no hyperparameter tuning was conducted and no validation set was required. Generalization error was therefore estimated directly using the repeated train/test procedure35. Results generated in this manner are similar to bootstrap estimates (difference being sampling with or without replacement) of model performance. For the exact process of model training, prediction, and evaluation see the script “actigraphy-frailty/scripts/analysis.R” (Code & Data Availability section in supplemental).
Test set predictions were averaged across the 50 replicates for final evaluation. Model evaluation was performed using the Spearman rank, Mann-Whitney U, or Kruskal-Wallis test: Application of these statistical tests were dependent on the target used for ML modeling being continuous, binary classes, or multiple classes respectively. All statistical tests are two-sided and reported as either -log10(P-value) or their associated rho value. For correlation network visualizations, pairwise correlations were calculated, and p-values were adjusted for multiple hypothesis testing using the Bonferroni correction method. Since each predictive model is an independent hypothesis trained and evaluated on only one outcome it does not form a single family of tests, therefore multiple hypothesis correction is unnecessary.
Correlation network visualizations
All datasets were visualized using correlation network structures. Each actigraphy feature is represented by a node and the network layout is determined by the t-SNE algorithm applied to the complete correlation network36. All correlation p-values are adjusted for multiple hypothesis testing using the Bonferroni correction method.
Results
XGBoost models accurately predicted FTS on test data (8.70×10−36 Spearman p-value and 0.55 Spearman rho) (Fig. 1); Similar analyses were done on Fried Score, Rockwood Index, FTS5 and FTS5 based frailty class (Supplemental Results and Supplemental Fig. 3, 4, 5, 6 and 7). Restricting training data to days of the week showed model performance was related to the day of measurement (Fig. 2a). Averaging data from each hour of the day across days exhibited a similar model behavior (Fig. 2b). Increasing the amount of data aggregated in one hour increments showed models built on 48 hours of data produced similar results to those built on all available data (Fig. 2c). Data aggregated from consecutive days showed a 48 hour window of collection sufficiently effective, with a slight improvement when highly predictive days aggregated together (Fig. 2d). Suggesting that two days of data collected in a targeted fashion are highly effective at determining FTS.
Fig. 1. Actigraphy Data used to predict FTS.
Model features, shown as t-SNE of pairwise correlation, with higher predictive importance visualized by larger nodes with a darker purple colour. Correlations are calculated on the 50 model repeats a. Model predictions of FTS on unseen data compared to actual measurements b. Comparison of actigraphy predicted FTS values vs actual FTS values as predictors of risk factors. Correlations are calculated on the entire cohort n = 437 for all 69 risk factors c.
Fig. 2. Daily and Hourly importance to prediction of FTS.
Model performance when data is collected at specific days for the FTS measure a. Similarly, averaging activity from a specific hour across each day shows how specific hours can determine frailty of individuals b. Through incremental inclusion of actigraphy data by the hour we can see the relationship between additional data and predictive performance c. Given the similarity in performance between the first 48 hours and the total data model aggregation of two consecutive days was performed to investigate performance given a more targeted data sampling strategy d. All statistical tests were two-sided Spearman’s Correlation. All results are reported on the full n = 437 cohort.
Additional ML models of various age related risk factors including the Mean Corpuscular Hemoglobin Concentration (MCHC), Maximum Corpuscular Volume (MVC), and grip strength (Fig. 3a) produced robust predictions. Comparing predictive capabilities of FTS with actigraphy derived ML models showed the ML approach superior for a large number of risk factors (Fig. 3b). Actigraphy predicted FTS and FTS5 as predictors of risk factors perform similarly (Fig. 1c & Supplemental Fig. 4c). Ultimately, actigraphy-based frailty predictions recapitulate frailty metrics, their associations with risk factors, and generally enhance predictions through direct analysis. Additionally, activity based prediction of long term outcomes, specifically hospitalization37 and mortality, have an observable improvement over FTS and FTS5 as predictors (Fig. 3c) once again showing the malleability of our analytic approach.
Fig. 3. Actigraphy models reproduce frailty scores and predict mortality and hospital admission.
The correlation network displays actigraphies predictive ability where node size and intensity of colour indicate p-value of model prediction (two-sided Spearman) a. Comparison of the 113 actigraphy risk predictions, and FTS as a predictor of risks via -log10 two-sided Spearman correlation p-values b. Predictive models for hospitalization c and mortality d, compare the predictive capabilities of actigraphy data with XGBoost (visualized curves) to FTS and FTS5. All AUC and correlation results are reported on the full n = 437 cohort.
Discussion
Frailty is a clinically relevant syndrome in the elderly population associated with increased risk for falls, incident disability, hospitalization, and mortality6,38–41. Consequently, reducing the burden of frailty is one of the most essential challenges faced by public health authorities12 as the global population ages. However, current methods developed for assessing frailty have not been widely adopted especially in clinical settings and often do not adequately reflect the biological condition of frailty. Wearable technologies present a unique solution which can assess frailty, risk factors, and long term health outcomes simultaneously.
This work demonstrates the feasibility and potential clinical application of wearable sensors and machine learning models to predict patient frailty, age related risk factors, and long term outcomes. Previous work using data derived from wearable devices to predict frailty have focused on frailty classes (frail vs non-frail)18–20 which can obfuscate variations of frailty within a given class, whereas this work directly estimates multiple frailty metrics. Additionally, studies using patient activity to determine mortality and hospitalization often focus on specific causes of mortality such as heart failure, or on specific at risk populations37,42–44, this work makes no distinction based on cause. The use of actigraphy to predict frailty, risk factors, and long term outcomes presented here offers a unified approach for determining characteristics of health and potential adverse events. The broad scope of targets presented show the capabilities and limitations of actigraphy data, particularly in its ability to characterize and assess frailty.
Using observational patient activity data, as opposed to tests of fitness or self reported data, allows for an easy to collect, objective, scalable, and biologically relevant evaluation of frailty metrics as well as an abundant data source for ML models. ML models to predict Fried score, Rockwood index, FTS, and FTS5 have varied performance and demonstrate how each metric relates to a patient’s physiological capacity for movement (Supplemental Fig. 2). Such models need a small amount of data to sufficiently recreate FTS values allowing for easy inclusion of frailty metrics in resource and time limited settings. Additional analysis of age related risk factors and long-term outcomes as targets show the versatility of the actigraphy measurements in multiple predictive settings (Fig. 3). Like frailty, predictive quality for risk factors and long term outcomes relate to an individual’s activity levels, with risk factors and outcomes such as MVC, MCHC (Fig. 3b), hospitalization, and mortality (Fig. 3c, d) being highly dependent on physiological state.
Some limitations of this work include the TSHA cohort, as ethnicity and socioeconomic status are not included impacting its generalizability. While the XGBoost method proved effective, other methods are worth considering for future work both modern over parameterized deep learning and traditional approaches. Deep neural networks for instance are a powerful ML paradigm with standout performance in numerous problem settings but offer little post-hoc inferential ability and require vast amounts of data. Alternatively, traditional time to event models such as Cox regression offer more interpretable results potentially at the cost of predictive ability. Specific considerations towards this would be that non-linearities need explicit specification in Cox regression, the handling of high dimensional data (important for incorporation of biological modalities), and data availability as other methods require significantly more. Future exploration of other analytical approaches between canonical time-to-event models, other statistical ML models, and deep neural networks would elucidate which approach is most appropriate. Appropriateness being a balance of the model’s ability to be interpreted and performance. Canonical models often allow for direct interpretation, XGboost like other ML models can be interpreted through application of additional methods like Shapley additive explanations, while deep learning approaches often require more experimental methods of interpretation like sparse autoencoders45,46. Finally, the lack of a consensus frailty metric limits the applicability of our model as some features of frailty are not reflected in activity measures but offer an important perspective; Such as the Romberg Balance Test (Supplemental Fig. 8). Ideally, a comprehensive approach incorporating physiological, biological, and clinical features is essential for developing a robust model of frailty and aging.
While determination of frailty, risk factors, and long term outcomes from actigraphy data has direct clinical applications47, further connection to biology is needed to better understand the biological determinants and consequences of aging. The immunological status of patients has been previously associated with actigraphy as a surrogate of recovery and a similar systems biology approach to the study of frailty and its metrics would elucidate the relationship between cellular function and frailty48. Additional biological modalities49 would allow for the investigation of domain specific frailty biomarkers and potential modifiable targets to relieve the burden of frailty and aging. The Stanford 1000 immunomes project and UCSF 10,000 immunome project are exemplary of such a comprehensive project, with the former being of particular interest due to its focus on aging50,51. Ultimately the integration of wearable technologies with multiple biological modalities would provide a robust, scalable, and objective perspective on frailty and a broad range of age-related characteristics.
Supplementary information
Description of Additional Supplementary files
Acknowledgements
We acknowledge funding from the NIH R35GM138353 (to N.A.).
Author contributions
A.C.—Conceptualisation, Methodology, Analysis Design, Analysis, Writing, Visualization, Data Curation. A.M.—Conceptualisation, methodology, Data Acquisition, Writing. K.S., F.J.G.-G., J.L.-R., L.M.A., and L.R.-M.—Writing, Data Interpretation. A.L.C., C.E., D.D.F., T.P., M.B., M.X., N.G.R., R.F., B.G., and M.A.—Review and Editing, Analysis Design. I.A., and N.A.—Conceptualisation, Review and Editing, Supervising, Project Administrating
Peer review
Peer review information
Communications Medicine thanks Björn Friedrich and Emanuele Seminerio for their contribution to the peer review of this work.
Data availability
Requests for original data can be handled through (www.ciberfes.es/) via a review committee. The hospital reviews and determines the purposes for the data requests and what data can be released. Data requests can be sent to: Research and teaching unit, Virgen del Valle Hospital Ctra. Cobisa S/N, 45071 Toledo – Spain, info@estudiotoledo.com. All procedures were approved by the Clinical Research Ethics Committee of the Toledo Hospital and were conducted in accordance with the Declaration of Helsinki for human studies.
Supplemental data can be accessed either as the included.csv files or as.rda files available at the github repository in the results folder https://github.com/tripodlaboratories/actigraphy-frailty/tree/main/results52. The included supplementary data files 1, 2, 4 and 8 contain results (rho, p-values, and RMSE) for each individual model repeat of FTS components, risk factors & outcomes, FTS5 components, and the various frailty metrics respectively. Supplementary Data 3, 5, 6 and 7 contain the average model repeat prediction evaluations for risk factors & outcomes, FTS5 components, FTS components, and the various frailty metrics respectively. Actigraphy activity data and descriptions of clinical outcomes, risks, and frailty components are provided as.csv’s in the data folder of the associated github https://github.com/tripodlaboratories/actigraphy-frailty/tree/main/data52.
Code availability
Code necessary to reproduce results, figures, and models can be found at https://github.com/tripodlaboratories/actigraphy-frailty52. This code was run using R 4.1.2 and xgboost 1.5.2.1 on a macOS 12.3.1 system. There are no restrictions on its access or use.
Competing interests
All authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Anthony Culos, Asier Manas.
These authors jointly supervised this work: Ignacio Ara, Nima Aghaeepour.
Supplementary information
The online version contains supplementary material available at 10.1038/s43856-026-01419-7.
References
- 1.Department of Economics and Social Affairs, U. N. World Population Ageing 2019. (2020).
- 2.Clegg, A., Young, J., Iliffe, S., Rikkert, M. O. & Rockwood, K. Frailty in elderly people. Lancet381, 752–762 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Fried, L. P. et al. The physical frailty syndrome as a transition from homeostatic symphony to cacophony. Nat. Aging1, 36–46 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Campbell, A. J. & Buchner, D. M. Unstable disability and the fluctuations of frailty. Age Ageing26, 315–318 (1997). [DOI] [PubMed] [Google Scholar]
- 5.Rodriguez-Mañas, L. & Fried, L. P. Frailty in the clinical scenario. Lancet385, e7–e9 (2015). [DOI] [PubMed] [Google Scholar]
- 6.Fried, L. P. et al. Frailty in older adults: evidence for a phenotype. J. Gerontol. A Biol. Sci. Med. Sci.56, M146–M156 (2001). [DOI] [PubMed] [Google Scholar]
- 7.Rockwood, K. et al. A global clinical measure of fitness and frailty in elderly people. CMAJ173, 489–495 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Li, G. et al. Comparison between frailty index of deficit accumulation and phenotypic model to predict risk of falls: data from the global longitudinal study of osteoporosis in women (GLOW) Hamilton cohort. PLoS ONE 10, e0120144 (2015). [DOI] [PMC free article] [PubMed]
- 9.Woo, J., Leung, J. & Morley, J. E. Comparison of frailty indicators based on clinical phenotype and the multiple deficit approach in predicting mortality and physical limitation. J. Am. Geriatr. Soc60, 1478–1486 (2012). [DOI] [PubMed] [Google Scholar]
- 10.García-García, F. J. et al. A new operational definition of frailty: the Frailty Trait Scale. J. Am. Med. Dir. Assoc.15, 371.e7–371.e13 (2014). [DOI] [PubMed] [Google Scholar]
- 11.García-García, F. J. et al. Frailty trait scale-short form: a frailty instrument for clinical practice. J. Am. Med. Dir. Assoc.21, 1260–1266.e2 (2020). [DOI] [PubMed] [Google Scholar]
- 12.Rodríguez-Artalejo, F. & Rodríguez-Mañas, L. The frailty syndrome in the public health agenda. J. Epidemiol. Community Health68, 703–704 (2014). [DOI] [PubMed] [Google Scholar]
- 13.Landi, F. et al. Moving against frailty: does physical activity matter?. Biogerontology11, 537–545 (2010). [DOI] [PubMed] [Google Scholar]
- 14.Peterson, M. J. et al. Physical activity as a preventative factor for frailty: the health, aging, and body composition study. J. Gerontol. A Biol. Sci. Med. Sci.64, 61–68 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.van Hees, V. T. et al. A Novel, Open Access Method to Assess Sleep Duration Using a Wrist-Worn Accelerometer. PLoS ONE10, e0142533 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Santos-Lozano, A. et al. Actigraph GT3X: validation and determination of physical activity intensity cut points. Int. J. Sports Med.34, 975–982 (2013). [DOI] [PubMed] [Google Scholar]
- 17.Dinh-Le, C., Chuang, R., Chokshi, S. & Mann, D. Wearable health technology and electronic health record integration: scoping review and future directions. JMIR Mhealth Uhealth7, e12861 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Schwenk, M. et al. Wearable sensor-based in-home assessment of gait, balance, and physical activity for discrimination of frailty status: baseline results of the Arizona frailty cohort study. Gerontology61, 258–267 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Razjouyan, J. et al. Wearable sensors and the assessment of frailty among vulnerable older adults: an observational cohort study. Sensors18, (2018). [DOI] [PMC free article] [PubMed]
- 20.Kim, B., McKay, S. M. & Lee, J. Consumer-grade wearable device for predicting frailty in Canadian home care service clients: prospective observational proof-of-concept study. J. Med. Internet Res.22, e19732 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Sayed, N. et al. An inflammatory aging clock (iAge) based on deep learning tracks multimorbidity, immunosenescence, frailty and cardiovascular aging. Nat. Aging1, 598–615 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Rahman, S. A. & Adjeroh, D. A. Deep learning using convolutional LSTM estimates biological age from physical activity. Sci. Rep.9, 11425 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Rockwood, K., McMillan, M., Mitnitski, A. & Howlett, S. E. A frailty index based on common laboratory tests in comparison with a clinical frailty index for older adults in long-term care facilities. J. Am. Med. Dir. Assoc.16, 842–847 (2015). [DOI] [PubMed] [Google Scholar]
- 24.Garcia-Garcia, F. J. et al. The prevalence of frailty syndrome in an older population from Spain. The Toledo Study for Healthy Aging. J. Nutr. Health Aging15, 852–856 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Mañas, A. et al. Dose-response association between physical activity and sedentary time categories on ageing biomarkers. BMC Geriatr19, 270 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Freedson, P. S., Melanson, E. & Sirard, J. Calibration of the Computer Science and Applications, Inc. accelerometer. Med. Sci. Sports Exerc.30, 777–781 (1998). [DOI] [PubMed] [Google Scholar]
- 27.Chinoy, E. D. et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep44, (2021). [DOI] [PMC free article] [PubMed]
- 28.Sivertsen, B. et al. A comparison of actigraphy and polysomnography in older adults treated for chronic primary insomnia. Sleep29, 1353–1358 (2006). [DOI] [PubMed] [Google Scholar]
- 29.Radtke, T., Rodriguez, M., Braun, J. & Dressel, H. Criterion validity of the ActiGraph and activPAL in classifying posture and motion in office-based workers: A cross-sectional laboratory study. PLoS ONE16, e0252659 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.An, H.-S., Kim, Y. & Lee, J.-M. Accuracy of inclinometer functions of the activPAL and ActiGraph GT3X + : A focus on physical activity. Gait Posture51, 174–180 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Chen, T. & Guestrin, C. XGBoost: A Scalable Tree Boosting System. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD ’16 785–794 (ACM Press, 10.1145/2939672.2939785.(2016).
- 32.Friedman, J. H. Greedy function approximation: a gradient boosting machine. Ann. Statist.29, 1189–1232 (2001). [Google Scholar]
- 33.Friedman, J., Hastie, T. & Tibshirani, R. Additive logistic regression: a statistical view of boosting (With discussion and a rejoinder by the authors). Ann. Statist.28, 337–407 (2000). [Google Scholar]
- 34.James, G., Witten, D., Hastie, T. & Tibshirani, R. An Introduction to StatisticalLearning. vol. 103 (Springer New York, (2013).
- 35.King, R. D., Orhobor, O. I. & Taylor, C. C. Cross-validation is safe to use. Nat. Mach. Intell.3, 276–276 (2021). [Google Scholar]
- 36.Maaten, L. van der & Hinton, G. Visualizing Data using t-SNE. Journal of Machine Learning Research (2008).
- 37.Kraus, W. E. et al. Relationship between baseline physical activity assessed by pedometer count and new-onset diabetes in the NAVIGATOR trial. BMJ Open Diabetes Res. Care6, e000523 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Bandeen-Roche, K. et al. Phenotype of frailty: characterization in the women’s health and aging studies. J. Gerontol. A Biol. Sci. Med. Sci.61, 262–266 (2006). [DOI] [PubMed] [Google Scholar]
- 39.Gill, T. M., Gahbauer, E. A., Allore, H. G. & Han, L. Transitions between frailty states among community-living older persons. Arch. Intern. Med.166, 418–423 (2006). [DOI] [PubMed] [Google Scholar]
- 40.Graham, J. E. et al. Frailty and 10-year mortality in community-living Mexican American older adults. Gerontology55, 644–651 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Ensrud, K. E. et al. A comparison of frailty indexes for the prediction of falls, disability, fractures, and mortality in older men. J. Am. Geriatr. Soc.57, 492–498 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Burnham, J. P., Lu, C., Yaeger, L. H., Bailey, T. C. & Kollef, M. H. Using wearable technology to predict health outcomes: a literature review. J. Am. Med. Inform. Assoc.25, 1221–1227 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Walsh, J. T., Charlesworth, A., Andrews, R., Hawkins, M. & Cowley, A. J. Relation of daily activity levels in patients with chronic heart failure to long-term prognosis. Am. J. Cardiol.79, 1364–1369 (1997). [DOI] [PubMed] [Google Scholar]
- 44.Yates, T. et al. Association between change in daily ambulatory activity and cardiovascular events in people with impaired glucose tolerance (NAVIGATOR trial): a cohort analysis. Lancet383, 1059–1066 (2014). [DOI] [PubMed] [Google Scholar]
- 45.Lundberg, S. M. & Lee, S.-I. A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems (2017).
- 46.Cunningham, H., Ewart, A., Riggs, L., Huben, R. & Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv10.48550/arxiv.2309.08600.(2023)
- 47.Howlett, S. E., Rutenberg, A. D. & Rockwood, K. The degree of frailty as a translational measure of health in aging. Nat. Aging1, 651–665 (2021). [DOI] [PubMed] [Google Scholar]
- 48.Fallahzadeh, R. et al. Objective Activity Parameters Track Patient-specific Physical Recovery Trajectories After Surgery and Link With Individual Preoperative Immune States. Ann. Surg.277, e503–e512 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Ghaemi, M. S. et al. Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancy. Bioinformatics35, 95–103 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Zalocusky, K. A. et al. The 10,000 immunomes project: building a resource for human immunology. Cell Rep25, 513–522.e3 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Furman, D., Davis, M. M., Dekker, C. L., Tibshirani, R. & Maecker, H. 1000 Immunomes Project. https://med.stanford.edu/1000immunomes.html.
- 52.tripodlaboratories/actigraphy-frailty: Initial Release. Zenodo10.5281/zenodo.18187284.(2026)
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Description of Additional Supplementary files
Data Availability Statement
Requests for original data can be handled through (www.ciberfes.es/) via a review committee. The hospital reviews and determines the purposes for the data requests and what data can be released. Data requests can be sent to: Research and teaching unit, Virgen del Valle Hospital Ctra. Cobisa S/N, 45071 Toledo – Spain, info@estudiotoledo.com. All procedures were approved by the Clinical Research Ethics Committee of the Toledo Hospital and were conducted in accordance with the Declaration of Helsinki for human studies.
Supplemental data can be accessed either as the included.csv files or as.rda files available at the github repository in the results folder https://github.com/tripodlaboratories/actigraphy-frailty/tree/main/results52. The included supplementary data files 1, 2, 4 and 8 contain results (rho, p-values, and RMSE) for each individual model repeat of FTS components, risk factors & outcomes, FTS5 components, and the various frailty metrics respectively. Supplementary Data 3, 5, 6 and 7 contain the average model repeat prediction evaluations for risk factors & outcomes, FTS5 components, FTS components, and the various frailty metrics respectively. Actigraphy activity data and descriptions of clinical outcomes, risks, and frailty components are provided as.csv’s in the data folder of the associated github https://github.com/tripodlaboratories/actigraphy-frailty/tree/main/data52.
Code necessary to reproduce results, figures, and models can be found at https://github.com/tripodlaboratories/actigraphy-frailty52. This code was run using R 4.1.2 and xgboost 1.5.2.1 on a macOS 12.3.1 system. There are no restrictions on its access or use.



