Abstract
Background
Deficit accumulation frailty indices (FIs) are widely used to characterize frailty. FIs vary in number and composition of items; the impact of this variation on reliability and clinical applicability is unknown.
Method
We simulated 12 000 studies using a set of 70 candidate deficits in 12 080 community-dwelling participants 65 years and older. For each study, we varied the number (5, 10, 15, 25, 35, 45) and composition (random selection) of items defining the FI and calculated descriptive and predictive estimates: frailty score, prevalence, frailty cutoff, mortality odds ratio, predicted probability of mortality for FI = 0.28 (prevalence threshold), and FI cutoff predicting 10% mortality over the follow-up. We summarized the estimates’ medians and spreads (0.025–0.975 quantiles) by number of items and calculated intraclass correlation coefficients (ICCs).
Results
Medians of frailty scores were 0.11–0.12 with decreasing spreads from 0.04–0.24 to 0.10–0.14 for 5-item and 45-item FIs. The median cutoffs identifying 15% as frail was 0.19–0.20 and stable; the spreads decreased with more items. However, medians and spreads for the prevalence of frailty (median: 11%–3%), mortality odds ratio (median: 1.24–2.19), predicted probability of mortality (median: 8%–17%), and FI cutoff predicting 10% mortality (median: 0.38–0.20) varied markedly. ICC increased from 0.19 (5-item FIs) to 0.84 (45-item FIs).
Conclusions
Variability in the number and composition of items of individual FIs strongly influences their reliability. Estimates using FIs may not be sufficiently stable for generalizing results or direct application. We propose avenues to improve the development, reporting, and interpretation of FIs.
Keywords: CLSA, Measurement error, Psychometrics, Regression dilution, Reliability
Frailty, conceptualized as a state of increased vulnerability to stressors, is often used to characterize heterogeneity in aging (1) in research and clinical practice (2,3). While there is currently no consensus definition for frailty, the deficit accumulation frailty index (FI) (4,5) is one of the 2 major frameworks underpinning the current understanding of frailty, along with the Fried frailty phenotype (6–8). The deficit accumulation FI conceptualizes frailty as the accumulation of age-related deficits, which is quantified using a standard procedure that calculates the proportion of deficits across a selection of 35 or more items measuring health states in multiple age-related domains (eg, physical function, medical conditions) (9). Since first described in 2001 (4), numerous other deficit accumulation FIs have been developed using different sources of data (eg, comprehensive geriatric assessments (10), electronic medical records (11), laboratory values (12), trial (13) or registry (14) data), number of items (from 5 to 92) (15,16), from different domains (17), and using various cutoffs (18,19).
To incorporate frailty measures into clinical practice, frailty scores or frailty status measures should identify the same individuals as frail from one context or study to another. In a recent review of the reliability of 35 frailty instruments, FIs have shown the highest agreement among frailty frameworks (19). However, only 6 well-constructed FIs were included in the analysis, whereas the myriad implementations of FIs currently in use may not demonstrate the same reliability. Although coding items as dichotomous or ordinal does not affect the performance of FIs (20), the influence of variation in the number and the composition of items on reliability is unknown. Moreover, classical measurement statistics of reliability, such as intraclass correlation coefficients (ICCs), kappa, or standard error of measurement (SEm) (21,22), may not translate easily to clinical correlates and implications when incorporating frailty in practice. For instance, how reliability influences estimates such as prevalence (19,23), association with outcomes, and cutoffs categorizing frailty status (frail vs non-frail) has not been well described. Clinical decision-making at the individual level may require more stringent conditions than population-level inferences (24,25).
In this study, we investigated the reliability of FIs and the stability of estimates when computed in a single community-dwelling older adult population as represented by participants from the Canadian Longitudinal Study on Aging (CLSA). We simulated 12 000 studies in which we varied the number of items and the specific composition of individual FIs to describe the “in vivo” implications of FIs for developing, reporting, and interpreting studies using the FI.
Method
Cohort
We used baseline cross-sectional data from the CLSA which enrolled a nationally representative sample of over 51 000 participants aged 45–85 years, at baseline, into a telephone-only cohort or an in-person cohort (Comprehensive) assessed from 2012 to 2015 (26,27). As frailty is an age-related concept and is typically measured in older adult populations, we restricted our cohort to adults 65 years and older with mortality data as of July 2019, and to the Comprehensive cohort since data from physical assessments were required. Exclusion criteria for the CLSA were people living in the Canadian territories, on a First Nations reserve, or in institutions, being full-time members of the Armed Forces, having cognitive impairment (as assessed by interviewers), and being unable to respond in French or English. There were 30 097 CLSA participants in the Comprehensive cohort of whom 12 080 met the inclusion criteria for our analyses.
Methodological Framework
The methodological framework for our study is depicted in Figure 1. To investigate the interindex reliability of the FI across multiple configurations, we generated and compared a total of 12 000 iterations, each representing a potential individual study. Each iteration (or individual study) applied a unique definition of the FI by varying (i) the number of health deficits composing the FI (we chose 6 configurations: 5, 10, 15, 25, 35, and 45 items, with 2 000 iterations per configuration for a total of 12 000 single studies) and (ii) by randomly selecting the specific deficits composing each. Configurations with few items (ie, 5, 10, and 15) were included since 2 highly cited and influential deficit accumulation–derived FIs in usage include 5 and 11 items (14,15,28).
Figure 1.
Methodological framework for appraising the clinical measurement properties of various configurations of the frailty index. ICC = intraclass correlation coefficient; SEm = standard error of measurement.
Measure, Outcome, and Analyses at the Single-Study Level
At the individual study level, we created FIs using the standard procedure outlined by Searle et al. (9). Briefly, the FI is calculated as the proportion of health deficits in an individual, with health deficits satisfying the following conditions for a single time-point study: (i) deficits must be associated with health status, (ii) their prevalence must generally increase with age, (iii) deficits should not saturate too early, and (iv) they should cover a range of systems as a group. We considered a collection of 70 health deficits, across 10 domains, as the basis for our FI definitions. Deficits were chosen to map closely to those in the most cited FI proposed by Rockwood et al. (5) in 2005 using items from the Canadian Study of Health and Aging (CSHA). Supplementary Table 1 reports the health deficits and domains, cutoffs for categorizing continuous deficits, missing data for deficits, and their mapping to CSHA items. Individual deficits were assessed in-person and by self-report; for details, see the CLSA Protocol (27). The CLSA updated the mortality data in July 2019 (median follow-up: 5.6 years, interquartile range [IQR]: 1.4 years). The mortality data came from the next of kin contacting CLSA directly, identification of death at the time of follow-up, and from linkage to provincial vital statistics. Because date of death was not available, mortality status was determined for all participants in July 2019.
Analyses
Within each single study, we first computed descriptive estimates: (i) mean population frailty score (in the 12 080 participants), (ii) prevalence of frailty (as defined as a FI >0.28, the mean of cutoffs for the FIs in our scoping review of 150 articles on frailty, 35 of which used FIs [range of cutoffs = 0.18–0.41; Q. D. Nguyen, manuscript under review, 2021]), and (iii) the FI threshold that categorizes 15% of the population as frail (ie, defined as the 85th percentile of FI) (29). We then computed predictive estimates between the FI and mortality using logistic regression: (iv) the odds ratio (OR) for an 0.1 increase in the FI, (v) the average predicted risk of mortality over the follow-up period for a FI of 0.28, and (vi) the FI cutoff predicting a 10% risk of mortality. Together, these 6 estimates account for common descriptive and clinical usages of frailty: describing a population, identifying those with frailty, estimating the association between frailty and an outcome, predicting risk for an outcome, and identifying a risk threshold for decision-making.
Analyses and Outcome of Interest at the Comparative Level
We compared the 6 different configurations of the FI using 2 000 iterations (ie, “individual studies”). At the comparative level, our primary outcome of interest was the stability of each descriptive and predictive estimate, as measured by the median and the 2.5 and 97.5 percentiles. We also computed classical reliability statistics for each configuration: the ICC for agreement (ICCA = σ 2individuals /[σ 2individuals + σ 2FIs + σ 2residual]) and the SEm for agreement (SEmA = √[σ 2FIs + σ 2residual]), which scales the ICCA on the scale of the FI (30). Using ICC in this situation is analogous to assessing interrater reliability where “raters” are the randomly defined FIs. Following guidelines presented by Nunnally, a threshold of 0.9 was considered acceptable for ICCA in this clinical context (22,25). As an exploratory analysis on the impact of domain coverage of FIs, we repeated the stability of estimates analyses above, stratifying by the number of domains included in 10, 25, and 45-item FI configurations (as determined by the inclusion of at least one item from that domain). We calculated the ICCA for all configurations and number of domains. Analyses were performed using R 4.0.3 (R Foundation).
Results
Among the 12 080 community-dwelling participants, the mean (SD) age was 73.0 (5.7) years, and 6 097 (50.5%) were male. Impairment in activities of daily living (ADL, n = 495 [4.1%]) and instrumental ADL (iADL, n = 982 [8.2%]) was relatively infrequent. Self-reported hypertension (n = 5 937 [49.3%]) and osteoarthritis (n = 4 167 [34.5%]) were the most prevalent conditions. Overall, participants had good physical performance measures: mean time for single chair rise was 2.9 (0.9) seconds and mean gait speed was 0.9 (0.2) m/s. Table 1 reports the baseline characteristics of our study sample. Between the baseline and mortality assessment (median follow-up of 5.6 years, IQR = 1.4]), 762 (6.3%) participants were known to have died.
Table 1.
Baseline Characteristics of Participants 65 Years and Older of the Canadian Longitudinal Study on Aging (n = 12 080)
| Age, mean (SD) | 73.0 (5.7) |
| Male sex (%) | 6 097 (50.5) |
| White race/ethnicity (%) | 11 673 (96.6) |
| Married or living with partner (%) | 7 051 (61.7) |
| Living location (%) | |
| House | 8 760 (72.6) |
| Apartment or condominium | 3 126 (25.9) |
| Seniors’ housing | 132 (1.1) |
| Other | 53 (0.4) |
| ADL impairment (%) | 495 (4.1) |
| IADL impairment (%) | 982 (8.2) |
| Chronic conditions (%) | |
| Hypertension | 5 937 (49.3) |
| Diabetes | 2 624 (21.8) |
| Heart disease | 2 772 (23.0) |
| Stroke or transient ischemic attack | 918 (7.6) |
| Lung disease | 2 017 (16.7) |
| Kidney disease | 499 (4.1) |
| Thyroid disease | 2 156 (18.2) |
| Osteoarthritis | 4 167 (34.5) |
| Osteoporosis | 1 701 (14.3) |
| Cancer | 2 803 (23.3) |
| Anxiety or depression | 2 053 (17.0) |
| Physical performance measures, mean (SD) | |
| Grip strength (kg) | 31.7 (10.6) |
| Chair rise time (s) | 2.9 (0.9) |
| Gait speed (m/s) | 0.9 (0.2) |
Note: ADL = activities of daily living; IADL = instrumental activities of daily living; SD = standard deviation.
Reliability and Stability of Frailty Index Measurements and Estimates
Complete results for the reliability of FIs and stability of estimates are reported in Table 2. Figure 2 provides a graphical summary of results for the stability and distribution of estimates.
Table 2.
Reliability and Stability of Frailty Index Measurements and Estimates for Descriptive and Predictive Uses, by Number of Items in 6 Configurations
| Number of Items in Frailty Index in Configuration (2 000 iterations for each configuration) | ||||||
|---|---|---|---|---|---|---|
| 5 | 10 | 15 | 25 | 35 | 45 | |
| Descriptive uses of the frailty index, median (0.025 and 0.975 quantiles) | ||||||
| Frailty index, meana | 0.11 (0.04, 0.24) | 0.12 (0.06, 0.19) | 0.12 (0.07, 0.18) | 0.12 (0.08, 0.16) | 0.12 (0.09, 0.15) | 0.12 (0.10, 0.14) |
| Prevalence of frailty (FI > 0.28), median | 0.11 (0.02, 0.33) | 0.10 (0.02, 0.27) | 0.06 (0.02, 0.16) | 0.04 (0.02, 0.10) | 0.04 (0.02, 0.07) | 0.03 (0.02, 0.06) |
| Frailty index cutoff identifying 15% as having frailty | 0.20 (0.13, 0.40) | 0.20 (0.11, 0.30) | 0.20 (0.13, 0.29) | 0.20 (0.16, 0.25) | 0.20 (0.16, 0.23) | 0.19 (0.16, 0.22) |
| Predictive uses of the frailty index (0.025 and 0.975 quantiles) | ||||||
| Odds ratio for mortality over the follow-up, per 0.1 FI increase | 1.24 (1.02, 1.44) | 1.44 (1.18, 1.69) | 1.60 (1.31, 1.88) | 1.87 (1.58, 2.18) | 2.04 (1.79, 2.34) | 2.19 (1.97, 2.47) |
| Predicted probability of mortality over the follow-up for FI = 0.28 | 0.08 (0.06, 0.12) | 0.10 (0.07, 0.15) | 0.11 (0.08, 0.17) | 0.14 (0.10, 0.19) | 0.15 (0.12, 0.21) | 0.17 (0.14, 0.21) |
| Frailty index cutoff predicting 10% mortality over the follow-up | 0.38 (0.18, 1.23) | 0.28 (0.18, 0.46) | 0.25 (0.17, 0.35) | 0.22 (0.17, 0.28) | 0.21 (0.18, 0.25) | 0.20 (0.18, 0.23) |
| Measurement statistics (95% confidence interval) | ||||||
| Intraclass correlation coefficient for agreement | 0.19 (0.18, 0.19) | 0.34 (0.33, 0.34) | 0.45 (0.44, 0.46) | 0.63 (0.62, 0.63) | 0.75 (0.75, 0.76) | 0.84 (0.84, 0.85) |
| Standard error of measurement for agreement | 0.13 (0.13, 0.13) | 0.09 (0.09, 0.09) | 0.07 (0.07, 0.07) | 0.05 (0.05, 0.05) | 0.04 (0.04, 0.04) | 0.03 (0.03, 0.03) |
Notes: FI = frailty index. Two thousand iterations (simulated studies) were performed for each configuration.
aThe median, 0.025, and 0.975 quantiles are reported for the distribution of mean frailty index scores in each of 2 000 simulations.
Figure 2.
Summary of simulation results for 6 configurations of the frailty index (FI): frailty score, prevalence, frailty cutoffs, odds ratio, and mortality prediction (6 configurations × 2 000 iterations × 12 080 participants). (A) Nonsmooth FI densities: the overall area under each curve represents the full distribution of FI values, and the height is the relative proportion of FI values; the vertical lines indicate the median and the shaded areas represent those with frailty (prevalence) defined as FI > 0.28. (B–F) Boxplots: middle line in box indicates the median, box indicates the 25th and 75th percentiles, and whiskers indicate the 2.5th and 97.5th percentiles. (E) Probability of death over follow-up for FI = 0.28 (prevalence cutoff).
Descriptive Uses of the FI
The medians of the mean FI for all 6 configurations (5, 10, 15, 25, 35, and 45 items) was 0.11–0.12 with decreasing spreads between the 2.5 and 97.5 percentiles from 0.04–0.24 to 0.10–0.14 for 5-item FIs and 45-item FIs, respectively. Figure 2A shows the distribution of FIs by configurations and the increasing narrowness and smoothness of distributions as the number of items increases. The prevalence of frailty varied widely between configurations: the median prevalence was 0.11 for 5-item FIs but decreased to 0.03 for 45-item FIs; within each configuration, the 0.025–0.975 quantiles spreads of prevalence ranged from 0.02–0.33 to 0.02–0.06. The median of the FI cutoff identifying 15% of participants as frail was stable at 0.19–0.20; the spreads of this cutoff progressively decreased from 0.20–0.40 to 0.16–0.22.
Predictive Uses of the FI and Measurement Statistics
The medians of the OR from regressions of mortality on the FI varied, increasing from 1.24 (95% CI: 1.02, 1.43) for 5-item FIs to 2.19 (1.97, 2.47) for 45-item FIs. Likewise, the predicted probabilities of mortality over the follow-up period in individuals with a FI of 0.28 increased from 0.08 (0.06, 0.12) up to 0.17 (0.14, 0.21). Both the medians of the FI and spreads identifying a 10% risk of mortality decreased from 0.38 (0.18, 1.23) for 5-item FIs to 0.20 (0.18, 0.23) for 35-item FIs. The ICCs were higher as the number of items included increased, starting at 0.19 (0.18, 0.19) and reaching 0.84 (0.84, 0.85) for 45-item FIs. The SEm for agreement was substantial at 0.13 (0.13, 0.13) for 5-item FIs and decreased progressively to 0.03 (0.03, 0.03) for 45-item FIs.
Reliability and Stability by Number of Domains
Exploratory analysis results for FIs stratified by the number of domains are presented in Supplementary Figure 1 for the stability of estimates and in Supplementary Figure 2 for ICCA. Although there were differences in the stability of estimates for 10-item FIs by the number of domains included, the variability between the number of items included outweighed the variability of the number of domains covered. The ICCA followed a similar trend, where most of the differences in reliability was due to the total number of items.
Discussion
In this study of a single older adult population, we generated FI definitions based on the same set of 70 health deficits but varied both the number of deficits and the specific deficits considered. We show notable variation in the descriptive and predictive estimates computed from FIs, for example, mean score, prevalence, categorical frailty status, OR, frailty cutoff, and risk prediction. The instability of estimates was 2-fold: within configurations and between configurations. Within each configuration using the same number of deficits, the instability of estimates is observed as the spread around each median value. When comparing between configurations, we demonstrate additional variability in the median values themselves for the prevalence (from 5-item FIs to 45-item FIs: 0.11–0.03), the OR (1.24–2.19), the FIs predicting 10% mortality (0.08–0.17), and the estimated risk of mortality for a FI of 0.28 (0.38–0.20).
Previous work has investigated the psychometric properties and comparability of major existing frailty frameworks (19,31–34), as well as of FIs specifically (16,20). Although variation in prevalence and low interchangeability between frameworks is attributable to different underlying concepts of frailty (19), FIs report the highest agreement among frailty frameworks. As a group, FIs have good criterion and construct validity (16), but misclassification and decreased predictive accuracy may occur with the reduction of domains composing FIs (35). Our findings add to this body of research by focusing on (i) the reliability of the FIs and stability of estimates; (ii) the clinical implications of the reliability of FIs, specifically for patient care; and (iii) the differences between their population and individual interpretations of studies.
Reliability of FIs, Stability of Estimates, and Regression Dilution
Psychometrically, an ideal frailty measurement, used repeatedly in a short timeframe where health status is stable, should yield similar scores and identify, as frail, the same individuals. Intraclass correlation for agreement, the ratio of the variation between individuals and the overall variation (including variation due to measurement), is a classical measure of reliability. We found that when using a 10-item FI (ICCA = 0.34), only 34% of the variation in frailty scores was due to differences between individuals. As expected, we show that reliability improves as the number of items included in computing FIs increased. Yet, even with 35 and 45 items, the ICCA is 0.75 and 0.84, respectively, which is less than the 0.90 threshold for clinical decision-making recommended by Nunnally (22,25).
The reliability of FIs directly informs prevalence estimates: the lower the reliability, the wider the spread of prevalence estimates. Interestingly, the median prevalence also varied by number of items included (from 0.11 to 0.03) due to the nonsmooth distribution of frailty scores in relation to location of the frailty cutoff, as shown in Figure 2A. Varying frailty prevalence has been attributed to differing frailty operationalizations, cutoffs, and populations under study; we show that variation in prevalence may also be due to measurement issues, even within the FI framework using a single cutoff.
Using a small number of items increases measurement error which has important consequences on predictive estimates. This is illustrated in Table 2 showing that the magnitude of the association between frailty and mortality lessens as the number of items decreases. When exposures are mismeasured, as is the case in FIs using a lower number of items, the well-known phenomenon of regression dilution can occur, whereby mismeasurement biases the true association towards the null (36,37). In a seminal paper outlining the standard procedure to create a FI, Searle et al. specify at least 30–40 items should be included and that estimates are unstable when the number of items (~10) is small (9). This cautionary recommendation has not been completely heeded; according to our scoping review (Q. D. Nguyen, manuscript under review, 2021), the median number of items included was 32 in a random subset of 35 studies using FIs published in 2017 and 2018, of which 13 used the 5- or 11-item modified FI (14,15,28). Whereas previous work has highlighted the role of domain coverage in optimizing the predictive accuracy of FIs (35), our findings suggest that the total number of items included, rather than domain coverage, is more influential for the reliability of FIs.
Clinical Implications of the Reliability of FIs
The reliability of FIs has direct implications on their application to clinical practice. Our findings indicate firstly that FIs, especially those with a small number of items, can vary in scores attributed to each individual from one constructed index to another. Identifying frailty status as a consistent basis for decision-making is challenging due to the large SEm, from 0.13 for 5 items to 0.03 for 45 items, a substantial variation when considering the distribution of frailty scores. Second, in addition to reliability issues, the accuracy of the FI, continuous or categorical, to predict outcomes such as mortality, may be decreased due to regression dilution. Third, establishing cutoffs of the FI may yield unexpected results because of the nonsmooth distribution of frailty when measured using fewer than 25 items. Fourth, generalizing or transporting results between FIs or to clinical settings can only be done with caution since the average “difficulty” of items comprising each FI may vary markedly. In a single population, the mean frailty score may be 0.10 for one FI and 0.20 in another, without the possibility of identifying where each FI is anchored relative to the other in the underlying (latent) frailty continuum. Comparatively, other frailty instruments, such as the Clinical Frailty Scale (5) and the components of the Fried physical frailty phenotype (6), have absolute anchors by way of descriptive rubrics or explicit cutoffs, respectively. It may be useful to note that the clinical implications we describe are not unique to the FI: other constructs in geriatric practice and aging research (eg, multimorbidity) do not have a well-defined set of components, cutoffs, or anchoring.
Differences Between Population and Individual Interpretation of FIs
Although we focused on the clinical application of FIs, their usage in research to draw inferences at the population level may not face the same limitations (24), Psychometric issues are critical when translating research findings to individuals in clinical practice (22), However, the lack of reliable anchoring of FIs does not invalidate their consistent associations with adverse outcomes. At the population level, frailty is indeed associated with falls, progression of disability, and mortality (3); but what specific level of frailty can reliability predict falls and the initiation of fall prevention program at the individual level remains undetermined and varies between different constructed indices. A continuum in applicability exists whereby the requirements for clinical applicability are higher than that for research and population-level interpretation. In light of our findings, Table 3 presents recommendations for clinical usage and research of FIs, when devising, reporting, and interpreting FIs.
Table 3.
Recommendations for Clinical Interpretation and Usage, and Research Using Frailty Indices
| Clinical interpretation and usage |
| • Frailty indices and categorical thresholds should be applied with considerable caution due to potentially low reliability. |
| • Frailty indices are not interchangeable. Clinical interpretation and application should carefully consider the specific composition of items and the thresholds used in each definition. |
| • For studies using frailty indices constructed using a low number of items, the interpretation of results should be restricted to the specific frailty index and deficits considered. |
| • Frailty indices should include at least 30 items as has been previously recommended by Searle et al. (9), and ideally ≥45 items. Measures using fewer items may be used but should not be considered comparable to other frailty indices. |
| • When interpreting results from frailty indices as a whole, regression dilution should be considered whereby frailty indices comprising a fewer number of items will bias the true magnitude of association toward the null. |
| Reporting and research |
| • Frailty indices and their constitutive items should be characterized in greater detail: specific items, item cutoffs, and cutoffs for frailty status should be systematically reported. |
| • Consider anchoring current and future frailty indices by describing their distribution in a freely available standard population of older adults, thus allowing comparisons. |
| • Consider using Item Response Theory methods (38) to further characterize the most frequently used deficits and cutoffs, to compare frailty indices, and potentially to devise a standard blueprint for frailty items to be included. |
Limitations
Our study has a few important limitations that should be recognized. First, measures of reliability are dependent on the sample in which they are assessed. We included older adults 65 years and older from the CLSA which excluded institutionalized and cognitively impaired adults at baseline; the findings from these analyses may not be transportable to a younger or a less healthy population. Second, our results may be sensitive to the 70 candidate items and cutoffs we selected as a basis for our FIs; however, we chose items to map to the original CSHA items (5), using standard or first-quintile cutoffs as recommended (9). Third, we examined FIs with a number of items below the recommended 30–35 which lowers reliability; nonetheless, a large proportion of FIs are defined using a small number of items (12,14,15). Fourth, because we used the same 70 candidate items and cutoffs for our FIs, our analysis may actually overestimate the reliability of FIs currently in usage which may use more diverse items and cutoffs. As electronic medical records may allow the calculation of more reliable FIs due to greater availability of items, their cutoff and “level of difficulty” will be increasingly important to consider. Finally, because date of death was not available, associations and predictions between frailty and mortality using logistic regressions should be interpreted with caution, although comparisons remain valid between the predictive estimates of each configuration.
Conclusion
By simulating 12 000 individual studies, we show that variability in the number and the composition of items of individual FIs strongly influences their overall reliability. Descriptive and predictive estimates using FIs may not be sufficiently stable for generalizing results or for direct application to clinical practice. Although reliability improves as the number of items is increased, we propose further avenues to improve the development, reporting, and interpretation of FIs.
Supplementary Material
Acknowledgments
This research was made possible using the data/biospecimens collected by the Canadian Longitudinal Study on Aging (CLSA). This research has been conducted using the Baseline CLSA Comprehensive V4.0 data set and the Follow-up 1 Comprehensive V1.0 data set, under Application Number 190209. The CLSA is led by Drs. Parminder Raina, Christina Wolfson, and Susan Kirkland. The opinions expressed in this manuscript are the author’s own and do not reflect the views of the CLSA.
Funding
Funding for the Canadian Longitudinal Study on Aging (CLSA) is provided by the Government of Canada through the Canadian Institutes of Health Research (CIHR) under grant reference: LSA 94473 and the Canada Foundation for Innovation.
Conflict of Interest
None declared.
Author Contributions
Q.D.N., E.M.M., and C.W. designed the study. Q.D.N. analyzed data and drafted the manuscript. All authors contributed significantly to the content, critically reviewed, and approved the final manuscript for publication.
References
- 1. Nguyen QD, Moodie EM, Forget MF, Desmarais P, Keezer MR, Wolfson C. Health heterogeneity in older adults: exploration in the Canadian Longitudinal Study on Aging. J Am Geriatr Soc. 2021;69(3):678–687. doi: 10.1111/jgs.16919 [DOI] [PubMed] [Google Scholar]
- 2. Dent E, Kowal P, Hoogendijk EO. Frailty measurement in research and clinical practice: a review. Eur J Intern Med. 2016;31:3–10. doi: 10.1016/j.ejim.2016.03.007 [DOI] [PubMed] [Google Scholar]
- 3. Hoogendijk EO, Afilalo J, Ensrud KE, Kowal P, Onder G, Fried LP. Frailty: implications for clinical practice and public health. Lancet. 2019;394(10206):1365–1375. doi: 10.1016/S0140-6736(19)31786-6 [DOI] [PubMed] [Google Scholar]
- 4. Mitnitski AB, Mogilner AJ, Rockwood K. Accumulation of deficits as a proxy measure of aging. Scientific World Journal. 2001;1:323–336. doi: 10.1100/tsw.2001.58 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Rockwood K, Song X, MacKnight C. A global clinical measure of fitness and frailty. Can Med Assoc J. 2005;173(5):489–495. doi: 10.1503/cmaj.050051 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Fried LP, Tangen CM, Walston J, et al. ; Cardiovascular Health Study Collaborative Research Group . Frailty in older adults: evidence for a phenotype. J Gerontol A Biol Sci Med Sci. 2001;56(3):M146–M156. doi: 10.1093/gerona/56.3.m146 [DOI] [PubMed] [Google Scholar]
- 7. Buta BJ, Walston JD, Godino JG, et al. Frailty assessment instruments: systematic characterization of the uses and contexts of highly-cited instruments. Ageing Res Rev. 2016;26:53–61. doi: 10.1016/j.arr.2015.12.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. de Vries NM, Staal JB, van Ravensberg CD, Hobbelen JS, Olde Rikkert MG, Nijhuis-van der Sanden MW. Outcome instruments to measure frailty: a systematic review. Ageing Res Rev. 2011;10(1):104–114. doi: 10.1016/j.arr.2010.09.001 [DOI] [PubMed] [Google Scholar]
- 9. Searle SD, Mitnitski A, Gahbauer EA, Gill TM, Rockwood K. A standard procedure for creating a frailty index. BMC Geriatr. 2008;8(1):24. doi: 10.1186/1471-2318-8-24 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Jones DM, Song X, Rockwood K. Operationalizing a frailty index from a standardized comprehensive geriatric assessment. J Am Geriatr Soc. 2004;52(11):1929–1933. doi: 10.1111/j.1532-5415.2004.52521.x [DOI] [PubMed] [Google Scholar]
- 11. Clegg A, Bates C, Young J, et al. Development and validation of an electronic frailty index using routine primary care electronic health record data. Age Ageing. 2016;45(3):353–360. doi: 10.1093/ageing/afw039 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Howlett SE, Rockwood MR, Mitnitski A, Rockwood K. Standard laboratory tests to identify older adults at increased risk of death. BMC Med. 2014;12:171. doi: 10.1186/s12916-014-0171-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Pajewski NM, Williamson JD, Applegate WB, et al. ; SPRINT Study Research Group . Characterizing frailty status in the systolic blood pressure intervention trial. J Gerontol A Biol Sci Med Sci. 2016;71(5):649–655. doi: 10.1093/gerona/glv228 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Farhat JS, Velanovich V, Falvo AJ, et al. Are the frail destined to fail? Frailty index as predictor of surgical morbidity and mortality in the elderly. J Trauma Acute Care Surg. 2012;72(6):1526–1530; discussion 1530. doi: 10.1097/TA.0b013e3182542fab [DOI] [PubMed] [Google Scholar]
- 15. Subramaniam S, Aalberg JJ, Soriano RP, Divino CM. New 5-factor modified frailty index using American College of Surgeons NSQIP data. J Am Coll Surg. 2018;226(2):173–181.e8. doi: 10.1016/j.jamcollsurg.2017.11.005 [DOI] [PubMed] [Google Scholar]
- 16. Drubbel I, Numans ME, Kranenburg G, Bleijenberg N, de Wit NJ, Schuurmans MJ. Screening for frailty in primary care: a systematic review of the psychometric properties of the frailty index in community-dwelling older people. BMC Geriatr. 2014;14(1):27. doi: 10.1186/1471-2318-14-27 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Sternberg SA, Wershof Schwartz A, Karunananthan S, Bergman H, Mark Clarfield A. The identification of frailty: a systematic literature review. J Am Geriatr Soc. 2011;59(11):2129–2138. doi: 10.1111/j.1532-5415.2011.03597.x [DOI] [PubMed] [Google Scholar]
- 18. Romero-Ortuno R. An alternative method for frailty index cut-off points to define frailty categories. Eur Geriatr Med. 2013;4(5):299–303. doi: 10.1016/j.eurger.2013.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Aguayo GA, Donneau AF, Vaillant MT, et al. Agreement between 35 published frailty scores in the general population. Am J Epidemiol. 2017;186(4):420–434. doi: 10.1093/aje/kwx061 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Peña FG, Theou O, Wallace L, et al. Comparison of alternate scoring of variables on the performance of the frailty index. BMC Geriatr. 2014;14(1):25. doi: 10.1186/1471-2318-14-25 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Mokkink LB, Terwee CB, Patrick DL, et al. The COSMIN checklist for assessing the methodological quality of studies on measurement properties of health status measurement instruments: an international Delphi study. Qual Life Res. 2010;19(4):539–549. doi: 10.1007/s11136-010-9606-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Streiner DL, Norman GR, Cairney J.. Health Measurement Scales: A Practical Guide to Their Development and Use. 5th ed. Oxford: Oxford University Press; 2014. doi: 10.1093/med/9780199685219.001.0001. [DOI] [Google Scholar]
- 23. Collard RM, Boter H, Schoevers RA, Oude Voshaar RC. Prevalence of frailty in community-dwelling older persons: a systematic review. J Am Geriatr Soc. 2012;60(8):1487–1492. doi: 10.1111/j.1532-5415.2012.04054.x [DOI] [PubMed] [Google Scholar]
- 24. Henderson R, Keiding N. Individual survival time prediction using statistical models. J Med Ethics. 2005;31(12):703–706. doi: 10.1136/jme.2005.012427 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Nunnally JC. Psychometric Theory. 2nd ed. McGraw-Hill; 1978. [Google Scholar]
- 26. Raina P, Wolfson C, Kirkland S, et al. Cohort profile: the Canadian Longitudinal Study on Aging (CLSA). Int J Epidemiol. 2019;48(6):1752–1753. doi: 10.1093/ije/dyz173 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Raina PS, Wolfson C, Kirkland SA. Canadian Longitudinal Study on Aging—Protocol.https://clsa-elcv.ca/doc/511. Accessed November 26, 2019.
- 28. Velanovich V, Antoine H, Swartz A, Peters D, Rubinfeld I. Accumulating deficits model of frailty and postoperative mortality and morbidity: its application to a national database. J Surg Res. 2013;183(1):104–110. doi: 10.1016/j.jss.2013.01.021 [DOI] [PubMed] [Google Scholar]
- 29. Bandeen-Roche K, Seplaki CL, Huang J, et al. Frailty in older adults: a nationally representative profile in the United States. J Gerontol A Biol Sci Med Sci. 2015;70(11):1427–1434. doi: 10.1093/gerona/glv133 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. de Vet HCW, Terwee CB, Mokkink LB, Knol DL.. Measurement in Medicine: A Practical Guide. Cambridge, UK: Cambridge University Press; 2011. doi: 10.1017/CBO9780511996214 [DOI] [Google Scholar]
- 31. Rockwood K, Andrew M, Mitnitski A. A comparison of two approaches to measuring frailty in elderly people. J Gerontol A Biol Sci Med Sci. 2007;62(7):738–743. doi: 10.1093/gerona/62.7.738 [DOI] [PubMed] [Google Scholar]
- 32. Walston JD, Bandeen-Roche K. Frailty: a tale of two concepts. BMC Med. 2015;13:185. doi: 10.1186/s12916-015-0420-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Rockwood K, Blodgett JM, Theou O, et al. A frailty index based on deficit accumulation quantifies mortality risk in humans and in mice. Sci Rep. 2017;7:43068. doi: 10.1038/srep43068 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Theou O, Brothers TD, Mitnitski A, Rockwood K. Operationalization of frailty using eight commonly used scales and comparison of their ability to predict all-cause mortality. J Am Geriatr Soc. 2013;61(9):1537–1551. doi: 10.1111/jgs.12420 [DOI] [PubMed] [Google Scholar]
- 35. Shi SM, McCarthy EP, Mitchell S, Kim DH. Changes in predictive performance of a frailty index with availability of clinical domains. J Am Geriatr Soc. 2020;68(8):1771–1777. doi: 10.1111/jgs.16436 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Keogh RH, White IR. A toolkit for measurement error correction, with a focus on nutritional epidemiology. Stat Med. 2014;33(12):2137–2155. doi: 10.1002/sim.6095 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Hutcheon JA, Chiolero A, Hanley JA. Random measurement error and regression dilution bias. Br Med J. 2010;340:c2289. doi: 10.1136/bmj.c2289 [DOI] [PubMed] [Google Scholar]
- 38. Lord FM, Novick MR, Birnbaum A.. Statistical Theories of Mental Test Scores. Reading, MA: Addison-Wesley; 1968. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.


