Abstract
Objective
The Knee Osteoarthritis Outcome Score for Joint Replacement (KOOS-JR) scale is commonly used to assess patient progress. Scale structural validity has not been completely assessed. The purpose of this study was to assess the internal consistency, structural validity, and multi-group invariance properties of the KOOS-JR in a large sample of patients receiving knee arthroplasty or non-operative care.
Methods
A cross-sectional study using the Surgical Outcome System (SOS) database. Patients receiving care for degenerative knee conditions were included in the study. Internal consistency was assessed using Cronbach's alpha and McDonald's Omega. A confirmatory factor analysis was conducted to confirm scale structure of the KOOS-JR using a priori cut-off values (Comparative Fit Index [CFI], Tucker-Lewis Index [TLI], Incremental Fit Index [IFI] ≥ 0.95, Root Mean Square Error of Approximation [RMSEA] ≤ 0.06 preferred and ≤0.08 acceptable). Multigroup invariance testing was conducted across sex, age, and intervention groups.
Results
Internal consistency was acceptable (alpha = 0.83; omega = 0.83). The unidimensional structure of the KOOS-JR exceeded most contemporary model fit recommendations (CFI = 0.976, TLI = 0.964, IFI = 0.976, RMSEA = 0.067). The KOOS-JR was invariant across groups, allowing for comparison of variances and means between sex, age, and intervention groups.
Conclusion
The KOOS-JR met or exceeded most of the recommendations for model fit. The scale can be used to assess differences between males and females, middle and older aged adults, and between baseline measures of patients who received total knee arthroplasty or non-operative care.
Keywords: Structural validity, Psychometric analysis, Knee pathology, Total knee arthroplasty
1. Introduction
Patient outcomes (e.g., pain, quality of life) are evaluated using various reporting method perspectives (e.g., clinician, self, physiologic) [1]. Although clinician and physiologic-reported outcomes are valuable, recent emphasis has been placed on understanding the patient's perspective of their injury or well-being using patient-reported outcome measures (PROMs) [[2], [3], [4]]. PROMs can positively inform patient care by providing valuable insight into clinical intervention effectiveness [[5], [6], [7], [8]]. Thus, researchers have developed health related (e.g., quality of life, pain, disablement) and joint-specific (e.g., knee, ankle) PROMs [5,7]. A commonly used unidimensional joint-specific PROM is the Knee Injury and Osteoarthritis Outcome Score for Joint Replacement (KOOS-JR).
The KOOS-JR was developed using Rasch analysis on patient responses to the 42 items of the Knee Injury and Osteoarthritis Outcome Score (KOOS) instrument [9]. The KOOS was originally designed as a multi-dimensional scale for patients with end-stage osteoarthritis scheduled to undergo total knee arthroplasty [9]. The KOOS-JR uses seven of the original items to assess patient perceptions of overall knee health using questions pertaining to functional limitations, pain, and symptomology [9]; however, the KOOS-JR has been proposed as a unidimensional scale that provides a single score to quantify knee health [9]. The reduced length makes the KOOS-JR an attractive option for clinicians and patients alike (e.g., reduced response burden, reduced administration times, etc.) [[10], [11], [12], [13], [14]].
Researchers have found evidence of construct validity between the KOOS-JR and the KOOS Pain and KOOS Activities of Daily Living subscales [10]. Additionally, minimal detectable change (MDC) and minimal clinically important difference (MCID) values were calculated for the scale [11,14], and researchers have reported the scale demonstrated good responsiveness [10,12,13]. Thus, some initial findings provided evidence the scale can be used to measure the intended constructs and change in knee health across time [10,12,13]. However, subsequent psychometric evaluations of the scale have varied, particularly across statistical analysis approaches [11,14]. Researchers [[15], [16], [17]] have identified potential concerns with the psychometric properties (e.g., poor person differentiation, etc.) of the KOOS-JR, and examinations of the internal consistency, structural validity, and multi-group invariance testing of the scale are lacking. Specifically, there is a need to assess the internal consistency, structural validity and invariance of the KOOS-JR in a large, heterogenous sample of participants who only responded to the KOOS-JR items [[18], [19], [20], [21], [22], [23], [24]]; prior assessment of the scale [9] used responses to all the KOOS items and it is possible that participant responses to the 7 retained items were influenced by the other KOOS items or the provided responses to those items not retained in the KOOS-JR.
Therefore, further analysis should be conducted to assess measurement properties of the KOOS-JR to guide scale use in clinical practice and research. Recommended scale development steps include assessing scale internal consistency (e.g., Cronbach's alpha, etc.) and conducting confirmatory factor analysis (CFA) and multi-group invariance testing in a large, heterogeneous sample of patients who only respond to the 7-items in the KOOS-JR [[18], [19], [20], [21], [22], [23], [24]]. Evaluating internal consistency is important to support dimensional structure (i.e., support the claim of a unidimensional structure) and provide evidence that redundant or parallel items are not included in the scale [[18], [19], [20], [21]]. Conducting CFA is necessary to assess the structural validity, which would fill a gap in the literature for confirming structural validity in a large heterogenous sample who completed the KOOS-JR as a unique PRO. Further, multi-group invariance testing ensures the scale can be used to assess group differences and would indicate items are being interpreted similarly and the latent construct is being operationalized similarly across groups (e.g., sex, age, etc.) [[22], [23], [24]]. Thus, multi-group invariance would provide evidence that group differences are outside of measurement bias or error and would support scale use to assess group differences and test hypotheses [22,23,25]. Therefore, the purpose of this study was to assess the internal consistency, structural validity, and multi-group invariance properties of the KOOS-JR in a large sample of patients who would receive knee arthroplasty or non-operative care.
2. Methods
The Surgical Outcome System (SOS) is an international deidentified patient-reported outcome database that adheres to the Health Insurance Portability and Accountability Act (HIPAA), has already received Institutional Review Board (IRB) approval, and allows for retrospective analysis of previously collected patient data. The university IRB indicated approval for this study was not required because analysis of the deidentified data set from the SOS database was not considered human subject research. However, IRB approval was granted by the Cedar-Sinai Office of Research Compliance and Quality Improvement as part of a larger research project using SOS data. The data set utilized in this study included patients who had completed the KOOS-JR prior to receiving treatment/surgical intervention (i.e., knee arthroplasty, non-operative care) within the SOS database. Patient consent was obtained prior to the completion of patient-reported outcomes to be included in the SOS database.
2.1. KOOS-JR
The KOOS-JR is a 7-item instrument with questions pertaining to stiffness, pain, and knee function [9]. Patients respond to the seven questions using a 5-point Likert scale (none = 0, mild = 1, moderate = 2, severe = 3, extreme = 4) [26]. Raw scores are summed and range from 0 to 28, with 0 indicating perfect patient-perceived knee health. Raw scores are then transformed into an interval score, ranging from 0 (raw score of 28) to 100 (raw score of 0) [9]. For the purpose of this study, scores were converted to the transformed score.
2.2. Data analysis
KOOS-JR data and relevant demographic information were exported from the SOS database for analysis in the Statistical Package for Social Sciences (SPSS V. 25.0, Armonk, NY), the Hayes OMEGA SPSS package [27], and Analysis of Moment Structures (AMOS V. 25.0, Chicago, IL) software. Skewness values, kurtosis values, and histograms were assessed to evaluate normality of the data. Univariate and multivariate outliers were identified using z-scores (±3.3) and Mahalanobis distance, identified using a chi-square table with degrees of freedom and a p-value > 0.001 [23]. Cases in violation of the univariate or multivariate outlier criteria were removed. Descriptive statistics and frequencies were calculated for relevant demographic variables (e.g., age, sex) for the full sample and subgroups, and corresponding percentages were reported as appropriate.
2.3. Internal consistency
Internal consistency was assessed by calculating Cronbach's alpha (α) and McDonald's maximum likelihood omega (ω) for the 7-item unidimensional solution. Values < 0.70 indicate inadequate internal consistency, while values ≥ 0.90 may indicate item redundancy [[20], [21], [22], [23],28]. An acceptable range for Cronbach's alpha (α) and McDonald's maximum likelihood omega (ω) was set at ≥ 0.70 but ≤ .89 [20,28].
2.4. Scale structure - confirmatory factor analysis
A confirmatory factor analysis (CFA) with maximum likelihood estimation was conducted in AMOS on the proposed 7-item unidimensional KOOS-JR model. The model fit indices evaluated were set at established a priori values: Comparative Fit Index (CFI) ≥ 0.95; Tucker-Lewis Index (TLI) ≥ 0.95; Bollen Incremental Fit Index (IFI) ≥ 0.95; and root mean square error of approximation (RMSEA) ≤ 0.06 preferred and ≤0.08 acceptable [29,30]. The likelihood ratio statistic (CMIN) was calculated but was not used to assess model fit due to how heavily influenced the statistic is by sample size [22,23].
2.5. Multi-group invariance testing
A series of multi-group invariance analyses were conducted across relevant groups (i.e., intervention group [i.e., knee arthroplasty vs. non-operative care], sex [i.e., male, female], age group [youth: < 18, emerging adult: 18–25, early adulthood: 26–40, middle age: 41–65, older adult: > 65]) [31] to assess whether items were being interpreted equivalently across subgroups. Multi-group invariance testing was completed by testing three different models: configural model (i.e., to assess equal structure), metric model (i.e., to assess equal loadings), and scalar model (i.e., to assess equal intercepts), with each model being progressively more restrained than the previous model [23]. The CFI difference test (CFIdiff) and Chi-square difference test (χ2diff; p ≤ 0.01) were used to assess model fit. Model fit was considered adequate if CFIdiff was ≤0.01 [22,23,29]. While χ2diff was also assessed, the sensitivity of the statistic in large sample sizes led to χ2diff results being weighed less heavily in determining model fit [22,23,29]. Thus, multi-group invariance testing was continued if the CFIdiff criterion was met but the χ2diff criterion was violated.
3. Results
A total of 13470 complete responses (i.e., all items of the KOOS-JR were answered at baseline) from patients classified into a knee arthroplasty surgery group (n = 11564) or a non-operative care group (n = 1906) were extracted for data cleaning. From the sample of 13470 cases, 120 (knee arthroplasty: n = 110; non-operative: n = 10) univariate and multivariate cases were removed during the data cleaning process. A total of 1896 nonoperative cases and 11454 arthroplasty cases remained. To have an equal number of cases per group, a random sample of 1896 arthroplasty cases were extracted. The random selection of cases was generated in a three-step process: 1) a random number generator created a unique and random number for each arthroplasty case in the data set; 2) the unique identifiers were then sorted in ascending order; 3) the first 1896 cases were selected. A final sample of 3792 total cases (female = 2127, 56.09%; male = 1527, 40.27%; not reported = 138, 3.64%; age = 63.31 ± 9.94 y [range = 12–89 y]) were used for analysis.
3.1. Internal consistency
Cronbach's alpha (α) and McDonald's maximum likelihood omega (ω) were conducted on the full sample of baseline scores for all items as a single dimension. Cronbach's alpha and McDonald's omega were acceptable (α = 0.83; ω = 0.83).
3.2. Scale structure - confirmatory factor analysis
The CFA conducted on the full sample (n = 3792) met most recommended model fit indices values (CFI = 0.976, TLI = 0.964, IFI = 0.976). The RMSEA value (0.067) exceeded the preferred fit criterion (≤0.06) but met the criterion for acceptable RMSEA (Fig. 1). Modification indices did not demonstrate meaningful cross-loadings but indicated covariance between error terms for items 6 and 7; however, this was not specified as model fit was adequate.
Fig. 1.
Confirmatory Factor Analysis Measurement Model with Standardized Loadings. CFI = Comparative Fit Index; TLI = Tucker-Lewis Index; IFI = Bollen's Incremental Fit Index; RMSEA = Root Mean Square Error of Approximation, df = degrees of freedom. Each circle represents an unobserved variable, and each rectangle represents an observed variable. The unobserved variables to the right of each rectangle observed variable (e.g., e1, e2, e3, etc.) represent variance not explained by the factor that the indicator variable was intended to measure. The numbers along the lines from the KOOS-JR latent unobserved variable to each observed variable (e.g., Pre_1, Pre_2, etc.) are standardized regression weights. The numbers on the top right of each rectangle observed variable (e.g., Pre_1, Pre_2, etc.) are squared multiple correlations.
4. Multi-group invariance testing
4.1. Demographic information is provided in Table 1
Table 1.
Multigroup invariance testing group demographics.
| Characteristics | N (%) | Mean Age (SD) | Males (%) | Females (%) |
|---|---|---|---|---|
| Intervention Group | 1896 (50.00) | 65.56 ± 8.51 | 741 (40.90) | 1072 (59.10) |
| Knee Arthroplasty | 1896 (50.00) | 61.07 ± 10.72 | 786 (42.70) | 1055 (57.30) |
| Non-Operative Care | ||||
| Characteristic | N (%) | Mean Age (SD) | Arthroplasty (%) | Non-Op Care (%) |
| Sex Group | 1527 (41.79) | 63.08 ± 10.20 | 741 (48.52) | 786 (51.47) |
| Males | 2127 (58.21) | 63.51 ± 9.72 | 1072 (50.40) | 1055 (49.60) |
| Females | ||||
| Age Group | 1940 (53.7) | 57.72 ± 5.71 | 837 (43.14) | 1103 (56.86) |
| Middle Age (41-65y) | 1672 (46.29) | 71.49 ± 4.58 | 996 (59.57) | 676 (40.43) |
| Older Adult (+66y) |
4.1.1. Intervention group
The full sample (n = 3792) was used for analysis. Most model fit indices were met for the baseline knee arthroplasty and non-operative care group models, except for RMSEA (knee arthroplasty = 0.072, non-operative care = 0.068). The initial multi-group model (i.e., configural/equal form) exceeded all recommended fit indices (Table 2), and multi-group invariance testing continued. The metric model met fit criteria, warranting assessment of an equal latent variance model. The equal latent variance model passed all fit criteria, indicating variances were equal across groups. The scalar model also met model fit criteria, warranting examination of an equal latent means model. The equal latent means model did not pass the fit criteria, indicating the latent means were not equal across groups. Mean assessment indicated the non-operative care group reported significantly higher scores (i.e., better perception of knee health) than the knee arthroplasty group at baseline examination (Table 2).
Table 2.
Goodness-of-fit indices for measurement invariance analyses across intervention group.
| χ2 | dfa | χ2 difference (df) | CFI | CFI difference | TLI | IFI | RMSEA | |
|---|---|---|---|---|---|---|---|---|
| Arthroplasty (n = 1896) | 151.102 | 14 | b | 0.964 | b | 0.946 | 0.964 | 0.072 |
| Non-Op (n = 1896) | 136.810 | 14 | b | 0.977 | b | 0.966 | 0.977 | 0.068 |
| Configural (equal form) | 287.911 | 28 | b | 0.972 | b | 0.958 | 0.972 | 0.049 |
| Metric (equal loadings) | 306.714 | 34 | 18.803 (6) | 0.970 | 0.002 | 0.964 | 0.970 | 0.046 |
| Equal factor variances | 325.503 | 35 | 37.592 (7) | 0.969 | 0.003 | 0.962 | 0.969 | 0.047 |
| Scalar (equal indicator intercepts) | 366.056 | 40 | 78.145 (12) | 0.965 | 0.007 | 0.963 | 0.965 | 0.046 |
| Equal latent means | 594.785 | 41 | 306.874 (13)c | 0.940 | 0.032c | 0.939 | 0.940 | 0.60 |
CFIdiff criterion exceeded.
df = degrees of freedom.
Indicates the value is not calculated at this step.
Indicates the model did not pass invariance criteria.
4.1.2. Sex
A total of 3654 individuals (96.36%) reported their sex and were used for analysis (Table 1). The baseline models for sex met all preferred model fit criteria except for RMSEA values; RMSEA values slightly exceeded the preferred cut-off value but met the criterion for adequate fit (Table 3). The initial multi-group model (i.e., configural/equal form) exceeded all recommended fit indices; thus, multigroup invariance testing proceeded. The metric model met fit criteria, indicating assessment of an equal latent variance model was warranted. The equal latent variance model met fit criteria, indicating equal latent variances between groups. The scalar model also met fit criteria; thus, an equal latent means model was tested. The equal latent means model did not meet model fit criteria; further assessment of the means revealed females reported significantly lower scores (i.e., perceptions of worse knee health) than males (Table 3).
Table 3.
Goodness-of-fit indices for measurement invariance analyses across sex.
| χ2 | dfa | χ2 difference (df) | CFI | CFI difference | TLI | IFI | RMSEA | |
|---|---|---|---|---|---|---|---|---|
| Males (n = 1527) | 110.768 | 14 | b | 0.976 | b | 0.964 | 0.976 | 0.067 |
| Females (n = 2127) | 142.679 | 14 | b | 0.976 | b | 0.964 | 0.976 | 0.066 |
| Configural (equal form) | 253.448 | 28 | b | 0.976 | b | 0.964 | 0.976 | 0.047 |
| Metric (equal loadings) | 261.666 | 34 | 8.218 (6) | 0.975 | 0.001 | 0.970 | 0.975 | 0.043 |
| Equal factor variances | 261.760 | 35 | 8.312 (7) | 0.976 | 0.000 | 0.971 | 0.976 | 0.042 |
| Scalar (equal indicator intercepts) | 322.327 | 40 | 68.879 (12) | 0.970 | 0.006 | 0.968 | 0.970 | 0.044 |
| Equal latent means | 388.349 | 41 | 134.901 (13)c | 0.963 | 0.013c | 0.962 | 0.963 | 0.048 |
CFIdiff criterion exceeded.
df = degrees of freedom.
Indicates the value is not calculated at this step.
Indicates the model did not pass invariance criteria.
4.1.3. Age group
A total of 3612 (95.25%) cases from the middle age and older adult age groups were utilized for analysis (Table 1). The other a priori established age groups (i.e., youth, emerging adult, young adult) [31] were not utilized in the analyses due to small sample sizes and substantially unequal sample sizes between groups. The sample size issues could unduly influence model fit estimation and including those subgroups in the analysis was deemed inappropriate [23]. The baseline models for age groups also met all preferred model fit criteria except for RMSEA values; RMSEA values exceeded the preferred cut-off value but met the criterion for adequate fit (Table 4). The initial multi-group model (i.e., configural/equal form) exceeded all recommended fit indices. Multigroup invariance testing proceeded and revealed the fit criteria were met for all subsequent testing steps: metric model, scalar model, equal latent variances model, and equal latent means model (Table 4). Thus, latent variances and latent means were not found to be statistically different between groups.
Table 4.
Goodness-of-fit indices for measurement invariance analyses across age group.
| χ2 | dfa | χ2 difference (df) | CFI | CFI difference | TLI | IFI | RMSEA | |
|---|---|---|---|---|---|---|---|---|
| Middle Age (n = 1940) | 143.239 | 14 | b | 0.976 | b | 0.964 | 0.976 | 0.069 |
| Older Adult (n = 1672) | 111.803 | 14 | b | 0.975 | b | 0.962 | 0.975 | 0.065 |
| Configural (equal form) | 255.042 | 28 | b | 0.975 | b | 0.963 | 0.975 | 0.047 |
| Metric (equal loadings) | 266.325 | 34 | 11.283 (6) | 0.975 | 0.000 | 0.969 | 0.975 | 0.044 |
| Equal factor variances | 273.272 | 35 | 18.230 (7) | 0.974 | 0.001 | 0.969 | 0.974 | 0.043 |
| Scalar (equal indicator intercepts) | 308.318 | 40 | 53.276 (12) | 0.971 | 0.004 | 0.696 | 0.971 | 0.043 |
| Equal latent means | 315.195 | 41 | 60.153 (13) | 0.970 | 0.005 | 0.969 | 0.970 | 0.043 |
CFIdiff criterion exceeded.
df = degrees of freedom.
Indicates the value is not calculated at this step.
5. Discussion
The purpose of our study was to assess internal consistency and structural and multi-group invariance properties of the KOOS-JR in a large sample of patients with degenerative joint disease of the knee. Scale structure of the KOOS-JR was assessed using contemporary classical test theory procedures [23,29]. Our results support the structural validity of the KOOS-JR and provide evidence that the scale can be utilized to assess group differences based on sex, age, or intervention group (i.e., knee arthroplasty or non-operative care). Furthermore, acceptable internal consistency indicates scale parsimony and supports a unidimensional scale structure [[20], [21], [22], [23],28].
Our CFA results reveal sound model fit exceeding most of the recommendations for preferred model fit criteria. Thus, our findings further support prior Rasch analysis findings [9] of a structurally valid unidimensional model. We did not identify overall model fit concerns or local model fit concerns (e.g., low path coefficient loadings, meaningful item cross-loadings), that would suggest other alternative specifications to the scale are necessary to maximize fit or parsimony.
Our study also provides novel insight into the multi-group invariance properties of the scale to support and guide use of the KOOS-JR in clinical practice and research. Multi-group invariance testing supports scale hypothesis testing (e.g., can the scale be used to determine if patients with more severe injuries report higher levels of dysfunction), which provides valuable insight to clinicians [22,23,29]. Multigroup invariance testing also provides evidence that scale items are being interpreted similarly across groups (e.g., females, males) and that underlying constructs (i.e., knee health) are being measured similarly across groups. Thus, invariance allows for the comparison of scores across groups and indicates score differences in knee health are true group differences as opposed to differences in how group members interpret an item or operationalize the latent construct [22,23,29]. To our knowledge, we are the first group to present multi-group invariance testing results on the KOOS-JR.
We found the KOOS-JR was invariant at baseline measure (i.e., intake physical examination) between the two care intervention groups (i.e., knee arthroplasty and non-operative care). We found statically significant latent mean differences between the knee arthroplasty and non-operative care groups, with the non-surgical group reporting higher mean scores (i.e., higher level of perceived knee function) than the surgical group. While caution should be used in interpreting KOOS-JR scores as diagnostic, our findings provide preliminary support scale validity because patients with greater severity of knee degeneration (i.e., those who warrant surgical intervention) should report lower scores (i.e., greater knee health impairment) on the KOOS-JR. We would expect greater levels of perceived knee health impairment on the KOOS-JR to correlate with greater levels of joint degeneration (e.g., more advanced osteoarthritis) that would warrant surgical intervention. Our findings may provide support that baseline KOOS-JR scores function in this manner as the group means of patients who reported more impaired knee health on the KOOS-JR were the group who ultimately received surgical intervention.
However, other explanations (e.g., patients who seek conservative care may have better coping strategies for impaired knee health, etc.) could help explain the differences between treatment groups and warrant future research. Specifically, assessing baseline KOOS-JR scores across treatment groups in comparison to pathology (e.g., stage of OA, etc.) and patient psychosocial variables (e.g., coping strategies, resilience, quality of life, etc.) would be valuable for establishing scale validity and helping to determine if baseline KOOS-JR scores could support surgical or non-surgical intervention decisions for patients. Further analysis (e.g., longitudinal analysis) of the KOOS-JR would also be beneficial to determine if the KOOS-JR is invariant across time and if diagnostic cut-off criteria could be developed for the KOOS-JR to guide intervention decisions.
Our results also provide evidence that the KOOS-JR is invariant between groups of older adult populations (i.e., 41 years or older) and across sexes, which indicates the scale can be used to assess differences in knee health across these groups. Significant latent variance and latent mean differences were not found between the age groups suggesting minimal differences in knee health were perceived between the groups. The data set available, however, was not sufficient for multi-group testing of all age groups (i.e., 40 years or younger); thus, caution is warranted if assessment of group differences is conducted in these patient groups until further research is conducted to establish multi-group invariance and structural validity.
Statistically significant latent mean differences between the males and females in our sample were found. Females accounted for a larger proportion of the knee arthroplasty subgroup, which could partially explain these findings. For example, females were a larger proportion of the knee arthroplasty group, which may mean that more females presented with severe pathology, which would correlate with greater knee health impairment scores on the KOOS-JR compared to males. Other potential sex differences, however, may also help explain the latent mean sex differences given the percentage of females was similar in the knee arthroplasty (59.10%) and non-operative care (57.30%) groups. For example, musculoskeletal pain sex differences have been reported for condition prevalence, pain experiences, treatment responses, and age [[32], [33], [34], [35]]. Researchers have also suggested women have reduced tolerance to painful stimuli [36] and a poorer capacity to cope with musculoskeletal pain due to higher levels of emotional distress and disability [37]. Others have indicated women report higher activity levels, pain acceptance, and social support than male counterparts who report more mood disturbances, higher kinesephobia, and lower activity levels with similar levels of perceived pain severity [35]. Thus, further research is needed to understand the underlying mechanisms for sex differences in knee health captured by the KOOS-JR.
While this study has several strengths, including a large, heterogenous sample of patients seeking care, limitations do exist. First, all possible subgroups (e.g., younger populations, athletes, different surgical procedures) were not analyzed due to sample size limitations or the lack of information present in the SOS database. Thus, caution is warranted when examining KOOS-JR score differences in groups not analyzed. Further, the SOS dataset utilized did not include responses from healthy participants, nor did it include details on pathology (e.g., diagnosis, symptom duration), intervention utilized (e.g., specifics of non-operative care, surgical approach utilized), complete demographic information (e.g., ethnicity, physical activity level), or longitudinal data (e.g., baseline, post-surgery). Thus, many valuable analyses could not be performed, such as assessing test-retest reliability or calculating minimal detectable change and minimal clinically important differences (MCIDs).
Future research should determine if the KOOS-JR is invariant across younger age groups or across different levels of physical activity if the scale is to be used in those populations. Additionally, longitudinal invariance testing should be performed to ensure the measurement properties of the scale are maintained across repeated testing. Longitudinal testing also provides insight into if the KOOS-JR can be used to monitor recovery and guide patient care decisions (e.g., rehabilitation progression, discharge). Additionally, further longitudinal invariance and latent growth modeling analyses with more complete data (e.g., injury type, surgical approach, patient activity level) would allow for testing of substantive clinical questions (e.g., differences in treatment outcomes, rate of recovery between interventional techniques) that are important for guiding clinical practice.
6. Conclusion
The KOOS-JR met or exceeded most of the recommendations for model fit. Our findings support the structural validity of the KOOS-JR and the use of the scale to assess differences between males and females, middle and older aged adults, and between baseline measures of patients who received total knee arthroplasty or non-operative care. Further psychometric testing is necessary to establish reliability precision estimates (e.g., MCIDs) and longitudinal invariance to guide clinical use of the scale to assess patient progress over time.
Ethics approval and consent to participate
Institutional Review Board (IRB) approval for the project was granted by the Cedar-Sinai Office of Research Compliance and Quality Improvement as part of a larger research project using SOS data. University IRB was not required because the deidentified data set was not considered human subject research and the SOS adheres to the Health Insurance Portability and Accountability Act (HIPAA).
Availability of data and materials
The datasets analyzed during the study are not publicly available per study protocol; however, deidentified data may be available from the corresponding author with permission from Cedar-Sinai Office of Research Compliance and Quality Improvement, the Kerlan-Jobe Institute, and the University of Idaho upon reasonable request.
Competing interests
The authors declare they have no competing interests.
Funding
This publication was supported by an Institutional Development Award (IDeA) from the National Institute of General Medical Sciences of the National Institutes of Health under Grant #P20GM103408 and an Idaho WWAMI Research Training Support Award.
Author contributions
CA – concept/design, data interpretation, manuscript drafting, revisions, final edits/approval. AJR – data analysis/interpretation, manuscript drafting, revisions, final edits/approval. MPC –data acquisition, data analysis/interpretation, manuscript drafting, revisions, final edits/approval. ACC – concept/design, data acquisition, manuscript drafting, revisions, final edits/approval. RTB – concept/design, data acquisition, data analysis/interpretation, manuscript drafting, revisions, final edits/approval.
Acknowledgements
Not applicable.
Contributor Information
Caleb Allred, Email: cmallred@uw.edu.
Ashley J. Reeves, Email: reevesa@uidaho.edu.
Madeline P. Casanova, Email: mcasanova@uidaho.edu.
Adam C. Cady, Email: adamccady@gmail.com.
Russell T. Baker, Email: russellb@uidaho.edu.
References
- 1.Chin R., Lee B.Y. Elsevier Inc; London, UK: 2008. Principles and Practice of Clinical Trial Medicine. [Google Scholar]
- 2.Cano S., Pendrill L., Melin J., Fisher W., Jr. Towards consensus measurement standards for patient-centered outcomes. Measurement. 2019;141:62–69. [Google Scholar]
- 3.Institute of Medicine (US) Crossing The Quality Chasm-A New Health System For the 21stCentury. first ed. National Academies Press; Washington, DC: 2001. Committee on quality of health care in America. [PubMed] [Google Scholar]
- 4.Ware J.E., Jr., Snyder M.K., Wright W.R., Davies A.R. Defining and measuring patient satisfaction with medical care. Eval. Progr. Plann. 1983;6(3–4):247–263. doi: 10.1016/0149-7189(83)90005-8. [DOI] [PubMed] [Google Scholar]
- 5.Farnsworth J.L., II, Evans T., Binkley H., Kang M. Evaluation of knee-specific patient-reported outcome measures using Rasch analysis. J. Sport Rehabil. 2020;30(2):278–285. doi: 10.1123/jsr.2019-0263. [DOI] [PubMed] [Google Scholar]
- 6.Jette A.M. Outcomes research: shifting the dominant research paradigm in physical therapy. Phys. Ther. 1995;75(11):965–970. doi: 10.1093/ptj/75.11.965. [DOI] [PubMed] [Google Scholar]
- 7.Lam K.C., Harrington K.M., Cameron K.L., Snyder Valier A.R. Use of patient-reported outcome measures in athletic training: common measures, selection considerations, and practical barriers. J. Athl. Train. 2019;54(4):449–458. doi: 10.4085/1062-6050-108-17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Reeves M., Lisabeth L., Williams L., Katzan I., Kapral M., Deutsch A., et al. Patient-reported outcome measures (PROMs) for acute stroke: rationale, methods and future directions. Stroke. 2018;49(6):1549–1556. doi: 10.1161/STROKEAHA.117.018912. [DOI] [PubMed] [Google Scholar]
- 9.Lyman S., Lee Y.Y., Franklin P.D., Li W., Cross M.B., Padgett D.E. Validation of the KOOS, JR: a short-form knee arthroplasty outcomes survey. Clin. Orthop. Relat. Res. 2016;474(6):1461–1471. doi: 10.1007/s11999-016-4719-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Buller L.T., McLawhorn A.S., Lee Y.Y., Cross M., Haas S., Lyman S. The Short Form KOOS, JR is valid for revision knee arthroplasty. J. Arthroplasty. 2020;35(9):2542–2549. doi: 10.1016/j.arth.2020.04.016. [DOI] [PubMed] [Google Scholar]
- 11.Hung M., Bounsanga J., Voss M.W., Saltzman C.L. Establishing minimum clinically important difference values for the Patient-Reported Outcomes Measurement Information System Physical Function, hip disability and osteoarthritis outcome score for joint reconstruction, and knee injury and osteoarthritis outcome score for joint reconstruction in orthopaedics. World J. Orthoped. 2018;9(3):41–49. doi: 10.5312/wjo.v9.i3.41. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hung M., Saltzman C.L., Greene T., Voss M.W., Bounsanga J., Gu Y., et al. Evaluating instrument responsiveness in joint function: the HOOS JR, the KOOS JR, and the PROMIS PF CAT. J. Orthop. Res. 2018;36(4):1178–1184. doi: 10.1002/jor.23739. [DOI] [PubMed] [Google Scholar]
- 13.Khalil L.S., Darrith B., Franovic S., Davis J.J., Weir R.M., Banka T.R. Patient-reported outcomes measurement information system (PROMIS) global health short forms demonstrate responsiveness in patients undergoing knee arthroplasty. J. Arthroplasty. 2020;35(6):1540–1544. doi: 10.1016/j.arth.2020.01.032. [DOI] [PubMed] [Google Scholar]
- 14.Lyman S., Lee Y.Y., McLawhorn A.S., Islam W., MacLean C.H. What are the minimal and substantial improvements in the HOOS and KOOS and JR versions after total joint replacement? Clin. Orthop. Relat. Res. 2018;476(12):2432–2441. doi: 10.1097/CORR.0000000000000456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Hunnicutt J.L., Hand B.N., Gregory C.M., Slone H.S., McLeod M.M., Pietrosimone B., et al. KOOS-JR demonstrates psychometric limitations in measuring knee health in individuals after ACL reconstruction. Sport Health. 2019;11(3):242–246. doi: 10.1177/1941738118812454. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Padilla J.A., Rudy H.L., Gabor J.A., Friedlander S., Iorio R., Karia R.J., et al. Relationship between the patient-reported outcome measurement information system and traditional patient-reported outcomes for osteoarthritis. J. Arthroplasty. 2019;34(2):265–272. doi: 10.1016/j.arth.2018.10.012. [DOI] [PubMed] [Google Scholar]
- 17.Eckhard L., Munir S., Wood D., Talbot S., Brighton R., Walter B., et al. The ceiling effects of patient reported outcome measures for total knee arthroplasty. Orthop. Traumatol. 2021;107(3) doi: 10.1016/j.otsr.2020.102758. [DOI] [PubMed] [Google Scholar]
- 18.Panayides P. Coefficient alpha: interpret with caution. Eur. J. Psychol. 2013;9(4):687–696. [Google Scholar]
- 19.Pesudovs K., Burr J.M., Harley C., Elliott D.B. The development, assessment, and selection of questionnaires. Optom. Vis. Sci. 2007;84(8):663–674. doi: 10.1097/OPX.0b013e318141fe75. [DOI] [PubMed] [Google Scholar]
- 20.Streiner D.L. Starting at the beginning: an introduction to coefficient alpha and internal consistency. J. Pers. Assess. 2003;80(1):99–103. doi: 10.1207/S15327752JPA8001_18. [DOI] [PubMed] [Google Scholar]
- 21.Taber K.S. The use of Cronbach's alpha when developing and reporting research instruments in science education. Res. Sci. Educ. 2018;48(6):1273–1296. [Google Scholar]
- 22.Brown T.A. second ed. Guilford Publications; New York, NY: 2014. Confirmatory Factor Analysis for Applied Research. [Google Scholar]
- 23.Kline R.B. fourth ed. Guilford Press; New York, NY: 2016. Principles and Practice of Structural Equation Modeling. [Google Scholar]
- 24.Mokkink L.B., Prinsen C.A.C., Patrick D.L., Alonso J., Bouter L.M., de Vet H.C.W., et al. COSMIN study design checklist for patient reported outcome measurement instruments. https://www.cosmin.nl/wp-content/uploads/COSMIN-study-designing-checklist_final.pdf Available at: Accessed.
- 25.Chen F.F. Sensitivity of goodness of fit indexes to lack of measurement invariance. Struct. Equ. Model. 2007;14(3):464–504. [Google Scholar]
- 26.Roos E.M., Stefan Lohmander L. The knee injury and osteoarthritis outcome score (KOOS): from joint injury to osteoarthritis. Health Qual. Life Outcome. 2003;1:64. doi: 10.1186/1477-7525-1-64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Hayes A.F., Coutts J.J. Use omega rather than Cronbach's alpha for estimating reliability. But…. Commun. Methods Measur. 2020;14(1):1–24. [Google Scholar]
- 28.Leech N.L., Barrett K.C., Morgan G.A. fifth ed. Routledge; New York, NY: 2015. IBM SPSS for Intermediate Statistics: Use and Interpretation. [Google Scholar]
- 29.Byrne B.M. third ed. Routledge; New York, NY: 2016. Structural Equation Modeling with AMOS: Basic Concepts, Applications, and Programming. [Google Scholar]
- 30.Hu L.T., Bentler P.M. Cutoff criteria for fit indexes in covariance structural analysis: conventional criteria versus new alternatives. Struct. Equ. Model. 1999;6(1):1–55. [Google Scholar]
- 31.Sigelman C.K., Rider E.A. ninth ed. Cenage Learning; Boston, MA: 2018. Life-Span Human Development. [Google Scholar]
- 32.Breivik H., Collett B., Ventafridda V., Cohen R., Gallacher D. Survey of chronic pain in Europe: prevalence, impact on daily life, and treatment. Eur. J. Pain. 2006;10(4):287–333. doi: 10.1016/j.ejpain.2005.06.009. [DOI] [PubMed] [Google Scholar]
- 33.Gerdle B., BjoÈrk J., Henriksson C., Bengtsson A. Prevalence of current and chronic pain and their influences upon work and healthcare-seeking: a population study. J. Rheumatol. 2004;31(7):1399–1406. [PubMed] [Google Scholar]
- 34.Stubbs D., Krebs E., Bair M., Damush T., Wu J., Sutherland J., et al. Sex differences in pain and pain-related disability among primary care patients with chronic musculoskeletal pain. Pain Med. 2010;11(2):232–239. doi: 10.1111/j.1526-4637.2009.00760.x. [DOI] [PubMed] [Google Scholar]
- 35.Rovner G.S., Sunnerhagen K.S., Bjorkdahl A., Gerdle B., Borsbo B., Johansson F., et al. Chronic pain and sex differences; women accept and move, while men feel blue. PLoS One. 2017;12(4) doi: 10.1371/journal.pone.0175737. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Queme L.F., Jankowski M.P. Sex differences and mechanisms of muscle pain. Curr. Opin. Phsyiol. 2019;11:1–6. doi: 10.1016/j.cophys.2019.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Grossi G., Soares J.F., Lundberg U. Gender differences in coping with musculoskeletal pain. Int. J. Behav. Med. 2000;7(4):305–321. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets analyzed during the study are not publicly available per study protocol; however, deidentified data may be available from the corresponding author with permission from Cedar-Sinai Office of Research Compliance and Quality Improvement, the Kerlan-Jobe Institute, and the University of Idaho upon reasonable request.

