Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Feb 1.
Published in final edited form as: Arthritis Care Res (Hoboken). 2022 Aug 31;75(2):381–390. doi: 10.1002/acr.24760

Comparison of PROMIS® Computerized Adaptive Testing versus Fixed Short Forms in Juvenile Myositis

Ruchi N Patel 1, Valeria G Esparza 2, Jin-Shei Lai 3, Elizabeth L Gray 4, Bryce B Reeve 5, Rowland W Chang 6, David Cella 7, Kaveh Ardalan 8,9
PMCID: PMC8800940  NIHMSID: NIHMS1728785  PMID: 34328696

Abstract

Objective:

Patient-Reported Outcomes Measurement Information System® (PROMIS®) measures can be administered via computerized adaptive testing (CAT) or fixed short forms (FSF), but the empirical benefits of CAT versus FSF are unknown in juvenile myositis (JM). This study assesses if PROMIS CAT is feasible, precise, correlated with FSF, and less prone to respondent burden and floor/ceiling effects than FSF in JM.

Methods:

Patients 8-17 yo (self-report and parent proxy) and parents of patients 5-7 yo (only parent proxy) completed PROMIS Fatigue, Pain Interference, Upper Extremity Function, Mobility, Anxiety and Depressive Symptoms measures. Pearson correlations, paired t-tests, and Cohen’s D were calculated between PROMIS CAT and FSF. McNemar test assessed floor/ceiling effects between CAT and FSF. Precision and respondent burden were examined across the T-score range.

Results:

Data from 67 patient-parent dyads were analyzed. CAT and FSF mean scores did not significantly differ except in parent proxy Anxiety and Fatigue (effect size: 0.23 and 0.19, respectively). CAT had less pronounced floor/ceiling effects at the less symptomatic extreme in all domains except self-report Anxiety. Increased item burden and higher standard errors were seen in less symptomatic scorers for CAT. Modified stopping rules limiting CAT item administration did not decrease precision.

Conclusion:

PROMIS CAT appears to be feasible and correlated with FSF. CAT had less pronounced floor/ceiling effects, allowing detection of individual differences in less symptomatic patients. Modified stopping rules for CAT may decrease respondent burden. CAT can be considered for long-term follow-up of JM patients.

Introduction

Despite more effective treatments, children with juvenile myositis (JM) continue to have poorer physical and psychosocial health-related quality of life (HRQoL) when compared to healthy children (1,2). Even after remission, JM patients often experience persistent deconditioning and other deficits in HRQoL which can continue into adulthood (1,3-7). Patients with JM often experience intense emotional distress even during periods of low disease activity leading to worse psychosocial HRQoL (8,9). Published guidelines call for inclusion of HRQoL as part of core sets of outcomes in JM and parents of children with JM have designated HRQoL as the most important outcome to assess in routine care (10-12). Hence, there has been an increased focus on understanding and measuring HRQoL outcomes in JM both during active disease and remission. However, existing HRQoL surveys are often limited by the choice between high respondent burden and imprecise measurement (13). Existing HRQoL measures used in JM and other pediatric rheumatic conditions are often limited by ceiling and floor effects that can lead to underestimation of HRQoL deficits and may reduce the ability to assess change reliably over time especially for those in the tails of the HRQoL distribution (7,14-17). Measures with strong psychometric evidence that can capture subtler, yet still important, deficits in function are needed given that patients with JM experience adverse HRQoL impact even in low disease activity states and during clinical remission (1,3-6).

Patient-Reported Outcomes Measurement Information System® (PROMIS®) instruments assess a broad range of HRQoL domains such as Mobility, Upper Extremity Function, Anxiety, Depressive Symptoms, Fatigue, and Pain Interference (18). PROMIS measures have been developed using rigorous mixed methods approaches employing focus groups, expert item review, cognitive interviewing, and psychometric evaluation and calibration in large populations of individuals with and without health conditions (19,20). PROMIS pediatric measures have seen support in studies of children with various pediatric rheumatic diseases, including JM (21-23).

PROMIS measures are comprised of large pools of items, known as item banks, that can be administered as fixed short forms (FSF) or computerized adaptive tests (CAT) (18). FSF present the same set of items derived from the item banks to each respondent, while CAT is an individualized assessment that selects items based on participants’ response to the prior items. In CAT, an initial item that assesses an intermediate level of the construct of interest is selected and administered from the item bank. Based on the response to this initial item, a score and a standard error is estimated, and the most informative items at the estimated scores will be chosen by a pre-set algorithm (18). CAT will continue to administer items until meeting criteria for the ‘stopping rule’, defined by a preferred precision level or pre-defined maximum number of items (whichever comes first) (24,25). By offering more individualized assessments, CAT has been hypothesized to decrease the number of items needed to reach a desired precision level in comparison with FSF (18). However, few studies have been conducted comparing CAT with FSF and available evidence has been mixed (21,24-29).

One of the most important benefits of CAT is its potential to better differentiate respondents at the floor or ceiling of a HRQoL measure (24,28). A study of healthy children as well as children with up to 2 medical conditions supported that CAT could significantly reduce ceiling and floor effects by increasing score ranges compared to FSF (28). The potential for CAT to minimize floor/ceiling effects is relevant to the JM population as it would allow for better characterization of HRQoL among JM patients with low disease activity or persistent deconditioning despite inactive disease.

Given the need for more robust measurement of HRQoL in JM, this study compares PROMIS CAT and FSF in JM patients and their parents (as proxy) in order to assess whether CAT is: 1) feasible; 2) precise; 3) correlated with FSF without large differences in mean scores; 4) less burdensome than FSF; and 5) less prone to floor/ceiling effects than FSF.

Patients and Methods

Patient Selection and Data Collection

This cross-sectional study of patients with JM was conducted at the Ann & Robert H. Lurie Children’s Hospital of Chicago Cure JM Center of Excellence in Juvenile Myositis Care and Research, a national referral center for patients with JM. Children and adolescents with JM ranging in age from 5-17 years old as well as their parents were enrolled at routine clinic visits. Patient/parent participants were excluded if they had significant cognitive impairment or developmental delay, or if they were unable to complete English language questionnaires. The study received approval from the Ann & Robert H. Lurie Children's Hospital of Chicago Institutional Review Board (IRB# 2016-438) and study activities were in compliance with the Declaration of Helsinki.

For patients 8-17 years old, both self-report and parent proxy measures were administered. For children ranging 5-7 years old, only parent proxy measures were administered. PROMIS CAT and FSF measures assessed the following domains: Mobility, Upper Extremity Function, Anxiety, Depressive Symptoms, Fatigue, and Pain Interference. Participants prospectively completed CAT and FSF via REDCap by using a tablet computer (i.e. iPad) during the clinic visit (30). The CAT measures were administered first as these measures are the focus of the present study, followed by FSF versions of the same measures. A PROMIS domain T-score of 50 represents the average for the original calibration sample that included children with and without a health condition (31). The standard deviation was set at 10 points on the T-score metric (32). Our study used the default PROMIS CAT stopping rule of either a standard error of less than 0.4 (4.0 on a T-score metric) or 12 items administered (but no less than 4 items), whichever occurred first (33). Higher scores in the PROMIS Pediatric Anxiety, Depressive Symptoms, Fatigue, and Pain Interference domains indicate worse symptom burden, while higher scores for PROMIS Pediatric Mobility and Upper Extremity Function indicate better physical function. PROMIS pediatric and parent proxy CAT and FSF measures, including but not limited to domains assessed in our study, can be readily accessed online (https://www.healthmeasures.net/).

Measures of disease severity with acceptable psychometric properties in JM patients were collected at the same study visit at which PROMIS measures were administered in order to describe the study sample, including: Physician Global Assessment of Disease Activity (PGA); Disease Activity Score Muscle and Skin domains (DAS-Muscle, DAS-Skin); muscle enzyme levels (CK, AST, ALT, LDH, Aldolase); and Childhood Myositis Assessment Scale (CMAS), performed and scored by a physical therapist present at all JM clinic visits (13). All other clinical and demographic data was extracted from routine clinical documentation at the study visit.

Statistical Analysis

Descriptive statistics were calculated for all clinical, demographic, and PROMIS data. We calculated the percentage of respondents who were able to complete all self-report and parent proxy PROMIS CAT and FSF in order to ensure that both methods are feasible.

CAT and FSF scores were compared using published methods and criteria (24,34). Paired t-tests (criterion p < 0.05) were used to assess whether T-scores produced by CAT and FSF systematically differ. Pearson correlations (criterion: > 0.7) were used to evaluate the strength of concordance between CAT and FSF scores. Effect size was assessed by calculating Cohen’s D (criterion: < 0.2) in order to quantify the magnitude of any differences between mean CAT and FSF scores while accounting for within subjects variation (34). Pearson correlations were calculated to evaluate agreement between self-report and parent proxy versions of each measure.

In addition to visual presentations of the data, McNemar test was used to compare proportions of extreme scores in CAT versus FSF to determine floor/ceiling effects. PROMIS Mobility and Upper Extremity Function scores at the maximum of the scoring range (indicating best function detectable) were considered as being at the ceiling; PROMIS Anxiety, Depressive Symptoms, Fatigue, and Pain Interference scores at the minimum of the scoring range (indicating lowest detectable symptom intensity) were considered as being at the floor.

To assess precision, the standard errors of each PROMIS self-report and parent proxy domain were plotted against T-scores to generate three curves: CAT, FSF, and “abbreviated CAT”. Abbreviated CAT refers to CAT data where the responses used are truncated such that any administration length of CAT will be less than or equal to its FSF counterpart (e.g. comparison of an abbreviated 8-item CAT with an 8-item FSF). These plots were created to evaluate precision of CAT and FSF across the scoring range and to determine if shortening CAT lengths would impact measurement precision. The dotted line in these plots represents a reliability score of approximately 0.9 (standard error (SE) 3.2), considered a threshold for precise measurement; reliability estimates range from 0 to 1 and are calculated as 1-SE2, with higher values corresponding to higher reliability and therefore greater measurement precision (20,21).

Respondent burden was determined by comparing the number of items administered by CAT with FSF length. The number of items administered by CAT was plotted against T-scores for each PROMIS self-report and parent proxy domain to visualize differences in the number of items administered at different score ranges. These plots were examined in order to determine if at certain T-score ranges more items were administered to reach the standard error specified by the stopping rule.

Results

Data from 67 JM patient-parent dyads was analyzed. Enrolled patients were demographically similar to previously reported JM cohorts with most patients diagnosed with juvenile dermatomyositis (n=61; 91%), female (n=56; 84%), and white (n=51; 76%) and a median age of 11.8 years (IQR 7.4, 15.2) (35,36). Further clinical and demographic data is shown in Table 1.

Table 1.

Patient demographic and clinical characteristics (N=67)

n (%) or median (IQR)
Diagnosis
   Juvenile Dermatomyositis 61 (91)
   Overlap Syndrome 5 (7.5)
   Necrotizing Myopathy 1 (1.5)
Female Sex 56 (83.6)
White Race 51 (76.1)
Age at Disease Onset 5.2 [3.9, 7.1]
Age at First Study Visit 11.8 [7.4, 15.2]
Patients Less Than 8 Years Old 18 (26.9)
Duration of Untreated Disease 125[62, 326.5]
Clinical and Lab Assessments
PGA 0.5 [0, 1.5]
DAS-Total 2 [0, 5]
   DAS-Muscle 0 [0, 1]
   DAS-Skin 1 [0, 5]
CMAS 52 [50, 52]
Muscle Enzymes
   CK 92 [69, 130]
   AST 28.5 [23.0, 32.8]
   LDH 257 [223, 307]
   Aldolase 5.5 [4.8, 6]

IQR= interquartile range; PGA= Physician Global Assessment of Disease Activity; DAS= Disease Activity Score (Muscle and Skin domains); CMAS= Childhood Myositis Assessment Scale; CK = Creatine kinase; AST = Aspartate aminotransferase; ALT = Alanine aminotransferase; LDH = Lactate dehydrogenase

CMAS performed and scored on a scale from 0-52 by a physical therapist present at all JM clinic visits

Feasibility

CAT had 100% completion and all but one patient-parent dyad completed FSF, demonstrating feasibility of both methods.

Comparison of CAT and FSF scores

As shown in Table 2, PROMIS CAT and FSF were highly correlated, with Pearson’s correlation coefficients ranging 0.87-0.92 and 0.79-0.9 for self-report and parent proxy respectively (Table 2). Significant differences between parent proxy CAT and FSF on Anxiety and Fatigue with small effect sizes (0.23 and 0.19, respectively) were noted. However, the magnitude of differences in mean CAT and FSF Anxiety and Fatigue scores was smaller than clinically meaningful differences defined in JM and other pediatric populations (37,38). No significant differences were found for any other comparisons.

Table 2.

Comparison of PROMIS Computerized Adaptive Testing-Administered Item Banks and Fixed Short Forms

CAT
FSF
CAT vs. FSF
PROMIS
Measure:
# Items,
Mean
(SD)
T-Score*
Mean (SD)
#
Items
T-Score*
Mean (SD)
Corr t-Test
(p-value)
Cohen’s D
(95% CI)§
SELF-REPORT
Anxiety 9.6 (2.8) 41.2 (10.2) 8 40.9 (9.7) 0.91 0.786 0.02
(−0.11, 0.14)
Depressive Symptoms 8.2 (3.3) 45 (11.5) 8 44.9 (10.6) 0.89 0.949 0.004
(−0.13, 0.14)
Fatigue 9.4 (3) 39.7 (13.4) 10 39.4 (10.8) 0.92 0.697 0.02
(−0.09, 0.13)
Pain Interference 9.3 (3.4) 39.2 (8.9) 8 38.7 (7.2) 0.87 0.405 0.06
(−0.08, 0.2)
Mobility 9.8 (3.2) 53.5 (8.6) 8 53.8 (7.1) 0.92 0.594 −0.03
(−0.14, 0.08)
Upper Extremity 10.6 (2.6) 51.7 (8.5) 8 52.5 (8) 0.92 0.078 −0.1
(−0.21, 0.01)
PARENT-PROXY
Anxiety 8.1 (3.4) 44.4 (10.1) 8 42.1 (10.2) 0.90 < 0.001 0.23
(0.12, 0.34)
Depressive Symptoms 7.2 (3.2) 45.1 (10.2) 6 43.4 (9.9) 0.79 0.056 0.15
(0, 0.31)
Fatigue 7.7 (3.3) 43.7 (10.8) 10 41.6 (9.2) 0.83 0.012 0.19
(0.04, 0.33)
Pain Interference 8.6 (3.5) 43.4 (8.4) 8 42.9 (7.9) 0.82 0.124 0.11
(−0.03, 0.26)
Mobility 9.2 (3.4) 51.3 (8.9) 8 51.1 (7.8) 0.84 0.533 0.04
(−0.09, 0.18)
Upper Extremity 9.4 (3.3) 47.1 (8.9) 8 47.9 (9.1) 0.86 0.281 −0.07
(−0.2, 0.06)
*

Higher score for the Anxiety, Depressive Symptoms, Fatigue and Pain Interference domains indicates worse symptoms while higher score for Mobility and Upper Extremity Function indicates better function

For CAT, sample size for self-report was n=49 for all domains and n=67 for all parent proxy domains

For FSF, sample size for self-report was n=49 for all domains except Anxiety and Depressive Symptoms (n=48) and n=66 for all parent-proxy domains

Pearson’s Correlations between CAT and FSF

§

Measurement of mean effect size (95% confidence interval) of differences between CAT and FSF

Both CAT and FSF demonstrated moderate-to-high patient-parent correlations. Pearson’s correlations between CAT self-report and parent proxy were: Anxiety 0.65, Depressive Symptoms 0.67, Fatigue 0.57, Mobility 0.85, Pain Interference 0.59, and Upper Extremity 0.75. Pearson’s correlations between FSF self-report and parent proxy were: Anxiety 0.53, Depressive Symptoms 0.63, Fatigue 0.60, Mobility 0.72, Pain Interference 0.54, and Upper Extremity 0.55.

Respondent Burden

Table 2 presents the mean number of items administered by CAT compared with the length of each FSF. CAT did not administer fewer items than FSF on average for any domains except for Fatigue for which the mean number of items administered by CAT was 0.6 and 2.3 less than that for FSF for self-report and parent proxy respectively. The remaining CAT self-report measures administered between 0.2 and 2.6 more items on average (Table 2). For other CAT parent proxy measures (besides Fatigue), between 0.1 and 1.4 more items were administered on average compared with FSF (Table 2). As shown in Supplementary Figures 1-3, respondents with extreme scores (i.e., least symptomatic or best functioning) frequently endorsed the maximum pre-set number of items (n = 12) as specified by the current PROMIS CAT stopping rule.

Floor and Ceiling Effects

Table 3 shows the percent of respondents at ceiling or floor for each PROMIS Pediatric domain. CAT had lower floor/ceiling effects than FSF in all domains and this difference is statistically significant in all self-report and parent proxy domains except for self-report Anxiety.

Table 3.

Floor and Ceiling Effects*

FSF
CAT
FSF vs CAT
Domains n % floor/ceiling n % floor/ceiling p-value
SELF-REPORT
  Anxiety 48 25 (52%) 49 21 (43%) 0.289
  Depressive Symptoms 48 20 (42%) 49 14 (29%) 0.041
  Fatigue 49 22 (45%) 49 13 (27%) 0.016
  Pain Interference 49 31(63%) 49 25 (51%) 0.041
  Mobility 49 32 (65%) 49 18 (37%) 0.001
  Upper Extremity 49 37 (76%) 49 31 (63%) 0.041
PARENT-PROXY
  Anxiety 66 36 (55%) 67 22 (33%) 0.001
  Depressive Symptoms 66 36 (55%) 67 17 (25%) <0.001
  Fatigue 66 32 (49%) 67 14 (21%) <0.001
  Pain Interference 66 43 (65%) 67 33 (49%) 0.009
  Mobility 66 42 (64%) 67 27 (40%) 0.001
  Upper Extremity 66 40 (61%) 67 31 (46%) 0.008
*

Anxiety, Depressive Symptoms, Fatigue, and Pain Interference domains demonstrate floor effects at the less symptomatic extremes, while Mobility and Upper Extremity Function demonstrate ceiling effects at higher functioning extremes

Measurement Precision

When plotting standard error against T-score for each self-report and parent proxy domain, higher standard error is noted at the less symptomatic extremes (Figures 1-3). Standard error for FSF is typically lower in the middle T-score range, but at the low symptom/high function extremes CAT demonstrates lower standard error in most domains (Figures 1-3). Plots of standard errors across the T-score range generally form a ‘U-shape’ (Figures 2-3), with the exceptions of the Mobility and Upper Extremity Function measures which form a ‘J-shape’ (Figure 1), consistent with prior studies (21,28,29). Abbreviated CAT (i.e. item administration capped at number of items in FSF) standard error curves for all domains are nearly indistinguishable from CAT (Figures 1-3).

Figure 1. Plotting standard error (y-axis) against T-score (x-axis) for PROMIS self-report and parent proxy Mobility and Upper Extremity Function measures.

Figure 1.

Standard error rises as PROMIS Mobility and Upper Extremity T-scores increase, forming a ‘J-shape’ curve. Abbreviated CAT standard error curve approximates that of CAT, suggesting similar measurement precision across the scoring range.

Figure 3. Plotting standard error (y-axis) against T-score (x-axis) for PROMIS self-report and parent proxy Anxiety and Depressive Symptoms measures.

Figure 3.

Standard error is highest for PROMIS Anxiety and Deperessive Symptoms in the low T-score range, lowest in the middle T-score range, and rises modestly at the higher T-score range, forming a ‘U-shape’ curve. Abbreviated CAT standard error curve approximates that of CAT, suggesting similar measurement precision across the scoring range.

Figure 2. Plotting standard error (y-axis) against T-score (x-axis) for PROMIS self-report and parent proxy Fatigue and Pain Interference measures.

Figure 2.

Standard error is highest for PROMIS Fatigue and Pain Interference in the low T-score range, lowest in the middle T-score range, and rises modestly at the higher T-score range, forming a ‘U-shape’ curve. Abbreviated CAT standard error curve approximates that of CAT, suggesting similar measurement precision across the scoring range.

CAT administered more items in the low symptom/high function extremes of the T-score range for all self-report and parent proxy domains (Supplementary Figures 1-3). For domains such as self-report Depressive Symptoms, Pain Interference, and Upper Extremity and parent proxy Fatigue, Depressive Symptoms and Pain Interference, up to 12 items were usually administered for the low symptom/high function T-score range while 5 items were typically administered across most of the remaining T-score range, indicating a bimodal distribution of items administered in relation to T-score range (Supplementary Figures 1-3).

Discussion

Our study provides support for the use of PROMIS CAT measures, suggesting that these measures are feasible, do not demonstrate large differences in mean scores when compared to most FSF measures, and are less prone to floor/ceiling effects than FSF. This study has several notable strengths. Our study was conducted at a national referral center including patients outside the immediate geographic region, allowing for enrollment of a relatively large sample that was representative of other JM study cohorts despite the overall rarity of JM (35,36). The study sample was well-characterized using prospectively collected, psychometrically strong outcome measures scored by experienced clinicians and physical therapists. Below we discuss some promising aspects of CAT as well as considerations for further optimizing its use in HRQoL measurement for patients with JM and other pediatric rheumatic diseases.

We found both CAT and FSF had very high completion rates, providing support that both methods are feasible. While investigators can customize items to be included in FSFs based on specific needs in their study population, in this study, we used the default FSFs and found only modest differences between CAT and FSF scores on parent proxy Anxiety and Fatigue. As seen in other patient populations, correlations between CAT and FSF are high, reassuring that CAT and FSF generate scores measuring similar constructs (21,24,28,39). Some modest differences in mean PROMIS CAT versus FSF scores were noted in parent proxy Anxiety and Fatigue domains but these differences are unlikely to be clinically meaningful as they are smaller than previously described minimal clinically important differences in PROMIS pediatric measures (37,38). Both CAT and FSF demonstrated comparably high correlations between self-report and parent proxy measures; however, self-report and parent proxy data did not completely correlate, consistent with studies in other pediatric rheumatic diseases suggesting that parent proxy and patient perspectives are different and parent proxy should not replace child self-report or vice versa (21,40,41).

We enrolled a convenience sample of prevalent, rather than incident, patients with JM into our study, with relatively low disease activity and long disease duration, in whom CAT measures were hypothesized to better detect subtle deficits (7). Unsurprisingly, PROMIS FSF measures demonstrated high rates of floor/ceiling effects, often greater than or equal to floor/ceiling effects noted in nationally representative pediatric cohorts (42). However, PROMIS CAT measures (except PROMIS pediatric self-report Anxiety) demonstrated significantly reduced floor/ceiling effects compared with FSF. In particular, the floor/ceiling effects for PROMIS self-report and parent proxy Mobility as well as PROMIS parent proxy Fatigue and Upper Extremity Function measures in low symptom/high function JM patients were dramatically less pronounced than floor/ceiling effects for FSF in both our study and those reported in nationally representative cohorts (42). In general, these findings are consistent with published studies demonstrating CAT administrations produce a broader scoring range with corresponding decrease in floor/ceiling effects when compared with FSF (28,29,43). However, there were some exceptions, with PROMIS Anxiety (self-report and parent-proxy) FSF and Pain Interference (self-report) FSF demonstrating greater floor effects than those noted in nationally representative pediatric cohorts; for these domains, CAT did decrease floor effects but they remained greater than or equal to published rates of PROMIS FSF floor effects (42). Additionally, differences in floor effects for PROMIS Anxiety (self-report) measures did not reach statistical significance, possibly due to inadequate statistical power given the somewhat higher standard errors noted for this measure than for other domains (Figure 3) and the relatively small pediatric self-report sample size. Nevertheless, for most key outcomes (e.g. physical function, fatigue) in JM, CAT’s ability to detect subtler deficits may be preferable when monitoring long-term JM patient outcomes.

Our findings suggest that modification of the default PROMIS pediatric CAT stopping rule may be worth considering in order to minimize respondent burden among patients with low disease activity. The default stopping rule is met when participants’ responses reach a pre-specified precision level or a maximum of 12 items has been administered. Self-report and parent proxy CAT and FSF in our study produced reliable measurements, often reaching standard errors low enough to meet the stopping rule’s pre-specified precision level in patients with higher symptom/lower function responses (Figures 1-3). However, since our study recruited a prevalent (not incident) sample of patients with JM, most participants had low disease activity as they had been treated for longer periods. Consistent with prior literature, respondents with PROMIS T-scores in the low symptom/high function ranges were frequently administered the maximum 12 items (Supplementary Figures 1-3) since patients in this portion of the T-score range tended to have the highest standard errors (Figures 1-3) (21,26,28,39). We also found that abbreviated CAT (i.e. number of items administered is capped at the number of items in FSF) could successfully decrease respondent burden without diminishing measurement precision, replicating findings in other samples (Figures 1-3) (25-27,29,44). Further studies are recommended in order to optimize stopping rules for children with JM, especially for less symptomatic/higher functioning patients who may nevertheless be experiencing ongoing deficits.

Limitations of this study may provide directions for future research. One limitation of this study was that participants were recruited at a single referral center, potentially limiting generalizability. While the sample size was relatively large for this rare disease population, in absolute terms it was nevertheless small and this limited statistical power to detect some group differences. Therefore, we call for future multicenter studies to assess whether our findings can be replicated in larger, more geographically dispersed samples. The low disease activity in our study sample could represent a potential limitation. Hypothesizing that CAT may be particularly beneficial for overcoming floor/ceiling effects, we purposely recruited a convenience sample of prevalent (not only incident) JM cases with low average disease activity and longer disease duration similar to other described cohorts (7,36). However, the relative performance and benefits of CAT versus FSF may differ in incident cases of JM and/or those with higher disease activity, so future studies should also evaluate CAT and FSF in inception cohorts and refractory patients. Additionally, abbreviated CAT scores were constructed and compared post-hoc with CAT and FSF; future studies should prospectively assess the impact of modified stopping rules on respondent burden and psychometric performance of CAT compared with FSF. Finally, since PROMIS measures utilize item response theory, it is possible to create customized FSF (i.e. selecting items tailored to a particular patient population); our study utilized default FSFs and future studies should evaluate the relative merits of CAT versus customized FSFs.

In summary, this study suggests that CAT decreases the ceiling/floor effects among low symptom/high function JM patients when compared to FSF. Our study supports modification of stopping rules to decrease respondent burden, particularly for less symptomatic patients. CAT appears to be feasible and can be used for longitudinal assessment of JM patients to help detect subtler deficits in HRQoL than FSF. We recommend that CAT be considered for use in clinical and research settings as a tool to better monitor JM patients as we continue to strive for improved HRQoL outcomes for this population.

Supplementary Material

fS1
fS2
fS3

Significance and Innovations.

  • PROMIS CAT appears feasible in JM patients and PROMIS CAT scores are correlated with FSF scores.

  • PROMIS CAT appears to have less pronounced ceiling and floor effects compared with FSF in JM patients.

  • PROMIS CAT may administer more items to less symptomatic scorers but modified CAT stopping rules can decrease respondent burden without sacrificing precision.

  • PROMIS CAT may be better suited in discerning subtle disease activity or changes in JM patients during long-term monitoring in clinical and research settings.

Acknowledgements:

We are grateful to the patients and parents who participated in this study for generously lending their time and insight. We would also like to thank Cure JM Foundation and the Rheumatology Research Foundation for support of this study.

This publication was made possible with support from the Rheumatology Research Foundation Medical Student Preceptorship (R Patel, K Ardalan, EL Gray) and Cure JM Foundation (K Ardalan, EL Gray). Research reported in this publication was supported, in part, by the National Institutes of Health’s National Center for Advancing Translational Sciences, Grant Number UL1TR001422, and National Institutes of Health, Grant Number U2CCA186878 (D Cella). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Dr. Ardalan reports receiving grant funding from Cure JM Foundation to assist with clinical trial design with ReveraGen Biopharma, however he receives no funding or other compensation from ReveraGen Biopharma.

References

  • 1.Huber A, Feldman BM. Long-term outcomes in juvenile dermatomyositis: how did we get here and where are we going? Current Rheumatology Reports 2005;7:441–446. [DOI] [PubMed] [Google Scholar]
  • 2.Apaz MT, Saad-Magalhães C, Pistorio A, Ravelli A, De Oliveira Sato J, Marcantoni MB, et al. Health-related quality of life of patients with juvenile dermatomyositis: Results from the Paediatric Rheumatology International Trials Organisation multinational quality of life cohort study. Arthritis Care & Research 2009;61:509–517. [DOI] [PubMed] [Google Scholar]
  • 3.Huber AM, Lang B, LeBlanc CM, Birdi N, Bolaria RK, Malleson P, et al. Medium- and long-term functional outcomes in a multicenter cohort of children with juvenile dermatomyositis. Arthritis & Rheumatism 2000;43:541–549. [DOI] [PubMed] [Google Scholar]
  • 4.Sanner H, Kirkhus E, Merckoll E, Tollisen A, Røisland M, Lie BA, et al. Long-term muscular outcome and predisposing and prognostic factors in juvenile dermatomyositis: a case-control study. Arthritis Care & Research 2010;62:1103–1111. [DOI] [PubMed] [Google Scholar]
  • 5.Tsaltskan V, Aldous A, Serafi S, Yakovleva A, Sami H, Mamyrova G, et al. Long-term outcomes in juvenile myositis patients. Seminars in Arthritis and Rheumatism 2020;50:149–155. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Tollisen A, Sanner H, Flatø B, Wahl AK. Quality of life in adults with juvenile-onset dermatomyositis: a case-control study. Arthritis Care & Research 2012;64:1020–1027. [DOI] [PubMed] [Google Scholar]
  • 7.Ravelli A, Trail L, Ferrari C, Ruperto N, Pistorio A, Pilkington C, et al. Long-term outcome and prognostic factors of juvenile dermatomyositis: a multinational, multicenter study of 490 patients. Arthritis Care & Research 2010;62:63–72. [DOI] [PubMed] [Google Scholar]
  • 8.Livermore P, Gray S, Mulligan K, Stinson JN, Wedderburn LR, Gibson F. Being on the juvenile dermatomyositis rollercoaster: a qualitative study. Pediatric Rheumatology 2019;17:30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ardalan K, Adeyemi O, Wahezi DM, Caliendo AE, Curran ML, Neely J, et al. Parent perspectives on addressing emotional health for children and young adults with juvenile myositis. Arthritis Care & Research 2021;73:18–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Tory HO, Carrasco R, Griffin T, Huber AM, Kahn P, Robinson AB, et al. Comparing the importance of quality measurement themes in juvenile idiopathic inflammatory myositis between patients and families and healthcare professionals. Pediatric Rheumatology 2018;16:28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lazarevic D, Pistorio A, Palmisani E, Miettunen P, Ravelli A, Pilkington C, et al. The PRINTO criteria for clinically inactive disease in juvenile dermatomyositis. Annals of the Rheumatic Diseases 2013;72:686–693. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Rider LG, Aggarwal R, Pistorio A, Bayat N, Erman B, Feldman BM, et al. 2016 American College of Rheumatology/European League Against Rheumatism Criteria for minimal, moderate, and major clinical response in juvenile dermatomyositis: an International Myositis Assessment and Clinical Studies Group/Paediatric Rheumatology International Trials Organisation collaborative initiative. Arthritis & Rheumatology 2017;69:911–923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Rider LG, Werth VP, Huber AM, Alexanderson H, Rao AP, Ruperto N, et al. Measures of adult and juvenile dermatomyositis, polymyositis, and inclusion body myositis: Physician and Patient/Parent Global Activity, Manual Muscle Testing (MMT), Health Assessment Questionnaire (HAQ)/Childhood Health Assessment Questionnaire (C-HAQ), Childhood Myositis Assessment Scale (CMAS), Myositis Disease Activity Assessment Tool (MDAAT), Disease Activity Score (DAS), Short Form 36 (SF-36), Child Health Questionnaire (CHQ), Physician Global Damage, Myositis Damage Index (MDI), Quantitative Muscle Testing (QMT), Myositis Functional Index-2 (FI-2), Myositis Activities Profile (MAP), Inclusion Body Myositis Functional Rating Scale (IBMFRS), Cutaneous Dermatomyositis Disease Area and Severity Index (CDASI), Cutaneous Assessment Tool (CAT), Dermatomyositis Skin Severity Index (DSSI), Skindex, and Dermatology Life Quality Index (DLQI). Arthritis Care & Research 2011;63:S118–S157. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Huber AM, Hicks JE, Lachenbruch PA, Perez MD, Zemel LS, Rennebohm RM, et al. Validation of the Childhood Health Assessment Questionnaire in the juvenile idiopathic myopathies. Juvenile Dermatomyositis Disease Activity Collaborative Study Group. The Journal of Rheumatology 2001;28:1106–1111. [PubMed] [Google Scholar]
  • 15.Groen W, Ünal E, Nørgaard M, Maillard S, Scott J, Berggren K, et al. Comparing different revisions of the Childhood Health Assessment Questionnaire to reduce the ceiling effect and improve score distribution: Data from a multi-center European cohort study of children with JIA. 2010;8:16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Dempster H, Porepa M, Young N, Feldman BM. The clinical meaning of functional outcome scores in children with juvenile arthritis. Arthritis & Rheumatism 2001;44:1768–1774. [DOI] [PubMed] [Google Scholar]
  • 17.Varni JW, Seid M, Smith Knight T, Burwinkle T, Brown J, Szer IS. The PedsQL in pediatric rheumatology: reliability, validity, and responsiveness of the Pediatric Quality of Life Inventory Generic Core Scales and Rheumatology Module. Arthritis & Rheumatism 2002;46:714–725. [DOI] [PubMed] [Google Scholar]
  • 18.Khanna D, Krishnan E, Dewitt EM, Khanna PP, Spiegel B, Hays RD. The future of measuring patient-reported outcomes in rheumatology: Patient-Reported Outcomes Measurement Information System (PROMIS®). Arthritis Care & Research 2011;63:S486–S490. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.DeWalt DA, Rothrock N, Yount S, Stone AA. Evaluation of item candidates: the PROMIS qualitative item review. Medical Care 2007;45:S12–S21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Reeve BB, Hays RD, Bjorner JB, Cook KF, Crane PK, Teresi JA, et al. Psychometric evaluation and calibration of health-related quality of life item banks: plans for the Patient-Reported Outcomes Measurement Information System (PROMIS). Medical Care 2007;45:S22–S31. [DOI] [PubMed] [Google Scholar]
  • 21.Brandon TG, Becker BD, Bevans KB, Weiss PF. Patient-Reported Outcomes Measurement Information System tools for collecting patient-reported outcomes in children with juvenile arthritis. Arthritis Care & Research 2017;69:393–402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Jones JT, Carle AC, Wootton J, Liberio B, Lee J, Schanberg LE, et al. Validation of Patient-Reported Outcomes Measurement Information System short forms for use in childhood-onset systemic lupus erythematosus. Arthritis Care & Research 2017;69:133–142. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Ardalan K, Cella D, Pachman LM, Gray EL, Lee J, Fahey K, et al. Initial validation of Patient-Reported Outcomes Measurement Information System (PROMIS) in children with juvenile myositis [abstract]. Arthritis & Rheumatology 2017;69. [Google Scholar]
  • 24.Lai J-S, Beaumont JL, Nowinski CJ, Cella D, Hartsell WF, Han-Chih Chang J, et al. Computerized adaptive testing in pediatric brain tumor clinics. Journal of Pain and Symptom Management 2017;54:289–297. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kallen MA, Cook KF, Amtmann D, Knowlton E, Gershon RC. Grooming a CAT: customizing CAT administration rules to increase response efficiency in specific research and clinical settings. Quality of Life Research 2018;27:2403–2413. [DOI] [PubMed] [Google Scholar]
  • 26.Choi SW, Reise SP, Pilkonis PA, Hays RD, Cella D. Efficiency of static and computer adaptive short forms compared to full-length measures of depressive symptoms. Quality of Life Research 2010;19:125–136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Segawa E, Schalet B, Cella D. A comparison of computer adaptive tests (CATs) and short forms in terms of accuracy and number of items administrated using PROMIS profile. Quality of Life Research 2020;29:213–221. [DOI] [PubMed] [Google Scholar]
  • 28.Varni JW, Magnus B, Stucky BD, Liu Y, Quinn H, Thissen D, et al. Psychometric properties of the PROMIS® pediatric scales: precision, stability, and comparison of different scoring and administration options. Quality of Life Research 2014;23:1233–1243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Luijten MAJ, Terwee CB, Van Oers HA, Joosten MMH, Van Den Berg JM, Schonenberg-Meinema D, et al. Psychometric properties of the pediatric Patient-Reported Outcomes Measurement Information System item banks in a Dutch clinical sample of children with juvenile idiopathic arthritis. Arthritis Care & Research 2020;72:1780–1789. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)—A metadata-driven methodology and workflow process for providing translational research informatics support. Journal of Biomedical Informatics 2009;42:377–381. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Irwin DE, Stucky BD, Thissen D, DeWitt EM, Lai JS, Yeatts K, et al. Sampling plan and patient characteristics of the PROMIS pediatrics large-scale survey. Qual Life Res 2010;19:585–594. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Anon. PROMIS pediatric and proxy profile scoring manual. 2021. Available at: http://www.healthmeasures.net/images/PROMIS/manuals/PROMIS_Pediatric_and_Proxy_Profile_Scoring_Manual.pdf. Accessed January 31, 2021.
  • 33.Anon. Computer Adaptive Tests (CATs). Health Measures. Available at: https://www.healthmeasures.net/index.php?option=com_content&view=category&layout=blog&id=164&Itemid=1133. Accessed January 14, 2021. [Google Scholar]
  • 34.Cohen J. Statistical power analysis for the behavioral sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. [Google Scholar]
  • 35.Phillippi K, Hoeltzel M, Robinson AB, Kim S, for the Childhood Arthritis and Rheumatology Research Alliance (CARRA) Legacy Registry Investigators. Race, income, and disease outcomes in juvenile dermatomyositis. The Journal of Pediatrics 2017;184:38–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Robinson AB, Hoeltzel MF, Wahezi DM, Becker ML, Kessler EA, Schmeling H, et al. Clinical characteristics of children with juvenile dermatomyositis: the Childhood Arthritis and Rheumatology Research Alliance Registry. Arthritis Care & Research 2014;66:404–410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Thissen D, Liu Y, Magnus B, Quinn H, Gipson DS, Dampier C, et al. Estimating minimally important difference (MID) in PROMIS pediatric measures using the scale-judgment method. Quality of Life Research 2016;25:13–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Wolfe M, Robinson A, Lai JS, Coles T, Gray E, Chang R, et al. Estimation of clinically important differences in Patient-Reported Outcomes Measurement Information System (PROMIS) measures in juvenile myositis [abstract]. Arthritis & Rheumatology 2020;72. [Google Scholar]
  • 39.Amtmann D, Bamer AM, Kim J, Bocell FD, Chung H, Park R, et al. A comparison of computerized adaptive testing and fixed-length short forms for the Prosthetic Limb Users Survey of Mobility (PLUS-M ™ ). Prosthetics and Orthotics International 2018;42:476–482. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Garcia-Munitis P, Bandeira M, Pistorio A, Magni-Manzoni S, Ruperto N, Schivo A, et al. Level of agreement between children, parents, and physicians in rating pain intensity in juvenile idiopathic arthritis. Arthritis Care & Research 2006;55:177–183. [DOI] [PubMed] [Google Scholar]
  • 41.Gaultney AC, Bromberg MH, Connelly M, Spears T, Schanberg LE. Parent and child report of pain and fatigue in JIA: does disagreement between parent and child predict functional outcomes? Children 2017;4:11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Carle AC, Bevans KB, Tucker CA, Forrest CB. Using nationally representative percentiles to interpret PROMIS pediatric measures. Quality of Life Research 2020;Online ahead of print. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Craig J, Feldman BM, Spiegel L, Dover S. Comparing the measurement properties and preferability of patient reported outcome measures in pediatric rheumatology: PROMIS versus CHAQ. Journal of Rheumatology 2020;Online ahead of print. [DOI] [PubMed] [Google Scholar]
  • 44.Choi SW, Grady MW, Dodd BG. A new stopping rule for computerized adaptive testing. Educational and Psychological Measurement 2011;71:37–53. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

fS1
fS2
fS3

RESOURCES