Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2022 Apr 1.
Published in final edited form as: J Pain. 2020 Oct 24;22(4):386–399. doi: 10.1016/j.jpain.2020.10.003

III. Detecting treatment effects in clinical trials with different indices of pain intensity derived from ecological momentary assessment

Stefan Schneider 1, Doerte U Junghaenel 1, Masakatsu Ono 1, Joan E Broderick 1, Arthur A Stone 1,2
PMCID: PMC8043984  NIHMSID: NIHMS1640545  PMID: 33172597

Abstract

Pain intensity represents the primary outcome in most pain clinical trials. Identifying methods to measure aspects of pain that are most sensitive to treatment may facilitate discovery of effective interventions. In this third of three articles examining alternative indices of pain intensity derived from ecological momentary assessments (EMA), we compare treatment effects based on Average Pain, Maximum Pain, Minimum Pain, Pain Variability, Time in High Pain, Time in Low Pain, and Pain After Wake-Up. We also examine which indices contribute to Patient Global Impressions of Change (PGIC). Data came from two randomized, double-blind, placebo-controlled trials examining the efficacy of milnacipran for fibromyalgia treatment; 2084 patients provided >1 million EMA pain intensity ratings over 24 (Study 1) or 26 (Study 2) treatment weeks. Pain Variability and Time in High Pain produced significantly smaller treatment effects than Average Pain; other pain indices showed effects that were numerically smaller, but not significantly different from Average Pain. Changes in all pain indices were significantly associated with PGIC, with improvements in Maximum Pain and in Pain Variability offering small incremental contributions to understanding PGIC over Average Pain. Results suggest that different pain indices could be used to detect treatment effects in pain clinical trials.

Perspective

Alternative summary measures of pain intensity derived from EMA may broaden the scope of outcomes useful in pain clinical trials. In this analysis of a pharmacological treatment for fibromyalgia, most pain summary measures indicated similar effects; improvements in Maximum Pain and Pain Variability contributed to understanding PGIC over Average Pain.

Keywords: Pain intensity, Ecological Momentary Assessment, Pain indices, Treatment effects, Clinical trials, Patient Global Impression of Change

Introduction

Pain intensity is the primary outcome in most clinical trials of pain disorders [14]. Although a substantial body of research has been devoted to determining reliable and valid methods of self-reported pain assessment [25], an important question is whether and how the information obtained from pain intensity measures could be improved to enhance detection of treatment effects [23,24,47,48,53]. The overall amount of pain (typically conceptualized as the average pain level over a day or week) has served as the most common pain intensity outcome in many clinical trials. However, an exclusive focus on patients’ average pain level may miss important effects of treatment on patients’ pain experiences in daily life. For example, the Initiative on Methods, Measurement, and Pain Assessment in Clinical Trials (IMMPACT) suggests that measures of the temporal aspects of pain, including the duration of pain states and the variability in pain intensity have not received adequate attention as clinical trial outcomes [14]. In addition, the United States Food and Drug Administration (FDA) draft guidance on analgesic drug development recommends assessing patients’ worst pain levels rather than changes in average pain as primary outcome in trials for analgesic medications [57]. To date, however, these calls for alternative pain intensity measures are not supported by a systematic evidence base.

In the present research, we present a strategy to construct new outcome measures capturing different aspects of pain intensity from ecological momentary assessments (EMA) [46] and test whether the new measures are sensitive to detecting treatment effects. The use of EMA as a method to collect outcomes in clinical trials has not been widely considered, but its potential benefits rest on two central aspects. First, momentary measurement reduces recall biases and error, capturing patients’ current pain and not their retrospective memories or beliefs about their pain [7,8]. Second, ratings of momentary pain are obtained multiple times per day. This makes it possible to obtain a range of different pain summary measures that can serve as endpoints in clinical trials [30,51].

In addition to average pain levels derived from EMA, we evaluate whether indices of maximum (worst) and minimum (least) pain are sensitive to detecting treatment effects, given that they have been recommended as outcomes in clinical trials [4,14,48,57]. The amount of time patients spent either in low pain or in high pain considers alternative indices of potential importance that have garnered much lesser attention and these emphasize the frequency and duration of pain at low or high levels. The available evidence suggests that pain frequency represents a distinctive feature of pain [14,26,44,49], and our previous work suggests that an index of pain frequency may be especially sensitive to change due to treatment [51]. Additionally, we explore indices capturing the amount of variability in pain [1,22,42] and diurnal aspects (pain after wake-up) [5] of patients' pain intensity to examine whether they can provide useful information on treatment efficacy.

This is the third in a series of three articles examining alternative pain indices derived from EMA to determine those with the greatest potential for chronic pain research. The first and second article explored which indices are most preferred by stakeholders and which contribute most to patient functioning [41,50]. Here, we re-examine two large analgesic clinical trials for the treatment of fibromyalgia to evaluate the utility of different pain indices for detecting intervention effects. In addition, we examine the extent to which changes in the different pain indices relate to Patient Global Impressions of Change (PGIC) over the treatment period, given that PGIC ratings are commonly used as anchors for evaluating clinically significant change [16,17,19]. To the extent that alternative pain indices relate to PGIC above changes in "average" pain, they may contribute additional information on intervention effects that patients perceive as clinically meaningful.

Methods

Study design and sample

This study is based on secondary analyses using data from two published Phase 3 clinical trials. Both studies were randomized, double-blind, placebo-controlled trials examining the efficacy and safety of milnacipran for the treatment of fibromyalgia. Detailed information about study design and study entry criteria is presented in the original published articles [9,29]; a Cochrane systematic review presents summaries of the original trial results and methodology [12]. Briefly, and most relevant to the present analyses, both studies involved the comparison of two doses of milnacipran – 100 or 200 milligram per day – to a placebo control group, in adults between 18 and 70 years of age who met the American College of Rheumatology criteria for fibromyalgia. After an initial medication washout period of 1 to 4 weeks, patients entered a pre-treatment assessment period of 2 weeks during which they received training in the use of an electronic diary for EMA and daily data collection. To continue in the study, participants needed to complete at least 70% of the random EMA prompts during the pre-treatment period. They were then randomized and entered a 3-week dose escalation period, followed by a period of stable dose treatment that lasted 24 weeks (in study 1; [29]) or 26 weeks (in study 2; [9]). Primary end-points in the original reports included daily 24-hour recall pain ratings and questionnaires of patient global impression of change and physical functioning. To conduct the present analyses using momentary pain ratings, we obtained electronic copies of the de-identified primary data files and annotated case report forms from the sponsor. The University of Southern California Institutional Review Board approved the secondary analysis project.

Participants in study 1 were on average 49.4 (SD = 10.7) years of age, predominantly female (95.6%) and White (93.6%), with an average duration of fibromyalgia of 5.7 (SD = 5.4) years. Of 1639 patients screened, 888 were randomized into the 100 mg/day (n = 224), 200 mg/day (n = 441), or placebo (n = 223) groups. Of those who were randomized, 57.1% (100 mg/day group), 54.2% (200 mg/day group), and 65% (placebo group) completed the 27-week treatment period.

Participants in study 2 had a mean age of 50.2 (SD = 10.6) years, were 96.2% female and 93.4% White, with an average disease duration of 9.7 (SD = 8.2) years. Of 2270 patients who were screened, 1207 were initially randomized but 11 were withdrawn from the analysis sample, leaving an intent-to-treat sample of n = 1196 assigned to the 100 mg/day (n = 399), 200 mg/day (n = 396), or placebo control (n = 401) groups. Of those, 39.1% (100 mg/day group), 35.4% (200 mg/day group), and 40.4% (placebo group) completed the full 29-week period.

Collection of Patient Global Impression of Change (PGIC)

PGIC scores were assessed at with a single question administered at treatment weeks 3, 7, 11, 15, 19, 23, and 27 after randomization. The question was worded “Since the start of the study, overall my fibromyalgia is”, with categorical response options 1= very much improved, 2 = much improved, 3 = minimally improved, 4 = no change, 5=minimally worse, 6 = much worse, 7 = very much worse.

Collection of momentary pain intensity ratings

Momentary pain intensity ratings were collected on an electronic diary with proprietary software (invivodata, inc., Scotts Valley, CA) in the morning after participants woke up, in response to approximately 3-4 random prompts throughout the day, and in the evening (at about 8 PM), for a total of approximately 5-6 measurement time-points per day. EMA pain ratings were assessed each day over the full course of the baseline and treatment period. At each EMA measurement occasion, participants were asked to “rate your current level of pain” using a visual analogue scale with response anchors 0 = no pain to 100 = worst possible pain.

Construction of summary measures (“pain intensity indices”) from momentary pain data

The different pain intensity indices derived from the momentary pain data for the purposes of the present study are illustrated in Figure 1. In constructing the different pain indices, we had to make a decision about the time interval over which the momentary data were to be summarized to create an outcome measure. We decided to use 7-day intervals, given that a week is a common reporting period for pain clinical trials [8,27], and because creating the measures for shorter (e.g., 24-hr) periods would likely have generated unreliable indices that would have needed to be aggregated (e.g., averaged) over multiple days (e.g., a week).

Figure 1:

Figure 1:

Illustration of indices derived from momentary pain ratings, using data from a single participant at the baseline week.

A measure of each patient’s Average Pain Level was created by taking the arithmetic mean of all pain intensity ratings over the course of a given week. The average pain level is the most common pain outcome measure, and it served as a comparison for the treatment effects obtained from the other pain summary measures. The Maximum Pain Level was defined as each patient’s highest momentary pain rating over the week, and the Minimum Pain Level as the lowest momentary pain rating over the week, in accordance with previous conceptualizations of patients’ worst and least pain levels [52]. To obtain an indicator of Pain Variability, we calculated the intraindividual standard deviation (SD) of patients’ momentary pain ratings over the week, in line with previous studies on pain variability [18,22,54]. The intraindividual SD is arguably the most commonly used measure of variability [30]. It directly reflects the magnitude of fluctuations in pain regardless of their temporal ordering, and therefore can be readily computed if pain ratings are unequally spaced (as is the case in EMA protocols). Compared to other variability indices (e.g., mean squared successive differences, autocorrelation), the intraindividual SD has also been shown to require relatively fewer repeated observations to be reliably captured [13]. An index representing the amount of Time in High Pain was derived by calculating the percentage of momentary pain ratings that were at a level of ≥ 75 on the 0-100 scale for each patient and week, and a corresponding index of Time in Low Pain was calculated as the percentage of momentary ratings ≤ 34 on the 0-100 scale. The thresholds for high (≥ 75) and low pain (≥ 34) were selected based on previous work that has established cut-off points for severe and mild pain on the visual analogue scale for patients with chronic musculoskeletal pain [6]. It should be noted that whereas the indices of maximum and minimum pain capture the worst (or least) intensity of pain, the measures of time in high (or low) pain are intended to capture the frequency of pain experiences that could be characterized as severe (or mild). Finally, a measure of the Average Pain After Wake-up was constructed by averaging the first momentary pain rating of each day (i.e., selecting only the pain ratings after wake-up) across the 7 days for each patient; morning pain has been shown to be a hallmark feature in patients with fibromyalgia [34,39,45].

All pain intensity indices were represented as continuous variables. Increasing values conceptually reflect worsening of patients’ pain experience for all measures except for the percent of time in low pain, where increases reflect improvements in the pain experience.

Data analysis

Descriptive analyses

Initial descriptive analyses examined participant completion rates of EMA prompts over the course of the studies. Because completion rates have previously been shown to decline over the duration of EMA study protocols [33,35], changes in EMA completion rates were also examined. Multilevel growth models were used for this purpose [36], where the number of daily completed EMA prompts served as dependent variable, and day of study served as predictor variable, allowing for random effects (individual differences) in the (linear) time trend of completed daily prompts.

We next examined descriptive statistics (means, SDs) and correlations among the different pain indices at baseline. If some of the alternative indices were near perfectly correlated with each other, this would suggest that they do not capture different aspects of pain in daily life. Additionally, we examined the test-retest reliability of each pain index with intraclass correlation coefficients (ICCs) computed between the first and second pre-treatment weeks. The rationale was that pain indices showing a low test-retest reliability would not be good candidates for measuring the impact of treatment because the reliability of a measure sets an upper limit for its validity.

Analyses of treatment effects

The goal of the primary analyses was to examine the treatment effects obtained for each pain index and compare them with the treatment effects on Average Pain. Because some of the indices were not scaled in the same way (e.g., the SD metric of the Pain Variability index differs from indices tapping average, high, or low pain levels), we compared standardized effect sizes (ES) of the treatment effects, defined as the difference in change from baseline between an active treatment group and the placebo control group relative to the pooled standard deviation of changes from baseline in each group. The second week of the 2-week pre-treatment assessment period was selected as the baseline week because it was most proximal to treatment (alternatively, we could also have created a baseline score by averaging the scores for the first and second pre-treatment weeks for each pain index, but we decided against this because baseline assessment periods of multiple weeks are not available in all trials, limiting generalizability; this decision did not impact the results). The 3-week dose escalation period was not included in the analyses because treatment effects were likely to change during this period. Thus, the analyses considered treatment effects from baseline to each of the weeks of stable dose treatment.

Repeated measures ANOVA models were used to estimate the standardized group differences in change for each pain index, contrasting each of the two doses (100 or 200 mg per day) to the placebo control group. Because each trial involved multiple weeks of stable dose treatment (24 weeks in study 1 and 26 weeks in study 2), it was possible to obtain multiple repeated treatment ES estimates per trial. Rather than testing a separate ANOVA model for each treatment week, we fitted combined models for all treatment weeks simultaneously, whereby the (repeated) changes from baseline were predicted from treatment group, effect-coded treatment week, and the group × week interaction terms. Using effect coding for the categories of treatment week, a main effect for treatment group on standardized change in a given pain index is interpreted as the ES of the treatment averaged over all the treatment weeks, and the week × treatment group interaction effects represent deviations from the average treatment effect for a given week. This allowed us to obtain the pooled average ES, as well as the week-to-week variance in treatment effects, for each pain index. Whereas the pooled average ES represents our best estimate of the overall effect of the intervention on a given pain index, the week-to-week variance in effect size is a potentially useful measure to evaluate the consistency of the treatment effects over time. Initial ANOVAs were conducted for each index separately. To investigate the differences in ES estimates comparing the Average Pain index and each of the alternative indices, we fitted a multivariate extension of the repeated-measures ANOVA models, in which two indices (i.e., Average Pain and one of the alternatives) were treated as correlated outcomes. ES differences between the indices for each week (as well as pooled average ES differences across weeks, treatment doses, and trials) were obtained using the delta method [37].

In secondary analyses, we estimated the Number Needed to Treat (NNT) for each pain index. The NNT is a clinically intuitive effect size measure to evaluate the importance of treatment effects in clinical trials. NNTs for the efficacy of milnacipran in these samples have been previously reported for several clinical outcome measures (daily pain; global impressions of change; responder status on composite outcomes comprised of pain, global impressions of change, physical functioning) [12], such that estimating the NNT for each pain index also provided a means to evaluate the magnitude of the effects found for the indices in the context of previous findings for primary trial outcomes. The NNT is defined as the number of patients who would need to receive the active treatment in order to have one more success (or one less failure) than if treated with the placebo [11,16]. The NNT is based on the number of “responders” in each group and is calculated as NNT = 1/[responder rate in treatment group – responder rate in control group]. Consistent with the original trial analyses [9,29] and with recommendations for the reporting of core outcome measures in pain clinical trials [14], we defined responders as those obtaining reductions from baseline of at least 30%, separately for each pain index. To obtain a pooled average NNT across all treatment weeks for each pain index, we used generalized estimating equations (GEE) in which each individual’s weekly responder status was the binary dependent variable and in which treatment group, week (effect coded), and the treatment by week interaction served as independent variables (the resulting log odds from this model were transformed into NNTs using the formula described above).

Analyses of Patient Global Impression of Change

Additional analyses examined the extent to which changes in each pain index related to PGIC over the treatment period. For each pain index, we computed change scores for each patient by subtracting the baseline score from the score obtained for the week preceding a given PGIC assessment. The change scores were then used as predictor variables in ordinal logistic regressions treating PGIC scores as ordered categorical (ordinal) outcome variables. Because each participant contributed up to 7 PGIC ratings (assessed every 4 months over the course of the treatment period) we used cluster-robust standard errors to accommodate the nesting of multiple observations within individuals. Effect sizes quantifying the strength of association between changes in each pain index and the PGIC scores were calculated using Nagelkerke’s coefficient of determination (pseudo R2) [32]. A first set of analyses examined each pain index individually as a predictor in separate models. To evaluate whether any of the alternative pain indices showed an incremental contribution to understanding PGIC over the Average Pain index, hierarchical ordinal logistic regressions were estimated in which changes in the Average Pain index were controlled in a first step to obtain incremental effects (partial odds ratios and changes in pseudo R2) of an alternative pain index entered in the second step.

Missing data

Consistent with intent-to-treat (ITT) principles, all patients who were randomized into one of the treatment or control conditions and were considered part of the ITT sample in the original studies were included in the analyses. Missing values were accommodated using full information maximum likelihood parameter estimation, which has the desirable feature of using all available information in the analysis irrespective of whether data for some weeks are missing from a given participant because they dropped out of the study [40]. Analyses were conducted using SAS version 8.4 (Cary, NJ) and Mplus version 8.1 [31].

Maximum likelihood estimation assumes that missing data are at least missing at random (MAR), which means that the missingness mechanism is ignorable given the observed data. This MAR assumption does not hold if participants drop out for reasons that are not represented in the observed data, in which case the missingness becomes nonignorable. We used the “index of local sensitivity to nonignorability” (ISNI) as implemented in the R package ISNI [58] to evaluate the extent to which the results (i.e., the estimated treatment effects) would be impacted by violations of the MAR assumption [28,56]. An ISNI analysis was performed for all parameters of the repeated measures ANOVA models, taking into account both intermittent nonresponses and dropout from the study (for details, see Xie et al.[58]). The potential impact of nonignorable missingness on the parameter estimates was evaluated using the c statistic, where small values of c < 1 indicate that a parameter is sensitive to nonignorability [56]. The c statistic exceeded the critical value of 1.0 for all treatment effect parameters of both studies and pain indices (cs > 4.90 in study 1 and cs > 3.05 in study 2), suggesting that nonignorable missingness would have little impact on the results and that analyses assuming MAR carried little risk of bias.

Results

EMA completion rates

Across the two studies, a total of 1,137,907 EMA pain intensity ratings (507,872 EMA ratings for n = 888 patients in study 1 and 630,035 EMA ratings for n = 1196 patients in study 2) were analyzed. Patients completed a mean of 4.74 (SD = 1.27) EMA pain ratings per day (across 107,121 days) in study 1, and a mean of 4.52 (SD = 1.45) EMA ratings per day (across 139,280 days) in study 2 (these completion rates include intermittent missing days but exclude time periods during which participants did not complete EMA ratings because they dropped out of the study). The first EMA rating (pain after wake-up) occurred between 6 and 10 AM on 80% of the days in both studies. Multilevel growth models indicated that EMA completion rates showed a significant linear decrease over time (study 1: t[887] = −17.6, p <.001; study 2: t[1195] = −20.8, p <.001); however, these effects were modest with the number of observed daily EMA ratings declining in rates of 2.3% (0.11 ratings, study 1) and 3.2% (0.15 ratings, study 2) per month.

Descriptive characteristics of pain indices at baseline

Descriptive statistics (means, SDs) of the EMA pain indices at baseline are shown in Table 1. Across the two studies, baseline pain ratings showed had mean Average Pain levels of around 63 to 67, mean Maximum pain levels of around 83 to 86, and mean Minimum pain levels of around 42 to 48, on the 0-100 scale. The mean Pain Variability was around 10, indicating that patients’ momentary pain levels on average fluctuated by about 10 points across EMA assessments. In terms of pain frequency indices, patients spent around 30 to 36 percent of Time in High Pain, and they spent around 4 to 8 percent of Time in Low Pain at baseline. The mean level of Pain After Wake-up was comparable to the Average Pain level (see Table 1).

Table 1:

Descriptive statistics and correlations among the pain indices at baseline.

Average pain
level
Maximum
pain level
Minimum
pain level
Pain
variability
Percent time
in high pain
Percent time
in low pain
Average pain
after wakeup
Correlations
 Average pain level -- .707 .866 −.518 .903 −.601 .888
 Maximum pain level .779 -- .397 .134 .750 −.172 .632
 Minimum pain level .877 .507 -- −.818 .726 −.644 .790
 Pain variability −.424 .130 −.745 -- −.327 .632 −.473
 Percent time in high pain .878 .751 .739 −.279 -- −.317 .794
 Percent time in low pain −.704 −.408 −.682 .492 −.378 -- .536
 Average pain after wakeup .884 .703 .807 −.404 .766 .645 --
Mean (SD)
 Study 1 67.41 (13.38) 85.60 (10.06) 47.51 (19.37) 9.79 (5.03) 36.17 (33.93) 4.27 (10.01) 67.43 (14.78)
 Study 2 63.31 (15.07) 82.98 (12.05) 42.26 (19.32) 10.53 (4.90) 30.20 (33.28) 7.92 (15.17) 62.94 (16.75)
Test-retest reliability (ICC)
 Study 1 .744 .654 .753 .715 .792 .496 .723
 Study 2 .809 .688 .774 .685 .826 .616 .798

Note: Correlations among pain indices for study 1 are above the main diagonal and for study 2 below the main diagonal. ICC = intraclass correlation coefficient. All correlations and ICCs in the table are significant at p < .001.

Test-retest correlations (ICCs) suggested moderate to high temporal stability of all pain indices across the two baseline weeks in both trials. ICCs of 0.70 or above were evident for measures of Average Pain, Minimum Pain, the percent Time in High Pain, and Pain After Wake-up. The Maximum Pain level and Pain Variability indices showed ICCs of about 0.65 to 0.70. The percent Time in Low Pain index was the least reliable with ICCs of about 0.50 to 0.60. In terms of the correlations among the indices at baseline, higher Average Pain levels were associated with higher Minimum Pain levels, more Time in High Pain, and higher Pain After Wake-up, where between 75% and 82% of the variance was commonly shared, and with higher Maximum Pain levels, where about 50% of the variance was shared. Higher Average Pain levels were associated with lower Pain Variability and less Time in Low Pain, with shared variances of 18% to 50% (see Table 1). Thus, even though several indices showed a high degree of communality, each of them also contained unique variance components.

Effects of treatment on the different pain indices

Table 2 shows the treatment ES estimates by study for the comparison between each of the active treatment groups (100 mg/day or 200 mg/day) versus placebo, for each pain index. The mean ES across all treatment weeks (first column in Table 2) ranged between .25 and .35 (ps < .01) when the Average Pain level served as the outcome measure. With few exceptions, the alternative pain indices also indicated significant treatment effects, even though the magnitude of effects was numerically smaller than the ES on Average Pain in most instances (the ES estimates for Pain After Wake-up were slightly higher than those on Average Pain in Study 2). Only two of the alternative pain indices did not show consistently significant treatment effects: the percent Time in High Pain index did not show a significant effect of the 100 mg/d dose and was only significant for the 200 mg/d dose in both trials, and the effect on Pain Variability was never significant (because the measurement of Pain Variability can be impacted by trends in the data, sensitivity analyses also examined effects on Pain Variability constructed from the residuals of EMA ratings after linear trends were removed for each week; the results were virtually identical). The week-to-week variation of the ES estimates (in SD units) for each index is also shown in Table 2, where a smaller variation in ES suggests greater consistency (or replicability) of treatment effects across weeks. While the Average Pain index generally showed the smallest variability of effect sizes, the ES variation was similar for all indices.

Table 2:

Treatment effect size statistics for the pain indices

Mean ES across weeks Variation of ES across
weeks
NNT


Estimate (95% CI) SD Range
Study 1 (100mg dose)
 Average pain level .290** (.090 ; .491) .036 .205 - .352 8.246
 Maximum pain level .276** (.085 ; .466) .051 .178 - .365 9.299
 Minimum pain level .220* (.024 ; .416) .047 .120 - .310 9.253
 Pain variability .131 (−.058 ; .319) .072 −.051 - .229 23.441
 Percent time in high pain .191* (−.004 ; .386) .047 .094 - .253 9.154
 Percent time in low pain .261** (.072 ; .450) .044 .173 - .353 9.003
 Avg. pain after wakeup .237* (.037 ; .437) .032 .158 - .298 9.156
Study 1 (200mg dose)
 Average pain level .346*** (.177 ; .515) .039 .279-.416 7.638
 Maximum pain level .330*** (.169 ; .491) .040 .259 - .400 8.471
 Minimum pain level .309*** (.140 ; .479) .044 .227 - .408 7.859
 Pain variability .083 (−.074 ; .240) .055 −.002 - .202 51.159
 Percent time in high pain .202* (.031 ; .373) .041 −.072 - .077 10.323
 Percent time in low pain 274*** (.114 ; .434) .038 .211 - .379 9.138
 Avg. pain after wakeup .312*** (.141 ; .482) .039 .241 - .369 8.428
Study 2 (100mg dose)
 Average pain level .249** (.097 ; .401) .038 .150 - .324 10.917
 Maximum pain level .232** (.084 - .381) .056 .094 - .316 12.549
 Minimum pain level .237** (.087 - .386) .038 .177 - .324 9.749
 Pain variability −.017 (−.161 ; .128) .045 −.110 - .074 −1989.57
 Percent time in high pain .130 (−.019 ; .280) .033 .077 - .191 13.749
 Percent time in low pain .225** (.078 ; .372) .040 .151 - .306 11.391
 Avg. pain after wakeup .291*** (.138 ; .443) .037 .201 - .355 9.532
Study 2 (200mg dose)
 Average pain level .329*** (.172 ; .485) .035 .264 - .397 7.534
 Maximum pain level .303*** (.153 ; .454) .049 .206 - .390 8.763
 Minimum pain level .274** (.119 ; .428) .047 .197 - .393 7.526
 Pain variability .001 (−.146 ; .148) .054 −.142 - .087 −86.840
 Percent time in high pain .176* (.022 ; .330) .039 .106 - .241 8.038
 Percent time in low pain .323*** (.170 ; .475) .053 .232 - .413 8.112
 Avg. pain after wakeup .334*** (.178 ; .489) .036 .257 - .395 7.864

Notes: ES = effect size; NNT = number needed to treat (average across all treatment weeks).

*

p <.05;

**

p < .01;

***

p < .001.

A direct comparison of effect sizes obtained from the Average Pain index and the alternative indices is provided in Figure 2, which shows a forest plot of ES differences contrasting the Average Pain index with each of the alternatives, for all treatment weeks. The estimated grand means of ES differences (weighted averages across all treatment weeks) are shown at the bottom of the forest plots; they were ΔES = −.019 (95% CI = −.060 to .023) for Maximum pain, ΔES = −.036 (CI = −.081 to .009) for Minimum Pain, ΔES = −.036 (CI = −.094 to .022) for Time in Low Pain, ΔES = −.007 (CI = −.038 to .023) for Pain After Wake-up, indicating nonsignificantly smaller effect sizes for each of these indices. The mean ES were significantly smaller than the ES obtained from Average Pain for Pain Variability: ΔES = −.263 (CI = −.396 to −.129), and for Time in High Pam: ΔES = −.129 (CI = −.213 to −.044).

Figure 2:

Figure 2:

Forrest plots of differences in standardized effect sizes between the Average Pain index and each of the alternative pain indices. Negative values indicate that the effect size of the alternative index is lower than the effect size of the Average Pain index. All effect sizes are based on the comparison between an active treatment group and the placebo conrol group, and are organized by study (1 and 2), treatment dose (100 mg/d or 200 mg/d), and treatment week (4-27 in study 1, and 4-29 in study 2). Error bars represent 95% confidence intervals. Abbreviation: ES = effect size.

Results from the secondary analyses examining treatment effects via the NNT are shown in the last column of Table 2. Across studies and treatment doses, the NNT ranged between 7.6 and 10.9 when the Average Pain level was used as the basis for defining treatment responders. To provide a benchmark for comparison of these effect sizes, primary efficacy outcomes with the same doses of milnacipran in this patient population previously have yielded similar NNTs of 9.0 to 10.0 (≥ 30% improvement in daily pain), 7.7 to 7.8 (global impressions of change ratings of “much” or “very much” improved), and 11.0 to 14.0 (responder on composite outcomes) [12].

In study 1, NNTs for the alternative indices exceeded the NNT for Average Pain in magnitudes between 3% (Minimum Pain, 200 mg/day) and 35% (Time in High Pain, 200 mg/day) except for Pain Variability, which showed NNTs exceeding that for Average Pain by more than 180%. In study 2, the Minimum Pain index showed the lowest NNTs (0.1% and 10% lower than Average Pain for the 100 mg/day and 200 mg/day doses, respectively), and the NNT for Pain After Wake-up was 13% lower than the NNT of Average pain for the 200 mg/day dose. The NNTs for the other indices exceeded those based on the Average Pain index by 4% to 16% (except for Pain Variability, which showed a negative NNT in study 2).

Implications of treatment effect sizes for sample size and statistical power

Given that greater effect sizes imply a greater statistical power to detect treatment effects if they actually exist in the population, the question how effect size differences between the pain indices translate into differences in statistical power is also of interest. To that end, we conducted power calculations, in which we (somewhat arbitrarily) assumed a standardized difference in change between two groups on the Average Pain index of ES = 0.3 (a small to medium effect as per Cohen [10]), and a correlation from baseline to post-treatment of 0.5. We then determined the sample sizes required to detect this treatment effect with 80% power if the Average Pain index or an alternative pain index was used as the outcome measure. The results of the power calculations are shown in Figure 3. To ensure 80% power, the required sample sizes per treatment group were n = 176 (Average Pain), n =200 (Maximum Pain), n =232 (Minimum Pain), n =11,468 (Pain Variability), n =538 (Percent Time in High Pain), n =227 (Percent Time in Low Pain), and n =184 (Pain after Wake-up).

Figure 3:

Figure 3:

Estimated power to detect a significant treatment effect (a difference in change between 2 groups) based on each pain index, assuming an effect size of 0.3 for the Average Pain level and a pre-post correlation of 0.5.

Relationships between changes in pain indices and Patient Global Impression of Change

The final set of analyses examined the extent to which changes in the pain indices corresponded with PGIC scores taken over the course of the treatment period. As shown in Table 3, when the pain indices were considered individually in bivariate logistic regressions predicting PGIC, the change scores for each of the indices were significantly associated with PGIC scores in the expected direction (such that improvements on each index corresponded with perceived improvements reflected in PGIC). The numerically largest effect sizes were evident for changes in Average Pain, which explained 43.2% and 43.9% of the variance in PGIC ratings across the two studies. The smallest, albeit significant, effect sizes were found for changes in Time in High Pain (explaining 8.5% and 14.1% of the variance) and for changes in Pain Variability (explaining 1.0% and 2.8% of the variance in PGIC). In hierarchical logistic regressions, several alternative pain indices showed small but significant incremental contributions to PGIC beyond Average Pain (incrementally explaining between 0.4% and 1.3% of the variance, see Table 3). In study 1, improvements in Maximum Pain and in Pain Variability uniquely predicted improvements in PGIC; in study 2, improvements in Maximum Pain, Pain Variability, and the Percent Time in Low Pain uniquely predicted improvements in PGIC, whereas increases in Minimum Pain predicted improvements in PGIC after controlling for Average Pain.

Table 3:

Regression results for changes in pain indices as predictors of Patient Global Impressions of Change

Bivariate logistic regressions Logistic regressions controlling for
Average Pain level
OR (95% CI) Pseudo
R2
Partial
OR
(95% CI) ΔPseudo
R2
Study 1
 Average pain level 5.43*** (4.68 ; 6.31) .439
 Maximum pain level 4.60*** (4.03 ; 5.25) .388 1.57*** (1.26 ; 1.95) .007
 Minimum pain level 3.69*** (3.24 ; 4.20) .331 0.90 (0.71 ; 1.30) .000
 Pain variability 1.20** (1.07 ; 1.35) .010 1.17** (1.05 ; 1.30) .004
 Percent time in high pain 2.10*** (1.87 ; 2.36) .141 0.98 (0.85 ; 1.12) .000
 Percent time in low pain 3.66*** (3.22 ; 4.15) .320 1.18 (0.97 ; 1.42) .001
 Avg. pain after wakeup 4.84*** (4.21 ; 5.57) .409 1.17 (0.81 ; 1.70) .000
Study 2
 Average pain level 5.28*** (4.66 ; 5.97) .432
 Maximum pain level 4.82*** (4.32 ; 5.39) .405 1.88*** (1.57 ; 2.26) .013
 Minimum pain level 3.35*** (3.01 ; 3.73) .298 0.70*** (0.57 ; 0.85) .005
 Pain variability 1.38*** (1.24 ; 1.53) .028 1.33*** (1.21 ; 1.46) .012
 Percent time in high pain 1.74*** (1.61 ; 1.88) .085 0.93 (0.85 ; 1.01) .001
 Percent time in low pain 3.90*** (3.52 ; 4.32) .346 1.55*** (1.34 ; 1.80) .010
 Avg. pain after wakeup 4.29*** (3.83 ; 4.81) .376 0.89 (0.67 ; 1.18) .000

Notes: Standardized changes in each pain index served as predictor variables. OR = odds ratio.

**

p < .01;

***

p < .001.

Discussion

Identifying methods to measure aspects of pain intensity that are more or less sensitive to a given treatment may facilitate a better understanding of therapeutic benefits and contribute to the development and refinement of effective treatment options. In this paper, we examined the use of different pain intensity measures derived from EMA as endpoints in two pharmacological clinical trials. Our results showed test-retest reliabilities approaching or exceeding a level of 0.7 for most of the examined pain indices, suggesting that different outcome measures of pain intensity can be reliably captured from EMA data collected over the course of a week. The Average Pain EMA index was associated with the numerically greatest treatment effect sizes and was most closely associated with PGIC. This finding also corresponds with our results from stakeholder interviews with clinical trialists (reported in the first paper of this series), who ranked changes in patients’ average pain levels as the most useful index of successful treatment in clinical trials [50]. We emphasize that the Average Pain index assessed with EMA should not be viewed as interchangeable with commonly used measures of average pain that are based on retrospective (e.g., 7-day recall) reports [7,8]. Whereas indices derived from EMA may be deemed more ecologically valid, it is possible that recall reports of average pain capture additional information of clinical relevance for understanding treatment effects, which was not investigated here.

EMA measures of Average, Maximum, and Minimum Pain yielded similar and closely correlated treatment effect sizes. This finding is in line with previous research that has compared the assay sensitivity of patients’ reports of average, worst, and least pain. In a recent meta-analytic review, Smith et al. [48] synthesized the results from 23 pharmacological treatment studies that reported effect sizes for both average pain and worst pain reports (using 24-hr or 1-week recall items) across various chronic pain conditions and found no significant difference in efficacy estimates obtained from average and worst pain ratings. Notably, the difference in effect sizes between average and worst pain outcomes was very small at d = 0.02, nearly identical to the present results. Similarly, a previously reported pooled analysis of 4 clinical trials with fibromyalgia patients testing the effects of duloxetine showed very similar treatment effect estimates from 24-hour recall ratings of average, worst, and least pain[2].

Even though the treatment effect sizes based on these measures were very similar, the present study showed that changes in Maximum Pain (in both studies) and Minimum Pain (in study 2) were uniquely related to PGIC after controlling for EMA Average Pain. This suggests that they offer an alternative and potentially valuable perspective on patients’ pain experience and that there may be utility in using one or more of these alternative indices to more fully capture the range of effects that are deemed important by patients in pain clinical trials. Paralleling the meta-analytic results on the relationships between pain indices and patient functioning outcomes described in the second paper of this series, patients (in study 2) perceived greater improvement in their overall condition when Maximum Pain levels decreased more and Minimum Pain levels decreased less (or increased more) than would be expected by changes in Average Pain. This suggests that reductions in the differences between worst and least pain levels (e.g., due to reductions in short-term symptom flares) may be especially desirable for patients with fibromyalgia and especially important for how they judge improvements in their overall status [43]. This, in turn, may have clinical consequences because global retrospective impressions have been shown to play a central role in patient decision making and patients’ willingness to continue a therapeutic regimen [38].

The Percent of Time in Low Pain states was measured somewhat less reliably from momentary pain ratings compared to other indices. Nevertheless, it captured treatment effects of similar magnitude to Average Pain, and (in study 2) uniquely predicted PGIC ratings. From a clinical perspective, spending more time in low or “tolerable” pain is a desirable treatment outcome in that it reflects not just the overall pain intensity, but incorporates the frequency or duration of time over which low pain levels are achieved. Our results are in line with a prior study in which we found that an EMA-derived index of pain frequency captured beneficial effects of treatment regimen changes among rheumatology patients, with effect sizes slightly exceeding those for average pain [51].

An interesting finding was that indices of the percent of Time in High Pain and of Pain Variability showed mostly non-significant treatment effect sizes. This could be taken to mean that these indices are not sensitive to detecting treatment benefits. However, it could also suggest that these indices capture distinct aspects of pain that were not affected by the treatment. Milnacipran is a dual reuptake inhibitor, which decreases pain by increasing the neurotransmission of norepinephrine and serotonin [3,20]. It is not specifically designed to reduce momentary fluctuations in patients’ pain levels (as would be the case for rescue medications) or to specifically target states of severe pain, which may explain the small effect sizes for these two indices. The notion that Pain Variability captures meaningful information in the context of clinical trial outcomes is supported by the finding that changes in Pain Variability explained incremental variance in PGIC ratings beyond Average Pain, which suggests that impressions of improvement as perceived by patients could be further augmented by treatments that target reductions in pain variability.

This study has several strengths. First, the number of EMA pain intensity ratings that were the basis for the present analyses was substantial. Patients collected momentary pain data for several months and contributed over 1 million EMA pain intensity ratings over the course of the active treatment period, providing insights into the reproducibility of results across repeated assessments and across two separate studies. Second, both studies used a stringent clinical trial design, in that both were randomized, double-blind, placebo-controlled, parallel group studies. Third, using EMA ratings as the basis of the present analysis made it possible to evaluate and compare treatment effects obtained for a relatively broad range of outcome measures characterizing different aspects of pain.

Several limitations also need to be considered. First, the study was based on secondary data analyses of two existing trials whose goals were not to compare the sensitivity of different pain intensity measures to evaluate treatment benefits. Second, our study sample consisted exclusively of patients with fibromyalgia who were mostly female (as expected for this diagnosis), and White. The study findings could have been different for patients with other chronic pain conditions and demographic characteristics. It should be emphasized that it may be problematic to draw generalizing conclusions from studies of fibromyalgia, which represents a very complex pain disorder with many unique characteristics. Third, the results of the present analyses are limited to one specific pharmacological treatment and the conclusions drawn from this study cannot speak to other types of pain treatment. It is important to examine whether the present results generalize to other studies, and the utility of alternative pain intensity indices for other types of treatments should continue to be explored. These include rescue medications for pain flares, spinal cord stimulation, and psychosocial approaches including cognitive-behavioral therapy [48]. As these treatments might target different aspects of patients’ pain experience, the tailored use of a given pain index that can best reflect the therapeutically desired outcome could be a fruitful strategy for future research. Finally, the present research focused on pain outcome measures that capture basic distributional characteristics of real-time pain experiences. Additional measures capturing temporal dynamics of pain can be derived from EMA but these were not considered here because they require more specialized analyses. For example, novel applications of time-series analyses have been shown to capture unique temporal features of pain intensity, including the persistence (e.g., autocorrelation) of pain states and the amplitude of shifts between elevated and reduced pain states, which may provide important avenues for future efforts to optimize the detection of efficacious treatments [30,43,55].

There are also other ways in which pain intensity indices derived from EMA could facilitate the detection of treatment effects that were not considered in this study. For example, recommendations from IMMPACT for improving assay sensitivity in chronic pain clinical trials mention pain intensity patterns (e.g., pain variability, pain constancy) at the baseline assessment period that could affect the ability to detect treatment effects [15]. Supporting this hypothesis, Farrar et al. [18] found that pain variability assessed from 7 daily diaries at the baseline period moderated observed treatment effect sizes, in that higher baseline pain variability was associated with a greater likelihood of treatment response in placebo-control groups but not in active medication-treated groups. It is possible that other aspects of baseline pain such as those used in this study could yield similar moderating effects or that baseline levels of different pain indices could be useful for treatment tailoring: these are open questions for future research.

Conclusions

Alternative summary measures of pain intensity derived from EMA have the potential to broaden the scope of outcome measures that could be useful as endpoints in pain clinical trials and may contribute to detecting changes from treatment that are deemed relevant by patients. It is important for our findings to be extended to other chronic pain diagnoses and treatment approaches. Comparative effectiveness trials may especially benefit from including multiple summary measures of pain intensity to determine whether different types of treatment affect different aspects of pain intensity.

Series Concluding Remarks

Our primary goal in this series of three papers was to determine from several perspectives the characteristics of different indices of pain intensity that can be derived from EMA protocols in addition to the traditional assessment of Average Pain. We argued throughout the series that these indices provide a new viewpoint on patients’ experience of pain. We were pleased to find that it was possible to reliably measure various temporal aspects of pain intensity with momentary assessments. Perhaps most strikingly, no single pain index emerged consistently as “superior” for understanding patients’ pain experience. Instead, the results indicated that the validity and potential usefulness of different pain indices depend on the context in which they were evaluated: whether they were preferred by stakeholders, whether they enhanced understanding of patient functioning, or whether they indicated treatment effects in clinical trials. We view this variation in pain index performance as pertinent information for informing decisions about the selection of pain intensity measures for pain research and clinical practice.

We recognize that this series of studies does not provide a definitive answer on which aspects of pain intensity are most important to measure in a particular setting. We cannot be definitive for many reasons, including limited sample sizes, examining clinical trial data for a single drug and a particular pain diagnosis, and our dependence on secondary data analyses. We also recognize additional sources of information that we did not access, such as the ability of different pain intensity indices to discriminate between pain diagnoses [21]. Nevertheless, we hope the series makes a compelling case that broadening the scope of pain intensity measures with indices based on momentary data is useful while at the same time acknowledging that many decisions need to be made about which aspects of pain are most pertinent to measure in different contexts.

The three studies in the series each offered a unique perspective on the indices. Results from the stakeholder opinions, empirical associations with functional outcomes, and treatment detection do not squarely align in a manner that unambiguously leads one to select a single pain index. Although maybe not all of the pain indices might be compelling as primary pain outcomes, they could serve to supplement primary measures and expand our understanding of the pain experience. Finally, we are not in a position to recommend a formula for integrating the different sources of information about the indices as it will undoubtedly vary by the reason for collecting the data. Choosing the “right” combination of pain intensity outcomes will demand much thoughtful effort from researchers and clinicians. Our hope is that this series will generate heightened interest in the measurement of alternative indices of pain intensity that can be derived from EMA to ultimately facilitate evidence-based decision making regarding the most suitable measures of pain intensity for research and practice.

HIGHLIGHTS.

  • Treatment effects based on alternative EMA indices of pain intensity were compared

  • Over 1 million EMA pain intensity ratings over 20+ weeks of treatment were analyzed

  • Multiple EMA pain indices detected treatment effects in pain clinical trials

  • Changes in all EMA indices were associated with patient global impression of change

  • Alternative pain indices may inform understanding of clinical trial outcomes.

Acknowledgments

The data for the present analyses were provided by Allergan plc. The authors would like to thank Dr. John Edwards and Raffaele Migliore for making the data available.

This work was supported by a grant from the National Institute of Arthritis and Musculoskeletal and Skin Diseases (R01 AR066200; A.A.S. and S.S., principal investigators).

Footnotes

Conflict of interest statement

A.A.S. is a Senior Scientist with the Gallup Organization and a consultant with IQVIA and Adelphi Values, Inc. The remaining authors have no conflict of interest to declare.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  • [1].Allen KD: The value of measuring variability in osteoarthritis pain. J Rheumatol 34:2132–2133,2007 [PubMed] [Google Scholar]
  • [2].Arnold LM, Clauw DJ, Wohlreich MM, Wang F, Ahl J, Gaynor PJ, Chappell AS: Efficacy of duloxetine in patients with fibromyalgia: pooled analysis of 4 placebo-controlled clinical trials. Primary care companion to the Journal of clinical psychiatry 11:237, 2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Arnold LM, Gendreau RM, Palmer RH, Gendreau JF, Wang Y: Efficacy and safety of milnacipran 100 mg/day in patients with fibromyalgia: results of a randomized, double-blind, placebo-controlled trial. Arthritis and rheumatism 62:2745–2756, 2010 [DOI] [PubMed] [Google Scholar]
  • [4].Atkinson TM, Mendoza TR, Sit L, Passik S, Scher HI, Cleeland C, Basch E: The Brief Pain Inventory and its "pain at its worst in the last 24 hours" item: clinical trial endpoint considerations. Pain medicine 11:337–346, 2010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [5].Bellamy N, Sothern R, Campbell J, Buchanan W: Rhythmic variations in pain, stiffness, and manual dexterity in hand osteoarthritis. Annals of the rheumatic diseases 61:1075–1080, 2002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [6].Boonstra AM, Preuper HRS, Balk GA, Stewart RE: Cut-off points for mild, moderate, and severe pain on the visual analogue scale for pain in patients with chronic musculoskeletal pain. PAIN® 155:2545–2550, 2014 [DOI] [PubMed] [Google Scholar]
  • [7].Broderick JE, Schwartz JE, Schneider S, Stone AA: Can End-of-day reports replace momentary assessment of pain and fatigue? The Journal of Pain 10:274–281, 2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Broderick JE, Schwartz JE, Vikingstad G, Pribbernow M, Grossman S, Stone AA: The accuracy of pain and fatigue items across different reporting periods. Pain 139:146–157, 2008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Clauw DJ, Mease P, Palmer RH, Gendreau RM, Wang Y: Milnacipran for the treatment of fibromyalgia in adults: a 15-week, multicenter, randomized, double-blind, placebo-controlled, multiple-dose clinical trial. Clinical therapeutics 30:1988–2004, 2008 [DOI] [PubMed] [Google Scholar]
  • [10].Cohen J: Statistical Power Analysis for the Behavioral Sciences. Hillsdale, N.J.: Lawrence Erlbaum, 1988. [Google Scholar]
  • [11].Cook RJ, Sackett DL: The number needed to treat: a clinically useful measure of treatment effect. BMJ: British Medical Journal 310:452, 1995 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Cording M, Derry S, Phillips T, Moore RA, Wiffen PJ: Milnacipran for pain in fibromyalgia in adults. Cochrane Database of Systematic Reviews, 2015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Du H, Wang L: Reliabilities of intraindividual variability indicators with autocorrelated longitudinal data: Implications for longitudinal study designs. Multivar Behav Res 53:502–520, 2018 [DOI] [PubMed] [Google Scholar]
  • [14].Dworkin RH, Turk DC, Farrar JT, Haythornthwaite JA, Jensen MP, Katz NP, Kerns RD, Stucki G, Allen RR, Bellamy N, Carr DB, Chandler J, Cowan P, Dionne R, Galer BS, Hertz S, Jadad AR, Kramer LD, Manning DC, Martin S, McCormick CG, McDermott MP, McGrath P, Quessy S, Rappaport BA, Robbins W, Robinson JP, Rothman M, Royal MA, Simon L, Stauffer JW, Stein W, Tollett J, Wernicke J, Witter J: Core outcome measures for chronic pain clinical trials: IMMPACT recommendations. Pain 113:9–19, 2005 [DOI] [PubMed] [Google Scholar]
  • [15].Dworkin RH, Turk DC, Peirce-Sandner S, Burke LB, Farrar JT, Gilron I, Jensen MP, Katz NP, Raja SN, Rappaport BA, Rowbotham MC, Backonja MM, Baron R, Bellamy N, Bhagwagar Z, Costello A, Cowan P, Fang WC, Hertz S, Jay GW, Junor R, Kerns RD, Kerwin R, Kopecky EA, Lissin D, Malamut R, Markman JD, McDermott MP, Munera C, Porter L, Rauschkolb C, Rice ASC, Sampaio C, Skljarevski V, Sommerville K, Stacey BR, Steigerwald I, Tobias J, Trentacosti AM, Wasan AD, Wells GA, Williams J, Witter J, Ziegler D: Considerations for improving assay sensitivity in chronic pain clinical trials: IMMPACT recommendations. Pain 153:1148–1158, 2012 [DOI] [PubMed] [Google Scholar]
  • [16].Dworkin RH, Turk DC, Wyrwich KW, Beaton D, Cleeland CS, Farrar JT, Haythornthwaite JA, Jensen MP, Kerns RD, Ader DN: Interpreting the clinical importance of treatment outcomes in chronic pain clinical trials: IMMPACT recommendations. The Journal of Pain 9:105–121, 2008 [DOI] [PubMed] [Google Scholar]
  • [17].Farrar JT, Pritchett YL, Robinson M, Prakash A, Chappell A: The clinical importance of changes in the 0 to 10 numeric rating scale for worst, least, and average pain intensity: analyses of data from clinical trials of duloxetine in pain disorders. The Journal of Pain 11:109–118, 2010 [DOI] [PubMed] [Google Scholar]
  • [18].Farrar JT, Troxel AB, Haynes K, Gilron I, Kerns RD, Katz NP, Rappaport BA, Rowbotham MC, Tierney AM, Turk DC, Dworkin RH: Effect of variability in the 7-day baseline pain diary on the assay sensitivity of neuropathic pain randomized clinical trials: an ACTTION study. Pain 155:1622–1631, 2014 [DOI] [PubMed] [Google Scholar]
  • [19].Farrar JT, Young JP Jr., LaMoreaux L, Werth JL, Poole RM: Clinical importance of changes in chronic pain intensity measured on an 11-point numerical pain rating scale. Pain 94:149–158, 2001 [DOI] [PubMed] [Google Scholar]
  • [20].Fields HL, Heinricher MM, Mason P: Neurotransmitters in nociceptive modulatory circuits. Annu Rev Neurosci 14:219–245, 1991 [DOI] [PubMed] [Google Scholar]
  • [21].Fillingim RB, Loeser JD, Baron R, Edwards RR: Assessment of chronic pain: Domains, methods, and mechanisms. The Journal of Pain 17:T10–T20, 2016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Harris RE, Williams DA, McLean SA, Sen A, Hufford M, Gendreau RM, Gracely RH, Clauw DJ: Characterization and consequences of pain variability in individuals with fibromyalgia. Arthritis and Rheumatism 52:3670–3674, 2005 [DOI] [PubMed] [Google Scholar]
  • [23].Jensen MP, Hu X, Potts SL, Gould EM: Single vs composite measures of pain intensity: relative sensitivity for detecting treatment effects. Pain 154:534–538, 2013 [DOI] [PubMed] [Google Scholar]
  • [24].Jensen MP, Hu X, Potts SL, Gould EM: Measuring outcomes in pain clinical trials: The importance of empirical support for measure selection. The Clinical journal of pain 30:744–748, 2014 [DOI] [PubMed] [Google Scholar]
  • [25].Jensen MP, Karoly P: Self-report scales and procedures for assessing pain in adults, in Turk DC, Melzack R (eds): Handbook of pain assessment. New York: Guilford Press, 2001, pp 15–34 [Google Scholar]
  • [26].Krueger AB, Stone AA: Assessment of pain: a community-based diary survey in the USA. Lancet 371:1519–1525, 2008 [DOI] [PubMed] [Google Scholar]
  • [27].Litcher-Kelly L, Martino SA, Broderick JE, Stone AA: A systematic review of measures used to assess chronic musculoskeletal pain in clinical and randomized controlled clinical trials. The Journal of Pain 8:906–913, 2007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Ma G, Troxel AB, Heitjan DF: An index of local sensitivity to nonignorable drop- out in longitudinal modelling. Stat Med 24:2129–2150, 2005 [DOI] [PubMed] [Google Scholar]
  • [29].Mease PJ, Clauw DJ, Gendreau RM, Rao SG, Kranzler J, Chen W, Palmer RH: The efficacy and safety of milnacipran for treatment of fibromyalgia. A randomized, double-blind, placebo-controlled trial. The Journal of rheumatology 36:398–409, 2009 [DOI] [PubMed] [Google Scholar]
  • [30].Mun CJ, Suk HW, Davis MC, Karoly P, Finan P, Tennen H, Jensen MP: Investigating intraindividual pain variability: methods, applications, issues, and directions. Pain 160:2415–2429, 2019 [DOI] [PubMed] [Google Scholar]
  • [31].Muthén LK, Muthén BO: Mplus user's guide. Los Angeles, CA: Muthén & Muthén, 1998–2017. [Google Scholar]
  • [32].Nagelkerke NJ: A note on a general definition of the coefficient of determination. Biometrika 78:691–692, 1991 [Google Scholar]
  • [33].Okifuji A, Bradshaw DH, Donaldson GW, Turk DC: Sequential analyses of daily symptoms in women with fibromyalgia syndrome. The Journal of Pain 12:84–93, 2011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Older SA, Battafarano DF, Danning CL, Ward JA, Grady EP, Derman S, Russell IJ: The effects of delta wave sleep interruption on pain thresholds and fibromyalgia-like symptoms in healthy subjects; Correlations with insulin-like growth factor I. J Rheumatol 25:1180–1186, 1998 [PubMed] [Google Scholar]
  • [35].Ono M, Schneider S, Junghaenel DU, Stone AA: What Affects the Completion of Ecological Momentary Assessments in Chronic Pain Research? An Individual Patient Data Meta-Analysis. Journal of medical Internet research 21:e11398, 2019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Raudenbush SW, Bryk AS: Hierarchical linear models. Thousand Oaks, CA: Sage, 2002. [Google Scholar]
  • [37].Raykov T, Marcoulides GA: Using the delta method for approximate interval estimation of parameter functions in SEM. Structural Equation Modeling 11:621–637, 2004 [Google Scholar]
  • [38].Redelmeier DA, Katz J, Kahneman D: Memories of colonoscopy: a randomized trial. Pain 104:187–194, 2003 [DOI] [PubMed] [Google Scholar]
  • [39].Roizenblatt S, Moldofsky H, Benedito- Silva AA, Tufik S: Alpha sleep characteristics in fibromyalgia. Arthritis & Rheumatism: Official Journal of the American College of Rheumatology 44:222–230, 2001 [DOI] [PubMed] [Google Scholar]
  • [40].Schafer JL, Graham JW: Missing data: our view of the state of the art. Psychol Methods 7:147–177, 2002 [PubMed] [Google Scholar]
  • [41].Schneider S, Junghaenel DU, Broderick JE, Ono M, May M, Stone AA: Indices of pain intensity derived from ecological momentary assessments and their relationships with patient functioning: an indivdiual patient data meta-analysis. Journal of Pain, (under review) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Schneider S, Junghaenel DU, Keefe FJ, Schwartz JE, Stone AA, Broderick JE: Individual differences in the day-to-day variability of pain, fatigue, and well-being in patients with rheumatic disease: Associations with psychological variables. Pain 153:813–822, 2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [43].Schneider S, Junghaenel DU, Ono M, Stone AA: Temporal dynamics of pain: an application of regime-switching models to ecological momentary assessments in patients with rheumatic diseases. Pain 159:1346–1358, 2018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].Schneider S, Stone AA: Distinguishing between frequency and intensity of health-related symptoms from diary assessments. J Psychosom Res 77:205–212, 2014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [45].Shaver JL, Lentz M, Landis CA, Heitkemper MM, Buchwald DS, Woods NF: Sleep, psychological distress, and stress arousal in women with fibromyalgia. Research in Nursing & Health 20:247–257, 1997 [DOI] [PubMed] [Google Scholar]
  • [46].Shiffman S, Stone AA, Hufford MR: Ecological momentary assessment. Annu Rev Clin Psychol 4:1–32, 2008 [DOI] [PubMed] [Google Scholar]
  • [47].Smith SM, Amtmann D, Askew RL, Gewandter JS, Hunsinger M, Jensen MP, McDermott MP, Patel KV, Williams M, Bacci ED: Pain intensity rating training: results from an exploratory study of the ACTTION PROTECCT system. Pain 157:1056–1064, 2016 [DOI] [PubMed] [Google Scholar]
  • [48].Smith SM, Jensen MP, He H, Kitt R, Koch J, Pan A, Burke LB, Farrar JT, McDermott MP, Turk DC: A Comparison of the Assay Sensitivity of Average and Worst Pain Intensity in Pharmacologic Trials: An ACTTION Systematic Review and Meta-Analysis. The Journal of Pain 19:953–960, 2018 [DOI] [PubMed] [Google Scholar]
  • [49].Smith WR, Bauserman RL, Ballas SK, McCarthy WF, Steinberg MH, Swerdlow PS, Waclawiw MA, Barton BA, Hydro IMS: Climatic and geographic temporal patterns of pain in the Multicenter Study of Hydroxyurea. Pain 146:91–98, 2009 [DOI] [PubMed] [Google Scholar]
  • [50].Stone AA, Broderick JE, Goldman RE, Junghaenel DU, Bolton A, May M, Schneider S: Indices of pain intensity derived from ecological momentary assessments: rationale and stakeholder interviews. Journal of Pain, (under review) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [51].Stone AA, Broderick JE, Schneider S, Schwartz JE: Expanding options for developing outcome measures from momentary assessment data. Psychosom Med 74:387–397, 2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [52].Stone AA, Broderick JE, Schwartz JE: Validity of average, minimum, and maximum end-of-day recall assessments of pain and fatigue. Contemporary clinical trials 31:483–490, 2010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53].Stone AA, Schneider S, Broderick JE, Schwartz JE: Single-day pain assessments as clinical outcomes: not so fast. The Clinical journal of pain 30:739–743, 2014 [DOI] [PubMed] [Google Scholar]
  • [54].Stone AA, Schwartz JE, Broderick JE, Shiffman SS: Variability of momentary pain predicts recall of weekly pain: a consequence of the peak (or salience) memory heuristic. Personality and Social Psychology Bulletin 31:1340–1346, 2005 [DOI] [PubMed] [Google Scholar]
  • [55].Tighe PJ, Bzdega M, Fillingim RB, Rashidi P, Aytug H: Markov chain evaluation of acute postoperative pain transition states. Pain 157:717–728, 2016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [56].Troxel AB, Ma G, Heitjan DF: An index of local sensitivity to nonignorability. Statistica Sinica 1221–1237, 2004 [Google Scholar]
  • [57].US Department of Health and Human Services Food and Drug Administration: Guidance for Industry: Analgesic Indications: Developing Drug and Biological Products. Available at https://www.fda.gov/downloads/drugs/guidancecomplianceregulatoryinformation/guidances/ucm384691.pdf. Accessed April 26, 2019.
  • [58].Xie H, Gao W, Xing B, Heitjan DF, Hedeker D, Yuan C: Measuring the impact of nonignorable missingness using the R package isni. Computer methods and programs in biomedicine 164:207–220, 2018 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES