Abstract
Response bias characterized by decreases in self-reported subjective states when measured repeatedly over short time-frames is a potential concern in social science. Recent work suggests that this initial elevation bias (IEB) is pronounced among young adult students’ self-reports of affect when using ambulatory methods, but it is unclear if such bias extends broadly across samples, designs, and constructs. We examined the conditions wherein reliable and robust IEB may manifest by conducting a coordinated analysis of seven ecological momentary assessment (EMA) studies with diverse lifespan samples to test the generalizability of IEB across study designs and affective constructs (momentary negative and positive affect). Overall, evidence for substantial IEB across studies was weak. No reliable evidence emerged for IEB in negative affect, with evidence for a small magnitude IEB for positive affect when comparing initial reports with reports made 1 week later, although the latter was not evident in other comparisons and was attenuated to nonsignificance when controlling for temporal factors. The magnitude and direction of IEB varied, but in mostly nonsystematic ways, as a function of study design and affective valence. Meta-analytic summary revealed consistently low combined effect sizes (Cohen’s ds ranging from −.05 to .14). We found little evidence that IEB in momentary affect is sufficiently reliable, robust, or generalizable across designs and constructs to pose broad and/or serious concerns for EMA studies. Nonetheless, we recommend systematically examining the potential for IEB across study designs and constructs to help identify the conditions/contexts where IEB may or may not manifest.
Keywords: initial elevation bias, affect, generalizability, ecological momentary assessment, intensive longitudinal designs
Response bias characterized by decreases in self-reported subjective states when repeatedly assessed over short timescales is a recognized concern in social science (French & Sutton, 2010). Recent work using ambulatory methods suggests that there may be bias characterized by pronounced initial elevation in self-reported momentary affect. Initial elevation bias (IEB) is defined as reporting higher levels or severity of a construct than participants experience in the initial survey (Shrout et al., 2018). We know very little, however, about the contexts and conditions under which IEB may emerge. It is thus important to determine the conditions wherein reliable and robust IEB may manifest.
We conducted a coordinated analysis of intensive longitudinal studies that utilized multiple daily reports assessing momentary affective states (i.e., ecological momentary assessment [EMA] studies) to systematically evaluate the generalizability of IEB across seven studies that differ in sample composition, study design, and measurement of self-reported momentary affective states. Coordinated analysis involves conducting separate analyses across independent studies (with both similar and systematically varying characteristics), followed by a comparison of results (Hofer & Piccinin, 2009). If results replicate across the independent studies, conclusions drawn from analyses can be considered more robust and generalizable across the different features of each study. Alternatively, results that hold in only some studies may provide evidence as to the contexts that do and do not support a particular relation or finding.
Response Patterns Across Repeated Assessments of Momentary Affective States
Response patterns in affective states have been observed in microlongitudinal data where self-reported levels and severity of (typically negative) affective states diminish unexpectedly over repeated assessments (e.g., Knowles et al., 1996; Sharpe & Gilbert, 1998). This response pattern may manifest from numerous sources of bias, including IEB (e.g., Shrout et al., 2018) and attenuation effects (e.g., Lucas et al., 1999), as well as participant burden or familiarity with assessment protocols/surveys (e.g., Arslan et al., 2021). Such a pattern may thus produce biased assessments that threaten the validity of conclusions drawn from self-report data, but little is known about the circumstances in which it manifests (French & Sutton, 2010).
Prior empirical work has evaluated evidence for the phenomenon in protocols with a single follow-up or widely spaced intervals over time (e.g., once per week), examining self-reports of negative and positive affective states (NA and PA, respectively). Knowles et al. (1996) showed that college undergraduates reported significantly higher scores on the first assessment versus the second assessment 1 week later on an anxiety scale, with a small-to-medium sized effect. Sharpe and Gilbert (1998) similarly showed that college undergraduates reported significantly higher scores on scales of depression, tension, anger, fatigue, confusion, and total mood disturbance on the first assessment compared to the second and third assessments completed 1 week and 2 weeks later, respectively. No evidence of these effects emerged for measures of PA (Sharpe & Gilbert, 1998). In a sample of 65 women attending a breast clinic, Johnston (1999) showed that anxiety and depression measure scores were significantly higher for anxiety, but not depression, on the first assessment versus the second assessment completed on the following day, with a medium-to-large sized effect.
IEB in Affective States
Building on these earlier findings, work by Shrout et al. (2018) suggests that there may be pronounced IEB in affect among younger adult student samples using intensive longitudinal methods. Although these earlier findings were based on studies with few repeated assessments with varying reporting periods, Shrout et al. focused on evaluating IEB in densely repeated measurement contexts with momentary reports (i.e., reporting on mood “right now”). Results from Shrout et al. (2018) four studies suggest there is greater IEB in affect than in behavior (i.e., study time), and in NA (i.e., anxiety/anxious mood) than in PA (i.e., vigor/vigorous mood). Shrout et al. (2018) note the difference in affect levels between early and later observations was likely due to IEB (rather than other sources such as selection bias or regression to the mean) because (a) the timing of the report was disentangled from the serial position of the report within the series of repeated assessments; and (b) participants were randomized to different start dates. Study 1 is of particular relevance to the present coordinated analysis because of examination of IEB within persons across intensive repeated momentary assessments. Studies 2–4 emphasized between-person analyses outside the scope of the present study. Therefore, we focus on a broad replication of the within-person analyses presented in Shrout et al’s Study 1; this study collected self-reports of anxious and vigorous mood among recent law school graduates (mean age = 29.7 years) in a time period of preparation for the bar exam (i.e., a condition of potential threat/evaluation; Shrout et al., 2018).
Participants completed three items assessing anxious mood (on edge, uneasy, anxious) and three items assessing vigorous mood (cheerful, lively, vigorous) from the Profile of Mood States (POMS; McNair et al., 1992), and were instructed to rate the extent to which they were feeling or experiencing these states “right now” (see Gleason et al., 2008, Shrout et al., 2010; cf. Cranford et al., 2006). Reports were obtained once in the morning upon awakening and once in the evening before going to bed. Reliable small-to-medium effect sizes were reported for IEB in anxious mood such that scores were significantly higher on day 1 compared to day 8 in momentary reports in the morning (d = .47, 95% CI 0.36–0.58) and evening (d = .26, 95% CI .15–.37). Less pronounced, marginally significant effect sizes were reported for IEB in vigorous mood such that scores were higher on day 1 compared to day 8 in momentary reports in the morning (d = .09, 95% CI −.03 to .21) and evening (d = .10, 95% CI −.01 to .21).
If IEB in reports of momentary affect is pervasive and of meaningful magnitude, as suggested by data from Shrout et al. (2018), social science researchers would have cause for concern. Intensive longitudinal studies could be susceptible to potentially misleading conclusions about, for example, changes in affect across the study protocol. These concerns might also have practical consequences; for instance, by potentially confounding the impacts of mental health interventions with effects attributable to IEB, as well as by misinterpreting clinical cutoffs for mood disorder screening tools that may result in individuals being inadvertently miscategorized based on IEB. Therefore, it is important to clarify if such bias extends broadly in diverse adult lifespan samples across different EMA study designs and components of momentary affect. Indeed, Shrout et al. (2018) noted the need for additional work to assess the potential for IEB in other survey-based research and to determine whether it may vary as a function of other study characteristics (p. E22).
The Present Study
We attempted to approximate Shrout et al. (2018) within-person analytic approach, leveraging seven EMA studies in coordinated analysis.
Different EMA Study Designs
Examining the characteristics of different study designs may help evaluate whether the amount of repeated sampling impacts IEB. Shrout et al. (2018) examined data from a 44-day intensive repeated measurement design where momentary reports of anxious mood and vigorous mood (“right now”) were collected in the morning and evening of each day using paper-and-pencil forms. Determining whether the findings suggested by Shrout et al. (2018) are observed in EMA designs that differ in the frequency of assessments repeated within days, the number of assessment days (study duration), and other design features (e.g., capturing data on a smartphone) would contribute to understanding the extent, magnitude, and potential boundary conditions for IEB effects in momentary affect reports.
Components of Affect
IEB was more pronounced in anxious mood than vigorous mood in Shrout et al. (2018) study. These affective states broadly align to the negative valence-high arousal and positive valence-high arousal quadrants of the circumplex model of affect (Feldman, 1995), respectively. Assessment of a wider array of affect items (i.e., both high and low arousal) and more general composites of momentary NA and PA would help to clarify whether IEB in high arousal affect generalizes to broader composites of affect incorporating both high and low arousal affective states.
Research Questions
We tested the generalizability of IEB in diverse adult lifespan samples across EMA study designs and constructs (NA, PA). Our study addresses two research questions (Figure 1).
Figure 1. Schematic Diagram for Primary Research Questions.

Note. EMA = Ecological momentary assessment. Each shape reflects an assessment of negative affect or positive affect. The solid black arrow at the top of the diagram represents the comparison of first beeped EMA on day 1 versus day 8 (or day 7 for studies with 7 days of assessment; Research Question 1). The dotted black arrow in the middle of the diagram represents the comparison of last beeped EMA on day 1 versus day 8 (or day 7 for studies with 7 days of assessment; Research Question 1). The last beeped EMA is labeled “Beep k” due to differences in the total number of beeped assessments throughout the day across the studies. The square dotted arrow represents the secondary comparison of aggregated EMA beeps on day 1 versus day 8 (or day 7 for studies with 7 days of assessment; Research Question 2). The right-most column, “n” additional days, signifies the additional days of assessment included in some of the studies, but not included in analyses addressing the research questions.
Is there IEB in self-reported momentary affective states? We examine the potential for an elevation in affect in the initial surveys, followed by a decline to lower values in subsequent surveys. We further explore if IEB differs by valence of affect to see if affect exhibits divergent patterns of IEB characterized by more pronounced IEB in NA compared to PA. In each EMA study, we used the first and last momentary assessment of affect for each day. We judged this approach to most closely approximate the momentary morning and evening assessments examined by Shrout et al. (2018) based on the consistency in item reporting timeframe (i.e., “right now”) and completion time in the A.M. and P.M., respectively.
Is there IEB when aggregating reports across momentary assessments? We extend Research Question 1 to reports that are aggregated across momentary assessments measured during the day to determine if findings differ from those obtained using the last momentary assessment of affect for each day. Aggregating across all momentary assessments within the day offers an opportunity to incorporate more data points than the single data points taken on each day in Research Question 1.
Additional Analyses
In addition to the research question-driven replication and extension of Shrout et al. (2018), we pursued additional analyses that explored a range of additional comparisons within and across study days to further test the generalizability of findings. Specifically, we also explore the potential for IEB: (a) in studies with run-in or practice-related data prior to collection of primary data, (b) that manifests over the course of the earliest momentary assessments on the first day of study protocols, (c) across consecutive days early in the study, (d) in anxiety-related affect, and (e) following adjustment for time of day and day of week.
Method
We assess the generalizability of IEB in diverse adult lifespan samples across EMA study designs and constructs (NA, PA) in the following seven studies: Stress, Health, and Daily Experiences in individuals with arthritis (SHADE Arthritis) and asthma (SHADE Asthma), Stress and Working Memory (SAWM), Ecological Validity in Patient Chronic Disease (EVIP), Momentary Reactivity and Chronic Pain (MRCP), Treatment Changes in Pain Patients (TCIP), and Effects of Stress on Cognitive Aging, Physiology, and Emotion (ESCAPE). All studies received institutional review board approval. Informed consent was obtained from all participants. Table 1 provides descriptions of participants, measures, and procedures for each study. Table 2 provides design details and sample demographics for each study.
Table 1.
Descriptions of Participants, Measures, and Procedures by Study
| EMA studies without run-in |
||||
| Study detail | SHADE Arthritis (N = 19), Asthma (N = 56) | EVIP (N =106) | MRCP (N = 65) | TCIP (N = 116) |
|
| ||||
| Population Sampled | Adult patients with rheumatoid arthritis or asthma | Adult patients with chronic pain or fatigue | Adult patients with chronic pain | Adult patients with chronic pain |
| Negative Affect Items | Depressed, unhappy, angry, frustrated, worried | Frustrated | Depressed/blue, angry/hostile, frustrated, worried/anxious | Depressed, angry, frustrated, worried |
| Positive Affect Items | Happy, joyful, enjoyment, pleased | Happy, full of life, energetic | Happy, enjoying/having fun, pleased | Happy, enjoying/having fun, pleased |
| Directions for Reporting | right now | right now | right now | right now |
| Response Option | Likert-type: 0 (not at all) to 6 (extremely) | VAS: 0 (not at all) to 100 (extremely) | VAS: 0 (not at all) to 100 (extremely) | VAS: 0 (not at all) to 100 (extremely) |
| Run-in Period | N/A | N/A | N/A | N/A |
| Procedure | • One training session | • One training session | • One training session | • One training session |
| • Study-provided palmtop computer | • Study-provided electronic device | • Study-provided electronic device | • Study-provided electronic device | |
| • Devices beeped 5 times per day for 7 consecutive days | • Devices beeped an average of 7 randomly scheduled times per day for 30 consecutiv days based on self-reported wake time | • Devices beeped either 3, 6, or 12 times per day for 14 consecutive days using stratified random sampling based on sampling density | • Devices beeped 9 times per day for 7 consecutive days using stratified random sampling that generated a random beep with uniform probability between 35 and 177 min after previous beep | |
| Reference | Smyth et al. (2014) | Broderick et al. (2008) | Stone et al. (2003) | Schneider et al. (2018) |
|
| ||||
| EMA Studies With Run-In |
||||
| Study Detail | SAWM (N = 156) | ESCAPE (N = 234) | ||
|
| ||||
| Population Sampled | Community-dwelling adults | Community-dwelling adults | ||
| Negative Affect Items | Tense, sad, upset, disappointed | Tense/anxious, angry/hostile, depressed/blue, frustrated, unhappy | ||
| Positive Affect Items | Happy, enthusiastic, content, excited | Happy, pleased, enjoyment/fun, joyful | ||
| Directions for Reporting | right now | right now | ||
| Response Option | Likert-type: 1 (not at all) to 7 (extremely) | VAS: 0 (not at all) to 100 (extremely) | ||
| Run-in Period | 2 days | 2 days | ||
| Procedure | • One training session | • One training session | ||
| • Study-provided palmtop computer | • Study-provided smartphone | |||
| • Devices beeped 5 times per day for 7 consecutive days based on self-reported wake time | • Devices beeped 5 times per day for 14 consecutive days based on self-reported wake time | |||
| Reference | Mogle et al. (2019) | Scott et al. (2015) | ||
Note. SHADE Arthritis = Stress, health, and daily experiences in individuals with arthritis; SHADE Asthma = Stress, health, and daily experiences in individuals with asthma; VAS = Visual analog scale. For all studies with multiple items, negative affect and positive affect were operationalized as the average of responses. Higher scores reflected greater affective experience for each component of affect. For ecological momentary assessment (EMA), participants were instructed to complete a momentary survey following each beep. See Scott et al. (2015) for more information on effects of stress on cognitive aging, physiology, and emotion (ESCAPE’S) systematic probability sampling of registered voter lists for zip code 10475 in the Bronx, NY. See Smyth et al. (2014) for more information on the SHADE study examining how daily experiences are associated with health and well-being for adults with rheumatoid arthritis and asthma. Stress and working memory (S AWM) participants were recruited from the Syracuse, NY area to participate via advertisements and flyers (see Mogle et al., 2019 for additional details). Ecological validity in patient chronic disease (EVIP) participants were recruited to participate from two offices of a community rheumatology practice that informed patients in the waiting room that they may be eligible to participate in the study if they experience chronic pain or fatigue (see Broderick et al., 2008, 2009 for additional details). Momentary reactivity and chronic pain (MRCP) participants had at least one of four chronic pain disorders (fibromyalgia, rheumatoid arthritis, osteoarthritis, or ankylosing spondylitis) and were recruited to participate via advertisements and flyers (see Stone et al., 2003 for additional details). Only 23 participants had data on day 1 of the study. To preserve the capacity to assess initial elevation for participants without data on day 1, we assessed data from the first observed reports and 7 days later for these participants. Treatment changes in pain patients (TCIP) participants had at least of three chronic pain disorders (fibromyalgia, rheumatoid arthritis, or osteoarthritis) and were recruited to participate following a discussion with their rheumatologist (see Schneider et al., 2018 for additional details).
Table 2.
Design Details and Descriptive Statistics by Study
| EMA studies without run-in |
EMA studies with run-in |
||||||
|---|---|---|---|---|---|---|---|
| Design detail and descriptive statistic | SHADE Arthritis | SHADE Asthma | EVIP | MRCP | TCIP | SAWM | ESCAPE |
| Study duration (days) | 7 | 7 | 30 | 14 | 7 | 7 | 14 |
| Beeped Assessments per day | 7 | 7 | 6 | 3/6/12 | 9 | 5 | 5 |
| Participants (N) | 19 | 56 | 104 | 65 | 116 | 156 | 234 |
| Mean age (SD) | 50.81(12.92) | 42.76(13.00) | 56.20(11.10) | 51.00(10.60) | 57.37(13.12) | 46.81(16.94) | 46.77(10.88) |
| Age range | 26–80 | 18–78 | 28–88 | 25–75 | 25–75 | 21–80 | 25–65 |
| Sex (Female %) | 75% | 73% | 86% | 86% | 85% | 52% | 67% |
| Education (%) High school diploma/GED or less |
16% | 21% | 29% | 32% | 25% | 44% | 22% |
| Bachelor’s degree/some college | 62% | 50% | 55% | 56% | 53% | 46% | 60% |
| Beyond Bachelor's degree | 22% | 29% | 16% | 12% | 22% | 10% | 18% |
| Race European-American (%) | 90% | 88% | 92% | 91% | 95% | 59% | 10% |
| Mean negative affect (SD) | 0.94(1.04) | 1.08(1.17) | 27.84(26.46) | 25.35(22.84) | 18.54(19.35) | 1.97(1.22) | 22.74(21.90) |
| Negative affect range | 0–6 | 0–6 | 0–100 | 0–100 | 0–100 | 1–7 | 0–100 |
| Mean positive affect (SD) | 2.49(1.24) | 2.88(1.39) | 49.08(19.59) | 51.23(21.18) | 58.66(20.94) | 4.36(1.22) | 60.90(25.31) |
| Positive affect range | 0–6 | 0–6 | 0–100 | 0–100 | 0–100 | 1–7 | 0–100 |
Note. EMA = Ecological momentary assessment; SHADE Arthritis = Stress, health, and daily experiences in individuals with arthritis; SHADE Asthma = Stress, health, and daily experiences in individuals with asthma; EVIP = Ecological validity in patient chronic disease; MRCP = momentary reactivity and chronic pain; TCIP = Treatment changes in pain patients; SAWM = Stress and working memory; ESCAPE = Effects of stress on cognitive aging, physiology, and emotion.
Analytic Plan
Data for each study were cleaned and prepared based on protocol descriptions (see Table 1 for source paper references for each study). Notably, studies differed as to the exposure participants had to data collection methods prior to the start of the actual study; that is, two studies (SAWM and ESCAPE) had explicit “run-in” (e.g., surveys completed prior to the primary study protocol that involve evaluation of compliance a participant needed to meet to qualify for participation in the actual study) or practice periods prior to the formal data collection. This issue is important, as if we were to conduct analyses beginning on “day 1” of the study (i.e., at the onset of formal data collection) there is the possibility that the IEB would be washed out during the run-in period for studies that had them, and thus missed—biasing our overall estimate of IEB toward zero. Although it would have been desirable to include data from run-in and/or practice periods, studies were inconsistent in the quality and availability of this data. That is, some studies did not explicitly instruct participants to take responses seriously (e.g., during practice, the goal was to familiarize participants with the response options and interface not to collect veridical data), and several studies did not retain data from practice periods. As such, our primary analyses were conducted on the five studies without run-in and practice data (SHADE Arthritis, SHADE Asthma, EVIP, MRCP, TCIP); thereby excluding the two studies with run-in and practice data to ensure there was no prestudy exposure to the assessments. Recognizing the importance of this issue, however, we conducted sensitivity analyses with SAWM and ESCAPE to explore whether a run-in period might eliminate or reduce IEB.
Analyses were completed in SAS (SAS Institute, 2013). In our efforts to replicate and extend the findings of Shrout et al. (2018), our initial approach was to follow their analytic structure in our first set of analyses. We then extended these analyses with a range of additional comparisons within and across study days to further test the generalizability of findings.
For the first set of analyses (Research Question 1), we compared the first momentary assessment of the day (i.e., in the morning) completed on days 1 and 8 to evaluate changes in affect ratings (as this was most comparable to the momentary morning reports in Shrout et al., 2018). For studies with 7 days of observations (SHADE Arthritis, SHADE Asthma, TCIP, SAWM), day 7 was used in place of day 8. To condense the presentation of information, we use “Day 8” throughout the manuscript when referring to analyses comparing momentary affect scores 1 week apart, although this reflects day 7 for 1-week studies. Next, we compared the last momentary assessment of the day (i.e., in the evening) completed on days 1 and 8 as this was most comparable to the momentary evening reports in Shrout et al. (2018). Instances where the last momentary assessment occurred before 12:00 p.m. on a given day (e.g., due to missing afternoon and evening reports) were not used in the analyses to ensure the observation aligned with evening analyses in Shrout et al. (2018). For the extension to IEB when aggregating across momentary beeped assessments (Research Question 2), we averaged all momentary ratings obtained for each person during the first and eighth days and proceeded with analyses outlined in Research Question 1.
Across all EMA studies, the median time stamp of the first beeped assessment within a given day was 9:42 a.m. (medians ranged from 8:50 a.m. to 10:51 a.m. across individual studies), confirming that the timing of the first beeped assessments in our coordinated analysis was a close approximation to the momentary morning assessment obtained by Shrout et al. (2018). Further, the median time stamp of the last beeped assessment within a given day was 8:37 p.m. (medians ranged from 7:36 p.m. to 9:20 p.m. across individual studies), confirming that the timing of the last beeped assessments was a close approximation to the momentary evening assessment obtained within an hour of participants going to bed in the evening examined by Shrout et al. (2018). Following the difference calculations by Shrout et al. (2018), we subtracted day 8 from day 1 scores such that positive effect size values reflect an initial elevation in the variable (day 1 score was higher than the day 8 score); one-sample t-tests comparing difference scores to 0 were conducted to evaluate statistical significance. To assess the magnitude of effects, Cohen’s d effect sizes were calculated (as mean within-subject change divided by the pooled between-person standard deviation).
Meta-analytic summaries of each set of analyses were conducted to determine combined effects across all studies. Fixed-effects meta-analysis models were used due to the relatively limited number of studies included in the analyses and the associated risk of imprecise estimation of between-studies variance (Borenstein et al., 2010). Combined effects were weighted using the inverse-variance from each study (Neyeloff et al., 2012). Data, study materials, and study analysis code may be made available for appropriate use upon emailed request to the corresponding authors. This study was not preregistered.
Results
Table 2 provides descriptive statistics for NA and PA next to demographic characteristics across all studies and designs. Figures 2, 3, and 4 provide panels of individual NA and PA plots across all available days for each of the five included studies for first beep, last beep, and aggregated EMA, respectively. Appendix A Supplemental Materials provides Z-scores of affect in each study plotted in the same graph for additional graphical representation of affect across the first 7 or 8 study days in each study. Overall, the evidence for IEB across studies was very weak. The magnitude and direction of IEB varied, but in mostly nonsystematic ways, as a function of study design and affective valence. Cohen’s d effect sizes were small and meta-analytic summaries of analyses revealed consistently low combined effect sizes (d’s ranging from −0.05 to 0.14). Below, we report results for the replication and extension of Shrout et al. (2018), as well as additional evaluations performed as sensitivity analysis.
Figure 2. First Beeped EMA of Negative and Positive Affect at Each Study Day.

Note. EMA = Ecological momentary assessment
Figure 3. Last Beeped EMA of Negative and Positive Affect at Each Study Day.

Note. EMA = Ecological momentary assessment.
Figure 4. Aggregated EMA of Negative and Positive Affect at Each Study Day.

Note. EMA = Ecological momentary assessment.
Replication and Extension of Shrout et al. (2018)
First Beeped EMA on Day 1 Versus Day 8 (Research Question 1)
Negative Affect.
In our replication of Shrout et al. (2018) examining first momentary assessments of days 1 and 8, no evidence of significant IEB emerged in NA (see Table 3, “First Beeped EMA” column). One of the five studies showed a significant negative effect (d = −.15; i.e., values were lower on assessment 1 compared to assessment 2). All remaining effects were nonsignificant, small, and ranged from positive to negative (ds = .13 to −.09). Meta-analytic weighting of the effect sizes revealed a nonsignificant small combined effect (d = −.05, 95% CI −0.14 to 0.03). Notably, this effect is in the opposite direction of what would be expected from an IEB. Figure 5A provides a forest plot depicting each study’s effect juxtaposed to the combined effect and referent findings from Shrout et al. (2018).
Table 3.
Effect Size Estimates of Initial Elevation Bias in Negative Affect
| Study | First beeped EMA |
Last beeped EMA |
Aggregated EMA |
|||
|---|---|---|---|---|---|---|
| N | Effect size [95% CI] | N | Effect size [95% CI] | N | Effect size [95% CI] | |
| SHADE Arthritis | 19 | .13 [−.31, .56] | 19 | .56 [.02, 1.11]* | 19 | .43 [.03, .83]* |
| SHADE Asthma | 56 | .10 [−.09, .10] | 53 | .19 [−.08, .46] | 56 | .13 [−.05, .30] |
| EVIP | 104 | −.15 [−.30, .00]* | 101 | −.06 [−.23, .10] | 104 | −.13 [−.24, −.01]* |
| MRCP | 65 | −.02 [−.22, .17] | 59 | −.06 [−.25, .13] | 65 | −.02 [−.18, .14] |
| TCIP | 116 | −.09 [−.24, .06] | 90 | .02 [−.17, .21] | 116 | −.06 [−.20, .08] |
| Combined effect | W | −.05 [−.14, .03] | W | .01 [−.08, .11] | W | −.03 [−.10, .04] |
| Cochran’s Q(DF) | Q | 5.34(4), p = .25 | Q | 7.62(4), p = .11 | Q | 12.06(4), p = .02 |
| Shrout et al. (2018) A.M.a | 286 | .47 [.36, .58]* | 286 | .47 [.36, .58]* | 286 | .47 [.36, .58]* |
| Shrout et al. (2018) P.M.b | 285 | .26 [.15, .37]* | 285 | .26 [.15, .37]* | 285 | .26 [.15, .37]* |
Note. EMA = Ecological momentary assessment; SHADE Arthritis = Stress, health, and daily experiences in individuals with arthritis; SHADE Asthma = Stress, health, and daily experiences in individuals with asthma; EVIP = Ecological validity in patient chronic disease; MRCP = momentary reactivity and chronic pain; TCIP = Treatment changes in pain patients.
= Referent findings from momentary A.M. and P.M. assessments in Shrout et al. (2018), respectively. Effect Size = Cohen’s d metric. W = Inverse-variance weighted combined effect size across all studies. Cochran’s Q(df) = Test of heterogeneity among studies included in each meta-analytic comparison.
p < .05.
Figure 5. Examining Initial Elevation Bias in First Beeped EMA.

Note. EMA = Ecological momentary assessment.
Positive Affect.
Evidence of significant IEB in PA emerged in only one of the five studies (d = .33; see Table 4, “First Beeped EMA” column). One study showed a marginally significant effect (d = .23). All remaining effects were nonsignificant and small (ds = .06–.40) or negative (d = −.06; i.e., values were lower on day 1 compared to day 8). Meta-analytic weighting of the effect sizes revealed a small, but significant, combined effect (d = .14,95% CI .05–0.24; see Figure 5B).
Table 4.
Effect Size Estimates of Initial Elevation Bias in Positive Affect
| Study | First beeped EMA |
Last beeped EMA |
Aggregated EMA |
|||
|---|---|---|---|---|---|---|
| N | Effect size [95% CI] | N | Effect size [95% CI] | N | Effect size [95% CI] | |
| SHADE Arthritis | 19 | .40 [−.13, .93] | 19 | −.03 [−.57, .52] | 19 | .05 [−.47, .58] |
| SHADE Asthma | 56 | .06 [−.17, .28] | 53 | −.02 [−.29, .25] | 56 | .14 [−.08, .36] |
| EVIP | 104 | −.06 [−.23, .11] | 101 | −.01 [−.15, .12] | 104 | −.07 [−.20, .06] |
| MRCP | 65 | .23 [−.02, .49]† | 59 | .33 [.12, .55]*** | 65 | .13 [−.07, .33] |
| TCIP | 116 | .33 [.16, .49]*** | 90 | .12 [−.06, .31] | 116 | .27 [.13, .42]*** |
| Combined effect | W | .14 [.05, .24]* | W | .10 [.01, .19]* | W | 0.10 [.02, .18]* |
| Cochran’s Q(DF) | Q | 12.67 (4), p = .01 | Q | 7.93 (4), p = .09 | Q | 12.76 (4), p = .01 |
| Shrout et al. (2018) A.M.a | 286 | .09 [−.03, .21] | 286 | .09 [−03, .21] | 286 | .09 [−.03, .21] |
| Shrout et al. (2018) P.Mb | 285 | .10 [−.01, .21]† | 285 | .10 [−.01, .21]† | 285 | .10 [−.01, .21]† |
Note. EMA = Ecological momentary assessment; SHADE Arthritis = Stress, health, and daily experiences in individuals with arthritis; SHADE Asthma = Stress, health, and daily experiences in individuals with asthma; EVIP = Ecological validity in patient chronic disease; MRCP = momentary reactivity and chronic pain; TCIP = Treatment changes in pain patients.
= Referent findings from momentary A.M. and P.M. assessments in Shrout et al. (2018), respectively. Effect Size = Cohen’s d metric. W = Inverse-variance weighted combined effect size across all studies. Cochran’s Q(df) = Test of heterogeneity among studies included in each meta-analytic comparison.
p < .10.
p < .05.
p < .001.
Last Beeped EMA on Day 1 Versus Day 8 (Research Question 1)
Negative Affect.
In our replication analyses examining last momentary assessments of days 1 and 8, evidence of significant IEB in NA emerged in only one of the five studies (d = .56; see Table 3, “Last Beeped EMA” column). All remaining effects were nonsignificant and small (ds = .01–.19) or negative (ds = −.06; i.e., values were lower on day 1 compared to day 8). Meta-analytic weighting of the effect sizes revealed a nonsignificant small combined effect (d = .01, 95% CI −0.08 to 0.11; see Figure 6A).
Figure 6. Examining Initial Elevation Bias in Last Beeped EMA.

Note. EMA = Ecological momentary assessment.
Positive Affect.
Evidence of significant IEB in PA similarly emerged in only one of the five studies (d = .33; see Table 4, “Last Beeped EMA” column). All remaining effects were nonsignificant and small (d = .12) or negative (ds = −.01 to −.03; i.e., values were lower on day 1 compared to day 8). Meta-analytic weighting of the effect sizes revealed a small, but significant, combined effect (d = .10, 95% CI 0.01–0.19; see Figure 6B).
Aggregated EMA on Day 1 Versus Day 7 or 8 (Research Question 2)
Negative Affect.
In analyses examining aggregated beeped assessments, evidence of significant IEB in NA emerged in only one of the five studies (d = .43; see Table 3, “Aggregated EMA” column). One study also showed a significant effect in the opposite direction (d = −.13; i.e., values were lower on day 1 compared to day 8). All remaining effects were nonsignificant and small (d = .13) or negative (ds = −.02 to −.03). Meta-analytic weighting of the effect sizes revealed a nonsignificant small combined effect (d = −.03, 95% CI −0.10 to 0.04; see Figure 7A).
Figure 7. Examining Initial Elevation Bias in Aggregated EMA.

Note. EMA = Ecological momentary assessment
Positive Affect.
Evidence of significant IEB in PA emerged in one of the five studies (d = .27; see Table 4, “Aggregated EMA” column). All remaining effects were nonsignificant and small (ds = .05–.14) or negative (d = −.07; i.e., values were lower on day 1 compared to day 8). Meta-analytic weighting of the effect sizes revealed a small, but significant, combined effect (d = .10, 95% CI .02–0.18; see Figure 7B).
Additional Analyses
Sensitivity Analysis in Studies With a Run-in Period
Two of the seven studies, SAWM and ESCAPE, implemented a 2-day run-in period to identify participants that demonstrated good compliance. Thus, these studies were not included in the primary analyses because of initial exposure to assessments of affect during run-in periods that were not experienced by participants from the other five studies. It may be that run-in periods may provide habituation to the measures that result in eliminated or reduced IEB. If a run-in period reduces IEB in the primary study protocol, then we should see systematically smaller magnitude of weighted combined effects across comparisons in analyses of SAWM and ESCAPE. For PA, the direction of combined effects maintained, with the magnitude of combined effects attenuated to statistical nonsignificance (weighted combined effects ranged from .01 to .08). For NA, there was again no evidence of IEB (i.e., IEB inconsistently observed and small in magnitude; weighted combined effects ranged from −.02 to .04). Appendix B provides supplemental tables with all weighted combined effects for each comparison.
Extension to Consecutive Days
Comparing day 1 scores with day 7 or 8 scores establishes whether there is evidence of IEB in momentary affect scores 1 week apart. It is possible that IEB involves “faster” processes that emerge across consecutive days. We therefore extended the analyses to explore whether the difference between day 1 and day 2 scores was significant. Findings for NA and PA were similarly inconsistent and meta-analytic summaries revealed weighted combined effects that were very small in magnitude (ranging from −.04 to .08). A small, but statistically significant, weighted combined effect (d = .08) emerged in the comparison of last beeped EMA for NA. All other weighted combined effects were nonsignificant. Appendix C provides supplemental tables with full reports on effect size estimates for each study and weighted combined effects.
Extension to the First Assessment Compared With the Next Two Assessments on Day 1
In addition to IEB that emerges across days, it is also possible that IEB emerges within the first day of the protocol across the earliest momentary assessments. We thus examined the potential for an initial elevation in the first assessment, followed by a decline to lower values in the second and third assessments completed that same day. Findings for both NA and PA were inconsistent and meta-analytic summaries revealed weighted combined effects that were quite small in magnitude (ranging from −.10 to .07). A small, but statistically significant, weighted combined effect in the opposite direction of IEB (d = −.10) emerged in the comparison of the first and third assessments on day 1 for NA. All other weighted combined effects were nonsignificant. Appendix D provides supplemental tables with full reports on effect size estimates for each study and weighted combined effects.
Extension to Anxiety-Related Affect Only
The components of affect assessed in this coordinated analysis are general and encompass both high arousal (e.g., anxious, angry) and low arousal (e.g., sad, depressed) affect, reflecting a broader operationalization of NA than the high arousal anxious mood scale Shrout et al. (2018) used. We thus conducted analyses utilizing a subset of items that were more closely aligned with those analyzed by Shrout et al. (2018). NA high arousal subscales for each study were computed based on the mean of items aligned with the circumplex model of affect’s (Feldman, 1995) negative valence-high arousal quadrant (See Appendix E Supplemental Materials for list of items). Patterns of results were generally similar between the high arousal anxiety-related NA subscale (weighted combined effects ranged from −.07 to −.01) and the general NA composite (weighted combined effects ranged from −.05 to .01). Appendix E provides supplemental tables with full reports on effect size estimates for each study and weighted combined effects.
Sensitivity Analysis Adjusting for Time of Day and Day of Week
We recognize that pursuing analyses that compare scores 1 week apart, on consecutive days, and on multiple assessments within the same day does not account for (and may introduce bias associated with) the potential influences of time of day and day of week. To formally test these possible confounds, we also conducted additional analyses on residuals from multilevel models (SAS proc mixed) that regressed NA and PA on a categorical time of day variable (Range: 1–4; 1 = morning, 2 = afternoon, 3 = evening, 4 = overnight), a categorical day of week variable (Range: 1–7; Monday = 1, Sunday = 7), and their interaction. Random effects allowed for individual differences in time of day and day of week effects on NA and PA. Sensitivity analysis revealed that findings were slightly attenuated following adjustments for the influences of time of day and day of week. The direction of effects generally maintained with slight attenuation in magnitude for comparisons made 1 week apart (day 1 vs. day 7 or 8; weighted combined effects ranged from −.04 to .07), comparisons made on consecutive days (day 1 vs. day 2; weighted combined effects ranged from −.05 to .06), and comparisons made on the earliest momentary assessments on the first day of the study protocol (assessment 1 vs. 2 and 3 on day 1; weighted combined effects ranged from −.12 to .07). The combined effects were attenuated to statistical nonsignificance for comparison of last beeped assessments of NA across consecutive days (day 1 vs. day 2), as well as comparisons of first, last, and aggregated assessments of PA made 1 week apart (day 1 vs. day 7 or 8). Appendix F provides supplemental tables with full reports on residual effect size estimates for each study and weighted combined effects.
Discussion
We conducted a coordinated analysis of intensive longitudinal studies to test the generalizability of IEB in diverse adult lifespan samples across EMA study designs and NA and PA constructs. We included study designs with different numbers of assessments and study days to extend beyond Shrout et al. (2018) seminal study. Descriptively, results suggest that IEB in momentary affect scores may not be sufficiently reliable, robust, or generalizable across EMA study designs to pose broad and/or serious concerns for the field. We discuss the findings in comparison and extension to Shrout et al. (2018) report of consistently small-to-medium effect sizes in student samples under what appear to be conditions of threat/evaluation. Considerations for analyses of IEB in EMA studies with self-report data are discussed in the context of the designs and constructs in which IEB may or may not be expected to manifest.
Generalizability of IEB Across Different EMA Studies
Overall, concerns about IEB would be serious if we observed evidence of IEB that was both reliable and of meaningful magnitude. We compared affect reports from the first and last momentary assessments for each day made 1 week apart as the closest approximation to the morning and evening reports of momentary affect examined by Shrout et al. (2018), respectively. To ensure broad evaluation of the potential for IEB beyond the replication analyses, we also aggregated assessments for each day made 1 week apart and conducted additional analyses comparing affect reports from consecutive days and across the first three assessments of the first day of the protocols.
No consistent evidence of IEB emerged for NA from the combined analyses, based either on statistical significance or on effect sizes. This was the case when comparing momentary affect reports made 1 week apart (i.e., in analyses comparing the first, last, and aggregated assessments on day 1 vs. day 8), which was the construct where Shrout et al. (2018) report the most pronounced IEB effects. Indeed, the present study’s three meta-analytic effects, and 13/15 individual study effects for first beeped, last beeped, and aggregated EMA of NA, were statistically nonsignificant and/or in the opposite direction of IEB. We similarly did not observe consistent evidence for IEB in our supplemental comparisons made over consecutive days and within the first day of assessment. Further, results from our supplemental examination of NA high arousal subscales did not strengthen the correspondence between the present study’s findings for NA and the IEB effects derived from the POMS anxiety scale (on edge, uneasy, anxious) used in Shrout et al. (2018).
Turning to PA, we observed statistically significant combined effects consistent with IEB for first beeped EMA, last beeped EMA, and aggregated EMA when comparing momentary affect reports made 1 week apart, although each of these effects was of small magnitude. Interestingly, the significant combined effects for PA were not replicated in supplementary analyses evaluating IEB across consecutive days (i.e., all three of the meta-analytic effects attenuated to nonsignificant when comparing first, last, and aggregated assessments on day 1 vs. day 2; Appendix C Supplemental Materials). Similarly, no combined effect for PA was statistically significant when comparing the earliest momentary assessments on the first day of EMA protocols. Finally, IEB effects on PA also were no longer significant when controlling for temporal factors (time of day and day of week).
In general, our findings from multiple EMA studies using diverse samples and varying sampling intensities/frequencies did not consistently replicate the patterns of IEB findings reported by Shrout et al. (2018). If sampling density was systematically related to IEB, we would expect to see some pattern across the five studies characterized by more or less IEB based on how many assessments are completed in each study (e.g., more IEB in studies with fewer assessments per day as one possibility). However, significant IEB effects in comparisons made 1 week apart emerged sporadically across the five studies; isolated instances of significant IEB emerged in studies with sampling frequencies ranging from seven assessments per day (NA in SHADE Arthritis) to up to 12 assessments per day (PA in MRCP). As noted, these significant effects were not observed across the other momentary affect (i.e., no IEB in PA for SHADE Arthritis, no IEB in NA for MRCP) and across the other EMA studies. Further, these significant effects were not replicated in supplementary analyses across shorter-time frames and following adjustment for time of day and day of week. Thus, these findings suggest that IEB in EMA designs may not strongly relate to sampling density; that is, the evidence for IEB (when it emerged) was not systematically related to sampling density—with as little as two (from Shrout et al., 2018) and as many as 12 assessments per day (MRCP).
Overall, we found evidence that IEB was not reliably present for NA, and only present for PA in some circumstances, and the magnitude of effects was consistently small in combined analyses. We thus conclude that IEB does not pose a systematic, widespread concern for momentary affect reporting in EMA studies. IEB did not emerge in combined analyses over the primary timeframe evaluated (i.e., week) for NA, although there was some indication of small magnitude IEB in NA in isolated instances included in supplementary analyses (i.e., last beeped EMA day 1 vs. day 2). In contrast, for PA, there was more evidence for consistent, albeit small in magnitude, IEB effects with comparisons of PA made 1 week apart. However, these effects were not replicated when comparing across consecutive days, assessments made within the first day of the protocol, or following adjustment for time of day and day of week. Thus, IEB in EMA designs seems to manifest in small effects for PA that may not be robust (e.g., emerging for comparisons 1-week apart, but not for comparisons over shorter-time frames, and not present when controlling for temporal factors), and not consistently emerging for NA. Thus, based on these analyses, IEB for momentary affect reports may be associated with relatively inconsistent and somewhat idiosyncratic effects that are smaller in magnitude than the effects reported by Shrout et al. (2018). Further, we did not identify a systematic pattern of findings that would suggest a particular number of assessments or duration of assessment period differentially impact the likelihood of IEB manifesting.
Considerations for Analyses of IEB in Intensive Longitudinal Self-Report Data
In the interest of replicating and extending the findings of Shrout et al. (2018), we aligned our analytic plan to their approach in replication as well as a range of additional comparisons within and across days. This approach enabled tests of whether initial scores were significantly higher than subsequent scores. With multiple assessments nested within multiple study days throughout week(s) nested within individuals, however, intensive longitudinal designs offer numerous analytic opportunities at multiple microtimescales. For research questions involving trajectories of affect, modeling techniques are available that use all available data. Multilevel modeling, for example, handles the nested structure of the intensive longitudinal data, enabling a potentially fuller evaluation of assessment-level and day-level trends that could result from IEB. These methods can also leverage random effects to allow individuals to vary in their possible IEB and can incorporate predictors of individual differences in IEB.
Importantly, however, application of these modeling approaches to address IEB would raise conceptual issues that are not straightforward. For example, model specification would require a priori knowledge and careful consideration of the assumptions being made on the nature of the change processes that would uniquely reflect IEB rather than true changes in people’s affective experiences over time, and how best to define changes due to IEB (e.g., do you expect linear or nonlinear declines? At what time frame(s)?). Although this set of questions was outside the scope of the present study (largely as we were attempting a replication and extension of prior work and had no a priori expectations for specific temporal patterns that would guide the application of such modeling approaches), we recognize the potential value of analytic approaches that model all available data.
Limitations and Future Directions
Several additional limitations of this study should be considered. The studies included in the coordinated analysis did not have a context/condition of threat/evaluation similar to the context of bar exam preparation included in study 1 of Shrout et al. (2018). Although we were unable to test the influence of an evaluative or threatful context in the present study, we remain open to the possibility that it could contribute to the elicitation of IEB. Determining the mechanisms that underlie IEB and other manifestations of response biases and/or habituation effects (e.g., gradual adjustment of response criterion, alterations of variability in responses) could further inform other elements of intensive longitudinal designs (e.g., recruitment and sampling procedures, training protocols, response options, item order) that may be relevant, but were not evaluated in the present study.
Evidence for IEB (as well as other time-related differences not reflective of IEB) did occasionally emerge and were of varying effect size. Accordingly, although we do not view this evidence as supporting IEB as a consistent and pervasive problem for intensive momentary ambulatory assessment designs, we believe that researchers should look at evidence for potential IEB as a routine analysis step in the initial evaluation of such data. Most broadly, a dataset could be analyzed using the procedures applied in the present study to determine the extent to which, if at all, IEB in self-report is apparent, and adjust accordingly. If no systematic evidence of IEB is found, researchers can be somewhat more confident in their use of the dataset without adjustment. In this way, evaluating the possible presence of IEB on a case-by-case basis encourages researchers to avoid summarily removing unbiased observations from analyses (e.g., by routinely discarding the first few days of data), and may also help contribute to a clearer understanding of the conditions wherein IEB may or may not manifest. If IEB is systematically observed and researchers are interested in adjusting for this initial rise in scores, a variety of approaches are available to consider. From a study design perspective, one could design a run-in period prior to the primary study protocol. As noted, our sensitivity analysis showed that the small magnitude IEB combined effects for PA attenuated to statistical nonsignificance in studies with a run-in period. The time- and resource-intensive nature of run-in periods should be considered, however, when deciding if and how to implement a run-in period to attenuate potential small and inconsistent IEB effects. From an analytic perspective, detrending by extracting residuals or covarying for repeated responding are some examples of accounting for time-related trends in the data. Critically, however, specific procedures to adjust for IEB need to be carefully selected to ensure that they do not undermine the possibility to address substantive research questions at hand.
Sample sizes for the seven studies included in this report varied widely, with several having small numbers of participants. Although a coordinated analysis approach across multiple studies was implemented, we recognize that additional research with larger samples would be informative. We included studies with diverse community-dwelling and clinical samples to ensure broad coverage of ages (21–88 years), education level (High school diploma, Bachelor’s degree, beyond Bachelor’s degree) and health status (e.g., healthy general population, patients with chronic conditions) across the adult lifespan. Consequently, we had a more diverse set of samples in our coordinated analysis than the college educated younger adult sample in Shrout et al. (2018) study 1. Descriptively, there did not appear to be any clear pattern of sample characteristics wherein IEB may be more likely to emerge in the present study. We do not yet understand, however, the extent to which individual differences in sociodemographic and health characteristics (e.g., younger adults vs. older adults; community-dwelling vs. clinical samples) may influence the likelihood of IEB to manifest and/or its magnitude.
Finally, this coordinated analysis tested the generalizability of IEB in self-reported momentary affect across different EMA study designs. The extent to which IEB may permeate across other intensive longitudinal designs (e.g., end-of-day daily diary [EOD], measurement burst studies) with varying reporting frames for affect (e.g., “right now” vs. a summary rating at EOD) is not yet understood and is an important next step for future research. Others have begun to explore this issue. For example, Arslan et al. (2021) examined evidence for IEB in a large (N = 1,345 women) daily diary study collecting summary ratings at EOD of psychosocial factors over 70 days, using a planned missingness design that randomized the days on which participants first saw specific items. The study reported only inconsistent and small effects of IEB across six negative and positive subjective states (i.e., loneliness, irritability, stress, self-esteem, risk taking, mood; each assessed with a single item). Continued research attention on IEB across different intensive longitudinal study designs is warranted to determine if IEB manifests at different timescales. Testing the generalizability of IEB in measurement burst studies (e.g., Sliwinski, 2008; Stawski et al., 2015), for example, would facilitate examination of the extent and magnitude of IEB at both shorter timescales (e.g., days) and longer timescales (e.g., every 6–12 months).
Conclusion
This coordinated analysis examined if IEB is consistently observed in momentary reports of affect, and if any such bias was of meaningful magnitude. IEB was not reliably observed for NA, although IEB for PA was observed for momentary reports taken 1 week apart (although these effects were not observed across different comparisons, such as across the first few days or over the first day of a study, and were not robust to inclusion of control variables). When present, IEB effects were of small magnitude. Thus, IEB may not be sufficiently reliable, robust, or generalizable across designs and constructs to pose broad and/or serious concerns for the assessment of momentary affect in EMA studies.
Supplementary Material
Public Significance Statement.
Evidence for response bias characterized by decreases in affect when repeatedly assessed over short timescales is weak, with bias inconsistently observed and typically small in magnitude. Analyses across seven different studies suggest that this type of response bias is not sufficiently reliable or robust to pose broad concerns for assessing affect via ecological momentary assessment designs. Nonetheless, it may be prudent to explicitly test for bias in future studies.
Acknowledgments
The content in this manuscript was supported by National Institute on Aging under Grants R37 AG057685, R01 AG042407, R01 AG039409 and R01 AG026728, National Heart, Lung, and Blood Institute under Grant R01 HL067990, National Institute of Arthritis and Musculoskeletal and Skin Diseases under Grants R01 AR066200 and U01 AR052170, and National Cancer Institute under Grant R01 CA085819. The first author was supported by National Institute on Aging under Grant T32 AG049676 to Pennsylvania State University.
Footnotes
All authors have no conflicts of interest to declare.
Data, study materials, and study analysis code may be made available for appropriate use upon emailed request to the corresponding authors. This study was not preregistered.
Supplemental materials: https://doi.org/10.1037/pas0001108.supp
References
- Arslan RC, Reitz AK, Driebe JC, Gerlach TM, & Penke L (2021). Routinely randomize potential sources of measurement reactivity to estimate and adjust for biases in subjective reports. Psychological Methods, 26(2), 175–185. 10.1037/met0000294 [DOI] [PubMed] [Google Scholar]
- Borenstein M, Hedges LV, Higgins JP, & Rothstein HR (2010). A basic introduction to fixed-effect and random-effects models for meta-analysis. Research Synthesis Methods, 1(2), 97–111. 10.1002/jrsm.12 [DOI] [PubMed] [Google Scholar]
- Broderick JE, Schwartz JE, Schneider S, & Stone AA (2009). Can End-of-day reports replace momentary assessment of pain and fatigue? Journal of Pain, 10(3), 274–281. 10.1016/j.jpain.2008.09.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Broderick JE, Schwartz JE, Vikingstad G, Pribbernow M, Grossman S, & Stone AA (2008). The accuracy of pain and fatigue items across different reporting periods. Pain, 139(1), 146–157. 10.1016/j.pain.2008.03.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cranford JA, Shrout PE, Iida M, Rafaeli E, Yip T, & Bolger N (2006). A procedure for evaluating sensitivity to within-person change: Can mood measures in diary studies detect change reliably? Personality and Social Psychology Bulletin, 32(7), 917–929. 10.1177/0146167206287721 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Feldman LA (1995). Variations in the circumplex structure of mood. Personality and Social Psychology Bulletin, 21(8), 806–817. 10.1177/0146167295218003 [DOI] [Google Scholar]
- French DP, & Sutton S (2010). Reactivity of measurement in health psychology: How much of a problem is it? What can be done about it? British Journal of Health Psychology, 15(3), 453–468. 10.1348/135910710X492341 [DOI] [PubMed] [Google Scholar]
- Gleason ME, Iida M, Shrout PE, & Bolger N (2008). Receiving support as a mixed blessing: Evidence for dual effects of support on psychological outcomes. Journal of Personality and Social Psychology, 94(5), 824–838. 10.1037/0022-3514.94.5.824 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hofer SM, & Piccinin AM (2009). Integrative data analysis through coordination of measurement and analysis protocol across independent longitudinal studies. Psychological Methods, 14(2), 150–164. 10.1037/a0015566 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Johnston M (1999). Mood in chronic disease: Questioning the answers. Current Psychology, 18(1), 71–87. 10.1007/s12144-999-1017-z [DOI] [Google Scholar]
- Knowles ES, Coker MC, Scott RA, Cook DA, & Neville JW (1996). Measurement-induced improvement in anxiety: Mean shifts with repeated assessment. Journal of Personality and Social Psychology, 71(2), 352–363. 10.1037/0022-3514.71.2.352 [DOI] [PubMed] [Google Scholar]
- Lucas CP, Fisher P, Piacentini J, Zhang H, Jensen PS, Shaffer D, Dulcan M, Schwab-Stone M, Regier D, & Canino G (1999). Features of interviews questions associated with attenuation of symptom reports. Journal of Abnormal Child Psychology, 27(6), 429–437. 10.1023/A:1021975824957 [DOI] [PubMed] [Google Scholar]
- McNair DM, Lorr M, & Droppleman LF (1992). Revised manual for the profile of mood states. Educational and Industrial Testing Service. [Google Scholar]
- Mogle J, Muñoz E, Hill NL, Smyth JM, & Sliwinski MJ (2019). Daily memory lapses in adults: Characterization and influence on affect. The Journals of Gerontology: Series B, 74(1), 59–68. 10.1093/geronb/gbx012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Neyeloff JL, Fuchs SC, & Moreira LB (2012). Meta-analyses and Forest plots using a microsoft excel spreadsheet: Step-by-step guide focusing on descriptive data analysis. BMC Research Notes, 5(1), Article 52. 10.1186/1756-0500-5-52 [DOI] [PMC free article] [PubMed] [Google Scholar]
- SAS Institute. (2013). SAS (university edition). SAS Institute. [Google Scholar]
- Schneider S, Junghaenel DU, Ono M, & Stone AA (2018). Temporal dynamics of pain: An application of regime-switching models to ecological momentary assessments in patients with rheumatic diseases. Pain, 159(7), 1346–1358. 10.1097/j.pain.0000000000001215 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Scott SB, Graham-Engeland JE, Engeland CG, Smyth JM, Almeida DM, Katz MJ, Lipton RB, Mogle JA, Munoz E, Ram N, & Sliwinski MJ (2015). The effects of stress on cognitive aging, physiology and emotion (ESCAPE) project. BMC Psychiatry, 15(1), Article 146. 10.1186/s12888-015-0497-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sharpe JP, & Gilbert DG (1998). Effects of repeated administration of the Beck Depression Inventory and other measures of negative mood states. Personality and Individual Differences, 24(4), 457–463. 10.1016/S0191-8869(97)00193-1 [DOI] [Google Scholar]
- Shrout PE, Bolger N, Iida M, Burke C, Gleason ME, & Lane SP (2010). The effects of daily support transactions during acute stress: Results from a diary study of bar exam preparation. In Sullivan K & Davila J (Eds.), Support Processes in Intimate Relationships (pp. 175–199). Oxford University Press. [Google Scholar]
- Shrout PE, Stadler G, Lane SP, McClure MJ, Jackson GL, Clavél FD, Iida M, Gleason MEJ, Xu JH, & Bolger N (2018). Initial elevation bias in subjective reports. Proceedings of the National Academy of Sciences of the United States of America, 115(1), E15–E23. 10.1073/pnas.1712277115 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sliwinski MJ (2008). Measurement-burst designs for social health research. Social and Personality Psychology Compass, 2(1), 245–261. 10.1111/j.1751-9004.2007.00043.x [DOI] [Google Scholar]
- Smyth JM,Zawadzki MJ,Santuzzi AM,&Filipkowski KB (2014). Examining the effects of perceived social support on momentary mood and symptom reports in asthma and arthritis patients. Psychology & Health, 29(7), 813–831. 10.1080/08870446.2014.889139 [DOI] [PubMed] [Google Scholar]
- Stawski RS, MacDonald SWS, & Sliwinski MJ (2015). Measurement burst design. In Whitbourne SK (Ed.), The encyclopedia of adulthood and aging (pp. 1–5).Wiley. 10.1002/9781118521373.wbeaa313 [DOI] [Google Scholar]
- Stone AA, Broderick JE, Schwartz JE, Shiffman S, Litcher-Kelly L, & Calvanese P (2003). Intensive momentary reporting of pain with an electronic diary: Reactivity, compliance, and patient satisfaction. Pain, 104(1–2), 343–351. 10.1016/S0304-3959(03)00040-X [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
