Abstract
Workplace wellness programs aim to improve employee health and lower health care spending. Recent randomized studies have found modest short-run effects on health behaviors, but longer-run effects remain poorly understood. We analyzed a clustered randomized trial of a workplace wellness program implemented at a large multisite US employer. Twenty-five randomly selected treatment worksites received the program, with five of the worksites added at the trial’s midpoint, and 135 randomly selected control worksites did not. The program included modules on nutrition, physical activity, and stress reduction, implemented by registered dietitians. The effects of program availability and participation were assessed. At the end of three years, employees at the treatment worksites had better self-reported health behaviors, including a higher rate of actively managing their weight. No significant differences were found in self-reported health; clinical markers of health; health care spending or use; or absenteeism, tenure, or job performance. Improvements in health behaviors after three years were similar to those at eighteen months, but the longer follow-up did not yield detectable improvements in clinical, economic, or employment outcomes.
Employers have increasingly turned to workplace wellness programs to try to lower health care spending and improve employee health and productivity. Half of small employers and 84 percent of large employers in the US offered workers a wellness program in 2019, focused on topics such as weight loss, smoking cessation, and stress management—with particularly notable growth in the incorporation of personal health assessments and biometric screenings.1 Despite the growth of workplace wellness programs, causal evidence of their effects remains scarce.
Two recent reports of randomized controlled trials, including a predecessor to this longer-run study, found that workplace wellness programs improved some self-reported health behaviors in the first eighteen to twenty-four months, but with little evidence of reduced health care spending, improved objective measures of health, or changes in other outcomes.2–4 Longitudinal observational studies have also shown small or no measurable returns on investment from the lifestyle management component of wellness programs.5–7 These studies contrast with a larger prior literature, which often suggested substantial positive financial returns of wellness programs in the first year or so.8–10 In contrast to randomized controlled trials, these observational studies were more susceptible to selection bias and potential confounding; many also lacked a credible counterfactual or other robust methodologies to establish causal connections. Indeed, although these studies varied in design and rigor, even a carefully selected more rigorous subset of them showed a positive, albeit smaller, financial return on investment in the short run.8
The more recent findings of limited short-run effects raise a key question: Do short-term improvements in health behaviors translate into improvements in other outcomes over a longer period? This question may be particularly important to US employers, who retain workers for a median of 4.2 years, as well as to policy makers and clinicians aiming to improve chronic health conditions.11 In this study we evaluated the effect of a clustered randomized controlled trial of a longitudinal workplace wellness program through three years.
Study Data and Methods
Setting And Intervention
During the period 2015–17 a multicomponent workplace wellness program was implemented at a randomly selected subset of worksites within BJ’s Wholesale Club, a large warehouse retail company that employs about 26,000 workers across more than 200 worksites in the eastern United States. The program comprised twelve content-based modules, spanning February 2015 through June 2017, which focused on key topics in prevention and wellness, including nutrition, physical activity, and stress reduction, each lasting about four to eight weeks (see online appendix exhibits A1 and C1).12 Their content combined individual activities and coaching with group activities, often with modest individual incentives such as a $25 BJ’s Wholesale Club gift card for completing a particular module or modestly larger team incentives. Total potential incentives across the program for an employee averaged about $300.
The intervention was designed and run by an established wellness vendor, Wellness Workdays, which tailored the modules to the BJ’s environment and workforce. Program implementation was led by registered dietitians assigned to treatment worksites, who deployed the modules drawing on common materials; added customized elements for the facilities and employees; taught individual sessions within the modules; and coordinated team-based activities and challenges. BJ’s leadership actively supported the program and encouraged both regional leadership and worksite managers to assist in implementation. Leadership disseminated information about the program to employees, allowed common spaces or break rooms to accommodate program modules, and provided materials and resources in kind to support program activities. Managers encouraged employees through messages and announcements to participate in the modules and on-site surveys and clinical assessments during paid working hours.
The research protocol was approved by the Institutional Review Board at Harvard University. Written informed consent was obtained from all participants before primary data collection. Statistical analyses were prespecified and publicly archived at ClinicalTrials.gov and the American Economic Association’s registry for randomized controlled trials.
Randomization
Among 160 eligible worksites that were geographically accessible by the on-site dietitians and similar to one another in health insurance coverage, the wellness program was implemented in a subset of worksites selected using simple randomization (appendix exhibit B1).12 Worksites, rather than individuals, were randomized, as wellness programs often use team-based interventions that may influence all employees at a worksite. This is one important distinction between our trial and the other recent trial at the University of Illinois, which randomized individuals within a workplace.2,3
Our interim evaluation spanned eighteen months and included twenty treatment worksites that offered the program and collected in-person data. Among 140 control worksites, 20 were randomly selected as “primary controls” with in-person data collection but no wellness program, and 120 were “secondary controls” with no wellness program and collection of only administrative data. At the eighteen-month mark, given additional funding available, 5 additional worksites were randomly selected from the 120 secondary control sites to join the treatment group and another 5 to join the primary control group. This three-year evaluation thus comprised 25 treatment, 25 primary control, and 110 secondary control worksites (appendix exhibit B2).12 In addition, we separately assessed the 20 continuous treatment sites versus controls in subgroup analyses. In-person data collection occurred at the midpoint of the study (eighteen months) and toward the end of the three-year study period, whereas administrative data were collected continuously from all worksites.
We assigned treatment or control status to workers based on their worksites at the time of randomization or initial employment. This ensured that subsequent changes in worksites, which could in theory be related to the wellness program, did not affect assignment. All workers in treatment worksites were eligible but not required to participate, and those who chose to participate could stop at any time.13 Similarly, the in-person data collection comprising a survey and clinical screening was voluntary.
Outcomes
We collected four domains of prespecified outcomes (appendix exhibits C2 and C3).12 Two were gathered in person in the twenty-five treatment and twenty-five primary control sites. “Self-reported Health and Behaviors” were collected from personal health assessment surveys and included self-reported exercise, food choices, smoking, and other information.14 “Clinical Measures of Health” were obtained from biometric screenings by registered nurses and included cholesterol, blood glucose, blood pressure, and body mass index. No imputation was done for any unanswered survey items or unmeasured biometrics. We assessed potential selection into in-person data collection by comparing the baseline characteristics of employees who did and did not participate in surveys and biometrics.
Two domains of administrative data across all 160 worksites were gathered continuously from employment records and health insurance claims. “Health Care Spending and Utilization” were gathered for workers enrolled in employer-sponsored plans through Cigna, the third-party administrator for this self-insured firm. About half of stably employed workers (defined below) and a third of all workers were enrolled in Cigna. Claims-based outcomes were annualized to account for any part-year coverage. “Employment Outcomes” were gathered from employment records and included absenteeism, tenure, and work performance evaluations. There were no missing employment data, although only 66 percent of employees received an evaluation during the study period. Power calculations were conducted before the trial was launched (appendix exhibit A2).12
Statistical Analyses
In primary analyses, we estimated the effect of working at a treatment worksite on outcomes, regardless of participation in the wellness program. For administrative outcomes, we compared all employees at treatment sites with all employees at control sites employed any time during the study period (an intention-to-treat design); for in-person primary data outcomes, we analyzed data for employees who were available during the collection of data (analogous to intention-to-treat, assessed in the available population).
We used an individual-level linear model with an indicator for treatment site and a dose-response measure of the effect of exposure to the intervention (treatment indicator interacted with time worked during the study period) as the key independent variables. Thus, the sum of the coefficient on the treatment indicator and the coefficient on the interaction term estimated at the average level of participation captured the overall effect of randomization into exposure to the treatment. The model used a set of weights that balanced idiosyncratic differences in age, sex, and race/ethnicity between treatment and control. These weights were constructed to tolerate a difference in the standardized means in age, sex, and race/ethnicity up to 0.001 standard deviations between treatment and control while being calibrated to match the characteristics of the entire study population. This strategy has been shown to perform better than a model-based approach that fits a propensity score.15–17 To improve the precision of our estimates, we also controlled for age, sex, age-sex interactions, race/ethnicity, and initial employment characteristics (variables not plausibly affected by the program).
Because multiple measures within a domain may reflect a common outcome, we prespecified standardized treatment effects within the domains of self-reported health, behaviors, and clinical measures of health. Standardized treatment effects are summary measures of related outcomes and capture the mean change across all components in the domain in units of standard deviations (that is, the estimated effect for an outcome relative to the standard deviation of that outcome, averaged across all of the outcomes in the domain).
We also adjusted for multiple inference within domains of outcomes and produced both standard, per comparison p values and adjusted, “family-wise” p values, using a conservative approach of grouping related outcomes following the Westfall-Young method with 1,000 bootstrap replications.18 Standard errors were clustered by worksite. Two-tailed tests were used with a significance level of p < 0.05. Detailed methods are in appendix exhibit A3.12
In addition to the effect of working in a treatment site, we also evaluated the effect of participating in the program. Because participation was voluntary (potentially related to underlying health or prior health behaviors), a simple comparison of participants with nonparticipants risks producing biased estimates. We used a two-stage least squares instrumental variables approach to estimate the local average treatment effect of program participation, with randomization into treatment and exposure to the treatment as the instruments for participation. In our base analysis we defined participation as the completion of any module, but we also tested robustness to other definitions (the number of modules completed or completion of at least three modules). Our base estimates thus indicate the causal effect of participating in at least one module.
We prespecified several secondary analyses to assess heterogeneity and the robustness of our findings. First, we examined heterogeneity of program effects by age, sex, and employment type to assess whether the intervention had different effects for different groups of employees. Second, because effects might be different for workers at the five worksites added to the treatment group at the eighteen-month mark, we analyzed the 20 initial treatment worksites that were fully exposed to three years of the program and the 20 primary and 110 secondary control worksites that analogously were continuously in the control group. Third, given that our prior eighteen-month study used a different sample and empirical specification, we analyzed outcomes through three years using the prior model and weights to facilitate comparison. Fourth, because employee retention could in principle be affected by the wellness program, we estimated effects on a cohort of workers who were stably employed for at least thirteen consecutive weeks before the program began. Fifth, because employers and others might wish to gauge program effectiveness based on worksite-wide outcomes and independent of employee turnover, we assessed aggregate employment and claims outcomes at the worksite level. Sixth, because linear models might not always fit binary outcomes well, we estimated logistic models for binary outcomes to test sensitivity to functional form.
A voluntary program may attract participants who differ from nonparticipants in ways that introduce bias. As described above, our primary analysis was designed to be independent of such selection, but because selective participation is of interest in and of itself, we assessed biased selection into program participation by comparing the baseline characteristics of participants with those of nonparticipants in treatment sites. To assess potential bias in selection into participation in primary data collection, we compared the baseline characteristics of workers who provided survey or biometric data with those of workers who did not.
Last, to examine the role that confounding factors such as biased program participation might play in observational studies compared with our randomized trial approach, we generated observational estimates of program effects using ordinary least squares to compare program participants with nonparticipants (instead of using the variation generated by randomization). Specifically, we conducted two observational analyses: first, we compared program participants in the treatment group with control-group members, and second, within the treatment group, we compared program participants with nonparticipants.
Limitations
Our study had several limitations. First, our evaluation assessed only one wellness program at one company, although it is a multifaceted program fielded at a multisite employer. Similarly, because the same modules were fielded at all treatment sites, we were not able to separately estimate the effect of any particular module. Other programs that differ in content, delivery, work environment, or worker population may generate different effects. Similarly, effects might be different during different periods. For example, effects might accumulate over a longer window if behavior change takes time, or they might attenuate over a longer window if behavior change is hard to maintain or if turnover limits program continuity.
Second, the program may have generated effects that were smaller than we were able to detect, given our sample size and outcomes analyzed. For outcomes such as health care spending, we were likely underpowered despite strategies to maximize power, such as augmenting the sample size during the study period and focusing our analysis on the outcomes and populations most likely affected by the intervention. We implemented a relatively conservative strategy for multiple-inference adjustment, grouping related outcomes within domains. This further eroded statistical significance relative to single-comparison p values.
Third, employees at the original primary control worksites participated in in-person data collection comprising surveys and biometrics at the study midpoint, which may have themselves generated an effect. To the extent that the surveys and biometrics did, our differences between treatment and primary controls would be biased toward the null. This potential bias would be much attenuated in analyses including secondary control worksites, which have a sample size 4.6 times that of the primary controls.
Fourth, not all data were available for all employees, with samples differing across outcomes. Survey and biometric data were available only for people employed during the midpoint and end of the study who chose to participate in the data collection and were present on the days of data collection. Similarly, employees with different tenures had different lengths of exposure to the program. Seventeen percent of the study population and 51 percent of the stably employed subpopulation were employed for the full three years. Without preintervention in-person data, the study largely relied on randomization to produce balanced samples, and causal inference was further aided by balancing weights and covariate adjustment to improve precision. Reassuringly, mean employment duration was similar across treatment and control groups, with the program having no effect on employee tenure, suggesting that entry and exit from the sample were not affected by the intervention, and our stably employed subpopulation was not susceptible to differential entry. We saw some differential selection by age into completing the survey and screening, although not by other characteristics. We controlled for age and these other characteristics. Claims data were available only for those with Cigna coverage. All employees were represented in the absenteeism and tenure outcomes. Overall, data availability was similar between the treatment and control groups, suggesting little bias in the sample representation.
Study Results
Study Population
The study population included 7,288 employees at the 25 treatment worksites, 7,377 at the 25 primary control worksites, and 33,999 at the 110 secondary control worksites. Worksites had a mean of 116 employees at a given time and 304 unique employees during the three-year study period. About 25 percent of the population was Black, and 17 percent was Hispanic. Full-time workers constituted about 50 percent of the study population. Full-time salaried workers earned slightly less than $50,000 per year, and full-time hourly workers earned about half that amount. Demographic and employment characteristics, along with balance between treatment and control groups, are in appendix exhibits C4 and C5.12
Program Exposure and Participation
Program exposure averaged more than 800 days for the employees represented in the individual-level data, as described below. Among all workers ever employed during the study period, including short-term and seasonal workers, the average number of days employed was more than 400.
Program participation increased from 12.3 percent in the first module to 25.6–37.6 percent in the subsequent modules. Overall, 28.4 percent of people ever employed in treatment worksites completed at least one module, 16.7 percent completed at least three modules, and 1.2 modules were completed, on average. Among those completing at least one module, 58.8 percent completed at least three modules, with a mean of 4.1 modules completed. In the stably employed subsample, 50.9 percent completed at least one module, among whom 70.8 percent completed at least three modules, with a mean of 5.2 modules completed (appendix exhibit C6).12
Participation in the personal health assessment survey and biometric screening was 18.2 percent and 18.1 percent, respectively, among people ever employed in the twenty-five treatment or twenty-five primary control worksites during the study period. This reflects in part the fact that only about 44 percent of the employees ever employed during the study period were employed during the primary data collection period. Among the 6,538 people employed at these fifty worksites during the data collection period, mean participation in both surveys and screenings exceeded 40 percent (exhibits 1 and 2).
Exhibit 1:
Effect on self-reported health and behaviors
| Variable | Unadjusted | Effect of availability of wellness program (intention to treat) | Effect of participation in wellness program modules (local average treatment effect) | |||
|---|---|---|---|---|---|---|
| Treatment group mean | Control group mean | Effect | (95% CI) | Effect | (95% CI) | |
| Health behaviors | ||||||
| Screenings and exams | ||||||
| Recommended tests received (%) | 59.9 | 58.0 | 3.49* | (1.13, 5.85) | 4.23 | (0.85, 7.61) |
| Physical activity | ||||||
| Regular exercise (%) | 59.8 | 53.8 | 5.92** | (2.00, 9.84) | 8.69* | (2.78, 14.61) |
| Hours sitting per day (n) | 3.7 | 3.7 | −0.12 | (−0.26, 0.02) | −0.21 | (−0.42, 0.00) |
| Nutrition | ||||||
| Nonzero calorie drinks per day (n) | 0.7 | 0.8 | −0.01 | (−0.13, 0.12) | −0.05 | (−0.21, 0.12) |
| Read the nutrition facts panel (%) | 62.6 | 58.9 | 4.25 | (0.27, 8.24) | 5.78 | (0.70, 10.85) |
| Consume at least 2 cups of fruit and 2.5 cups of vegetables per day (%) | 56.3 | 54.1 | 2.94 | (−0.87, 6.74) | 4.52 | (−1.06, 10.11) |
| Weight management | ||||||
| Actively managing weight (%) | 68.2 | 61.5 | 6.94** | (2.78, 11.10) | 11.05*** | (5.55, 16.55) |
| Tobacco use | ||||||
| Smoker (%) | 13.0 | 14.7 | −1.61 | (−4.99, 1.76) | −1.83 | (−6.75, 3.09) |
| Self-reported health | ||||||
| SF-8 score–physical summary scorea | 51.7 | 51.6 | 0.05 | (−0.54, 0.65) | 0.1 | (−0.72, 0.93) |
| SF-8 score–mental summary score1 | 52.2 | 51.9 | 0.35 | (−0.41, 1.12) | 0.51 | (−0.55, 1.57) |
| Unmanaged stress (%) | 29.8 | 30.7 | −0.78 | (−4.05, 2.49) | −1.73 | (−6.62, 3.15) |
| Unmanaged depression (%) | 30.1 | 30.6 | −0.40 | (−3.36, 2.57) | 0.33 | (−4.38, 5.05) |
| Stress at work (%) | 38.5 | 39.7 | 0.31 | (−4.09, 4.72) | −0.82 | (−7.17, 5.53) |
| Standardized treatment effect 2 | ||||||
| Health behaviors | —b | —b | 0.077**** | (0.04, 0.11) | 0.122**** | (0.06, 0.18) |
| Self-reported health | —b | —b | 0.015 | (−0.04, 0.07) | 0.023 | (−0.06, 0.10) |
SOURCE Authors’ analysis. NOTES All regressions included demographic and employment controls and clustered standard errors at the worksite level. Regressions and means use weights that balance treatment and control samples on age, sex, and race. 95% confidence intervals associated with unadjusted p values are reported in parentheses, and family-wise p values are denoted by asterisks. Because survey questions varied in their number of respondents, sample sizes of the regressions ranged between 2,729 and 2,793 (41.7–42.7 percent of people employed at the twenty-five treatment and twenty-five primary control worksites from August to October 2017). For variables measured as percentages in group means, the program effect is expressed in percentage points.
Scores on the Medical Outcomes Study 8-Item Short-Form Health Survey (SF-8) range from 0 to 100, normalized to a mean of fifty and a standard deviation of ten in the general US population. Higher scores designate better self-reported health-related quality of life.
For standardized treatment effects, the unadjusted p value is denoted by asterisks. The standardized treatment effects were calculated using the outcomes within their subdomains.
p < 0.10
p < 0.05
p < 0.01
p < 0.001
Exhibit 2:
Effect on clinical measures of health
| Variable | Treatment group mean | Control group mean | Effect of availability of wellness program (intention to treat) | Effect of participation in wellness program (local average treatment effect) | ||
|---|---|---|---|---|---|---|
| Effect | (95% CI) | Effect | (95% CI) | |||
| Total cholesterol (mg/dL) | 180.1 | 179.9 | 2.19 | (−1.49, 5.87) | 2.95 | (−2.39, 8.28) |
| High-density lipoprotein cholesterol (mg/dL) | 50.0 | 49.5 | 0.49 | (−0.64, 1.62) | 0.53 | (−1.15, 2.22) |
| Glucose (mg/dL) | 102.7 | 104.8 | −1.57 | (−4.39, 1.24) | −1.86 | (−6.16, 2.43) |
| Blood pressure, systolic (mmHg) | 123.2 | 124.8 | −1.10 | (−2.75, 0.55) | −1.78 | (−4.06, 0.50) |
| Body mass index (kg/m2) | 30.1 | 29.8 | 0.43 | (−0.13, 1.00) | 0.74 | (−0.03, 1.51) |
| Standardized treatment effect1 | —a | —a | 0.004 | (−0.04, 0.05) | 0.006 | (−0.07, 0.08) |
SOURCE Authors’ analysis. NOTES All regressions included demographic and employment controls and clustered standard errors at the worksite level. Regressions and means use a weight that balances treatment and control samples on age, sex, and race. 95% confidence intervals associated with unadjusted p values are reported in parentheses, and family-wise p values were all greater than 0.10. Because biometric measurements varied in their number of participants, sample sizes of the regressions ranged between 2,709 and 2,770 (41.4–42.4 percent of people employed at the twenty-five treatment and twenty-five primary control worksites from August to October 2017). These numbers exceeded the clinical biometrics sample in online appendix exhibit B2 (see note 12 in text) because some people from secondary control worksites unexpectedly took part in the biometrics. For variables measured as percentages in group means, the program effect is expressed in percentage points.
The standardized treatment effect was calculated using all of the clinical outcomes.
Not applicable.
Exhibits 1–4 show the effects of working at a treatment worksite and of participating in the wellness program on outcomes. Full results across domains and for alternative populations are in appendix exhibits C7–C10.12
Intention-Treat Effects
SELF-REPORTED HEALTH AND BEHAVIORS:
The number of participants represented in the data on these outcomes ranged between 2,729 and 2,793 (41.7–42.7 percent of people employed at the treatment and primary control worksites during data collection) (exhibit 1). These people were employed for an average of 813 days, with 50.3 percent employed the full three years (data not shown).
The proportion of employees who reported engaging in regular exercise was 5.9 percentage points higher in treatment sites (control-group mean, 53.8 percent; adjusted p = 0:048), and the proportion who reported actively managing their weight was 6.9 percentage points higher (control-group mean, 61.5 percent; adjusted p = 0.03) (exhibit 1).
Randomization into treatment led to higher self-reports of receiving recommended tests and reading nutrition facts panels that were statistically significant using traditional p values, but not with multiple inference adjustment. Receipt of recommended tests was 3.5 percentage points higher in treatment sites (unadjusted p = 0.004; adjusted p = 0.054). Reading the nutrition facts panel was 4.3 percentage points higher (unadjusted p = 0.04; adjusted p = 0.22). Other outcomes in this domain were not significantly affected by randomization into treatment (all p values > 0.05).
In the standardized treatment effect, health behaviors improved by 0.08 standard deviations p < 0.001). (As a single index, standardized treatment effects do not have adjusted p values.) The standardized treatment effect for self-reported health was not statistically significant.
CLINICAL MEASURES OF HEALTH:
The number of people represented in the data on these outcomes ranged between 2,709 and 2,770 (41.4–42.4 percent of people employed at the treatment and primary control worksites during data collection) (exhibit 2). They were employed for an average of 819 days, with 50.8 percent employed for three years (data not shown). The program did not significantly affect clinical measures of health (all p values > 0.05) or their standardized treatment effect (exhibit 2).
HEALTH CARE SPENDING AND USE:
Employees covered by Cigna represented about one-third of all employees, and they were substantially more stably employed than those not covered by Cigna. They were employed, on average, for twenty-eight months of the three years, during which they had Cigna coverage for twenty-four months, in both the 25 treatment and 135 total control worksites. Claims data included primary and secondary insurer payments (coordination of benefits), as well as cost sharing. There should not, therefore, be systematically missing claims data for these Cigna enrollees.
Randomization into a treatment worksite did not have a detectable effect on total medical spending (−$298 per employee per year; adjusted p = 0.91; control-group mean $4,801) or pharmaceutical spending ($89 per year; adjusted p = 0.93; control-group mean, $1,238) (exhibit 3). Program exposure led to 0.02 fewer hospitalizations per employee per year (control-group mean, 0.09), which was statistically significant under the traditional p value (unadjusted p = 0.02), but not after multiple inference adjustment (adjusted p = 0.12).
Exhibit 3:
Effect on annual health care spending and utilization
| Variable | Treatment group mean | Control group mean | Effect of availability of wellness program (intention to treat) | Effect of participation in wellness program (local average treatment effect) | ||
|---|---|---|---|---|---|---|
| Effect | (95% CI) | Effect | (95% CI) | |||
| Medical spending ($) | ||||||
| Total spending | 4,532 | 4,801 | −297.90 | (−1,201, 605) | −457.54 | (−1,656, 741) |
| Medical utilization | ||||||
| Outpatient visits | 3.96 | 3.90 | 0.06 | (−0.26, 0.38) | 0.07 | (−0.42, 0.56) |
| Hospitalizations | 0.06 | 0.09 | −0.02 | (−0.04, 0.00) | −0.03 | (−0.05, 0.00) |
| Emergency department visits | 0.30 | 0.28 | 0.02 | (−0.03, 0.08) | 0.03 | (−0.04, 0.10) |
| Pharmaceutical spending ($) | ||||||
| Total spending | 1,331 | 1,238 | 88.78 | (−268, 446) | 202.51 | (−362, 767) |
| Pharmaceutical utilization | ||||||
| Distinct medications | 5.65 | 5.49 | 0.21 | (−0.13, 0.56) | 0.41 | (−0.17, 0.99) |
| Medication months | 11.63 | 11.20 | 0.31 | (−0.94, 1.57) | 0.82 | (−1.21, 2.85) |
SOURCE Authors’ analysis. NOTES All regressions included demographic and employment controls and clustered standard errors at the worksite level. Regressions in this exhibit do not control for Cigna coverage status as the data come from claims of these employees with Cigna coverage. Regressions and means use a weight that balances treatment and control samples on age, sex, and race. 95% confidence intervals associated with unadjusted p values are reported in parentheses, and family-wise p values were all greater than 0.10. All employees with Cigna coverage in the twenty-five treatment worksites (1,385 people) and 135 total control worksites (7,174 people) were included. Spending included primary and secondary insurer payments (coordination of benefits) and cost-sharing. Spending outcomes are reported as dollars per person per year, adjusted for inflation to 2016 dollars.
EMPLOYMENT OUTCOMES:
Exhibit 4 shows results for absenteeism, work performance, and job tenure, derived from the full sample of 48,664 employees for absenteeism and tenure and 31,988 for work performance. The program had no significant effects on these outcomes (all p values > 0.05).
Exhibit 4:
Effect on employment outcomes
| Variable | Treatment group mean | Control group mean | Effect of availability of wellness program (intention to treat) | Effect of participation in wellness program (local average treatment effect) | ||
|---|---|---|---|---|---|---|
| Effect | (95% CI) | Effect | (95% CI) | |||
| Absenteeism (% of scheduled hours missed) | 2.1 | 2.2 | −0.08 | (−0.22, 0.06) | −0.21 | (−0.46, 0.04) |
| Performance review (% with good performance)1 | 46.6 | 46.6 | 0.62 | (−4.69, 5.93) | 0.55 | (−9.92, 11.01) |
| Tenure (days employed during the treatment period) | 411.0 | 416.7 | 5.43 | (−6.44, 17.29) | 14.93 | (−6.90, 36.76) |
SOURCE Authors’ analysis. NOTES All regressions included demographic and employment controls and clustered standard errors at the worksite level. Regressions and means use a weight that balances treatment and control samples on age, sex, and race. 95% confidence intervals associated with unadjusted p values are reported in parentheses, and family-wise p values were all greater than 0.10. These data were collected from the twenty-five treatment and 135 total control worksites across the study period. Because of variation in the number of performance reviews that employees received, including those who did not receive a performance review during the study period, the sample sizes were 48,664 (100% of all employees) for absenteeism and tenure and 31,988 (65.7% of all employees) for performance reviews. For variables measured as percentages in group means, the program effect is expressed in percentage points. The maximum tenure during the study period was 1,093 days.
Performance reviews were scored from 1 (best) to 5 (poorest). Given variation in the number of performance reviews received during the study period, this outcome averaged available performance review scores for each person (weighted by the duration over which a score was held). In this binary outcome measure, a score of 1–2 was good performance and 3–5 was poor performance.
Local Average Treatment Effects
Exhibits 1–4 (right-hand columns) also show the effect of participation in the program, defined as completing at least one module. Participating led to a 11.1-percentage-point higher share of employees who reported actively managing weight compared with control (adjusted p = 0.005). Participation also led to an 8.7-percentage-point higher share of employees who reported regular exercise, but this was not significant after multiple inference adjustment (adjusted p = 0.07). Other self-reported health behaviors were not affected by program participation, but health behaviors were 0.12 standard deviations better (p < 0:001) among participants overall relative to control. Participation had no detectable effects on self-reported health or clinical measures of health (exhibits 1 and 2).
Participation led to 0.03 fewer hospitalizations per employee per year (unadjusted p = 0.02), but this was not significant after multiple inference adjustment (adjusted p = 0.13). Medical spending and employment outcomes were not affected by program participation (exhibits 3 and 4).
Heterogeneity Analyses
There were no significant differences in the effect of program exposure and participation on active weight management on the basis of age, sex, or employment type, although the effect was more statistically significant for full-time than part-time workers. Similar effects on regular exercise were seen across age groups and employment types, although they were more significant for female workers. Despite there being no significant effects on clinical measures of health in the sample overall, program exposure led to 2.9 mmHg lower systolic blood pressure among older workers (unadjusted p = 0.009) and a 1.1 kg/m2 higher body mass index among younger workers (unadjusted p = 0.004). Although employment outcomes were unaffected overall, program exposure led to lower absenteeism of 0.2 percentage points among full-time workers (unadjusted p = 0.005). Participation led to similar heterogeneous effects (appendix exhibit C11).12
Analyses of the 20 initial treatment worksites fully exposed to the program during the three years compared with the 20 primary and 110 secondary control worksites fully unexposed through the three years yielded similar results (appendix exhibit C12).12
Secondary and Sensitivity Analyses
Alternative definitions of program participation produced similar results (appendix exhibit C13).12 Effects from the stably employed subsample were similar to those from the full sample (appendix exhibits C7–10).12 Effects at the worksite level were similar to those at the individual level for claims-based and employment outcomes (appendix exhibits C9–10).12
Relative to program effects at eighteen months, estimates after three years were attenuated for some outcomes, including active weight management and regular exercise; reproducing our main results at three years using the statistical model and sample from the earlier study produced similar results (appendix exhibit C14).12 Logistic regressions for binary outcomes also produced similar results (appendix exhibit C15).12
Selection into Program Participation
Comparisons of preintervention characteristics between workers in treatment worksites showed differences in some employee characteristics between those who participated in the program and those who did not. Participants were more likely than nonparticipants to be female, White, full-time workers, and in sales positions. Preintervention health care spending and earnings did not differ significantly between participants and nonparticipants, with participation defined as completion of any module. With participation defined as completion of at least three modules, participants were more likely than nonparticipants to have incurred any medical spending preintervention (appendix exhibit C16).12 Younger workers were more likely to select into participating in survey and biometric data collection in the treatment worksites than in control worksites, but no other evidence of differential selection into in-person data collection was evident (appendix exhibit C17).12
Analyses using two observational strategies (comparing participants with the control group and comparing participants with nonparticipants in the treatment group, both capturing potentially biased participation decisions) found multiple significant effects that were not found using the randomized design (appendix exhibit C18).12 For example, observational analyses found the program to be associated with lower rates of smoking, higher rates of screenings and exams, and better job performance. Although differences between observational and randomized controlled results were variable, these differences underscore the importance of randomization to obtain unbiased estimates.
Discussion
This three-year randomized clinical trial of a multicomponent workplace wellness program in a middle- and lower-income worker population found that exposure to and participation in the program led to better self-reported health behaviors, notably including active weight management. These improvements were substantial in magnitude, with about an 11 percent increase in the share of employees who reported actively managing their weight in treatment relative to control sites, or about an 18 percent increase among those participating in the wellness program. The program did not, however, produce significant differences in clinical measures of health, health care spending and use, or employment outcomes between treatment and control. We note, however, that 95% confidence intervals were large for some estimates, so we could not rule out substantial effects for some outcomes.
Improvements in self-reported health behaviors attenuated somewhat after the eighteen-month assessment, although they generally continued to be observed at three years. However, at both assessments, the program did not produce measurably better health, savings on medical or prescription drug spending, or improved absenteeism or job performance (although some heterogeneous effects by age, sex, and worker type were found). Thus, behavior change, as measured through self-reports, may be easier to affect or sustain than durable health or employment outcomes. Although it is difficult to price the unobserved benefits of behavior change, we did not observe a financial return on investment in the employment and claims measures we examined.
These results are broadly consistent with initial results from a randomized evaluation of a wellness program implemented at the University of Illinois.2,3 However, they differ from much of the prior literature on workplace wellness programs, which generally used nonrandomized designs and often found positive and large returns on investment.8–10 Given selection bias and other methodological concerns in observational studies, randomized trials likely provide more reliable estimates of program effects. However, results from any specific wellness program, employer, or worker population might not generalize to other programs or settings, and rigorous observational studies with empirical strategies to minimize bias can provide important complementary evidence.
Conclusion
In this randomized controlled trial of a workplace wellness program, exposure to the program for up to three years led to higher shares of employees reporting better health behaviors at the end of the study. However, there were no significant differences in clinical measures of health, health care spending and use, or employment outcomes between treatment and control groups. To the extent that these results are representative of other wellness programs, they temper expectations of substantial improvements in health outcomes or financial returns on investment from wellness programs up to a three-year horizon.
Supplementary Material
Acknowledgements
This work was supported by the National Institute on Aging (Grant No. R01AG050329 and Grant No. P30AG012810 through the National Bureau of Economic Research), Robert Wood Johnson Foundation (Grant No. 72611), and Abdul Latif Jameel Poverty Action Lab North America. BJ’s Wholesale Club provided in-kind logistical and personnel support for the fielding of the wellness program. The funders had no role in the design or conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the manuscript; or decision to submit the manuscript for publication. The findings and conclusions expressed are solely those of the authors, and do not represent the views of any of the funders. Zirui Song has served as a consultant on Medicare risk adjustment for the Research Triangle Institute, given a guest lecture for GV and the International Foundation of Employee Benefit Plans educational program outside of this work, and provided consultation in legal cases outside of this work. Katherine Baicker serves on the board of directors of Eli Lilly, Mayo Clinic, NORC, and the Chicago Council on Global Affairs and on the advisory boards for the Congressional Budget Office and the National Institute for Health Care Management Foundation. The authors thank José Zubizarreta and Yige Li (Harvard Medical School) for statistical guidance and contributions to the sample weights without financial compensation. They also thank Sherri Rose (Harvard Medical School) for statistical guidance on randomization without financial compensation; David Molitor and Julian Reif (University of Illinois at Urbana-Champaign) for guidance on the statistical software for multiple inference adjustment they created in the University of Illinois wellness trial, which was used in this study, without financial compensation; Ozlem Barin and Erica Paulos (Harvard Medical School) and Kathryn Clark and Bethany Maylone (Harvard T. H. Chan School of Public Health at the time of this study) for research assistance and project management; Josephine Fisher, Jack Huang, Harlan Pittell, and Artemis (Yuanxiaoyue) Yang (Harvard T. H. Chan School of Public Health at the time of this study) for research assistance; and Luke Sonnet (University of California Los Angeles) for replicating the midpoint study results through the Abdul Latif Jameel Poverty Action Lab’s Research Transparency and Reproducibility Initiative, without financial compensation. The authors also thank the study partners, BJ’s Wholesale Club and Wellness Workdays, for collaboration and assistance in the design and fielding of the workplace wellness program. Finally, they thank seminar participants at the American Society of Health Economists, Harvard T. H. Chan School of Public Health, and Harvard Medical School for comments and suggestions made without financial compensation.
Contributor Information
Zirui Song, assistant professor of health care policy and medicine at Harvard Medical School, general internist at Massachusetts General Hospital, and faculty member in the Center for Primary Care at Harvard Medical School, in Boston, Massachusetts..
Katherine Baicker, dean of and the Emmett Dedmon Professor in the Harris School of Public Policy, University of Chicago, in Chicago, Illinois..
Endnotes
- 1.Henry J Kaiser Family Foundation. 2019Employer Health Benefits Survey [Internet]. San Francisco (CA): KFF; 2019September25 [cited 2021 Mar 26]. Available from: https://www.kff.org/health-costs/report/2019-employer-health-benefits-survey/ [Google Scholar]
- 2.Jones D, Molitor D, Reif J. What do workplace wellness programs do? Evidence from the Illinois Workplace Wellness Study. Q J Econ. 2019; 134(4):1747–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Reif J, Chan D, Jones D, Payne L, Molitor D. Effects of a workplace wellness program on employee health, health beliefs, and medical use: a randomized clinical trial. JAMA Intern Med. 2020;180(7): 952–60. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Song Z, Baicker K. Effect of a workplace wellness program on employee health and economic outcomes: a randomized clinical trial. JAMA. 2019;321(15):1491–501. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Nyman JA, Abraham JM, Jeffery MM, Barleen NA. The effectiveness of a health promotion program after 3 years: evidence from the University of Minnesota. Med Care. 2012; 50(9):772–8. [DOI] [PubMed] [Google Scholar]
- 6.Caloyeras JP, Liu H, Exum E, Broderick M, Mattke S. Managing manifest diseases, but not health risks, saved PepsiCo money over seven years. Health Aff (Millwood). 2014;33(1):124–31. [DOI] [PubMed] [Google Scholar]
- 7.Mattke S, Liu HH, Caloyeras JP, Huang CY, Van Busum KR, Khodyakov D, et al. Workplace wellness programs study [Internet]. Santa Monica (CA): RAND Corporation; 2013. [cited 2021 Apr 14]. (Publication No. RR-254-DOL). Available from: https://www.rand.org/pubs/research_reports/RR254.html [PMC free article] [PubMed] [Google Scholar]
- 8.Baicker K, Cutler D, Song Z. Workplace wellness programs can generate savings. Health Aff (Millwood). 2010;29(2):304–11. [DOI] [PubMed] [Google Scholar]
- 9.Goetzel RZ, Ozminkowski RJ. The health and cost benefits of work site health-promotion programs. Annu Rev Public Health. 2008;29:303–23. [DOI] [PubMed] [Google Scholar]
- 10.Pelletier KR. A review and analysis of the clinical and cost-effectiveness studies of comprehensive health promotion and disease management programs at the worksite: update VIII 2008 to 2010. J Occup Environ Med. 2011;53(11):1310–31. [DOI] [PubMed] [Google Scholar]
- 11.Bureau of Labor Statistics [Internet]. Washington (DC): BLS. Press release, Employee tenure in 2018; 2018September20 [cited 2021 Mar 16]. Available from: https://stats.bls.gov/news.release/archives/tenure_09202018.pdf [Google Scholar]
- 12.To access the appendix, click on the Details tab of the article online.
- 13.Mello MM, Rosenthal MB. Wellness programs and lifestyle discrimination—the legal limits. N Engl J Med. 2008;359(2):192–9. [DOI] [PubMed] [Google Scholar]
- 14.Ware JE, Kosinski M, Dewey JE, Gandek B. How to score and interpret single-item health status measures: a manual for users of the SF-8 Health Survey. Lincoln (RI): QualityMetric; 2001. [Google Scholar]
- 15.Zubizarreta JR. Stable weights that balance covariates for estimation with incomplete outcome data. J Am Stat Assoc. 2015;110(511):910–22. [Google Scholar]
- 16.Wang Y, Zubizarreta JR. Minimal dispersion approximately balancing weights: asymptotic properties and practical considerations. Biometrika. 2020;107(1):93–105. [Google Scholar]
- 17.Hirshberg DA, Zubizarreta JR. On two approaches to weighting in causal inference. Epidemiology. 2017;28(6):812–6. [DOI] [PubMed] [Google Scholar]
- 18.Westfall PH, Young SS. Resamplingbased multiple testing: examples and methods for p-value adjustment. First edition. New York (NY): Wiley & Sons; 1993. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
