Skip to main content
Annals of Behavioral Medicine: A Publication of the Society of Behavioral Medicine logoLink to Annals of Behavioral Medicine: A Publication of the Society of Behavioral Medicine
. 2025 Apr 21;59(1):kaaf026. doi: 10.1093/abm/kaaf026

Improving the design and analytic methods used in NIH-funded clinical trials involving behavioral interventions

David M Murray 1,, Jane M Simoni 2
PMCID: PMC12010243  PMID: 40257118

Abstract

Behavioral interventions are widely used in clinical trials supported by the National Institutes of Health (NIH). When behavioral interventions involve group-formatted components and/or shared interventionists, they require special design and analytic methods not needed in trials that do not involve these features. The NIH Office of Disease Prevention (ODP) and the NIH Office of Behavioral and Social Sciences Research (OBSSR) offer resources to make it easier for investigators to use appropriate methods to evaluate these interventions. This commentary draws attention to these issues and highlights the ODP and OBSSR resources available to investigators. We urge investigators to take advantage of these resources to learn about and adopt appropriate sample size and analytic methods for trials to evaluate behavioral interventions so that their results will be reliable and reproducible. That is the best way to advance the science of behavioral interventions to improve health.

Keywords: behavioral intervention, study design, analytic methods, sample size methods


Investigators are urged to take advantage of resources from NIH and identified in this commentary to learn about and adopt appropriate sample-size and analytic methods for trials to evaluate behavioral interventions so that their results will be reliable and reproducible.


In fiscal year 2023, the National Institutes of Health (NIH) funded 380 new clinical trials through its extramural research program. Of these, 239 (63%) proposed to evaluate a behavioral intervention. Such interventions often randomize groups or clusters rather than individuals or use group-formatted intervention components and/or shared interventionists; when they do, they require special design and analytic methods not needed in trials that do not involve these features.1 The purpose of this commentary is to draw attention to this issue and to resources provided by the NIH Office of Disease Prevention (ODP) and Office of Behavioral and Social Sciences Research (OBSSR) to help investigators use appropriate methods for these trials.

When individuals are randomized to study arms, participants receive the intervention appropriate for their arm, and participants remain independent for the duration of the trial, investigators can use standard methods for the design, sample size estimation, and analysis of data (eg,2). This randomized controlled trial (RCT) is the strongest and most efficient design for evaluating a behavioral intervention. An example would be a trial evaluating a mHealth intervention where the intervention is delivered entirely by an app or website and participants have no interaction post-randomization with each other or with a shared interventionist.

When individuals are randomized to study arms but logistical or other considerations require delivery of intervention components in a group format (in person or virtual) or by shared interventionists (eg, therapists, facilitators, trainers, clinicians), responses from participants will become correlated over the course of the trial (eg,3). Failure to reflect the expected positive intraclass correlation (ICC) in the sample-size calculation methods will result in insufficient power for a valid analysis while failure to reflect the expected ICC in the analysis will result in an invalid analysis with an inflated type 1 error rate. These are Individually Randomized Group-Treatment (IRGT) trials.1 An example would be a weight loss or smoking cessation trial where participants are randomized to study arms but those in the intervention arm receive some portion of their intervention in a group format or with a shared interventionist. Literature reviews suggest that investigators are usually unaware of the need to address IRGT trial issues in planning their study, with only 5%–8% of published IRGT trials reporting appropriate methods (eg,4). Analytic and sample size methods are readily available for IRGT trials1; as such, investigators who wish to evaluate a behavioral intervention that includes group-formatted components and/or shared intervention agents should become familiar with the appropriate methods for IRGT trials.

When the intervention manipulates the physical or social environment, involves group processes, or cannot be delivered to individuals without serious risk of contamination, groups or clusters may be randomized to study arms and members of those groups or clusters measured to assess the impact of the intervention. This design is a parallel group- or cluster-randomized trial (GRT).1 It is common for GRTs to involve randomization of existing groups or clusters such as schools, families, or dyads. An example would be a trial to evaluate an intervention to prevent alcohol, tobacco, or other drug use among adolescents. Fearing contamination if adolescents in the same school are randomized to intervention and control arms, schools are randomized to study arms and all students in the school receive either the intervention or the control experience. The proportion of parallel GRT designs that report appropriate analytic and sample-size methods (57%) has been much higher than for IRGT trials, but there remain far too many trials that employ a GRT design but use inappropriate methods. It would be important for investigators who wish to evaluate a behavioral intervention that manipulates the physical or social environment, involves group processes, or cannot be delivered to individuals without serious risk of contamination to become familiar with the appropriate methods for parallel GRT designs.

If the investigators believe that it is important for all participants to receive the intervention before the end of the trial, groups or clusters may be randomized to sequences and transition from control to intervention on a staggered schedule according to their sequence. At the beginning, all groups or clusters will be in the control condition but by the end of the trial, all groups or clusters will be in the intervention condition. This design is a stepped-wedge group- or cluster-randomized trial (SWGRT).1 An example would be a trial to evaluate a change to the way that care for a specific condition is delivered. Practices would be randomized to sequences and transition from control to intervention in a staggered order based on their sequence. Patients within each practice receive the control condition initially, but once the practice transitions to the intervention arm, they may receive the intervention. SWGRTs face more serious threats to internal validity than parallel GRTs, often because they are long and involve many repeated measures on each group or cluster; under those conditions, they are at greater risk of the effect of external events and it is important to accurately model the pattern of correlation over time, They also face serious risks for inaccurate effect estimates and standard errors in the presence of time-varying intervention effects (eg,5). Because SWGRTs face more serious risks than the other 3 designs, it is particularly important for investigators planning a new SWGRT to evaluate a behavioral intervention to become familiar with the appropriate design, analytic, and sample size methods for this design.

IRGT trials, GRTs, and SWGRTs will usually require a larger sample size than RCTs, with the magnitude of the increase dependent on the patterns of correlation expected in the data, particularly the ICC. The increase is often greatest for GRTs and least for IRGTs compared to RCTs. Even so, the choice of the design should be made based on the methods planned for the assignment of participants to study arms and for the delivery of interventions to participants, not just on the required sample size.1

These trials require analytic methods that accommodate the expected patterns of correlation in the data. Mixed or hierarchical models are the most common but generalized estimating equations, latent growth models, and related methods can be used.1

ODP has made a substantial investment over the last 12 years in developing resources to make it easier for investigators to use appropriate methods in research supported by NIH, including studies testing behavioral interventions for prevention and treatment. ODP created and supports the Research Methods Resources website to provide information, key references, sample size calculators, and other resources for parallel GRTs, SWGRTs, and IRGTs. The website was recently expanded to include parallel material for group or cluster regression discontinuity designs. This website would be an excellent starting point for investigators who wish to become familiar with the appropriate analytic and sample size estimation methods for these designs. Additionally, ODP hosts the Methods: Mind the Gap webinar series to provide a forum for presentation and discussion of recent methodological research relevant to clinical trials supported by NIH, including those involving behavioral interventions. ODP supports a 4-part online course on the Design and Analysis of Pragmatic and Group-Randomized Trials in Public Health and Medicine. ODP created a new web page to draw attention to Research Methods for Multilevel Interventions to Reduce Health Disparities after sponsoring a Supplemental Issue in Prevention Science focused on this topic. Investigators can find good examples of studies designed to evaluate behavioral interventions to reduce health disparities in this Supplemental Issue.

OBSSR also has a substantial interest in promoting the use of the most appropriate methods in clinical trials involving behavioral interventions. Since 2000, OBSSR has supported the annual Summer Institute on Randomized Behavioral Clinical Trials, providing training in state-of-the-science methods for such trials, including methods for parallel GRTs, SWGRTs, and IRGTs. The Office also supports the Good Clinical Practice for Social and Behavioral Research eLearning Course, which helps investigators and clinical trial staff protect the rights, safety, and well-being of human subjects; it ensures that clinical trials are conducted according to approved plans, and that the data collected are reliable.

We recognize that the sample size and cost of an IRGT, GRT, or SWGRT are usually greater than for an RCT. Even so, using the wrong design because it allows a smaller and less expensive study could lead to a false-positive finding and mislead other investigators and policymakers into pursuing a behavioral intervention that is not supported by good science. Instead, the research design should be selected to reflect the research question and the methods planned for the assignment of participants to study arms and for delivery of the behavioral intervention to participants, following current guidance for selecting from among these designs.1 We urge investigators to take advantage of the resources offered by ODP and OBSSR to learn about and adopt appropriate sample sizes and analytic methods for trials to evaluate behavioral interventions so that their results will be reliable and reproducible. That is the best way to advance the science of behavioral interventions to improve health.

Contributor Information

David M Murray, Associate Director for Prevention, Director, Office of Disease Prevention, National Institutes of Health, 6705 Rockledge Drive, Room 734B, Bethesda, MD 20892, United States.

Jane M Simoni, Associate Director for Behavioral and Social Sciences Research, Director, Office of Behavioral and Social Sciences Research, National Institutes of Health, 31 Center Drive, Building 31, Room B1C19, Bethesda, MD 20892, United States.

Author contributions

Dr. Murray conceptualized the article and wrote the original draft. Dr. Simoni reviewed and edited the draft. Both Dr. Murray and Dr. Simoni approved the final version.

Funding

There was no funding for this work beyond the salaries provided by NIH to the authors.

Human subjects and animals

This article does not contain any studies with human participants or animals performed by any of the authors.

References

  • 1. Murray  DM, Taljaard  M, Turner  EL, George  SM.  Essential ingredients and innovations in the design and analysis of group-randomized trials. Annu Rev Public Health.  2020;41:1-19. https://doi.org/ 10.1146/annurev-publhealth-040119-094027 [DOI] [PubMed] [Google Scholar]
  • 2. Friedman  LM, Furberg  CD, DeMets  DL, Reboussin  DM, Granger  CB.  Fundamentals of Clinical Trials. 5th ed. Springer; 2015. [Google Scholar]
  • 3. Roberts  C, Roberts Stephen  A, Roberts  SA.  Design and analysis of clinical trials with clustering effects due to treatment. Clin Trials.  2005;2:152-162. https://doi.org/ 10.1191/1740774505cn076oa [DOI] [PubMed] [Google Scholar]
  • 4. Pals  SL, Murray  DM, Alfano  CM, Shadish  WR, Hannan  PJ, Baker  WL.  Individually randomized group treatment trials: a critical appraisal of frequently used design and analytic approaches. Am J Public Health.  2008;98:1418-1424. https://doi.org/ 10.2105/AJPH.2007.127027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Hughes  JP, Lee  WY, Troxel  AB, Heagerty  PJ.  Sample size calculations for stepped wedge designs with treatment effects that may change with the duration of time under intervention. Prev Sci.  2024;25:348-355. https://doi.org/ 10.1007/s11121-023-01587-1 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Annals of Behavioral Medicine: A Publication of the Society of Behavioral Medicine are provided here courtesy of Oxford University Press

RESOURCES