Skip to main content
Trials logoLink to Trials
. 2025 Dec 18;27:70. doi: 10.1186/s13063-025-09352-1

Development of a consensus extension of the estimands framework for cluster randomised trials (CRT-estimands): results from an international Delphi study

Brennan C Kahan 1,✉, Melanie Bahti 2, Dongquan Bi 1, Frank Bretz 3,4, Gary S Collins 5,6, Andrew Copas 1, Michael O Harhay 7, Fan Li 8,9, Catherine L Auriemma 2,10
PMCID: PMC12822161  PMID: 41413910

Abstract

Background

Estimands are increasingly used in randomised trials to clarify research objectives. The ICH E9(R1) addendum sets out five attributes necessary to describe a well-defined estimand. However, the addendum was primarily developed for individually randomised trials. There is growing recognition that estimand descriptions for cluster randomised trials, where groups of individuals are randomised, may require specification of additional considerations. We conducted a Delphi study to assess stakeholder views on additional items for inclusion in a consensus extension of the ICH E9(R1) for cluster randomised trials.

Methods

We invited experts in estimands and cluster randomised trials to participate in a modified Delphi process to identify critical items for describing estimands in cluster randomised trials. The research team generated an initial list of eight items and definitions. Across three Delphi rounds, panellists scored items, suggested additional items, and provided open-ended rationales for responses. The consensus threshold was set as ≥ 70% of respondents rating an attribute as “essential” (i.e. score of ≥ 7 on a 9-point Likert scale) and < 15% of respondents rating the item as “not important” (i.e. a score of ≤ 3).

Results

Seventy-three (52%) invited individuals participated in Round 1. Response rates were 85% in Round 2 and 95% in Round 3. Panellists included largely statisticians (62, 85%) and clinical trialists (18, 25%). After Round 1, one additional item was added for Round 2 inclusion. After Round 3, five items met consensus criteria: how individuals and clusters are weighted, population of clusters, exposure time of clusters and individuals to the intervention, whether treatment effects are marginal or cluster-specific, and handling of cluster-level intercurrent events.

Conclusions

This Delphi identified expert consensus around the importance of several key items for defining estimands in cluster randomised trials. These results can inform the development of consensus guidance outlining the set of attributes to describe when defining estimands for cluster randomised trials.

Supplementary Information

The online version contains supplementary material available at 10.1186/s13063-025-09352-1.

Keywords: Cluster randomised trials, Estimands, ICH E9(R1), Delphi study, Consensus guidelines

Background

Since publication of the ICH E9(R1) addendum [1], the use of estimands in randomised trials has rapidly increased [2]. An estimand is a precise description of the treatment effect a trial sets out to quantify and can be used to enhance clarity around the trial objectives and ensure alignment between a trial’s methods and its objectives [1, 3–8]. The ICH E9(R1) addendum sets out five attributes that should be described in order to have a well-defined estimand (population, treatment conditions, endpoint, population-level summary measure, and strategies to handle intercurrent events) [1].

However, the ICH E9(R1) addendum was primarily developed for individually randomised trials. Cluster randomised trials involve randomising groups of individuals, such as hospitals, schools, or villages, between treatments [9–11]. With the increased use of estimands, there is growing recognition that while the five attributes specified in the addendum will be applicable for all trials, cluster randomised trials may require additional considerations in order to have a fully defined estimand [12–20]. For instance, a notable example highlighted in the literature is how individuals and clusters are weighted in the estimand definition [12, 13, 19–22]. Commonly used estimators for cluster randomised trials use different weighting schemes (e.g. some give equal weight to individuals, some to clusters, and some weight by the inverse-variance of the cluster). This is separate from how individuals and clusters are weighted in the estimand; however, it does have implications for which estimand is being targeted by the estimator. Therefore, the choice of which estimator is used can implicitly lead to a different estimand being targeted and hence a different size of effect [12, 13, 20].

We therefore convened the CRT-Estimands executive committee (BCK, MB, DB, FB, GSC, AC, MOH, FL, CA), with the aim of developing a consensus extension of the ICH E9(R1) addendum for cluster randomised trials. The objective of this extension is to provide guidance on which attributes should be described when defining estimands in cluster randomised trials. As part of developing this guidance, we have conducted a review of published cluster randomised trials which motivated the need for new guidance [23] and a scoping review which identified potential additional items that may be included in the guidance [24].

In this article, we present results from a three-stage Delphi study conducted to elicit stakeholder perspectives and assess consensus on potential items to be included in the final guidance. The results from this Delphi study will be used to inform a consensus meeting, in which the final guidance will be decided.

Methods

We conducted a three-round modified Delphi consensus process [25–27] using similar methodology to previous international consensus studies our team has led or been involved in [28–36]. The modified Delphi process was an iterative approach that allowed stakeholders to (1) rate items for inclusion in the guidance and explain their ratings; (2) suggest additional items to include; (3) review the ratings and explanations of other participants; and (4) subsequently revise responses based upon information learned in a prior round. This study was approved as exempt by the University of Pennsylvania IRB (protocol #856778). The protocol was posted prospectively on the Open Science Framework on 16 September 2024 (https://osf.io/a5kdy/). This study is reported following the ACCORD guidelines for consensus-based research (Supplement) [37].

Participant recruitment

We convened a panel of stakeholders with expertise in cluster randomised trials and/or estimands, including statisticians, methodologists, clinicians, and other healthcare professionals, as well as researchers from related fields. Potential experts were identified through the authorship team’s professional networks, review of author lists from a recent scoping review [24] and through a process of “snowballing” in which invited experts could recommend additional colleagues for participation in the study [38]. Members of the executive committee were invited to participate in the Delphi. Prospective panellists were invited to participate directly by the authors via email and were informed that completing the study surveys indicated informed consent to participate.

Survey design and administration

The overarching objective of this Delphi study was to identify items with expert consensus for consideration of inclusion in a consensus extension of the ICH E9(R1) addendum for cluster randomised trials. An initial list of items was identified through a scoping review of literature addressing estimands in cluster randomised trials [24]. Descriptions and explanations of each item were developed by the executive committee iteratively through internal piloting for readability of explanations (Table 1).

Table 1.

All Items and definitions. All items and explanations relate to the estimand definition. However, certain elements also need to be defined for the analysis method (estimator). For instance, how individuals and clusters are weighted needs to be defined for the estimand, but a weighting scheme also needs to be defined for the estimator (e.g. some statistical estimators, such as unweighted independence estimating equations, give equal weight to each individual; some estimators, such as the unweighted analysis of cluster-level summaries, give equal weight to each cluster; and some estimators, such as mixed-effects models, weight by the inverse-variance of the cluster)

graphic file with name 13063_2025_9352_Tab1_HTML.jpg

Anonymity of individual panellists’ responses was maintained throughout the Delphi process. The REDCap online survey platform hosted at the University of Pennsylvania was used to collect demographic information as well as to administer the Delphi surveys [39, 40]. Members of the executive group piloted the surveys to ensure functionality and clarity of questions. Participants received up to three reminder emails to complete each survey round. The Delphi study was conducted between 16 October 2024 and 10 February 2025.

Round 1

In Round 1, participants reviewed the initial list of eight items and the definitions prepared by the executive committee. Participants were asked to rate the importance of each item and, optionally, explain their ratings for each in open-ended text. After rating the importance of the initial eight items, participants were invited to suggest additional items to be considered for inclusion in the guidance. Round 1 was conducted between 16 and 31 October 2024.

Round 2

Following the completion of Round 1, participant responses were compiled and summarised in an executive summary distributed to panellists by email along with an invitation to complete Round 2 (Supplement). Summary ratings were presented graphically. A selection of comments from participants’ free-text explanations was summarised by the executive committee to ensure that comments across a range of ratings were represented and to communicate any substantive concerns or reasoning raised by participants. A full list of all comments was available to participants by hyperlink. Participants were asked to review the summary prior to initiating the Round 2 survey.

The Round 2 survey included the initial eight items that were rated in Round 1 and an additional ninth item that was developed based on suggestions made by multiple participants in Round 1. Round 2 included an importance-rating question for all nine items. For the first eight items, participants were able to view their own rating and the average rating from Round 1, a graphical representation of Round 1 ratings, and the selected comments from the executive summary embedded within the survey platform (in addition to their inclusion in the executive summary). For the ninth item, which was rated for the first time, participants could optionally provide an explanation of their importance rating in open-ended text. Only individuals who completed Round 1 were invited to complete Round 2. Round 2 was conducted between 20 November and 9 December 2024.

Round 3

Following Round 2, participant responses were again compiled and summarised in an executive summary distributed to panellists by email along with an invitation to complete Round 3 (Supplement). The summary again included a graphical representation of ratings for the item and a representative selection of participant comments from Round 2. A full list of comments was available to participants by hyperlink. Participants were asked to review the summary prior to completing the Round 3 survey.

The Round 3 survey included only the new item from Round 2, so it could be rated a second time. Participants were able to view their own rating and the average rating from Round 2, a graphical representation of Round 2 ratings, and the selected comments from the executive summary embedded within the survey platform (in addition to their inclusion in the executive summary). Participants were asked to rate the importance of item 9 and invited to provide any additional comments about the consensus extension of the ICH E9(R1) addendum for cluster randomised trials in open-ended text. Only individuals who completed Round 2 were invited to complete Round 3. Round 3 was conducted between 24 January and 10 February 2025.

Statistical reporting and analysis

Response rates were defined as the proportion of invited panellists who completed each survey. Quantitative responses were summarised using descriptive statistics. The importance of an item was rated on a 9-point Likert scale ranging from 1 (“not important”) to 9 (“critical”). Participants could also select “unable to rate”. The consensus threshold for considering an item as essential was set a priori as ≥ 70% of participants rating an item as “critical” (i.e., rating ≥ 7) and < 15% of respondents rating the item as “not important” (i.e., rating of ≤ 3). The consensus threshold for considering an item as not essential for inclusion in the guidance was set a priori as ≥ 70% of participants rating an item as “not important” (i.e. rating ≤ 3) and < 15% rating it as “critical” (i.e. rating ≥ 7). Any other combination of ratings indicated no consensus about whether or not to include the item. Similar consensus definitions have been used in prior studies to ensure that an item will not achieve consensus if a subset of stakeholders commonly rates it as not important [28, 29, 41, 42]. Final assessments of consensus were determined after an item had been rated twice (after Round 2 for Items 1–8 and after Round 3 for Item 9). Free-text responses were analysed using comparison techniques across the range of observed ratings within an item and across items by concerns or reasoning raised by participants.

Results

Invitations to participate in the Delphi were initially sent to 114 individuals (Fig. 1). From snowball sampling, 21 participants made a total of 50 suggestions for additional people to invite to the Delphi. Of the 50 suggestions, there were 42 unique individuals recommended, 13 of whom (31%) had already been included in the initial round of invitations. Twenty-seven of the suggested 29 additional unique individuals were subsequently invited to participate (two of the unique suggestions were made too close to the survey deadline for invitations to be sent). A total of 141 individuals received invitations to complete Round 1 of the Delphi.

Fig. 1.

Fig. 1

Participant recruitment and retention. Only individuals who completed Round 1 were invited to participate in Round 2 and only individuals who completed Round 2 were invited to participate in Round 3

Seventy-three individuals participated in Round 1 of the Delphi (response rate 52%) and are therefore considered members of the panel (Table 2). Most participants identified as a statistician (62, 85%) and/or a clinical trialist (18, 25%) and had prior experience in cluster randomised trials (60, 82%). Half or more of the participants reported at least 10 years of experience in clinical trials (37, 50%) and had been involved in six or more clinical trials (37, 51%). Participants reported residing in North America (33, 45%); Europe (26, 36%); and Australia/Oceania (11, 15%).

Table 2.

Participant characteristics

Characteristic (n (%)) Participants (N = 73)
Job rolea
 Statistician 62 (85)
 Clinical trialist 18 (25)
 Journal editor 6 (8)
 Healthcare professional 4 (5)
 Other 4 (5)
 Health economist 1 (1)
Type of expertiseb
 Cluster randomised trials 60 (82)
 Estimands 32 (44)
 Guideline development 14 (19)
 Other 4 (5)
 Prefer not to say 1 (1)
Number of clinical trials involved in
 0 6 (8)
 1–2 11 (15)
 3–5 17 (23)
 6 or more 37 (50)
 Missing 2 (3)
Number of years of experience in clinical trials
 No experience 4 (5)
 Less than a year 2 (3)
 1–5 years 16 (22)
 6–10 years 12 (16)
 More than 10 years 37 (51)
 Missing 2 (3)
Race
 White 55 (75)
 Prefer not to say 7 (10)
 Asian 6 (8)
 Black 3 (4)
 Missing 2 (3)
Gender
 Man 38 (52)
 Woman 27 (37)
 Prefer not to say 6 (8)
 Missing 2 (23)
Geographic area of residence
 North America 33 (45)
 Europe 26 (36)
 Australia/Oceania 11 (15)
 Prefer not to say 3 (4)

aParticipants could select more than one option. “Other” included: biostatistician/faculty (1); epidemiologist (1); analytically inclined epidemiologist (1); funder (1)

bParticipants could select more than one option. “Other” included: trials methodology (1); clinical trials, crossover trials (1); power and sample size for multilevel and longitudinal data, especially continuous (1); did not specify (1)

Round 1

In Round 1 of the Delphi, participants rated eight items for potential inclusion in an extension of the ICH E9(R1) addendum for cluster randomised trials. The highest-rated items were “weighting of individuals and clusters in the estimand” (mean rating 8.22, standard deviation (SD) 1.36); “whether treatment effects are marginal or cluster-specific” (7.73, 1.55); and “handling of cluster-level intercurrent events” (7.51, 1.69).

The lowest-rated items were “handling of interference or spillover effects” (6.46, 2.18); “exposure time of clusters and individuals to the intervention” (6.77, 2.12); and “handling of individuals who leave or change clusters” (6.80, 1.73) (Table 3). Review of open-ended responses suggested that Delphi participants felt these items were relevant to only a minority of cluster randomised trials that they were already covered within the existing five ICH E9(R1) attributes, or that they were not directly related to estimands (Table S1).

Table 3.

Importance ratings and consensus assessments across rounds. *n = number of ratings provided in a given round (excludes missing data due to non-response and respondents who selected “unable to rate”). n/a, not assessed; SD, standard deviation. Consensus Thresholds: “Essential”: ≥ 70% of participants rating an item 7–9 and < 15% rating 1–3. “Not Essential”: ≥ 70% of participants rating an item 1–3 and < 15% rating 7–9. “No consensus”: any other combination of ratings

Item Round 1 Round 2 Round 3
n* Mean (SD) Not important (1–3)
n (%)
Critical (7–9)
n (%)
Consensus n* Mean (SD) Not important (1–3)
n (%)
Critical (7–9)
n (%)
Consensus n* Mean (SD) Not important (1–3)
n (%)
Critical (7–9)
n (%)
Consensus
1. Weighting of individuals and clusters in the estimand 73 8.22 (1.36) 2 (3) 69 (95) Essential 62 8.53 (0.78) 0 (0) 59 (95) Essential n/a
2. Population of clusters 71 7.25 (1.95) 3 (4) 49 (69) No consensus 62 7.56 (1.54) 0 (0) 46 (74) Essential n/a
3. Population of individuals under selection or recruitment bias 71 7.07 (1.76) 4 (6) 49 (69) No consensus 62 6.66 (1.66) 3 (5) 36 (58) No consensus n/a
4. Exposure time of clusters and individuals to the intervention 73 6.77 (2.12) 9 (12) 50 (68) No consensus 62 7.21 (1.59) 3 (5) 46 (74) Essential n/a
5. Whether treatment effects are marginal or cluster-specific 70 7.73 (1.55) 2 (3) 57 (81) Essential 62 7.95 (1.32) 0 (0) 52 (84) Essential n/a
6. Handling of cluster-level intercurrent events 72 7.51 (1.69) 4 (6) 55 (76) Essential 62 7.97 (1.21) 0 (0) 55 (89) Essential n/a
7. Handling of interference or spillover effects 70 6.46 (2.18) 10 (14) 40 (57) No consensus 61 6.28 (1.76) 4 (7) 35 (57) No consensus n/a
8. Handling of individuals who leave or change clusters 70 6.80 (1.73) 3 (4) 40 (57) No consensus 62 6.84 (1.52) 2 (3) 39 (63) No consensus n/a
9. Handling of clusters that split, merge, or are empty n/a 61 5.84 (1.89) 8 (13) 25 (41) No consensus 59 5.19 (1.41) 3 (5) 11 (19) No consensus

Three items met the threshold for consensus as essential for inclusion after Round 1: “weighting of individuals and clusters in the estimand”; “whether treatment effects are marginal or cluster-specific”; and “handling of cluster-level intercurrent events” (Table 3).

Participants’ open-ended comments were reviewed for suggestions of topics not already included in the initial list of items. From these suggestions, a single new item was proposed: “handling of clusters that split, merge, or are empty” (Table 1).

Round 2

Round 2 was completed by 62 individuals (85% response rate). In Round 2, participants rated all items from Round 1 as well as one newly added item, “handling of clusters that split, merge, or are empty”. Five items officially met consensus as essential after Round 2: “how individuals and clusters are weighted in the estimand”; “population of clusters”; “exposure time of clusters and individuals to the intervention”; “whether treatment effects are marginal or cluster-specific”; and “strategies for handling cluster-level intercurrent events” (Table 3).

The new item rated for the first time in Round 2, “handling of clusters that split, merge, or are empty”, received a mean rating of 5.84 (SD 1.89), with 25 (41%) rating it as “critical” and 8 (13%) rating it as “not important”. In open-ended comments, multiple participants noted that this was an uncommon occurrence (Table S1).

Round 3

Round 3 was completed by 59 individuals (95% response rate). The only item rated in this round was “handling of clusters that split, merge, or are empty”. Eleven (19%) participants rated this item as “critical” and three (5%) rated it as “not important”; there was no consensus as to whether it was essential or not essential to include in the guidance (Table 3).

Discussion

Though precise definitions of estimands are increasingly adopted in randomised trials, concerns have been raised that the five attributes outlined in the ICH E9(R1) addendum may not be sufficient for a clear and comprehensive estimand definition in cluster randomised trials. Based on recent work evaluating the use of estimands in cluster randomised trials, there is a clear need for updated guidance in this area [23]. To fill this gap, we recently convened a group to develop a consensus extension of the ICH E9(R1) addendum for cluster randomised trials.

In a previous scoping review, we identified eight potential additional items that could be used to help define the estimand for cluster randomised trials [24]. In the Delphi survey described in this article, we asked expert stakeholders to rate the importance of each of these proposed items, as well as provide feedback to explain their rating or suggest additional items for consideration in the guidance.

Because one additional item was added during the Delphi, based on participant suggestions, stakeholders rated nine items in total. Of these, five items achieved consensus at the end of the Delphi (how individuals and clusters are weighted; population of clusters; exposure time of individuals and clusters; whether effects are marginal or cluster-specific; and strategies for handling cluster-level intercurrent events). Key themes that emerged from participant feedback on these items were that they were viewed both as applicable to a large percentage of cluster randomised trials and as essential for proper interpretation of trial results.

The remaining four items did not achieve consensus, either for their inclusion or their exclusion. Common reasons given by Delphi participants for low scores were that items were only relevant to a small percentage of cluster randomised trials; that items were already covered by the existing ICH E9(R1) attributes; or that items were not relevant to the estimand and would be better handled elsewhere (e.g. as part of the description of planned statistical methods).

These results may indicate a preference from stakeholders for briefer guidance which contains only the most essential items, rather than a more inclusive approach which expands the number of items included in the guidance. Notably, all items assessed in the Delphi will be carried forward to the consensus meeting, as no items reached consensus as not important. However, the results from this Delphi will be used to inform discussions at this project’s consensus meeting, in which the final guidance will be decided. For instance, item 9 (handling of clusters that split, merge, or are empty) received generally low importance scores, with a mean rating of 5.19 in the final round, where only 19% of respondents rated it as critical. Many respondents felt this item was not sufficiently applicable for inclusion in the guidance, and this view, along with the associated ratings, will form a key discussion point during the consensus meeting.

A notable strength of this study was the high response rate of the Delphi. Seventy-three of 141 invited individuals participated in Round 1. Moreover, retention in subsequent rounds exceeded 80%, indicating a high degree of interest in this project. Furthermore, Delphi participants were on average highly experienced (e.g. the majority had > 10 years’ experience in clinical trials and had been involved in > 5 trials), indicating that the Delphi results can be seen to represent the views of experts in this area. An additional strength is our inclusion of snowball sampling as a method to expand the representativeness of the experts invited to participate. Notably, 30% of those individuals recommended by participants were already included in our invitation lists, suggesting substantial coverage of the experts in this field.

A limitation of this work is that, like with all Delphi surveys, the results may not be generalisable beyond those individuals who participated. Furthermore, a number of the items in the Delphi involve complex or technical concepts—although we are confident in the expertise of the participants included in this Delphi, we did not formally assess for comprehension of each proposed item. Although we took steps to increase clarity—for instance, iterative piloting and updating of definitions amongst the study team, and inclusion of detailed examples illustrating the concepts—we cannot rule out that some participants may have understood items differently from one another.

In conclusion, the results from this Delphi provide strong evidence regarding the importance of multiple key items when defining estimands for cluster randomised trials. These results will be used in the development of consensus guidance outlining the set of attributes that should be described when defining estimands for cluster randomised trials.

Supplementary Information

13063_2025_9352_MOESM1_ESM.docx (150KB, docx)

Additional file 1: ACCORD Guidelines. Checklist for reporting consensus methods. Round 1 Executive Summary. Round 2 Executive Summary. Table S1. Selected comments from executive summaries.

Acknowledgements

We thank the Delphi panellists who participated in this study for their time and perspectives: Agnès Caille; Rebecca Andridge; Laura B. Balzer; Andrew W Brown; Ashley Buchanan; Mike Campbell; Siobhan Creanor; Catherine M. Crespi; Anurika De Silva; Andrew Forbes; Bruno Giraudeau; Deborah H. Glueck; Richard J Hayes; Jennifer Hellier; Karla Hemming; Joanna Hindley; James P. Hughes; Lee Kennedy-Shaffer; Kenneth M. Lee; Clémence Leyrat; David P. MacKinnon; Joanne McKenzie; Lynne Moore; Tim P. Morris; Keith E. Muller; David M. Murray; Stephen Nash; Kate A. Nelson; Joshua R. Nugent; Sherri L. Pals; Pierre Poupin; Joseph S. Ross; Anca Chis Ster; Lehana Thabane; Elizabeth L. Turner; Obioha Ukoumunne; Rebecca Walwyn; Bingkai Wang; Samuel I. Watson; Lisa Yelland.

Authors’ contributions

BCK, MB, DB, FB, GSC, AC, MOH, FL, and CLA conceived of the study. BCK, MB, DB, FB, GSC, AC, MOH, FL, and CLA participated in the design of the study. BCK, MB, and CA conducted data acquisition. BCK, MB, DB, FB, GSC, AC, MOH, FL, and CLA analysed and interpreted the data, revised the manuscript critically for important intellectual content, approved the final manuscript, and agreed to be accountable for its overall content.

Funding

BCK, DB, and AC are funded by the UK Medical Research Council (grant nos. MC_UU_00004/07 and MC_UU_00004/09). CA is supported by a National Institute of Health (NIH), National Heart, Lung, and Blood Institute (NHLBI) grant (K23-HL163402). FL and MOH are supported by a Patient-Centered Outcomes Research Institute Award® (PCORI® Award ME-2022C2- 27676). The statements presented in this article are solely the responsibility of the authors and do not necessarily represent the official views of PCORI®, its Board of Governors, or the Methodology Committee. GSC is a National Institute for Health and Care Research (NIHR) Senior Investigator. The views expressed in this article are those of the author(s) and not necessarily those of the NIHR, or the Department of Health and Social Care.

Data availability

The datasets used and/or analysed during the current study are available from the senior author upon reasonable request (catherine.auriemma@pennmedicine.upenn.edu).

Declarations

Ethics approval and consent to participate

This study was approved as exempt by the University of Pennsylvania IRB (protocol #856,778).

Consent for publication

Not applicable.

Competing interests

The authors declare that they have no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.ICH E9 (R1) addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials https://www.ema.europa.eu/en/documents/scientific-guideline/ich-e9-r1-addendum-estimands-sensitivity-analysis-clinical-trials-guideline-statistical-principles_en.pdf.
  • 2.International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH Guideline Implementation, https://www.ich.org/page/efficacy-guidelines#9-2. Accessed 21/06/2025.
  • 3.Akacha M, Bretz F, Ruberg S. Estimands in clinical trials-broadening the perspective. Stat Med. 2017;36:5–19 2016/07/21. [DOI] [PubMed] [Google Scholar]
  • 4.Akacha M, Frank B, David O, et al. Estimands and their role in clinical trials. Stat Biopharm Res. 2017;9:268–71. [Google Scholar]
  • 5.Kahan BC, Cro S, Li F, Harhay MO. Eliminating ambiguous treatment effects using estimands. Am J Epidemiol. 2023;192:987–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Kahan BC, Hindley J, Edwards M, et al. The estimands framework: a primer on the ICH E9(R1) addendum. BMJ. 2024;384:e076316. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Kahan BC, Morris TP, White IR, et al. Estimands in published protocols of randomised trials: urgent improvement needed. Trials. 2021;22:686 20211009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Cro S, Kahan BC, Rehal S, et al. Evaluating how clear the questions being investigated in randomised trials are: systematic review of estimands. BMJ. 2022;378:e070146 20220823. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Campbell MK, Piaggio G, Elbourne DR, et al. Consort 2010 statement: extension to cluster randomised trials. BMJ. 2012;345:e5661 20120904. [DOI] [PubMed] [Google Scholar]
  • 10.Hayes RJ, Moulton LH. Cluster randomised trials. Boca Raton, FL. Chapman and Hall/CRC, 2017.
  • 11.Murray DM. Design and Analysis of Group-Randomized Trials. Monographs in Epidemiology and Biostatistics. New York, NY. Oxford University Press, 1998.
  • 12.Kahan BC, Blette BS, Harhay MO, et al. Demystifying estimands in cluster-randomised trials. Stat Methods Med Res. 2024;33(7): pp. 1211–1232. 10.1177/09622802241254197 [DOI] [PMC free article] [PubMed]
  • 13.Kahan BC, Li F, Copas AJ, Harhay MO. Estimands in cluster-randomized trials: choosing analyses that answer the right question. Int J Epidemiol. 2023;52(1):107–118 [DOI] [PMC free article] [PubMed]
  • 14.Kenny A, Voldal EC, Xia F, et al. Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect. Stat Med. 2022;41:4311–39 20220630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Lee KM, Cheung YB. Cluster randomized trial designs for modeling time-varying intervention effects. Stat Med. 2024;43:49–60 2023/11/10. [DOI] [PubMed] [Google Scholar]
  • 16.Li F, Tian Z, Bobb J, et al. Clarifying selection bias in cluster randomized trials. Clin Trials. 2021. 10.1177/17407745211056875. [DOI] [PubMed] [Google Scholar]
  • 17.Li F, Tian Z, Tian Z, Li F. A note on identification of causal effects in cluster randomized trials with post-randomization selection bias. Commun Stat Theory Methods. 2022;53(5):1–13
  • 18.Maleyeff L, Li F, Haneuse S, Wang R. Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials. Biometrics. 2023. 10.1111/biom.13803. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.McKenzie JE, Taljaard M, Hemming K, et al. Reporting of cluster randomised crossover trials: extension of the CONSORT 2010 statement with explanation and elaboration. BMJ. 2025;388:e080472. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wang X, Turner EL, Li F, et al. Two weights make a wrong: cluster randomized trials with variable cluster sizes and heterogeneous treatment effects. Contemp Clin Trial. 2022;114:106702 2022/02/06. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.A Benitez ML Petersen MJ Laan van der et al Defining and estimating effects in cluster randomized trials: a methods comparison Stat Med. 2023;42(19):3443–3466. 10.1002/sim.9813 [DOI] [PMC free article] [PubMed]
  • 22.Bugni F, Canay I, Shaikh A and Tabord-Meehan M. Inference for cluster randomized experiments with non-ignorable cluster sizes. J Political Econ Micro. 3:255–88.
  • 23.Bi D, Copas A, Kahan BC. Use of estimands in cluster randomised trials: a review. medRxiv. 2025:2025.2006.2010.25329251. [DOI] [PMC free article] [PubMed]
  • 24.Bi D, Copas A, Li F and Kahan B. A scoping review identified additional considerations for defining estimands in cluster randomised trials. J Clin Epidemiol. 2026;189:112015 [DOI] [PubMed]
  • 25.Barrett D, Heale R. What are Delphi studies? Evid Based Nurs. 2020;23:68–9 2020/05/21. [DOI] [PubMed] [Google Scholar]
  • 26.Fink A, Kosecoff J, Chassin M, Brook RH. Consensus methods: characteristics and guidelines for use. Am J Public Health. 1984;74:979–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Sinha IP, Smyth RL, Williamson PR. Using the Delphi technique to determine which outcomes to measure in clinical trials: recommendations for the future based on a systematic review of existing studies. PLoS Med. 2011;8:e1000393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Auriemma CL, Butt MI, Bahti M, et al. Measuring quality-weighted hospital-free days in acute respiratory failure: a modified Delphi study. Ann Am Thorac Soc. 2024;21:928–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Bahti M, Kahan BC, Li F, et al. Prioritizing attributes of approaches to analyzing patient-centered outcomes that are truncated due to death in critical care clinical trials: a Delphi study. Trials. 2025;26:15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Chan AW, Boutron I, Hopewell S, et al. SPIRIT 2025 statement: updated guideline for protocols of randomised trials. BMJ. 2025;389:e081477. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD). Circulation. 2015;131:211–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Hopewell S, Chan AW, Collins GS, et al. CONSORT 2025 statement: updated guideline for reporting randomised trials. BMJ. 2025;389:e081123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Kahan BC, Hall SS, Beller EM, et al. Reporting of factorial randomized trials: extension of the CONSORT 2010 statement. JAMA. 2023;330:2106–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kahan BC, Hall SS, Beller EM, et al. Consensus statement for protocols of factorial randomized trials: extension of the SPIRIT 2013 statement. JAMA Netw Open. 2023;6:e2346121-e2346121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kahan BC, Juszczak E, Beller E, et al. Guidance for protocol content and reporting of factorial randomised trials: explanation and elaboration of the CONSORT 2010 and SPIRIT 2013 extensions. BMJ. 2025;388(e080785):20250204 20250204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Gattrell WT, Logullo P, van Zuuren EJ, et al. ACCORD (ACcurate COnsensus Reporting Document): a reporting guideline for consensus methods in biomedicine developed via a modified Delphi. PLoS Med. 2024;21:e1004326. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Sadler GR, Lee HC, Lim RS, Fullerton J. Recruitment of hard-to-reach population subgroups via adaptations of the snowball sampling strategy. Nurs Health Sci. 2010;12:369–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Harris PA, Taylor R, Minor BL, et al. The REDCap consortium: building an international community of software platform partners. J Biomed Inform. 2019;95:103208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Harris PA, Taylor R, Thielke R, et al. Research electronic data capture (REDCap)- -a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42:377–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Needham DM, Sepulveda KA, Dinglas VD, et al. Core outcome measures for clinical research in acute respiratory failure survivors. An international modified Delphi consensus study. Am J Respir Crit Care Med. 2017;196:1122–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Turnbull AE, Sepulveda KA, Dinglas VD, et al. Core domains for clinical research in acute respiratory failure survivors: an international modified Delphi consensus study. Crit Care Med. 2017;45:1001–10. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

13063_2025_9352_MOESM1_ESM.docx (150KB, docx)

Additional file 1: ACCORD Guidelines. Checklist for reporting consensus methods. Round 1 Executive Summary. Round 2 Executive Summary. Table S1. Selected comments from executive summaries.

Data Availability Statement

The datasets used and/or analysed during the current study are available from the senior author upon reasonable request (catherine.auriemma@pennmedicine.upenn.edu).


Articles from Trials are provided here courtesy of BMC

RESOURCES