Skip to main content
BMJ Open logoLink to BMJ Open
. 2026 Jun 22;16(6):e114008. doi: 10.1136/bmjopen-2025-114008

Do we consider scalability from the outset? A methodological review of pilot randomised controlled health trials

Matthew Mclaughlin 1,2,3, Kaylie Toll 1,2,4, Bronwyn Myers 2,5, Hugh Riddell 1,2, Joanna Moullin 1,2, Christopher M Reid 1, Eleanor Quested 1,2, Matthew D McDonald 1,2,✉
PMCID: PMC13288692  PMID: 42331572

Abstract

Abstract

Objectives

Health interventions should be designed to be appropriate for scaling from the outset, but the extent to which factors that are important for scalability are reported on in pilot randomised trials is unclear. This review assesses the extent to which pilot randomised trials report on 15 domains of intervention scalability.

Design

Methodological review.

Data sources

Four journals were searched: BMJ Open, BMC Pilot and Feasibility Studies, BMC Trials and PLoS One for articles published between January 2023 and October 2024.

Eligibility criteria

We included pilot randomised trials of health interventions.

Data extraction and synthesis

Data relevant to 15 scalability domains, derived from the Intervention Scalability Assessment Tool and wider implementation science literature, were extracted. Data were double-extracted for 20% of the included studies. Two authors scored all studies from 0 to 3 (0=Not at all; 1=Small extent; 2=Moderate extent; 3=Large extent) on the extent to which each of the 15 scalability domains was reported. For each scalability domain, we calculated the mean score and the frequency of each categorical score across the included studies.

Results

Titles and abstracts screening (521 publications) resulted in 132 full-text publications for review. Of the 104 eligible studies, a random sample of 50 studies were selected for detailed review. Through snowballing, an additional 49 associated publications were identified (eg, protocols), resulting in 99 publications across 50 studies. Most studies reported the Problem (30/50, 60%; mean=2.4 ± 0.8) and the Intervention (37/50, 74%; mean=2.7 ± 0.4) to a large extent (ie, scored 3), with the Delivery Setting and Workforce domain most often receiving a score of 2 (28/50, 56%; mean=2.2 ± 0.7). For eight of the scalability domains, the majority of studies scored 0.

Conclusions

The extent of scalability domain reporting in pilot trials of health interventions is limited. Explicit consideration of scalability in pilot trials could improve the design of fully powered randomised controlled trials and enhance the potential for effective interventions to be translated into practice. Future research should consider if and how to incorporate scalability considerations into pilot trials, and whether pilot trial and intervention reporting guidelines should be expanded to include scalability considerations.

Keywords: Feasibility Studies, Randomized Controlled Trial, Implementation Science


STRENGTHS AND LIMITATIONS OF THIS STUDY.

  • This review provides a methodological snapshot of scalability reporting in recently published pilot randomised trials of health interventions rather than an exhaustive review.

  • Data extraction and scalability scoring protocols were developed a-priori and based on established methods.

  • Two authors independently scored all included studies across scalability domains.

Background

If evidence-based health interventions are not scaled-up, the desired health impact of these interventions will not be achieved. Evidence-based interventions are often initially tested in pilot and feasibility trials (here, pilot trials). The primary objective of pilot trials is to assess the feasibility of conducting a future definitive trial that is powered to assess intervention effectiveness.1 2 Typically, a-priori progression criteria (eg, related to recruitment, intervention acceptability and retention) support researchers to decide if, and how, to proceed to a future definitive trial.3 However, focusing on establishing feasibility for progression to effectiveness trials may inadvertently result in the research and intervention design neglecting to consider factors critical for the future scaling of interventions.

Most public health4 and clinical5 6 interventions are not scaled-up or effectively integrated into practice.7 To address this, researchers and policymakers frequently suggest designing interventions with implementation and scaling in-mind8,12 and thinking politically13—to improve the real-world suitability of interventions.2 For example, designing an intervention with the end in mind is a key component of the Designing for Dissemination and Sustainability logic model.14 Scalability can be defined as ‘the ability of a health intervention shown to be efficacious on a small scale and/or under controlled conditions to be expanded under real-world conditions to reach a greater proportion of the eligible population, while retaining effectiveness’15 (pg.289). Without an explicit focus on scalability from the outset, interventions could be designed and shown to be effective in trials that may not be suitable for delivery outside of trial conditions.

Several scalability domains have been identified based on factors deemed important for scaling health interventions to benefit more people while retaining intervention effectiveness.16 The Intervention Scalability Assessment Tool (ISAT)16 outlines 10 scalability domains including: the problem being addressed; the intervention approach; the strategic and political context; intervention effectiveness; intervention costs and benefits; fidelity and adaptation; acceptability and target population reach; intended delivery setting and workforce; sustainability (long-term effectiveness); and implementation infrustructure.16 This tool is designed to support policymakers and practitioners to make decisions about the suitability of interventions for scale-up after they have been tested. However, if some of these scalability domains are also considered during the design and early stage evaluation of interventions, this may increase the real-world relevance of new health interventions and their potential for scaling.16

Despite the potential value of considering scalability during early stage trials, scalability is not explicit consideration in the Consolidated Standards of Reporting Trials (CONSORT) extension reporting guidelines for pilot and feasibility trials.1 However, some domains within the CONSORT extension for pilot and feasibility trials are also important for scaling interventions into practice (eg, intervention acceptability).12 Equally, some domains that are important for scaling interventions into practice are less relevant for pilot trials. For example, testing intervention effectiveness is not recommended in pilot trials,1 but evidence of effectiveness is critical to inform policymaker and practitioner decisions for scaling interventions to achieve positive health impact.17 18 Overall, little is currently known about the extent to which domains that are important for scaling interventions are reported in pilot trials of health interventions, and thus the extent to which scalability is explicitly considered from the outset in early-stage trials is unclear.

Aim

This methodological review assesses the extent to which recent pilot randomised trials report on intervention scalability across 15 scalability domains.

Methods

Protocol and registration

A protocol for this methodological review is registered on the Open Science Framework.19 This review is reported following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement.20

Eligibility criteria

Inclusion criteria were external randomised controlled pilot and/or feasibility trials of health interventions (including but not limited to pharmaceutical, surgical, public health and allied health interventions), with no restriction on study feasibility objectives or outcomes. Studies published in English between 1 January 2023 and 18 October 2024 were eligible. Exclusion criteria were non-randomised pilot or feasibility studies, protocol papers and pilot and feasibility studies embedded into a larger randomised controlled trial.

Information sources

Drawing on methods employed by Mellor et al,21 four journals were searched through PubMed: British Medical Journal (BMJ) Open, Pilot and Feasibility Studies (PAFS), Trials and Public Library of Science (PLoS) One. These four journals were selected because they have the most PubMed indexed publications that included the terms ‘pilot’ or ‘feasibility’ and ‘trial’ in their title in 2023 and 2024 and because they direct authors to align their reporting against the CONSORT statement.22 BMJ Open and Trials advise authors to use the most appropriate statement extension, and PAFS directs authors to the CONSORT Extension for Pilot and Feasibility Trials.1 The search terms are detailed in online supplemental file 1. Search terms included ‘pilot’ or ‘feasibility’ in the title and ‘trial’ or ‘study’ in the title or abstract.

Snowballing for associated papers

Three methods were used to identify additional publications associated with the included studies, such as publications related to intervention development, trial protocol and additional outcomes. First, the trial registration webpage for the study was searched for associated publications. Second, we hand-searched reference lists of included articles for related papers. Third, for outcomes published subsequently, a forward citation search was conducted. Where two or more publications were found for the same study (eg, a protocol paper and one or more outcome paper), we assessed the extent to which scalability was considered across all associated papers and report these as a single study.

Study selection

The review process, from title and abstract screening through to data extraction, was conducted in Covidence.23 Titles and abstracts of identified publications were screened against inclusion criteria by MMcL. Those that met inclusion criteria proceeded to full text review, conducted by MMcL. Of the studies that met inclusion criteria during full text review, we selected a random sample of 50 studies from which to extract data for detailed review. A researcher (HR) who was not involved in screening, data extraction or analysis, used Microsoft Excel’s random number generator to randomly assign each eligible study a unique number from 1 to 104. Studies numbered 1–50 progressed to data extraction.

Data extraction

Data extraction forms (see online supplemental file 2) produced in Covidence23 were prepiloted on the first 10 trials ordered alphabetically to ensure usability and completeness. One researcher (MMcL) extracted the data for all included publications, with support from a student. Another team member (KT) conducted a second data extraction for the first 20% of studies (n=10 studies). As we found minimal differences between the two data extractions, we decided not to conduct double data extraction for the remaining studies.

We extracted first author last name; study year; study title; study country; study aims and objectives; intervention setting; intervention type; target population; and intervention duration. We also extracted data relating to the 15 scalability domains outlined in table 1. The first 10 scalability domains were derived from the ISAT.16 Scalability domains included in the ISAT were identified via a literature review, expert input and iterative engagement with policymakers and practitioners.16 Application of the ISAT involves reviewing the evidence for each domain and assigning a score that indicates the extent to which that domain is addressed. Scores range from 0 (not at all) to 3 (to a large extent), with higher scores indicating greater readiness for intervention scaling.16 Results are summarised in a radar plot to visually represent relative strengths and domains that may require strengthening for intervention scaling.16

Table 1. Scalability domain descriptions and inclusion in reporting guidelines.

Scalability domain* Description Inclusion in CONSORT-Extension for pilot and feasibility trials1
The Problem A description of the impact of the problem on the health of the population. Aligns to Checklist Item 2a ‘Scientific background and explanation of rationale for future definitive trial, and reasons for randomised pilot trial’
The Intervention A description of the proposed programme/intervention to address the problem described, including the key elements and modes of delivery. Checklist Item 5. ‘The interventions for each group with sufficient details to allow replication, including how and when they were actually administered’
Also mentions ‘The template for intervention description and replication (TIDieR) guidelines should be followed and the checklist completed’.29
Political Context Any reporting or discussion of the intervention’s strategic alignment with national, state or regional policy directions or priorities. Any discussion of how scaling-up the health intervention will be strategically useful to funders. Not explicitly mentioned.
Possibly aligns to Checklist Item 2a ‘Scientific background and explanation of rationale for future definitive trial, and reasons for randomised pilot trial.’
Evidence of effectiveness Any reporting of impact on health outcomes. Methodological consideration and principle 7: ‘Formal hypothesis testing for effectiveness (or efficacy) is not recommended. The aim of a pilot trial is not to assess effectiveness (or efficacy) and it will usually be underpowered to do this’
Costs versus Benefits Any reporting or discussion of costs of development, delivery or maintenance (including income, if applicable) and/or cost-effectiveness. Not explicitly mentioned.
Fidelity and Adaptation Any reporting of adherence to intervention delivery protocol, or adaptations to the intervention that may affect fidelity to the original intervention. Not explicitly mentioned.
Possibly aligns to Checklist Item 3b: ‘Important changes to methods after pilot trial commencement (such as eligibility criteria), with reasons’.
The TIDieR intervention reporting template is recommended to use, which refers to Fidelity and adaptation.29
Acceptability Any reporting of the extent of stakeholder (including participant) satisfaction with the intervention, including discussion of cognitive and emotional responses to the intervention. Mentioned in several examples provided in the guidelines.
Links to Checklist Item 2b ‘Specific objectives or research questions for pilot trial’
Reach Any reporting of extent of intervention uptake by participants, and/or discussion of potential intervention uptake by participants at scale. Checklist Item 7a ‘Rationale for numbers in the pilot trial’ and 13a ‘For each group, the numbers of participants who were approached and/or assessed for eligibility, randomly assigned, received intended treatment, and were assessed for each objective’
Delivery Setting and Workforce Any reporting of who delivered the intervention and/or where it was delivered and/or in what setting. Checklist Item 5: ‘Details should include who administered the treatment, as well as what it comprised and how often and where it was delivered.’
The TIDieR intervention reporting template is recommended to use, which refers to the items ‘Where’ and ‘Who Provided’.29
Implementation Infrastructure Any reporting or discussion of pathways or considerations for real-world delivery/ongoing implementation and resources required, including feasibility of real-world delivery, ‘practicality’, ‘ease of delivery’ and ‘possible to undertake’. Not explicitly mentioned.
Relates to Checklist Item 4C: ‘…it is important to know of any specific aspects that might not be easy to implement in the future definitive RCT.’
Sustainability (long-term effectiveness) Any reporting of longer-term outcomes of health outcomes, beyond the end of the intervention and/or beyond 12 months of intervention (ie, medium-to-long term).
Not explicitly mentioned.
Implementation Barriers and Enablers Any reporting of barriers and facilitators to implementation or scalability. Not explicitly mentioned.
End-user Engagement Any reporting or discussion of engagement of end-users in the development process of the intervention and/or implementation support strategies (eg, co-design, advisory groups). Not explicitly mentioned.
Stakeholder partnerships (excluding end-users) Any reporting or discussion of project stakeholders or partners engaged, including discussions and engagement with stakeholders, through to listed in-kind and cash partners. Not explicitly mentioned.
Recruitment strategies (real-world) Any reporting or discussion of the applicability of the intervention recruitment methods for use in the real world. Recruitment mentioned throughout, but in relation to future trial and no explicit mention of real world.
*

Definitions of each scalability domain adapted from the ISAT16 and outcomes for implementation research.24

CONSORT, Consolidated Standards of Reporting Trials; ISAT, Intervention Scalability Assessment Tool; RCT, randomised controlled trial.

In this review, we split the ISAT domain of reach and acceptability into two separate domains. Reach and acceptability are distinct concepts,24 with acceptability more akin to satisfaction and reach relating to uptake of the intervention and population coverage. Further, we included four additional domains, which according to the implementation science literature are important for scaling interventions: understanding implementation barriers and enablers,25 end-user engagement (eg, co-design),26 stakeholder partnerships27 and recruitment strategies relevant in the real world.28

Scoring of scalability domains

Using an approach adapted from the ISAT,16 we rated the extent to which studies reported on each scalability domain on a scale from 0 to 3 (0=Not at all; 1=To a small extent; 2=To a moderate extent; 3=To a large extent). Studies were scored on all 15 domains (acknowledging that some domains such as effectiveness and sustainability may be less relevant to pilot randomised trials). Data scoring was completed independently by two authors (MMcL and KT). An operational definition for each domain was generated (see online supplemental file 3) to provide guidance and ensure consistency of scoring. Discrepancies of one or more points between scorers (eg, KT scored 1 and MMcL scored 2) were discussed between MMcL and KT to reach consensus. If consensus was not achieved, a third author (MDMcD) provided a score. Then, the final score was derived by consensus discussion between three authors.

Synthesis of results

Descriptive statistics (frequencies and mean, median and IQR for trial sample sizes) were produced to describe trial characteristics and scores for each scalability domain. Data were analysed in Microsoft Excel. As scores were not normally distributed, categorical frequency data for each score (ie, 0, 1, 2, 3) per domain across included studies were summarised and a stacked bar chart generated showing the percentage distribution of scores across scalability domains visually. To align with the instructions for using the ISAT tool16 mean score per scalability domain across all studies was also calculated and a radar plot with visual representation of reporting across scalability domains produced. We also conducted a narrative synthesis of the findings for each domain.

Patient and public involvement

Patients or the public were not involved in the design, conduct, reporting or dissemination plans of this methods-focused review.

Results

Study selection

Our search strategy identified 521 publications. We screened their titles and abstracts, then assessed the full texts of 132 publications for eligibility, resulting in 104 eligible studies. A random sample of 50 (of the 104) studies were selected for inclusion. From these 50 studies, an additional 49 related publications were identified, resulting in 99 publications for inclusion across 50 studies. Online supplemental file 4 provides a full list of included publications and their corresponding studies. Figure 1 shows the PRISMA flow chart of publications included and excluded at each stage. The main reason for exclusion at full text screening was no control group (n=16).

Figure 1. PRISMA flow chart. PRISMA, Preferred Reporting Items for Systematic Reviews and Meta-Analyses; RCT, randomised controlled trial.

Figure 1

Study characteristics

Table 2 summarises the characteristics of the included studies. Most studies were two-arm (47/50, 94%), non-industry-funded (36/50, 72%) and the most common intervention types were ‘public health’ (15/50, 30%).

Table 2. Characteristics of the included randomised pilot trials.

N (%)
Journal
 British Medical Journal Open 11 (22)
 Pilot and Feasibility Studies 18 (36)
 Trials 0 (0)
 PLoS One 21 (42)
Country
 Australia 5 (10)
 Canada 4 (8)
 UK 13 (26)
 USA 8 (16)
 Other 20 (40)
Funder
 Industry 0 (0)
 Non-industry 36 (72)
 A combination 4 (8)
 Trial did not receive funding 9 (18)
 Unknown 1 (2)
Intervention type
 Surgical 0 (0)
 Pharmaceutical 5 (10)
 Public health 15 (30)
 Mental health and psychological well-being 13 (26)
 Chronic disease management 8 (16)
 Allied health and rehabilitation 9 (18)
Sample size
 Median (IQR) 48 (33–76)
 Range (min–max) 610 (10–620)
Number of arms
 2 47 (94)
 >2 3 (6)
Intention to proceed to a full trial
 Yes (including with adaptations) 40 (80)
 No (do not proceed) 1 (2)
 Unsure (unclear) 9 (18)

Percentages may not sum up to 100 due to rounding.

Scalability domain scores across studies

Table 3 summarises the extent to which included studies reported on each of the scalability domains. The majority of studies received a score of 3 for the Problem (30/50; 60%, mean=2.4 ± 0.8) and the Intervention (37/50, 74%; mean=2.7 ± 0.4), indicating that information relating to these domains was reported to a large extent. Delivery Setting and Workforce was the next highest scoring domain with most studies receiving a score of 2 (28/50, 56%; mean=2.2 ± 0.7). Four domains scored an average of more than one but less than two. These were: Effectiveness; Fidelity and Adaptation; Acceptability; and Reach. The scores for the effectiveness domain reflect pilot trial recommendations not to undertake effectiveness testing and our scoring method with no trials reporting powered inferential statistics on health outcomes (ie, scoring 3). All included studies considered reach to a small extent (eg, how many people received the intervention) and scored 1, but it was usually in the context of trial recruitment rather than in the context of potential intervention reach if implemented and scaled in usual care settings.

Table 3. Extent of reporting of scalability domains across studies, in randomised pilot trials.

Reporting of scalability domains
Scalability domain Score=0
n (%)
Score=1
n (%)
Score=2
n (%)
Score=3
n (%)
Mean (SD) score if reported
(ie, 1–3)
Mean (SD) overall score
(ie, 0–3)
The Problem 0 (0) 8 (16) 12 (24) 30 (60) 2.4 (0.8) 2.4 (0.8)
The Intervention 0 (0) 0 (0) 13 (26) 37 (74) 2.7 (0.4) 2.7 (0.4)
Political Context 35 (70) 11 (22) 4 (8) 0 (0) 1.3 (0.4) 0.4 (0.6)
Effectiveness 7 (14) 14 (28) 29 (58) 0 (0) 1.7 (0.5) 1.4 (0.7)
Costs versus Benefits 35 (70) 9 (18) 2 (4) 4 (8) 1.7 (0.9) 0.5 (0.9)
Fidelity and Adaptation 21 (42) 15 (30) 8 (16) 6 (12) 1.7 (0.8) 1.0 (1.0)
Acceptability 18 (36) 6 (12) 15 (30) 11 (22) 2.2 (0.7) 1.4 (1.2)
Reach 0 (0) 47 (94) 3 (6) 0 (0) 1.1 (0.2) 1.1 (0.2)
Delivery Setting and Workforce 1 (2) 5 (10) 28 (56) 16 (32) 2.2 (0.6) 2.2 (0.7)
Implementation Infrastructure 37 (74) 12 (24) 1 (2) 0 (0) 1.1 (0.3) 0.3 (0.5)
Sustainability 47 (94) 2 (4) 1 (2) 0 (0) 1.3 (0.5) 0.1 (0.3)
Implementation Barriers and Enablers 32 (64) 7 (14) 8 (16) 3 (6) 1.8 (0.7) 0.6 (1.0)
End-user Engagement 29 (58) 9 (18) 8 (16) 4 (8) 1.8 (0.7) 0.7 (1.0)
Stakeholder Partnerships 42 (84) 7 (14) 1 (2) 0 (0) 1.1 (0.3) 0.2 (0.4)
Recruitment Strategies (Real World) 41 (82) 6 (12) 3 (6) 0 (0) 1.3 (0.5) 0.2 (0.5)

Percentages may not sum up to 100 due to rounding.

Figure 2 visually depicts the percentage distribution of scores across scalability domains. The majority of the included pilot trials did not report information (ie, scored 0) for eight scalability domains (Political Context; Costs versus Benefits; Implementation Infrastructure; Sustainability; Implementation Barriers and Enablers; End-user Engagement; Stakeholder Partnerships; and Recruitment Strategies (Real World)). For example, most studies (35/50; 70%) did not report information relating to Political Context, including alignment of the health intervention with government policy directions, priority health areas or broader strategic context. Although reporting intervention cost-effectiveness is beyond the scope of pilot trials, the majority did not report any information related to the cost to develop or deliver the intervention (35/50; 70%). Most did not report information related to End-user Engagement (29/50; 58%), Implementation Infrastructure (37/50; 74%) or Stakeholder Partnerships (42/50; 84%). The low score for the scalability domain Sustainability (long-term effectiveness) reflects that this is beyond the scope of pilot trials.

Figure 2. Distribution of scores across 15 scalability domains.

Figure 2

Figure 3 presents a radar plot of mean scalability domain scores across included studies. The plot visually depicts variation in the extent of reporting across domains, with higher scores (>2) observed for the Intervention, the Problem, and Delivery Setting and Workforce, and lower scores (<1) for domains such as Stakeholder Partnerships, Recruitment Strategies (Real World), Political Context and Costs versus Benefits.

Figure 3. Radar plot showing 15 scalability domains and mean overall scores for extent of reporting across included studies.

Figure 3

Discussion

Summary of main findings

Our study provides the first assessment of the extent to which pilot and feasibility randomised controlled trials of health interventions report on key domains that are important for intervention scalability. Findings from this review suggest that intervention scalability domains are not well reported in pilot and feasibility trials, with included studies largely limiting their consideration of intervention scalability to three of the 15 domains, namely: the Intervention, the Problem, and Delivery Setting and Workforce. These findings suggest that some domains that are important for the future scalability of health interventions could be more explicitly considered from the outset in pilot trials and reported in associated publications.

Comparison to other studies

The goal of pilot and feasibility trials is to assess the potential for progression to a definitive full trial,1 whereas scalability domains are factors important for scaling health interventions in real world contexts to benefit more people while retaining intervention effects.12 16 There is overlap between some scalability domains16 and what is recommended in CONSORT-pilot trial reporting guidelines.1 The CONSORT checklist of information to include when reporting a pilot trial1 includes describing the problem that the trial proposes to address in the background; describing the intervention in accordance with the Template for intervention description and replication (TIDieR)29; and providing details of who administered the intervention and where it was delivered. Respectively, these checklist items recommend pilot trials report information related to the scalability domains the Problem, the Intervention, and the Delivery Workforce and Setting. Correspondingly, the pilot trials included in our analysis largely reported these scalability domains comprehensively.

Recent research by Pfledderer et al30 developed a consolidated set of 20 considerations for the design, conduct, implementation and reporting of pilot and feasibility studies (all study design types) of behavioural interventions. Some of the scalability domains considered in the current review strongly align with these considerations including ‘Stakeholder engagement and co-production’, ‘Adaptations and tailoring’, ‘Fidelity’ and ‘Cost and Resources’.30 In our analysis of pilot trials, reporting of the linked scalability domains for Stakeholder Partnerships, End-user Engagement and Costs and Benefits was limited. Two (of six) core elements that underpin each phase outlined in the Medical Research Council framework for developing and evaluating complex interventions are ‘stakeholder engagement’ and ‘economic considerations’.31 This suggests that while these domains are critical for developing, evaluating and scaling interventions, they are not well reported from the outset in pilot trials.

Our analysis showed limited reporting of concepts related to Fidelity and Adaptation. Recognising the extent to which interventions are delivered as intended (ie, fidelity), and if any changes are made to the intervention (ie, adaptation and modifications), is important to understand what was delivered.25 There can be tensions between traditional trial approaches that emphasise controlled settings for intervention delivery, and less controlled research designs that embrace evaluating interventions delivered in diverse real-world contexts. Adaptations during randomised controlled trials are common32 and can be considerable, so understanding the extent, type and reasons for adaptation can inform scale-up.33 These tensions mirror broader challenges in bridging gaps between evidence generation and real-world implementation.

A recent JAMA Special Issue highlighted that the two limbs of knowledge generation and implementation largely operate as a ‘house divided’, and that better integration between clinical trials and practice is needed to accelerate improvements in the delivery of health interventions.6 Our findings reflect this disconnect, with the Implementation Infrastructure scalability domain given little consideration in the reporting of pilot trials. However, even where implementation infrastructure is robust, and evidence for intervention effects are well-established in later stage research, the decision to adopt or scale-up an intervention is largely determined by the health policy process and political context.8 11 For example, there has been a proliferation of research trials of physical activity interventions,34 but many fail to articulate the political context13 (eg, congruence with health policy priorities). The limited (reported) considerations by researchers positioning the intervention in political context could be a factor in why so few of these interventions are translated into practice.35 Most pilot trials in our analysis did not report information related to Political Context.

Taken together, our findings show that reporting of scalability domains in pilot trials is mixed but overall limited to three domains that correspond with information captured in the CONSORT extension for Pilot and Feasibility Trials. Implementation science has grown as a discipline to address failures in the translation of effective health interventions in a systematic and coordinated manner25 36 and funders internationally have increasingly mandated explicit translation and impact planning from the outset.37 However, findings from our review suggest that such principles have not been explicitly integrated into pilot trial reporting guidelines. As articulated in the recent elaboration and explanation publication for the CONSORT 2025 guideline update for reporting randomised trials,38 reporting issues is one aspect of research waste that is relatively easy to fix, and updating reporting guidelines periodically is essential to ensure their contemporary value and usefulness. To ensure that the potential for future scaling and translation of health interventions into practice is explicitly and systematically considered in the development, conduct and reporting of pilot trials, we suggest that the CONSORT extension for Pilot and Feasibility Trial reporting guidelines is updated.

The anticipated future costs of interventions in the real-world, stakeholder engagement, political context and the implementation infrastructure are some examples of scalability domains that could be reported from the outset in pilot trials of health interventions. However, the relevance and relative importance of specific scalability domains may differ by intervention stage, trial design and research field, and views on their incorporation into pilot trial reporting guidelines are likely to vary across methodological specialisations. Thus, we are unable to make strong recommendations on which specific domains could or should be reported, and in what ways. Any specific recommendations should be developed through international and interdisciplinary consensus and align with best practice for the development of reporting guidelines.39

A second suggestion is for explicit consideration of scalability within recommendations for progression criteria in randomised pilot trials.3 This would ensure that progression decisions are informed not only by the feasibility of conducting a full trial to test intervention effectiveness, but also by the wider context surrounding the intervention’s potential for future translation into practice. A final suggestion is to include explicit consideration of scalability within the TIDieR intervention reporting template.29 While the primary purpose of TIDieR is to support the completeness and clarity of intervention reporting for replication, we suggest that explicitly capturing key scalability considerations within the template would support the systematic consideration of scalability for interventions spanning pilot trial to later stage research. Ultimately, such changes would help support researchers to plan for scalability from the outset and position interventions optimally for future health impact, as recommended by implementation researchers.8,11

Strengths and limitations of the study

The scoring of scalability domains was underpinned by the implementation science literature on domains important for scalability25 27 28 including the ISAT tool.16 The ISAT tool has previously been used to assess the extent to which scalability is considered in publications related to certain types of interventions40,43 and as a reflective tool for researchers to assess the potential scalability of their own interventions.44 However, no reviews have explored the extent of consideration across scalability domains for pilot and feasibility trials.

An exhaustive review was not feasible due to the high number of published pilot trials—the CONSORT extension for pilot and feasibility studies has been cited over 4000 times.1 In line with guidance for methodological reviews,45 our approach to selecting studies for inclusion balanced decisions about feasibility, analytical depth and the usefulness of information obtained. Meaningful conclusions have been drawn from reviews with smaller sample sizes that assess scalability across interventions.43 46 A strength of this review is that information was extracted from all available linked published studies of the included pilot trials, including intervention development and protocol papers, as well as outcomes publications. This approach allowed for a comprehensive assessment of whether the included studies reported scalability.

This study investigated external randomised pilot trial publications; therefore, it is unclear if the findings can be generalised to other feasibility study designs such as non-randomised pilot trials. The scoring of the extent of scalability domain reporting was informed by clear definitions with example scoring interpretations, developed a priori for each scalability domain. However, inevitably the process involves subjectivity, hence why two researchers independently scored 100% of papers. We only included recent pilot trials published since January 2023. While this is a strength, it may also be a limitation that some associated papers (eg, subsequent process evaluations) that may further increase the scores for each scalability domain assessed, may not yet be published. Finally, we also acknowledge that the inclusion of all types of health interventions led to heterogeneity of interventions. It is possible that the reporting of some scalability domains may vary by intervention type.

Conclusion

Despite calls for interventions to be designed for scale from the outset, our findings highlight that the extent to which scalability is reported in recent pilot trials is limited. One way to help ensure that the potential for the future scaling and translation of health interventions into practice is explicitly and systematically considered in the development, conduct and reporting of pilot trials, is to update the CONSORT extension for Pilot and Feasibility Trial reporting guidelines. Updated guidance to this effect may benefit researchers, peer reviewers, journal editors and funders of external randomised pilot trials, and inform the design of subsequent definitive trials and eventual real-world delivery of interventions found to be effective. Future research should consider what scalability domains are most important and relevant for pilot trials to report, and how pilot trial or intervention reporting guidelines could be expanded to include scalability considerations. Any specific recommendations should be developed with international and interdisciplinary consensus, reflecting a diversity of expertise, contexts, research agendas and intervention types. In the meantime, researchers conducting pilot and feasibility trials may increase the real-world relevance of their studies by considering their specific intervention context and the potential relevance of reporting information relating to specific scalability domains such as political context, stakeholder partnerships, end-user engagement and potential intervention costs.

Supplementary material

online supplemental file 1
bmjopen-16-6-s001.docx (15.4KB, docx)
DOI: 10.1136/bmjopen-2025-114008
online supplemental file 2
bmjopen-16-6-s002.docx (24.1KB, docx)
DOI: 10.1136/bmjopen-2025-114008
online supplemental file 3
bmjopen-16-6-s003.docx (27KB, docx)
DOI: 10.1136/bmjopen-2025-114008
online supplemental file 4
bmjopen-16-6-s004.docx (139.6KB, docx)
DOI: 10.1136/bmjopen-2025-114008

Acknowledgements

We would like to thank Laura Taylor, an undergraduate student on placement with MDMcD, for her support with data management.

Footnotes

Funding: This work was supported by Curtin University School of Population Health Incentive Scheme—Interdisciplinary Projects 2024 (no grant number, unencumbered grant). The funder did not influence the outcomes of the study despite author affiliations with the funder.

Prepub: Prepublication history and additional supplemental material for this paper are available online. To view these files, please visit the journal online (https://doi.org/10.1136/bmjopen-2025-114008).

Provenance and peer review: Not commissioned; externally peer reviewed.

Patient consent for publication: Not applicable.

Ethics approval: Not applicable.

Data availability free text: In accordance with best practices for transparency and reproducibility, all data and materials related to this methodological review will be made available upon reasonable request. This includes the full dataset, analysis code and any supplementary materials used in the review. Requests for data sharing should be directed to the corresponding author.

Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.

Data availability statement

Data are available upon reasonable request.

References

  • 1.Eldridge SM, Chan CL, Campbell MJ, et al. CONSORT 2010 statement: extension to randomised pilot and feasibility trials. BMJ. 2016;355:i5239. doi: 10.1136/bmj.i5239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Indig D, Lee K, Grunseit A, et al. Pathways for scaling up public health interventions. BMC Public Health. 2017;18:68. doi: 10.1186/s12889-017-4572-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Mellor K, Albury C, Dutton SJ, et al. Recommendations for progression criteria during external randomised pilot trial design, conduct, analysis and reporting. Pilot Feasibility Stud. 2023;9:59. doi: 10.1186/s40814-023-01291-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wolfenden L, Milat AJ, Lecathelinais C, et al. A bibliographic review of public health dissemination and implementation research output and citation rates. Prev Med Rep. 2016;4:441–3. doi: 10.1016/j.pmedr.2016.08.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Curfman G. Integrating Clinical Trials and Practice: A New JAMA Series and Call for Papers. JAMA. 2024;332 doi: 10.1001/jama.2024.10266. [DOI] [Google Scholar]
  • 6.Angus DC, Huang AJ, Lewis RJ, et al. The Integration of Clinical Trials With the Practice of Medicine: Repairing a House Divided. JAMA. 2024;332:153–62. doi: 10.1001/jama.2024.4088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sutherland RL, Jackson JK, Lane C, et al. A systematic review of adaptations and effectiveness of scaled-up nutrition interventions. Nutr Rev. 2022;80:962–79. doi: 10.1093/nutrit/nuab096. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lee K, van Nassau F, Grunseit A, et al. Scaling up population health interventions from decision to sustainability - a window of opportunity? A qualitative view from policy-makers. Health Res Policy Syst. 2020;18:118. doi: 10.1186/s12961-020-00636-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Koorts H, Ma J, Swain CTV, et al. Systems approaches to scaling up: a systematic review and narrative synthesis of evidence for physical activity and other behavioural non-communicable disease risk factors. Int J Behav Nutr Phys Act. 2024;21:32. doi: 10.1186/s12966-024-01579-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.McKay HA, Kennedy SG, Macdonald HM, et al. The Secret Sauce? Taking the Mystery Out of Scaling-Up School-Based Physical Activity Interventions. J Phys Act Health. 2024;21:731–40. doi: 10.1123/jpah.2024-0274. [DOI] [PubMed] [Google Scholar]
  • 11.Murphy J, Mansergh F, O’Donoghue G, et al. Factors related to the implementation and scale-up of physical activity interventions in Ireland: a qualitative study with policy makers, funders, researchers and practitioners. Int J Behav Nutr Phys Act. 2023;20:16. doi: 10.1186/s12966-023-01413-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Klaic M, Kapp S, Hudson P, et al. Implementability of healthcare interventions: an overview of reviews and development of a conceptual framework. Implement Sci. 2022;17:10. doi: 10.1186/s13012-021-01171-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Mclaughlin M, McCue P, Swelam B, et al. Physical activity-the past, present and potential future: a state-of-the-art review. Health Promot Int. 2025;40 doi: 10.1093/heapro/daae175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Kwan BM, Luke DA, Adsul P, et al. Dissemination and implementation research in health: Translating science to practice. Oxford University Press; Designing for dissemination and sustainability: principles, methods, and frameworks for ensuring fit to context. Available. [DOI] [Google Scholar]
  • 15.Milat AJ, King L, Bauman AE, et al. The concept of scalability: increasing the scale and potential adoption of health promotion interventions into policy and practice. Health Promot Int. 2013;28:285–98. doi: 10.1093/heapro/dar097. [DOI] [PubMed] [Google Scholar]
  • 16.Milat A, Lee K, Conte K, et al. Intervention Scalability Assessment Tool: A decision support tool for health policy makers and implementers. Health Res Policy Syst. 2020;18:1. doi: 10.1186/s12961-019-0494-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wolfenden L, Hall A, Bauman A, et al. Research outcomes informing the selection of public health interventions and strategies to implement them: A cross-sectional survey of Australian policy-maker and practitioner preferences. Health Res Policy Syst. 2024;22:58. doi: 10.1186/s12961-024-01144-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.World Health Organization Practical guidance for scaling up health service innovations. 2009.
  • 19.Mclaughlin M, McDonald M, Riddell H, et al. Protocol: do we consider scalability from the outset?: a methodological review of pilot and feasibility studies of health interventions. 2024. Available. [DOI] [PMC free article] [PubMed]
  • 20.Moher D, Liberati A, Tetzlaff J, et al. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. Ann Intern Med. 2009;151:264–9. doi: 10.7326/0003-4819-151-4-200908180-00135. [DOI] [PubMed] [Google Scholar]
  • 21.Mellor K, Eddy S, Peckham N, et al. Progression from external pilot to definitive randomised controlled trial: a methodological review of progression criteria reporting. BMJ Open. 2021;11 doi: 10.1136/bmjopen-2020-048178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Schulz KF, Altman DG, Moher D, et al. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials. BMC Med. 2010;8:18. doi: 10.1186/1741-7015-8-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Covidence Systematic review tool. veritas health innovation. 2024. https://www.covidence.org Available.
  • 24.Proctor E, Silmere H, Raghavan R, et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm Policy Ment Health. 2011;38:65–76. doi: 10.1007/s10488-010-0319-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.McDonald M, McLaughlin M, Moullin JC. The routledge international handbook of health psychology. Routledge; 2025. Implementation strategies from lab to field.https://www.taylorfrancis.com/reader/download/747a468f-b07c-474a-aab9-7fa7365465f1/chapter/pdf?context=ubx Available. [Google Scholar]
  • 26.Greenhalgh T, Hinton L, Finlay T, et al. Frameworks for supporting patient and public involvement in research: Systematic review and co-design pilot. Health Expect. 2019;22:785–801. doi: 10.1111/hex.12888. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Indig D, Grunseit A, Greig A, et al. Development of a tool for the evaluation of obesity prevention partnerships. Health Promot J Austr. 2019;30:18–27. doi: 10.1002/hpja.10. [DOI] [PubMed] [Google Scholar]
  • 28.Andrade C. Real World Studies: What They Are and What They Are Not. Indian J Psychol Med. 2023;45:537–8. doi: 10.1177/02537176231188563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Hoffmann TC, Glasziou PP, Boutron I, et al. Better reporting of interventions: template for intervention description and replication (TIDieR) checklist and guide. BMJ. 2014;348 doi: 10.1136/bmj.g1687. [DOI] [PubMed] [Google Scholar]
  • 30.Pfledderer CD, von Klinggraeff L, Burkart S, et al. Consolidated guidance for behavioral intervention pilot and feasibility studies. Pilot Feasibility Stud. 2024;10:57. doi: 10.1186/s40814-024-01485-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Skivington K, Matthews L, Simpson SA, et al. A new framework for developing and evaluating complex interventions: update of Medical Research Council guidance. BMJ. 2021;374:n2061. doi: 10.1136/bmj.n2061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Mackie TI, Ramella L, Schaefer AJ, et al. Multi-method process maps: An interdisciplinary approach to investigate ad hoc modifications in protocol-driven interventions. J Clin Transl Sci. 2020;4:260–9. doi: 10.1017/cts.2020.14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Mclaughlin M, Campbell E, Sutherland R, et al. Extent, Type and Reasons for Adaptation and Modification When Scaling-Up an Effective Physical Activity Program: Physical Activity 4 Everyone (PA4E1) Front Health Serv. 2021;1 doi: 10.3389/frhs.2021.719194. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Bauman A, Lee KC, Pratt M. Understanding the Increases in Physical Activity Publications From 1985 to 2022: A Global Perspective. J Phys Act Health. 2025;22:175–81. doi: 10.1123/jpah.2024-0050. [DOI] [PubMed] [Google Scholar]
  • 35.Pratt M, Varela AR, Bauman A. The Physical Activity Policy to Practice Disconnect. J Phys Act Health. 2023;20:461–4. doi: 10.1123/jpah.2023-0071. [DOI] [PubMed] [Google Scholar]
  • 36.Grimshaw JM, Eccles MP, Lavis JN, et al. Knowledge translation of research findings. Implement Sci. 2012;7:50. doi: 10.1186/1748-5908-7-50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.McLean RKD, Graham ID, Tetroe JM, et al. Translating research into action: an international study of the role of research funders. Health Res Policy Syst. 2018;16:44. doi: 10.1186/s12961-018-0316-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Hopewell S, Chan A-W, Collins GS, et al. CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials. BMJ. 2025;389:e081124. doi: 10.1136/bmj-2024-081124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Moher D, Schulz KF, Simera I, et al. Guidance for developers of health research reporting guidelines. PLoS Med. 2010;7:e1000217. doi: 10.1371/journal.pmed.1000217. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Grady A, Jackson J, Wolfenden L, et al. Assessing the scalability of evidence-based healthy eating and physical activity interventions in early childhood education and care: A cross-sectional study of end-user perspectives. Aust N Z J Public Health. 2024;48:100122. doi: 10.1016/j.anzjph.2023.100122. [DOI] [PubMed] [Google Scholar]
  • 41.Lee K, Milat A, Grunseit A, et al. The Intervention Scalability Assessment Tool: a pilot study assessing five interventions for scalability. Public Health Res Pract. 2020;30:3022011. doi: 10.17061/phrp3022011. [DOI] [PubMed] [Google Scholar]
  • 42.Osinaike J, Myers A, Lowe A, et al. Implementation and Scalability of Physical Activity Interventions Delivered Within Primary Care: A Narrative Review. Lifestyle Medicine. 2024;5:e113. doi: 10.1002/lim2.113. [DOI] [Google Scholar]
  • 43.Lum M, Turon H, Keenan S, et al. A rapid review describing the scalability of early childhood education and care-based programs targeting children’s social and emotional learning. Ment Health Prev. 2024;35:200349. doi: 10.1016/j.mhp.2024.200349. [DOI] [Google Scholar]
  • 44.Calnan S, Lee K, McHugh S. Assessing the scalability of an integrated falls prevention service for community-dwelling older people: a mixed methods study. BMC Geriatr. 2022;22:17. doi: 10.1186/s12877-021-02717-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Aguinis H, Ramani RS, Alabduljader N. Best-Practice Recommendations for Producers, Evaluators, and Users of Methodological Literature Reviews. Organ Res Methods. 2023;26:46–76. doi: 10.1177/1094428120943281. [DOI] [Google Scholar]
  • 46.Grady A, Jackson J, Wolfenden L, et al. Assessing the scalability of healthy eating interventions within the early childhood education and care setting: secondary analysis of a Cochrane systematic review. Public Health Nutr. 2023;26:3211–29. doi: 10.1017/S1368980023002550. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    online supplemental file 1
    bmjopen-16-6-s001.docx (15.4KB, docx)
    DOI: 10.1136/bmjopen-2025-114008
    online supplemental file 2
    bmjopen-16-6-s002.docx (24.1KB, docx)
    DOI: 10.1136/bmjopen-2025-114008
    online supplemental file 3
    bmjopen-16-6-s003.docx (27KB, docx)
    DOI: 10.1136/bmjopen-2025-114008
    online supplemental file 4
    bmjopen-16-6-s004.docx (139.6KB, docx)
    DOI: 10.1136/bmjopen-2025-114008

    Data Availability Statement

    Data are available upon reasonable request.


    Articles from BMJ Open are provided here courtesy of BMJ Publishing Group

    RESOURCES