Skip to main content
JAMA Network logoLink to JAMA Network
. 2026 Jul 28;9(7):e2625935. doi: 10.1001/jamanetworkopen.2026.25935

Randomized Clinical Trials Using a Hierarchical Composite Primary End Point

A Scoping Review

Selina Ehrenzeller 1, Amos J de Jong 2, Yonas Martin 3, Marlieke E A de Kraker 4,5, Holly Jackson 4,5, Nasreen Hassoun-Kheir 4,5, Michaela Schumacher 6, Nadine Saskia Bernasconi 6, Marjolein P M Hensgens 2,7, Christian Appenzeller-Herzog 8, Natalie Rose 6,9, Steven Y C Tong 10,11, Nina Khanna 6, Richard Kuehl 6,12, Matthias Briel 1, Sean W X Ong 10,11,13,14, Benjamin Speich 1,✉
PMCID: PMC13416905  PMID: 42518231

Key Points

Question

How are hierarchical composite end points (HCEs) currently being used and analyzed as primary end points in randomized clinical trials across medical disciplines?

Findings

This scoping review of 92 randomized clinical trials (with 79 567 planned participants) using an HCE as a primary end point showed substantial heterogeneity and shortcomings in the planning and analysis of HCEs, as well as the reporting of the hierarchy construction, sample size assumptions, and handling of ties.

Meaning

These findings suggest that the current use of HCEs in randomized clinical trials lacks methodological consistency and transparent reporting, limiting interpretability and comparability across studies and underscoring the need for clear methodological and reporting guidelines.


This scoping review of randomized clinical trials using hierarchical composite end points examines characteristics of their use, including how they are constructed, analyzed, and reported.

Abstract

Importance

Hierarchical composite end points (HCEs) are a promising tool used in randomized clinical trials (RCTs) to integrate multiple outcomes of varying clinical relevance into a single measure.

Objective

To describe how often HCEs are used as primary outcomes in RCTs, and how they are constructed, analyzed, and reported.

Evidence Review

MEDLINE, Embase, CENTRAL, and Web of Science were searched on December 9, 2024, complemented by a forward citation search of methodological papers on HCEs. RCTs that used an HCE as their primary outcome were included, defined either by self-declaration (hierarchical composite or outcome ranking) or through use of an HCE-specific analytical approach (win ratio, win odds, probabilistic index, or generalized pairwise comparison). Pilot studies, post hoc analyses, and hierarchically tested coprimary end points were excluded. Data were independently screened and extracted in duplicate. Trial characteristics, end point composition, hierarchy justification, and analytical methods were summarized descriptively.

Findings

Among 5188 screened records, 92 RCTs were included, with 79 567 planned participants. The use of an HCE as a primary end point has increased, with 72 trials (78.3%) initiating recruitment within the past decade. Most RCTs were drug trials (43.5% [40 of 92]), in cardiology (43.5% [40 of 92]), multicenter (91.3% [84 of 92]), and non–industry sponsored (67.4% [62 of 92]). The 92 HCEs had a median (IQR) of 4 (3-5) components and the highest ranked component was usually mortality (80.4% [74 of 92]). The last hierarchical component was most often a continuous component (64.1% [59 of 92]). The majority of RCTs did not report any information on how the hierarchy was established (82.2% [60 of 73]; excluding RCTs where only information from registries was available). Generalized pairwise comparison was the most frequent analysis approach (57.3% [26 of 45]) among the 45 published RCTs, yet no standardized way for presenting results was observed.

Conclusions and Relevance

This scoping systematic review of 92 RCTs using HCEs found that their use has increased across medical fields, but their construction, analytical approaches, and reporting of results remained highly heterogeneous. By systematically mapping how HCEs are currently implemented, further review is essential for the development of much needed methodological and reporting standards.

Introduction

While randomized clinical trials (RCTs) are the criterion standard to assess the efficacy of new medical interventions, their conduct is often time consuming and expensive.1,2,3,4,5 In cases where the events of interest are rare and when treatment effect sizes are likely to be modest, large sample sizes are required. This may not always be feasible due to limited resources.6,7 To facilitate adequately powered RCTs capable of detecting small, but clinically important differences, many trials adopt composite end points, combining multiple individual outcomes into a single metric.8 This strategy increases event frequency and reduces required sample sizes. However, conventional composite end points have a significant limitation: they typically treat all component events equally, regardless of their relative clinical importance. This may obscure important changes in higher priority outcomes. For instance, when considering a composite end point of mortality and hospitalization, an intervention that substantially reduces mortality may appear to have limited benefit if survivors experience more nonfatal events such as rehospitalization.9,10,11 Hierarchical composite end points (HCEs) are designed to address this limitation by ranking 2 or more outcome components by clinical importance (eg, mortality is ranked of higher importance than hospitalization).

The concept of generalized pairwise comparisons, which is often used to analyze HCEs, was described in detail by Finkelstein and Schoenfeld in 1999.12,13 Building upon this, Pocock and colleagues proposed the win ratio in 2011.14 In these methods, all participants in the intervention group are compared pairwise with control participants to assess whether the intervention group has an overall better outcome. Winners and losers are determined starting from the highest-priority component, with ties carried forward to the next component (unresolved ties remain possible on the lowest priority component).15 Using this concept, multiple components can be integrated in a hierarchical outcome assessment, which might be more reflective of clinical decision-making.16

The win ratio and other related HCE analytic methods (eg, win odds, probabilistic index) have been increasingly used in cardiovascular trials17 as well as in infectious disease RCTs where HCEs are often labeled as desirability of outcome ranking (DOOR).16,18,19 A systematic review focusing specifically on DOOR end points in infectious disease studies has been conducted,15 but found only 2 infectious disease RCTs using DOOR-based HCEs as primary outcomes, one of them a pilot trial.20,21

Despite growing interest in HCEs, key questions remain unanswered, including how to define, justify, and analyze these end points, how prevalent these hierarchical end points are in different medical areas, what terminology is used for their description, what processes are used to justify their ranking, and how sample size calculations and results should be reported. We therefore conducted a scoping review to identify all RCTs which used an HCE as their primary end point to provide an overview of current research practice.

Methods

This scoping review is reported according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) reporting guideline extension for scoping reviews.22,23 A detailed protocol was prospectively registered on Open Science Framework.24 No ethical approval was obtained as this scoping review used only data from publicly available sources and did not involve the additional collection or use of primary data from human participants.

Systematic Search and Eligibility Criteria

An initial pilot search was conducted in Embase using generic terms for HCE to identify potential pertinent vocabulary. Following this, an information specialist composed the final search for RCTs (covering published studies, protocols and trial registries) using an HCE as a primary outcome. The systematic search was conducted on December 9, 2024, in Embase, Medline (Ovid), Web of Science Core Collection, and the Cochrane Central Register of Controlled Trials (CENTRAL) (detailed search strategies in eAppendix 1 in Supplement 1). Records were identified and exported to EndNote 21 (Clarivate) and deduplicated using the Automated Systematic Search Deduplicator.25

References were uploaded into Covidence and screened by titles and abstracts, followed by full-text screening in duplicate by 2 independent reviewers (S.E., A.J.dJ., Y.M., M.E.A.dK., H.J., N.H.K., M.S., N.S.B., N.R., S.W.X.O., and B.S.). Discrepancies were resolved by discussion, or, if no consensus could be reached, by consulting a third reviewer (S.E., S.W.X.O., M.B., and B.S.).

Trials were included independently of the trial status (ie, we also included ongoing trials based on their trial registration or protocol) if they used an HCE as their primary outcome. HCEs were defined as end points that combine multiple distinct clinical components into a single outcome, with a specified ranked or prioritized order for evaluating these components. Because differentiating between an HCE and an end point based on an ordinal scale can be challenging, we used the following definition for operational purposes: the primary composite end point was considered hierarchical if it was either (1) explicitly self-declared as a “hierarchical” end point or as an “outcome ranking” by the authors, or (2) if it was analyzed using a method specific for an HCE (eg, win ratio, win odds, probabilistic index, generalized pairwise comparisons). Studies were excluded if they were pilot or feasibility studies, were post hoc or secondary analyses, used only hierarchical secondary end points, used hierarchically tested coprimary end points, or used a hierarchy that was primarily developed to account for missing data or hierarchical testing to account for multiple testing. No restrictions on language or publication date were applied. From the included RCTs, we extracted cited methodological papers on HCEs, deduplicated these records, and conducted forward citation searches to identify additional potentially eligible trials.

Data Extraction

Data extraction was performed in duplicate by 2 researchers (S.E., A.J.dJ., Y.M., M.E.A.dK., H.J., N.H.K., M.S., N.S.B., N.R., S.W.X.O., and B.S.) using REDCap26,27 (agreement 88% on extracted variables). Conflicts were consolidated by a third reviewer (S.E., S.W.X.O., M.B., and B.S.) and discussed with the 2 researchers who performed extraction. We extracted detailed trial characteristics using a standardized data extraction form (eAppendix 2 in Supplement 1). Extracted information included trial identifiers, trial status, available sources of information (ie, publications, protocols, registry, other), patient population, medical field, intervention, sponsorship, planned and achieved sample size, study design, masking status of care provider, patients and outcome assessors, dates of first patient inclusion and publication of results, and details of the hierarchical composite primary end point (including number and type of components, process used to establish the hierarchy, reported strengths and limitations of HCEs, handling of continuous outcomes, tie-breaking methods, statistical analysis methods, and reporting of overall summary measures). Sample size assumptions and justification for expected improvements were also collected.

Statistical Analysis

We summarized the characteristics of the included RCTs descriptively using frequencies and percentages for categorical variables. When reporting proportions, we included only RCTs in the total denominator for which a result could have been expected (eg, when assessing reported limitations of HCE, we excluded RCTs for which we only had information from trial registries). Data cleaning, coding, and analysis were conducted in R version 4.5.1. (R Foundation for Statistical Computing).

Results

A total of 5188 records were screened (4200 from the initial literature search, plus 988 from the forward citation search). The 175 eligible identified records using an HCE as primary outcome were allocated to a total of 92 unique randomized clinical trials (RCTs could have multiple records such as results publication, published protocol, and registry record) (Table 1; flowchart available in the eFigure and overview of individual studies in eTable 1 in Supplement 1). These RCTs had an estimated total of 79 567 planned participants. Most included trials were drug trials (40 of 92 [43.5%]), conducted in the fields of cardiology (40 of 92 [43.5%]), were non–industry sponsored (62 of 92 [67.4%]), and conducted in multiple centers (84 of 92 [91.3%]) (stratified by sponsor type in eTable 2 in Supplement 1). The Figure shows the number of trials by medical field initiating patient recruitment each year, from the earliest included trial starting recruitment in 1975 to the most recent in 2025. While cardiology remained the dominant field, there was an increase in trials related to other medical fields (eg, infectious diseases), alongside a strong overall increase in the use of HCEs in the past decade (Figure). A total of 47 out of 92 included trials were completed (51.1%), with 45 having results published; 34 trials (37.0%) were recruiting or had ongoing follow-up, 6 trials (6.5%) were prematurely discontinued, and 4 trials (4.3%) had not started recruitment yet.

Table 1. General Characteristics of Included Randomized Clinical Trials (RCTs) Using a Hierarchical Composite End Point as Primary Outcome.

Characteristic Total RCTs, No (%) (N = 92)
Planned sample size, median (IQR) 400 (240-750)
Sponsor
Nonindustry 62 (67.4)
Industry 30 (32.6)
Medical field
Cardiology 40 (43.5)
Infectious diseases 10 (10.9)
Intensive care 9 (9.8)
Neurology 7 (7.6)
Othera 26 (28.3)
Trial status
Completed 47 (51.1)
Recruiting or ongoing follow-up 34 (37.0)
Premature discontinuation 6 (6.5)
Not yet recruiting 4 (4.3)
Unclear 1 (1.1)
Type of intervention
Drug 40 (43.5)
Health service delivery including diagnostics 15 (16.3)
Medical device 13 (14.1)
Surgical or invasive procedure 10 (10.9)
Otherb 14 (15.2)
No. of centers
Multicenter 84 (91.3)
Single center 8 (8.7)
International vs national
International 47 (51.1)
National 45 (48.9)
Design
Parallel group RCT 86 (93.5)
Crossover RCT 2 (2.2)
Factorial design 2 (2.2)
Otherc 2 (2.2)
Randomization unit
Individual 89 (96.7)
Cluster 3 (3.3)
Treatment groups
2 81 (88)
3 6 (6.5)
4 5 (5.4)
Trial objective
Superiority 90 (97.8)
Noninferiority 2 (2.2)
Patient masking
No 55 (59.8)
Yes 37 (40.2)
Physician or health care practitioner masking
No 59 (64.1)
Yes 33 (35.9)
Outcome assessor masking
No 30 (32.6)
Yes 52 (56.5)
Not reported 10 (10.9)
a

Total of 26 trials, including: 4 orthopedics; 3 gastroenterology; 3 cardiothoracic surgery; 2 pneumology; 2 nephrology; 2 oncology; 2 emergency medicine; 2 angiology; and 1 each of rheumatology, hematology, endocrinology, pediatrics, urology, and geriatrics.

b

Total of 14 trials, including: 4 biological interventions; 3 airway management; 2 behavioral, lifestyle, or psychological intervention; 2 mixed interventions (drug and procedure, behavioral and drug); and 1 each radiation, complementary therapy, or pharyngeal electrical stimulation.

c

Two trials: ENABLE was 2 separate parallel group RCTs analyzed combined, and LIVE-SMART sequential randomization.

Figure. Frequency of Recruitment Start of Randomized Clinical Trials Using a Hierarchical Composite Primary End Point.

Stacked bar chart of trials by year of first patient included and medical field. Single stacked bar chart with a white background and light gray horizontal gridlines. Vertical axis label at left: Trials, No., with tick marks from 0 up to 14 in steps of 2. Horizontal axis label at bottom: Year of first patient included, with labeled ticks at 1975, 1980, 1990, 2000, 2010, 2020, and 2025. A boxed legend in the upper left titled Medical field contains five color keys: Cardiology in medium blue gray; Infectious diseases in orange; Intensive care in light blue gray; Neurology in dark teal; Other with a superscript a in pale peach. Bars appear at selected years and are stacked by the legend categories, with black outlines around each stacked segment. Early years contain isolated single-segment bars: a pale peach bar near 1975 at about 1 trial; medium blue gray bars near 1989 and 1990 at about 1 trial each; a pale peach bar near 1995 at about 1 trial; a medium blue gray bar near 1999 at about 1 trial; a pale peach bar near 2005 at about 1 trial; and small dark teal or medium blue gray bars around 2007 to 2011 at about 1 trial each. From roughly 2012 to 2017, several short stacks appear, including a taller stack near 2012 reaching about 5 trials with multiple blue gray and dark teal segments, and additional stacks around 2013 to 2017 ranging about 2 to 4 trials, often including pale peach segments. From about 2019 onward, stacks increase markedly: a stack near 2019 around 7 trials; the tallest stacks occur around 2020 to 2024, reaching approximately 10 to 13 trials, with visible combinations of medium blue gray, orange, light blue gray, dark teal, and pale peach segments; a final shorter stack near 2025 reaches about 3 trials with pale peach and dark teal segments.

aA total of 26 trials, including: 4 orthopedics; 3 gastroenterology; 3 cardiothoracic surgery; 2 pneumology; 2 nephrology; 2 oncology; 2 emergency medicine; 2 angiology; and 1 each of rheumatology, hematology, endocrinology, pediatrics, urology, and geriatrics. For 4 trials the anticipated recruitment data was used; 1 did not report start date (2 years prior to publication was estimated).

Domains of the HCEs

The highest ranked component of the 92 HCEs was mortality for 74 trials (80.4%) (Table 2). The remaining 19.6% (18 of 92) could be grouped into the following 4 categories: disease-specific clinical outcome events (6 events; eg, asthma exacerbations), functional or recovery (6 events; eg, return to work, fracture union or clinical recovery), treatment failure or necessity of therapy escalation (4 events), and safety or adverse events (2 events).

Table 2. Characteristics of Components Included in Hierarchical Composite End Points.

Characteristic RCTs, No (%) (N = 92)
Highest ranked HCE component
Mortalitya 74 (80.4)
Disease specific clinical outcomeb 6 (6.5)
Functional or recoveryc 6 (6.5)
Treatment failure or necessity of therapy escalationd 4 (4.3)
Safety or adverse eventse 2 (2.2)
Handling of ties when both had an event on the highest ranked componentf
Taking into account lower component 21 (22.8)
Time to event taken into account 22 (23.9)
Counted as tie 19 (20.7)
Nonparametric rank methods (Wilcoxon or Mann-Whitney) 3 (3.3)
Unclear or not mentioned 35 (38.0)g
Lowest ranked HCE component
Continuous 59 (64.1)
Binary 21 (22.8)
Time-to-event 4 (4.3)
Otherh 8 (8.7)
Handling continuous componentsf
Direct comparison of the values 73 (71.6)
Predefined minimum threshold difference 11 (10.8)
Unclear 17 (16.7)
Other method usedi 1 (1.1)
No. of components, median (IQR) 4 (3-5)
No. of components per HCE
2 22 (23.9)
3 22 (23.9)
4 23 (25.0)
5 12 (13.0)
6 4 (4.3)
7 9 (9.8)

Abbreviations: HCE, hierarchical composite end point; RCT, randomized clinical trial.

a

Including 53 with all-cause mortality, 8 time to death, 7 cardiovascular death, 2 hospital mortality, 2 fatal perioperative or periprocedural bleeding, 1 mortality due to liver failure, and 1 death due to thrombosis.

b

Including asthma exacerbation, change in seizure frequency, major amputation, total rate of heart failure admissions, and swallowing safety.

c

Including clinical outcome, radiographic fracture union, efficacy plus safety, return to work, clinical, recovery, and adequate clinical response.

d

Including necessity of intensive care unit care during hospitalization and failure of noninvasive ventilation.

e

Including safety assessment and serious adverse events.

f

Multiple answers possible from a single hierarchical composite end point.

g

For 17 of these studies (48.6%), information was only available from registries.

h

Including 3 categorical, 3 ordinal, and 2 unclear.

i

Model-based approach.

The lowest ranked component of the HCE was continuous in 59 trials (64.1%) and binary in 21 trials (22.8%). The 92 HCEs had a total of 345 components, with a median (IQR) of 4 (3-5) components per HCE (range, 2-7) (Table 2). Of all 345 components, 184 components (53.3%) were binary, 95 (27.5%) continuous, 41 (11.9%) time-to-event, 11 (3.2%) were ordinal or ranking, and 11 (3.2%) were categorical (eTable 3 in Supplement 1). Handling of tie breaks when both pairs had an event on the highest ranked component of the HCE (eg, both patients died) was either not reported (35 of 92 [38.0%]) or handled by taking into account the lower hierarchy component (21 of 92 [22.8%]), time-to-event (22 of 92 [23.9%]), or counted as ties (19 of 92 [20.7%]); multiple approaches could be used within a single trial.

Terminology, Development, and Rationale

The terminology “hierarchical composite end point” or “hierarchical composite outcome” was used in 62 of 92 identified trials (67.4%). Other terms used were “rank composite” (12 of 92 [13.0%]), and DOOR (10 of 92 [10.9%]) (Table 3). DOOR end points were primarily used by infectious disease trials and RCTs in neurology (key characteristics by medical field available in eTable 4 of Supplement 1). In more than 80% of the RCTs, authors did not provide any explanation of how they developed or selected the hierarchical order (60 of 73 [82.2%]; excluding 19 RCTs for which only information from a clinical trials registry was available). Six trials (8.2%) cited a reference to an established hierarchy, and for 2 trials each (2.7%) consulting patient representatives or using an expert committee to define the hierarchy were reported (Table 3). A justification for the selection of an HCE was reported in 39 of 73 RCTs (53.4%). The top 3 reported reasons were reporting general advantages compared with traditional end points (26 of 73 [35.6%]), that an HCE better reflects overall health (24 of 73 [32.9%]), and that the HCE increases statistical power (16 of 73 [21.9%]). Furthermore, 3 (4.1%) switched from a traditional end point to an HCE when realizing that they could not meet the originally targeted sample size due to recruitment issues. Only 10 RCTs (13.7%) discussed potential limitations of HCEs, most commonly noting difficulties in interpreting the effect size (5 of 73 [6.8%]) (Table 3).

Table 3. Development, Rationale, and Limitations of Hierarchical Composite End Points.

Characteristics RCTs, No./total No. (%)
Terminology used for end point
Hierarchical composite end point/outcome 62/92 (67.4)
Rank composite 12/92 (13.0)
DOOR 10/92 (10.9)
Othera 8/92 (8.7)
Development of hierarchy of HCEb
Nothing reported 60/73c (82.2)
Provided information on HCE development 13/73 (17.8)
Hierarchy already established (reference provided) 6/73 (8.2)
Patient and public involvement 2/73 (2.7)
Expert committee 2/73 (2.7)
Otherd 4/73 (5.5)
Justification for use of HCEb
No justification reported 34/73 (46.6)
Justification reported 39/73 (53.4)
Advantage vs traditional composite end points 26/73 (35.6)
Better reflects overall health 24/73 (32.9)
Increase powere 16/73 (21.9)
Otherf 8/73 (11.0)
Limitations of HCEb
No limitations reported 63/73 (86.3)
Limitations reported 10/73 (13.7)
Effect size difficult to interpret or unclear 5/73 (6.8)
Otherg 7/73 (9.6)

Abbreviations: DOOR, desirability of outcome ranking; HCE, hierarchical composite end point; RCT, randomized clinical trial.

a

Including 4 with win ratio approach, 2 self-declared as “composite” but analyzed with hierarchy, and 2 net treatment benefit.

b

Multiple answers could be reported and counted toward multiple categories.

c

Including 73 studies; excluding 19 studies with only registry entry available.

d

Including 2 with rationale provided by authors, 1 suggested by data monitoring committee, and 1 simulation.

e

Including 3 in which started trials faced recruitment issues and therefore changed the originally planned end point to a hierarchical composite end point during the trial conduct.

f

Including 5 to better reflect and handle missing data, 1 for safety aspects at higher hierarchy where no difference is expected, 1 technical advantage (no need for the proportionality of hazards), and 1 combining continuous and binary end points.

g

Including 2 with dominance of less important components, 2 due to compensatory nature, may mask worse results for individual end points, 1 method may overestimate effect, 1 adjustment not possible or difficult, and 1 limited generalizability.

Sample Size Estimation, Analysis, and Reporting of Results

The sample size assumptions for all components in the control group and the expected improvement for all components were reported in 43.8% (32 of 73) and 47.9% (35 of 73), respectively (Table 4). Approximately half of the trials (40 of 73 [54.8%]) adjusted the sample size for loss to follow-up or provided a rationale why it was not needed. Only 1 RCT (1.4%) reported how much the individual components contribute to the overall power of the trial.

Table 4. Sample Size Estimation, Analysis, and Reporting of Results for Hierarchical Composite End Points.

Characteristic RCTs, No. (%)a
Protocol, results, or statistical analysis plan available (n = 73)
Assumptions for control group reported for each component
For all components 32 (43.8)
Only for single components (ie, not for all) 12 (16.4)
Not at all 29 (39.7)
Expected improvement for each component reported
For all components 35 (47.9)
Only for single components (ie, not for all) 10 (13.7)
Not at all 28 (38.4)
α Reported (type I error) 64 (87.7)
Power reported 69 (94.5)
Each component’s contribution (eg, % or proportion) to overall trial power indicated 1 (1.4)
Correction for loss to follow-up (or providing rationale why not done) 40 (54.8)
Reporting results (45 RCTs)
Results presentationb
Table with results for each component 20 (44.4)
Bar charts 10 (22.2)
Flowcharts 8 (17.8)
Results without illustration 6 (13.3)
Otherc 14 (31.1)
Primary analysis of hierarchical composite end point
GPC 29 (64.4)
GPC using win ratiod 21 (46.7)
Rank order method 11 (24.4)
Probabilistic index model 2 (4.4)
Othere 3 (6.6)
Additional partial credit analysis 1 (2.2)
Handling of missing data
No imputation or negligible missing 14 (31.1)f
Multiple imputation 8 (17.8)
Censoring on last observation 3 (6.6)
Single imputation 2 (4.4)
Otherg 5 (11.1)
Unclear 13 (28.8)
Missing data for each component reported 4 (8.9)
Adjustment for covariates in analysis 8 (17.8)
Each component included as a secondary end point 28 (62.2)
Pairwise comparison conductedh 33 (73.3)
No. of patient pairs reported 18 (40.0)
Win ratio/difference for each component reported 16 (35.6)
No. of ties for each component reported 11 (24.4)
Weight of each component on overall effect reported 6 (13.3)

Abbreviations: GPC, generalized pairwise comparison; RCT, randomized clinical trial.

a

For the sample size estimation and planned analysis only, studies were included for which at least a protocol, results publication, or statistical analysis plan was available (73 studies); for the analysis, only studies with results publication were considered (45 studies).

b

Multiple answers could be reported and counted toward multiple categories.

c

Including 9 with tables, 3 with forest plots, 2 with Kaplan-Meier curves (1 of those with additional table).

d

Including 19 (90.5%) conducted unmatched win ratio and 1 (4.8%) matched win ratio. One study reported using stratified win ratio (4.8%).

e

Including 1 analysis of covariance, 1 nonlinear mixed-effects model, and 1 model-based approach with adjustment and pairwise comparison as sensitivity analysis.

f

Including 8 reporting a sensitivity analysis using imputed data (4 studies) or tipping point analysis (2 studies), and 1 using inverse probability weighted analysis or different scenarios.

g

Including 2 studies with next hierarchy, 1 last observation carried forward, 1 combination of imputing and censoring, and 1 decision by independent outcome committee in case of missing data.

h

Including 4 studies that conducted a pairwise comparison not as the primary analysis.

Published results were available for 45 trials. The presentation of published results consisted of tables with results for each component (20 of 45 [44.4%]), flowcharts (8 of 45 [17.8%]), bar charts (10 of 45 [22.2%]), and results without illustration (6 of 45 [13.3%]). Two-thirds did not include each component of the primary end point as a secondary end point (28 of 45 [62.2%]). Only 6 trials (13.3%) reported the weight of each component on the overall effect size, and 4 (8.9%) reported the proportion of missing data for each component of the HCE. Adjusting the analysis for covariates (eg, stratification factors) was rarely conducted (8 of 45 [17.8%]). The most frequently used analysis methods were generalized pairwise comparisons (29 of 45 [64.4%]) and rank order method (11 of 45 [24.4%]) (Table 4). The most commonly reported summary measure was the win ratio (21 of 45 [46.7%]; 19 unmatched win ratios, 1 matched win ratio, and 1 stratified win ratio). A partial-credit analysis assigning varying weights to different outcome components was reported in 1 RCT (Table 4).

Among the 29 trials conducting a pairwise comparison, we identified a total of 11 RCTs that reported the overall wins, losses, and unresolved ties in the results publication. The proportion of remaining unresolved ties among these 11 trials ranged from 0.3% up to 77.3% (median [IQR] 23.8% [3.1%-31.8%]) (eTable 5 in Supplement 1). One reported win odds and win ratio, the remaining 10 reported win ratio.

Discussion

Our scoping review revealed that RCTs with HCEs as their primary end point have become more frequent, primarily in cardiovascular trials, but with growing adoption in other medical fields. This is in agreement with a recently published systematic review focusing on win ratio which showed an increased use beyond cardiovascular trials.28 The main arguments reported by trialists for using an HCE were the general advantages over traditional composite end points, that HCEs can better reflect overall health of patients, and the increase of power when using an HCE. Our scoping review reveals substantial heterogeneity and shortcomings in how investigators plan, analyze, and report the use of HCEs in RCTs. Given the potential methodological pitfalls associated with HCEs, as highlighted in the current research,17,29,30 the development of clear, standardized guidelines is essential to mitigate these risks. For example, when developing an HCE, Ajufo and colleagues17 recommend that several stakeholders, including patients, trialists and regulators, should be included to ensure that the HCE is patient-centered and clinically relevant. However, our scoping review has shown that for most RCTs (over 80%), it was not reported at all how the hierarchy was developed. Furthermore, reporting of the sample size calculation, so readers can understand the underlying assumptions for each component,31 was often incomplete. Only about half of the trials provided the expected event rates or values in the control group for each component, the anticipated improvement in the intervention group, and details on how the sample size was adjusted for missing data.

Our findings have several implications for clinicians, trialists, and policymakers. These stakeholders should be aware when assessing effect estimates of an HCE that the overall effect may be strongly driven by its lowest component.17 Therefore, all included components should be clinically meaningful to patients, and the contribution of each component to the overall effect size (ie, proportion of overall decisions made on this component) should be reported. Among the RCTs included in our study, only 13% reported the weight of each component on the overall effect size. Another frequently discussed limitation of HCEs (in particular for the win ratio) is that the win ratio does not include unresolved ties in the analysis. This is especially concerning when the proportion of unresolved ties is high.29 In the 11 RCTs that performed pairwise comparisons and reported unresolved ties, this proportion was substantial, with a median of 23.8% and a maximum of 77.3% in one of the trials. To account for this problem, the use of win odds, which allocates half the ties to the intervention and the other half to the control group, is often proposed.32,33 In our review, we identified only 1 RCT that reported win odds, and this trial had a negligible proportion of unresolved ties (0.6%). It is noteworthy that Pocock and colleagues argue that “it is a myth that ties are a problem using the win ratio,” suggesting that the win difference can be calculated for each component to quantify the absolute benefit.30 Yet, none of the trials we identified reported the win difference. We suggest that these different target parameters (win ratio, win odds, win difference or net treatment benefit, and probabilistic index) be reported in parallel to provide complementary measures of both the absolute and relative treatment effect.34,35

One additional remaining question is how to best present results of HCEs. Our review revealed a wide range of approaches, many of which lacked transparency. We suggest that a flowchart clearly illustrating the hierarchy and decision-making process may offer the most intuitive and comprehensive solution. The flowchart should not only include the overall wins and losses but be complemented by detailed numerical information for each component (such as the counts of wins, ties, and losses; eg, as presented by Mentz and colleagues36), along with its weight on the overall effect and corresponding win difference, enabling readers to clearly understand the contribution of each element to the overall effect. The optimal approach to integrating covariate adjustment into both the analysis and presentation of results remains uncertain and warrants further investigation.30

Limitations

While our scoping review provides a broad overview of RCTs employing HCEs as a primary end point, we acknowledge several limitations. First, because HCEs combine features from ordinal and composite end points,19 clearly defining an end point as an HCE is not always straightforward. For operational purposes we used a combination of either self-declaration or the use of analytical approaches typical for HCEs. Consequently, we may have included studies whose primary end point is not universally recognized as an HCE. In addition, we may have missed trials using HCEs which are not clearly labeled as such. Second, our scoping review provides only a snapshot of the current trial landscape. A large proportion of trials was still ongoing, with only limited information (eg, only from trial registries) and no published results available. It is therefore possible that certain aspects (eg, reporting quality of study results) may change. Third, our study did not allow assessment of other important considerations, such as resource savings compared with traditional end points, differences in effect sizes between composite end points and HCEs, or the level of acceptance of HCEs among decision-makers. Alternative designs such as economic evaluations, statistical simulations, and qualitative studies are needed to answer these remaining questions.

Conclusions

In this scoping review, we showed that the use of HCEs is expanding across medical fields in both academic and industry-sponsored trials, offering multiple potential advantages for evaluating clinically meaningful, patient-centered outcomes in RCTs. At the same time, our review showed considerable heterogeneity and shortcomings in how HCEs are developed, analyzed, and reported, with key steps such as hierarchy justification or handling of ties often being insufficiently described. This limits the interpretability of trial findings and complicates evidence synthesis. Establishing clear methodological standards and reporting recommendations for RCTs using HCEs would support more consistent implementation and enhance the validity and interpretability of HCE-based evidence for clinicians, patients, researchers, and decision-makers.

Supplement 1.

eAppendix 1. Full search strategy conducted in different databases

eAppendix 2. Data extraction form

eTable 1. Table of all included studies

eTable 2. Characteristics of trials stratified by sponsor

eTable 3. Categories of components used in the 92 identified hierarchical composite endpoints

eTable 4. Sub-group analysis of key characteristics by medical field

eTable 5. Studies that conducted a pairwise comparison and reported winners, losers, and ties for the hierarchical endpoint

eFigure. Flow chart

Supplement 2.

Data Sharing Statement

References

  • 1.Duley L, Antman K, Arena J, et al. Specific barriers to the conduct of randomized trials. Clin Trials. 2008;5(1):40-48. doi: 10.1177/1740774507087704 [DOI] [PubMed] [Google Scholar]
  • 2.Collins R, MacMahon S. Reliable assessment of the effects of treatment on mortality and major morbidity, I: clinical trials. Lancet. 2001;357(9253):373-380. doi: 10.1016/S0140-6736(00)03651-5 [DOI] [PubMed] [Google Scholar]
  • 3.Collier R. Rapidly rising clinical trial costs worry researchers. CMAJ. 2009;180(3):277-278. doi: 10.1503/cmaj.082041 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Larson GS, Carey C, Grarup J, et al. ; INSIGHT Group . Lessons learned: infrastructure development and financial management for large, publicly funded, international trials. Clin Trials. 2016;13(2):127-136. doi: 10.1177/1740774515625974 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Griessbach A, Speich B, Amstutz A, et al. ; MAking Randomized Trials Affordable (MARTA) Group . Resource use and costs of investigator-sponsored randomized clinical trials in Switzerland, Germany, and the United Kingdom: a metaresearch study. J Clin Epidemiol. 2024;176:111536. doi: 10.1016/j.jclinepi.2024.111536 [DOI] [PubMed] [Google Scholar]
  • 6.Sedgwick P. Randomised controlled trials: understanding effect sizes. BMJ. 2015;350:h1690. doi: 10.1136/bmj.h1690 [DOI] [PubMed] [Google Scholar]
  • 7.Ialongo C. Understanding the effect size and its measures. Biochem Med (Zagreb). 2016;26(2):150-163. doi: 10.11613/BM.2016.015 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.McCoy CE. Understanding the use of composite endpoints in clinical trials. West J Emerg Med. 2018;19(4):631-634. doi: 10.5811/westjem.2018.4.38383 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Freemantle N, Calvert M, Wood J, Eastaugh J, Griffin C. Composite outcomes in randomized trials: greater precision but with greater uncertainty? JAMA. 2003;289(19):2554-2559. doi: 10.1001/jama.289.19.2554 [DOI] [PubMed] [Google Scholar]
  • 10.Ferreira-González I, Permanyer-Miralda G, Busse JW, et al. Methodologic discussions for using and interpreting composite endpoints are limited, but still identify major concerns. J Clin Epidemiol. 2007;60(7):651-657. doi: 10.1016/j.jclinepi.2006.10.020 [DOI] [PubMed] [Google Scholar]
  • 11.Montori VM, Permanyer-Miralda G, Ferreira-González I, et al. Validity of composite end points in clinical trials. BMJ. 2005;330(7491):594-596. doi: 10.1136/bmj.330.7491.594 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Finkelstein DM, Schoenfeld DA. Combining mortality and longitudinal measures in clinical trials. Stat Med. 1999;18(11):1341-1354. doi: 10.1002/(SICI)1097-0258(19990615)18:11<1341::AID-SIM129>3.0.CO;2-7 [DOI] [PubMed] [Google Scholar]
  • 13.Tang R, Chen WC, Li H, Lu N, Zhao Y. The Finkelstein-Schoenfeld test: a note on some overlooked issues concerning power. Ther Innov Regul Sci. 2024;58(3):465-472. doi: 10.1007/s43441-023-00608-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Pocock SJ, Ariti CA, Collier TJ, Wang D. The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities. Eur Heart J. 2012;33(2):176-182. doi: 10.1093/eurheartj/ehr352 [DOI] [PubMed] [Google Scholar]
  • 15.Ong SWX, Petersiel N, Loewenthal MR, Daneman N, Tong SYC, Davis JS. Unlocking the DOOR-how to design, apply, analyse, and interpret desirability of outcome ranking endpoints in infectious diseases clinical trials. Clin Microbiol Infect. 2023;29(8):1024-1030. doi: 10.1016/j.cmi.2023.05.003 [DOI] [PubMed] [Google Scholar]
  • 16.Doernberg SB, Tran TTT, Tong SYC, et al. ; Antibacterial Resistance Leadership Group . Good studies evaluate the disease while great studies evaluate the patient: development and application of a desirability of outcome ranking endpoint for staphylococcus aureus bloodstream infection. Clin Infect Dis. 2019;68(10):1691-1698. doi: 10.1093/cid/ciy766 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ajufo E, Nayak A, Mehra MR. Fallacies of using the win ratio in cardiovascular trials: challenges and solutions. JACC Basic Transl Sci. 2023;8(6):720-727. doi: 10.1016/j.jacbts.2023.05.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Evans SR, Rubin D, Follmann D, et al. Desirability of outcome ranking (DOOR) and response adjusted for duration of antibiotic risk (RADAR). Clin Infect Dis. 2015;61(5):800-806. doi: 10.1093/cid/civ495 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ong SWX, Mahar RK, Selman CJ, et al. Making sense of hierarchical composite endpoints in randomized clinical trials–a primer for infectious disease clinicians and researchers. Clin Infect Dis. 2025;81(5):e319-e329. doi: 10.1093/cid/ciaf314 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Williams DJ, Creech CB, Walter EB, et al. ; The DMID 14-0079 Study Team . Short- vs Standard-Course Outpatient Antibiotic Therapy for Community-Acquired Pneumonia in Children: the SCOUT-CAP randomized clinical trial. JAMA Pediatr. 2022;176(3):253-261. doi: 10.1001/jamapediatrics.2021.5547 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Manning L, Metcalf S, Dymock M, et al. ; Australasian Society for Infectious Diseases Clinical Research Network . Short- versus standard-course intravenous antibiotics for peri-prosthetic joint infections managed with debridement and implant retention: a randomised pilot trial using a desirability of outcome ranking (DOOR) endpoint. Int J Antimicrob Agents. 2022;60(1):106598. doi: 10.1016/j.ijantimicag.2022.106598 [DOI] [PubMed] [Google Scholar]
  • 22.Tricco AC, Lillie E, Zarin W, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018;169(7):467-473. doi: 10.7326/M18-0850 [DOI] [PubMed] [Google Scholar]
  • 23.Mbuagbaw L, Lawson DO, Puljak L, Allison DB, Thabane L. A tutorial on methodological studies: the what, when, how and why. BMC Med Res Methodol. 2020;20(1):226. doi: 10.1186/s12874-020-01107-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Speich B. Randomized trials using a hierarchical composite primary endpoint: a systematic scoping review—study protocol wiki. Updated January 16, 2025. Accessed December 24, 2025. https://osf.io/mjc52/overview
  • 25.Hair K, Bahor Z, Macleod M, Liao J, Sena ES. The Automated Systematic Search Deduplicator (ASySD): a rapid, open-source, interoperable tool to remove duplicate citations in biomedical systematic reviews. BMC Biol. 2023;21(1):189. doi: 10.1186/s12915-023-01686-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Harris PA, Taylor R, Minor BL, et al. ; REDCap Consortium . The REDCap consortium: building an international community of software platform partners. J Biomed Inform. 2019;95:103208. doi: 10.1016/j.jbi.2019.103208 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)–a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42(2):377-381. doi: 10.1016/j.jbi.2008.08.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Li Q, Zhu Y, Zhao Y, et al. The win ratio in contemporary clinical trials: growing adoption and interpretation challenges—a meta-epidemiological study. J Clin Epidemiol. 2026. doi: 10.1016/j.jclinepi.2026.112315 [DOI] [PubMed] [Google Scholar]
  • 29.Kotanidis Christos P, Gorey S, Müller D, et al. How win ratios work. NEJM Evidence. 2025;4(8). doi: 10.1056/EVIDstat2500187 [DOI] [PubMed] [Google Scholar]
  • 30.Pocock SJ, Gregson J, Collier TJ, Ferreira JP, Stone GW. The win ratio in cardiology trials: lessons learnt, new developments, and wise future use. Eur Heart J. 2024;45(44):4684-4699. doi: 10.1093/eurheartj/ehae647 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Redfors B, Gregson J, Crowley A, et al. The win ratio approach for composite endpoints: practical guidance based on previous experience. Eur Heart J. 2020;41(46):4391-4399. doi: 10.1093/eurheartj/ehaa665 [DOI] [PubMed] [Google Scholar]
  • 32.Song J, Verbeeck J, Huang B, et al. The win odds: statistical inference and regression. J Biopharm Stat. 2023;33(2):140-150. doi: 10.1080/10543406.2022.2089156 [DOI] [PubMed] [Google Scholar]
  • 33.Brunner E, Vandemeulebroecke M, Mütze T. Win odds: an adaptation of the win ratio to include ties. Stat Med. 2021;40(14):3367-3384. doi: 10.1002/sim.8967 [DOI] [PubMed] [Google Scholar]
  • 34.Hardy MJ, Ong SWX, Paterson DL. The win ratio should be complemented by other win statistics to provide a comprehensive picture of relative and absolute treatment effects: the case study of the REPRIEVE trial. Clin Infect Dis. 2025;82(2):e413-e414. doi: 10.1093/cid/ciaf587 [DOI] [PubMed] [Google Scholar]
  • 35.Dong G, Huang B, Verbeeck J, et al. Win statistics (win ratio, win odds, and net benefit) can complement one another to show the strength of the treatment effect on time-to-event outcomes. Pharm Stat. 2023;22(1):20-33. doi: 10.1002/pst.2251 [DOI] [PubMed] [Google Scholar]
  • 36.Mentz RJ, Garg J, Rockhold FW, et al. ; HEART-FID Investigators . Ferric carboxymaltose in heart failure with iron deficiency. N Engl J Med. 2023;389(11):975-986. doi: 10.1056/NEJMoa2304968 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1.

eAppendix 1. Full search strategy conducted in different databases

eAppendix 2. Data extraction form

eTable 1. Table of all included studies

eTable 2. Characteristics of trials stratified by sponsor

eTable 3. Categories of components used in the 92 identified hierarchical composite endpoints

eTable 4. Sub-group analysis of key characteristics by medical field

eTable 5. Studies that conducted a pairwise comparison and reported winners, losers, and ties for the hierarchical endpoint

eFigure. Flow chart

Supplement 2.

Data Sharing Statement


Articles from JAMA Network Open are provided here courtesy of American Medical Association

RESOURCES