Abstract
Advances in oncology drug development are driving the emergence of novel therapies, challenging traditional dose-efficacy assumptions in dose-finding oncology trials. Traditional trial designs aim to identify a maximum tolerated dose (MTD) by assessing patients’ dose-limiting toxicities (DLTs) – adopting traditional dose-efficacy paradigms that efficacy increases with treatment dose. With these new therapies in mind, emphasis should shift toward methodological advancements in trial designs aimed at identifying optimal doses, rather than solely determining MTDs. Incorporating patient-reported outcomes (PROs) within dose-finding oncology trials is increasingly recommended to better understand treatments’ tolerability profiles, especially given the extended tolerability assessment windows for novel immunotherapies and targeted therapies. This article introduces PRO-ADD (Patient-Reported Outcomes Aided Dose-optimisation Design), a modular trial design framework for dose-optimisation. We leverage this framework to optimise dosage with respect to three key outcomes – clinician-assessed DLTs, PROs and preliminary efficacy. PRO-ADD performs well at identifying the optimal dose (both efficacious and tolerable) – successfully identifying the most tolerable effective dose and avoiding escalation to larger, safe doses offering no additional efficacy benefit. As the field evolves, patient-centric dose-finding approaches incorporating PROs are crucial in advancing our understanding of treatment tolerability, and in turn, shaping the future landscape of dose-finding oncology trials.
Keywords: Patient-reported outcomes, dose-finding trials, dose-optimisation, optimal biological dose, phase I
1. Introduction
Phase I dose-finding oncology trials typically look to assess the safety of new clinical treatments emerging from drug development across a range of doses. Traditional cytotoxic treatments often demonstrate a positive correlation between dosage, toxicity, and efficacy – with a higher dose exhibiting greater activity and an increased probability of toxicity. Consequently, traditional dose-finding oncology trial designs look to identify a maximum tolerated dose (MTD) by estimating the probability a patient experiences a dose-limiting toxicity (DLT) during the trial. Trial designs often look to recommend the largest dose with a DLT probability closest to the target level – recommending the (assumed) most efficacious dose whilst safeguarding patient’s safety.
Advances in oncology drug development are driving the emergence of novel therapies that challenge these traditional assumptions. 1 Research suggests the implicit monotonicity assumption embedded in MTD determination may be violated for these new treatments, particularly in the case of immunotherapies and molecularly targeted agents (MTAs). Unlike cytotoxic treatments, these novel agents may exhibit a plateau in efficacy beyond a certain dose, so increasing the dose does not necessarily increase efficacy. Recent recommendations therefore suggest that dose-finding oncology trials should focus on determining minimum effective doses or optimal biologically active dose rather than MTDs for these therapies. 2
Whilst MTDs are almost always identified for cytotoxic treatments in phase I trials, more than one third of phase I trials investigating MTAs observe few DLTs and fail to establish the MTD. 3 The rapid development of immune-oncology agents emerging from drug development has motivated the rise of seamless Phase I/II trial designs to assess both preliminary efficacy and toxicity endpoints within early phase trials. 4 Whilst such designs look to satisfy this specific dose-efficacy relationship for novel therapies, the majority of these trial designs continue to assess toxicity solely using clinician-assessed DLTs.5,6
For cytotoxic agents, treatment is often administered over a relatively short period of time – traditionally guided by DLTs that occur in the first cycle of treatment (usually 28 days). Such assessment windows are deemed appropriate for the investigation of cytotoxic chemotherapies as DLTs are often observable soon after treatment commencement.7,8 Conversely, MTAs and immunotherapies are administered for a longer period of time, often until disease progression or resistance is observed. 9 These treatments may induce low-grade toxicities 1 which, though not definable as a DLT, become intolerable to patients over prolonged dose administration periods. Immune-related adverse events associated with immunotherapy regimens are often observed beyond the traditional DLT assessment window, 2 and research has suggested that for MTAs, approximately 57% of grade 3 or 4 toxicities associated with treatment administration occur beyond the first cycle of treatment. 9 For example, Durvalumab, an approved immunotherapy treatment, and targeted cancer drug Ibrutinib have been associated with early discontinuation of treatment due to treatment adverse events.10,11 Trial designs with a time-to-event component have been recommended to account for such late-onset toxicities. 2
The utilisation of patient-reported outcomes (PROs) is being increasingly endorsed for use within dose-finding trial designs to refine our understanding of a treatment’s tolerability profile.12–15 As defined by the FDA, a PRO is a report of the status of a patient’s health condition coming directly from the patient, without interpretation of the patient’s response by a clinician or anyone else. 16 Standardised measures (PROMs), such as the PRO-CTCAE, allow patients to evaluate the tolerability of up to 78 symptomatic adverse events to assess the safety and tolerability of a novel therapy from a patient’s perspective. 17 The inclusion of PROs within dose-finding oncology trials has been advocated to support the assessment of subjective toxicities such as fatigue, pain and anxiety,13,18,19 which may lead to a conflicting assessment of treatment tolerability between patient and clinician.20,21
The use of PROs in dose-finding oncology trials has significantly increased over recent years, but is still emerging, with only 5.3% of dose-finding oncology trials including PROs as an endpoint. 22 A review of published dose-finding trials incorporating PROs found that PROs informed dose recommendations in just 11.4% of eligible trials (4/35). 23 Whilst it is recognised that PRO data in early phase trials may influence subsequent clinical development, 24 existing trial designs focus solely on binary patient-informed DLT endpoints observed within a DLT assessment window.25–27
The FDA are encouraging trialists to consider the inclusion of PROs for the assessment of tolerability in dose-finding and subsequent trials. 15 The integration of electronic systems for ePRO 28 collection facilitates tolerability assessment beyond the DLT assessment window, reducing the patient and site burden associated with paper-based data collection and handling. 29 Such advancements, alongside encouragement from regulators to assess PROs regularly within cancer clinical trials, 30 set the stage for longitudinal analysis of PRO data across multiple time points. Such analysis is increasingly motivated for immunotherapies and MTAs administered over prolonged periods, 24 where adverse events may be less clinically severe but increasingly bothersome to the patient over treatment duration. 31
Existing dose-finding trial designs do not incorporate longitudinal patient-reported outcome (PRO) assessments into dose-decision making. This represents a critical limitation of current approaches, with the longitudinal modelling of PROs more sensitively capturing patient-experienced toxicities over time compared to existing approaches. In this article, we introduce the first dose-finding trial design to integrate longitudinal PRO endpoints alongside preliminary efficacy and clinician-assessed toxicity. In doing so, we strengthen PRO-integrated trial designs’ ability to robustly evaluate tolerability and further align PRO-integrated trial designs with patient priorities for tolerability and therapeutic benefit. 32
Specifically, we introduce PRO-ADD (PRO-Aided Dose-optimisation Design) which combines clinician-assessed DLTs with PRO and preliminary efficacy data to recommend dose(s) which optimise patient-assessed tolerability to treatment and efficacy, and safeguards clinician-assessed tolerability. A beta regression mixed effect model is used to longitudinally model the novel PRO-nAE (PRO-normalised Adverse Event) burden score over time.
In Section 2 we introduce the PRO-ADD modular dose-optimisation trial design. Section 3 details the simulation study set-up. Results from the simulation studies are detailed in Section 4 and discussed in Section 5.
2. Methods
2.1. Modular trial design framework
PRO-ADD is a modular trial design framework that separates the trial into three key modules: initial dose-escalation, dose-optimisation, and final dose selection. The components of each module can be customised, allowing flexibility in selecting the dose-escalation method, endpoints, and interim analyses to tailor the design to the goals of a given trial. As illustrated in Figure 1, each component functions as an individual jigsaw piece, which can be assembled to build a tailored trial design.
Figure 1.
A figurative representation of the PRO-ADD modular trial design.
The PRO-ADD framework is intentionally conceptually simple, with modularity allowing additional elements where appropriate and feasible, supporting implementation across dose-finding trials of varying size and complexity whilst maintaining robustness and interpretability. This modular trial design firstly focuses upon identifying the MTD, before optimising dose with respect to additional key endpoints, and finally recommending a dose(s) for further evaluation. In the subsequent sections, we describe the endpoints and estimation approaches used in our illustrative implementation of PRO-ADD. Section 2.6 then details the trial design, randomisation procedure and final dose-decision criterion utilised for this example. Rather than prescribing a specific design, this example illustrates one possible implementation of PRO-ADD. Simpler alternatives may also be used if deemed appropriate.
2.2. Patient-reported outcome endpoints
Several binary, ordinal and continuous endpoints have been proposed to capture and summarise PRO data.25,33
2.2.1. Patient dose-limiting toxicity
The patient dose-limiting toxicity (P-DLT), first proposed by Lee et al. 25 is a binary endpoint similar to the clinician assessed DLT typically used to assess toxicities within dose-finding oncology trials. Presently, PROs have only been incorporated within dose-finding trial designs as a binary endpoint, with current work focussing solely upon the determination of a MTD.25–27,34
The limitations of condensing the multidimensional nature of tolerability assessment into a binary DLT have motivated the development of novel endpoints aimed at more comprehensively accounting for the overall severity and clinical relevance of multiple toxicities within tolerability assessments. 35 Limitations, such as overlooking moderate toxicities which fall short of the DLT definition, and the number of DLTs which occur, 35 impact both clinician-assessed DLT and P-DLT endpoints.
2.2.2. Normalised PRO-adverse event (PRO-nAE) burden score
A number of continuous burden scores have been proposed to summarise adverse events utilising CTCAE33,35–39 and many can be readily adapted for PRO measures including the PRO-CTCAE. A more in-depth discussion of an exemplar summary measure, the toxicity index score, along with its applicability to dose-finding trials, is provided in Section 1.1 of the Supplementary Materials.
In this article, we present and utilise the normalised PRO-Adverse event (PRO-nAE) burden score, adapted from the Adverse Event (AE) burden score of Le-Rademacher et al. 38 to provide a single quantitative summary measure that captures both the frequency and severity of toxicities experienced by a patient. Although originally developed for clinician-assessed adverse events, we have adapted this approach for patient-assessed toxicities using instruments such as the PRO-CTCAE. We refer to this endpoint as the PRO-nAE burden score from henceforth. describes the severity of toxicity experienced by patient at dose at time .
For severity grades , we define the normalised PRO-nAE burden score, for patient at time as follows:
A discussion linking the AE burden score and Total Toxicity Profile score can be found in Section 1.2 of the Supplementary Materials.
In this article, we illustrate the use of PRO-nAE burden score with the PRO-CTCAE severity attribute, which mirrors the CTCAE, with values 0–4 corresponding to an observed toxicity with grade {‘None’, ‘Mild’, ‘Moderate’, ‘Severe’, ‘Very severe’}. Alternatively, the PRO-CTCAE composite score 40 or similar instruments measuring adverse event burden can also be considered.
In Figure 2 we present an exemplar PRO questionnaire completed by a patient. Supposing a patient is asked about three adverse events and reports two moderate toxicities and one severe toxicity, their PRO-nAE burden score is equal to .
Figure 2.

Exemplar short PRO questionnaire filled in by a patient.
2.2.3. Interpreting PRO-nAE burden score
One may view the PRO-nAE burden score as the normalised average severity of side effects experienced by a patient during the trial. We can translate the normalised PRO-nAE burden score to the average mean severity grade experience by a patient as per Table 1. For example, a PRO-nAE score of 0.25 indicates that the patient is experiencing a level of severity equivalent to mild across all assessed side effects. If a patient has a PRO-nAE burden score of 0.60, this indicates the patient is experiencing a level of severity equivalent to moderate to severe side effects across all assessed side effects. For trialists interested in the exact mean distribution of side-effect severity associated with a given PRO-nAE burden score, please refer to Section 2 of the supplementary materials which provides examples for PRO-nAE burden scores of 0.06, 0.18, 0.24, 0.39, 0.52, and 0.59. Other PRO-nAE scores can be translated in the same manner using the code provided in this article.
Table 1.
Relationship between average (mean) side effect severity experienced by a patient and PRO-nAE burden score.
| Average (mean) side effect severity experienced by a patient | None | Mild | Moderate | Severe | Very severe |
|---|---|---|---|---|---|
| Normalised PRO-nAE burden score | 0.00 | 0.25 | 0.50 | 0.75 | 1.00 |
2.3. Estimating PRO-nAE burden score using generalised linear mixed effect models
We model PRO-nAE burden score longitudinally across multiple treatment cycles or timepoints. A generalised linear mixed effect model is employed, defined by a set of covariates with fixed effects coefficients and random effects for some covariance matrix . The conditional expected response is linked to the linear predictors via a link function, , 41 such that
| (1) |
To predict the PRO-nAE burden score, we use a Bayesian mixed-effect beta regression model 42 with a logit link function and beta likelihood. This approach ensures that model predictions remain bounded in [0,1], aligning with the properties of the PRO-nAE burden score. 42 In this framework, we apply uninformative priors to the model parameters to allow the data to predominantly inform the posterior distribution. The model assumes,
| (2) |
and
| (3) |
Bounded PRO-nAE responses are observed for patient at time point given their dose . The conditional expected response is predicted by a patient’s dosage , the current time point and respective interaction . The parameter defines the precision of the data. The intra-correlation between repeated measures for individual patients is captured by the random effects .
The uninformative priors placed on each covariate are,
| (4) |
Trialists may also wish to consider utilising weakly informative priors to incorporate modest prior knowledge.
Aligning with the OPTIMISE-ROR guidance, 43 which identifies key PRO research objectives for early phase dose-finding oncology trials, PROs are analysed within PRO-ADD to inform final dose-selection decision-making. We utilise this model to estimate , the PRO-nAE burden score at the final time point for each dose . The longitudinal analysis of PROs strengthens the estimation of by modelling correlated repeated measures over time at a patient-level.
Prior research has demonstrated that PRO score can vary non-linearly over the course of treatment. 44 To account for this, the inclusion of the quadratic time covariate allows the model to capture plateauing of PRO-nAE burden score. Methodological standards for PRO-CTCAE usage in oncology trials indicates that most adverse events occur during the initial cycles of treatment, with symptom severity typically plateauing during longer term treatment administration as symptomatology stabilises. 45 Prospective longitudinal studies of patient-reported symptoms within oncology trials similarly identified plateauing symptom severity toward the later stages of they study period. 46 As with any model, it is important to evaluate the trade-off between a model’s predictive performance and its parsimony to ensure computational stability. We consider this model particularly suitable for larger, dose-optimisation trials where reliable estimation of longitudinal symptom trajectories can support informed dose selection.
2.4. Estimating preliminary efficacy signals using iPIPE
Many dose-finding trial designs which integrate preliminary efficacy into dose recommendations estimate the probability of a binary response outcome, often defined by conventional criteria such as Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST v1.1). 47 A common approach employs a simple beta-binomial conjugate model to estimate the probability of response at dose .
Research indicates that more dosage is not necessarily correlated with greater treatment efficacy. 48 Brock et al’s review of dose-finding oncology trials found that whilst some cancer treatments (particularly cytotoxic treatments) do exhibit monotonic dose-response relationships, plateauing efficacy beyond critical doses or consistent efficacy response across all doses are possible dose-response relationships within cancer trials. 48 The FDA’s Project Optimus initiative has recognised that for targeted drugs, higher doses beyond a critical dose may not increase efficacy response. 49 Whilst immunotherapies such as Ipilimumab have shown increasing dose-response relationships, 50 other immunotherepies such as anti-PD-1/PD-L1 therapies have alternatively exhibited flat dose-response relationships. 51 Similar paradigms have been observed for CAR-T cell therapies. 52 Under these schemas, we utilise iPIPE 53 (inverse product-of-independent-probability-escalation) to estimate the probability of response for each dose . iPIPE allows for increasing and plateauing dose-response relationships by constraining the beta-binomial conjugate model under the assumption that for .
PIPE was first presented to identify MTDs within a dose-escalation trial design. 54 Cheung et al. 53 extends this methodology such that, rather than classifying amongst doses to identify the MTD, the inverse PIPE-classifier (iPIPE) estimates response rates with an implicit monotonicity/plateauing assumption. 53
For data , we first define such that,
| (5) |
For and ,
| (6) |
where indicates some weight and represents some quantile of the posterior distribution to estimate. We set parameter . The probability of response for dose is estimated by,
| (7) |
To estimate the probability of binary response, we utilise the beta-binomial conjugate model with denoting total patients allocated to dose with responders. Prior and the posterior distribution given data ,
| (8) |
is used to determine In this article, we utilise the uninformative prior .
2.5. Posterior sampling with iPIPE
iPIPE gives us access to the inverse cumulative distribution function for the posterior distribution of . For fixed , we can determine the -quantile of the posterior distribution, that is for fixed , we can find such that . For example, setting provides a posterior median point estimate.
As we can evaluate the inverse cumulative distribution function for the posterior distribution of at each , we can utilise inverse transform sampling to sample posterior estimates . Thus to sample using iPIPE, we sample and calculate at this value of . We repeat this process multiple times to recover a large number of posterior samples of .
2.6. Dose-optimisation using a modular trial design
2.6.1. Illustrative example of a PRO-ADD design
Figure 3 illustrates an example of a PRO-ADD modular trial design. The design initially identifies the MTD based on clinician- reported DLTs encompassing unacceptable toxicities and treatment related death, followed by randomisation among admissible doses. Early stopping rules for futility or excessive toxicity are incorporated. Final dose selection is guided by evaluating the trade-off between patient-reported tolerability and preliminary efficacy outcomes. Our proposed trial utilises a Bayesian Optimal Interval (BOIN) 55 backbone to identify the dose with a probability of DLT closest to an elicited target DLT rate . BOIN designs 56 are increasingly popular model-assisted designs and designated fit-for-purpose for dose-finding trials by the FDA. 57 Whilst the traditional BOIN design is utilised for dose-escalation in the PRO-ADD illustrative example presented, other trial designs and BOIN extensions can be utilised and discussed in further detail in the discussion.
Figure 3.
An illustration of a PRO-ADD trial design.
Within PRO-ADD, repeated measure modelling is utilised to inform estimated PRO-nAE burden score as per Section 2.3 and iPIPE is employed to estimate the probability of a binary response endpoint as per Section 2.4.
2.6.2. Stage 1
Stage 1 follows exactly the dose-escalation routine of BOIN, 55 and is set forth as follows. Dose is escalated by evaluating the empirical DLT rate ( ) for dose , where
For pre-specified dose escalation ( ) and de-escalation ( ) boundaries, dose is escalated if and de-escalated if . If neither condition is satisfied, the next cohort of patients are allocated to dose .
At each dose allocation, the safety stopping rule is applied to remove doses with an unacceptably high probability of DLT. 58 For the true probability of DLT at dose , threshold , and maximal admissible DLT rate , the admissible dose is defined as the set of all doses as:
| (9) |
is calculated using a beta-binomial model with Beta(0.5, 0.5) prior and . If dose is deemed unsafe under this stopping rule, all doses higher than are also eliminated.
Patients are escalated toward the MTD as per Stage 1 of the trial design until the same dose is recommended consecutively for a pre-specified number of cohorts or the first futility interim analysis occurs at the enrolment of a specific cohort.
2.6.3. Stage 2
At Stage 2, we introduce an additional futility stopping rule to remove dose(s) which have shown inadequate activity based on the complete response data collected up to that point. The admissible dose set must not only be safe but also show sufficient activity, and is defined as:
| (10) |
where is the complete response data collected up to time point , is the futility threshold, the minimally admissible efficacy rate, and the true probability of response at dose . The futility stopping rule uses a Beta(0.1, 0.9) prior and .
For a null response rate of 10% (deemed futile) and an alternative response rate of 25% (deemed effective), the futility stopping rule protects Type I error rate at 20% and power at 74%. 59 For trialists wishing to consider other criterion, PRO-ADD provides trialists flexibility to consider other priors aligning with their own criteria for futility.
Note that a patient with partial response data are excluded from futility stopping rule decision making. Thus a dose stopped for futility may be reinstated later if the partial data, once fully collected, suggests that the dose may in fact be efficacious. To ensure sufficient sample size for the assessment of futility, the stopping rule is evaluated only twice within the trial – once we have at least six cohorts of complete response data and only assessing doses with at least six patients on treatment. This approach allows for a simpler futility rule which avoids the need for methods that incorporate partial data into the futility decision making. Futility stopping rules incorporating partial data are particularly valuable when the futility rule is evaluated at each patient enrolment. For trialists who wish to utilise partial response data, a review of methods is provided by Zhou et al. 60
Within Stage 2, patients are allocated to doses using the following criteria:
For the largest investigated dose, if then escalate to dose . If all doses have been investigated or , continue to step 2.
Evaluate . If , end the trial early. Otherwise randomise the next cohort of patients to an admissible dose inversely proportional to the number of patients previously assigned to each admissible dose.
Repeat accordingly until the pre-specified sample size has been reached.
2.6.4. Termination of the trial
The trial is terminated once a maximum trial size is reached or if no doses are deemed admissible. The final admissible set contains all admissible doses at most as large as the MTD as identified by the BOIN design using isotonic regression. 55
2.6.5. Final dose recommendation
At this point, the longitudinal PRO data are utilised within this PRO-ADD design. The final recommended dose is assessed by quantifying the trade-off between dose response and PRO tolerability data using a loss function amongst admissible doses . Specifically, we wish to quantify the trade-off between PRO-nAE burden score at the final assessment time point ( ) and probability of response at each dose ( ). This approach allows flexibility to decide how to weight benefit versus tolerability.
In PRO-ADD we investigate -Loss,
| (11) |
where is the PRO-nAE burden score at the final PRO assessment time point. This approach produces a single summary measure to guide the selection of the optimal dose.
2.6.6. Quantifying uncertainty and final dose selection
Due to the small sample sizes of early phase dose-finding trials, it is to be expected that estimates of response and PRO-nAE burden scores are subject to significant uncertainty. Bayesian estimation valuably provides posterior distributions for parameters, capturing the uncertainty of response and PRO-nAE burden score posterior estimates. We recommend that we take advantage of this property by evaluating the expected (empirical) loss for each dose .
We estimate PRO-nAE burden score at the final assessment timepoint. Given the posterior distribution for the PRO-nAE burden score and probability of response using iPIPE, we use Monte Carlo to compute the expected -Loss by simulating samples from their respective posterior distributions: and .
Figure 4 shows potential samples and from posterior distributions of PRO-nAE burden score and preliminary efficacy respectively with a sample size of 60 patients.
Figure 4.
1000 sampled points and densities of the posterior distribution of iPIPE efficacy estimates and estimated PRO-nAE burden score for 5 doses under Scenario 5 of the simulation study, with true preliminary efficacy rates and PRO-nAE scores marked with a star. The unacceptable loss region (with a loss greater than 0.9) is coloured in dark blue.
Using these samples, the empirical loss for each dose is defined as,
| (12) |
2.6.7. Optimal dose
The optimal dose recommendation is the admissible dose which minimises the -Loss,
| (13) |
We recommend dose for investigation in a Phase II trial if the expected loss is smallest amongst admissible doses and lies below a maximum loss threshold, . If no dose has an estimated loss below , the trial does not recommend any dose for further testing.
This decision rule prioritises the three endpoints by first identifying admissible doses using DLT data only, before choosing the optimal dose using preliminary efficacy and patient-reported tolerability endpoints. This ensures any recommended dose is suitable from a safety perspective first, before it is then evaluated in terms of its preliminary efficacy and tolerability.
2.6.8. Choice of loss function and eliciting the acceptable loss
Clinically, defines the worst tolerability-efficacy trade-off a clinician would be willing to accept when recommending a dose. For a given loss value, any combination of preliminary efficacy and PRO-nAE burden scores that yields the same loss is considered equivalent. For example, by utilising a -Loss as per equation (11), a dose with a preliminary efficacy rate of 0.3 and no PRO-nAE burden score is equivalent to another dose with preliminary efficacy rate of 0.38 and burden score of 0.5 (indicating on average each side effect reported by a patient is deemed of moderate severity). Both doses in this scenario have a loss of 0.8 and are deemed equivalent. As such, it is important that a suitable loss function is utilised so that doses with numerically equivalent loss are also deemed clinically equivalent. Such frameworks are common in many loss- or utility-based dose-finding trial designs5,60,61 and considerations to support trialists to identify suitable loss functions are provided in Section 5.1.
Once a loss function has been chosen, the choice of can be implicitly determined by this loss and the selection of a minimally acceptable combination of preliminary efficacy rate and PRO-nAE burden score. For simplicity, we may ask “If a patient has no side effects whatsoever, what would be the minimal preliminary efficacy rate, , which you would consider acceptable for a dose?”. With the answer to this question and defined loss function, we can now define . Under -Loss, . As the loss function is constructed such that any dose with a numerically equivalent loss function is also clinically equivalent, this elicitation approach defines a minimally acceptable loss across all possible preliminary efficacy rates.
2.6.9. Acceptable dose(s)
Trialists may also wish to identify a set of acceptable doses. The set of acceptable doses are such that for acceptable loss ,
In the subsequent simulation study, we set an acceptable dose as one with a loss of no more than 0.90 (equivalent to a 10% response rate and 0 PRO-nAE burden score). This loss region is highlighted in Figure 4.
3. Simulation study
3.1. Case study
Our simulation study is motivated by KEYNOTE-001, a first-in-human Phase I dose-escalation study of Pembrolizumab in patients with advanced solid tumours investigating toxicity and activity across three treatment doses (ClinicalTrials.gov identifier: NCT01295827). 62
Whilst KEYNOTE-001 was initially designed as a dose-finding trial using the design, generally well-tolerated doses and promising anti-tumour activity led to adaptations of the trial during interim analyses, including the introduction of multiple expansion cohorts. 63 With the addition of multiple expansion cohorts, recruitment for the KEYNOTE-001 study ended in July 2014, with the treatment of 1,235 patients. As motivation for our trial, we specifically focus on the initial dose-finding trial originally executed in KEYNOTE-001. In the initial dose-finding trial, 32 patients were enrolled onto three doses until unacceptable toxicity or disease progression occurred. Primary objectives of the study were to evaluate the safety, pharmacokinetics, and pharmacodynamics of Pembrolizumab. Identification of the MTD of Pembrolizumab was an additional objective. 62 Whilst preliminary efficacy was an un-powered objective, objective response rate (ORR) was defined as the proportion of patients who demonstrated a partial response or complete response 51 as per Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST v1.1). 47 Patients were sequentially enrolled in cohorts of three to doses administered on days 1, 28 and every 14 days thereafter. DLTs were assessed during the first cycle of treatment (28 days) and tumour response was assessed every two months for the first year.
Subsequently published case reports on immune-related adverse events associated with Pembrolizumab indicated that adverse events including severe colitis, pneumonitis, and organ damage appeared on average of 5 to 15 weeks after commencement of treatment. 64 This reiterates the motivation for the assessment of adverse events beyond clinician DLT assessment windows. Follow up findings published three years after the KEYNOTE-001 study indicated that there was no association between the dose of pembrolizumab (2 mg/kg or 10 mg/kg every 3 weeks or 10 mg/kg every 2 weeks) and the activity nor toxicity of the immunotherapy. 51
3.2. Data generation
3.2.1. Binary endpoints
Binary clinician-assessed DLT observation and response outcomes are sampled from a Bernoulli distribution.
3.2.2. PRO-nAE burden score
The synthesis of PRO-CTCAE data has thus far has been limited to the simulation of a binary Patient-DLT. As such, within this article we take particular care with the simulation of the continuous PRO-nAE burden score.
The normalised adverse event score is bounded on [0,1], naturally lending itself to simulation via Beta sampling. However, to assess whether the distributional properties of the AE burden score abides by such a data-generating scheme, AE scores were first generated for each patient toxicity-by-toxicity using a non-parametric data generating schema. The synthesis of PRO-nAE data is motivated by a dataset of 219 patient’s PRO-CTCAE scores with advanced solid tumours who were enrolled on Phase I clinical trials at Princess Margaret Cancer Centre in Canada from 1 May 2017 to 1 January 2019. 65
The synthesised PRO-CTCAE data was generated under the following assumptions:
Dose-dependent toxicity: For a toxicity type , the probability of observing a more severe toxicity increases as dose level increases.
Time-dependent toxicity: For a toxicity type , the probability of observing a more severe toxicity increases as the time a patient remains on treatment increases.
Correlation between toxicities: For each patient , the severity of two or more toxicities may be correlated. If so, this correlation is assumed to remain constant over time.
A patient’s PRO-nAE burden score was synthesised non-parametrically using latent Beta random variables. Data generation therefore relied on a dimension matrix defining how a patient’s dosage ( ) and treatment cycle ( ) interplays with the grade ( ) of toxicity ( ) they experience. More details on this non-parametric data generating approach are presented in Section 3 of the Supplementary materials.
Following this investigation, we conclude that the simulated PRO-nAE burden scores can be appropriately sampled using a Beta distribution. Thus, for ease of computation, PRO-nAE scores are sampled from a Beta distribution, with shape and rate parameters determined via the maximum likelihood estimation of the non-parametric data generation using package ‘ ’ in R. 66
Correlation between binary clinician-DLT and the repeated continuous PRO-nAE scores for each patient is induced using a Gaussian copula. Figure 5 demonstrates how DLT is correlated with PRO-nAE burden score within this simulation study. 67 Whilst evidence indicates that DLT occurrence is more likely as dose increases, research suggests the probability of response may not increase similarly. 48 As such, in the main manuscript we suppose DLT and efficacy response are independent. We subsequently investigate PRO-ADD performance when DLT and efficacy responses are mildly correlated in sensitivity analyses.
Figure 5.
Simulated PRO-nAE burden scores for Scenario 1 with observed DLT and no DLT across 5 doses.
3.3. Fixed simulation scenario trial parameters
In the following simulation study, sample size is fixed at 60 patients, enrolled in cohorts of 3. We evaluate the first two tumour responses at week 8 and week 16. Patients are sequentially enrolled every 4 weeks. Whilst PROs were not collected in the original KEYNOTE-001 study, we schedule PRO collection every two weeks to align with current literature which recommends regular (weekly or fortnightly) PRO data collection over initial cycles of treatment to accurately capture the high number of adverse events likely to occur at the start of a trial. 45 Figure 6 highlights this schedule.
Figure 6.
Scheduling timetable for the simulation study, inspired by the KEYNOTE-001 dose-finding trial. 62
Patients are escalated toward the MTD as per Stage 1 of the trial design until the same dose is recommended consecutively for two cohorts or the first futility interim analysis takes place at the enrolment of the eleventh cohort. A futility stopping rule is used to remove doses with insufficient response rate at the end of the tenth cohort and sixteenth cohort (using the complete response data for cohorts 1–6 and cohorts 1–12, respectively).
The target DLT rate deemed admissible is 0.25. Dose (de-)escalation boundaries and are defined as 0.197 and 0.298, respectively, as per the BOIN shiny app implemented on the www.trialdesign.org platform. 68
3.4. Simulation scenarios
Eight relationships between PRO-nAE burden score and preliminary efficacy are investigated in this simulation study. These scenarios are labelled in Table 2. Of particular note, Scenario 7 explores a unimodal dose-efficacy curve and Scenario 8 explores a setting where two doses have an equivalent loss (both with a loss of 0.75). iPIPE is only an appropriate analysis approach to analyse preliminary efficacy if the dose-efficacy relationship is monotonic or plateauing, however we present Scenario 7 as a sensitivity analysis to illustrate PRO-ADD’s performance in the unlikely event that a treatment presumed to have a monotonic or plateauing dose-efficacy relationship instead exhibits a unimodal relationship. Should trialists suspect a treatment has a unimodal dose-efficacy relationship, an alternate analysis approach (such as a beta-binomial model) should be alternatively used.
Table 2.
PRO-nAE burden score and efficacy simulation scenarios investigated within this simulation study.
| PRO-nAE score | TIME TREND (across cycle) | Increasing | Plateauing | Increasing | |||||
| DOSE TREND (across dose) | Increasing | Plateauing | Increasing | Increasing | |||||
| Efficacy | DOSE TREND (across dose) | Increasing | Plateauing | Increasing | Plateauing | Increasing | Plateauing | Unimodal | Plateauing |
| Scenario 1 | Scenario 2 | Scenario 3 | Scenario 4 | Scenario 5 | Scenario 6 | Scenario 7 | Scenario 8 | ||
Two potential tolerability scenarios for PRO-nAE burden score are investigated. We consider the monotonic increasing of PRO-nAE burden score over treatment administration, which may reflect the growing symptom burden and accumulating moderate toxicities occurring over extended treatment windows (Scenarios 1–4 and 7–8). Secondly, we consider a scenario where a patient’s symptomatology stabilises over treatment administration. This is characterised by a plateau in observed PRO-nAE burden score over treatment cycles (Scenarios 5–6). 45
An example of the time trends to be investigated in this simulation study are presented in Figure 7. Section 4 of the Supplementary materials graphically displays simulation scenarios 1–6.
Figure 7.
The PRO-nAE burden score time trends investigated in the simulation study. (a) Monotonically increasing time trend and (b) Plateauing time trend.
These simulation scenarios were investigated under two MTD scenarios, the first where dose 5 is the MTD and another where dose 3 is the MTD. The DLT rates for both scenarios are shown below, with MTD indicated in bold:
DLT Scenario M5: 0.01, 0.05, 0.10, 0.15, 0.20,
DLT Scenario M3: 0.06, 0.13, 0.25, 0.40, 0.50.
3.5. Comparison to U-BOIN design
In the subsequent simulation study, we compare PRO-ADD to U-BOIN. 60 Similarly to PRO-ADD, U-BOIN is a two stage design – firstly identifying the MTD using DLT data alone before incorporating an efficacy endpoint into decision-making to identify the optimal biological dose. What’s more, like PRO-ADD, U-BOIN utilises a trade-off framework. Importantly, in U-BOIN’s case, decision making is guided by evaluating the utility of each dose based on its DLT and response rate, without incorporating patient-reported tolerability measures. Whilst U-BOIN can take ordinal toxicity and response endpoints, we compare PRO-ADD to the design with binary DLT and efficacy responses. U-BOIN simulation performance is assessed using the online web app www.trialdesign.org. The exact inputs used to run the U-BOIN simulation study are provided in Section 5 of the Supplementary materials. In line with current practice, to ensure fair comparison between U-BOIN and PRO-ADD, in the main manuscript we evaluate PRO-ADD performance supposing complete data collection. As a sensitivity analysis, we evaluate the performance of PRO-ADD in light of intercurrent events.
3.6. Investigated sensitivity analyses
We explore many sensitivity analyses for PRO-ADD in Section 6 and 7 of the Supplementary Materials. Investigations include,
Performance of PRO-ADD when dose 1 is the MTD or no doses are safe.
Performance of PRO-ADD when DLT and efficacy response are mildly correlated on a patient level. In this instance, patient’s responses are correlated using a Gaussian copula with covariance 0.15.
Performance of PRO-ADD with a different acceptable loss to results presented in the main manuscript, with .
Performance of PRO-ADD in light of intercurrent events including dose discontinuation due to DLT and death unrelated to treatment by utilising a hypothetical handling strategy. 69 Performance is evaluated both when patient DLT and activity responses are independent of one another or mildly correlated.
4. Results
For PRO-ADD, we present the probability of selecting the single optimal dose (i.e. the dose with the smallest loss amongst admissible doses) and the probability of identifying an acceptable dose based on -Loss (i.e. correctly determining whether a dose has a loss below a predefined maximum acceptable threshold). Note that more than one dose can be classified as acceptable within each trial simulation. In this simulation study, we use a maximum acceptable loss of 0.9 which, for example, could be achieved by a dose with a PRO-nAE burden score of 0 and a probability of efficacy of 0.10. All other equivalent combinations of PRO-nAE burden score and efficacy are implicitly defined by the choice of -Loss. Additional details on the severity of toxicities associated with each PRO-nAE burden score are provided in Section 2 of the Supplementary materials. For U-BOIN, we present the probability of selecting the single optimal dose. In the following simulation study, for each of their trade-off approaches, PRO-ADD and U-BOIN recommend the same dose as the optimal dose. The utilities associated with each dose for U-BOIN is presented in Table S2 of the Supplementary materials. The mean number of patients allocated to treatment within Stage 1 and Stage 2 of PRO-ADD is 23 and 37, respectively.
Table 3 shows PRO-ADD and U-BOIN performance across eight scenarios when the MTD is dose 5 and all doses are admissible for safety. The probability of correct selection of the optimal dose for PRO-ADD ranges from 43% to 77%. The optimal dose is also recognised as an acceptable dose (with a loss of at most 0.9) 71%–96% of the time. PRO-ADD performs better or approximately as well as U-BOIN in all scenarios. PRO-ADD most significantly improves upon U-BOIN performance (by at least 18%) in Scenarios 2, 4, 6, 7, and 8 where efficacy plateaus. In these scenarios, the inclusion of PRO-nAE burden score successfully constrains the decision criterion to choose the most effective dose with smallest tolerability burden. In these scenarios, U-BOIN often identifies higher, more intolerable doses with no improved efficacy. In Scenario 3 where doses 3–5 have equal PRO-nAE burden score, utilisation of iPIPE estimation of efficacy ensures PRO-ADD correctly identifies dose 5 as the best dose 13% more often than U-BOIN, which relies on a beta-binomial model to estimate response rate.
Table 3.
Proportion of optimal dose recommendations (bold) for each dose level under 8 scenarios for 5,000 simulated trials dose 5 is the MTD (Scenario M5), with probability no dose selected also indicated for PRO-ADD and U-BOIN.
| Dose 1 | Dose 2 | Dose 3 | Dose 4 | Dose 5 | No dose | ||
|---|---|---|---|---|---|---|---|
| DLT Scenario: M5 | 0.01 | 0.05 | 0.10 | 0.15 | 0.20 | ||
| Scenario 1 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.24 | 0.42 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.86 | 0.78 | 0.89 | ||
| PRO-ADD | Optimal recommendation | 0.02 | 0.06 | 0.32 | 0.50 | 0.07 | |
| Acceptable recommendation | 0.17 | 0.33 | 0.72 | 0.85 | 0.46 | 0.03 | |
| U-BOIN | Optimal recommendation | 0.00 | 0.01 | 0.14 | 0.49 | 0.36 | 0.00 |
| Scenario 2 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.42 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.91 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.03 | 0.74 | 0.19 | 0.02 | |
| Acceptable recommendation | 0.1 | 0.34 | 0.96 | 0.90 | 0.49 | 0.01 | |
| U-BOIN | Optimal recommendation | 0.00 | 0.01 | 0.48 | 0.31 | 0.19 | 0.00 |
| Scenario 3 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.39 | 0.39 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.28 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.83 | 0.82 | 0.68 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.04 | 0.11 | 0.22 | 0.60 | |
| Acceptable recommendation | 0.17 | 0.32 | 0.76 | 0.84 | 0.71 | 0.03 | |
| U-BOIN | Optimal recommendation | 0.01 | 0.02 | 0.29 | 0.21 | 0.47 | 0.00 |
| Scenario 4 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.69 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.27 | 0.27 | ||
| Loss | 0.97 | 0.95 | 0.83 | 1.01 | 1.01 | ||
| PRO-ADD | Optimal recommendation | 0.05 | 0.14 | 0.60 | 0.00 | 0.04 | |
| Acceptable recommendation | 0.18 | 0.32 | 0.75 | 0.05 | 0.11 | 0.16 | |
| U-BOIN | Optimal recommendation | 0.02 | 0.04 | 0.42 | 0.32 | 0.21 | 0.00 |
| Scenario 5 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.42 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.83 | 0.78 | 0.89 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.05 | 0.41 | 0.43 | 0.06 | |
| Acceptable recommendation | 0.15 | 0.32 | 0.76 | 0.82 | 0.39 | 0.04 | |
| U-BOIN | Optimal recommendation | 0.01 | 0.01 | 0.19 | 0.47 | 0.32 | 0.00 |
| Scenario 6 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.42 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.91 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.04 | 0.76 | 0.16 | 0.02 | |
| Acceptable recommendation | 0.16 | 0.34 | 0.95 | 0.88 | 0.39 | 0.01 | |
| U-BOIN | Optimal recommendation | 0.00 | 0.01 | 0.48 | 0.32 | 0.19 | 0.00 |
| Scenario 7 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.37 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.94 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.03 | 0.77 | 0.17 | 0.01 | |
| Acceptable recommendation | 0.19 | 0.34 | 0.96 | 0.90 | 0.39 | 0.01 | |
| U-BOIN | Optimal recommendation | 0.00 | 0.01 | 0.52 | 0.34 | 0.13 | 0.00 |
| Scenario 8 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.27 | 0.29 | 0.29 | 0.29 | 0.29 | ||
| Loss | 0.75 | 0.75 | 0.81 | 0.88 | 0.99 | ||
| PRO-ADD | Optimal recommendation | 0.52 | 0.32 | 0.11 | 0.05 | 0.00 | |
| Acceptable recommendation | 0.92 | 0.93 | 0.92 | 0.73 | 0.15 | 0.00 | |
| U-BOIN | Optimal recommendation | 0.30 | 0.29 | 0.21 | 0.13 | 0.08 | 0.00 |
Note: Acceptable doses are highlighted in green and have a maximum loss of 0.90. Individual patient DLT and activity responses are simulated to be independent and full patient data are collected.
Scenario 1 highlights the strength of PRO-ADD. Whilst all doses are considered safe, patient tolerability to treatment decreases as dosage increases. However, the treatment’s efficacy begins to level off at dose 4, with a marginal improvement in the probability of response between doses 4 and 5 of 2%. Traditional trial designs that target the MTD would typically recommend dose 5 as the recommended Phase II dose, with U-BOIN identifying dose 5 as best 36% of the time. However, with the addition of the PRO-nAE endpoint, the PRO-ADD design effectively determines that dose 4 is optimal using our pre-specified -Loss criterion. For this loss, the marginal improvement in response exhibited by dose 5 does not justify its increased tolerability burden. As such, PRO-ADD successfully recognises that dose 5 has an increased tolerability burden but does not offer patient’s any additional, meaningful efficacy benefit.
What’s more, in Scenario 2 PRO-ADD reliably identifies that dose 3 is the optimal dose approximately 74% of the time. As such, PRO-ADD successfully concludes that doses 4 and 5 have an increased tolerability burden but do not offer patient’s any additional efficacy benefit. PRO-ADD does well estimating the PRO-nAE burden score when burden score monotonically increases across cycles (Scenarios 1–4) and when PRO-nAE burden score plateaus beyond a certain cycle of treatment (Scenarios 5–8).
Whilst iPIPE is not a suitable analysis method for unimodal dose-efficacy relationships, PRO-ADD can still perform well in this instance as per Scenario 7. In this case, though iPIPE would estimate the efficacy rate for dose 5 to be at least that of dose 4, the increased PRO-nAE burden score associated with dose 5 ensures that this dose is not recognised as optimal. Whilst U-BOIN utilises a beta-binomial model to estimate preliminary efficacy rate, PRO-ADD still performs 25% better than U-BOIN.
In Scenario 8, we regard doses 1 and 2 as equally optimal. PRO-ADD selects doses 1 and 2 84% of the time, an improvement of 25% over U-BOIN.
Table 4 demonstrates PRO-ADD and U-BOIN performance when dose 3 is the MTD across scenarios 1–8. Dose 3 is recognised as the optimal dose between 44%–69% of the time. Across all scenarios, at most 9% of trials recommend a dose higher than the true MTD and optimal dose – showcasing the designs effective safety overdosing control. PRO-ADD once again performs equally well or better than U-BOIN for each scenario.
Table 4.
Proportion of optimal dose recommendations (bold) for each dose level under 8 scenarios for 5,000 simulated trials dose 3 is the MTD (Scenario M3), with probability no dose selected also indicated for PRO-ADD and U-BOIN.
| Dose 1 | Dose 2 | Dose 3 | Dose 4 | Dose 5 | No dose | ||
|---|---|---|---|---|---|---|---|
| DLT Scenario: M3 | 0.01 | 0.05 | 0.10 | 0.15 | 0.20 | ||
| Scenario 1 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.24 | 0.42 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.86 | 0.78 | 0.89 | ||
| PRO-ADD | Optimal recommendation | 0.03 | 0.11 | 0.44 | 0.09 | 0.00 | |
| Acceptable recommendation | 0.10 | 0.25 | 0.55 | 0.13 | 0.01 | 0.33 | |
| U-BOIN | Optimal recommendation | 0.08 | 0.12 | 0.42 | 0.16 | 0.07 | 0.15 |
| Scenario 2 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.42 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.91 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.06 | 0.70 | 0.03 | 0.00 | |
| Acceptable recommendation | 0.11 | 0.27 | 0.75 | 0.15 | 0.01 | 0.20 | |
| U-BOIN | Optimal recommendation | 0.05 | 0.07 | 0.60 | 0.11 | 0.04 | 0.13 |
| Scenario 3 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.39 | 0.39 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.28 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.83 | 0.82 | 0.68 | ||
| PRO-ADD | Optimal recommendation | 0.03 | 0.11 | 0.47 | 0.09 | 0.01 | |
| Acceptable recommendation | 0.10 | 0.25 | 0.59 | 0.13 | 0.01 | 0.30 | |
| U-BOIN | Optimal recommendation | 0.08 | 0.11 | 0.48 | 0.13 | 0.06 | 0.14 |
| Scenario 4 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.69 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.27 | 0.27 | ||
| Loss | 0.97 | 0.95 | 0.83 | 1.01 | 1.01 | ||
| PRO-ADD | Optimal recommendation | 0.04 | 0.12 | 0.53 | 0.00 | 0.00 | |
| Acceptable recommendation | 0.11 | 0.25 | 0.59 | 0.01 | 0.00 | 0.32 | |
| U-BOIN | Optimal recommendation | 0.07 | 0.11 | 0.51 | 0.12 | 0.05 | 0.14 |
| Scenario 5 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.27 | 0.42 | 0.44 | ||
| Loss | 0.97 | 0.95 | 0.83 | 0.78 | 0.89 | ||
| PRO-ADD | Optimal recommendation | 0.03 | 0.10 | 0.49 | 0.08 | 0.00 | |
| Acceptable recommendation | 0.09 | 0.23 | 0.58 | 0.14 | 0.01 | 0.31 | |
| U-BOIN | Optimal recommendation | 0.08 | 0.10 | 0.46 | 0.17 | 0.07 | 0.13 |
| Scenario 6 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.42 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.91 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.05 | 0.69 | 0.03 | 0.00 | |
| Acceptable recommendation | 0.09 | 0.25 | 0.73 | 0.14 | 0.01 | 0.21 | |
| U-BOIN | Optimal recommendation | 0.04 | 0.07 | 0.62 | 0.12 | 0.05 | 0.11 |
| Scenario 7 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.05 | 0.08 | 0.42 | 0.42 | 0.37 | ||
| Loss | 0.97 | 0.95 | 0.70 | 0.78 | 0.94 | ||
| PRO-ADD | Optimal recommendation | 0.01 | 0.06 | 0.69 | 0.03 | 0.00 | |
| Acceptable recommendation | 0.11 | 0.27 | 0.73 | 0.14 | 0.01 | 0.21 | |
| U-BOIN | Optimal recommendation | 0.05 | 0.07 | 0.60 | 0.12 | 0.04 | 0.12 |
| Scenario 8 | PRO-nAE burden score | 0.18 | 0.24 | 0.39 | 0.52 | 0.69 | |
| Probability of efficacy | 0.27 | 0.29 | 0.29 | 0.29 | 0.29 | ||
| Loss | 0.75 | 0.75 | 0.81 | 0.88 | 0.99 | ||
| PRO-ADD | Optimal recommendation | 0.53 | 0.35 | 0.11 | 0.01 | 0.00 | |
| Acceptable recommendation | 0.93 | 0.95 | 0.72 | 0.12 | 0.00 | 0.01 | |
| U-BOIN | Optimal recommendation | 0.46 | 0.35 | 0.13 | 0.03 | 0.02 | 0.02 |
Note: Acceptable doses are highlighted in green and have a maximum loss of 0.90. Inadmissible doses with a DLT rate above the target are highlighted in red. Individual patient DLT and activity responses are simulated to be independent and full patient data are collected.
A trial can fail to identify an optimal dose for two reasons. Firstly, the trial may end prematurely as futility and safety stopping rules are activated and no doses are deemed safe or efficacious. Alternatively, the trial may reach completion without identifying any acceptable doses (i.e., no dose has an estimated loss below 0.9). In Tables S3 and S4 of the Supplementary materials we breakdown the reason for failure to identify an optimal dose into these two factors for PRO-ADD. A larger proportion of trials are halted when dose 3 is the MTD, driven by an increased likelihood of stopping for safety. What’s more, under the M3 scenario, it is more difficult to identify that the loss of dose 3 is below 0.9. The optimal dose’s loss under M3 ranges from 0.70–0.86 rather than 0.68–0.83 under scenario M5.
The safety and futility rules remove inadmissible doses throughout the trial. For PRO-ADD, the mean number of patients allocated to each dose for MTD scenario M5 is (11.16, 11.47, 13.92, 13.26, 12.32). The mean number of patients allocated to each dose for MTD scenario M3 is (15.40, 16.11, 18.64, 11.70, 8.63). Table S5 in the Supplementary materials details PRO-ADD patient allocation for each dose and each simulation scenario. For both MTD scenarios, the safety and futility stopping rules ensure fewer patients are allocated to futile doses (doses 1 and 2 in Scenario M5) and to unsafe doses (doses 4 and 5 in scenario M3). For Scenario 1 where dose 5 is the MTD, the first futility stopping rule identifies at least one dose as futile 35.6% of the time, increasing to 81.5% and 87.1% at the second and final futility assessment at the end of the trial respectively. Specific detail on the proportion of trials which recognise each dose as futile for Scenario 1 is presented in Table S7 in the Supplementary Materials. Further discussion on sensitivity to futility stopping rules is presented in Section 6 of the Supplementary Materials. Section 7 of the Supplementary Materials presents results for other sensitivity analyses referred to in Section 3.6.
5. Discussion
We present the novel PRO-ADD dose-finding trial design – a modular framework for dose-optimisation trials. This innovative approach integrates three key outcomes: clinician-assessed DLTs, patient-assessed tolerability, and efficacy to determine the optimal dose. PRO-ADD dynamically adjusts dosing based on clinician-assessed DLTs and efficacy, and subsequently incorporates PRO-nAE burden score at the final analysis to recommend the most appropriate dose. The inclusion of patient-assessed tolerability provides PRO-ADD with superior probability of correct selection compared to the U-BOIN design across the majority of simulation scenarios. To our knowledge, PRO-ADD is the first trial design to introduce these three critical endpoints, including the longitudinal assessment of PROs, within a dose-optimisation design, offering a comprehensive approach to determining the optimal dose.
5.1. Practical design considerations
5.1.1. Selection of final PRO assessment timepoint
When modelling longitudinal PRO data with PRO-ADD, trialists should carefully consider the timing of PRO assessments, including the selection of the final analysis timepoint. In line with OPTIMISE-ROR guidance, 43 timing should be defined at the commencement of the trial in collaboration with stakeholders including clinicians, statisticians, PRO methodologists, and patient partners. Selection of the final assessment time point should be informed by a range of considerations including the treatment’s tolerability profile, mechanism of action, and investigated administration schedules. 43 For treatments such as immunotherapies and MTAs1,9 which can produce low-grade toxicities that accumulate over time, extended tolerability assessments beyond the initial treatment cycles may be warranted. PRO-ADD provides trialists with the flexibility to select a final assessment timepoint that aligns with the requirements of the investigational treatment.
5.1.2. Selecting loss functions and acceptable losses
Selection of an appropriate loss function is essential to ensure numerical loss translates to a clinical equivalence. Practical considerations for elicitation of utility functions have been provided for dose-finding trials and can be similarly utilised for trials incorporating loss functions. 70 In particular, the elicitation of a loss function relies on identifying three points that are judged to be equally desirable, with this approach effectively utilised in practice. 71 In our PRO-ADD example using the PRO-nAE burden score, we may elicit two of these points simply by asking “Supposing a patient experiences no side effects, what is the minimal probability of preliminary efficacy which would be acceptable for a dose?” and “Supposing efficacy is guaranteed, what is the maximal PRO-nAE burden score which you would deem acceptable for a given dose?”. The final point to be elicited must be a compromise of both PRO-nAE burden score and probability of preliminary efficacy, with equal attractiveness to the other elicited points. To aid such decision making, the mean distribution of side effect severity for varying PRO-nAE burden scores may be presented to stakeholders, translating PRO-nAE burden score to a dose’s average tolerability profile (see Figure S2 of the Supplementary Materials). Loss functions should be elicited in collaboration with clinical teams and patients to ensure that decision making reflects both clinical and patient priorities.
5.1.3. PRO-ADD in the presence of intercurrent events
Within the supplementary materials of this manuscript, we investigate PRO-ADD design performance considering two intercurrent events though other intercurrent events may occur in practice. To maintain design performance in light of trial-specific intercurrent events which may occur, trialists are encouraged to identify and select handling strategies for relevant intercurrent events upfront and evaluate design performance in light of such approaches. General guidance for the handling of intercurrent events within dose-finding and dose-optimisation trials should also be followed.69,72,73
5.2. Additional design considerations
The modular framework of this proposed design allows trialists to tailor design characteristics to their individual needs.
5.2.1. Stages 1 and 2
To complete the dose-escalation and optimisation routines in Stages 1 and 2 of PRO-ADD, we may consider any dose-escalation/optimisation trial design including extensions to the original BOIN design. Since its original publication, BOIN has been extended to introduce efficacy endpoints within interim and final decision making, including BOIN12, 61 BOIN-ET 74 and U-BOIN. 60 BOIN12 and BOIN-ET consider binary DLT and binary efficacy responses within a one and two stage design respectively. U-BOIN extends toxicity and efficacy to categorical responses with a utility function introduced to inform interim and final dose decision making.
The PRO-ADD trial design was developed in line with recommendations for the incorporation of PROs within early phase trials13,43 which recommends that PROs be integrated at final analyses to inform dose-selection. Whilst ePROs are an effective way to capture PROs, 28 paper-based PROs are still common. The extended administrative process and data collection period for paper-based PROs may reduce the current potential for PROs to be effectively used in adaptive decision-making. However, as ePROs collection becomes increasingly common, future work may wish to introduce PROs within adaptive, interim decision making. 13 For example, a trialist may also wish to introduce the continuous PRO-nAE burden score within the defining of admissible doses in Stage 2 of PRO-ADD. By utilising conjugacy models for continuous data, a tolerability stopping rule can be introduced similarly to the futility and safety stopping rules presented in this paper.
5.2.2. Final dose recommendation
The decision criterion defined in this article utilises a loss to make a final dose recommendation. Loss functions are implicitly embedded within the majority of early phase trials. For example, the -Loss is utilised for decision making in single outcome trial designs such as the Continual Reassessment Method (CRM) 75 and for other trial designs assessing joint outcomes.76,77 Whilst -Loss has been investigated in this simulation study, trialists have flexibility to choose their own loss and acceptable loss regions for dosing decisions. This can include a tailoring of the trade-off between efficacy estimates and PRO-nAE scores, which are equally weighted in our simulation study. For trialists who may value one endpoint more than another, a weighted -Loss function could be utilised to better reflect different clinical priorities. This could also include the use of other loss functions, including the Huber loss function, 78 shown to perform well with the CRM design. Other methods to assess the benefit-risk trade-off between doses also include the elicitation of utility, which has been implemented within trial designs utilising both continuous 79 and ordinal 60 endpoints. 80
By utilising a loss for the determination of the recommended dose, PRO-ADD performance depends on good predictive accuracy for PRO-nAE burden score and response rate for each dose. To ensure good convergence of the linear mixed effect model for PRO-nAE burden score, the presented PRO-ADD formulation is recommended for larger dose-optimisation trials. Trialists may wish to draw on existing knowledge of dose–response and toxicity relationships for the novel therapy to inform model building. Without such prior knowledge, the PRO-ADD model presented here provides a generic framework for estimation.
Trialists who wish to utilise PRO-ADD with a smaller sample size should consider a more parsimonious model or simpler conjugacy models. In such instances, trialists should take particular care evaluating model performance in light of missing data as a vital sensitivity analysis.
The modular framework enables trialists to select the most appropriate analysis methods for their specific needs. This could include the analysis of efficacy endpoints within interim and final analyses. Whilst use of binary efficacy endpoints is currently common practice, there is flexibility to adapt and modify these methods as practices evolve – such as incorporating alternate conjugate models for the analysis of continuous response data rather than discrete data. Whilst the majority of response endpoints remain binary, there is increasing interest and guidance supporting the use of continuous biomarkers in early phase drug development. 4 Draft guidance has been developed by the FDA to support the inclusion of the Circulating Tumor DNA (CtDNA) biomarker within cancer clinical trials, 81 and this biomarker has previously been utilised within novel dose-finding trial designs. 82 Whilst iPIPE has been implemented in this paper to estimate the probability of binary response, it can be used similarly with a normal conjugate model for continuous response data. Further details of this approach are provided in Section 8 of the Supplementary materials.
5.3. Point estimation for endpoints
Although the estimate of expected loss is the primary focus for final dose recommendation in PRO-ADD, trialists may still find value in the point estimates for efficacy and PRO-nAE burden score estimated in the trial. Estimates of bias and MSE for iPIPE and beta regression PRO-nAE burden score estimates are presented in Section 9 of the Supplementary Materials.
5.4. PRO summary scores
Whilst we consider the PRO-nAE burden score within this manuscript, PRO-ADD can be alternatively utilised when PROs are summarised using the Total Toxicity Profile score. 35 In such a setting, collaboration with patients and key stakeholders can help customise a weight matrix to emphasise toxicities of greater concern. However, in cases where eliciting a weight matrix may be challenging, the proposed PRO-nAE burden score provides a straightforward alternative, with equal weighting of all toxicity types being sufficient.
Trialists may also consider using alternative PRO measures to define the PRO-nAE burden score. Whilst PRO-CTCAE has been employed as a case study in this article to summarise patients’ symptomatic adverse events, other PRO measures that assess additional tolerability concepts can also be explored. Examples could include EORTC QLQ-C30 83 which is used to assess quality of life and identified as one of the most common PROMs within early phase dose-finding oncology trials. 23
The longitudinal assessment of PROs will become increasingly important as we look to assess patients’ tolerability to treatment beyond the traditionally short DLT assessment periods.
5.5. Further considerations
Having demonstrated the core operating characteristics of the proposed PRO-ADD design, we are now building on this work to explore the design’s performance under additional practical considerations. This includes the evaluation of analytical strategies for handling intercurrent events (e.g. treatment discontinuation) and missing data – both of which can impact outcome interpretation and design performance. These ongoing efforts aim to support the framework’s robust application and facilitate its adoption in real-world early phase trial settings.
5.6. Conclusion
This article presents PRO-ADD, a new modular trial design framework for dose-optimisation integrating clinician-assessed DLTs, PROs, and preliminary efficacy. By incorporating PROs, PRO-ADD ensures the selected dose is not only active but also considered safe and tolerable from both clinician and patient perspectives, setting a new standard for patient-centred dose-optimisation. We illustrate an example application of PRO-ADD, demonstrating the design’s strong performance identifying the most active and tolerable dose whilst avoiding unnecessary escalation to higher doses offering no additional benefit. As clinical development advances, incorporating patient-centric outcomes such as PROs will be valuable for refining dose-finding strategies – ensuring that dose decisions balance clinical benefit, patient-experienced tolerability, and quality of life.
Supplemental Material
Supplemental material, sj-pdf-1-smm-10.1177_09622802261435969 for PRO-ADD: Patient-empowered dose-finding trials integrating safety, preliminary efficacy and patient-reported outcomes for optimal dose selection by Emily Alger, Sumithra J Mandrekar, Jun Yin and Christina Yap in Statistical Methods in Medical Research
Acknowledgements
The authors acknowledge Scientific Computing at The Institute of Cancer Research for providing HPC and software development support. This support was instrumental in achieving the results presented in this article. https://doi.org/10.5281/zenodo.14962107
Footnotes
ORCID iDs: Emily Alger https://orcid.org/0000-0002-5378-7439
Christina Yap https://orcid.org/0000-0002-6715-2514
Sumithra J Mandrekar https://orcid.org/0000-0002-7658-1134
Ethical approval and informed consent: Not applicable.
Funding: The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: EA has been supported to undertake this work as part of a PhD studentship from the Institute of Cancer Research within the MRC/NIHR Trials Methodology Research Partnership. CY receives programmatic infrastructure funding from Cancer Research UK (CTUQQR-Dec22/100004), which supported this work.
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement: The code and data presented in this article are publicly available in the following GitHub repository: https://github.com/alemily100/PRO-ADD.
Supplemental material: Supplemental material for this article is available online.
References
- 1.Wong KM, Capasso A, Eckhardt SG. The changing landscape of phase I trials in oncology. Nat Rev Clin Oncol 2016; 13: 106–117. [DOI] [PubMed] [Google Scholar]
- 2.Wages NA, Chiuzan C, Panageas KS. Design considerations for early-phase clinical trials of immune-oncology agents. J Immunother Cancer 2018; 6: 81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Dowlati A, Manda S, Gibbons J, et al. Multi-institutional phase I trials of anticancer agents. J Clin Oncol 2008; 26: 1926–1931. [DOI] [PubMed] [Google Scholar]
- 4.Jaki T, Burdon A, Chen X, et al. Early phase clinical trials in oncology: realising the potential of seamless designs. Eur J Cancer 2023 Aug; 189: 112916. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Thall PF, Cook JD. Dose-finding based on efficacy-toxicity trade-offs. Biometrics 2004; 60: 684–693. [DOI] [PubMed] [Google Scholar]
- 6.Wages NA, Tait C. Seamless phase I/II adaptive design for oncology trials of molecularly targeted agents. J Biopharm Stat 2015; 25: 903–920. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Le Tourneau C, Lee JJ, Siu LL. Dose escalation methods in phase I cancer clinical trials. J Natl Cancer Inst 2009; 101: 708–720. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Araujo D, Greystoke A, Bates S, et al. Oncology phase I trial design and conduct: time for a change-MDICT guidelines 2022. Ann Oncol 2023; 34: 48–60. [DOI] [PubMed] [Google Scholar]
- 9.Postel-Vinay S, Gomez-Roca C, Molife LR, et al. Phase I trials of molecularly targeted agents: should we pay more attention to late toxicities?. J Clin Oncol 2011; 29: 1728–1735. [DOI] [PubMed] [Google Scholar]
- 10.Bryant AK, Sankar K, Zhao L, et al. De-escalating adjuvant durvalumab treatment duration in stage III non-small cell lung cancer. Eur J Cancer 2022; 171: 55–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Hou JZ, Ryan K, S Du, et al. Real-world ibrutinib dose reductions, holds and discontinuations in chronic lymphocytic leukemia. Fut Oncol 2021; 17: 4959–4969. [DOI] [PubMed] [Google Scholar]
- 12.Basch E, Yap C. Patient-reported outcomes for tolerability assessment in phase I cancer clinical trials. J Natl Cancer Inst 2021 Feb; 113: 943–944. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Yap C, Aiyegbusi OL, E Alger, et al. Advancing patient-centric care: integrating patient reported outcomes for tolerability assessment in early phase clinical trials – insights from an expert virtual roundtable. EClinicalMedicine 2024; 76: 102838. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Lai-Kwon J, Vanderbeek AM, A Minchom, et al. Using patient-reported outcomes in dose-finding oncology trials: surveys of key stakeholders and the national cancer research institute consumer forum. Oncologist 2022; 27: 768–777. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Oncology Center of Excellence, Center for Drug Evaluation, Research, Center for Biologics Evaluation, and Research. Optimizing the dosage of human prescription drugs and biological products for the treatment of oncologic diseases, draft guidance. Technical report, U.S. Department of Health and Human Services. Available at: https://www.fda.gov/media/164555/download (2024, accessed 3 March 2025).
- 16.US Department of Health , et al. Guidance for industry: patient-reported outcome measures: use in medical product development to support labeling claims: draft guidance. Health Qual Life Outcomes 2006; 4: 79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Basch E, Reeve BB, Mitchell SA, et al. Development of the national cancer institute’s patient-reported outcomes version of the common terminology criteria for adverse events (PRO-CTCAE). J Natl Cancer Inst 2014; 106: dju244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Basch E, Iasonos A, McDonough T, et al. Patient versus clinician symptom reporting using the national cancer institute common terminology criteria for adverse events: results of a questionnaire-based study. Lancet Oncol 2006 Nov; 7: 903–909. [DOI] [PubMed] [Google Scholar]
- 19.Veitch ZW, Shepshelovich D, Gallagher C, et al. Underreporting of symptomatic adverse events in phase I clinical trials. J Natl Cancer Inst 2021; 113: 980–988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Henon C, Lissa D, Paoletti X, et al. Patient-reported tolerability of adverse events in phase 1 trials. ESMO Open 2017; 2: e000148. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Di Maio M, Basch E, Bryce J, et al. Patient-reported outcomes in the evaluation of toxicity of anticancer treatments. Nat Rev Clin Oncol 2016; 13: 319–325. [DOI] [PubMed] [Google Scholar]
- 22.Lai-Kwon J, Yin Z, Minchom A, et al. Trends in patient-reported outcome use in early phase dose-finding oncology trials—an analysis of clinicaltrials.gov. Cancer Med 2021; 10: 7943–7957. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Alger E, Minchom A, Aiyegbusi OL, et al. Statistical methods and data visualisation of patient-reported outcomes in early phase dose-finding oncology trials: a methodological review. eClinicalMedicine 2023; 64: 102228. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Thanarajasingam G, Hubbard JM, Sloan JA, et al. The imperative for a new approach to toxicity analysis in oncology clinical trials. J Natl Cancer Inst 2015; 107: djv216. [DOI] [PubMed] [Google Scholar]
- 25.Lee SM, Lu X, Cheng B. Incorporating patient-reported outcomes in dose-finding clinical trials. Stat Med 2020 Feb; 39: 310–325. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Alger E, Lee SM, Cheung YK, et al. U-PRO-CRM: designing patient-centred dose-finding trials with patient-reported outcomes. ESMO Open 2024; 9: 103626. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Wages NA, Lin R. Isotonic phase I cancer clinical trial design utilizing patient-reported outcomes. Stat Biopharm Res 2025; 17: 36–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Schwartzberg L. Electronic patient-reported outcomes: the time is ripe for integration into patient care and clinical research. Am Soc Clin Oncol Educ Book 2016; 36: e89–e96. [DOI] [PubMed] [Google Scholar]
- 29.Spencer K, Butenschoen H, Alger E, et al. Amplifying the patient’s voice in oncology early-phase clinical trials: solutions to burdens and barriers. Am Soc Clin Oncol Educ Book 2024; 44: e433648 [DOI] [PubMed] [Google Scholar]
- 30.US Food, Drug Administration , et al. Core patient-reported outcomes in cancer clinical trials: guidance for industry (draft guidance). 2021, 2022.
- 31.Kluetz PG, Chingos DT, Basch EM, et al. Patient-reported outcomes in cancer clinical trials: measuring symptomatic adverse events with the national cancer institute’s patient-reported outcomes version of the common terminology criteria for adverse events (pro-ctcae). In American society of clinical oncology educational book. American Society of Clinical Oncology. Meeting. Vol. 35, 2016, pp.67–73. [DOI] [PubMed]
- 32.Alger E, Van Zyl M, Aiyegbusi OL, et al. Patient and public involvement and engagement in the development of innovative patient-centric early phase dose-finding trial designs. Res Involv Engagem 2024; 10: 63. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Langlais B, Mazza GL, Thanarajasingam G, et al. Evaluating treatment tolerability using the toxicity index with patient-reported outcomes data. J Pain Symptom Manage 2022; 63: 311–320. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Andrillon A, Biard L, Lee SM. Incorporating patient-reported outcomes in dose-finding clinical trials with continuous patient enrollment. J Biopharm Stat 2025; 35: 839–850. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Ezzalfani M, Zohar S, Qin R, et al. Dose-finding designs using a novel quasi-continuous endpoint for multiple toxicities. Stat Med 2013; 32: 2728–2746. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Lee SM, Hershman DL, Martin P, et al. Toxicity burden score: a novel approach to summarize multiple toxic effects. Ann Oncol 2012; 23: 537–541. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Jordan J, Maki RG. The weighted toxicity score: confirmation of a simple metric to communicate toxicity in randomized trials of systemic cancer therapy. Oncologist 2024; 29: 67–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Le-Rademacher JG, Hillman S, Storrick E, et al. Adverse event burden score – a versatile summary measure for cancer clinical trials. Cancers 2020; 12: 3251. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Spreafico M, Ieva F, Arlati F, et al. Novel longitudinal multiple overall toxicity (MOTox) score to quantify adverse events experienced by patients during chemotherapy treatment: a retrospective analysis of the MRC BO06 trial in osteosarcoma. BMJ Open 2021; 11: e053456. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Basch E, Becker C, Rogak LJ, et al. Composite grading algorithm for the national cancer institute’s patient-reported outcomes version of the common terminology criteria for adverse events (PRO-CTCAE). Clin Trials 2021; 18: 104–114. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Gamerman D. Sampling from the posterior distribution in generalized linear mixed models. Stat Comput 1997; 7: 57–68. [Google Scholar]
- 42.Xu XS, Samtani MN, Dunne A, et al. Mixed-effects beta regression for modeling continuous bounded outcome scores using nonmem when data are not on the boundaries. J Pharmacokinet Pharmacodyn 2013; 40: 537–544. [DOI] [PubMed] [Google Scholar]
- 43.Alger E, Aiyegbusi OL, Dueck AC, et al. International consensus-driven recommendations for patient-reported outcome research objectives in early phase dose-finding oncology trials: optimise-ROR. J Clin Oncol 2026; 44: 709–719. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Van Cutsem E, Kato K, Ajani J, et al. Tislelizumab versus chemotherapy as second-line treatment of advanced or metastatic esophageal squamous cell carcinoma (RATIONALE 302): impact on health-related quality of life. ESMO Open 2022; 7: 100517. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Basch E, Thanarajasingam G, Dueck AC. Methodological standards for using the patient-reported outcomes version of the common terminology criteria for adverse events (PRO-CTCAE) in cancer clinical trials. Clin Trials 2022; 19: 274–276. [DOI] [PubMed] [Google Scholar]
- 46.Rosenthal DI, Mendoza TR, Fuller CD, et al. Patterns of symptom burden during radiotherapy or concurrent chemoradiotherapy for head and neck cancer: a prospective analysis using the University of Texas MD Anderson Cancer Center Symptom Inventory-Head and Neck Module. Cancer 2014; 120: 1975–1984. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Eisenhauer EA, Therasse P, Bogaerts J, et al. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). Eur J Cancer 2009; 45: 228–247. Response assessment in solid tumours (RECIST): Version 1.1 and supporting papers. [DOI] [PubMed] [Google Scholar]
- 48.Brock K, Homer V, Soul G, et al. Is more better? An analysis of toxicity and response outcomes from dose-finding clinical trials in cancer. BMC Cancer 2021; 21: 1–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Shah M, Rahman A, Theoret MR, et al. The drug-dosing conundrum in oncology – when less is more. N Engl J Med 2021; 385: 1445–1447. [DOI] [PubMed] [Google Scholar]
- 50.Wolchok JD, Neyns B, Linette G, et al. Ipilimumab monotherapy in patients with pretreated advanced melanoma: a randomised, double-blind, multicentre, phase 2, dose-ranging study. Lancet Oncol 2010; 11: 155–164. [DOI] [PubMed] [Google Scholar]
- 51.Leighl NB, Hellmann MD, Hui R, et al. Pembrolizumab in patients with advanced non-small-cell lung cancer (KEYNOTE-001): 3-year results from an open-label, phase 1 study. Lancet Respir Med 2019; 7: 347–357. [DOI] [PubMed] [Google Scholar]
- 52.Rotte A, Frigault MJ, Ansari A, et al. Dose–response correlation for CAR-T cells: a systematic review of clinical studies. J Immunother Cancer 2022; 10: e005678. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Cheung YK, Diaz KM. Monotone response surface of multi-factor condition: estimation and Bayes classifiers. J R Stat Soc Ser B: Stat Methodol 2023; 85: 497–522. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Mander AP, Sweeting MJ. A product of independent beta probabilities dose escalation design for dual-agent phase I trials. Stat Med 2015; 34: 1261–1276. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Yuan Y, Hess KR, Hilsenbeck SG, et al. Bayesian optimal interval design: a simple and well-performing design for phase I oncology trials. Clin Cancer Res 2016 Aug; 22: 4291–4301. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Liu S, Yuan Y. Bayesian optimal interval designs for phase I clinical trials. J R Stat Soc Ser C (Appl Stat) 2015; 64: 507–523. [Google Scholar]
- 57.US Food, Drug Administration , et al. Drug development tools: fit-for-purpose initiative. US Food & Drug Administration. Available at: https://www.fda.gov/drugs/development-approval-process-drugs/drug-development-tools-fit-purp ose-initiative (2022b).
- 58.Ananthakrishnan R, Lin R, He C, et al. An overview of the BOIN design and its current extensions for novel early-phase oncology trials. Contemp Clin Trials Commun 2022; 28: 100943. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Zhou H, Zhao Y, Kuo YW, et al. BOP2: Bayesian optimal phase ii design with simple and complex endpoints. Available at: https://trialdesign.org/one-page-shell.html#BOP2 (2025, accessed 12 January 2025).
- 60.Zhou Y, Lee JJ, Yuan Y. A utility-based Bayesian optimal interval (U-BOIN) phase I/II design to identify the optimal biological dose for targeted and immune therapies. Stat Med 2019; 38: S5299–S5316. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Lin R, Zhou Y, Yan F, et al. BOIN12: Bayesian optimal interval phase I/II trial design for utility-based dose finding in immunotherapy and targeted therapies. JCO Precis Oncol 2020; 4: 1393–1402. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Patnaik A, Kang SP, Rasco D, et al. Phase I study of pembrolizumab (MK-3475; anti–PD-1 monoclonal antibody) in patients with advanced solid tumors. Clin Cancer Res 2015 Sep; 21: 4286–4293. [DOI] [PubMed] [Google Scholar]
- 63.Kang SP, Gergich K, Lubiniecki GM, et al. Pembrolizumab KEYNOTE-001: an adaptive study leading to accelerated approval for two indications and a companion diagnostic. Ann Oncol 2017; 28: 1388–1398. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.González FS, Palacios CAQ, Gordón AM, et al. Delayed immune-related hepatitis after 24 months of pembrolizumab treatment: a case report and literature review. Anticancer Drugs 2024; 35: 284–287. [DOI] [PubMed] [Google Scholar]
- 65.Watson GA, Veitch ZW, Shepshelovich D, et al. Evaluation of the patient experience of symptomatic adverse events on phase I clinical trials using PRO-CTCAE. Br J Cancer 2022; 127: 1629–1635. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Delignette-Muller ML, Dutang C. fitdistrplus: an R package for fitting distributions. J Stat Softw 2015; 64: 1–34. [Google Scholar]
- 67.Marra G, Fasiolo M, Radice R, et al. A flexible copula regression model with Bernoulli and Tweedie margins for estimating the effect of spending on mental health. Health Econ 2023; 32: 1305–1322. [DOI] [PubMed] [Google Scholar]
- 68.Zhou Y, Lin R, Kuo YW, et al. BOIN suite: a software platform to design and implement novel early-phase clinical trials. JCO Clin Cancer Inform 2021; 5: 91–101. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Englert S, Mercier F, Pilling EA, et al. Defining estimands for efficacy assessment in single arm phase 1b or phase 2 clinical trials in oncology early development. Pharm Stat 2023; 22: 921–937. [DOI] [PubMed] [Google Scholar]
- 70.Thall PF, Cook JD, Estey EH. Adaptive dose selection using efficacy-toxicity trade-offs: illustrations and practical considerations. J Biopharm Stat 2006; 16: 623–638. [DOI] [PubMed] [Google Scholar]
- 71.Brock K, Billingham L, Copland M, et al. Implementing the EffTox dose-finding design in the Matchpoint trial. BMC Med Res Methodol 2017; 17: 112. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Mercier F, Homer V, Geng J, et al. Estimands in oncology early clinical development: assessing the impact of intercurrent events on the dose-toxicity relationship. Stat Biopharm Res 2025; 17: 78–86. [Google Scholar]
- 73.Mukherjee A, Moscovici JL, Liu Z. Estimands for early-phase dose optimization trials in oncology. Biom J 2025; 67: e70072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Takeda K, Taguri M, Morita S. BOIN-ET: Bayesian optimal interval design for dose finding based on both efficacy and toxicity outcomes. Pharm Stat 2018; 17: 383–395. [DOI] [PubMed] [Google Scholar]
- 75.O’Quigley J, Conaway M. Continual reassessment and related dose-finding designs. Stat Sci 2010; 25: 202. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Du Y, Yin J, Sargent DJ, et al. An adaptive multi-stage phase I dose-finding design incorporating continuous efficacy and toxicity data from multiple treatment cycles. J Biopharm Stat 2019; 29: 271–286. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Yin J, Qin R, Ezzalfani M, et al. A Bayesian dose-finding design incorporating toxicity data from multiple treatment cycles. Stat Med 2017; 36: 67–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Zhang L, Bayman EO, Zamba KD. A modified Huber loss function for continual reassessment methods in clinical trials. Seq Anal 2024; 43: 28–48. [Google Scholar]
- 79.Thall PF, Cook JD. Dose-finding based on efficacy–toxicity trade-offs. Biometrics 2004; 60: 684–693. [DOI] [PubMed] [Google Scholar]
- 80.Alger Emily, Regnault Antoine, Dueck Amylou C, et al. A practical toolkit with recommendations for analysing and visualising patient-reported outcomes in early phase dose-finding oncology trials (OPTIMISE-AR). The Lancet Oncology 2026; 27: e218–e230. [DOI] [PubMed] [Google Scholar]
- 81.Food, Drug Administration , et al. Use of circulating tumor dna for early-stage solid tumor drug development-guidance for industry 2022, 2023.
- 82.Mozgunov P, Jaki T. A flexible design for advanced phase I/II clinical trials with continuous efficacy endpoints. Biometr J 2019; 61: 1477–1492. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Scott N, Fayers P, Aaronson N, et al. EORTC QLQ-C30. In: Reference values. Brussels: EORTC, 2008.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplemental material, sj-pdf-1-smm-10.1177_09622802261435969 for PRO-ADD: Patient-empowered dose-finding trials integrating safety, preliminary efficacy and patient-reported outcomes for optimal dose selection by Emily Alger, Sumithra J Mandrekar, Jun Yin and Christina Yap in Statistical Methods in Medical Research






