Abstract
This paper reviews two types of geometric methods proposed in recent years for defining statistical decision rules based on 2-dimensional parameters that characterize treatment effect in a medical setting. A common example is that of making decisions, such as comparing treatments or selecting a best dose, based on both the probability of efficacy and the probability toxicity. In most applications, the 2-dimensional parameter is defined in terms of a model parameter of higher dimension including effects of treatment and possibly covariates. Each method uses a geometric construct in the 2-dimensional parameter space based on a set of elicited parameter pairs as a basis for defining decision rules. The first construct is a family of contours that partitions the parameter space, with the contours constructed so that all parameter pairs on a given contour are equally desirable. The partition is used to define statistical decision rules that discriminate between parameter pairs in term of their desirabilities. The second construct is a convex 2-dimensional set of desirable parameter pairs, with decisions based on posterior probabilities of this set for given combinations of treatments and covariates under a Bayesian formulation. A general framework for all of these methods is provided, and each method is illustrated by one or more applications.
Keywords: Bayesian statistics, Clinical trials, Dose-finding, Indifference set, Medical decision making, Phase II clinical trial, Trade-offs
1. Introduction
Most medical treatments have multiple effects on patients. Typically, a drug's desirable anti-disease effect is called “response” or “efficacy,” while undesirable effects are called “toxicity” or ”adverse events.” When choosing among established treatments for an individual patient, or when evaluating treatments in a clinical trial, it is useful to consider the rates of such effects jointly. The widespread use of aspirin for its analgesic effects and its ability to lower the risk of cardiovascular disease is based on the idea that the likelihoods of these beneficial effects are a desirable trade-off for the risk that aspirin also may cause an upset stomach. Beta blockers effectively reduce abnormally high blood pressure but also may cause dizziness, fatigue, insomnia or a reduced libido. How one classifies outcomes as being desirable or undesirable depends on one's treatment goals. For example, a particular drug may be characterized either as an antihistamine that causes drowsiness, or as a sleeping pill that dries up one's sinuses. Consideration of trade-offs between desirable and undesirable clinical outcomes is especially important in oncology, where the goal is ideally to cure or hopefully to slow the growth of the patient's cancer, but most treatments also carry a substantial risk of harmful effects. For example, treating a hematologic malignancy with transplanted allogeneic blood or marrow stem cells obtained from a human leukocyte antigen (HLA) matched donor may have the desirable effect of killing cancer cells but also may kill normal cells in the skin, liver or gastro-intestinal tract through an autoimmune reaction known as graft-versus-host disease. Rapid treatment of stroke with agents aimed at dissolving the blood clot that caused the stroke also carries the risk that the agent may cause bleeding in the brain. A general problem arising in treatment of life-threatening diseases is that some therapies may extend average survival time but also may reduce the patient's quality of life.
This paper will focus primarily on settings where it is sensible to characterize treatment effect as a 2-dimensional parameter. The following simple example will serve to motivate the general problem and illustrate some of the main ideas. Consider a clinical setting where the observed outcome includes the pair Y = (Y1, Y2) with Y1 the indicator of response to treatment, “efficacy,” and Y2 the indicator of a treatment-related severe adverse event (SAE), “toxicity.” Denoting θa,b = Pr(Y1 = a, Y2 = b) for a, b ∈ {0, 1}, a reasonable likelihood for Y is the product multinomial
with 3-dimensional parameter vector θ = (θ1,1, θ1,0, θ0,1), writing θ0,0 = 1 − θ1,1 − θ1,0 − θ0,1. In this setting, before constructing statistical decision rules it often is appropriate to first apply the mapping
to reduce θ to the 2-dimensional parameter ξ = (Pr(Y1=1), Pr(Y2=1)) consisting of the marginal probabilities of response and toxicity. Thus, ξ1 = Pr(efficacy) and ξ2 = Pr(toxicity), so it is more desirable to have larger ξ1, smaller ξ2 or both. An equivalent formulation may be constructed in terms of ξ2 = Pr(no toxicity) rather than Pr(toxicity), in order to make larger values more desirable for both parameters, if desired.
Dimension reduction is central to statistical thinking. The usual motivation for computing a statistic is that it estimates an important parameter of a probability model for the data. The principle of sufficiency, that for the purpose of making inferences under a given probability model no information is lost by reducing the raw data to the sufficient statistic, provides an intuitively appealing rationale for the statistical process. A common example is that in which a data set Dn = {Y1, …, Yn} of 1-dimensional real-valued observations is reduced to the sample mean and variance, (Ȳ, s2), which estimate the corresponding population values, μ and σ2. This relies on the idea that these two parameters are the most interesting properties of the population from which the data arose. Given this, the mapping Dn → (Ȳ, s2) makes sense. In the regimes described here, the analogous dimension reduction θ → ξ is done on the parameter vector itself, based on the idea that ξ characterizes what is most interesting about θ.
As a numerical example, suppose that the null historical mean outcome probabilities with standard treatment are ξ0 = (ξ0,1, ξ0,2) = (.40, .30) so that, considering each parameter individually, either a value of ξ1 > .40 or a value of ξ2 < .30 is an improvement over the historical mean. In the single-arm phase II setting, Bryant and Day (1995) proposed an optimal 2-stage frequentist test-based design that requires an alternative hypothesis with improvements in both entries of ξ0, such as {ξ : ξ1 ≥ .55, ξ2 ≤ .20} in the example. This type of goal, which aims for both an increase of .15 or more in the response rate ξ1 and a decrease of .10 or more in the toxicity rate ξ2 compared to ξ0, is rather optimistic in most clinical settings, however. A more general approach with rectangular alternatives allows one of the two parameters to be equal to or worse than the null value so that, for example, the alternative might be {ξ : ξ1 ≥ .55, ξ2 ≤ .40}. This allows a .10 increase in the toxicity probability as a trade-off for a .15 increase in the probability of response. General tests for this type of alternative are described by Jennison and Turnbull (1993, 2000).
To construct a Bayesian design, a tractable model is obtained by assuming that θ follows a conjugate Dirichlet prior where α = (α1,1, α1,0, α0,1, α0,0) denotes a vector of positive-valued hyperparameters and we denote θ ∼ Dir(α). Under this model, for count data X = (X1,1, X1,0, X0,1, X0,0) where the posterior is θ | X ∼ Dir(α + X) and marginally ξ1 and ξ2 follow beta priors and posteriors. For a single-arm phase II clinical trial with rules corresponding to the frequentist test described above, one may stop the trial and declare the treatment to be not promising for future study if either Pr(ξ1 > .55 | data) < p1L or Pr(ξ2 < .40 | data) < p2L, with the decision cut-offs p1L and p2L in these Bayesian stopping criteria calibrated to obtain a design in which the two rules applied together give a design with good frequentist properties. If (.40, .30) is the historical mean, corresponding to the null values ξ(0) in the frequentist tests, one may regard these two rules as allowing, on average, an increase of .10 in Pr(toxicity) as a trade-off for achieving an increase of .15 in Pr(response). In general, the fixed values .55 and .40 may be replaced by random parameters corresponding to historical treatment, although this is not strictly necessary. Details for constructing multiple stopping rules of this form are given by Thall, Simon and Estey (1995).
The problem becomes much more difficult if two or more equally desirable targets are specified, such as the three pairs (.55, .40), (.60, .45) and (.40, .15). The target (.60, .45) allows a .15 increase in ξ2 as a trade-off for an increase of .20 in ξ1, while the target (.40, .15) allows the response rate to remain at the historical value provided that on average ξ2 drops from .30 to .15. In such settings, constructing a frequentist test is no longer straightforward, since there no longer is a simple rectangular alternative. Obtaining multiple targets from a physician is not only easy but, in my experience, a physician asked for a single target is very likely to provide several. The problem is how to construct statistical decision rules that reflect several equally desirable 2-dimensional parameter values. In particular, some sort of dimension reduction is needed, and it is not obvious how this should be done.
In practice, the most common approach to this problem is to ignore one of the two parameters, typically by using the semantic device of calling Y2 a “secondary outcome,” and constructing a conventional decision rule or trial design based on ξ1 alone. If, for example, the stated goal is to achieve a 50% or higher response rate, one may use simple decision criteria such as Pr(ξ1 > .50 | data) or mean E(ξ1 | data) under a Bayesian model or frequentist criteria such as the sample mean response rate ξ̂1 or the approximately normal test statistic (ξ̂1 − .50)/(.50 × .50/n)1/2. Of course, this approach has many defects, including the inability to identify treatment differences in settings where, for treatments A and B, ξ1,A = ξ1,B but ξ2,A > ξ2,B. Many clinical trial protocols that take this sort of approach include a statement that toxicity will be monitored, but do not provide explicit rules for doing so. Since critical decisions pertaining to safety monotoring are made in a purely subjective manner, the operating charactistics of such designs cannot be computed. The point is that treatment differences may occur in terms of ξ1 alone, ξ2 alone, or both parameters.
This paper will review some geometric methods for constructing decision rules based on 2-dimensional parameters in clinical trials. In each method, treatment effect is characterized by a parameter pair ξ = (ξ1, ξ2), typically arising from a model of higher dimension, and a geometric representation in the parameter space constructed from a set of elicited ξ values is used as a basis for defining decision rules. These methods are described in three recent papers. Thall, Sung and Estey (2002) compare two-stage treatment strategies for patients with acute leukemia based on covariate-adjusted probabilities of efficacy and death under a generalized logistic model. Thall and Cook (2004) present a general method for dose-finding in phase I/II clinical trials that uses trade-offs between the probabilities of efficacy and toxicity as a basis for selecting doses. The geometric construct used in each of these papers is a family of contours that partitions the 2-dimensional parameter space, with the contours constructed so that all parameter pairs on a given contour are equally desirable. Thall, Wooten and Shpall (2006) present a Bayesian method for comparing treatments for rapidly fatal diseases based on the probability of response and the mean time to response among responders, computed under a competing risks regression model for the times to response and death. They construct decision rules based on posterior probabilities of a convex set of desirable ξ values. While these methods deal with a variety of data structures, medical settings and statistical goals, they share several essential properties. My goals in this paper are to provide a general framework for these methods and illustrate how each may be applied in practice.
Section 2 will provide a general notational framework for what is to follow. In section 3, methods for constructing families of trade-off contours that are used as a basis for decision making are described, and the methods are illustrated by applications to actual clinical trial designs. Section 4 describes geometric methods for evaluating and comparing treatment effects based on posterior probabilities of two-dimensional parameter sets, and a brief discussion is given in section 5.
2. A General Framework
Consider a clinical trial in which the data on each patient consist of a multivariate outcome Y and a vector Z containing some combination of covariates, treatments or doses. A Bayesian model will be assumed that consists of a likelihood f (Y | Z, θ) and prior p(θ), with parameter vector θ = (θ1, …, θp) of dimension p ≥ 2. In some cases, the Bayesian structure is not strictly necessary and only the likelihood and a consistent estimator of θ are needed. All of the methods rely on the key assumption that, from both a scientific and a medical viewpoint, it is appropriate to base treatment evaluation and decision making on a 2-dimensional parameter (ξ1, ξ2) = ξ ∈ Ξ ⊂ R2 defined by a mapping (θ, Z) → ξ. In the settings addressed here, ξ is readily interpretable, and in my experience it often is proposed by the physicians planning the trial. In contrast, the likelihood, prior and model parameter θ are essentially statistical objects used to make statements and decisions about ξ and Z. In the applications discussed below, Ξ is, respectively, the triangular sub-region of [0, 1]2 where 0 < ξ1 + ξ2 ≤ 1 for probabilities corresponding to trinary outcomes, the unit square [0, 1]2 for unconstrained pairs of probabilities, [0, ∞) × [0, 100] for ξ1 = mean survival time and ξ2 = mean QOL, and [0, 1] × [0, ∞) for ξ1 a probability and ξ2 a mean event time.
The basis for the first set of methods is the idea that a family of contours may be constructed to reflect the area expert's goals in such a way that, for each real number δ, all are equally desirable, and partitions Ξ. Formally, and for distinct δ1 and δ2. Thus, δ quantifies the desirability of all . In this sense, each contour may be regarded as an “indifference” set. Depending on the method used for constructing , for each real δ the contour is a curve in Ξ constructed so that it satisfies a desirability constraint expressed in terms of an explicit function of ξ and δ.
In the illustrations presented below, the partition is constructed in several different ways. In the first method, one elicits several parameter pairs and their corresponding desirabilities δ1, …, δK. To implement the other methods, one elicits a set of target pairs that the physician considers equally desirable, these are used to construct a target contour , and is generated by applying a homotopy to . In some settings, the targets may be elicited in relation to an established null or historical mean value ξ0. Using a graphical representation of the elicited points and contours greatly facilitates the elicitation process. In constructions where a target contour is first constructed from elicited values, for K = 2 or 3 one may determine so that it passes through all of the elicited points, whereas for K > 3 one may treat the elicited points like regression data and use least squares or some other curve-fitting method to determine . Although things will be described in the context of designing a clinical trial where the experts are physicians, the methods may be applied more generally, and the purpose of the exercise may be design, decision-making or data analysis.
Assume without loss of generality that larger values of ξ1 and smaller values of ξ2 are more desirable, corresponding to the earlier example where ξ1 = Pr(efficacy) and ξ2 = Pr(toxicity), so that the lower right portion of Ξ corresponds to the most desirable ξ pairs. For each method used to construct , the main formal requirement is that, for each real δ, considering ξ2 as a function of ξ1 on , ξ2 must increase with ξ1 on . This ensures that if either ξ1 is increased or ξ2 is decreased with the other parameter held constant, then the point must move to a contour having a larger desirability. Essentially, δ : Ξ → R is a function mapping each 2-dimensional parameter of interest into a real number that can be used as a basis for comparing treatments and making decisions.
3. Constructing A Family of Contours
3.1 Eliciting Both Targets and Desirabilities
The first method for constructing a family of contours, described in Thall, Sung and Estey (2002), requires one to elicit K = 2 or 3 target pairs and their corresponding numerical desirabilities, such that for each j = 1, …, K, given a suitably defined function f that determines the shape of the contours. The elicited values are used to estimate the parameters of f, and thus one may determine the contour for all given δ. The practical reason for constructing a partition of Ξ in this way is to map Ξ into a set of one-dimensional δ values in such a way that, if δ(ξ1) > δ(ξ2), this accurately reflects the subjective belief that ξ1 is more desirable than ξ2. As noted earlier, one may use either consistent frequentist estimates or Bayesian posterior means to estimate ξ. For example, if the estimates of ξ for patients given treatments A and B are and , one may compare the treatments using the estimated desirabilities and . For many applications, a model more elaborate than the multinomial-Dirichlet given earlier may be needed if one wishes to account for regression of the ξj's on treatments and patient covariates known to affect outcome.
The clinical trial that motivated this method was designed to compare several two-stage chemotherapy strategies for treating patients with acute leukemia who had very poor prognosis due to the fact that their disease had re-appeared after previously being brought into remission. Because CR is a necessary but not sufficient condition for long-term survival in this patient group, chemotherapy typically is very aggressive, and patients may die due to infections or other effects that are related to both their disease and chemotherapy regimen. Patient outcome in stage s = 1, 2 was characterized as Ys = (Ys,1, Ys,2) where Ys,1 indicates complete remission (CR), Ys,2 indicates death and Ys,3 = 1 − Ys,1 − Ys,2 indicates that neither of these events occurred, with Ys scored within one month from the start of chemotherapy in that stage. For a 2-stage strategy consisting of treatments (a, b), the patient was first given a, and a second stage of therapy with b was given if the patient neither achieved CR nor died with a in stage 1, that is, if Y1,1 = Y1,2 = 0 and Y1,3 = 1. This is an example of a dynamic treatment regime (Murphy, 2003; Lavori and Dawson, 2004) since the treatment given in stage 2 depends on the stage 1 outcome. For s = 1 let t denote the stage 1 treatment, while for s = 2 let t denote the 2-stage strategy. To account for stage s = 1,2, outcome j = 1,2 indexing CR and death, treatment or strategy t, and covariates Z = (previous remission duration, age), we assumed a generalized logistic regression model of the form
| (1) |
for s = 1, 2 and j = 1, 2, with πs,3(t, Z) = 1 − πs,1(t, Z) − πs,2(t, Z). The linear term ηs,j,t,Z was formulated to account for all effects up to second order interactions, under a Bayesian model. Additional details are given in Thall, Estey and Sung (2002).
The 2-dimensional parameters of interest were the probabilities of achieving CR or dying during the two stages of therapy with a given strategy (a, b) for a given Z, given by
| (2) |
for j = 1,2. In this application, the parameter set is the triangular region ΞT = {ξ : 0 ≤ ξ1, ξ2 ≤ 1 and 0 < ξ1 + ξ2 ≤ 1} shown in Figures 1a-c. The particular desirability function used was with values of (α1, α2, α3) obtained by solving the K = 3 equations f(.50, .15) = 1, f(.40, .40) = 0 and f(.30, 0) = 0. These equations reflected the physician's subjective assignment of δ = 1 to the desirable pair corresponding to a 50% CR rate coupled with a 15% death rate, and δ = 0 to the less desirable pairs and . This yielded solution α1 = 3.333, α2 = −2.548 and α3 = .707. The contour of ξ pairs with desirability δ is . Figure 1a illustrates the family used in this application. Since ξ(a, b, Z) is a function of treatment strategy and covariates including treatment-covariate interactions, under a Bayesian model δ(E{ξ(a, b, Z) | data}) was used to determine the desirability of each 2-stage strategy for each Z. The targets and desirabilities for the leukemia trial illustrated in Figure 1a were elicited to correspond to average historical values using this particular type of 2-stage strategy. Alternatively, the elicitation process may be carried out in terms of a reference covariate vector corresponding to a “typical” patient chosen by the physician.
Figure 1.
Trade-off contours for the probabilities of complete remission and death in the acute leukemia trial. Figure 1a (upper left) shows the convex contours based on the actual elicited values. Figure 1b (upper right) and Figure 1c (lower left) show hypothetical families of contours that might be obtained in practice.
Figures 1b and 1c illustrate some other possibilities. Figure 1b shows the contours for having δ1 = 1 and having δ2 = 0 in the case K = 2, where the solution to the linear function δ(ξ) = α1ξ1 + α2ξ2 is (α1, α2) = (2.8571, −2.8571). The linear contours in Figure 1b also correspond to the conventional approach of eliciting utilities u = (u1, u2, u3) for the three outcomes and identifying δ with the expected utility, so that δ = ξ1u1 + ξ2u2 + ξ3u3 = ξ1(u1 − u3) + ξ2(u2 − u3) + u3. This corresponds to α1 = u1 − u3 and α2 = u2 − u3. For example, if the utility of the outcome “neither CR nor death” is u3 = 0 then CR has utility u1 = α1 = 2.8571 and death has utility u3 = −2.8571, whereas if u3 = −1 then u1 = 1.8571 and u3 = -3.8571. In Figure 1c, considering ξ2 as a function of ξ1 the contours are concave rather than convex, and f is characterized by (α1, α2, α3) = (3.3333, 6.6273, 1.75). The practical difference between the three families of contours is illustrated by considering the relationship between (.30, 0) and (.50, .15). In Figure 1a, which is based on the actual elicited values, these two pairs are equally desirable. In Figure 1b, (50, .15) is slightly more desirable than (.30, 0). In Figure 1c, (50, .15) is much more desirable than (.30, 0). In all three illustrations, ξ2 increases as a function of ξ1, which ensures that increasing ξ1 for ξ2 fixed increases δ and increasing ξ2 for ξ1 fixed decreases δ.
3.2 Eliciting Equally Desirable Targets
The second and third methods for constructing contours, given in Thall and Cook (2004) and Thall, Cook and Estey (2006) use a slightly different heuristic procedure for eliciting target pairs. For these methods, Ξ is either [0, 1]2 or the triangular subset of [0, 1]2 where 0 ≤ ξ1 + ξ2 ≤ 1. The first pair is a desirable target specified by the physician in the interior of Ξ. After specifying , the physician is asked for the smallest response probability that would make the second target equally desirable if toxicity were impossible, so that . Similarly, the physician is then asked for the largest toxicity probability that would make the third target as desirable as the first two if the response probability were maximal, so that either if Ξ = [0, 1]2 or in the trinary outcome case where ξ1 + ξ2 ≤ 1. A target contour is then fit to these three points, is assigned a convenient value of δ, such as 0 or 1, and the other contours in are generated by applying a homotopy that generates a contour from for given δ ∈ R. This simplifies the elicitation process in that the physician is not asked to specify numerical desirabilities, but rather only equally desirable target probability pairs.
To implement these methods, a flexible 3-parameter function, such as , is fit to the three elicited pairs. After solving for the parameters that characterize τ, the target contour is defined as . The only general requirement is that dξ2/dξ1 > 0 for all ξ1 in the subset of [0, 1] where ξ2(ξ1) ∈ Ξ. Once is established, to generate , for any pair q ∈ Ξ, let L(q) denote the straight line passing through q and the ideal point (1,0), so that L(q) consists of the points in Ξ satisfying ξ2 = {q2/(1 − q1)}(1 − ξ1). Let be the point where L(q) and intersect. This is illustrated in Figure 2, which corresponds to an application that will be discussed below. In Figure 2, the three elicited equally desirable targets are shown as round points, is given by the thick line and q and r as described are given by triangular points. Denoting the Euclidean distance from any ξ ∈ Ξ to (1,0) by ∥ξ − (1, 0)∥ = {(ξ1 − 1)2 + (ξ2 − 0)2}1/2, the desirability of q is defined to be
Figure 2.
The family of trade-off contours used for the IA-tPA dose-finding trial.
| (3) |
Since for all , as a convenience the number 1 is subtracted from the ratio of norms in (3) in order to make the desirability of any point on equal to 0. Thus, points inside the contour and hence closer to (1,0) have desirability δ > 0 and points outside have δ < 0. Given q, the point is found by solving the two equations ξ2 = τ(ξ1) and ξ2 = {q2/(1 − q1)}(1 − ξ1). Given δ, the contour of points having this desirability is determined by the continuous function defined by
| (4) |
which maps each to a point such that ∥ξ − (1, 0)∥/∥fδ(ξ) − (1, 0)∥ = δ + 1. homotopy Thus, for given desirability δ, and the target contour is . Defining the family of contours to be , it follows that for distinct δ1 and δ2 and so that partitions Ξ.
3.3 Using An Lp Norm With Equally Desirable Targets
The third method also is based on three equally desirable elicited points, but it uses an Lp norm to define . For p > 0 and ξ ∈ Ξ, define the Lp norm distance from ξ to (1,0) to be
| (5) |
In particular, the norm is defined so that . Setting δ = 0 for points on as before, the contour with desirability δ is defined to be
It is easy to show (Cook, 2006) that, given and , p is uniquely determined by and may be found by a monotone search. The Lp norm method has the advantage that a solution always exists, whereas when using a 3-parameter function to construct contours there is the possibility that, for particular elicited targets, the admissibility requirement of monotonicity for ξ2 as a function of ξ1 may be violated. Another technical difference is that, if ξ is given and one wishes to solve for its desirability δ, the method based on requires solving a cubic equation in δ, whereas the Lp norm method gives the solution from the simple equation δ = 1 − ∥ξ − (1, 0)∥p.
These methods will be illustrated with a clinical trial in which the scientific goal is to determine the best dose of tissue plasminogen activator (tPA) administered intra-arterially to children who have suffered a stroke within the past 3 hours. The therapeutic goal is to dissolve the blood clot that caused the stroke and thus restore blood flow to the stroke site in the brain, that is, achieve “recanalization.” Response is defined as recanalization based on comparison of pre-infusion and immediate post-infusion angiograms. Toxicity is defined as death, intracranial hemorrhage or any other SAE within 48 hours of administering the tPA. The three target probability pairs considered equally desirable by the physician planning the trial were ξ(1) = (.50, .05), ξ(2) = (.40, 0), and ξ(3) = (1, .20). This says that the physician considered 50% response with a 5% toxicity rate, certain response with a 20% toxicity rate, and 40% response with toxicity impossible to be equally desirable.
To obtain contours using the Lp-norm method, we first use ξ(2) and ξ(3) to define
The equation ∥(.50, .05) − (1, 0)∥ p = 1 has unique solution p = 1.1829, the resulting contours are given by
and the desirability of a given parameter pair is δ(ξ) = 1 − ∥ξ − (1, 0)∥1.1829. The contours are illustrated in Figure 2, which was described earlier. Alternatively, using the quadratic function rather than the Lp norm to define , the three elicited pairs give the solution α1 = .4167, α2 = −.25, α3 = .0333 so that, given δ,
The contours obtained using the quadratic function are virtually identical in appearance to those in Figure 2 obtained using the Lp norm method, hence are not shown.
In the tPA trial, the primary objective is to determine an optimal dose among the four levels {0.6, 0.8, 1.0, 1.2} mg/kg tPA. Let Z denote the standardized dose obtained by taking logs and centering at the mean. The method relies on the underlying dose-outcome model given by marginals ξ1(Z) = Pr(Y1 = 1) = Pr(response) = logit−1(β1,0 + β1,1Z + β1,2Z2) and ξ2(Z) = Pr(Y2 = 1) = Pr(toxicity) = logit−1(β2,0 + β2,1Z) with joint probabilities Pr(Y = (a, b) | Z) for a, b ∈ {0, 1} determined by these marginals and one association parameter ψ under a Gumbel model. Thus, the model parameter vector is θ = (β1,0, β1,1, β1,2, β2,0, β2,1, ψ), and the mapping (θ, Z) → ξ(Z) = (ξ1(Z), ξ2(Z)) characterizes (θ, Z) in terms of the pair of outcome probabilities as functions of dose. Additional details of the general model and method are given in Thall and Cook (2004). The design specifies a maximum of 24 patients to be treated in cohorts of size 2, starting at the lowest dose and not skipping an untried dose when escalating. Each new cohort is given the dose having the largest posterior mean E{ξ(Z) | data} based on the most recent data from the trial, and this criterion determines the final recommended dose based on the data from all 24 patients. An important aspect of the trade-off contours (Figure 1a) that are the basis for choosing among the doses in this trial is that a small increase in ξ2(Z) = Pr(toxicity | Z) is allowed for a relatively large increase in ξ1(Z) = Pr(recanalization | Z) in order to maintain the same desirability.
4. Basing Decisions On Posterior Probabilities
The next method is different from those discussed above in that is now used to define a set of parameter pairs such that each is at least as desirable as a pair on , and decisions are based on posterior probabilities . Thus, this method requires a Bayesian model. For an application where ξ is defined so that larger values of both parameters are more desirable, once has been established the target set is defined to be
| (6) |
and is oriented toward the upper right portion of Ξ. If the parameters are defined so that larger ξ1 and smaller ξ2 are more desirable than
| (7) |
and is oriented toward the lower right portion of Ξ. This method allows a variety of general forms of , including polygonal boundaries. This is illustrated by Figure 3, which corresponds to the first application that we will discuss below.
Figure 3.
Four hypothetical target sets for ξ1 = mean survival time and ξ2 = mean quality of life (QOL) on a 100-point scale, based on historical means of 3 years survival and 50 on the QOL scale.
In the methods discussed in section 3, Ξ was a set of probability pairs, which facilitated the use of Euclidean or Lp norm distance to define a homotopy that was used to generate from either one or two contours. In the present regime, the parameters ξ1 and ξ2 may be defined on different domains, as in the illustration below where ξ1 is a probability and ξ2 is a mean time. To construct , one may use any of the methods described previously, define a polygonal boundary as in Conaway and Petroni (1995) or Thall and Cheng (1999), or apply the following method. Given K ≥ 2 elicited target parameter pairs, , the basic idea is to treat these like regression data, fit a curve to these values using some standard method such as least squares, and define to be the fitted curve. That is, the 's are treated as predictors and the 's as outcome variables in a parametric regression model, such as or some other functional form. If desired, one may use least absolute value regression, that is, solve for the parameters that minimize |ξ2 − τ(ξ1)|, so that ξ1 and ξ2 are treated more symmetrically. Denoting the fitted function by τ̂(ξ1) and the interval of ξ1 for which by [ξ1,L, ξ1,U], the target contour is
The physician may wish to modify this contour by imposing a lower bound on ξ1 so that the left-hand portion of is a vertical line, as in Figure 3, or similarly bound ξ2. While this type of construction is more flexible than the previous methods, it requires some care in implementation. Showing the physician a graph of the scattergram of elicited points and the fitted curve is essential, since this allows the physician to adjust any point that disagrees substantially with the others.
As a first illustration of this method, consider a hypothetical clinical setting based on Y1 = survival time and Y2 = quality of life (QOL), with ξj = E(Yj) for j = 1, 2. Thus, ξ1 = mean survival time and ξ2 = mean QOL, and larger values of both parameters are more desirable. Assume that QOL is scored on a 100-point scale where 0 corresponds to the worst possible QOL, 80 ≤ QOL ≤ 90 is a typical range for people who are disease-free, and 100 corresponds to perfect health and psychological well-being. In practice, as in the previous applications, these parameters typically will arise from an underlying model accounting for treatment and covariate effects, and (Y1, Y2) may be part of an observed vector of higher dimension. For simplicity, such additional structure will be ignored in this example. Suppose that, based on historical data, the means are ξ1,0 = 3 years and ξ2,0 = 50 with standard therapy. Given ξ0 = (3,50), several hypothetical target sets will be described where is constructed from three elicited pairs. The first pair corresponds to a case where a decrease in ξ1 from 3 years to ξ̱1 is a trade–off for a very high QOL ξ̄2. The second pair corresponds to an improvement in both parameters. The third pair is , so that ξ̱2 is the lowest QOL that would be acceptable regardless of how large survival time may be. Since the parameters are defined so that larger values of both ξ1 and ξ2 are more desirable, is oriented toward the upper right of the domain Ξ = [0, ∞) × [0, 100]. A simple function to determine is the hyperbolic form ξ2 = (ξ̄2 − ξ̱2) (ξ̱1/ξ1)α + ξ̱2 with α > 0. The numerical values of the three elicited pairs are summarized in Table 1, and the target sets are illustrated in Figure 3. The target set in the upper left plot allows a decrease in mean survival from 3 to 2.5 years as a trade-off for perfect QOL, and a small drop in QOL from 50 to 45 if survival time is very large. The target in the upper right is slightly more optimistic in that the lower bound on QOL is maintained at the historical mean of 50. The set in the lower left requires a QOL of 90 as a trade-off for a 1 year drop in survival, from 3 to 2 years. The set in the lower right is the most optimistic, demanding a QOL no smaller than 70, along with a perfect QOL of 100 if survival drops from 3 to 2 years. For a given , a test would conclude the experimental treatment is superior to standard therapy for which E(ξ) = (3,50) if . The decision cut-off pU and sample size could be calibrated so that, for example, the frequentist probability of concluding that the experimental treatment is superior to the standard is at most .05 when in fact ξ = (3,50) and at least .80 for all . Making decisions in this way may be contrasted with conventional tests based on rectangular alternative hypotheses, which cannot accommodate multiple equally desirable target pairs. Also, this particular application of the geometric structure provides an alternative approach to quality-adjusted survival analysis (Gelber, Gelman and Goldhirsch, 1989; Cole, Gelber and Anderson, 1994).
Table 1.
Four sets of elicited target pairs used to determine a desirable set for ξ1 = mean survival time and ξ2 = mean quality of life on scale of 0 to 100, as improvements compared to the historical averages E(ξ1) = 3 years and E(ξ2) = 50.
| Case | |||
|---|---|---|---|
| 1 | (2.5, 100) | (4, 45) | (∞, 40) |
| 2 | (2.5, 100) | (4, 55) | (∞, 50) |
| 3 | (2.0, 90) | (4, 55) | (∞, 50) |
| 4 | (2.0, 100) | (4, 75) | (∞, 70) |
A clinical trial design given in Thall, Wooten and Shpall (2006) will provide a final illustration. In allogeneic transplantation (allotx) to treat hematologic malignancies, the patient's marrow is first ablated using high dose chemotherapy to destroy all the cancer cells. Bone marrow or peripheral blood stem cells from an HLA matched donor are then given to the patient, with the aim that the transplanted cells will engraft and restore the patient's bone marrow. Engraftment occurs when an absolute neutrophil count (ANC) is achieved in the patient's circulating blood at a level indicating that the transplanted cells have begun functioning normally in the bone marrow. Because a low ANC carries a high risk of fatal infection, engraftment is an essential precursor to long term survival and ideally should be achieved as soon as possible. When a matched donor cannot be found, an alternative is to use umbilical cord blood. The major disadvantage of cord blood allotx is that fewer cells are available, so it takes longer time to achieve engraftment, about 4 weeks compared to about 2 weeks with conventional allotx. A common approach to this problem is to grow more cells from those obtained initially (“ex vivo expansion”) before performing the transplant.
In the trial, all patients receive a double cord blood allotx, including cells from two different placentae. Patients are randomized to receive either cells from two normal cords, or from one normal cord and one in which the cells have been expanded. Engraftment is defined as an ANC > 500 cells per mm3 within 6 weeks. Under a competing risks model for the times to engraftment and death without engraftment, accounting for age as a covariate, the two parameters are ξ1 = π = Pr(engraftment) and ξ2 = μ = mean time to engraftment among patients who engraft, both defined at the reference age of 38 years. Thus, larger ξ1 and smaller ξ2 are more desirable. The motivation for basing decisions on ξ = (π, μ) is that using either parameter alone ignores critical information. In comparison with the historical means (π0, μ0) = (.69, 30), the six equally desirable target pairs given by the round points in Figure 3 were elicited from the physician. The curved portion of in Figure 3 is a simple quadratic function μ = α1 + α2π + α3π2 fit to the six points, and the requirement that π ≥ .60 necessitated truncating on the left as shown.
Denote the parameter pairs in the two treatment arms by ξj = (πj, μj) for j = 1,2 and let ξH = (πH, μH) be corresponding historical values, so that E(πH, μH) = (π0, μ0). Define the pairs of differences Δ12 = (π1 − π2, μ1 − μ2) = −Δ21 and ΔjH = (πj − πH, μj − μH) for j = 1,2. The decision rules in the trial are to terminate arm j if
| (8) |
and, if neither arm is terminated early, select treatment 1 at the end of the trial if
| (9) |
with treatment 2 selected if this inequality holds with the indices 1 and 2 reversed. This formulation accommodates the differences Δ1H, Δ2H and Δ12 by first centering the target set, so that is an improvement over (0,0) in the same way that is an improvement over (π0, μ0). Since large corresponds to superiority of treatment j over the historical mean then, for example, large corresponds to superiority of treatment 1 over treatment 2. Details of how the prior and pL were calibrated to obtain a design with good properties are given in Thall, Wooten and Shpall (2006). If a specified improvement p* on the probability domain is desired in order to conclude that one treatment is superior to the other, analogous to a test of hypothesis, then a more stringent version of rule (9) may be applied with p* added to the right-hand side posterior probability.
5. Discussion
The methods discussed here were motivated by the desire to account for trade-offs between desirable and and undesirable clinical outcomes in a way that made sense to physicians who asked for clinical trial designs dealing with such outcomes. The range of possible application seems much more general, although this type of regime has not been applied by the author outside medical statistics. The methodology as described here leads to many possible questions, including how the geometric constructs may be extended to accommodate ξ of dimension > 2, how one may use this sort of structure to construct tests with given frequentist properties, and how such tests may compare with existing frequentist tests based on multiple outcomes.
Figure 4.
The target set used as a basis for decision making in the double cord blood transplantation trial.
Acknowledgments
This research was partially supported by NCI grant RO1 CA 83932. The author is grateful to Harry Whelan for giving his permission to use the IA-tPA trial as an illustration, to Xuemei Wang for assisting with preparation of the figures, and to two referees for their constructive comments.
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- Bryant J, Day R. Incorporating toxicity considerations into the design of two-stage phase II clinical trials. Biometrics. 1995;51:1372–1383. [PubMed] [Google Scholar]
- Cole BF, Gelber RD, Anderson KM. Parametric approaches to quality-adjusted survival analysis. Biometrics. 1994;50:621–631. [PubMed] [Google Scholar]
- Conaway MR, Petroni GR. Bivariate sequential designs for phase II trials. Biometrics. 1995;51:656–664. [PubMed] [Google Scholar]
- Cook JD. Technical Report UTMDABTR-003-06. Dept. of Biostatistics, M.D. Anderson Cancer Center; 2006. Efficacy-toxicity trade-offs based on Lp norms. [Google Scholar]
- Gelber RD, Gelman RS, Goldhirsch A. A quality-of-life-oriented endpoint for comparing therapies. Biometrics. 1989;45:781–795. [PubMed] [Google Scholar]
- Jennison C, Turnbull BW. Group sequential tests for bivariate response: Interim analyses of clinical trials with both efficacy and safety endpoints. Biometrics. 1993;49:741–752. [PubMed] [Google Scholar]
- Jennison C, Turnbull BW. Group Sequential Methods With Applications to Clinical Trials. New York: Chapman and Hall; 2000. [Google Scholar]
- Lavori PW, Dawson R. Dynamic treatment regimes: practical design considerations. Clinical Trials. 2004;1:9–20. doi: 10.1191/1740774s04cn002oa. [DOI] [PubMed] [Google Scholar]
- Murphy SA. Optimal dynamic treatment regimes (with discussion) Journal of the Royal Statistical Society, Series B. 2003;65:331–366. [Google Scholar]
- Thall PF, Cheng SC. Treatment comparisons based on two–dimensional safety and efficacy alternatives in oncology trials. Biometrics. 1999;55:746–753. doi: 10.1111/j.0006-341x.1999.00746.x. [DOI] [PubMed] [Google Scholar]
- Thall PF, Cook John D. Dose-finding based on efficacy-toxicity trade-offs. Biometrics. 2004;60:684–693. doi: 10.1111/j.0006-341X.2004.00218.x. [DOI] [PubMed] [Google Scholar]
- Thall PF, Cook John D, Estey EH. Adaptive dose selection using efficacy-toxicity trade-offs: illustrations and practical considerations. J Biopharmaceutical Statistics. 2006;16:623–638. doi: 10.1080/10543400600860394. [DOI] [PubMed] [Google Scholar]
- Thall PF, Simon RM, Shen Y. Approximate Bayesian evaluation of multiple treatment effects. Biometrics. 2000;56:213–219. doi: 10.1111/j.0006-341x.2000.00213.x. [DOI] [PubMed] [Google Scholar]
- Thall PF, Sung HG, Estey EH. Selecting therapeutic strategies based on efficacy and death in multi-course clinical trials. J American Statistical Association. 2002;97:29–39. [Google Scholar]
- Thall PF, Wooten LH, Shpall EJ. A geometric approach to comparing treatments for rapidly fatal diseases. Biometrics. 2006;62:193–201. doi: 10.1111/j.1541-0420.2005.00434.x. [DOI] [PubMed] [Google Scholar]




