PERFORMANCE GUARANTEES FOR INDIVIDUALIZED TREATMENT RULES

Min Qian; Susan A Murphy

doi:10.1214/10-AOS864

. Author manuscript; available in PMC: 2012 Apr 1.

Published in final edited form as: Ann Stat. 2011 Apr 1;39(2):1180–1210. doi: 10.1214/10-AOS864

PERFORMANCE GUARANTEES FOR INDIVIDUALIZED TREATMENT RULES

Min Qian ^1,^*, Susan A Murphy ^1,^*

PMCID: PMC3110016 NIHMSID: NIHMS266525 PMID: 21666835

Abstract

Because many illnesses show heterogeneous response to treatment, there is increasing interest in individualizing treatment to patients [11]. An individualized treatment rule is a decision rule that recommends treatment according to patient characteristics. We consider the use of clinical trial data in the construction of an individualized treatment rule leading to highest mean response. This is a difficult computational problem because the objective function is the expectation of a weighted indicator function that is non-concave in the parameters. Furthermore there are frequently many pretreatment variables that may or may not be useful in constructing an optimal individualized treatment rule yet cost and interpretability considerations imply that only a few variables should be used by the individualized treatment rule. To address these challenges we consider estimation based on l₁ penalized least squares. This approach is justified via a finite sample upper bound on the difference between the mean response due to the estimated individualized treatment rule and the mean response due to the optimal individualized treatment rule.

Keywords and phrases: decision making, l₁ penalized least squares, Value

1. Introduction

Many illnesses show heterogeneous response to treatment. For example, a study on schizophrenia [12] found that patients who take the same antipsychotic (olanzapine) may have very different responses. Some may have to discontinue the treatment due to serious adverse events and/or acutely worsened symptoms, while others may experience few if any adverse events and have improved clinical outcomes. Results of this type have motivated researchers to advocate the individualization of treatment to each patient [16, 24, 11]. One step in this direction is to estimate each patient’s risk level and then match treatment to risk category [5, 6]. However, this approach is best used to decide whether to treat; otherwise it assumes the knowledge of the best treatment for each risk category. Alternately, there is an abundance of literature focusing on predicting each patient’s prognosis under a particular treatment [10, 28]. Thus an obvious way to individualize treatment is to recommend the treatment achieving the best predicted prognosis for that patient. In general the goal is to use data to construct individualized treatment rules that, if implemented in future, will optimize the mean response.

Consider data from a single stage randomized trial involving several active treatments. A first natural procedure to construct the optimal individualized treatment rule is to maximize an empirical version of the mean response over a class of treatment rules (assuming larger responses are preferred). As will be seen, this maximization is computationally difficult because the mean response of a treatment rule is the expectation of a weighted indicator that is non-continuous and non-concave in the parameters. To address this challenge we make a substitution. That is, instead of directly maximizing the empirical mean response to estimate the treatment rule, we use a two-step procedure that first estimates a conditional mean and then from this estimated conditional mean derives the estimated treatment rule. As will be seen in Section 3, even if the optimal treatment rule is contained in the space of treatment rules considered by the substitute two-step procedure, the estimator derived from the two-step procedure may not be consistent. However if the conditional mean is modeled correctly, then the two-step procedure consistently estimates the optimal individualized treatment rule. This motivates consideration of rich conditional mean models with many unknown parameters. Furthermore there are frequently many pretreatment variables that may or may not be useful in constructing an optimal individualized treatment rule, yet cost and interpretability considerations imply that fewer rather than more variables should be used by the treatment rule. This consideration motivates the use of l₁ penalized least squares (l₁-PLS).

We propose to estimate an optimal individualized treatment rule using a two step procedure that first estimates the conditional mean response using l₁-PLS with a rich linear model and second, derives the estimated treatment rule from estimated conditional mean. For brevity, throughout, we call the two step procedure the l₁-PLS method. We derive several finite sample upper bounds on the difference between the mean response to the optimal treatment rule and the mean response to the estimated treatment rule. All of the upper bounds hold even if our linear model for the conditional mean response is incorrect and to our knowledge are, up to constants, the best available. We use the upper bounds in Section 3 to illuminate the potential mismatch between using least squares in the two-step procedure and the goal of maximizing mean response. The upper bounds in Section 4.1 involve a minimized sum of the approximation error and estimation error; both errors result from the estimation of the conditional mean response. We shall see that l₁-PLS estimates a linear model that minimizes this approximation plus estimation error sum among a set of suitably sparse linear models.

If the part of the model for the conditional mean involving the treatment effect is correct, then the upper bounds imply that, although a surrogate two-step procedure is used, the estimated treatment rule is consistent. The upper bounds provide a convergence rate as well. Furthermore in this setting the upper bounds can be used to inform how to choose the tuning parameter involved in the l₁-penalty to achieve the best rate of convergence. As a by-product, this paper also contributes to existing literature on l₁-PLS by providing a finite sample prediction error bound for the l₁-PLS estimator in the random design setting without assuming the model class contains or is close to the true model.

The paper is organized as follows. In Section 2, we formulate the decision making problem. In Section 3, for any given decision, e.g. individualized treatment rule, we relate the reduction in mean response to the excess prediction error. In Section 4, we estimate an optimal individualized treatment rule via l₁-PLS and provide a finite sample upper bound on the maximal reduction in optimal mean response achieved by the estimated rule. In Section 5, we consider a data dependent tuning parameter selection criterion. This method is evaluated using simulation studies and illustrated with data from the Nefazodone-CBASP trial [13]. Discussions and future work are presented in Section 6.

2. Individualized treatment rules

We use upper case letters to denote random variables and lower case letters to denote values of the random variables. Consider data from a randomized trial. On each subject we have the pretreatment variables X ∈ Inline graphic , treatment A taking values in a finite, discrete treatment space , and a real-valued response R (assuming large values are desirable). An individualized treatment rule (ITR) d is a deterministic decision rule from into the treatment space .

Denote the distribution of (X, A, R) by P. This is the distribution of the clinical trial data; in particular, denote the known randomization distribution of A given X by p(·|X). The likelihood of (X, A, R) under P is then f₀(x)p(a|x)f₁(r|x, a), where f₀ is the unknown density of X and f₁ is the unknown density of R conditional on (X, A). Denote the expectations with respect to the distribution P by an E. For any ITR d : Inline graphic → , let P^d denote the distribution of (X, A, R) in which d is used to assign treatments. Then the likelihood of (X, A, R) under P^d is f₀(x)1_a₌_d₍_x₎f₁(r|x, a). Denote expectations with respect to the distribution P^d by an E^d. The Value of d is defined as V (d) = E^d(R). An optimal ITR, d₀, is a rule that has the maximal Value, i.e.

d_{0} \in arg max_{d} V (d),

where the argmax is over all possible decision rules. The Value of d₀, V(d₀), is the optimal Value.

Assume P[p(a|X) > 0] = 1 for all a ∈ Inline graphic (i.e. all treatments in are possible for all values of X a.s.). Then P^d is absolutely continuous with respect to P and a version of the Radon-Nikodym derivative is dP^d/dP = 1_a₌_d₍_x₎/p(a|x). Thus the Value of d satisfies

V (d) = E^{d} (R) = \int {RdP}^{d} = \int R \frac{{d P}^{d}}{d P} d P = E [\frac{1_{A = d (X)}}{p (A ∣ X)} R] .

(2.1)

Our goal is to estimate d₀, i.e. the ITR that maximizes (2.1), using data from distribution P. When X is low dimensional and the best rule within a simple class of ITRs is desired, empirical versions of the Value can be used to construct estimators [21, 27]. However if the best rule within a larger class of ITRs is of interest, these approaches are no longer feasible.

Define Q₀(X,A) ≜ E(R|X,A) (Q₀(X,A) is sometimes called the “Quality” of treatment a at observation x). It follows from (2.1) that for any ITR d,

V (d) = E [\frac{1_{A = d (X)}}{p (A ∣ X)} Q_{0} (X, A)] = E [\sum_{a \in A} 1_{d (X) = a} Q_{0} (X, a)] = E [Q_{0} (X, d (X))] .

Thus V (d₀) = E[Q₀(X, d₀(X))] ≤ E[max_a_∈Q₀(X, a)]. On the other hand, by the definition of d₀,

V (d_{0}) \geq V (d) |_{d (X) \in arg {max}_{a \in A} Q_{0} (X, a)} = E [max_{a \in A} Q_{0} (X, a)] .

Hence an optimal ITR satisfies d₀(X) ∈ arg max_a_∈ Q₀(X, a) a.s.

3. Relating the reduction in Value to excess prediction error

The above argument indicates that the estimated ITR will be of high quality (i.e. have high Value) if we can estimate Q₀ accurately. In this section, we justify this by providing a quantitative relationship between the Value and the prediction error.

Because Inline graphic is a finite, discrete treatment space, given any ITR, d, there exists a square integrable function Q : × → ℝ for which d(X) ∈ arg max_a Q(X, a) a.s. Let L(Q) ≜ E[R − Q(X,A)]² denote the prediction error of Q (also called the mean quadratic loss). Suppose that Q₀ is square integrable and that the randomization probability satisfies p(a|x) ≥ S⁻¹ for an S > 0 and all (x, a) pairs. Murphy [23] showed that

V (d_{0}) - V (d) \leq 2 S^{1 / 2} {[L (Q) - L (Q_{0})]}^{1 / 2} .

(3.1)

Intuitively, this upper bound means that if the excess prediction error of Q (i.e. E(R − Q)² − E(R − Q₀)²) is small, then the reduction in Value of the associated ITR d (i.e. V(d₀) − V(d)) is small. Furthermore the upper bound provides a rate of convergence for an estimated ITR. For example, suppose Q₀ is linear, that is Q₀ = Φ(X, A)θ₀ for a given vector-valued basis function Φ on Inline graphic × and an unknown parameter θ₀. And Suppose we use a correct linear model for Q₀ (here “linear” means linear in parameters), say the model = {Φ(X, A)θ : θ → ℝ^dim^(Φ)} or a linear model containing with dimension of parameters fixed in n. If we estimate θ by least squares and denote the estimator by θ̂, then the prediction error of Q̂ = Φθ̂ converges to L(Q₀) at rate 1/n under mild regularity conditions. This together with inequality (3.1) implies that the Value obtained by the estimated ITR, d̂(X) ∈ arg max_a Q̂(X, a), will converge to the optimal Value at rate at least $1 / \sqrt{n}$ .

In the following theorem, we improve this upper bound in two aspects. First, we show that an upper bound with exponent larger than 1/2 can be obtained under a margin condition, which implicitly implies a faster rate of convergence. Second, it turns out that the upper bound need only depend on one term in the function Q; we call this the treatment effect term, T. For any square integrable Q, the associated treatment effect term is defined as T(X,A) ≜ Q(X,A) − E[Q(X,A)|X]. Note that d(X) ∈ arg max_a T(X, a) = arg max_a Q(X, a) a.s. Similarly, the true treatment effect term is given by

T_{0} (X, A) ≜ Q_{0} (X, A) - E [Q_{0} (X, A) ∣ X] .

(3.2)

T₀(x, a) is the centered effect of treatment A = a at observation X = x; d₀(X) ∈ arg max_a T₀(X, a).

Theorem 3.1

Suppose p(a|x) ≥ S⁻¹ for a positive constant S for all (x, a) pairs. Assume there exists some constants C > 0 and α ≥ 0 such that

P (max_{a \in A} T_{0} (X, a) - max_{a \in A \ arg {max}_{a \in A} T_{0} (X, a)} T_{0} (X, a) \leq ε) \leq C ε^{α}

(3.3)

for all positive ε. Then for any ITR d : Inline graphic → and square integrable function Q : × → ℝ such that d(X) ∈ arg max_a_∈ Q(X, a) a.s., we have

V (d_{0}) - V (d) \leq C^{'} {[L (Q) - L (Q_{0})]}^{(1 + α) / (2 + α)},

(3.4)

and

V (d_{0}) - V (d) \leq C^{'} {[E {(T (X, A) - T_{0} (X, A))}^{2}]}^{(1 + α) / (2 + α)} .

(3.5)

where C′ = (2^2+3αS^1+αC)^1/(2+α).

The proof of Theorem 3.1 is in Appendix A.1.

Remarks

We set the second maximum in (3.3) to −∞ if for an x, T₀(x, a) is constant in a and thus the set \arg max_a_∈ T₀(x, a) = ∅.
Condition (3.3) is similar to the margin condition in classification [25, 18, 32]; in classification this assumption is often used to obtain sharp upper bounds on the excess 0 − 1 risk in terms of other surrogate risks [1]. Here can be viewed as the “margin” of T₀ at observation X = x. It measures the difference in mean responses between the optimal treatment(s) and the best suboptimal treatment(s) at x. For example, suppose X ~ U[−1, 1], P(A = 1|X) = P(A = −1|X) = 1/2 and T₀(X, A) = XA. Then the margin condition holds with C = 1/2 and α = 1. Note the margin condition does not exclude multiple optimal treatments for any observation x. However, when α > 0, it does exclude suboptimal treatments that yield a conditional mean response very close to the largest conditional mean response for a set of x with nonzero probability.
For C = 1, α = 0, Condition (3.3) always holds for all ε > 0; in this case (3.4) reduces to (3.1).
The larger the α, the larger the exponent (1 + α)/(2 + α) and thus the stronger the upper bounds in (3.4) and (3.5). However the margin condition is unlikely to hold for all ε if α is very large. An alternate margin condition and upper bound are as follows.

Suppose p(a|x) ≥ S⁻¹ for all (x, a) pairs. Assume there is an ε > 0, such that
$P (max_{a \in A} T_{0} (X, a) - max_{a \in A \ arg {max}_{a \in A} T_{0} (X, a)} T_{0} (X, a) < ε) = 0.$ (3.6)

Then V(d₀) − V(d) ≤ 4S[L(Q) − L(Q₀)]/ε and V(d₀) − V(d) ≤ 4SE(T − T₀)²/ε.

The proof is essentially the same as that of Theorem 3.1 and is omitted. Condition (3.6) means that T₀ evaluated at the optimal treatment(s) minus T₀ evaluated at the best suboptimal treatment(s) is bounded below by a positive constant for almost all X observations. If X assumes only a finite number of values, then this condition always holds, because we can take ε to be the smallest difference in T₀ when evaluated at the optimal treatment(s) and the suboptimal treatment(s) (note that if T₀(x, a) is constant for all a ∈ for some observation X = x, then all treatments are optimal for that observation).
Inequality (3.5) cannot be improved in the sense that choosing T = T₀ yields zero on both sides of the inequality. Moreover an inequality in the opposite direction is not possible, since each ITR is associated with many non-trivial T-functions. For example, suppose X ~ U[−1, 1], P(A = 1|X) = P(A = −1|X) = 1/2 and T₀(X, A) = (X − 1/3)²A. The optimal ITR is d₀(X) = 1 a.s. Consider T (X, A) = θA. Then maximizing T(X, A) yields the optimal ITR as long as θ > 0. This means that the left hand side (LHS) of (3.5) is zero, while the right hand side (RHS) is always positive no matter what value θ takes.

Theorem 3.1 supports the approach of minimizing the estimated prediction error to estimate Q₀ or T₀ and then maximizing this estimator over a ∈ Inline graphic to obtain an ITR. It is natural to expect that even when the approximation space used in estimating Q₀ or T₀ does not contain the truth, this approach will provide the best (highest Value) of the considered ITRs. Unfortunately this does not occur due to the mismatch between the loss functions (weighted 0–1 loss and the quadratic loss). This mismatch is indicated by remark 5 above. More precisely, note that the approximation space, say Inline graphic for Q₀, places implicit restrictions on the class of ITRs that will be considered. In effect the class of ITRs is = {d(X) ∈ arg max_a Q(X, a) : Q ∈ }. It turns out that minimizing the prediction error may not result in the ITR in that maximizes the Value. This occurs when the approximation space Inline graphic does not provide a treatment effect term close to the treatment effect term in Q₀. In the following toy example, the optimal ITR d₀ belongs to , yet the prediction error minimizer over does not yield d₀.

A toy example

Suppose X is uniformly distributed in [−1, 1], A is binary {−1, 1} with probability 1/2 each and is independent of X, and R is normally distributed with mean Q₀(X, A) = (X −1/3)²A and variance 1. It is easy to see that the optimal ITR satisfies d₀(X) = 1 a.s. and V(d₀) = 4/9. Consider approximation space Inline graphic = {Q(X, A; θ) = (1, X, A, XA)θ : θ ∈ ℝ⁴} for Q₀. Thus the space of ITRs under consideration is = {d(X) = sign(θ₃+θ₄X) : θ₃, θ₄ ∈ ℝ}. Note that d₀ ∈ since d₀(X) can be written as sign(θ₃ + θ₄X) for any θ₃ > 0 and θ₄ = 0. d₀ is the best treatment rule in Inline graphic . However, minimizing the prediction error L(Q) over yields Q^*(X, A) = (4/9−2/3X)A. The ITR associated with Q^* is d^*(X) = arg max_a_∈{−1_,_1} Q^*(X, a) = sign(2/3 − X), which has lower Value than $d_{0} (V (d^{*}) = E [\frac{1_{A (2 / 3 - X) > 0} R}{1 / 2}] = 29 / 81 < V (d_{0}))$ .

4. Estimation via l₁-penalized least squares

To deal with the mismatch between minimizing the prediction error and maximizing the Value discussed in the prior section, we consider a large linear approximation space Inline graphic for Q₀. Since overfitting is likely (due to the potentially large number of pretreatment variables and/or large approximation space for Q₀) we use penalized least squares (see Section S.1 of the supplementary material for further discussion of the overfitting problem). Furthermore we use l₁ penalized least squares (l₁-PLS, [31]) as the l₁ penalty does some variable selection and as a result will lead to ITRs that are cheaper to implement (fewer variables to collect per patient) and easier to interpret. See Section 6 for the discussion of other potential penalization methods.

Let ${(X_{i}, A_{i}, R_{i})}_{i = 1}^{n}$ represent i.i.d. observations on n subjects in a randomized trial. For convenience, we use E_n to denote the associated empirical expectation (i.e. $E_{n} f = \sum_{i = 1}^{n} f (X_{i}, A_{i}, R_{i}) / n$ for any real-valued function f on Inline graphic × × ℝ). Let := {Q(X, A; θ) = Φ(X, A) θ, θ ∈ ℝ^J} be the approximation space for Q₀, where φ(X, A) = (φ₁(X, A), …, φ_J (X, A)) is a 1 by J vector composed of basis functions on ×, θ is a J by 1 parameter vector, and J is the number of basis functions (for clarity here J will be fixed in n, see Appendix A.2 for results with J increasing as n increases). The l₁-PLS estimator of θ is

{\hat{θ}}_{n} = arg min_{θ \in R^{J}} {E_{n} {[R - Φ (X, A) θ]}^{2} + λ_{n} \sum_{j = 1}^{J} {\hat{σ}}_{j} ∣ θ_{j} ∣},

(4.1)

where σ̂_j = [E_nφ_j(X, A)²]^1/2, θ_j is the j^th component of θ and λ_n is a tuning parameter that controls the amount of penalization. The weights σ̂_j’s are used to balance the scale of different basis functions; these weights were used in Bunea et al. [4] and van de Geer [33]. In some situations, it is natural to penalize only a subset of coefficients and/or use different weights in the penalty; see Section S.2 of the supplementary material for required modifications. The resulting estimated ITR satisfies

{\hat{d}}_{n} (X) \in arg max_{a \in A} Φ (X, a) {\hat{θ}}_{n} .

(4.2)

4.1. Performance guarantee for the l₁-PLS

In this section we provide finite sample upper bounds on the difference between the optimal Value and the Value obtained by the l₁-PLS estimator in terms of the prediction errors resulting from the estimation of Q₀ and T₀. These upper bounds guarantee that if Q₀ (or T₀) is consistently estimated, the estimator of d₀ will be consistent and will inherit a rate of convergence from the rate of convergence of the estimator of Q₀ (or T₀). Perhaps more importantly, the finite sample upper bounds provided below do not require the assumption that either Q₀ or T₀ is consistently estimated. Thus each upper bound includes approximation error as well as estimation error. The estimation error decreases with decreasing model sparsity and increasing sample size. An “oracle” model for Q₀ (or T₀) minimizes the sum of these two errors among suitably sparse linear models (see remark 2 after Theorem 4.3 for a precise definition of the oracle model). In finite samples, the upper bounds imply that d̂_n, the ITR produced by the l₁-PLS method, will have Value roughly as if the l₁-PLS method detects the sparsity of the oracle model and then estimates from the oracle model using ordinary least squares (see remark 3 below).

Define the prediction error minimizer θ^* ∈ ℝ^J by

θ^{*} \in arg min_{θ \in R^{J}} L (Φ θ) = arg min_{θ \in R^{J}} E {(R - Φ θ)}^{2} .

(4.3)

For expositional simplicity assume that θ^* is unique, and define the sparsity of θ ∈ ℝ^J by its l₀ norm, ||θ||₀ (see Appendix A.2 for a more general setting, where θ^* is not unique and a laxer definition of sparsity is used). As discussed above, for finite n, instead of estimating θ^*, the l₁-PLS estimator θ̂_n estimates a parameter, $θ_{n}^{* *}$ , possessing small prediction error but with controlled sparsity. For any bounded function f on Inline graphic × , let ||f||_∞ ≜ sup_x_∈,_a_∈|f(x,a)|. $θ_{n}^{* *}$ lies in the set of parameters Θ_n defined by

Θ_{n} ≜ {θ \in R^{J} : {| | Φ (θ - θ^{*}) | |}_{\infty} \leq η, max_{j = 1, \dots, J} | \frac{E [φ_{j} Φ (θ - θ^{*})]}{σ_{j}} | \leq 15 η \sqrt{\frac{log (J n)}{n}} and {| | θ | |}_{0} \leq \frac{β}{640 U} \sqrt{\frac{n}{log (n J)}}},

(4.4)

where $σ_{j} = {(E φ_{j}^{2})}^{1 / 2}$ , and η, β and U are positive constants that will be defined in Theorem 4.1.

The first two conditions in (4.4) restrict Θ_n to θ’s with controlled distance in sup norm and with controlled distance in prediction error via first order derivatives (note that $∣ E [φ_{j} Φ (θ - θ^{*})] / σ_{j} ∣ = ∣ \partial L (Φ θ) / \partial θ_{j} - \partial L (Φ θ^{*}) / \partial θ_{j}^{*} ∣ / 2 σ_{j})$ . The third condition restricts Θ_n to sparse θ’s. Note that as n increases this sparsity requirement becomes laxer, ensuring that θ^* ∈ Θ_n for sufficiently large n.

When Θ_n is non-empty, $θ_{n}^{* *}$ is given by

θ_{n}^{* *} = arg min_{θ \in Θ_{n}} [L (Φ θ) + 3 {| | θ | |}_{0} λ_{n}^{2} / β] .

(4.5)

Note that $θ_{n}^{* *}$ is at least as sparse as θ^* since by (4.3), $L (Φ θ) + 3 {| | θ | |}_{0} λ_{n}^{2} / β > L (Φ θ^{*}) + 3 {| | θ^{*} | |}_{0} λ_{n}^{2} / β$ for any θ such that ||θ||₀ > ||θ^*||₀.

The following theorem provides a finite sample performance guarantee for the ITR produced by l₁-PLS method. Intuitively, this result implies that if Q₀ can be well approximated by the sparse linear representation $θ_{n}^{* *}$ (so that both $L (Φ θ_{n}^{* *}) - L (Q_{0})$ and ${| | θ_{n}^{* *} | |}_{0}$ are small), then d̂_n will have Value close to the optimal Value in finite samples.

Theorem 4.1

Suppose p(a|x) ≥ S⁻¹ for a positive constant S for all (x, a) pairs and the margin condition (3.3) holds for some C > 0, α ≥ 0 and all positive ε. Assume

the error terms ε_i = R_i − Q₀(X_i, A_i), i = 1, …, n, are independent of (X_i, A_i), i …, n and are i.i.d. with E(ε_i) = 0 and $E [{∣ ε_{i} ∣}^{l}] \leq \frac{l!}{2} c^{l - 2} σ^{2}$ for some c, σ² > 0 for all l ≥ 2;
there exist finite, positive constants U and η such that max_j=1,…,J ||φ_j||_∞/σ_j ≤ U and ||Q₀ − Φθ^*||_∞ ≤ η; and
E[(φ₁/σ₁, …, φ_J/σ_J )^T (φ₁/σ₁,…, φ_J/σ_J)] is positive definite, and the smallest eigenvalue is denoted by β.

Consider the estimated ITR d̂_n defined by (4.2) with tuning parameter

λ_{n} \geq k \sqrt{\frac{log (J n)}{n}},

(4.6)

where k = 82 max{c, σ, η}. Let Θ_n be the set defined in (4.4). Then for any n ≥ 24U² log(Jn) and for which Θ_n is non-empty, we have, with probability at least 1 − 1/n, that

V (d_{0}) - V ({\hat{d}}_{n}) \leq C^{'} {[min_{θ \in Θ_{n}} (L (Φ θ) - L (Q_{0}) + 3 {| | θ | |}_{0} λ_{n}^{2} / β)]}^{\frac{1 + α}{2 + α}},

(4.7)

where C′ = (2^2+3αS^1+αC)^1/(2+α).

The result follows from inequality (3.4) in Theorem 3.1 and inequality (4.10) in Theorem 4.3. Similar results in a more general setting can be obtained by combining (3.4) with inequality (A.7) in Appendix A.2.

Remarks

Note that $θ_{n}^{* *}$ is the minimizer of the upper bound on the RHS of (4.7) and that $θ_{n}^{* *}$ is contained in the set { $θ_{n}^{*, (m)}$ : m ⊂ {1, …, J}}. Each $θ_{n}^{*, (m)}$ satisfies $θ_{n}^{*, (m)} = arg {min}_{{θ \in Θ_{n} : θ_{j} = 0 for all j \notin m}} L (Φ θ)$ ; that is, $θ_{n}^{*, (m)}$ minimizes the prediction error of the model indexed by the set m (i.e. model {Σ_j_∈_m φ_jθ_j : θ_j ∈ ℝ}) (within Θ_n). For each $θ_{n}^{*, (m)}$ , the first term in the upper bound in (4.7) (i.e. $L (Φ θ_{n}^{*, (m)}) - L (Q_{0})$ ) is the approximation error of the model indexed by m within Θ_n. As in van de Geer [33], we call the second term $3 {| | θ_{n}^{*, (m)} | |}_{0} λ_{n}^{2} / β$ the estimation error of the model indexed by m. To see why, first put $λ_{n} = k \sqrt{log (J n) / n}$ . Then, ignoring the log(n) factor, the second term is a function of the sparsity of model m relative to the sample size, n. Up to constants, the second term is a “tight” upper bound for the estimation error of the OLS estimator from model m, where “tight” means that the convergence rate in the bound is the best known rate. Note that $θ_{n}^{* *}$ is the parameter that minimizes the sum of the two errors over all models. Such a model (the model corresponding to $θ_{n}^{* *}$ ) is called an oracle model. The log(n) factor in the estimation error is the price paid for not knowing the sparsity of the oracle model. By using the l₁-PLS method, we pay by a factor of log(n) in the estimation error and as an exchange, the l₁-PLS estimator behaves roughly as if it knew the sparsity of the oracle model and as if is was estimated from the oracle model using OLS. Thus the log(n) factor can be viewed as the price paid for not knowing the sparsity of the oracle model and thus having to conduct model selection. See remark 2 after Theorem A.1 for the precise definition of the oracle model and its relationship to $θ_{n}^{* *}$ .
Suppose λ_n = o(1). Then in large samples the estimation error term $3 {| | θ | |}_{0} λ_{n}^{2} / β$ is negligible. In this case, $θ_{n}^{* *}$ is close to θ^*. When the model Φθ^* approximates Q₀ sufficiently well, we see that setting λ_n equal to its lower bound in (4.6) provides the fastest rate of convergence of the upper bound to zero. More precisely, suppose Q₀ = Φθ^* (i.e. L(Φθ^*) − L((Q₀) = 0). Then inequality (4.7) implies that V(d₀) − V(d̂_n) ≤ O_p ((log n/n)⁽¹⁺^α⁾^/⁽²⁺^α⁾). A convergence in mean result is presented in Corollary 4.1.
In finite samples, the estimation error $3 {| | θ | |}_{0} λ_{n}^{2} / β$ is nonnegligible. The argument of the minimum in the upper bound (4.7), $θ_{n}^{* *}$ , minimizes prediction error among parameters with controlled sparsity. In remark 2 after Theorem 4.3, we discuss how this upper bound is a tight upper bound for the OLS estimator from an oracle model in the step-wise model selection setting. In this sense, inequality (4.7) implies that decision rule produced by the l₁-PLS method will have a reduction in Value roughly as if it knew the sparsity of the oracle model and were estimated from the oracle model using OLS.
Assumptions 1–3 in Theorem 4.1 are employed to derive the finite sample prediction error bound for the l₁-PLS estimator θ̂_n defined in (4.1). Below we briefly discuss these assumptions.

Assumption 1 implicitly implies that the error terms do not have heavy tails. This condition is often assumed to show that the sample mean of a variable is concentrated around its true mean with a high probability. It is easy to verify that this assumption holds if each ε_i is bounded. Moreover, it also holds for some commonly used error distributions that have unbounded support, such as the normal or double exponential.

Assumption 2 is also used to show the concentration of the sample mean around the true mean. It is possible to replace the boundedness condition by a moment condition similar to Assumption 1. This assumption requires that all basis functions and the difference between Q₀ and its best linear approximation are bounded. Note that we do not assume to be a good approximation space for Q₀. However, if Φθ* approximates Q₀ well, η will be small, which will result in a smaller upper bound in (4.7). In fact, in the generalized result (Theorem A.1) we allow U and η to increase in n.

Assumption 3 is employed to avoid collinearity. In fact, we only need
$E {[Φ (θ^{'} - θ)]}^{2} {| | θ | |}_{0} \geq β {(\sum_{j \in M_{0} (θ)} σ_{j} ∣ θ_{j}^{'} - θ_{j} ∣)}^{2},$ (4.8)

for θ, θ′ belonging to a subset of ℝ^J (see Assumption A.3), where M₀(θ) ≜ {j = 1,…, J: θ_j ≠ 0}. Condition (4.8) has been used in van de Geer [33]. This condition is also similar to the restricted eigenvalue assumption in Bickel et al. [3] in which E is replaced by E_n, and a fixed design matrix is considered. Clearly, Assumption 3 is a sufficient condition for (4.8). In addition, condition (4.8) is satisfied if the correlation |Eφ_jφ_k|/(σ_jσ_k) is small for all k ∈ M₀(θ), j ≠ k and a subset of θ’s (similar results in a fixed design setting have been proved in Bickel et al. [3]. The condition on correlation is also known as “mutual coherence” condition in Bunea at al. [4]). See Bickel et al. [3] for other sufficient conditions for (4.8).

The above upper bound for V (d₀) − V (d̂_n) involves L(Φθ) − L(Q₀), which measures how well the conditional mean function Q₀ is approximated by Inline graphic . As we have seen in Section 3, the quality of the estimated ITR only depends on the estimator of the treatment effect term T₀. Below we provide a strengthened result in the sense that the upper bound depends only on how well we approximate the treatment effect term.

First we identify terms in the linear model Inline graphic that approximate T₀ (recall that T₀(X,A) ≜ Q₀(X,A) − E[Q₀(X,A)|X]). Without loss of generality, we rewrite the vector of basis functions as Φ(X, A) = (Φ⁽¹⁾(X), Φ⁽²⁾(X, A)), where Φ⁽¹⁾ = (φ₁(X), …, φ_J⁽¹⁾(X)) is composed of all components in Φ that do not contain A and Φ⁽²⁾ = (φ_J⁽¹⁾+1(X, A), …, φ_J (X, A)) is composed of all components in Φ that contain A. Since A takes only finite values and the randomization distribution p(a|x) is known, we can code A so that E[Φ⁽²⁾(X, A)^T|X] = 0 a.s. (see Section 5.2 and Appendix A.3 for examples). For any θ = (θ₁, …, θ_J)^T ∈ ℝ^J, denote θ⁽¹⁾ = (θ₁, …, θ_J⁽¹⁾)^T and θ⁽²⁾ = (θ_J⁽¹⁾+1, …, θ_J)^T. Then Φ⁽¹⁾θ⁽¹⁾ approximates E(Q₀(X, A)|X) and Φ⁽²⁾θ⁽²⁾ approximates T₀.

The following theorem implies that if the treatment effect term T₀ can be well approximated by a sparse representation, then d̂_n will have Value close to the optimal Value.

Theorem 4.2

Suppose p(a|x) ≥ S⁻¹ for a positive constant S for all (x, a) pairs and the margin condition (3.3) holds for some C > 0, α ≥ 0 and all positive ε. Assume E[Φ⁽²⁾(X, A)^T|X] = 0 a.s. Suppose Assumptions 1 – 3 in Theorem 4.1 hold. Let d̂_n be the estimated ITR with λ_n satisfying condition (4.6). Let Θ_n be the set defined in (4.4). Then for any n ≥ 24U² log(Jn) and for which Θ_n is non-empty, we have, with probability at least 1 − 1/n, that

V (d_{0}) - V ({\hat{d}}_{n}) \leq C^{'} {[min_{θ \in Θ_{n}} (E {(Φ^{(2)} θ^{(2)} - T_{0})}^{2} + 5 {| | θ^{(2)} | |}_{0} λ_{n}^{2} / β)]}^{\frac{1 + α}{2 + α}},

(4.9)

where C′ = (2^2+3αS^1+αC)^1/(2+α).

The result follows from inequality (3.5) in Theorem 3.1 and inequality (4.11) in Theorem 4.3.

Remarks

Inequality (4.9) improves inequality (4.7) in the sense that it guarantees a small reduction in Value of d̂_n as long as the treatment effect term T₀ is well approximated by a sparse linear representation; it does not require that the entire conditional mean function Q₀ be well approximated. In many situations Q₀ may be very complex, but T₀ could be very simple. This means that T₀ is much more likely to be well approximated as compared to Q₀ (indeed, if there is no difference between treatments, then T₀ ≡ 0).
Inequality (4.9) cannot be improved in the sense that if there is no treatment effect (i.e. T₀ ≡ 0), then both sides of the inequality are zero. This result implies that minimizing the penalized empirical prediction error indeed yields high Value (at least asymptotically) if T₀ can be well approximated.

The following asymptotic result follows from Theorem 4.2. Note that when E[Φ⁽²⁾(X, A)^T|X] = 0 a.s. (see Section 5 for examples), E(Φθ − Q₀)² = E(Φ⁽¹⁾θ⁽¹⁾ − [Q₀ − E(Q₀|X)])² + E(Φ⁽²⁾θ⁽²⁾ − T₀)². Thus the estimation of the treatment effect term T₀ is asymptotically separated from the estimation of the main effect term Q₀ − E(Q₀|X). In this case, Φ⁽²⁾θ^(2),* is the best linear approximation of the treatment effect term T₀, where θ^(2),* is the vector of components in θ^* corresponding to Φ⁽²⁾.

Corollary 4.1

Suppose p(a|x) ≥ S⁻¹ for a positive constant S for all (x, a) pairs and the margin condition (3.3) holds for some C > 0, α ≥ 0 and all positive ε. Assume E[Φ⁽²⁾(X, A)^T|X] = 0 a.s. In addition, suppose Assumptions 1 – 3 in Theorem 4.1 hold. Let d̂_n be the estimated ITR with tuning parameter $λ_{n} = k_{1} \sqrt{log (J n) / n}$ for a constant k₁ ≥ 82 max{c, σ, η}. If T₀(X, A) = Φ⁽²⁾θ^(2),*, then

V (d_{0}) - E [V ({\hat{d}}_{n})] = O ({(log n / n)}^{(1 + α) / (2 + α)}) .

This result provides a guarantee on the convergence rate of V (d̂_n) to the optimal Value. More specifically, it means that if T₀ is correctly approximated, then the Value of d̂_n will converge to the optimal Value in mean at rate at least as fast as (log n/n)⁽¹⁺^α^)/(2+^α⁾ with appropriate choice of λ_n.

4.2. Prediction error bound for the l₁-PLS estimator

In this section we provide a finite sample upper bound for the prediction error of the l₁-PLS estimator θ̂_n. This result is needed to prove Theorem 4.1. Furthermore this result strengthens existing literature on l₁-PLS method in prediction. Finite sample prediction error bounds for the l₁-PLS estimator in the random design setting have been provided in Bunea et al. [4] for quadratic loss, van de Geer [33] mainly for Lipschitz loss, and Koltchinskii [15] for a variety of loss functions. With regards quadratic loss, Koltchinskii [15] requires the response Y is bounded, while both Bunea et al. [4] and van de Geer [33] assumed the existence of a sparse θ ∈ ℝ^J such that E(Φθ − Q₀)² is upper bounded by a quantity that decreases to 0 at a certain rate as n → ∞ (by permitting J to increase with n so Φ depends on n as well). We improve the results in the sense that we do not make such assumptions (see Appendix A.2 for results when Φ, J are indexed by n and J increases with n).

As in the prior sections, the sparsity of θ is measured by its l₀ norm, ||θ||₀ (see the Appendix A.2 for proofs with a laxer definition of sparsity). Recall that the parameter, $θ_{n}^{* *}$ defined in (4.5) has small prediction error and controlled sparsity.

Theorem 4.3

Suppose Assumptions 1–3 in Theorem 4.1 hold. For any η₁ ≥ 0, Let θ̂_n be the l₁-PLS estimator defined by (4.1) with tuning parameter λ_n satisfying condition (4.6). Let Θ_n be the set defined in (4.4). Then for any n ≥ 24U² log(Jn) and for which Θ_n is non-empty, we have, with probability at least 1 − 1/n, that

L (Φ {\hat{θ}}_{n}) \leq min_{θ \in Θ_{n}} (L (Φ θ) + 3 {| | θ | |}_{0} λ_{n}^{2} / β) = L (Φ θ_{n}^{* *}) + 3 {| | θ_{n}^{* *} | |}_{0} λ_{n}^{2} / β .

(4.10)

Furthermore, suppose E[Φ⁽²⁾(X, A)^T|X] = 0 a.s. Then with probability at least 1 − 1/n,

E {(Φ^{(2)} {\hat{θ}}_{n}^{(2)} - T_{0})}^{2} \leq min_{θ \in Θ_{n}} (E {(Φ^{(2)} θ^{(2)} - T_{0})}^{2} + 5 {| | θ^{(2)} | |}_{0} λ_{n}^{2} / β),

(4.11)

The results follow from Theorem A.1 in Appendix A.2 with ρ = 0, γ= 1/8, η₁ = η₂ = η, t = log 2n and some simple algebra (notice that Assumption 3 in Theorem 4.1 is a sufficient condition for Assumptions A.3 and A.4).

Remarks

Inequality (4.11) provides a finite sample upper bound on the mean square difference between T₀ and its estimator. This result is used to prove Theorem 4.2. The remarks below discuss how inequality (4.10) contributes to the l₁-penalization literature in prediction.

The conclusion of Theorem 4.3 holds for all choices of λ_n that satisfy (4.6). Suppose λ_n = o(1), then $L (Φ θ_{n}^{* *}) - L (Φ θ^{*}) \to 0$ as n → ∞ (since ||θ||₀ is bounded). Then (4.10) implies that L(Φθ̂_n) − L(Φθ^*) → 0 in probability. To achieve the best rate of convergence, equal sign should be taken in (4.6).

Note that

θ_{n}^{* *}

minimizes

L (Φ θ) - L (Q_{0}) + 3 {| | θ | |}_{0} λ_{n}^{2} / β

. Below we demonstrate that the minimum of

L (Φ θ) - L (Q_{0}) + 3 {| | θ | |}_{0} λ_{n}^{2} / β

can be viewed as the approximation error plus a “tight” upper bound of the estimation error of an “oracle” in the stepwise model selection framework (when “=” is taken in (4.6)). Here “tight” means the convergence rate in the bound is the best known rate, and “oracle” is defined as follows. Let m denote a non-empty subset of the index set {1, …, J}. Then each m represents a model which uses a non-empty subset of {φ₁, …, φ_J} as basis functions (there are 2^J − 1 such subsets). Define

{\hat{θ}}_{n}^{(m)} = arg {min}_{{θ \in R^{J} : θ_{j} = 0 for all j \notin m}} E_{n} {(R - Φ θ)}^{2}

and θ^*,(^m⁾ = arg min_{{θ ∈ ℝ^J:θ_j=0} for all _j_∉ _m_}L(Φθ). In this setting, an ideal model selection criterion will pick model m^* such that

L (Φ {\hat{θ}}_{n}^{(m^{*})}) = {inf}_{m} L (Φ {\hat{θ}}_{n}^{(m)})

{\hat{θ}}_{n}^{(m^{*})}

is referred as an “oracle” in Massart [20]. Note that the excess prediction error of each

{\hat{θ}}_{n}^{(m)}

can be written as

L (Φ {\hat{θ}}_{n}^{(m)}) - L (Q_{0}) = [L (Φ θ^{*, (m)}) - L (Q_{0})] + [L (Φ {\hat{θ}}_{n}^{(m)}) - L (Φ θ^{*, (m)})],

where the first term is called the approximation error of model m and the second term is the estimation error. It can be shown that [2] for each model m and x_m > 0, with probability at least 1 − exp(−x_m),

L (Φ {\hat{θ}}_{n}^{(m)}) - L (Φ θ^{*, (m)}) \leq constant \times (\frac{x_{m} + ∣ m ∣ log (n / ∣ m ∣)}{n})

under appropriate technical conditions, where |m| is the cardinality of the index set m. To our knowledge this is the best rate known so far. Taking x_m = log n + |m| log J and using the union bound argument, we have with probability at least 1 − O(1/n),

\begin{array}{l} L (Φ_{n} {\hat{θ}}_{n}^{(m^{*})}) - L (Q_{0}) \\ = min_{m} ([L (Φ θ^{*, (m)}) - L (Q_{0})] + L (Φ {\hat{θ}}_{n}^{(m)}) - L (Φ θ^{*, (m)})) \\ \leq min_{m} ([L (Φ θ^{*, (m)}) - L (Q_{0})] + constant \times \frac{∣ m ∣ log (J n)}{n}) \\ = min_{θ} ([L (Φ θ) - L (Q_{0})] + constant \times \frac{{| | θ | |}_{0} log (J n)}{n}) . \end{array}

(4.12)

On the other hand, take λ_n so that condition (4.6) holds with “=”. (4.10) implies that, with probability at least 1 − 1/n,

L (Φ {\hat{θ}}_{n}) - L (Q_{0}) \leq min_{θ \in Θ_{n}} ([L (Φ θ) - L (Q_{0})] + constant \times \frac{{| | θ | |}_{0} log (J n)}{n}),

which is essentially (4.12) with the constraint of θ ∈ Θ_n. (The “constant” in the above inequalities may take different values.) Since $θ = θ_{n}^{* *}$ minimizes the approximation error plus a tight upper bound for the estimation error in the oracle model, within θ ∈ Θ_n, we refer to $θ_{n}^{* *}$ as an oracle.

The result can be used to emphasize that l₁ penalty behaves similarly as the l₀ penalty. Note that θ̂_n minimizes the empirical prediction error, E_n(R − Θθ)², plus an l₁ penalty whereas θ^**(u_n) minimizes the prediction error L(Φθ) plus an l₀ penalty. We provide an intuitive connection between these two quantities. First note that E_n(R − Φθ)² estimates L(Φθ) and σ̂_j estimates σ_j. We use “≈” to denote this relationship. Thus
$E_{n} {(R - Φ θ)}^{2} + λ_{n} \sum_{j = 1}^{J} {\hat{σ}}_{j} ∣ θ_{i} ∣ \approx L (Φ θ) + λ_{n} \sum_{j = 1}^{J} σ_{j} ∣ θ_{j} ∣ \leq L (Φ θ) + λ_{n} \sum_{j = 1}^{J} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣ + λ_{n} \sum_{j = 1}^{J} σ_{j} ∣ {\hat{θ}}_{n, j} ∣,$ (4.13)

where θ̂_n,j is the j_th component of θ̂_n. In Appendix B we show that for any θ ∈ Θ_n, $λ_{n} \sum_{j = 1}^{J} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣$ is upper bounded by ${| | θ | |}_{0} λ_{n}^{2} / β$ up to a constant with a high probability. Thus θ̂_n minimizes (4.13) and θ^**(u_n) roughly minimizes an upper bound of (4.13).
The constants involved in the theorem can be improved; we focused on readability as opposed to providing the best constants.

5. A Practical Implementation and an Evaluation

In this section we develop a practical implementation of the l₁-PLS method, compare this method to two commonly used alternatives and lastly illustrate the method using the motivating data from the Nefazodone-CBASP trial [13].

A realistic implementation of l₁-PLS method should use a data-dependent method to select the tuning parameter, λ_n. Since the primary goal is to maximize the Value, we select λ_n to maximize a cross validated Value estimator. For any ITR d, it is easy to verify that E[(R − V (d))1_A₌_d₍_X₎/p(A|X)] = 0. Thus an unbiased estimator of V (d) is

E_{n} [1_{A = d (X)} R / p (A ∣ X)] / E_{n} [1_{A = d (X)} / p (A ∣ X)]

[21] (recall that the randomization distribution p(a|X) is known). We split the data into 10 roughly equal-sized parts; then for each λ_n we apply the l₁-PLS based method on each 9 parts of the data to obtain an ITR, and estimate the Value of this ITR using the remaining part; the λ_n that maximizes the average of the 10 estimated Values is selected. Since the Value of an ITR is noncontinuous in the parameters, this usually results in a set of candidate λ_n’s achieving maximal Value. In the simulations below the resulting λ_n is nonunique in around 97% of the data sets. If necessary, as a second step we reduce the set of λ_n’s by including only λ_n’s leading to the ITR’s using the least number of variables. In the simulations below this second criterion effectively reduced the number of candidate λ_n’s in around 25% of the data sets, however multiple λ_n’s still remained in around 90% of the data sets. This is not surprising since the Value of an ITR only depends on the relative magnitudes of parameters in the ITR. In the third step we select the λ_n that minimizes the 10-fold cross validated prediction error estimator from the remaining candidate λ_n’s; that is, minimization of the empirical prediction error is used as a final tie breaker.

5.1. Simulations

A first alternative to l₁-PLS is to use ordinary least squares (OLS). The estimated ITR is d̂_OLS ∈ arg max_a Φ(X, a)θ̂_OLS where θ̂_OLS is the OLS estimator of θ. A second alternative is called “prognosis prediction” [14]. Usually this method employees multiple data sets, each of which involves one active treatment. Then the treatment that is associated with the best predicted prognosis [14] is selected. We implement this method by estimating E(R|X, A = a) via least squares with l₁ penalization for each treatment group (each a ∈ Inline graphic ) separately. The tuning parameter involved in each treatment group is selected by minimizing the 10-fold cross-validated prediction error estimator. The resulting ITR satisfies d̂_PP (X) ∈ arg max_a_∈ Ê(R|X, A = a) where the subscript “PP” denotes prognosis prediction.

For simplicity we consider binary A. All three methods use the same number of data points and the same number of basis functions but use these data points/basis functions differently. l₁-PLS and OLS use all J basis functions to conduct estimation with all n data points whereas the prognosis prediction method splits the data into the two treatment groups and uses J/2 basis functions to conduct estimation with the n/2 data points in each of the two treatment groups. To ensure the comparison is fair across the three methods, the approximation model for each treatment group is consistent with the approximation model used in both l₁-PLS and OLS (e.g. if Q₀ is approximated by (1, X, A, XA)θ in l₁-PLS and OLS, then in prognosis prediction we approximate E(R|X, A = a) by (1, X)θ_PP for each treatment group). We do not penalize the intercept coefficient in either prognosis prediction or l₁-PLS.

The three methods are compared using two criteria: 1) Value maximization; and 2) simplicity of the estimated ITRs (measured by the number of variables/basis functions used in the rule).

We illustrate the comparison of the three methods using 4 examples selected to reflect three scenarios; please see Section S.3 of the supplementary material for 4 further examples.

There is no treatment effect (i.e. Q₀ is constructed so that T₀ = 0; example 1). In this case, all ITRs yield the same Value. Thus the simplest rule is preferred.
There is a treatment effect and the treatment effect term T₀ is correctly modeled (example 4 for large n, and example 2). In this case, minimizing the prediction error will yield the ITR that maximizes the Value.
There is a treatment effect and the treatment effect term T₀ is misspecified (example 4 for small n, and example 3). In this case, there might be a mismatch between prediction error minimization and Value maximization.

The examples are generated as follows. The treatment A is generated uniformly from {−1, 1} independent of X and the response R. The response R is normally distributed with mean Q₀(X, A). In examples 1–3, X ~ U [−1, 1]⁵ and we consider three simple examples for Q₀. In example 4, X ~ U [0, 1] and we use a complex Q₀, where Q₀(X, 1) and Q(X, −1) are similar to the blocks function used in Donoho and Johnstone [8]. Further details of the simulation design are provided in Appendix A.3.

We consider two types of approximation models for Q₀. In examples 1–3, we approximate Q₀ by (1, X, A, XA)θ. In example 4, we approximate Q₀ by Haar wavelets. The number of basis functions may increase as n increases (we index J, Φ and θ^* by n in this case). Plots for Q₀(X, A) and the associated best wavelet fits $Φ_{n} (X, A) θ_{n}^{*}$ are provided in Figure 1.

Fig 1 — Plots for: the conditional mean function Q₀(X; A) (left), Q₀(X; A) and the associated best wavelet fit when J_n = 8 (middle), and Q₀(X; A) and the associated best wavelet fit when J_n = 128 (right) (example 4).

For each example, we simulate data sets of sizes n = 2^k for k = 5, …, 10. 1000 data sets are generated for each sample size. The Value of each estimated ITR is evaluated via Monte Carlo using a test set of size 10, 000. The Value of the optimal ITR is also evaluated using the test set.

Simulation results are presented in Figure 2. When the approximation model is of high quality, all methods produce ITRs with similar Value (see examples 1, 2 and example 4 for large n). However, when the approximation model is poor, the l₁-PLS method may produce highest Value (see example 3). Note that in example 3 settings in which the sample size is small, the Value of the ITR produced by l₁-PLS method has larger median absolute deviation (MAD) than the other two methods. One possible reason is that due to the mismatch between maximizing the Value and minimizing the prediction error, the Value estimator plays a strong role in selecting λ_n. The non-smoothness of the Value estimator combined with the mismatch results in very different λ_ns and thus the estimated decision rules vary greatly from data set to data set in this example. Nonetheless, the l₁-PLS method is still preferred after taking the variation into account; indeed l₁-PLS produces ITRs with higher Value than both OLS and PP in around 46%, 55% and 67% in data sets of sizes n = 32, 64 and 128, respectively. Furthermore, in general the l₁-PLS method uses much fewer variables for treatment assignment than the other two methods. This is expected because the OLS method does not have variable selection functionality and the PP method will use all variables that are predictive of the response R whereas the use of the Value in selecting the tuning parameter in l₁-PLS discounts variables that are only useful in predicting the response (and less useful in selecting the best treatment).

Fig 2 — Comparison of the l₁-PLS based method with the OLS method and the PP method (examples 1 – 4): Plots for medians and median absolute deviations (MAD) of the Value of the estimated decision rules (top panels) and the number of variables (terms) needed for treatment assignment (including the main treatment effect term, bottom panels) over 1000 samples versus sample size on the log scale. The black dash-dotted line in each plot on the first row denotes the Value of the optimal treatment rule, V (d₀), for each example. (n = 32; 64; 128; 256; 512; 1024. The corresponding numbers of basis functions in example 4 are J_n = 8; 16; 32; 64; 64; 128).

5.2. Nefazodone-CBASP trial example

The Nefazodone-CBASP trial was conducted to compare the efficacy of several alternate treatments for patients with chronic depression. The study randomized 681 patients with non-psychotic chronic major depressive disorder (MDD) to either Nefazodone, cognitive behavioral-analysis system of psychotherapy (CBASP) or the combination of the two treatments. Various assessments were taken throughout the study, among which the score on the 24-item Hamilton Rating Scale for Depression (HRSD) was the primary outcome. Low HRSD scores are desirable. See Keller et al. [13] for more detail of the study design and the primary analysis.

In the data analysis, we use a subset of the Nefazodone-CBASP data consisting of 656 patients for whom the response HRSD score was observed. In this trial, pairwise comparisons show that the combination treatment resulted in significantly lower HRSD scores than either of the single treatments. There was no overall difference between the single treatments.

We use l₁-PLS to develop an ITR. In the analysis the HRSD score is reverse coded so that higher is better. We consider 50 pretreatment variables X = (X₁,…, X₅₀). Treatments are coded using contrast coding of dummy variables A = (A₁, A₂), where A₁ = 2 if the combination treatment is assigned and −1 otherwise and A₂ = 1 if CBASP is assigned, −1 if nefazodone and 0 otherwise. The vector of basis functions, Φ(X, A), is of the form (1, X, A₁, XA₁, A₂, XA₂). So the number of basis functions is J = 153. As a contrast, we also consider the OLS method and the PP method (separate prognosis prediction for each treatment). The vector of basis functions used in PP is (1, X) for each of the three treatment groups. Neither the intercept term nor the main treatment effect terms in l₁-PLS or PP is penalized (see Section S.2 of the supplementary material for the modification of the weights σ̂_j used in (4.1)).

The ITR given by the l₁-PLS method recommends the combination treatment to all (so none of the pretreatment variables enter the rule). On the other hand, the PP method produces an ITR that uses 29 variables. If the rule produced by PP were used to assign treatment for the 656 patients in the trial, it would recommend the combination treatment for 614 patients and nefazodone for the other 42 patients. In addition, the OLS method will use all the 50 variables. If the ITR produced by OLS were used to assign treatment for the 656 patients in the trial, it would recommend the combination treatment for 429 patients, nefazodone for the 145 patients and CBASP for the other 82 patients.

6. Discussion

Our goal is to construct a high quality ITR that will benefit future patients. We considered an l₁-PLS based method and provided a finite sample upper bound for V (d₀) − V (d̂_n), the excess Value of the estimated ITR.

The use of an l₁ penalty allows us to consider a large model for the conditional mean function Q₀ yet permits a sparse estimated ITR. In fact, many other penalization methods such as SCAD [9] and l₁ penalty with adaptive weights (adaptive Lasso; [37]) also have this property. We choose the non-adaptive l₁ penalty to represent these methods. Interested readers may justify other PLS methods using similar proof techniques.

The high probability finite sample upper bounds (i.e. (4.7) and (4.9)) cannot be used to construct a prediction/confidence interval for V (d₀) − V (d̂_n) due to the unknown quantities in the bound. How to develop a tight computable upper bound to assess the quality of d̂_n is an open question.

We used cross validation with Value maximization to select the tuning parameter involved in the l₁-PLS method. As compared to the OLS method and the PP method, this method may yield higher Value when T₀ is misspecified. However, since only the Value is used to select the tuning parameter, this method may produce a complex ITR for which the Value is only slightly higher than that of a much simpler ITR. In this case, a simpler rule may be preferred due to the interpretability and cost of collecting the variables. Investigation of a tuning parameter selection criterion that trades off the Value with the number of variables in an ITR is needed.

This paper studied a one stage decision problem. However, it is evident that some diseases require time-varying treatment. For example, individuals with a chronic disease often experience a waxing and waning course of illness. In these settings the goal is to construct a sequence of ITRs that tailor the type and dosage of treatment through time according to an individual’s changing status. There is an abundance of statistical literature in this area [29, 30, 22, 23, 26, 17, 34, 35]. Extension of the least squares based method to the multi-stage decision problem has been presented in Murphy [23]. The performance of l₁ penalization in this setting is unclear and worth investigation.

Supplementary Material

NIHMS266525-supplement-2.pdf^{(136.5KB, pdf)}

NIHMS266525-supplement-supplement_1.pdf^{(99.3KB, pdf)}

Acknowledgments

The authors thank Martin Keller and the investigators of the Nefazodone-CBASP trial for use of their data. The authors also thank John Rush, MD, for the technical support and Bristol-Myers Squibb for helping fund the trial. The authors thank valuable comments from Eric B. Laber and Peng Zhang.

APPENDIX

A.1. Proof of Theorem 3.1

For any ITR d: Inline graphic → , denote ΔT_d(X) ≜ max_a_∈T₀(X, a) − T₀(X, d(X)). Using similar arguments to that in Section 2, we have V (d₀) − V (d) = E(ΔT_d). If V (d₀) − V (d) = 0, then (3.4) and (3.5) automatically hold. Otherwise, E(ΔT_d)² ≥ (EΔT_d)² > 0. In this case, for any ε > 0, define the event

Ω_{ε} = {max_{a \in A} T_{0} (X, a) - max_{a \in A \ arg {max}_{a \in A} T_{0} (X, a)} T_{0} (X, a) \leq ε} .

Then ΔT_d ≤ (ΔT_d)²/ε on the event $Ω_{ε}^{C}$ . This together with the fact that ΔT_d ≤ (ΔT_d)²/ε + ε/4 implies

\begin{array}{l} V (d_{0}) - V (d) = E (1_{Ω_{ε}^{C}} Δ T_{d}) + E (1_{Ω_{ε}} Δ T_{d}) \\ \leq \frac{1}{ε} E [1_{Ω_{ε}^{C}} {(Δ T_{d})}^{2}] + E [1_{Ω_{ε}} (\frac{{(Δ T_{d})}^{2}}{ε} + \frac{ε}{4})] \\ = \frac{1}{ε} E [{(Δ T_{d})}^{2}] + \frac{ε}{4} P (Ω_{ε}) \leq \frac{1}{ε} E [{(Δ T_{d})}^{2}] + \frac{C}{4} ε^{1 + α}, \end{array}

where the last inequality follows from the margin condition (3.3). Choosing ε = (4E(ΔT_d)²/C)^1/(2+^α⁾ to minimize the above upper bound yields

V (d_{0}) - V (d) \leq 2^{α / (2 + α)} C^{1 / (2 + α)} {[E {(Δ T_{d})}^{2}]}^{(1 + α) / (2 + α)} .

(A.1)

Next, for any d and Q such that d(X) ∈ max_a_∈ Q(X, a) and decomposition Q(X, A) into W(X) + T(X, A),

\begin{array}{l} E {(Δ T_{d})}^{2} = E [{(max_{a \in A} T_{0} (X, a) - max_{a \in A} T (X, a) + T (X, d (X)) - T_{0} (X, d (X)))}^{2}] \\ \leq 2 E [{(max_{a \in A} T_{0} (X, a) - max_{a \in A} T (X, a))}^{2} + {(T (X, d (X)) - T_{0} (X, d (X)))}^{2}] \\ \leq 4 E [max_{a \in A} {(T (X, a) - T_{0} (X, a))}^{2}], \end{array}

E {(Δ T_{d})}^{2} \leq 4 S E [\sum_{a \in A} {(T (X, a) - T_{0} (X, a))}^{2} p (a ∣ X)] = 4 S E {(T (X, A) - T_{0} (X, A))}^{2} .

(A.2)

Inequality (3.5) follows by substituting (A.2) into (A.1) and setting W(X, A) = E[Q(X, A)|X]. Inequality (3.4) follows by setting W(X) = 0 and noticing that ΔT_d(X) = max_a_∈ Q₀(X, a) − Q₀(X, d(X)).

A.2. Generalization of Theorem 4.3

In this section, we present a generalization of Theorem 4.3 where J may depend on n and the sparsity of any θ ∈ ℝ^J is measured by the number of “large” components in θ as described in Zhang and Huang [36]. In this case, J, F and the prediction error minimizer θ^* are denoted as J_n, Φ_n and $θ_{n}^{*}$ , respectively. All relevant quantities and assumptions are re-stated below.

Let |M| denote the cardinality of any index set M ⊆ {1,…, J_n}. For any θ ∈ ℝ^J_n and constant ρ ≥ 0, define

M_{ρ λ_{n}} (θ) \in arg min_{{M \subseteq {1, \dots, J_{n}} : \sum_{j \in {1, \dots, J_{n}} \ M} σ_{j} ∣ θ_{j} ∣ \leq ρ ∣ M ∣ λ_{n}}} ∣ M ∣ .

Then M_{ρλ_n} (θ) is the smallest index set that contains only “large” components in θ. |M_{ρλ_n} (θ)| measures the sparsity of θ. It is easy to see that when ρ = 0, M₀(θ) is the index set of nonzero components in θ and |M₀(θ)| = ||θ||₀. Moreover, M_{ρλ_n} (θ) is an empty set if and only if θ = 0.

Let [ $θ_{n}^{*}$ ] be the set of most sparse prediction error minimizers in the linear model, i.e.

[θ_{n}^{*}] = arg min_{θ \in arg {min}_{θ} L (Φ_{n} θ)} ∣ M_{ρ λ_{n}} (θ) ∣ .

(A.3)

Note that [ $θ_{n}^{*}$ ] depends on ρλ_n.

To derive the finite sample upper bound for L(Φ_nθ̂_n), we need the following assumptions.

Assumption A.1

The error terms ε_i, i = 1,…, n are independent of (X_i, A_i), i = 1,…, n and are i.i.d. with E(ε_i) = 0 and $E [{∣ ε_{i} ∣}^{l}] \leq \frac{l!}{2} c^{l - 2} σ^{2}$ for some c, σ² > 0 for all l ≥ 2.

Assumption A.2

For all n ≥ 1,

there exists an 1 ≤ U_n < ∞ such that max_{j=1,…,J_n} ||φ_j||_∞/σ_j ≤ U_n, where $σ_{j} ≜ {(E φ_{j}^{2})}^{1 / 2}$ .
there exists an 0 < η_1,n < ∞, such that ${sup}_{θ \in [θ_{n}^{*}]} {| | Q_{0} - Φ_{n} θ | |}_{\infty} \leq η_{1, n}$ .

For any 0 ≤ γ < 1/2, η₂_,n ≥ 0 (which may depend on n) and tuning parameter λ_n, define

Θ_{n}^{o} = {θ \in R^{J_{n}} : \exists θ^{o} \in [θ_{n}^{*}] s . t . {| | Φ_{n} (θ - θ^{o}) | |}_{\infty} \leq η_{2, n} and max_{j = 1, \dots, J_{n}} | E [Φ_{n} (θ - θ^{o}) \frac{φ_{j}}{σ_{j}}] | \leq γ λ_{n}} .

Assumption A.3

For any n ≥ 1, there exists a β_n > 0 such that

E {[Φ_{n} (\tilde{θ} - θ)]}^{2} ∣ M_{ρ λ_{n}} (θ) ∣ \geq β_{n} [{(\sum_{j \in M_{ρ λ_{n}} (θ)} σ_{j} ∣ {\tilde{θ}}_{j} - θ_{j} ∣)}^{2} - ρ^{2} {∣ M_{ρ λ_{n}} (θ) ∣}^{2} λ_{n}^{2}]

for all $θ \in Θ_{n}^{o} \ {0}$ , θ̃ ∈ ℝ^J_n and $\sum_{j \in {1, \dots, J_{n}} \ M_{ρ λ_{n}} (θ)} σ_{j} ∣ {\tilde{θ}}_{j} ∣ \leq \frac{2 γ + 5}{1 - 2 γ} (\sum_{j \in M_{ρ λ_{n}} (θ)} ∣ {\tilde{θ}}_{j} - θ_{j} ∣ + ρ ∣ M_{ρ λ_{n}} (θ) ∣ λ_{n})$ .

When $E (Φ_{n}^{(2)} {(X, A)}^{T} ∣ X) = 0$ a.s. ( $Φ_{n}^{(2)}$ is defined in Section 4.1), we need an extra assumption to derive the finite sample upper bound for the mean square error of the treatment effect estimator, $E [{Φ_{n}^{(2)} {\hat{θ}}_{n}^{(2)} - T_{0} (X, A)]}^{2}$ (recall that T₀(X,A) ≜ Q₀(X, A) − E[Q₀(X,A)|X]).

Assumption A.4

For any n ≥ 1, there exists a β_n > 0 such that

E {[Φ_{n}^{(2)} ({\tilde{θ}}^{(2)} - θ^{(2)})]}^{2} ∣ M_{ρ λ_{n}}^{(2)} (θ) ∣ \geq β_{n} [{(\sum_{j \in M_{ρ λ_{n}}^{(2)} (θ)} σ_{j} ∣ {\tilde{θ}}_{j} - θ_{j} ∣)}^{2} - ρ^{2} {∣ M_{ρ λ_{n}}^{(2)} (θ) ∣}^{2} λ_{n}^{2}]

M_{ρ λ_{n}}^{(2)} (θ) \in arg min_{{M \subseteq {J_{n}^{(1)} + 1, \dots, J_{n}} : \sum_{j \in (J_{n}^{(1)} + 1, \dots, J_{n}) \ M} σ_{j} ∣ θ_{j} ∣ \leq ρ ∣ M ∣ λ_{n}}} ∣ M ∣

is the smallest index set that contains only large components in θ⁽²⁾.

Note that here for simplicity, we assume that Assumptions A.3 and A.4 hold with the same value of β_n. And with out loss of generality, we can always choose a small enough β_n so that ρβ_n ≤ 1 for a given ρ.

For any t > 0, define

Θ_{n} = {θ \in Θ_{n}^{o} : ∣ M_{ρ λ_{n}} (θ) ∣ \leq \frac{{(1 - 2 γ)}^{2} β_{n}}{120} [\sqrt{\frac{1}{9} + \frac{n}{2 U_{n}^{2} [log (3 J_{n} (J_{n} + 1)) + t]}} - \frac{1}{3}]} .

(A.4)

Note that we allow U_n, η₁_,n, η₂_,n and $β_{n}^{- 1}$ to increase as n increases. However, if those quantities are small, the upper bound in (A.7) will be tighter.

Theorem A.1

Suppose Assumptions A.1 and A.2 hold. For any given 0 ≤ γ < 1/2, η_2,n > 0, ρ ≥ 0 and t > 0, let θ̂_n be the l₁-PLS estimator defined in (4.1) with tuning parameter

λ_{n} \geq \frac{8 max {3 c, 2 (η_{1, n} + η_{2, n})} U_{n} (log 6 J_{n} + t)}{(1 - 2 γ) n} + \frac{12 max {σ, (η_{1, n} + η_{2, n})}}{(1 - 2 γ)} \sqrt{\frac{2 (log 6 J_{n} + t)}{n}} .

(A.5)

Suppose Assumption A.3 holds with ρβ_n ≤ 1. Let Θ_n be the set defined in (A.4) and assume Θ_n is non-empty. If

\frac{log 2 J_{n}}{n} \leq \frac{2 {(1 - 2 γ)}^{2}}{27 U_{n}^{2} - 10 γ - 22},

(A.6)

then with probability at least $1 - exp (- k_{n}^{'} n) - exp (- t)$ , we have

L (Φ_{n} {\hat{θ}}_{n}) \leq min_{θ \in Θ_{n}} [L (Φ_{n} θ) + K_{n} \frac{∣ M_{ρ λ_{n}} (θ) ∣}{β_{n}} λ_{n}^{2}],

(A.7)

where $k_{n}^{'} = 13 {(1 - 2 γ)}^{2} [6 (27 U_{n}^{2} - 10 γ - 22)]$ and K_n = [40γ(12β_nρ + 2γ + 5)]/[(1 − 2γ)(2γ + 19)] + 130(12β_nρ + 2γ + 5)²/[9(2γ + 19)²].

Furthermore, suppose $E (Φ_{n}^{(2)} {(X, A)}^{T} ∣ X) = 0$ a.s. If Assumption A.4 holds with ρβ_n ≤ 1, then with probability at least $1 - exp (- k_{n}^{'} n) - exp (- t)$ , we have

E {(Φ_{n}^{(2)} {\hat{θ}}_{n}^{(2)} - T_{0})}^{2} \leq min_{θ \in Θ_{n}} [E {(Φ_{n}^{(2)} θ^{(2)} - T_{0})}^{2} + K_{n}^{'} \frac{∣ M_{ρ λ_{n}}^{2} (θ) ∣}{β_{n}} λ_{n}^{2}] .

where $K_{n}^{'} = 20 (12 β_{n} ρ + 2 γ + 5) {γ / [(1 - 2 γ) (7 - 6 β_{n} ρ)] + [3 (1 - 2 γ) β_{n} ρ + 10 (2 γ + 5)] / [9 {(2 γ + 19)}^{2}]}$ .

Remark

Note that K_n is upper bounded by a constant under the assumption β_nρ ≤ 1. In the asymptotic setting when n → ∞ and J_n → ∞, (A.7) implies that with probability tending to 1, $L (Φ_{n} {\hat{θ}}_{n}) - L (Φ_{n} θ_{n}^{*}) \to 0$ if (i) $∣ M_{ρ λ_{n}} (θ_{n}^{*}) ∣ λ_{n}^{2} / β_{n} = o (1)$ , (ii) $U_{n}^{2} log J_{n} / n \leq k_{1}$ and $∣ M_{ρ λ_{n}} (θ_{n}^{*}) ∣ \leq k_{2} β_{n} \sqrt{n / (U_{n}^{2} log J_{n})}$ for some sufficiently small positive constants k₁ and k₂, and (iii) $λ_{n} \geq k_{3} max {1, η_{1, n} + η_{2, n}} \sqrt{log J_{n} / n}$ for a sufficiently large constant k₃, where $θ_{n}^{*} \in [θ_{n}^{*}]$ (take t = log J_n).
Below we briefly discuss Assumptions A.2 – A.4.

Assumption A.2 is very similar to Assumption 2 in Theorem 4.1 (which is used to prove the concentration of the sample mean around the true mean), except that U_n and η₁_,n may increase as n increases. This relaxation allows the use of basis functions for which the sup norm max_j ||φ_j||_∞ is increasing in n (e.g. the wavelet basis used in example 4 of the simulation studies).

Assumption A.3 is a generalization of condition (4.8) (which has been discussed in remark 4 following Theorem Theorem 4.1)) to the case where J_n may increase in n and the sparsity of a parameter is measured by the number of “large” components as described at the beginning of this section. This condition is used to avoid the collinearity problem. It is easy to see that when ρ = 0 and β_n is fixed in n, this assumption simplifies to condition (4.8).

Assumption A.4 puts a strengthened constraint on the linear model of the treatment effect part, as compared to Assumption A.3. This assumption, together with Assumption A.3, is needed in deriving the upper bound for the mean square error of the treatment effect estimator. It is easy to verify that if $E [Φ_{n}^{T} Φ_{n}]$ is positive definite, then both A.3 and A.4 hold. Although the result is about the treatment effect part, which is asymptotically independent of the main effect of X (when $E [Φ_{n}^{(2)} (X, A) ∣ X] = 0$ a.s.), we still need Assumption A.3 to show that the cross product term $E_{n} [(Φ_{n}^{(1)} {\hat{θ}}_{n}^{(1)} - Φ_{n}^{(1)} θ^{(1)}) (Φ_{n}^{(2)} {\hat{θ}}_{n}^{(2)} - Φ_{n}^{(2)} θ^{(2)})]$ is upper bounded by a quantity converging to 0 at the desired rate. We may use a really poor model for the main effect part E(Q₀(X, A)|X) (e.g. $Φ_{n}^{(1)} \equiv 1$ ), and Assumption A.4 implies Assumption A.3 when ρ = 0. This poor model only effects the constants involved in the result. When the sample size is large (so that λ_n is small), the estimated ITR will be of high quality as long as T₀ is well approximated.

Proof

For any θ ∈ Θ_n, define the events

\begin{array}{l} Ω_{1} = \cap_{j = 1}^{J_{n}} {\frac{2 (1 + γ)}{3} σ_{j} \leq {\hat{σ}}_{j} \leq \frac{2 (2 - γ)}{3} σ_{j}} (where {\hat{σ}}_{j} ≜ {(E_{n} φ_{j}^{2})}^{1 / 2}), \\ Ω_{2} (θ) = {max_{j, k = 1, \dots, J_{n}} | (E - E_{n}) (\frac{φ_{j} φ_{k}}{σ_{j} σ_{k}}) | \leq \frac{{(1 - 2 γ)}^{2} β_{n}}{120 ∣ M_{ρ λ_{n}} (θ) ∣}}, \\ Ω_{3} (θ) = {max_{j = 1, \dots, J_{n}} | E_{n} [(R - Φ_{n} θ) \frac{φ_{j}}{σ_{j}}] | \leq \frac{4 γ + 1}{6} λ_{n}} . \end{array}

Then there exists a $θ^{o} \in [θ_{n}^{*}]$ such that

\begin{array}{l} L (Φ_{n} {\hat{θ}}_{n}) = L (Φ_{n} θ) + 2 E [(Φ_{n} θ^{o} - Φ_{n} θ) Φ_{n} (θ - {\hat{θ}}_{n})] + E {[Φ_{n} ({\hat{θ}}_{n} - θ)]}^{2} \\ \leq L (Φ_{n} θ) + 2 max_{j = 1, \dots, J_{n}} | E [Φ_{n} (θ^{o} - θ) \frac{φ_{j}}{σ_{j}}] | (\sum_{j = 1}^{J_{n}} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣) + E {[Φ_{n} ({\hat{θ}}_{n} - θ)]}^{2} \\ \leq L (Φ_{n} θ) + 2 γ λ_{n} (\sum_{j = 1}^{J_{n}} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣) + E {[Φ_{n} ({\hat{θ}}_{n} - θ)]}^{2}, \end{array}

where the first equality follows from the fact that E[(R − Φ_nθ^o)φ_j] = 0 for any $θ^{o} \in [θ_{n}^{*}]$ for j = 1,…, J_n and the last inequality follows from the definition of $Θ_{n}^{o}$ .

Based on Lemma A.1 below, we have that on the event Ω₁ ∩ Ω₂(θ) ∩Ω₃(θ),

L (Φ_{n} {\hat{θ}}_{n}) \leq L (Φ_{n} θ) + K_{n} \frac{∣ M_{ρ λ_{n}} (θ) ∣}{β_{n}} λ_{n}^{2} .

Similarly, when $E [Φ_{2}^{(2)} {(X, A)}^{T} ∣ X] = 0$ , by Lemma A.2, we have that on the event Ω₁ ∩ Ω₂(θ) ∩ Ω₃(θ),

\begin{array}{l} E {(Φ_{n}^{(2)} {\hat{θ}}_{n}^{(2)} - T_{0})}^{2} \leq E {(Φ_{n}^{(2)} θ^{(2)} - T_{0})}^{2} + 2 γ λ_{n} (\sum_{j = J_{n}^{(1)} + 1}^{J_{n}} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣) + E {[Φ_{n}^{(2)} ({\hat{θ}}_{n}^{(2)} - θ^{(2)})]}^{2} \\ \leq E {(Φ_{n}^{(2)} θ^{(2)} - T_{0})}^{2} + K_{n}^{'} \frac{∣ M_{ρ λ_{n}}^{(2)} (θ) ∣}{β_{n}} λ_{n}^{2} . \end{array}

The conclusion of the theorem follows from the union probability bounds of the events Ω₁, Ω₂(θ) and Ω₃(θ) provided in Lemmas A.3, A.4 and A.5.

Below we state the lemmas used in the proof of Theorem A.1. The proofs of the lemmas are given in Section S.3 of the supplementary material.

Lemma A.1

Suppose Assumption A.3 holds with ρβ_n ≤ 1. Then for any θ ∈ Θ_n, on the event Ω₁ ∩ Ω₂(θ) ∩ Ω₃(θ), we have

\sum_{j = 1}^{J_{n}} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣ \leq \frac{20 (12 ρ β_{n} + 2 γ + 5)}{(1 - 2 γ) (19 + 2 γ) β_{n}} ∣ M_{ρ λ_{n}} (θ) ∣ λ_{n}

(A.8)

and

E {[Φ_{n} ({\hat{θ}}_{n} - θ)]}^{2} \leq \frac{130 {(12 ρ β_{n} + 2 γ + 5)}^{2}}{9 {(19 + 2 γ)}^{2} β_{n}} ∣ M_{ρ λ_{n}} (θ) ∣ λ_{n}^{2}

(A.9)

Remark

This lemma implies that θ̂_n is close to each θ ∈ Θ_n on the event Ω₁ ∩ Ω₂(θ) ∩ Ω₃(θ). The intuition is as follows. Since θ̂_n minimizes (4.1), the first order conditions imply that max_j |E_n(R −Φ_nθ̂_n)φ_j/σ̂_j| ≤ λ_n/2. Similar property holds for θ on the event Ω₁ ∩ Ω₃(θ). Assumption A.3 together with event Ω₂(θ) ensures that there is no collinearity in the n × J_n design matrix ${(Φ_{n} (X_{i}, A_{i}))}_{i = 1}^{n}$ . These two aspects guarantee the closeness of θ̂_n to θ.

Lemma A.2

Suppose $E [Φ_{n}^{(2)} {(X, A)}^{T} ∣ X] = 0$ a.s. and Assumption A.4 holds with ρβ_n ≤ 1. Then for any θ ∈ Θ_n, on the event Ω₁ ∩ Ω₂(θ) ∩ Ω₃(θ), we have

\sum_{j = J_{n}^{(1)} + 1}^{J_{n}} σ_{j} ∣ {\hat{θ}}_{n, j} - θ_{j} ∣ \leq \frac{10 (12 β_{n} ρ + 2 γ + 5)}{(1 - 2 γ) (7 - 6 β_{n} ρ) β_{n}} ∣ M_{ρ λ_{n}}^{(2)} (θ) ∣ λ_{n}

(A.10)

and

E {[Φ_{n}^{(2)} ({\hat{θ}}_{n}^{(2)} - θ^{(2)})]}^{2} \leq \frac{20 (12 ρ β_{n} + 2 γ + 5) [3 (1 - 2 γ) β_{n} ρ + 10 (2 γ + 5)]}{9 {(2 γ + 19)}^{2} β_{n}} ∣ M_{ρ λ_{n}}^{(2)} (θ) ∣ λ_{n}^{2}

(A.11)

Lemma A.3

Suppose Assumption A.2(a) and inequality (A.6) hold. Then $P (Ω_{1}^{C}) \leq exp (- k_{n}^{'} n)$ , where $k_{n}^{'} = 13 {(1 - 2 γ)}^{2} / [6 (27 U_{n}^{2} - 10 γ - 22)]$ .

Lemma A.4

Suppose Assumption A.2(a) holds. Then for any t > 0 and θ ∈ Θ_n, P({Ω₂(θ)}^C) ≤ 2 exp(−t)/3.

Lemma A.5

Suppose Assumptions A.1 and A.2 hold. For any t > 0, if λ_n satisfies condition (A.5), then for any θ ∈ Θ_n, we have P({Ω₃(θ)}^C) ≤ 2 exp(−t)/3.

A.3. Design of simulations in Section 5.1

In this section, we present the detailed simulation design of the examples used in Section 5.1. These examples satisfy all assumptions listed in the Theorems (it is easy to verify that for examples 1–3. Validity of the assumptions for example 4 is addressed in the remark after example 4). In addition, Θ_n defined in (4.4) is non-empty as long as n is sufficiently large (note that the constants involved in Θ_n can be improved and are not that meaningful. We focused on a presentable result instead of finding the best constants).

In examples 1 – 3, X = (X₁,…, X₅) is uniformly distributed on [−1, 1]⁵. The treatment A is then generated independently of X uniformly from {−1, 1}. Given X and A, the response R is generated from a normal distribution with mean Q₀(X, A) = 1+2X₁ +X₂ +0.5X₃ +T₀(X, A) and variance 1. We consider the following three examples for T₀.

T₀(X, A) = 0 (i.e. there is no treatment effect).
T₀(X, A) = 0.424(1 − X₁ − X₂)A.
T₀(X, A) = 0.446sign(X₁)(1 − X₁)²A.

Note that in each example T₀(X, A) is equal to the treatment effect term, Q₀(X, A) − E[Q₀(X, A)|X]. We approximate Q₀ by Inline graphic = {(1, X, A, XA)θ: θ ∈ ℝ¹²}. Thus in examples 1 and 2 the treatment effect term T₀ is correctly modeled, while in example 3 the treatment effect term T₀ is misspecified.

The parameters in examples 2 and 3 are chosen to reflect a medium effect size according to Cohen’s d index. When there are two treatments, the Cohen’s d effect size index is defined as the standardized difference in mean responses between two treatment groups, i.e.

e s = \frac{E (R ∣ A = 1) - E (R ∣ A = - 1)}{{([Var (R ∣ A = 1) + Var (R ∣ A = - 1)] / 2)}^{1 / 2}} .

Cohen [7] tentatively defined the effect size as “small” if the Cohen’s d index is 0.2, “medium” if the index is 0.5 and “large” if the index is 0.8.

In example 4, X is uniformly distributed on [0, 1]. Treatment A is generated independently of X uniformly from {−1, 1}. The response R is generated from a normal distribution with mean Q₀(X, A) and variance 1, where $Q_{0} (X, 1) = \sum_{j = 1}^{8} ϑ_{(1), j} 1_{X < u_{(1), j}}, Q_{0} (X, - 1) = \sum_{j = 1}^{8} ϑ_{(- 1), j} 1_{X < u_{(- 1), j}}$ , and ϑ’s and u’s are parameters specified in (A.12). The effect size is small.

\begin{array}{l} (ϑ_{(1), 1}, \dots, ϑ_{(1), 8}) = (- 0.781, 0.730, 0.635, 0.512, - 2.278, 1.347, 1.155, - 0.030); \\ (ϑ_{(- 1), 1}, \dots, ϑ_{(- 1), 8}) = (- 2.068, 1.520, - 0.072, - 0.637, 1.003, - 0.611, - 0.305, 1.016); \\ (u_{(1), 1}, \dots, u_{(1), 8}) = (0.028, 0.144, 0.171, 0.298, 0.421, 0.443, 0.463, 0.758); \\ (u_{(- 1), 1}, \dots, u_{(- 1), 8}) = (0.061, 0.215, 0.492, 0.544, 0.6302, 0.650, 0.785, 0.909) . \end{array}

(A.12)

We approximate Q₀ by Haar wavelets

Q = {θ_{(0), 0} h_{0} (X) + \sum_{l k} θ_{(0), l k} h_{l k} (X) + (θ_{(0), 1} h_{0} (X) + \sum_{l k} θ_{(1), l k} h_{l k} (X)) A : θ ., . \in R},

where h₀(x) = 1_x_∈[0,1] and h_lk(x) = 2^l^/2 (1_{2^lx∈[k+1/2,k+1)} − 1_{2^lx∈[k,k+1/2)}) for l = 0,…, l̄_n. We choose l̄_n = ⌊3log₂ n/4⌋ − 2. For a given l and sample ${(X_{i}, A_{i}, R_{i})}_{i = 1}^{n}$ , k takes integer values from ⌊2^l min_i X_i⌋ to ⌈2^l max_i X_i⌉ − 1. Then J_n = 2^{⌊3 log₂n/4⌋} ≤ n^3/4.

Remark

In example 4, we allow the number of basis functions J_n to increase with n. The corresponding theoretical result can be obtained by combining Theorem 3.1 and Theorem A.1. Below we demonstrate the validation of the assumptions used in the theorems.

Theorem 3.1 requires that the randomization probability p(a|x) ≥ S⁻¹ for a positive constant for all (x, a) pairs and the margin condition (3.3) or (3.6) holds. According the generative model, we have that p(a|x) = 1/2 and condition (3.6) holds.

Theorem A.1 requires Assumptions A.1 - Assumptions A.4 hold and Θ_n defined in (A.4) is non-empty. Since we consider normal error terms, Assumption A.1 holds. Note that the basis functions used in Haar wavelet are orthogonal. It is also easy to verify that Assumptions A.3 and A.4 hold with β_n = 1 and Assumption A.2 holds with U_n = n^3/8/2 and $η_{1, n} \leq constant + constant \times {| | θ_{n}^{*} | |}_{0}$ (since each $∣ φ_{j} θ_{n, j}^{*} ∣ = ∣ φ_{j} E (φ_{j} R) ∣ \leq constant \times ∣ φ_{j} ∣ E ∣ φ_{j} ∣ \leq O (1)$ ). Since Q₀ is piece-wise constant, we can also verify that ${| | θ_{n}^{*} | |}_{0} \leq O (log n)$ . Thus for sufficiently large n, Θ_n is non-empty and (A.6) holds. The RHS of (A.5) converges to zero as n → ∞.

Contributor Information

Min Qian, Email: minqian@umich.edu.

Susan A. Murphy, Email: samurphy@umich.edu.

References

1.Bartlett PL, Jordan ML, McAuliffe JD. Convexity, classification, and risk bounds. Journal of the American Statistical Association. 2006;135(3):311–334. [Google Scholar]
2.Bartlett PL. Fast rates for estimation error and oracle inequalities for model selection. Econometric Theory. 2008;24(2):545–552. [Google Scholar]
3.Bickel PJ, Ritov Y, Tsybakov AB. Simultaneous analysis of lasso and dantzig selector. The Annals of Statistics. 2009;37(4):1705–1732. [Google Scholar]
4.Bunea F, Tsybakov AB, Wegkamp MH. Sparsity oracle inequalities for the Lasso. Electronic Journal of Statistics. 2007;1:169–194. [Google Scholar]
5.Cai T, Tian L, Lloyd-Jones DM, Wei LJ. Evaluating Subject-level Incremental Values of New Markers for Risk Classification Rule. Harvard University Biostatistics Working Paper Series. Working Paper 91. 2008a doi: 10.1007/s10985-013-9272-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
6.Cai T, Tian L, Uno H, Solomon SD, Wei LJ. Calibrating Parametric Subject-specific Risk Estimation. Harvard University Biostatistics Working Paper Series. Working Paper 92 2008b [Google Scholar]
7.Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2. Hillsdale, NJ: Lawrence Erlbaum Associates, Inc; 1988. [Google Scholar]
8.Donoho D, Johnstone I. Ideal spatial adaptation by wavelet shrinkage. Biometrika. 1994;81(3):425–455. [Google Scholar]
9.Fan J, Li R. Variable selection via nonconcave penalized likelihood and it oracle properties. Journal of the American Statistical Association. 2001;96:1348–1360. [Google Scholar]
10.Feldstein ML, Savlov ED, Hilf R. A statistical model for predicting response of breast cancer patients to cytotoxic chemotherapy. Cancer Research. 1978;38(8):2544–2548. [PubMed] [Google Scholar]
11.Insel TR. Translating scientific opportunity into public health impact: a strategic plan for research on mental illness. Archives of General Psychiatry. 2009;66(2):128–33. doi: 10.1001/archgenpsychiatry.2008.540. [DOI] [PubMed] [Google Scholar]
12.Ishigooka J, Murasaki M, Miura S The Olanzapine Late-Phase II Study Group. Olanzapine optimal dose: Results of an open-label multicenter study in schizophrenic patients. Psychiatry and Clinical Neurosciences. 2001;54(4):467–478. doi: 10.1046/j.1440-1819.2000.00738.x. [DOI] [PubMed] [Google Scholar]
13.Keller MB, McCullough JP, Klein DN, Arnow B, Dunner DL, Gelenberg AJ, Markowitz JC, Nemeroff CB, Russell JM, Thase ME, Trivedi MH, Zajecka J. A comparison of nefazodone, the cognitive behavioral-analysis system of psychotherapy, and their combination for the treatment of chronic depression. The New England Journal of Medicine. 2000;342(20):1462–1470. doi: 10.1056/NEJM200005183422001. [DOI] [PubMed] [Google Scholar]
14.Kent DM, Hayward RA, Griffith JL, Vijan S, Beshansky JR, Califf RM, Selker HP. An independently derived and validated predictive model for selecting patients with myocardial infarction who are likely to benefit from tissue plasminogen activator compared with streptokinase. The American Journal of Medicine. 2002;113(2):104–11. doi: 10.1016/s0002-9343(02)01160-9. [DOI] [PubMed] [Google Scholar]
15.Koltchinskii V. Sparsity in penalized empirical risk minimization. Annales de l’Institut Henri Poincaré Probabilitiés et Statistiques. 2009;45(1):7–57. [Google Scholar]
16.Lesko LJ. Personalized Medicine: Elusive Dream or Imminent Reality? Clinical Pharmacology and Therapeutics. 2007;81:807–816. doi: 10.1038/sj.clpt.6100204. [DOI] [PubMed] [Google Scholar]
17.Lunceford JK, Davidian M, Tsiatis AA. Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials. Biometrics. 2002;58:48–57. doi: 10.1111/j.0006-341x.2002.00048.x. [DOI] [PubMed] [Google Scholar]
18.Mammen E, Tsybakov A. Smooth discrimination analysis. The Annals of Statistics. 1999;27:1808–1829. [Google Scholar]
19.Massart P. Ecole d’Eté de Probabilités de Saint-Flour XXXIII, Concentration inequalities and model selection. Springer; 2003. [Google Scholar]
20.Massart P. A non asymptotic theory for model selection. Proceedings of the 4th European Congress of Mathematicians (Ed. Ari Laptev); European Mathematical Society; 2005. pp. 309–323. [Google Scholar]
21.Murphy SA, van der Laan MJ, Robins JM CPPRG. Marginal mean models for dynamic regimes. Journal of the American Statistical Association. 2001;96:1410–1423. doi: 10.1198/016214501753382327. [DOI] [PMC free article] [PubMed] [Google Scholar]
22.Murphy SA. Optimal Dynamic Treatment Regimes. Journal of the Royal Statistical Society, Series B (with discussion) 2003;65(2):331–366. [Google Scholar]
23.Murphy SA. A Generalization error for Q-Learning. Journal of Machine Learning Research. 2005;6:1073–1097. [PMC free article] [PubMed] [Google Scholar]
24.Piquette-Miller P, Grant DM. The Art and Science of Personalized Medicine. Clinical Pharmacology and Therapeutics. 2007;81:311–315. doi: 10.1038/sj.clpt.6100130. [DOI] [PubMed] [Google Scholar]
25.Polonik W. Measuring mass concentrations and estimating density contour clusters - an excess mass approach. The Annals of Statistics. 1995;23(3):855–881. [Google Scholar]
26.Robins JM. Optimal-regime structural nested models. In: Lin DY, Haegerty P, editors. Lecture notes in Stastitics; Proceedings of the Second Seattle Symposium on Biostatistics; New York: Springer; 2004. [Google Scholar]
27.Robins JM, Orellana L, Rotnitzky A. Estimation and extrapolation of optimal treatment and testing strategies. Statistics in Medicine. 2008;27(23):4678–4721. doi: 10.1002/sim.3301. [DOI] [PubMed] [Google Scholar]
28.Stoehlmacher J, Park DJ, Zhang W, Yang D, Groshen S, Zahedy S, Lenz HJ. A multivariate analysis of genomic polymorphisms: prediction of clinical outcome to 5-FU/oxaliplatin combination chemotherapy in refractory colorectal cancer. British Journal of cancer. 2004;91(2):344–354. doi: 10.1038/sj.bjc.6601975. [DOI] [PMC free article] [PubMed] [Google Scholar]
29.Thall PF, Millikan RE, Sung HG. Evaluating multiple treatment courses in clinical trials. Statistics in Medicine. 2000;19:1011–1028. doi: 10.1002/(sici)1097-0258(20000430)19:8<1011::aid-sim414>3.0.co;2-m. [DOI] [PubMed] [Google Scholar]
30.Thall PF, Sung HG, Estey EH. Selecting therapeutic strategies based on efficacy and death in multicourse clinical trials. Journal of the American Statistical Association. 2002;97:29–39. [Google Scholar]
31.Tibshirani R. Regression shrinkage and selection via the Lasso. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 1996;32:135–166. [Google Scholar]
32.Tsybakov AB. Optimal aggregation of classifiers in statistical learning. The Annals of Statistics. 2004;32:135–166. [Google Scholar]
33.van de Geer S. High-dimensional generalized linear models and the Lasso. The Annals of Statistics. 2008;36(2):614–645. [Google Scholar]
34.van der Laan MJ, Petersen ML, Joffe MM. History-Adjusted Marginal Structural Models and Statically-Optimal Dynamic Treatment Regimens. The International Journal of Biostatistics. 2005;1(1) Article 4. [Google Scholar]
35.Wahed AS, Tsiatis AA. Semiparametric efficient estimation of survival distribution for treatment policies in two-stage randomization designs in clinical trials with censored data. Biometrika. 2006;93:163–177. [Google Scholar]
36.Zhang CH, Huang J. The sparsity and bias of the lasso selection in high-dimensional linear regression. The Annals of Statistics. 2008;36(4):1567–1594. [Google Scholar]
37.Zou H. The Adaptive Lasso and its Oracle Properties. Journal of the American Statistical Association. 2006;101(476):1418–1429. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

NIHMS266525-supplement-2.pdf^{(136.5KB, pdf)}

NIHMS266525-supplement-supplement_1.pdf^{(99.3KB, pdf)}

[R1] 1.Bartlett PL, Jordan ML, McAuliffe JD. Convexity, classification, and risk bounds. Journal of the American Statistical Association. 2006;135(3):311–334. [Google Scholar]

[R2] 2.Bartlett PL. Fast rates for estimation error and oracle inequalities for model selection. Econometric Theory. 2008;24(2):545–552. [Google Scholar]

[R3] 3.Bickel PJ, Ritov Y, Tsybakov AB. Simultaneous analysis of lasso and dantzig selector. The Annals of Statistics. 2009;37(4):1705–1732. [Google Scholar]

[R4] 4.Bunea F, Tsybakov AB, Wegkamp MH. Sparsity oracle inequalities for the Lasso. Electronic Journal of Statistics. 2007;1:169–194. [Google Scholar]

[R5] 5.Cai T, Tian L, Lloyd-Jones DM, Wei LJ. Evaluating Subject-level Incremental Values of New Markers for Risk Classification Rule. Harvard University Biostatistics Working Paper Series. Working Paper 91. 2008a doi: 10.1007/s10985-013-9272-6. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R6] 6.Cai T, Tian L, Uno H, Solomon SD, Wei LJ. Calibrating Parametric Subject-specific Risk Estimation. Harvard University Biostatistics Working Paper Series. Working Paper 92 2008b [Google Scholar]

[R7] 7.Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2. Hillsdale, NJ: Lawrence Erlbaum Associates, Inc; 1988. [Google Scholar]

[R8] 8.Donoho D, Johnstone I. Ideal spatial adaptation by wavelet shrinkage. Biometrika. 1994;81(3):425–455. [Google Scholar]

[R9] 9.Fan J, Li R. Variable selection via nonconcave penalized likelihood and it oracle properties. Journal of the American Statistical Association. 2001;96:1348–1360. [Google Scholar]

[R10] 10.Feldstein ML, Savlov ED, Hilf R. A statistical model for predicting response of breast cancer patients to cytotoxic chemotherapy. Cancer Research. 1978;38(8):2544–2548. [PubMed] [Google Scholar]

[R11] 11.Insel TR. Translating scientific opportunity into public health impact: a strategic plan for research on mental illness. Archives of General Psychiatry. 2009;66(2):128–33. doi: 10.1001/archgenpsychiatry.2008.540. [DOI] [PubMed] [Google Scholar]

[R12] 12.Ishigooka J, Murasaki M, Miura S The Olanzapine Late-Phase II Study Group. Olanzapine optimal dose: Results of an open-label multicenter study in schizophrenic patients. Psychiatry and Clinical Neurosciences. 2001;54(4):467–478. doi: 10.1046/j.1440-1819.2000.00738.x. [DOI] [PubMed] [Google Scholar]

[R13] 13.Keller MB, McCullough JP, Klein DN, Arnow B, Dunner DL, Gelenberg AJ, Markowitz JC, Nemeroff CB, Russell JM, Thase ME, Trivedi MH, Zajecka J. A comparison of nefazodone, the cognitive behavioral-analysis system of psychotherapy, and their combination for the treatment of chronic depression. The New England Journal of Medicine. 2000;342(20):1462–1470. doi: 10.1056/NEJM200005183422001. [DOI] [PubMed] [Google Scholar]

[R14] 14.Kent DM, Hayward RA, Griffith JL, Vijan S, Beshansky JR, Califf RM, Selker HP. An independently derived and validated predictive model for selecting patients with myocardial infarction who are likely to benefit from tissue plasminogen activator compared with streptokinase. The American Journal of Medicine. 2002;113(2):104–11. doi: 10.1016/s0002-9343(02)01160-9. [DOI] [PubMed] [Google Scholar]

[R15] 15.Koltchinskii V. Sparsity in penalized empirical risk minimization. Annales de l’Institut Henri Poincaré Probabilitiés et Statistiques. 2009;45(1):7–57. [Google Scholar]

[R16] 16.Lesko LJ. Personalized Medicine: Elusive Dream or Imminent Reality? Clinical Pharmacology and Therapeutics. 2007;81:807–816. doi: 10.1038/sj.clpt.6100204. [DOI] [PubMed] [Google Scholar]

[R17] 17.Lunceford JK, Davidian M, Tsiatis AA. Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials. Biometrics. 2002;58:48–57. doi: 10.1111/j.0006-341x.2002.00048.x. [DOI] [PubMed] [Google Scholar]

[R18] 18.Mammen E, Tsybakov A. Smooth discrimination analysis. The Annals of Statistics. 1999;27:1808–1829. [Google Scholar]

[R19] 19.Massart P. Ecole d’Eté de Probabilités de Saint-Flour XXXIII, Concentration inequalities and model selection. Springer; 2003. [Google Scholar]

[R20] 20.Massart P. A non asymptotic theory for model selection. Proceedings of the 4th European Congress of Mathematicians (Ed. Ari Laptev); European Mathematical Society; 2005. pp. 309–323. [Google Scholar]

[R21] 21.Murphy SA, van der Laan MJ, Robins JM CPPRG. Marginal mean models for dynamic regimes. Journal of the American Statistical Association. 2001;96:1410–1423. doi: 10.1198/016214501753382327. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R22] 22.Murphy SA. Optimal Dynamic Treatment Regimes. Journal of the Royal Statistical Society, Series B (with discussion) 2003;65(2):331–366. [Google Scholar]

[R23] 23.Murphy SA. A Generalization error for Q-Learning. Journal of Machine Learning Research. 2005;6:1073–1097. [PMC free article] [PubMed] [Google Scholar]

[R24] 24.Piquette-Miller P, Grant DM. The Art and Science of Personalized Medicine. Clinical Pharmacology and Therapeutics. 2007;81:311–315. doi: 10.1038/sj.clpt.6100130. [DOI] [PubMed] [Google Scholar]

[R25] 25.Polonik W. Measuring mass concentrations and estimating density contour clusters - an excess mass approach. The Annals of Statistics. 1995;23(3):855–881. [Google Scholar]

[R26] 26.Robins JM. Optimal-regime structural nested models. In: Lin DY, Haegerty P, editors. Lecture notes in Stastitics; Proceedings of the Second Seattle Symposium on Biostatistics; New York: Springer; 2004. [Google Scholar]

[R27] 27.Robins JM, Orellana L, Rotnitzky A. Estimation and extrapolation of optimal treatment and testing strategies. Statistics in Medicine. 2008;27(23):4678–4721. doi: 10.1002/sim.3301. [DOI] [PubMed] [Google Scholar]

[R28] 28.Stoehlmacher J, Park DJ, Zhang W, Yang D, Groshen S, Zahedy S, Lenz HJ. A multivariate analysis of genomic polymorphisms: prediction of clinical outcome to 5-FU/oxaliplatin combination chemotherapy in refractory colorectal cancer. British Journal of cancer. 2004;91(2):344–354. doi: 10.1038/sj.bjc.6601975. [DOI] [PMC free article] [PubMed] [Google Scholar]

[R29] 29.Thall PF, Millikan RE, Sung HG. Evaluating multiple treatment courses in clinical trials. Statistics in Medicine. 2000;19:1011–1028. doi: 10.1002/(sici)1097-0258(20000430)19:8<1011::aid-sim414>3.0.co;2-m. [DOI] [PubMed] [Google Scholar]

[R30] 30.Thall PF, Sung HG, Estey EH. Selecting therapeutic strategies based on efficacy and death in multicourse clinical trials. Journal of the American Statistical Association. 2002;97:29–39. [Google Scholar]

[R31] 31.Tibshirani R. Regression shrinkage and selection via the Lasso. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 1996;32:135–166. [Google Scholar]

[R32] 32.Tsybakov AB. Optimal aggregation of classifiers in statistical learning. The Annals of Statistics. 2004;32:135–166. [Google Scholar]

[R33] 33.van de Geer S. High-dimensional generalized linear models and the Lasso. The Annals of Statistics. 2008;36(2):614–645. [Google Scholar]

[R34] 34.van der Laan MJ, Petersen ML, Joffe MM. History-Adjusted Marginal Structural Models and Statically-Optimal Dynamic Treatment Regimens. The International Journal of Biostatistics. 2005;1(1) Article 4. [Google Scholar]

[R35] 35.Wahed AS, Tsiatis AA. Semiparametric efficient estimation of survival distribution for treatment policies in two-stage randomization designs in clinical trials with censored data. Biometrika. 2006;93:163–177. [Google Scholar]

[R36] 36.Zhang CH, Huang J. The sparsity and bias of the lasso selection in high-dimensional linear regression. The Annals of Statistics. 2008;36(4):1567–1594. [Google Scholar]

[R37] 37.Zou H. The Adaptive Lasso and its Oracle Properties. Journal of the American Statistical Association. 2006;101(476):1418–1429. [Google Scholar]

PERMALINK

PERFORMANCE GUARANTEES FOR INDIVIDUALIZED TREATMENT RULES

Min Qian

Susan A Murphy

Abstract

1. Introduction

2. Individualized treatment rules

3. Relating the reduction in Value to excess prediction error

Theorem 3.1

Remarks

A toy example

4. Estimation via l1-penalized least squares

4.1. Performance guarantee for the l1-PLS

Theorem 4.1

Remarks

Theorem 4.2

Remarks

Corollary 4.1

4.2. Prediction error bound for the l1-PLS estimator

Theorem 4.3

Remarks

5. A Practical Implementation and an Evaluation

5.1. Simulations

Fig 1.

Fig 2.

5.2. Nefazodone-CBASP trial example

6. Discussion

Supplementary Material

Acknowledgments

APPENDIX

A.1. Proof of Theorem 3.1

A.2. Generalization of Theorem 4.3

Assumption A.1

Assumption A.2

Assumption A.3

Assumption A.4

Theorem A.1

Remark

Proof

Lemma A.1

Remark

Lemma A.2

Lemma A.3

Lemma A.4

Lemma A.5

A.3. Design of simulations in Section 5.1

Remark

Contributor Information

References

Associated Data

Supplementary Materials

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases

4. Estimation via l₁-penalized least squares

4.1. Performance guarantee for the l₁-PLS

4.2. Prediction error bound for the l₁-PLS estimator