Summary
We construct optimal designs for group testing experiments where the goal is to estimate the prevalence of a trait using a test with uncertain sensitivity and specificity. Using optimal design theory for approximate designs, we show that the most efficient design for simultaneously estimating the prevalence, sensitivity, and specificity requires three different group sizes with equal frequencies. However, if estimating prevalence as accurately as possible is the only focus, the optimal strategy is to have three group sizes with unequal frequencies. Based on a Chlamydia study in the United States, we compare performances of competing designs and provide insights into how the unknown sensitivity and specificity of the test affect the performance of the prevalence estimator. We demonstrate that the proposed locally D- and Ds-optimal designs have high efficiencies even when the prespecified values of the parameters are moderately misspecified.
Keywords: D-optimality, Ds-optimality, Group testing, Sensitivity, Specificity
1. Introduction
In a group testing study, the goal is often to estimate the prevalence of a rare disease or a particular trait (Hughes-Oliver & Swallow, 1994; Hughes-Oliver & Rosenberger, 2000). Group testing is frequently used in studies where testing individuals for a trait is expensive and individual samples are relatively plentiful. In a group testing study, samples from individuals are pooled and tested as a single unit. Fewer tests or trials are therefore required so the cost of the study is reduced. (Throughout, we use the terms trials and tests interchangeably.) The key assumption is that the test result from a group of individuals is positive if and only if at least one individual in the group has the trait.
Applications of group testing abound in the literature, especially for case diagnosis (McMahan et al., 2012) and prevalence estimation of diseases (Shipitsyna et al., 2007). Several researchers have shown that group testing techniques can be used to efficiently estimate the prevalence of a rare disease (Hughes-Oliver, 2006). Dorfman (1943) gave an early perspective and an overview. From the design point of view, some key research questions are to determine how many distinct group sizes should be used, and how many trials at each such group size should be run. In practice, group testing studies usually have only one group size, but this may be because the potential advantages of using unequal group sizes have not been well explored.
Much of the work on group testing makes a simplifying assumption that there is no testing error, see for example, Hughes-Oliver & Rosenberger (2000). In practice, the test in most group testing experiments is likely to have nonzero error rates, with sensitivity and specificity being less than 100%. We recall that sensitivity is the probability that the test will be positive given that the tested sample is truly positive, and specificity is the probability that the test will be negative given that the tested sample is truly negative.
A few authors have considered application of group testing when there may be testing errors, but they assumed that the sensitivity and specificity are known and do not depend on the group size. When the total number of individuals is fixed and all group sizes are equal, Tu et al. (1995) obtained the maximum likelihood estimator of the prevalence and determined graphically the common group size that minimizes the variance of the prevalence estimator. For a given group size, Liu et al. (2012) determined conditions under which group testing performs better than individual testing, when either the total number of trials or the total number of individuals is fixed.
The sensitivity and specificity can be directly estimated if samples are tested using both the test of interest and a “gold standard” test that has no testing errors. Such a gold standard test may not exist however, and even if one does exist, it may be too expensive or complex for routine use. Lacking a gold standard test, investigators may prespecify sensitivity and specificity values. However, nominal values may be optimistic, and values estimated from different populations could be biased. Therefore it is more realistic in applications to assume that the sensitivity and specificity of the test are both unknown. Under this assumption, Zhang et al. (2014) used numerical methods to obtain group testing designs with or without a gold standard test. Their proposed designs are not fully supported by theory as being optimal.
When the cost of a test is much greater than the cost of collecting an individual sample, it may be reasonable to design the study taking the total number of trials, rather than the total number of individuals, to be fixed. Our goal is to apply optimal design theory to construct optimal group testing designs when we have a fixed number of trials, and when both the sensitivity and specificity of the test are unknown. We consider two situations: (i) we want to estimate the prevalence, sensitivity and specificity, all with equal interest; and (ii) only the prevalence is of interest, but both testing error rates are treated as nuisance parameters in the design problem.
In the framework of optimal experimental design (Atkinson et al., 2007; Berger & Wong, 2009), the first setting requires estimating all unknown parameters, and the second setting requires one to estimate only one of the three unknown parameters in the model. D-optimality as a design criterion is appropriate for the first situation and Ds-optimality is appropriate for the second situation. The resulting optimal designs are called D-optimal and Ds-optimal respectively. These optimal designs minimize the generalized variance of all or a subset of the model parameters of interest among all designs, so they provide the most accurate inference for all three or just the prevalence parameter. Other design criteria, such as A-optimality or c-optimality, can also be used.
Our setup is different from earlier studies for finding efficient or optimal designs for group testing experiments, such as Tu et al. (1995) and Zhang et al. (2014). The main differences are that: (i) our optimal designs are sought among all possible designs without any restriction on the number of different group sizes; and (ii) our optimal designs are justified by theory. Tu’s results are more restrictive because their designs are sought among all designs with a common group size and cannot incorporate uncertainty in testing error rates. Zhang’s designs are obtained by numerical optimization with dimension as large as the upper bound of the group sizes; therefore their global optimality cannot be guaranteed, and one may not be able to interpret how prespecified parameter values affect their optimal designs. To our best knowledge, our work is the first one that uses optimal design theory to provide a theoretical justification for designing group testing experiments when testing error rates are uncertain.
The organization of our paper is as follows. Section 2 describes the statistical background. Section 3 constructs optimal group testing designs for the two situations described above. In Section 4, we evaluate the finite sample properties and robustness properties of our proposed optimal designs when the parameter vector is misspecified. The setup of the simulation is based on a Chlamydia study conducted in the United States. We end the paper with a discussion in Section 5 with remarks on alternative methods and ideas for further research directions in this area.
All computations in this paper were done using our code called gtest.nb developed using Mathematica 10.0. The package can be freely downloaded from http://www.math.nsysu.edu.tw/~lomn/huangsh/.
2. Preliminaries
Throughout, we denote the prevalence, sensitivity and specificity of the test by p0, p1 and p2, respectively. We assume that p0 ∈ (0, 1), p1, p2 ∈ (0.5, 1], and the two types of testing errors occur randomly with testing error rates 1 − p1 and 1 − p2 respectively. Letting θ = (p0, p1, p2)T, a direct calculation shows that the probability of a positive response (either a true positive or a false positive) for a trial with group size x is
| (1) |
In practice, there are generally limits on the group sizes so we consider designs subject to a known group size constraint 1 ≤ xL ≤ x ≤ xU < ∞. We note that π(x) is a convex combination of p1 and 1 − p2. For a study consisting of n trials at k distinct group sizes x1 < x2 ⋯ < xk, the log-likelihood function of θ is
| (2) |
apart from an unimportant additive constant. Here each yi is the number of positive responses among the ni trials at group size xi and . We assume that each yi has a Binomial distribution with parameters (ni, π (xi|θ)), and the are mutually independent.
In what follows, we work with approximate designs advocated by Kiefer (1974). An approximate group testing design denoted by contains k distinct group sizes (or support points) x1 < x2 < ⋯ < xk and corresponding weights w1, w2, …, wk, where and each wi > 0 is interpreted as the proportion of trials at group size xi. When the total number of trials n is fixed a priori by either cost or practical considerations, the group testing design ξ takes nwi groups of size xi, i = 1, …, k subject to nw1 + … + nwk = n. In reality, the group sizes and the numbers of trials at these sizes must be positive integers. Therefore, optimal approximate group testing designs are rounded before implementation so that each group size and its frequency are positive integers. We call such designs rounded optimal designs. Simple rounding procedures to obtain a design for implementation have been shown to produce little loss in efficiency. More efficient rounding apportionment introduced in section 12.5 of Pukelsheim (2006) can also be applied to obtain the optimal number of trials at the rounded group size.
Under model (1), the maximum likelihood estimator (MLE) of θ is obtained by maximizing (2). If design ξ were used to obtain the data, a direct calculation shows the covariance matrix of the estimated θ is asymptotically proportional to the inverse of the 3 × 3 Fisher information matrix given by
| (3) |
Here
| (4) |
and
The design we seek for estimating θ is a D-optimal design that maximizes the determinant of the information matrix (3), or equivalently a design that minimizes the generalized variance of all parameter estimators, by choice of k, x1, …, xk, w1, …, wk. In other words, a D-optimal design maximizes
| (5) |
among all possible designs under which θ is estimable. When we are only interested in estimating a subset of the model parameters, we use the Ds-optimality criterion. Here we are only interested in accurately estimating the prevalence, and treat sensitivity and specificity as nuisance parameters. The Ds-criterion here is to minimize only the asymptotic variance of the prevalence estimator. Such a Ds-optimal design maximizes
| (6) |
among all designs under which p0 is estimable, where M− is a generalized inverse of matrix M and we use the notation O11 to denote the (1, 1) entry of the matrix O. Note that the Ds-optimality above is equivalent to c-optimality with c = (1, 0, 0)T (Fedorov, 1972).
The solutions to our optimization problems require solving the equations in Lemma 1 below. Its proof mainly requires calculus and the intermediate value theorem and is therefore omitted. Before stating Lemma 1, we introduce further notation. Let
For a ∈ (r, 1), let
and let
The above definitions imply that r < a < 1 < c which is used in the proofs without further mention.
Lemma 1
Each of the following equations has a root a ∈ (r, 1):
| (7) |
and
| (8) |
3. The D- and Ds-optimal designs
Before characterizing the D- and Ds-optimal designs, we first characterize a small class of designs guaranteed to contain D- and Ds-optimal designs, and then obtain the optimal designs from the class. Such an approach is based on so-called “essentially complete classes” used in Cheng (1995) and Yang & Stufken (2009), among others, to identify optimal designs in certain settings. Recently, a series of papers, including Yang & Stufken (2012) and Hu et al. (2015), proposed a framework to obtain essentially complete classes, and provided several important properties of optimal designs identified in the class. We follow such an approach.
A class 𝒞 is called an essentially complete class if for an arbitrary design ξ1 ∉ 𝒞, there exists a design ξ2 ∈ 𝒞 such that M(ξ2) − M(ξ1) is nonnegative definite. Therefore, it guarantees that there exists an optimal design in 𝒞 when the design criterion has certain properties such as concavity, isotonicity, and smoothness. For example, Assumption A in Hu et al. (2015) includes criteria like D- and Ds-optimality, and other criteria based on the information matrix. Let
Note that by allowing wi = 0 at some points, this class contains all one- and two-point designs, and all three-point designs with a support point at xL. By applying Theorem 2(a) in Yang & Stufken (2012), we have the following theorem which allows us to search for a D- and a Ds-optimal group testing design in 𝒞0
Theorem 1
The class 𝒞0 is essentially complete.
The justifications of this result and others are given in the Appendix.
3.1. The unique D-optimal design
When we wish to estimate prevalence, sensitivity, and specificity simultaneously, we need a design with at least three points to estimate the three parameters. The sought design is D-optimal for estimating the three parameters of the model. From Theorem 1, such a D-optimal group testing design exists in 𝒞0 and the following theorem ensures the uniqueness of the D-optimal design and characterizes it.
Theorem 2
The D-optimal design ξD for estimating the prevalence and the two testing error rates is unique. It has three group sizes xL, , and xU with equal weights, where and A1 uniquely solves (7).
Theorem 2 asserts that the D-optimal design requires equal numbers of trials at three distinct group sizes, with two of them being the extreme sizes xL and xU. No design with four or more group sizes can estimate the three parameters more accurately than the D-optimal design. In addition, our proof of Theorem 2 implies that (i) the root of equation (7) must be unique in (r, 1), or otherwise it contradicts the uniqueness of the D-optimal design, and (ii) either having a smaller value of xL or having a larger value of xU always strictly reduces the generalized variance of the MLE of θ. Thus, we suggest the lower bound xL to be one whenever possible. However, too large a value of xU may cause a dilution effect (Zenios & Wein, 1998), impacting the sensitivity or the specificity of the test. Thus, the upper bound xU should be set carefully.
When p0 and xL are small and xU is large, we observe that π(xL) is close to 1 − p2 and π(xU) is close to p1, which implies that trials at size xL and xU mainly provide information for p2 and p1 respectively. We may conclude that contributes most for estimating p0. In practice, p1 and p2 for a given test may be estimated by comparing the test results with a gold standard test having no testing errors. In the absence of a gold standard, the maximum likelihood approach can be seen as comparing the test results from the intermediate group size to those from the two extreme group sizes, to acquire information about the error parameters p1, p2, and the prevalence p0. Accordingly, we may regard the test results from the two extreme group sizes as playing a similar role as a gold standard test for providing information on p1 and p2.
3.2. The unique Ds-optimal design
Theorem 2 assumes that there is equal interest in estimating prevalence, sensitivity and specificity. In practice, the interest is usually to accurately estimate the prevalence alone, and the sensitivity and specificity of the test are treated as unknown nuisance parameters. When the main goal is only to estimate the prevalence as accurately as possible, the Ds-optimality criterion (6) is the most suitable. A Ds-optimal design minimizes the (1, 1) element of the inverse of the information matrix (3), which is proportional to the asymptotic variance of the prevalence estimator.
To determine the optimal weights under the Ds-criterion for a fixed number of group sizes, suppose that there are three group sizes x1 < x2 < x3 in the allowable range [xL, xU]. Let
for i = 1, 2, 3, where and are the two group sizes of {x1, x2, x3} \ {xi} such that . To be more specific, , and . For any group testing design with three group sizes, the next lemma prescribes the optimal weights under the Ds-criterion.
Lemma 2
For any group testing design with three group sizes, x1, x2, and x3 ∈ [xL, xU], the optimal weights under the Ds-criterion for only estimating the prevalence are unique and given by
| (9) |
for i = 1, 2, and 3, and
| (10) |
Our next result gives a complete description of the Ds-optimal design for estimating the prevalence when the testing error rates are unknown.
Theorem 3
The Ds-optimal design ξs for estimating the prevalence alone is unique and requires three group sizes xL, , and xU, where and As uniquely solves (8). The proportions of groups with such sizes are given in (9) in Lemma 2.
We observe that (8) approximates (7) when Δ1 and Δ2 are both close to zero, which may occur in practice when π(xL) tends to zero and π(xU) tends to one. Consequently, A1 ≈ As, and the intermediate support point of the Ds-optimal design for estimating the prevalence is close to the intermediate support of the corresponding D-optimal design.
Theorem 3 shows that the Ds-optimal design for estimating prevalence with unknown testing error rates has similar properties as the D-optimal design for estimating the three parameters. In particular, (i) no design with four or more different group sizes can estimate the prevalence more accurately than the Ds-optimal design with three different group sizes, (ii) the root of equation (8) must be unique in (r, 1), (iii) either having a smaller value of xL or having a larger value of xU strictly improves the accuracy of prevalence estimation, and (iv) groups with size xL and xU mainly contribute information about the specificity and sensitivity, respectively, and the groups with the intermediate size mainly contribute information about the prevalence.
When xL is small and xU is large, which implies that π(xL) is much smaller than 0.5 and π(xU) is much larger than 0.5 by (1), then is much larger than and . Consequently, the proportion of groups with the intermediate size is much larger than the other two group sizes. This observation is not surprising partly because the Ds-optimal design criterion seeks information for the prevalence only and does so by having more groups with the intermediate size.
4. Design performance
The optimal designs discussed in the previous section are locally optimal approximate designs, and their group sizes and the number of trials at each group size must be rounded to positive integers for implementation. In addition, their information matrices are based on the asymptotic distribution of the MLE of the parameter vector. In what follows, we study the properties of the rounded optimal designs when we have moderate sample size (total number of trials), and their sensitive to the prespecified value of the parameter vector.
Recall that θ = (p0, p1, p2)T is the true but unknown vector of parameters in our model, [xL, xU] is the user-specified allowable range of group sizes, and we are interested in either estimating θ or only the prevalence parameter p0. For an approximate design with integer-valued group sizes, denote its exact design with n trials as , where n1, …, nk are integers obtained by applying the efficient rounding apportionment (Pukelsheim, 2006, section 12.5) to {nw1, …, nwk}. Let θ̂ξ(n) be the MLE of the parameter vector θ under the exact design ξ(n). We define the mean squared error matrix of θ̂ξ(n) (scaled by sample size n) by
| (11) |
Note that mse(θξ(n)) converges to M(ξ)−1 as n goes to infinity by the asymptotic property of MLE.
The analytical form of mse(θ̂ξ(n)) is quite complicated unless the sample size n is very small. We therefore study mse(θ̂ξ(n)) by simulation as follows. Let N be the number of simulation replications. For t = 1, …, N, we simulate a sample of n trials generated from ξ(n), consisting of ni binary outcomes with response probability π(xi|θ) for i = 1 … k, and compute , the MLE of θ from the sample. The simulation-based mean squared error matrix of ξ is defined by
| (12) |
We use two measures to assess the finite sample performance of ξ(n). They are the (simulation-based) D- and Ds-efficiencies, defined respectively, by
| (13) |
where ξD and ξs are the D- and Ds-optimal designs obtained from Theorems 2 and 3.
In the following, we set the number of simulation replications to N = 10000, and set the prespecified values of the test parameters to match the Chlamydia study from Nebraska, the United States (McMahan et al., 2012, Table 1). They prespecified the prevalence to be 0.07, and the sensitivity and specificity to be 0.93 and 0.96 for female’s swab specimens. Suppose we have a budget for n = 3000 trials, and the allowable range of the group size is [xL, xU] = [1, 61].
Table 1.
The D- and Ds-efficiencies of the rounded D- and Ds-optimal designs, and the k-point uniform designs for k = 3, …, 6.
| Design | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| effD | 0.9904 | 0.8193 | 0.7988 | 0.8456 | 0.8133 | 0.7805 | ||||||
| effs | 0.6977 | 1.0052 | 0.3170 | 0.4851 | 0.5274 | 0.5509 |
4.1. Finite sample performance of the rounded optimal designs
First, we construct the D- and Ds-optimal designs, and show that the effects of rounding (both the group sizes and their number of support points) and the use of asymptotic approximations are small. Under θ = (0.07, 0.93, 0.96)T and [xL, xU] = [1, 61], the D-optimal design ξD from Theorem 2 is equally supported on {1, 16.79, 61}. For n = 3000, its rounded design has group sizes {1, 17, 61} with 1000 groups per size.
Similarly, the Ds-optimal design ξs obtained from Theorem 3 is supported on {1, 15.68, 61}. Its rounded design is therefore supported on {1, 16, 61}, and has {393, 1884, 723} trials at {1, 16, 61} respectively. These numbers of trials are based on the Ds-optimal weights on {1, 16, 61} by Lemma 2, which are slightly different from the weights of ξs. The simulation results in Table 1 show that the D-efficiency of and the Ds-efficiency of are both close to one. This indicates that the performance of the rounded D- (and Ds-) optimal design is similar to the asymptotic performance of the D- (and Ds-) optimal approximate design.
Next, we compare the finite sample performance of the rounded D- and Ds-optimal designs to k-point uniform designs , supported on the positive integers , for k = 3, …, 6. For example, has 750 trials at each group size of {1, 21, 41, 61}. Table 1 also shows the performance of these k-point uniform designs. We observe that although the optimal designs ξD and ξs have exactly three support points, is not the best one among these uniform designs regarding D- or Ds-efficiency. On the other hand, performs best for estimating θ among these uniform designs, but its D-efficiency is about 0.85 and is less than . Furthermore, the design has highest Ds-efficiency among these uniform designs, but its Ds- efficiency is only about 0.55 and is significantly less than . This indicates that when focusing on estimating θ (or p0), the rounded D- (or Ds-) optimal design under the true parameter vector performs much better than the uniform designs above.
4.2. Robustness of the rounded optimal designs
Now we consider the situation where the prespecified value of θ used in the local design construction is incorrect. Specifically, suppose that the prespecified parameter vector θ̃ = (p̃0, p̃1, p̃2)T is a point in the region Θ = [0.01, 0.10] × [0.9, 1]2, containing the true value of the parameter vector θ = (0.07, 0.93, 0.96)T.
Let be the rounded D-optimal design under the prespecified parameter vector θ̃. Note that in the D-optimality, the prespecified value of the parameter vector only affects the intermediate group size in Theorem 2, but does not affect the other two group sizes and the weights on the three support points. Hence, is equally supported on {1, x̃, 61}, where x̃ depends on θ̃. Figure 1 shows the corresponding x̃ for selected θ̃ ∈ Θ. We observe that when the specified θ̃ is within Θ, x̃ shifts from 12 to 25, where x̃ = 17 for θ̃ = θ; x̃ increases as p̃0 or p̃2 decreases, or as p̃1 increases. When the prevalence is small such as 0.01, the specificity seems to play a dominant role on the D-optimal design. While as the prevalence gets larger, the dominant factor shifts from the specificity to sensitivity; and the sensitivity and specificity have similar impact on the design when the prevalence is close to 0.04.
Fig. 1.
The intermediate group size x̃ of the rounded D-optimal designs under a prespecified θ̃ = (p̃0, p̃1, p̃2)T. Here p̃0 = 0.01, 0.04, 0.07, and 0.10; p̃1 and p̃2 = 0.90, 0.91, …, 0.99, and 1.00.
Figure 2 shows the D-efficiency of for θ̃ ∈ Θ in terms of x̃, where x̃ = 12 to 25. Generally, these rounded D-optimal designs under misspecified θ̃ ∈ Θ have D-efficiencies greater than 90%. Note that the rounded D-optimal design is equally supported on three group sizes including two extreme ones. In contrast, the design has the same form but its D-efficiency is only about 0.8. This indicates that the rounded D-optimal design under a misspecified parameter vector is quite robust for estimating θ.
Fig. 2.
The D-efficiency of the design that is equally supported on {1, x̃, 61} with n = 3000, for x̃ = 12 to 25. These designs arise as rounded D-optimal designs for θ̃ ∈ Θ as defined in the text.
Similar to the parameter misspecification setting for the D-optimal design, we obtain the rounded Ds-optimal designs designs supported on {1, x̃s, 61} under a possibly misspecified parameter vector θ̃, where x̃s depends on θ̃ ∈ Θ. A direct calculation shows that the intermediate group size x̃s is between 11 and 25, and the weights of group sizes 1, x̃s, and 61 are respectively within 0.140 ± 0.05, 0.67 ± 0.14, and 0.19 ± 0.12.
Figure 3 shows the Ds-efficiencies of for selected θ̃ ∈ Θ. We observe that is larger than or close to 0.8 when the prespecified θ̃ belongs in the region [0.04, 0.10] × [0.90, 0.99] × [0.90, 1.00] ⊂ Θ, which limits the distance from θ̃ to the true value of θ = (0.07, 0.93, 0.96)T. However, when p̃0 = 0.01 (not near the true prevalence p0 = 0.07) or p̃1 = 1.00 (not near the true sensitivity p1 = 0.93), drops to as low as 0.6. We observe that the Ds-optimal design is slightly more sensitive than the D-optimal design when there is substantial parameter misspecification, but still performs well when the prespecified parameter vector is not far from its true value.
Fig. 3.
The Ds-efficiency of the rounded Ds-optimal design under prespecified θ̃ = (p̃0, p̃1, p̃2)T. Here p̃0 = 0.01, 0.04, 0.07, and 0.10; p̃1 = 0.90 (dashed), 0.93 (thick), 0.96 (dotted), 0.99 (dot-dashed), and 1.00 (thin); p̃2 = 0.90, 0.93, 0.96, 0.99, and 1.00.
Finally, comparing Figures 2 and 3 to Table 1, we observe that when the prespecified values of parameters are moderately misspecified, a rounded D- (or Ds-) optimal design still has at least 0.9 D-efficiency (or 0.8 Ds-efficiency), but those uniform designs only have at most 0.85 D-efficiency (or 0.55 Ds-efficiency). Our conclusion is that when we wish to estimate θ (or p0), we would recommend implementing a rounded D- (or Ds-) optimal design over any of the uniform design that we had investigated.
5. Discussion
Our work addresses several design issues for group testing studies using optimal design theory. Under uncertain testing error rates, we provide locally optimal designs for estimating prevalence alone or jointly with sensitivity and specificity. Our designs are justified by rigorous theoretical analysis and they provide the most accurate parameter estimates among all group testing designs. By contrasting the D- and Ds-optimal designs, we observe that relatively fewer trials at extreme group sizes are required in the Ds-optimal design, where the sensitivity and specificity are treated as nuisance parameters in the Ds-criterion. In the optimal approximate design, the theoretical group sizes and the frequencies of the group sizes are not guaranteed to be integers. We show by simulations that the rounded D- and Ds-optimal designs still have very high efficiencies. In addition, our simulation results show that our designs are still quite robust when the prespecified values of parameters are not far from their true values.
Although our theoretical results apply in general, our simulation studies focused on settings with 3000 trials, and had a maximum group size of approximately 60. This can be compared to two published studies using group testing: one in Russia estimating Chlamydia prevalence (Shipitsyna et al., 2007), which had three designs, involving 150 to 1500 trials, and one in the United States focusing on Chlamydia diagnosis (McMahan et al., 2012), which used 7000 trials. Thus the number of trials in our simulation study is similar to what has been used in practice, where the estimation burden is lighter because the sensitivity and specificity are treated as known there. We also note that the number of individual samples in our simulation example is larger than what were used in these two real-world studies. Since our designs are optimal, this reflects the unavoidable cost of estimating the prevalence as well as sensitivity and specificity from a single dataset with no prior knowledge about any of the parameters.
Our locally optimal designs can be used in the multi-stage adaptive approach as in Hughes-Oliver & Swallow (1994), where they assumed no testing errors. For example, a D- or Ds-optimal design may be implemented based on any available prior knowledge from the first stage. In the subsequent stages, we construct a D- or Ds-optimal design based on the information about the parameters obtained in the previous stages. Alternatively, a Bayesian optimal design approach or a minimax approach (Dette et al., 2014) can be implemented. For example, the experimenters may obtain a design that minimizes Ds-optimality criterion averaged over the parameters with respect to their prior distribution, or a design that minimizes the largest possible variance of the prevalence estimator among different values for the parameters. Such Bayesian or minimax optimization problems are difficult to address theoretically and even numerically, but are of interest for future studies.
The results presented here are appropriate for a setting in which the testing errors occur randomly and are unrelated to the group size within a prespecified range. In practice, the plausible range of group sizes comes from cost considerations or the physical constraints imposed on the study. A careful decision about the range of values for the group sizes is important because a group size that is too large may cause dilution effects, which in turn, may reduce the sensitivity and the specificity of the test. Our group testing framework can potentially be extended to accommodate dilution effects and other complicating issues, such as stratified populations. For example, we may wish to estimate the prevalence of a rare trait in multiple geographic regions with potentially different disease prevalences, or when the sensitivity and specificity of the test vary among the sub-populations. Such issues may be interesting directions for future research.
Acknowledgments
All authors would like to thank the reviewers for their helpful and thoughtful comments on earlier versions of this paper. The first two authors were supported by MOST 101-2118-M-110-002-MY2. The third author was supported by the National Center of Theoretical Sciences, Taiwan. The research of Wong reported in this publication was partially supported by the National Institute of General Medical Sciences of the National Institutes of Health under Award Number R01GM107639. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Appendix
Before giving the proofs of Theorem 1, Theorem 2, Lemma 2, and Theorem 3, we introduce additional notation. Let
Proof of Theorem 1
We use Theorem 2(a) in Yang & Stufken (2012) to show that 𝒞0 is essentially complete. First we rewrite the information matrix (3) as
| (14) |
Let fi1(a) = dΨi(a)/da for i = 1, …, 5 and let
Then for a ∈ [r, 1] ⊂ (0, 1], we have
Therefore, for a ∈ [r, 1]. Hence, the group testing model satisfies the conditions for Theorem 2(a) in Yang & Stufken (2012), and therefore the designs having at most two points together with the designs having three points including xL (which coincides with a = 1) form an essentially complete class. In addition, according to the proof of Theorem 2 in Yang & Stufken (2012), it can be shown that there exists s0, …, s5 = ±1 such that
| (15) |
This property will be used in the proof of Theorem 3.
Proof of Theorem 2
There are 3 steps in the proof. In step (a) we show that the D-optimal design is unique and requires three group sizes with equal frequencies, including one being the smallest group size xL. In step (b), we show that the upper bound xU is a support point of the D-optimal design. In step (c), we determine the intermediate group size.
By Theorem 1 and the fact that θ is only estimable under a design with at least 3 group sizes, there exists a D-optimal design ξD belonging to 𝒞0 with exactly 3 group sizes and one of them is xL. Because the number of support points of the D-optimal design ξD is equal to the number of unknown parameters, ξD must be equally supported (see for example, section 8.12 of Pukelsheim (2006)). By Theorem 2.5 in Hu et al. (2015) and the fact that D- and Φ0-criteria are equivalent, it follows that ξD is the unique D-optimal design.
-
Following (a), we denote the unique D-optimal design as , where . Recall that and r = (1 − p0)xU − xL. It suffices to show that for all a3 ∈ [r, a2), i.e., the D-criterion is a strictly decreasing function of a3. Thus the criterion is maximized by taking a3 as small as possible, which implies .
A direct calculation shows
and the following statements are true by calculus:- ;
- ;
- .
The last assertion holds because a2 log(a2) + a3(1 − a2)) < a2 log(a2) + a2(1 − a2)) < 0 and (− log(a3) + (1 − a3)a3) > 0 for all 0 < a3 < a2 < 1. Consequently,
and the desired result holds. - Our remaining goal is to determine . We do this by showing that ∂ΦD[M(ξD)]/∂a2 = 0 is equivalent to (7), and hence where A1 is the solution of (7). In (b) we have shown that the optimal choice for a3 is . By arguments similar to step (b), we have that
which is equivalent to (7), and hence the proof is complete.
Proof of Lemma 2
Let ξ be a design supported on {x1, x2, x3} with x1 < x2 < x3. Our problem here is to find the vector of the positive weights at these three given points that minimizes (M(ξ)−)11, where M(ξ) = Mf (ξ) · Diagλ(ξ) · Mf (ξ)T, Mf (ξ) = [f(x1) f(x2) f(x3)] is nonsingular and is positive-definite. Let e1 = (1, 0, 0)T. Then
and we have
| (16) |
Since Qi(x1, x2, x3) > 0 for i = 1, 2, 3 and |Mf (ξ)|2 > 0, we apply the method of Lagrange multipliers directly to minimize the value in (16) subject to the constraints on the weights. The resulting solution is displayed in (9).
Proof of Theorem 3
We prove this theorem by similar steps as those in the proof of Theorem 2. In Step (a), we show that (a.1) p0 is only estimable under a design with at least three distinct group sizes, and (a.2) the Ds-optimal design is unique. Therefore, by Theorem 1, the unique Ds-optimal design has exactly three group sizes, and one of them is the smallest allowable group size xL. Steps (b) and (c) of this theorem’s proof use similar arguments as those in the proof of Theorem 2 and are therefore omitted.
- The result is shown by contradiction. Without loss of generality, suppose there exists a design such that is estimable under ξ, where xL ≤ x1 < x2 ≤ xU, w1,w2 ≥ 0, w1 + w2 = 1, and e1 = (1, 0, 0)T. Therefore, e1 belongs to the range of M(ξ) = Mf (ξ) · Diagλ(ξ) · Mf (ξ)T, where Mf (ξ) = [f(x1) f(x2)] and Diagλ(ξ) = diag(w1λ(x1),w2λ(x2)). Hence, e1 belongs to the range of Mf (λ), or equivalently, the determinant of [f(x1) f(x2) e1] = 0. However,
for arbitrary x1 < x2 and p0 ∈ (0, 1). This contradiction shows that p0 is only estimable under a design with at least three points. - Suppose that there are two different Ds-optimal designs, ξ1 and ξ2, with at least 3 points. Let . Therefore, ξ must have at least 4 different support points due to ξ1 ≠ ξ2 and Lemma 2. By the concavity of Φs, we have that Φs[M(ξ)] ≥ Φs[M(ξ1)] = Φs[M(ξ2)], and hence ξ is also Ds-optimal. By the equivalence theorem (see for example, section 2.7 of Fedorov (1972) or section 10.3 of Atkinson et al. (2007)), we must have
with equality at each support point of ξ, where fs(x) = (f1(x), f2(x))T and Ms(ξ) is the 2 × 2 submatrix of M(ξ) deleting its first row and first column. Therefore, there exists a small enough ε > 0 such that ϕs(x, ξ)−ε has at least 4×2−2 = 6 roots. On the other hand, by direct calculation, for some di ∈ ℝ with , where a = (1 − p0)x− xL ∈ [r, 1] and is a Chebyshev system on [r, 1] by (15). Therefore, Theorem 4.1 on page 22 of Karlin & Studden (1966) shows that ϕs(x, ξ) − ε has at most 5 roots. This contradiction shows that the Ds-optimal design is unique.
Contributor Information
Shih-Hao Huang, National Sun Yat-sen University, Kaohsiung, Taiwan.; The University of Michigan, Ann Arbor, U.S.A.
Mong-Na Lo Huang, National Sun Yat-sen University, Kaohsiung, Taiwan..
Kerby Shedden, The University of Michigan, Ann Arbor, U.S.A..
Weng Kee Wong, University of California at Los Angeles, Los Angeles, U.S.A..
References
- Atkinson AC, Donev AN, Tobias RD. Optimum experimental designs, with SAS. Oxford: Oxford University Press; 2007. [Google Scholar]
- Berger MP, Wong WK. An introduction to optimal designs for social and biomedical research. Chichester: Wiley; 2009. [Google Scholar]
- Cheng C-S. Complete class results for the moment matrices of designs over permutation-invariant sets. The Annals of Statistics. 1995;23:41–54. [Google Scholar]
- Dette H, Kiss C, Benda N, Bretz F. Optimal designs for dose finding studies with an active control. Journal of the Royal Statistical Society: Series B. 2014;76:265–295. [Google Scholar]
- Dorfman R. The detection of defective members of large populations. The Annals of Mathematical Statistics. 1943;14:436–440. [Google Scholar]
- Fedorov VV. Theory of Optimal Experiments. New York: Academic Press; 1972. [Google Scholar]
- Hu L, Yang M, Stufken J. Saturated locally optimal designs under differentiable optimality criteria. The Annals of Statistics. 2015;43:30–56. [Google Scholar]
- Hughes-Oliver JM. Pooling experiments for blood screening and drug discovery. In: Dean A, Lewis S, editors. Screening: Methods for Experimentation in Industry, Drug Discovery, and Genetics. New York: Springer; 2006. pp. 46–68. [Google Scholar]
- Hughes-Oliver JM, Rosenberger WF. Efficient estimation of the prevalence of multiple traits. Biometrika. 2000;87:315–327. [Google Scholar]
- Hughes-Oliver JM, Swallow WH. A two-stage adaptive group-testing procedure for estimating small proportions. Journal of the American Statistical Association. 1994;89:982–993. [Google Scholar]
- Karlin S, Studden WJ. Tchebycheff Systems with Applications in Analysis and Statistics. New York: Wiley; 1966. [Google Scholar]
- Kiefer J. General equivalence theory for optimum designs (approximate theory) The Annals of Statistics. 1974;2:849–879. [Google Scholar]
- Liu A, Liu C, Zhang Z, Albert PS. Optimality of group testing in the presence of misclassification. Biometrika. 2012;99:245–251. doi: 10.1093/biomet/asr064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McMahan CS, Tebbs JM, Bilder CR. Informative Dorfman Screening. Biometrics. 2012;68:287–296. doi: 10.1111/j.1541-0420.2011.01644.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pukelsheim F. Optimal design of experiments. Philadelphia: SIAM; 2006. [Google Scholar]
- Shipitsyna E, Shalepo K, Savicheva A, Unemo M, Domeika M. Pooling samples: the key to sensitive, specific and cost-effective genetic diagnosis of Chlamydia trachomatis in low-resource countries. Acta Derm Venereol. 2007;87:140–143. doi: 10.2340/00015555-0196. [DOI] [PubMed] [Google Scholar]
- Tu XM, Litvak E, Pagano M. On the informativeness and accuracy of pooled testing in estimating prevalence of a rare disease: Application to HIV screening. Biometrika. 1995;82:287–297. [Google Scholar]
- Yang M, Stufken J. Support points of locally optimal designs for nonlinear models with two parameters. The Annals of Statistics. 2009;37:518–541. [Google Scholar]
- Yang M, Stufken J. Identifying locally optimal designs for nonlinear models: A simple extension with profound consequences. The Annals of Statistics. 2012;40:1665–1681. [Google Scholar]
- Zenios SA, Wein LM. Pooled testing for HIV prevalence estimation: Exploiting the dilution effect. Statistics in Medicine. 1998;17:1447–1467. doi: 10.1002/(sici)1097-0258(19980715)17:13<1447::aid-sim862>3.0.co;2-k. [DOI] [PubMed] [Google Scholar]
- Zhang Z, Liu C, Kim S, Liu A. Prevalence estimation subject to misclassification: the mis-substitution bias and some remedies. Statistics in Medicine. 2014;33:4482–4500. doi: 10.1002/sim.6268. [DOI] [PMC free article] [PubMed] [Google Scholar]



