Skip to main content
Applied Psychological Measurement logoLink to Applied Psychological Measurement
. 2018 Sep 18;43(5):402–414. doi: 10.1177/0146621618798664

Improved Wald Statistics for Item-Level Model Comparison in Diagnostic Classification Models

Yanlou Liu 1, Björn Andersson 2, Tao Xin 3,, Haiyan Zhang 4, Lingling Wang 3
PMCID: PMC6572908  PMID: 31235985

Abstract

Diagnostic classification models (DCMs) have been widely used in education, psychology, and many other disciplines. To select the most appropriate DCM for each item, the Wald test has been recommended. However, prior research has revealed that this test provides inflated Type I error rates. To address this problem, the authors propose to replace the asymptotic covariance matrix from the original version of the Wald statistic with a matrix obtained from improved computation methods. In this study, the Wald test based on the observed information matrix and the Wald test based on the sandwich-type matrix are proposed for item-level model comparisons and a simulation study is conducted to investigate their empirical behavior. Simulation results indicate that when the sample size is reasonably large (N1,000), the Type I error rates of the Wald test based on the sandwich-type matrix are accurate with adequate or excellent power under most of the simulation conditions.

Keywords: model comparison, diagnostic classification models, Wald test, asymptotic covariance matrix, likelihood ratio test

Introduction

Diagnostic classification models (DCMs) characterize relationships between responses of an individual obtained from questionnaire/test items and his or her latent attributes. DCMs have great promise to provide multidimensional discrete diagnostic information on a set of fine-grained attributes (such as cognitive processes, skills, abilities or knowledge representations) for instructional interventions (Rupp, Templin, & Henson, 2010). To date, DCMs have been widely used in educational, psychological, and psychiatric measurement, and also in many other disciplines (e.g., Greeno, 1980; Leighton & Gierl, 2007; Sorrel et al., 2016; Templin & Henson, 2006).

As a class of statistical models, a variety of DCMs have been developed (Rupp et al., 2010) based on different assumptions about how the probability of a correct response conditional on the Q-matrix and an examinee’s latent attribute mastery pattern is defined. The different DCMs reflect specific assumptions about how latent attributes influence test performance in a specific domain, where the Q-matrix specifies the relationship between items and latent attributes (Tatsuoka, 1990). For example, the deterministic inputs, noisy, “and” gate model (DINA; Haertel, 1989; Junker & Sijtsma, 2001) represents a noncompensatory assumption in that all required attributes are expected to have been mastered to obtain a correct response to an item. Hence, the absence of any of the required attributes cannot be compensated by the presence of other attributes. The deterministic inputs, noisy, “or” gate model (DINO; Templin & Henson, 2006) and the additive cognitive diagnostic model (A-CDM; de la Torre, 2011) are compensatory models, in which the absence of some attributes can be compensated by the presence of other attributes. The DINA, DINO, and A-CDM are specific DCMs, in other words these models can be shown to be constrained special cases of the general DCMs, or saturated DCMs, such as the general diagnostic model (von Davier, 2005, 2008), the log-linear cognitive diagnosis model (Henson, Templin, & Willse, 2009) and the generalized DINA model (GDINA; de la Torre, 2011).

As many DCMs are available, one critical step is to select the most appropriate DCM for response data at the item level. Compared with the saturated DCMs, the selected specific DCMs are more suitable for practical applications (de la Torre & Lee, 2013; Ma, Iaconangelo, & de la Torre, 2016; Rojas, de la Torre, & Olea, 2012; Sorrel, Abad, Olea, de la Torre, & Barrada, 2017; Sorrel, de la Torre, Abad, & Olea, 2017). In DCMs, given that the Q-matrix is correctly specified and that the saturated model provides the best model-data fit (Liu, Tian, & Xin, 2016), there are at least five different statistics to use for item-level model selection: the Wald test, the SX2 test, the likelihood ratio (LR) test, the Lagrange multiplier test, and the two-step likelihood ratio (2LR) test (Sorrel, Abad, et al., 2017; Sorrel, de la Torre, et al., 2017). De la Torre (2011) originally proposed that the Wald (Wd) test can be used to select the most interpretable reduced model from the saturated DCM. De la Torre and Lee (2013) investigated the empirical Type I error rate control and power of Wd by taking DINA, DINO, and A-CDM as examples. Furthermore, Ma et al. (2016) evaluated the performance of Wd in selecting the most appropriate model at the item level for additive models under the log and logit link functions (Hartz, 2002; Maris, 1999). Sorrel, Abad, et al. (2017) examined the performance of Wd, SX2, LR, and the Lagrange multiplier test and found that, compared with other statistics, only the empirical Type I error of SX2 was close to the nominal level. However, SX2 had low power to reject false models. Sorrel, de la Torre, et al. (2017) compared the performance of the 2LR, LR, and Wd in selecting the best reduced DCM. Although Wd has some advantages to other methods, such as ease of computation and slightly higher power, all of the previous studies consistently found that the Wd test that was computed using the item-wise information matrix proposed by de la Torre (2011) had inflated Type I error rates. This suggests that better computation methods of the covariance matrix or the item parameter standard errors (SEs) are needed (Ma et al., 2016) for the test to maintain the correct significance level.

Liu et al. (2016) have noted that since not only the item parameters but also the structural model parameters are estimated from the examinees’ response data (Rupp et al., 2010), both the item parameters and structural model parameters should be considered in the computation of the information matrix. Philipp, Strobl, de la Torre, and Zeileis (2018) have analytically shown that the incomplete information matrix (de la Torre, 2009) or the item-wise information matrix (de la Torre, 2011) lead to underestimated item SEs. Furthermore, in DCMs two new methods to directly estimate the item parameter covariance matrix, the inverse of the observed information matrix, and the sandwich-type estimator that take into account both the item and structural parameters have been proposed (Liu, Xin, Andersson, & Tian, 2018; Xin, Liu, & Tian, 2016). Liu et al. (2018) compared the performance of the two new methods, the empirical cross-product information matrix, and the two approaches proposed by de la Torre (2009, 2011) in providing consistent SEs of item parameter estimates and found that with correctly specified or slightly misspecified models the observed information matrix and the sandwich-type matrix had good performance. However, with substantial model misspecification only the SEs based on the sandwich-type covariance matrix were robust. Hence, in this study two new ways to calculate the Wald statistic for item-level model comparison will be introduced. That is, the Wald test based on the observed information matrix (WObs) or the Wald test based on the sandwich-type matrix (WSW).

The main interest of this study is to investigate the empirical behavior of the two newly proposed methods together with the Wd and LR statistics for item-level model comparisons under the GDINA framework via simulations. The remainder of this article is organized as follows: First, useful background on the saturated GDINA and its reduced models are briefly described, followed by an introduction of the statistical tests for item-level model comparisons that will be investigated in this study, and a discussion of several methods that can be used to obtain the asymptotic covariance matrix of the item parameter estimates. Then, results from a simulation study that systematically evaluates the performance of Wd, WObs, WSW, and LR were reported. The article concludes with discussion, concluding remarks, and some directions for further research.

The Saturated GDINA and the Reduced DCMs

The GDINA is a general framework and many commonly used DCMs such the DINA, DINO, or A-CDM can be viewed as special cases of the GDINA model (de la Torre, 2011). Suppose that K binary attributes are diagnosed in a DCM test with J dichotomous items. Define α=(α1,,αl,,αL) as the discrete attribute mastery profile matrix with a total number of possible attribute mastery patterns equal to L=2K. The structural parameter vector η=(η1,,ηl,,ηL-1) describes the probability p(αl|η)of a randomly selected examinee being a member of the lth attribute mastery profile. The number of freely estimated structural parameters is thus L1 (Liu et al., 2016; Rupp et al., 2010). A J×K binary Q-matrix specifies which attributes are measured by each item, for example, qj=(qj1,,qjk,,qjK) is the jth row of the Q-matrix. In the Q-matrix, if attribute k is required for solving item j, qjk=1, and qjk=0 otherwise. Let x=(x1,xi,xN) be the observed response matrix of N examinees, where xi=(xi1,,xiJ) is the ith examinee’s response vector.

In the saturated GDINA, the probability of a correct response on item j for examinee i with attribute mastery pattern αl=(αl1,,αlK) is given by de la Torre (2011).

Pj(αl)=λj,0+k=1Kλj,1,(k)αlkqjk+k=1K1k=k+1Kλj,2,(k,k)αlkαlkqjkqjk+, (1)

where λj,0 is the intercept parameter, λj,1,(k) is the main effect parameter associated with attribute k, and λj,2,(k,k) is the two-way interaction effect parameter associated with attribute k and attribute k. For example, if the first, the second, and the third attributes are required by item j, qj=(1,1,1) and the ith examinee has mastered all the required attributes by the item αi=(1,1,1), then according to Equation 1 the saturated GDINA can be expressed as follows,

Pj(αi)=λj,0+λj,1,(1)+λj,1,(2)+λj,1,(3)+λj,2,(1,2)+λj,2,(1,3)+λj,2,(2,3)+λj,3,(1,2,3). (2)

By setting some of the parameters of the saturated model to fixed values, the commonly used reduced DCMs can be obtained, see de la Torre (2011). In this example, the DINA model can be obtained from GDINA by setting the main and the lower order interaction effect terms to zero and thus the model becomes

Pj(αi)=λj,0+λj,3,(1,2,3). (3)

To obtain the A-CDM, all the interaction effect parameters in the saturated GDINA are fixed to zero and the item response function of the A-CDM can be written as

Pj(αi)=λj,0+λj,1,(1)+λj,1,(2)+λj,1,(3). (4)

The mathematical function for the DINO model is given by

Pj(αi)=λj,0+λj,1,(1), (5)

by setting λj,1,(1)=λj,1,(2)=λj,1,(3)=λj,2,(1,2)=λj,2,(1,3)=λj,2,(2,3)=λj,3,(1,2,3).

Information Matrix Estimation Procedures for DCMs

In DCMs, the item parameters λ and the structural parameters η are simultaneously estimated from the observed data, typically by using the MMLE/EM (marginal maximum likelihood estimation via expectation–maximization) algorithm. Denote γ^=(λ^,η^) as the MLEs of the model parameters that maximizes the log likelihood function of the observed response data (γ|x).

To directly estimate the asymptotic covariance matrix of the item parameter estimates in DCMs, there are at least five procedures available, namely, the incomplete information matrix in which only the item parameters are considered, the item-wise information matrix, the empirical cross-product information matrix, the observed information matrix, and the sandwich-type covariance matrix (Liu et al., 2018; Xin et al., 2016).

Philipp et al. (2018) have proved that the incomplete information matrix (de la Torre, 2009) and the item-wise information matrix (de la Torre, 2011) underestimates item SEs. They also studied the property of the empirical cross-product information matrix in providing consistent item parameter SEs, in which the item and structural parameters were simultaneously incorporated. The empirical cross-product information matrix is the cross-product of the derivative of the log likelihood of the observed data with respect to the model parameters,

IXPD=(γ|x)γ(γ|x)γ. (6)

They found that compared with the incomplete or the item-wise information matrix, the SEs based on the empirical cross-product information matrix showed smaller bias.

The observed information matrix and the sandwich-type estimator that included both the item and structural parameters for the estimation of the asymptotic covariance matrix were proposed by Liu et al. (2018). The observed information matrix is the negative of the second derivatives of the log likelihood of the observed data matrix with respect to the model parameters

IObs=2(γ|x)γγ, (7)

while the sandwich-type covariance matrix can be expressed as

ISW=IObs1IXPDIObs1. (8)

Their Monte Carlo results revealed that, under the assumption of a correctly specified DCM and a correctly specified Q-matrix, compared with the empirical cross-product, the incomplete, or the item-wise information matrix, the observed information matrix IObs and the sandwich-type matrix ISW had good performance in providing consistent item parameter SEs. Meanwhile, with model misspecification only the SEs based on the sandwich-type covariance matrix ISW were robust.

The Wald and LR Statistics

The Wald Test

By setting particular item parameters of the saturated GDINA to fixed values, the reduced DCMs can be obtained. Under the condition that the saturated GDINA provides a good fit to the observed data, the Wald test can be used to evaluate the statistical significance that a set of parameters are equal to some value (de la Torre & Lee, 2013). In DCMs, the original version of Wd is defined as (de la Torre, 2011),

Wd=(Rλ^j)(R^jR) 1(Rλ^j), (9)

where λ^j is the maximum likelihood estimate vector of item parameters for item j, R is a restriction matrix for the Wald test that can be used to impose a restriction on the item parameters of the saturated model and to obtain the corresponding item parameter covariance matrix ^j. Under the null hypothesis that the saturated DCM is equivalent to the specific DCM of interest, the Wald statistic asymptotically follows a chi-square distribution with degrees of freedom equal to the number of item parameters involved in the saturated DCM minus the number of item parameters in the reduced DCM.

Using the same example as above, here the authors also suppose that qj=(1,1,1) and αi=(1,1,1). In the DINA model, the null hypothesis is that the main and the lower order interaction effect parameters in the saturated GDINA are simultaneously equal to zero. Then, the corresponding restriction matrix R is

R=(010000000010000000010000000010000000010000000010). (10)

According to Equation 10, Rλ^j=(λ^j,1,(1),λ^j,1,(2),λ^j,1,(3),λ^j,2,(1,2),λ^j,2,(1,3),λ^j,2,(2,3)) is used to set the main and the lower order interaction effect parameters to zero. The restriction matrix R for the A-CDM to set all the interaction effect parameters to zero and to compute the corresponding covariance matrix is

R=(00001000000001000000001000000001). (11)

Note that for the A-CDM, Rλ^j=(λ^j,2,(1,2),λ^j,2,(1,3),λ^j,2,(2,3),λ^j,3,(1,2,3)). To impose restriction on the saturated GDINA to obtain the DINO model and to test the hypothesis that λ^j,1,(1)=λ^j,1,(2)=λ^j,1,(3)=λ^j,2,(1,2)=λ^j,2,(1,3)=λ^j,2,(2,3)=λ^j,3,(1,2,3), the following restriction matrix is defined,

R=(011000000011000001111000001111000110011001001101). (12)

Close inspection of Equation 9 reveals that the Wald statistic is dependent on the accuracy of the item parameter estimates λ^j and the corresponding asymptotic covariance matrix ^j. However, in the original implementation of the Wald statistic Wd for selecting the most suitable model, the item parameter covariance matrix was computed based on the item-wise information matrix, which leads to underestimated item parameter SEs (Philipp et al., 2018).

In this study, the authors propose that the Wald test for item-level model comparison is computed using IObs or ISW. As previous studies (Liu et al., 2018; Xin et al., 2016) have shown that the covariance matrices ^j of the item parameter estimates based on the inverse of IObs, or ISW are more accurate than those based on the empirical cross-product, the incomplete, or the item-wise information matrix, in this study, the Wald statistic in which the computation of the item parameter covariance matrix ^j is based on IObs or ISW, will be denoted by WObs or WSW. In theory, it is reasonable to expect that the performance of WObs or WSW would be better than the original version Wd.

The LR Test

In DCMs, another way to address the item-level model comparison question is through use of the LR test. The LR statistic is computed as

LR=2log(LrLs)=2(sr), (13)

where Lr is the likelihood for the reduced DCM and Ls is the likelihood for the saturated DCM. Suppose that the model parameter estimates γ^ lie in the interior of the parameter space. Then, under the null hypothesis, this statistic is also asymptotically chi-square distributed with degrees of freedom equal to the number of item parameters that are constrained in the reduced DCM at the item level. By comparing the log likelihoods of the saturated and the reduced models sr, the LR can be used to test whether the difference between these two models is statistically significant. The calculation of the LR is straightforward and does not require the estimation of the information matrix.

Sorrel, Abad, et al. (2017) found that the Type I error control of the LR test was more robust than Wd. Hence, it is of interest to empirically compare the behavior of the WObs, WSW, and LR under the same conditions.

Simulation Study

The primary interest of the current study was to examine the empirical performance of the WObs, WSW, Wd, or LR via Monte Carlo simulations under conditions varying by sample size, attribute correlation and fitted DCM. The purpose of the first part of the simulation was to investigate the Type I error rate control of these statistics. The binary item response data matrix was generated from DINA, A-CDM, or DINO separately. The second part of the simulation was conducted to illustrate the empirical power of these statistics when the fitted model was misspecified. The saturated GDINA model was used to generate the response data.

Three factors were manipulated in a 5×3×3 factorial design: (a) To systemically investigate the impact of sample size on the use of asymptotic sampling distributions, five sample sizes were considered, 500, 1,000, 2,000, 3,000, and 5,000. (b) To better reflect the reality of the correlations between attributes in applied research (Kunina-Habenicht, Rupp, & Wilhelm, 2012), the correlation coefficients between the attributes were ρ=0, ρ=.5, and ρ=.8. (c) The fitted DCMs were the DINA, A-CDM, and DINO models. This gave a total of 45 simulation conditions for each method. To further assess the effect of the manipulated between-subjects factors, the fitted DCM, sample size, and attribute correlation on the performance of the WObs, WSW, Wd, or LR method with respect to the Type I error, the marginal means of the Type I error rates for each methods were analyzed using ANOVAs, and computing the effect size, omega-squared index (ω2). Kirk (1996) offered the following guidelines for interpreting the magnitude of effect size ω2: small (0.010), medium (0.059), and large (0.138), consequently, only those factors whose ω2 index is greater than 0.138 were focused.

The true probabilities of a correct response for examinees who mastered none and all of the required attributes were fixed to .20 and .90. The main and/or interaction effect parameters were fixed to (.9.2)/(.9.2)njnj, where nj is the number of the main and/or interaction effect parameters for item j. The Q-matrix with five attributes and 30 items is provided in the online appendix Table A1. For each simulation condition, 1,000 replications were performed to provide stable results. The 95% confidence interval pepmical±1.96pepmical(1pepmical)/R of the empirical Type I error rate of these statistics which cover .05 is considered to be accurate (Bradley, 1978), pepmical is the empirical Type I error rate, R=1,000 is the number of replications in each condition. As with de la Torre and Lee (2013) a power higher than or equal to .80 is considered adequate, at least .9 is considered excellent. The code was written in R (R Core Team, 2017) where the R functions for model parameter estimation, Wald test calculation and item parameter covariance matrix estimation were modified from the R packages CDM (Robitzsch, Kiefer, George, & Uenlue, 2017) and dcminfo (Liu & Xin, 2017). The R codes in the current study are available upon request from the corresponding author.

Results

Type I Error Rate Control

The empirical Type I error rate control results of the Wd, WObs, WSW, and LR statistics across simulation conditions for DINA, A-CDM, and DINO at the 5% significance level are summarized in Tables 1 to 3, respectively. Results were averaged over items with the same number of required attributes (Kj). The 95% confidence interval of the empirical Type I error rate which do not cover .05 is marked in boldface.

Table 1.

Type I Error Rates for DINA (α=.5).

Wd
WObs
WSW
LR
ρ N Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3
0 500 .066 .110 .060 .096 .044 .070 .055 .063
1,000 .065 .072 .061 .068 .053 .058 .054 .053
2,000 .055 .065 .053 .062 .050 .058 .053 .053
3,000 .054 .055 .052 .054 .050 .051 .050 .049
5,000 .048 .051 .047 .050 .046 .049 .053 .049
.5 500 .082 .139 .071 .111 .046 .069 .061 .069
1,000 .064 .089 .060 .081 .051 .065 .057 .057
2,000 .054 .069 .052 .066 .047 .059 .053 .055
3,000 .056 .061 .053 .059 .051 .055 .047 .055
5,000 .054 .061 .053 .060 .051 .057 .050 .048
.8 500 .112 .131 .102 .121 .037 .050 .072 .097
1,000 .076 .135 .068 .111 .048 .074 .058 .070
2,000 .063 .090 .060 .084 .052 .069 .054 .055
3,000 .056 .072 .052 .067 .047 .060 .052 .057
5,000 .053 .064 .051 .061 .049 .057 .054 .051

Note. Wd is the Wald test based on the item-wise information matrix; WObs is Wald test based on the observed information matrix; WSW is Wald test based on the sandwich-type covariance matrix; LR is the likelihood ratio test; Kj is the number of required attributes by the jth item. The 95% confidence interval of the value which do not cover .05 is marked in boldface. DINA = deterministic inputs, noisy, “and” gate model; ρ = the attribute correlation; N = the sample size.

Table 3.

Type I Error Rates for DINO (α=.5).

Wd
WObs
WSW
LR
ρ N Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3
0 500 .079 .108 .071 .087 .044 .059 .061 .075
1,000 .062 .100 .058 .094 .050 .080 .058 .057
2,000 .057 .072 .055 .070 .051 .066 .052 .051
3,000 .054 .064 .052 .063 .049 .059 .053 .056
5,000 .053 .058 .052 .057 .051 .055 .048 .048
.5 500 .091 .071 .077 .063 .039 .035 .081 .083
1,000 .065 .128 .061 .115 .052 .094 .061 .057
2,000 .062 .090 .060 .084 .054 .076 .055 .054
3,000 .055 .072 .053 .071 .050 .066 .053 .051
5,000 .055 .064 .054 .063 .052 .060 .052 .057
.8 500 .091 .035 .089 .084 .029 .040 .227 .209
1,000 .093 .136 .081 .095 .055 .067 .195 .179
2,000 .072 .138 .069 .128 .060 .101 .146 .142
3,000 .065 .109 .062 .105 .057 .093 .119 .108
5,000 .057 .083 .056 .081 .053 .076 .080 .093

Note. Wd is the Wald test based on the item-wise information matrix; WObs is Wald test based on the observed information matrix; WSW is Wald test based on the sandwich-type covariance matrix; LR is the likelihood ratio test; Kj is the number of required attributes by the jth item. The 95% confidence interval of the value which do not cover .05 is marked in boldface. DINO = deterministic inputs, noisy, “or” gate model; ρ = the attribute correlation; N = the sample size.

As shown in Table 1, for the DINA model among the four statistics under investigation, the performance of the Wd or WObs to select the most appropriate DCM tended to inflate the Type I error rates, and although the WObs was slightly better than that of Wd, the WSW or LR test had much better Type I error rates under the same conditions. The empirical Type I error rates presented in Table 2 tell a different story. For the A-CDM, the performance of the WObs and WSW was better than Wd and LR in almost all cases. The Type I error rates of the LR were inflated under most conditions. The Type I error rates for the DINO model are reported in Table 3. It can be observed that only the WSW demonstrated reasonably robust Type I error behavior across all conditions. Under the ρ=.8 condition the LR rejected too many well-fitting items (as high as .227).

Table 2.

Type I Error Rates for A-CDM (α=.5).

Wd
WObs
WSW
LR
ρ N Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3
0 500 .068 .098 .057 .075 .034 .035 .094 .094
1,000 .056 .067 .051 .059 .040 .042 .083 .069
2,000 .053 .062 .050 .058 .047 .050 .075 .065
3,000 .052 .059 .050 .056 .047 .050 .074 .064
5,000 .052 .055 .051 .053 .048 .050 .067 .062
.5 500 .073 .116 .063 .090 .030 .036 .078 .085
1,000 .058 .077 .054 .068 .045 .047 .071 .072
2,000 .055 .060 .052 .056 .048 .047 .068 .064
3,000 .056 .060 .053 .056 .049 .051 .066 .061
5,000 .049 .057 .046 .055 .044 .051 .062 .059
.8 500 .087 .176 .083 .162 .018 .036 .084 .111
1,000 .065 .098 .056 .078 .038 .045 .066 .072
2,000 .062 .072 .056 .064 .050 .049 .062 .065
3,000 .059 .063 .054 .059 .050 .051 .058 .060
5,000 .056 .053 .051 .049 .049 .043 .059 .058

Note. Wd is the Wald test based on the item-wise information matrix; WObs is Wald test based on the observed information matrix; WSW is Wald test based on the sandwich-type covariance matrix; LR is the likelihood ratio test; Kj is the number of required attributes by the jth item. The 95% confidence interval of the value which do not cover .05 is marked in boldface. A-CDM = additive cognitive diagnostic model; ρ = the attribute correlation; N = the sample size.

ANOVAs were conducted to further assess the effect of the manipulated between-subjects factors on the performance of the WObs, WSW, Wd, or LR methods. For each method, the mean empirical Type I error was used as the ANOVA dependent variable. The ω^2 values forWd, WObs, WSW, and LR presented in Table 4 can be summarized as follows: (a) The Type I error of Wd was largely affected by sample size; (b) The main effects of sample size and attribute correlation for WObs were large; (c) The performance of WSW was substantially affected by the fitted DCM, and sample size; (d) The main effects of the fitted DCM and attribute correlation factors, and their interaction for LR were large. In considering the ω^2 values for Wd, WObs, WSW, and LR in Table 4 and the empirical Type I error rates in Tables 1 to 3, the authors can note that (a) as the sample size increased, the performance of all statistics were considerably better; (b) the attribute correlation strongly affected the performance of WObs and increasing the attribute correlation leads to inflated Type I error rates; (c) the Type I error rates of the WSW were well controlled for the DINA and A-CDM under most of the simulation conditions, but in general slightly inflated for the DINO model; and (d) for A-CDM, the Type I error of LR was closer to the nominal level as the attribute correlation increased, however, for the DINO model, the Type I error was inflated when the attribute correlation was high (ρ=.8).

Table 4.

ω^2 Values for Wd, WObs, WSW, and LR.

DCM N ρ DCM×N DCM×ρ N×ρ
Wd 0.015 0.467 0.125 0.092 −0.031 −0.041
WObs 0.066 0.429 0.179 0.058 −0.022 0.010
WSW 0.343 0.260 −0.002 0.117 0.004 0.075
LR 0.142 0.127 0.198 0.006 0.334 0.027

Note. Wd is the Wald test based on the item-wise information matrix; WObs is Wald test based on the observed information matrix; WSW is Wald test based on the sandwich-type covariance matrix; LR is the likelihood ratio test; DCM is the fitted diagnostic classification model; N = the sample size; ρ = the attribute correlation; The ω^2 value that is greater than 0.138 is marked in boldface.

Close inspection of the simulation results in Tables 1 to 4 reveals that, as expected, the Wald test Wd was improved by using the observed information matrix or the sandwich-type covariance matrix, and among the four tests, the newly proposed WSW had the best performance in providing reliable Type I error control across most of the simulation conditions.

Power

For the DINA and DINO models, the correct detection rates of the Wd, WObs, WSW, and LR tests were all above .831 under all of the simulation conditions, with one exception: The empirical power of WSW was .524 (Kj=2) and .729 (Kj=3) for DINA model, when ρ=.8 and N=500. Thus, only the correct detection rates for A-CDM are provided in Table 5. The correct detection rates of Wd, WObs, WSW, and LR in Table 5 are shown in strikeout font for conditions with inflated or deflated Type I errors. As shown in Table 5, for A-CDM, when N>500, among the four tests under investigation in this study, only the WSW had accurate Type I error rate control and adequate or excellent power (higher than .863) under all conditions with one exception (ρ=.8 and N=1,000). Sample size and attribute correlation influenced the power of WSW, the A-CDM tended to have less power to correctly reject the null hypothesis when the sample size was small (N=500), and attribute correlation was high (ρ=.8). This is not surprising, given that the previous study showed that the omission of interaction effect terms had little impact on classification accuracy (Kunina-Habenicht et al., 2012). Although the correct detection rates of the Wd, WObs, and LR tests are similar to the patterns of the power results of the WSW, it should be pointed out that due to the inaccurate Type I error rates in the simulation, these results are not necessarily indicative of good performance in real applications.

Table 5.

The Correct Detection Rates for A-CDM (α=.5).

Wd
WObs
WSW
LR
ρ N Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3 Kj=2 Kj=3
0 500 .699 .728 .673 .683 .582 .560 .714 .733
1,000 .928 .950 .928 .950 .916 .936 .942 .961
2,000 1.000 1.000 1.000 1.000 1.000 1.000 .999 1.000
3,000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
5,000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
.5 500 .624 .833 .592 .800 .456 .697 .634 .823
1,000 .884 .985 .877 .983 .863 .978 .893 .989
2,000 .994 1.000 .994 1.000 .994 1.000 .995 1.000
3,000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
5,000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000
.8 500 .495 .760 .453 .682 .179 .440 .469 .721
1,000 .754 .955 .731 .946 .667 .922 .747 .953
2,000 .958 .999 .955 .999 .951 .999 .965 1.000
3,000 .994 1.000 .994 1.000 .994 1.000 .995 1.000
5,000 1.000 1.000 1.000 1.000 1.000 1.000 1.000 1.000

Note. Wd is the Wald test based on the item-wise information matrix; WObs is Wald test based on the observed information matrix; WSW is Wald test based on the sandwich-type covariance matrix; LR is the likelihood ratio test; Kj is the number of required attributes by the jth item. The value that is smaller than .80 in each condition is in boldface. The condition with inflated or deflated Type I error is in strikeout font. A-CDM = additive cognitive diagnostic model; ρ = the attribute correlation; N = the sample size.

Discussion

Many DCMs with different assumptions or theories about how an examinee’s latent attributes influence test performance are now available in education, psychology, and many other disciplines. Selecting the most appropriate DCM for each item is crucial in theory and actual applications, since no particular DCM suits all the diagnostic test items in many, if not all, diagnostic assessments. For item-level model selection, a Wald test based on a simplified statistic Wd has been commonly used. However, according to several studies the Wd statistic yielded inflated Type I error rates (de la Torre & Lee, 2013; Ma et al., 2016; Sorrel, Abad, et al., 2017; Sorrel, de la Torre, et al., 2017).

As an alternative, the item parameter asymptotic covariance matrix ^j from the original version of the Wald statistic Wd was replaced with a matrix obtained from the observed information IObs or the sandwich-type matrix ISW to address the inflation of Type I error in this study, building on work by Liu et al. (2016), Xin et al. (2016), Philip et al. (2018), and Liu et al. (2018). Evidence from the Monte Carlo simulation suggests that when the sample size was reasonably large (N1,000) for all the three fitted models, the DINA, DINO, and A-CDM, the Type I error rate control of WSW was accurate or robust. When N1,000, for the DINA, DINO, or A-CDM, the power of WSW was adequate or excellent under all the simulation conditions with only one exception. Although the Type I error rate control of WObs was better than that of Wd, the WSW had the best performance under all simulation conditions. In this study, the LR statistic demonstrated good Type I error behavior under some conditions, however, its good behavior is not guaranteed. Similar to previous studies (Sorrel, Abad, et al., 2017; Sorrel, de la Torre, et al., 2017), the authors also found that the computation of the LR was time consuming. For example, with the DINO model the LR and Wald took approximately 1 min and less than 1 s to compute in one replication, respectively, when N=5,000 and ρ=.8. According to the simulation results in this study, WSW is the recommended tool for item-level model comparisons in empirical applications.

Although, as expected, WObs or WSW had Type I error rates close to the nominal level and had good power under most of simulation conditions, and the simulation results indicate that the WSW may serve as a practical tool for empirically selecting the most appropriate model at item level in DCMs, further research is needed. Many other factors that could potentially affect the behavior of these statistics were not manipulated in this study. For example, previous studies (Ma et al., 2016; Sorrel, Abad, et al., 2017; Sorrel, de la Torre, et al., 2017) have found that item quality, attribute distribution, the fitted model and test length affected the performance of Wd, and the authors believe it would be instructive to investigate how these factors affect the statistical properties of WObs or WSW.

Supplemental Material

Supplementary_Material_R2.2 – Supplemental material for Improved Wald Statistics for Item-Level Model Comparison in Diagnostic Classification Models

Supplemental material, Supplementary_Material_R2.2 for Improved Wald Statistics for Item-Level Model Comparison in Diagnostic Classification Models by Yanlou Liu, Björn Andersson, Tao Xin, Haiyan Zhang and Lingling Wang in Applied Psychological Measurement

Acknowledgments

The authors would like to thank Dr. Hua-Hua Chang, Dr. John Donoghue, and the two anonymous reviewers for their valuable comments on this manuscript.

Footnotes

Declaration of Conflicting Interests: The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Funding: The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The first author’s research was supported by a grant from Social Science Foundation of Shandong Province, China (Grant No. 18CJYJ16).

Supplemental Material: Supplemental material is available for this article online.

References

  1. Bradley J. (1978). Robustness? British Journal of Mathematical and Statistical Psychology, 31, 144-152. [Google Scholar]
  2. de la Torre J. (2009). DINA model and parameter estimation: A didactic. Journal of Educational and Behavioral Statistics, 34, 115-130. [Google Scholar]
  3. de la Torre J. (2011). The generalized DINA model framework. Psychometrika, 76, 179-199. [Google Scholar]
  4. de la Torre J., Lee Y.-S. (2013). Evaluating the Wald test for item-level comparison of saturated and reduced models in cognitive diagnosis. Journal of Educational Measurement, 50, 355-373. [Google Scholar]
  5. Greeno J. G. (1980). Trends in the theory of knowledge for problem solving. In Tuma D. T., Reif F. (Eds.), Problem solving and education: Issues in teaching and research (pp. 9-23). Hillsdale, NJ: Lawrence Erlbaum. [Google Scholar]
  6. Haertel E. H. (1989). Using restricted latent class models to map the skill structure of achievement items. Journal of Educational Measurement, 26, 301-323. [Google Scholar]
  7. Hartz S. M. (2002). A Bayesian framework for the unified model for assessing cognitive abilities: Blending theory with practicality (Unpublished doctoral dissertation), University of Illinois, Urbana–Champaign. [Google Scholar]
  8. Henson R. A., Templin J. L., Willse J. T. (2009). Defining a family of cognitive diagnosis models using log-linear models with latent variables. Psychometrika, 74, 191-210. [Google Scholar]
  9. Junker B. W., Sijtsma K. (2001). Cognitive assessment models with few assumptions, and connections with nonparametric item response theory. Applied Psychological Measurement, 25, 258-272. [Google Scholar]
  10. Kirk R. E. (1996). Practical significance: A concept whose time has come. Educational and Psychological Measurement, 56, 746-759. [Google Scholar]
  11. Kunina-Habenicht O., Rupp A. A., Wilhelm O. (2012). The impact of model misspecification on parameter estimation and item-fit assessment in log-linear diagnostic classification models. Journal of Educational Measurement, 49, 59-81. [Google Scholar]
  12. Leighton J., Gierl M. (2007). Cognitive diagnostic assessment for education: Theory and applications. Cambridge, UK: Cambridge University Press. [Google Scholar]
  13. Liu Y., Tian W., Xin T. (2016). An application of M2 statistic to evaluate the fit of cognitive diagnostic models. Journal of Educational and Behavioral Statistics, 41, 3-26. [Google Scholar]
  14. Liu Y., Xin T. (2017). dcminfo: Information matrix for diagnostic classification models (R package version 0.1.6). Retrieved from https://CRAN.R-project.org/package=dcminfo
  15. Liu Y., Xin T., Andersson B., Tian W. (2018). Information matrix estimation procedures for cognitive diagnostic models. British Journal of Mathematical and Statistical Psychology. doi: 10.1111/bmsp.12134 [DOI] [PubMed] [Google Scholar]
  16. Ma W., Iaconangelo C., de la Torre J. (2016). Model similarity, model selection, and attribute classification. Applied Psychological Measurement, 40, 200-217. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Maris E. (1999). Estimating multiple classification latent class models. Psychometrika, 64, 187-212. [Google Scholar]
  18. Philipp M., Strobl C., de la Torre J., Zeileis A. (2018). On the estimation of standard errors in cognitive diagnosis models. Journal of Educational and Behavioral Statistics, 43, 88-115. [Google Scholar]
  19. R Core Team. (2017). R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing; Retrieved from https://www.R-project.org/ [Google Scholar]
  20. Robitzsch A., Kiefer T., George A. C., Uenlue A. (2017). CDM: Cognitive diagnosis modeling (R package version 5.9-27). Retrieved from http://CRAN.R-project.org/package=CDM
  21. Rojas G., de la Torre J., Olea J. (2012, April). Choosing between general and specific cognitive diagnosis models when the sample size is small. Paper presented at the meeting of the National Council on Measurement in Education, Vancouver, British Columbia, Canada. [Google Scholar]
  22. Rupp A. A., Templin J., Henson R. A. (2010). Diagnostic measurement: Theory, methods, and applications. New York, NY: Guilford Press. [Google Scholar]
  23. Sorrel M. A., Abad F. J., Olea J., de la Torre J., Barrada J. R. (2017). Inferential item-fit evaluation in cognitive diagnosis modeling. Applied Psychological Measurement, 41, 614-631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Sorrel M. A., de la Torre J., Abad F. J., Olea J. (2017). Two-step likelihood ratio test for item-level model comparison in cognitive diagnosis models. Methodology: European Journal of Research Methods for the Behavioral and Social Sciences, 13, 39-47. [Google Scholar]
  25. Sorrel M. A., Olea J., Abad F. J., de la Torre J., Aguado D., Lievens F. (2016). Validity and reliability of situational judgement test scores: A new approach based on cognitive diagnosis models. Organizational Research Methods, 19, 506-532. [Google Scholar]
  26. Tatsuoka K. K. (1990). Toward an integration of item-response theory and cognitive error diagnosis. In Frederiksen N., Glaser R., Lesgold A., Safto M. (Eds.), Monitoring skills and knowledge acquisition (pp. 453-488). Hillsdale, NJ: Lawrence Erlbaum. [Google Scholar]
  27. Templin J. L., Henson R. A. (2006). Measurement of psychological disorders using cognitive diagnosis models. Psychological Methods, 11, 287-305. [DOI] [PubMed] [Google Scholar]
  28. von Davier M. (2005). A general diagnostic model applied to language testing data (ETS Research Report RR-05-16). Princeton, NJ: Educational Testing Service. [DOI] [PubMed] [Google Scholar]
  29. von Davier M. (2008). A general diagnostic model applied to language testing data. British Journal of Mathematical and Statistical Psychology, 61, 287-307. [DOI] [PubMed] [Google Scholar]
  30. Xin T., Liu Y., Tian W. (2016, April). Information matrix estimation procedures for cognitive diagnostic model. Paper presented at the meeting of the National Council on Measurement in Education, Washington, DC. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary_Material_R2.2 – Supplemental material for Improved Wald Statistics for Item-Level Model Comparison in Diagnostic Classification Models

Supplemental material, Supplementary_Material_R2.2 for Improved Wald Statistics for Item-Level Model Comparison in Diagnostic Classification Models by Yanlou Liu, Björn Andersson, Tao Xin, Haiyan Zhang and Lingling Wang in Applied Psychological Measurement


Articles from Applied Psychological Measurement are provided here courtesy of SAGE Publications

RESOURCES