Abstract
This study introduces two new statistics for measuring the score comparability of computerized adaptive tests (CATs) based on comparing conditional standard errors of measurement (CSEMs) for examinees that achieved the same scale scores. One statistic is designed to evaluate score comparability of alternate CAT forms for individual scale scores, while the other statistic is designed to evaluate the overall score comparability of alternate CAT forms. The effectiveness of the new statistics is illustrated using data from grade 3 through 8 reading and math CATs. Results suggest that both CATs demonstrated reasonably high levels of score comparability, that score comparability was less at very high or low scores where few students score, and that using random samples with fewer students per grade did not have a big impact on score comparability. Results also suggested that score comparability was sometimes higher when the bottom 20% of scorers were used to calculate overall score comparability compared to all students. Additional discussion related to applying the statistics in different contexts is provided.
Keywords: computerized adaptive testing, score comparability, conditional standard error of measurement, equity, vertical scales
An important consideration when developing multiple forms of a test is the score comparability of the forms. In contrast to fixed-form tests that often have a limited number of forms and a single base form, computerized adaptive tests (CATs) can have thousands of forms with each examinee typically seeing a unique form. Because items are dynamically selected based on performance, there is not a single base form to which all forms are equated, instead items are selected from a calibrated item pool and examinees often see harder or easier items based on their ability. The existence of a large number of forms and the fact that forms often differ in difficulty as a function of ability can make it hard to evaluate score comparability of alternate CAT forms. CATs with vertical scales present further challenges due to shifts in ability and test blueprints across grades. However, evaluating score comparability of CATs remains an important concern. Standard 5.16 from the Standards of Educational and Psychological Testing (AERA, NCME, & APA, 2014) emphasizes the need to provide documentation showing score comparability for CATs and multistage tests. Similarly, federal peer review guidelines and National Center on Intensive Intervention (NCII) and National Commission for Certifying Agency (NCCA) standards all look for evidence of score comparability for alternate forms.
A key question is how to evaluate the score comparability of alternate forms. A few strategies have been suggested in the literature (see Kolen, 1999; Wang & Shin, 2010). One set of methods looks at observed scores on different forms and evaluates the similarity of the full distribution of observed scores or specific moments of these distributions. For example, if test forms were available in paper and computer-based versions, one may compare the means for paper and computer-based tests to see if they were comparable. Many score comparability studies of this type exist (see Wang et al., 2007, 2008; Zeng et al., 2015). In the context of evaluating score comparability of alternate CAT forms, comparing observed score distributions or moments of these distributions is challenging because few examinees typically take each form and performance on alternate forms is expected to differ as a function of ability. Therefore, methods based on observed scores are not a practical way to evaluate score comparability of alternate CAT forms.
A second set of methods uses item level statistics, such as differential item functioning (DIF) statistics, item parameter drift statistics, distributions of students choosing each answer choice, or item p-values to evaluate score comparability. For example, Keng et al. (2008) used DIF analyses, tests of item p-values, and response distributions for each item to evaluate comparability of online and paper versions of a statewide assessment. Similarly, Kim and Huynh (2010) used differential item and bundle analyses to evaluate the score comparability of computer and online versions of an English test. While item level statistics are useful for evaluating score comparability when one wants to compare how items function across different groups or over time, these statistics are not designed to answer the question of whether scores on alternate forms that are composed of different items are comparable, as is common in CATs.
A third set of methods involves performing simulation studies to look at score comparability. In these contexts, one often looks to see that similar scores are obtained on different forms and that low bias and root mean square error are found in different conditions. Several studies of this type also appear in the literature (see Davey & Thomas, 1996; Eignor et al., 1993; Harris et al., 2021; Thompson & Way, 2007; Wang & Kolen, 2001). For example, Thompson and Way (2007) simulated three CAT designs and compared them to a paper version of a statewide assessment to see to what extent score comparability was achieved. Harris et al. (2021) used a simulation study to investigate comparability of CAT scores when content specifications were or were not maintained. Simulations are very useful for evaluating score comparability of CATs as they allow one to test how the CAT functions and compare scores produced in different conditions. However, simulations only provide insight into the conditions included in the simulation design. These conditions often come close to matching situations in operational CATs, but they may not fully represent real CATs or all operationally administered forms. The fact that simulations may not fully represent real CATs or all administered forms means that a simulation study may not give precise estimates of score comparability observed in practice. Simulations also do not provide a way to estimate score comparability for observed operational CAT data.
A fourth set of methods is based on Lord’s (1980) equity property, where distributions of observed scores for examinees with the same ability are compared. Lord (1980) showed that in the strictest sense it was impossible for these conditional distributions to be equal unless the alternate forms were strictly parallel or perfectly reliable. Neither situation happens in practice, so weaker forms of Lord’s equity property are often used instead (see Kolen & Tong, 2005; Wyse & Reckase, 2011). Two weaker forms of equity are first-order equity and second-order equity. First-order equity requires that the means are equivalent conditional on ability (Divgi, 1981), while second-order equity requires that the conditional standard errors of measurement are equivalent conditional on ability (Morris, 1982). First and second-order equity are commonly used to evaluate score comparability of fixed-form tests when the tests are scored using item response theory (IRT) models and there are a limited number of forms. Part of the reason for their widespread use in these cases is because simple graphs and indices related to first-order and second-order equity can be created. Figure 1 shows three graphs often used in these contexts. The top left panel shows a plot of test characteristic curves (TCCs), the top right panel shows a plot of test information functions (TIFs), and the bottom left panel shows a plot of conditional standard error of measurement (CSEM) curves. The goal is that the curves for the new and base form are as close as possible. In fact, van der Linden (2005) outlined several automated test assembly algorithms that minimize differences between these curves when generating alternate forms. One can look at the deviations between the base and new forms based on these curves to evaluate score comparability. For example, Tong and Kolen (2005) proposed two indices, including one based on first-order equity and one based on second-order equity, and applied these indices to compare equipercentile, IRT observed score, and IRT true score equating methods. Similarly, Wyse and Reckase (2011) suggested a first-order equity index based on comparing TCCs and used this method to compare six IRT equating methods for the Multistate Bar Exam. Simple comparisons of curves for CATs are difficult because there are thousands of forms, item selection for each form is tailored based on ability, and there is not a single base form to which all forms are equated. These features of CATs imply that it is often impossible to compare each form to a single base form, that the desire is not to compare CAT forms over the full range of ability, and that curves should be expected to shift as a function of ability.
Figure 1.
Test characteristic curves, test information functions, conditional standard error of measurement (CSEM) curves for a new and base form.
The purpose of this article is to introduce two new statistics to evaluate the score comparability of CATs. In the next section of this article, we introduce the new statistics and provide a rationale for using these statistics to evaluate score comparability of operational CATs. The new statistics are then demonstrated using grade 3–8 reading and math CAT data. The study concludes with additional discussion and considerations related to applying the new statistics.
New Statistics to Evaluate Score Comparability of CATs
The new statistics compare the conditional standard errors of measurement (CSEMs) for test takers that achieved the same scale scores to evaluate score comparability of CATs. First, we propose calculating the median CSEM for all examinees that achieved the same scale and then looking at what percentage of the CSEMs for these examinees are within some number of scale score points of the median CSEM. We call this number of scale score points the distance from the median CSEM. Algebraically, this statistic can be defined as
| (1) |
where is the number of examinees with CSEMs that fall within the distance of the median CSEM for scale score i, and is the number of examinees that achieved scale score i. The percent within median CSEM (PWMC) statistic gives a measure of score comparability for each scale score. One can graph the values of Equation (1) to evaluate score comparability across the range of scale scores. Equation (1) ranges from 0% to 100% with values close to 0% observed when very few CSEM values are within the distance of the median CSEM value and values close to 100% observed when nearly all CSEM values fall within the distance of the median CSEM value. PWMC values close to 100% are desirable and indicate greater score comparability.
While Equation (1) is useful for providing an individual level index of score comparability, it does not provide an overall measure of score comparability for a CAT. This can be found by aggregating Equation (1) across all scale scores. Algebraically, this overall percent within median CSEM (OPWMC) statistic can be defined as
| (2) |
where is the minimum possible scale score, is the maximum possible scale score, is the number of examinees with CSEMs that fall within the distance of the median CSEM for scale score i, and is the number of examinees that achieved scale score i. Like the individual level PWMC statistic, the OPWMC statistic in Equation (2) ranges from 0% to 100%. OPWMC values close to 100% are desirable and signal CATs with greater score comparability.
A key part of using Equations (1) and (2) to evaluate score comparability of CATs is defining the distance from the median CSEM values. In the context of evaluating equating results, Dorans and Feigenbaum (1994) introduced the concept of the difference that matters (DTM) and defined the DTM as the difference in scores that would practically impact the score reported to a test taker. The DTM changes as a function of the scale used to report scores. For some tests, the DTM might be .5 or 1 scale score points, while for other tests the DTM might be 5 or 10 scale score points. We propose using the DTM as the primary method for selecting the distance from the median CSEM values. The distance can be made larger than the DTM to help researchers evaluate the closeness of the observed CSEMs to the median CSEM.
There are several other aspects of using Equations (1) and (2) that warrant further discussion. First, one may wonder why the PWMC and OPWMC statistics are based on CSEMs instead of TIFs or TCCs. The rationale for this choice stems from common goals of CATs. One common goal of CATs is to be equally precise for examinees with the same ability (Weiss, 1982; Weiss & McBride, 1984). Another common goal of CATs is to stop them when the CSEM is small enough to yield a classification decision with an acceptable level of accuracy (Thompson, 2009). This type of termination rule implies that examinees with the same scale score should have similar CSEMs to yield similarly high classification accuracy. Equations (1) and (2) relate to these common CAT design goals and Lord’s second-order equity property because they evaluate score comparability by looking at how similar CSEMs are for examinees with the same ability.
One may also wonder why the statistics are based on the median CSEM instead of the mean CSEM or taking the difference between the maximum and minimum CSEM. We suggest using the median CSEM to reduce the impact of outliers. The median CSEM is less influenced by outliers than the mean CSEM or taking the difference between maximum and minimum CSEM, especially in cases where they might be only a few observations at each scale score. Such situations are possible in CATs that span hundreds of points and are vertically scaled.
Another important aspect of using Equations (1) and (2) is grouping examinees with the same scale score together when computing the statistics. Each examinee’s scale score is an imperfect proxy of their ability. However, in the absence of knowing true ability, examinees with the same scale scores are the most like each other of any examinees in the observed data and practically they are often viewed as being equivalent in ability. In CATs, there tends to be greater measurement precision at each scale score than in fixed-form tests because item selection is tailored to each examinee. This fact provides further justification for viewing these examinees as similar in ability and grouping them together to evaluate score comparability of CATs.
Data and Methods
Data for this study come from grade 3 to 8 reading and math CATs. These 34-item CATs are used to measure student progress and growth throughout the United States. Each CAT utilizes the Rasch (1960) model and reports scores on a vertical scale ranging from 600 to 1400 in 1-point increments. The reading CATs consist of vocabulary-in-context and multiple-choice items associated with reading passages, and items cover five content domains. The math CATs consist of multiple-choice items with items covering six content domains. Both the math and reading CATs utilize a 67% correct item selection rule, item exposure control based on the randomesque technique (Kingsbury & Zara, 1989), and grade-specific content constraints where each item is selected from the content domain with the largest difference from its target number in the test blueprint (see Kingsbury & Zara, 1989, 1991). The CATs use maximum a posteriori (MAP) estimation to estimate ability until a student gets at least one item correct and one item incorrect. The MAP estimator assumes a normal prior with a mean equal to a grade-specific starting ability and a standard deviation of 2. Maximum likelihood estimation (MLE) is used to estimate ability once at least one item is correct and one item is incorrect. CSEMs are calculated by multiplying the slope of the linear equation used to create scale scores by 1 over the square root of the Fisher test information function. Each CAT item pool includes over 5000 items. We use all grade 3 to 8 reading and math CAT data from one school year in our analyses.
Table 1 presents summary statistics for the reading and math CATs, including the N counts, mean, median, and standard deviation of the scale scores and CSEMs. For each test, the mean and median of the scale scores increased with grade, while the standard deviations of the scale scores increased with grade for the math CATs but did not increase with grade for the reading CATs. For the reading CATs, the mean CSEMs was 16 in grade 3 and 17 in all other grades, the median CSEMs were 16 in every grade, and the standard deviations of the CSEMS ranged from 1.3 to 1.7. For the math CATs, the mean and median CSEMs were 17 in every grade and the standard deviations ranged from 1.4 to 1.9. There were several million test takers that took each test in each grade with higher sample sizes observed for the reading CATs.
Table 1.
Descriptive Statistics for Math and Reading CATs.
| Scale score | CSEM | |||||||
|---|---|---|---|---|---|---|---|---|
| Subject | Grade | N | Mean | Median | SD | Mean | Median | SD |
| Reading | 3 | 4864484 | 952 | 959 | 74 | 16 | 16 | 1.3 |
| 4 | 4409706 | 991 | 999 | 71 | 16 | 16 | 1.3 | |
| 5 | 4041638 | 1021 | 1028 | 71 | 16 | 16 | 1.3 | |
| 6 | 2977353 | 1044 | 1051 | 72 | 16 | 16 | 1.4 | |
| 7 | 2380728 | 1062 | 1069 | 74 | 16 | 16 | 1.5 | |
| 8 | 2219531 | 1079 | 1086 | 74 | 17 | 16 | 1.7 | |
| All grades | 20893440 | 1013 | 1019 | 84 | 16 | 16 | 1.4 | |
| Math | 3 | 2172997 | 948 | 953 | 64 | 17 | 17 | 1.4 |
| 4 | 2043227 | 995 | 1001 | 67 | 17 | 17 | 1.5 | |
| 5 | 1900277 | 1033 | 1038 | 70 | 17 | 17 | 1.6 | |
| 6 | 1522242 | 1056 | 1066 | 71 | 17 | 17 | 1.8 | |
| 7 | 1296745 | 1073 | 1084 | 74 | 17 | 17 | 1.9 | |
| 8 | 1260012 | 1087 | 1100 | 74 | 17 | 17 | 1.9 | |
| All grades | 10195500 | 1022 | 1026 | 84 | 17 | 17 | 1.6 | |
Note. CSEM is the conditional standard of error of measurement and SD is standard deviation.
Since each CAT reports scores on a scale from 600 to 1400 in 1-point increments and scale scores are created by linearly transforming ability estimates and then using truncation rounding to keep the whole number part of the transformed ability estimates, the DTM for these CATs is 1-point and we compute the PWMC and OPWMC statistics using a distance of 1. However, we also calculate the statistics using distances of 2 and 5 to investigate how changing the distance from the median CSEM may impact the statistics. The value of 2 is twice the DTM and larger than the standard deviation of the CSEMs in any grade, which was greater than 1 but less than 2 in all cases. The value of 5 is five times the DTM and slightly greater than the average monthly scale score growth across grades for both CATs, which was 4.42 for reading and 4.24 for the math. In practice, these two CATs have been submitted for reviews by the National Center for Intensive Intervention (NCII), which asks for evidence of score comparability for all test takers and the bottom 20% of scorers. Therefore, results are reported for all test takers as well as the bottom 20% of scorers. Further, many testing programs assess smaller samples than the math and reading CATs. To investigate the impact of using smaller samples, analyses are run with random samples of 5000, 25,000, and 100,000 students per grade. These sample sizes mimic small, moderate, and large samples that one might observe in K-12 state testing programs. For the random samples, we select 1000 samples for each sample size. The random samples provide real data-based simulations for two CATs and three sample sizes. We calculate the mean, minimum, maximum, and standard deviation over the samples to summarize performance of the new statistics. Throughout the analyses, the goal is that the PWMC and OPWMC statistics are high and close to 100%, signaling that the CAT forms have high degrees of score comparability.
Results
Reading CATs
Figure 2 shows the median CSEMs, the PWMC statistic with distances of 1, 2, and 5 points, and the distribution of scale scores for the grade 3 to 8 reading CATs. The median CSEM and the PWMC curves across grades were smoothed using a cubic spline function to reduce jaggedness in the lines and to show overall trends in estimates. The median CSEMs were very similar and below 20 for the scale scores between 700 and 1250 where most examinees score. The PWMC statistics with a distance of 1 tended to be over 80% in this score range, with a slight dip for scores between 1000 and 1200. The PWMC statistic, as expected, increased as the distance used with the statistic increased and was nearly 100% with a distance of 5 between scores of 700 and 1250. Some lower PWMC statistics were observed with scale scores below 700 and above 1250 for all three distances. Scale scores in these ranges are quite rare. It makes sense that PWMC statistics would be lower in these ranges since very few grade 3 through 8 items have difficulties in this range and as a result items are less well matched to ability when students receive very high or low scores. Observing less well targeted items at the extremes of the ability scale is common in K-12 tests like this reading CAT (see Wyse & McBride, 2021).
Figure 2.
Median CSEM, percent within median CSEM (PWMC), and frequency distribution of scale scores for all students taking the Grade 3 through 8 reading CATs.
Figure 3 shows the OPWMC statistic for each grade for all students, the bottom 20% of scorers, and the random samples of 5000, 25,000, and 100,000 students. The OPWMC statistic with distance of 1 was in the low to mid-80s for all students with higher values for the bottom 20% of scorers. The figure makes it clear that using smaller sample sizes did not have much impact on results as the mean OPWMC across samples for different sample sizes were within 1%–2% of the full sample value. The OPWMC statistic with a distance of 2 was in the low to mid-90s and the OPWMC statistic with distance of 5 was nearly 100% in all conditions.
Figure 3.
Overall percent within median CSEM (OPWMC) statistics for the grade 3 through 8 reading CATs.
Figures and tables with summaries of performance for the PWMC and OPWMC for the random samples are presented in the online supplement in Appendix A. Figure A.1 provides the mean PWMC statistic over replications for the three sample sizes. The figure shows lines with similar shapes to Figure 2 for scores of 750–1200 with more pronounced dips between scores of 1000–1200 for larger sample sizes. Below scores of 750, the mean PWMC statistics was higher than in Figure 2, while PWMC statistics above 1200 could not be estimated in many cases because very few students were selected with those scores and statistics for those scores do not consistently appear across replications. Tables A.1 to A.3 provide the OPWMC statistics and illustrate variation over samples. Higher standard deviations over samples were found with smaller sample sizes. The tables also show that the mean value of the OPWMC statistic dropped slightly with distances of 1 and 2 as sample size increased but were generally close to each other.
Math CATs
Figure 4 shows the median CSEMs, the PWMC statistics, and the distribution of scale scores for the grade 3 to 8 math CATs. There are some similarities and differences compared to the reading results. The median CSEMs were again below 20 for the scale scores between 700 and 1250 where most examinees score. Some lower PWMC statistics were observed with scale scores below 700 and above 1300. As expected, the PWMC statistics were again higher with a distance of 5 compared to distances of 1 and 2. Compared to the reading CATs, the PWMC curves had different shapes and often had lower values. The PWMC statistic with a distance of 1 showed the most different values and was in the mid-70s between scores of 800–1200.
Figure 4.
Median CSEM, percent within median CSEM (PWMC), and frequency distribution of scale scores for all students taking the Grade 3 through 8 Math CATs.
Figure 5 shows the OPWMC statistic for each grade for the math CATs. The OPWMC statistic with a distance of 1 was in the high-70s and low 80s in all conditions. Somewhat different than reading, the OPWMC statistic was slightly higher in grade 3 for the bottom 20% of scorers but was close to the values for all students at the other grades. Like was observed with reading, using smaller sample sizes did not have much impact on the results as the mean OPWMC statistics across samples for different sample sizes were again within 1–2% of the full sample value. The OPWMC statistic with a distance of 2 was in the high-80s and low-90s and the OPWMC statistic with a distance of 5 was again nearly 100% in all conditions.
Figure 5.
Overall percent within median CSEM (OPWMC) statistics for the grade 3 through 8 math CATs.
Figures and tables with summaries of performance for the PWMC and OPWMC for the random samples for math CAT are presented in the online supplement in Appendix B. Figure B.1 provides the mean PWMC statistics and shows lines with similar shapes to Figure 4 for scores below 1200. High scores again presented some challenges and statistics could not be estimated in many cases because students with those scores do not consistently appear in the random samples. Tables B.1 to B.3 present the OPWMC statistics and again show some variation over samples. Higher standard deviations were again found with smaller sample sizes. Compared to reading, the standard deviations were higher in all conditions. These results are somewhat expected since there were higher standard deviations for the CSEMs for math than reading in Table 1. The tables show that the mean value of the OPWMC statistic rose slightly with distances of 1 and 2 as sample size increased but were generally close to each other.
Discussion
Despite calls in the Standards to provide documentation of score comparability for CATs, a widely accepted method has not been presented in the research literature. This study introduces two new statistics for these purposes based on examining the percent of CSEM values that fall within some distance of the median CSEM observed at each scale score. Two real data examples and simulation studies were provided to show the utility of the new statistics for evaluating the score comparability of CATs. The examples illustrated that the reading CATs had greater score comparability than the math CATs, that the PWMC and OPWMC statistics were higher when larger distances were used, that score comparability was often less at very high or low scale scores where few students score, and that using random samples with fewer students per grade did not have a large impact other than sometimes creating situations where the PWMC statistic could not be estimated for high scores. In addition, score comparability was sometimes higher when the bottom 20% of scorers were used to calculate the OPWMC statistic.
While the new statistics appeared to work well for the math and reading CATs, different results may be found in other situations. The math and reading CATs had large item banks and moderate test lengths. It is possible that different results may be found if the size of the item bank or the length of the test was changed. Different results may also be found if the score distributions or sample sizes are different from the math and reading CATs. Other testing programs may test fewer examinees than the smallest sample size of 5000 examinees per grade and subject in our examples. Using smaller samples of students than in our examples could impact how the statistics perform. For other tests, the reporting scale will differ from the math and reading CATs, which may impact the chosen distance and the PWMC and OPWMC statistics. Other CATs may also use different constraints, exposure controls, scoring procedures, and termination rules, which also may impact results. Other CATs may also apply different IRT models. Additional research should explore how some of these factors impact the statistics.
Another factor that may require some further investigation is how the level of adaptivity and variability of CSEMs impacts the statistics. In our analyses, we found that the math CATs tended to have lower PWMC and OPWMC statistics than the reading CATs. We also found that math CATs had more variability in CSEM values. In addition, PWMC statistics tended to be lower at the top and bottom of the scales. Both the lower statistical values for the math CATs and for scores at the top and bottom of the scale seem to be related to cases where selected items where less well matched to the ability level of the student and the CAT algorithm appeared to be less adaptive. These results make sense because CSEMs tend to be higher and more variable when items are less well targeted to ability. However, additional research could be done to more directly evaluate how the level of adaptivity impacts the values of the statistics.
An important practical consideration when using the statistics relates to selecting the distance from the median CSEM. In our examples, we explored distances of 1, 2, and 5, which corresponded to the DTM, twice the DTM, and five times the DTM. As expected, when the distance from the median CSEM was increased, the PWMC and OPWMC statistics increased. The OPWMC statistic with a distance of 1 was in the mid-70s to mid-80s, while the OPWMC statistic with a distance of 2 was in the high-80s to low-90s, and the OPWMC statistic with a distance of 5 was nearly 100% in all cases. Practically speaking, setting the distance equal to the DTM seems reasonable as this requires that the change in CSEM is below the value that would change the reported score. Depending on the test and the rounding rule used to create scale scores, the DTM may be as low as .5 or 1.0 points but could be higher if scores are reported in 5- or 10-point increments. One could choose a distance other than the DTM depending on the uses of the test. For example, if the test is used for classification purposes and the CAT is stopped using a CSEM rule, one may set the distance from the median CSEM by looking at the change in CSEM that would change classification accuracy by some percentage.
When the PWMC and OPWMC statistics were introduced, it was noted that the ideal value for them was a value as close to 100% as possible. However, in practice a range of values may be observed. For the math and reading CATs, the OPWMC statistic with a distance of 1 and the PWMC statistic with a distance of 1 in the region where most examinees score were in the mid-70s to mid-80s. Are these acceptable levels of score comparability? Some guidance to answer this question may come from research literature on reliability. In these contexts, a value of .70 has been suggested as acceptable with higher values being desirable (Nunnally & Bernstein, 1994). We suggest using 70% (i.e., .70 * 100) as a starting threshold for the new statistics since these statistics look at how close CSEMs are to a measure of central tendency. Looking at how close CSEMs are to a measure of central tendency has some parallels to looking at consistency of ratings for a group of raters when computing interrater reliability. Of course, just like in reliability analyses, using thresholds above .70 might make sense depending on the stakes and uses of the test. The fact that the OPWMC and PWMC statistics in our examples often exceeded 70% makes sense. When the item bank is large and test length is moderate, a good match between the selected items and the test taker’s ability and similar CSEMs should be expected for most test takers. Less similar CSEMs would be expected when selected items are less well matched to ability. We observed some of these situations with the PWMC statistics when students scored at the low or high ends of the scale. Even though 70% seems like a reasonable threshold for indicating acceptable levels of score comparability for the PWMC and OPWMC statistics with a distance equal to the DTM, additional research is needed to evaluate how this threshold works in other situations.
The PWMC and OPWMC statistics have several advantages over previously suggested statistics. First, the PWMC and OPWMC statistics can be used in tandem to look at score comparability both at an individual and aggregate level. Most previous suggested statistics are not designed for use at both an individual and aggregate level. Second, the PWMC and OPWMC statistics do not require a single base form to which all other forms are compared. Comparing new forms to a single base is the typical strategy in most existing indices. Instead, PWMC and OPWMC evaluate score comparability by looking at CSEMs for examinees with the same score. Comparing CSEMs for examinees with the same score has direct links to second-order equity and how CATs are commonly designed. It also allows for forms to differ in difficulty as a function of ability. Finally, the new statistics report results in an intuitive metric, the percentage of CSEMs within a certain distance of the median CSEM. Previously suggested indices do not report score comparability in a percentage scale with a defined range. For example, the indices suggested by Tong and Kolen (2005) have a minimum of 0, but no defined maximum.
It is also important to point out that the PWMC and OPWMC statistics are designed to evaluate score comparability of CATs in a statistical sense by looking at the values of CSEMs for different examinees with the same scale score. They are not designed to capture other form differences that may lead one to question the comparability of test forms. In particular, K-12 CATs with vertical scales often allow items to be selected above and below a student’s grade as long as the items fall within the content domains in the test blueprint. The PWMC and OPWMC statistics do not account for content differences or the number out of grade items a student may have seen. The statistics also do not account for different ways students may progress through a CAT or the difficulty of the items that the student saw when testing. If one views these factors as important aspects of score comparability, additional methods would be needed to capture them.
The new statistics proposed in this article provide simple and intuitive indices that can be used to present evidence of score comparability for CATs and multistage tests in response to requests for such evidence in the Standards, federal peer review, NCII, and NCCA guidelines. The new statistics have both practical and theoretical appeal. The new statistics are directly linked to what it means for tests to be parallel, how CATs are often designed, and are easy to apply in a variety of contexts. In addition, the new statistics can be used to evaluate comparability at individual and aggregate levels and are presented in a percentage-based metric that is easy to understand and interpret. When utilized by looking at the distance within the DTM, the new statistics express the percentage of CSEMs that are sufficiently close to each other that the values reported to test takers would be the same. This type of logic is easy to communicate to a variety of stakeholders, including those without advanced training in statistics.
Supplemental Material
Supplemental Material for Two Statistics for Measuring the Score Comparability of Computerized Adaptive Tests by Adam E. Wyse in Applied Psychological Measurement
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding: The author(s) received no financial support for the research, authorship, and/or publication of this article.
Author Note: Any opinions, findings, conclusions, or recommendations expressed in this manuscript are those of the author and are not necessarily the official position of Renaissance.
Supplemental Material: Supplemental material for this article is available online.
ORCID iD
Adam E. Wyse https://orcid.org/0000-0002-1719-9461
References
- American Educational Research Association, American Psychological Association, and National Council for Measurement in Education (2014). Standards for educational and psychological testing. American Educational Research Association. [Google Scholar]
- Davey T., Thomas L. (1996). April 8–12). Constructing adaptive tests to parallel conventional programs. American Educational Research Association Annual Meeting. [Google Scholar]
- Divgi D. R. (1981). April 13 – 17). Two procedures for scaling and equating tests with item response theory. American Educational Research Association Annual Meeting. [Google Scholar]
- Dorans N. J., Feigenbaum M. D. (1994). Equating issues engendered by changes to the SAT and PSAT/NMSQT. In Lawrence I. M., Dorans N. J., Feigenbaum M. D., Feryok N., Schmitt A. P., Wright N. K. (Eds.), Technical issues related to the introduction of the new SAT and PSAT/NMSQT (ETS RM-94-10). Educational Testing Service. [Google Scholar]
- Eignor D. R., Stocking M. L., Way W. D., Steffen M. (1993). Case studies in computer adaptive test design through simulation (ETS-RR-93-56). Educational Testing Service. [Google Scholar]
- Harris D. J., Fang Y., Li D. (2021). Examining comparability across CAT assessments. Educational Measurement: Issues and Practice, 40(4), 18–20. 10.1111/emip.12473 [DOI] [Google Scholar]
- Keng L., McClarty K. L., Davis L. L. (2008). Item-level comparative analysis of online and paper administrations of the Texas assessment of knowledge and skills. Applied Measurement in Education, 21(3), 207–226. 10.1080/08957340802161774 [DOI] [Google Scholar]
- Kim D. H., Huynh H. (2010). Equivalence of paper-and-pencil and online administration modes of the statewide English test for students with and without disabilities. Educational Assessment, 15(2), 107–121. 10.1080/10627197.2010.491066 [DOI] [Google Scholar]
- Kingsbury C. G., Zara A. R. (1991). A comparison of procedures for content-sensitive item selection in computerized adaptive tests. Applied Measurement in Education, 4(3), 241–261. 10.1207/s15324818ame0403_4 [DOI] [Google Scholar]
- Kingsbury G. G., Zara A. R. (1989). Procedures for selecting items for computerized adaptive tests. Applied Measurement in Education, 2(4), 359–375. 10.1207/s15324818ame0204_6 [DOI] [Google Scholar]
- Kolen M. J. (1999). Threats to score comparability with applications to performance assessments and computerized adaptive tests. Educational Assessment, 6(2), 73–96. 10.1207/S15326977EA0602_01 [DOI] [Google Scholar]
- Lord F. M. (1980). Applications of item response theory to practical testing problems. Erlbaum. [Google Scholar]
- Morris C. N. (1982). On the foundations of test equating. In Holland P. W., Rubin D. B. (Eds.), Test equating (pp. 169–191). Academic Press. [Google Scholar]
- Nunnally J. C., Bernstein I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill. [Google Scholar]
- Rasch G. (1960). Probabilistic models for some intelligence and attainment tests. Denmark Paedogogiske Institut. [Google Scholar]
- Thompson N. A. (2009). Item selection in computerized classification testing. Educational and Psychological Measurement, 69(5), 778–793. 10.1177/0013164408324460 [DOI] [Google Scholar]
- Thompson T., Way D. (2007, June 7–8). Investigating CAT designs to achieve comparability with a paper test. GMAC® conference on computerized adaptive testing. [Google Scholar]
- Tong Y., Kolen M. J. (2005). Assessing equating results on different equating criteria. Applied Psychological Measurement, 29(6), 418–432. 10.1177/0146621606280071 [DOI] [Google Scholar]
- van der Linden W. J. (2005). Linear models for optimal test design. Springer-Verlag. [Google Scholar]
- Wang H., Shin C. D. (2010). Comparability of computerized adaptive and paper-pencil tests. Test, Measurement, and Research Services Bulletin, 13, 1–7. Pearson. [Google Scholar]
- Wang S., Jiao H., Young M. J., Brooks T., Olson J. (2007). A meta-analysis of testing mode effects in grade K-12 mathematics tests. Educational and Psychological Measurement, 67(2), 219–238. 10.1177/0013164406288166 [DOI] [Google Scholar]
- Wang S., Jiao H., Young M. J., Brooks T., Olson J. (2008). Comparability of computer-based and paper-and-pencil testing in K-12 reading assessments: A meta-analysis of testing mode effects. Educational and Psychological Measurement, 68(1), 5–24. 10.1177/0013164407305592 [DOI] [Google Scholar]
- Wang T., Kolen M. J. (2001). Evaluating comparability in computerized adaptive testing: Issues, criteria, and an example. Journal of Educational Measurement, 38(1), 19–49. 10.1111/j.1745-3984.2001.tb01115.x [DOI] [Google Scholar]
- Weiss D. J. (1982). Improving measurement quality and efficiency with adaptive testing. Applied Psychological Measurement, 6(4), 473–492. 10.1177/014662168200600408 [DOI] [Google Scholar]
- Weiss D. J., McBride J. R. (1984). Bias and information of Bayesian adaptive testing. Applied Psychological Measurement, 8(3), 273–285. 10.1177/014662168400800303 [DOI] [Google Scholar]
- Wyse A. E., McBride J. R. (2021). A framework for measuring the amount of adaptation of Rasch-based computerized adaptive tests. Journal of Educational Measurement, 58(1), 83–103. 10.1111/jedm.12267 [DOI] [Google Scholar]
- Wyse A. E., Reckase M. D. (2011). A graphical approach to evaluating equating using test characteristic curves. Applied Psychological Measurement, 35(3), 217–234. 10.1177/0146621610377082 [DOI] [Google Scholar]
- Zeng J., Yin P., Shedden K. A. (2015). Does matching quality matter in mode comparison studies. Educational and Psychological Measurement, 75(6), 1045–1062. 10.1177/0013164414565006 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplemental Material for Two Statistics for Measuring the Score Comparability of Computerized Adaptive Tests by Adam E. Wyse in Applied Psychological Measurement





