Skip to main content
Applied Psychological Measurement logoLink to Applied Psychological Measurement
. 2025 Jun 12;49(7):367–378. doi: 10.1177/01466216251350342

Increase of Uncertainty in Summed-Score-Based Scoring in Non-Rasch IRT

Eisuke Segawa 1,
PMCID: PMC12162545  PMID: 40520435

Abstract

Summed-score (SS)-based scoring in non-Rasch IRT allows for pencil-and-paper administration and is used in the Patient-Reported Outcomes Measurement Information System (PROMIS) alongside response-pattern-based scoring. However, this convenience comes with an increase in uncertainty (the increase) associated with SS scoring. The increase can be quantified through the relationship between Bayesian SS and RP scoring. Given an SS of s, the SS posterior is a weighted sum of RP posteriors, with weights representing the marginal probabilities of RPs. From this mixture, the SS score (SS posterior mean) is a weighted sum of RP posterior means, and its uncertainty (variance of the SS posterior) is decomposed into the uncertainty of RP scoring (the weighted sum of RP posterior variances) and the increase (variance of RP posterior means). Without quantifying the increase, PROMIS recommends RP scoring for greater accuracy, suggesting SS scoring as a second option. Using variance decomposition, we quantified the increases for two short forms (SFs). In one, the increase is very small, making SS scoring as accurate as RP scoring, while in the other, the increase is large, indicating SS scoring may not be a viable second option. The increase varies widely, influencing scoring decisions, and should be reported for each SF when SS scoring is used.

Keywords: summed score, item response theory, patient-reported outcomes

Introduction

Summed-score-based (SS) scoring is an attractive administrative method in Health-Related Patient-Reported Outcome (HR-PRO) studies within the context of non-Rasch item response theory (IRT) because it allows for paper-and-pencil administration, which is a widely used approach in clinical settings. Recognizing this importance, the Patient-Reported Outcomes Measurement Information System (PROMIS), which offers over 200 HR-PRO short forms (SFs) (126, 27, 16, and 39 SFs for Adult, Pediatrics, Early Childhood Parent-Report [ages 1–5], and Parent Proxy for Pediatric Patients [ages 5–17], respectively, as of 5/16/2024), has integrated SS scoring along with response-pattern-based (RP) scoring (Cella et al., 2007, 2010; “HealthMeasures,” 2024). Three administrative methods are available in PROMIS: Computer Adaptive Testing with RP scoring, fixed-length testing with RP scoring, and fixed-length testing with SS scoring. Earlier research and this manuscript compare the first two methods and the latter two methods, respectively (Segawa et al., 2020).

Users of SS scoring experience greater uncertainty in scores compared to RP scoring, as SS scoring only uses partial information from response patterns (RPs). Prior studies have shown that this increase in uncertainty (the increase) is small—no more than 10% of the standard deviation (SD)—as illustrated through tables and figures in several works (Steinberg & Thissen, 2013; Thissen et al., 1995, 2001; Thissen and Orlando 2001).

Based on these studies, PROMIS guidance recommends RP scoring as the first option due to its higher accuracy, with SS scoring as a good second option (Appendix A. All appendices are available as online supplements). Following these studies, we employ Bayesian scoring, where posterior means and variances represent scores and their associated uncertainties (Ayala, 2009). The increase is calculated by subtracting RP posterior variance from the corresponding SS posterior variance. We use posterior variance instead of posterior standard deviation (SD) due to the mathematical relationship between the two variances, which does not hold for SDs. The percent increase in SD is roughly half of that in variance (Appendix B).

We computed the increase of 34 PROMIS SFs using a different method than what is presented in this paper. Item parameters of 34 PROMIS SFs were available to us. Most were publicly available, with a few sourced from a previous project (Segawa et al., 2020). The PROMIS guidance may mislead users for certain SFs. The increases in 6 SFs were very small (less than 2% in SD), while one SF exhibited a large increase (more than 20% in SD). For SFs with very small (and negligible) increases, framing SS scoring as “a good second option” might lead increase-sensitive users to prioritize accuracy over convenience, potentially causing unnecessary shifts from paper-and-pencil scoring to avoid the negligible increase. For the SF with a large increase, presenting SS scoring as “a good second option” may mislead users into believing their increases are small, even though they are not.

SFs with neither very small nor large increases must be identified earlier. Previous studies examined only a few SFs, and their findings were applied solely to those SFs. However, PROMIS generalized these findings to more than 200 SFs, which we see as a generalization fallacy. To avoid this, the guidance should include the increase for each SF. However, the previous approach cannot accurately measure this increase because it obtained the maximum uncertainty difference between SS and RP scoring but not the average difference. To address this limitation, we propose a new method. We calculate the increase by considering the nested structure of the data: RPs are nested within SSs, and SSs are nested within SFs. Since scoring decisions occur at the SF-level, the increase must be aggregated accordingly (SF-level increase). In the first nesting, we apply a mathematical relationship: SS posterior is a weighted sum of RP posteriors with weights being the marginal probabilities of RPs (SS posterior is a mixture of RP posteriors, proof is presented later).

SF-level RP posterior variance is computed in two steps. First, SS-level RP posterior variances are obtained for all SSs by computing weighted sums of RP posterior variances, with weights being the marginal probabilities of RPs (RP WEIGHTS). Second, SF-level RP posterior variance is obtained by computing a weighted sum of SS-level RP posterior variances. The weights (SS WEIGHTS) reflect the expected distribution of latent symptoms among patients in the target population. SF-level SS posterior variance is a weighted sum of SS posterior variances using SS WEIGHTS. Finally, the SF-level increase is computed by subtracting SF-level RP posterior variance from SF-level SS posterior variance.

To better contextualize the evaluation of the increase, we introduce the increase in posterior uncertainty due to shortening SFs, a well-known consideration in HR-PRO research. Shortening is particularly relevant because reducing patient burden is a key objective in SF development. PROMIS provides multiple versions of SFs—8-item, 6-item, and 4-item—allowing users to balance measurement precision and respondent burden. Segawa et al. (2020) examined the relationship between SF length and the increase in uncertainty, highlighting its practical implications. For example, if shortening an 8-item SF by one and two items results in average increases of 10% and 20%, respectively, PROMIS users can better interpret the magnitude of the increase. A 5% increase in SS scoring would be considered minor, as it is half the effect of a 1-item reduction, whereas a 40% increase would be substantial, doubling the impact of a 2-item reduction.

In the following sections, we describe SS and RP scoring, establish the mathematical relationship between SS and RP posteriors, validate this relationship, present our approach to measuring the increase, and examine the variation in RP scores that induces the increase in SS scoring. These analyses are conducted using two SFs, one with a very small increase (one of the 34 SFs with very small increase) and another with a large increase (one with the highest increase among the 34 SFs). To ease the interpretation of the increase, we reference the increase in uncertainty associated with shortening SFs.

Bayesian RP Scoring and SS Scoring

Let there be a total of i=1,,I ordinal items. Let Ti(kθ) be the i -th item’s trace line for category k=1,,K . Due to the assumption of independence of item responses, conditional on the latent θ , the likelihood for response pattern y=(y1,y2,,yI) is given by:

L(yθ)=i=1ITi(yiθ). (1)

Denoting Ys={yi=1Iyi=s} as the set of RPs that yield a sum score of s , the likelihood of sum score s=1,2,,S is:

L(sθ)=yYsL(yθ)=yYsi=1ITi(yiθ). (2)

Given a prior distribution g(θ) , the posterior distribution is:

p(θs)=L(sθ)g(θ)p(s) (3)

where p(s) is the marginal probability:

p(s)=L(sθ)g(θ)dθ (4)

Therefore, the posterior mean and posterior variance are:

E(θs)=1p(s)θL(sθ)g(θ)dθ (5)

and

V(θs)=E(θ2s)E2(θs)=1p(s)θ2L(sθ)g(θ)dθE2(θs) (6)

The posterior mean and variance are a point estimate for θ (score) and a score uncertainty, respectively. The RP posterior p(θy) , marginal probability p(y) , posterior mean E(θy) , and posterior variance V(θy) are obtained by replacing the SS likelihood in equation (2) with the RP likelihood in equation (1).

Calculating the SS likelihood involves evaluating all KI possible RP likelihoods, which becomes computationally intractable for large I . To address this, Lord and Wingersky (1984) proposed a recursive algorithm for computing the SS likelihoods. Subsequent work by Thissen et al. (1995) and Cai (2015) extended the algorithm to accommodate ordinal items and multidimensional IRT, enabling practical computation for large I .

1PL Model in SS Scoring

This section demonstrates that SS and RP scoring are identical under the 1PL model. To illustrate this, we demonstrate that the likelihoods of SS and RP scoring are equivalent up to a constant factor, ensuring identical posterior distributions under Bayesian scoring. Consider the likelihood of a response pattern y* with a summed score s , denoted as L(y*θ) . Due to the sufficiency property of the 1PL model—where all information about a person’s ability is fully encapsulated in their SS—the likelihood of any response pattern y with the same SS is identical up to a constant:

L(y*θ)=aL(yθ) (7)

for all y*,yYs (Appendix C). Applying equation (7) to the SS likelihood in equation (2), we obtain:

L(sθ)=yYsL(yθ)=i=1maiL(y*θ)=(i=1mai)L(y*θ) (8)

where m is the number of response patterns that yield s , and ai is a constant specific to each pattern. Equation (8) shows that SS likelihood is equivalent to RP likelihood up to a constant i=1mai . Consequently, when Bayesian scoring is applied with identical priors (e.g., standard normal), the resulting posterior distributions—and thus the scores—are the same for both SS and RP scoring.

SS Posterior as a Weighted Sum of RP Posteriors

By substituting equation (2) into equation (3) and using L(yθ)g(θ)=p(θy)p(y) , the SS posterior can be expressed as a weighted sum of RP posteriors for SSs equal to s :

p(θs)=yYsL(yθ)g(θ)p(s)=yYsp(θy)p(y)p(s)=yYsp(θy)p(y)p(s). (9)

Since p(s) , the marginal probability of SS being s , is the sum of marginal probabilities of RPs y that yield s :

p(s)=L(sθ)g(θ)dθ=yYsL(yθ)g(θ)dθ=yYsL(yθ)g(θ)dθ=yYsp(y),

Equation (9) shows that the SS posterior is a weighted sum of RP posteriors with RP WEIGHTS. Using the properties of the mean and variance of mixture, the mean and variance of the SS posterior are:

E(θs)=Ey(E(θy)) (10)
V(θs)=Ey(V(θy))+Vy(E(θy)) (11)

where the outer expectation and variance are on the probability space of RPs y (Frühwirth-Schnatter, 2006). Equation (10) indicates that the SS posterior mean is a weighted sum of RP posterior means, and equation (11) indicates that the variance of the SS posterior is composed of two parts: the expected (SS-level) RP posterior variance and the variance of RP posterior means. Details of the derivations of equations (10) and (11) are found in Appendix D.

The Increase of Uncertainty in SS Scoring

The first and second terms in equation (11) are the SS-level RP posterior variance and the increase, respectively. Weighted sums of SS-level RP posterior variances and SS posterior variances, using SS WEIGHTS, provide SF-level RP and SS posterior variances, respectively. The increase is the difference between SF-level SS and RP posterior variances.

Methods

Following PROMIS, the latent variable θ is re-scaled by multiplying by 10 and adding by 50. We present methods for four analyses: verification of the above mathematical relations numerically (VERIFICATION analysis), an example analysis of our approach to examining the increase in SS scoring (INCREASE analysis), an examination of variation of RP scores within SS which induces the increase (variance of RP scores analysis), and comparisons of the increase of uncertainties in SS scoring with shortened SF (SHORTENED SF analysis). The analyses were conducted using the R software and R-packages of mirt, tidyverse, and kableExtra (Chalmers, 2012; R Core Team, 2021; Wickham et al., 2019; Zhu, 2024).

Instruments

Samejima’s graded model was used in all analyses (Samejima, 1969). For the VERIFICATION analysis, we use the Science SF, which includes four 4-category ordinal items, and Pediatric Life Satisfaction 8b, which consists of eight 5-category items (Ayala, 2009; Forrest et al., 2018). For the INCREASE analysis and variance of posterior scores analysis, we used Pediatric Life Satisfaction 8b, which exhibits a large increase, and the Adult Anxiety 8a, which has a very small increase (Forrest et al., 2018; Pilkonis et al., 2011). For the SHORTENED SF analysis, we used the PROMIS Adult Profile, which measures seven domains of basic health-related quality of life using 7 SFs. For each domain, we used two versions: 8-item and 6-item version. The 6-item version was created by excluding the last two items from the 8-item version. In addition, we create a 7-item version, specific to this analysis, by excluding the last item from the 8-item version. Thus, three versions—8-item, 7-item, and 6-item—were used to measure each domain. Each SF consisted of 5-category ordinal items. The item parameters for the SCIENCE, Anxiety 8a, and Pediatric Life Satisfaction 8b are publicly available (Ayala, 2009; Forrest et al., 2018; Pilkonis et al., 2011).

Procedures

VERIFICATION Analysis

Using the Science SF, we plotted the SS posterior when SS equals 5, using both the original definition in equation (3) and the relation in equation (9). There are four RPs: [1112],[1121],[1211], and [2111] , which yield a score of 5. Using PROMIS Pediatric Life Satisfaction 8b, we verified equations (10) and (11).

We employed a simulation-based approach instead of full enumeration for verification. Initially, we considered full enumeration to be computationally demanding and its implementation complex. While full enumeration is feasible for short forms with a small number of items (e.g., 58 = 390,625 possible RPs), it does not scale to longer PROMIS scales with 30 items, where the number of possible RPs grows exponentially. In contrast, our simulation-based approach remains practical for larger questionnaires. Importantly, our approach successfully verified equations (10) and (11), as demonstrated in the results section, making full enumeration unnecessary for our purposes.

Ten million latent values were sampled from the standard normal distribution. These latent values were converted into response patterns using known item parameters, and their RP posterior means and variances were computed, which were then grouped by SSs. The weighted sums of RP posterior means by SS served as approximations of SS posterior means (equation (10)). The sum of the weighted RP posterior variances (approximating SS-level RP posterior variance) and the weighted sum of squared differences between RP posterior means and SS posterior mean (approximating the increase in SS scoring) gave an approximation of the SS posterior variance (equation (11)). Alternatively, we directly computed SS posterior means and variances using the recursive method (Lord & Wingersky, 1984). We compute the mean and maximum of the absolute errors. This verification of equations (10) and (11) also serves as an assessment of the accuracy of the computations used in later analyses.

INCREASE Analysis

We computed SS posterior variances using the recursive method and SS-level RP posterior variances for Anxiety 8a, just as we did for Life Satisfaction 8b in the second part of the VERIFICATION analysis. We then plotted SS posterior variances and SS-level RP posterior variances for both SFs, producing posterior variance curves.

SF-level statistics are computed as weighted sums of SS-level counterparts, with weights determined by the proportions of SSs in the simulated data. However, when latent values are sampled from a standard normal distribution (mean = 0), as in the VERIFICATION analysis, healthier SSs tend to be oversampled. This occurs because PROMIS SFs are designed so that the mean of the general population, which is healthier than the target population, aligns with the mean of the standard normal distribution. As a result, using a standard normal distribution oversamples the healthiest SS (all item responses are the healthiest) and undersamples the unhealthiest SS (all item responses are the least healthy). While both extreme SSs yield the increase of 0, the sum of their proportions is greater when the healthiest SS is oversampled than when the two extremes are more balanced. Consequently, using a standard normal distribution skews the weighting in favor of SS scoring by increasing the proportion of SSs with the increase of 0.

To mitigate this issue, we shift the mean of the standard normal distribution by m to reduce the overrepresentation of the healthiest SS. Specifically, we compute the proportions of the smallest and largest SSs using one million latent values sampled from a normal distribution with its mean shifted by 1 in the unhealthy direction. We then iteratively adjust the mean in increments of 0.1 in the direction that decreases the sum of extreme SS proportions. This process continues until we determine the optimal m , which minimizes the total proportion of extreme SSs. Using the resulting SS proportions as weights, we compute SF-level SS posterior variance, RP posterior variance, the increase, and the percent increase.

Variance of RP Scores Analysis

We examined the variation of RP scores, which drives the increase in SS scoring, by plotting RP and SS scores for Anxiety 8a and Life Satisfaction 8b (equation (11)). To prevent overcrowding due to the large number of RPs ( 58 ), we selected one SS value with a relatively small number of RPs for each SF. Further, we randomly selected one-tenth of the RP scores to ensure clarity.

SHORTENED SF Analysis

Two factors influence the increases in 1- and 2-item shortenings. To ensure reliable reference values, we controlled for these factors. The first is which items within an SF are excluded, and the second is the selection of SFs. To control the first factor, as described in the Instrument section, we excluded the last two items (present in the 8-item but not the 6-item version) for the 2-item shortening. The last item was excluded for the 1-item shortening. To control the second factor, we could have computed weighted averages of the increases in 1- and 2-item shortenings across all PROMIS SFs, using SF usage frequencies as weights. However, due to the unavailability of item parameters and usage frequencies, we instead used the average across the seven SFs in the PROMIS Profile, which spans a broad range of HR-PRO domains and is frequently used—a reasonable representative of PROMIS SFs.

We randomly sampled one million latent values from a normal distribution with a mean m (determined as described in the INCREASE analysis) and a standard deviation of 1. Using these samples, we computed the increases in SS scoring across the 7 domains as in the INCREASE analysis. Since SSs for the 8-item SF are not available for the 6- and 7-item SFs, we computed the SF-level RP posterior variances for the 7- and 6-item SFs by averaging their respective RP posterior variances. Using the SF-level posterior variances for the 8-, 7-, and 6-item SFs, we calculated the increases in posterior variances for the 7- and 6-item SFs and then computed the percent increases.

Results

VERIFICATION Analysis

We confirmed equation (9) by showing that the weighted mean of the four RP posteriors perfectly overlapped with the directly computed SS posterior (Figure 1). Equations (10) and (11) were also validated, as the approximated posterior means and variances closely matched those obtained through the recursive method. The means and maximums of absolute errors were 0.011% and 0.01% for the SS posterior means and 0.14% and 0.67% for the SS posterior variances.

Figure 1.

Figure 1.

SS posterior when summed score is 5, four weighted RP posteriors which yield 5, and a weighted sum of the four RP posteriors for Science SF.

INCREASE Analysis

Excluding the endpoints where SS and RP scoring are identical and SS posterior variances (circles) completely overlap with SS-level RP posterior variances (Xs), we observed that SS posterior variances (circles) were slightly larger in Anxiety 8a and clearly larger in Life Satisfaction 8b compared to their corresponding SS-level RP posterior variances (Xs) (Figure 2). Consequently, the SF-level increases were 2.55% for Anxiety 8a and 39.56% for Life Satisfaction 8b.

Figure 2.

Figure 2.

Comparison of SS posterior variances and SS-level RP posterior variances in Anxiety 8a and Life Satisfaction 8b.

Variance of RP Scores Analysis

SS and SS-level RP posterior variances for the selected SS values (12 and 35 for Anxiety 8a and Life Satisfaction 8b, respectively) were highlighted with arrows in Figure 2. The rectangle in Life Satisfaction 8b represents the smallest region containing all RP scores, and the same-sized rectangle was applied to Anxiety 8a. These rectangles are expanded in Figure 3, showing 58,868 and 54,441 RP scores for Anxiety 8a and Life Satisfaction 8b, respectively.

Figure 3.

Figure 3.

Individual RP points for Anxiety 8a and Life Satisfaction 8b.

RP-level increases, defined as (SS posterior variance) - (RP posterior variance), varied across RPs for both SFs. In some cases, RP scores were less precise (i.e., had negative increases) compared to SS scores. The proportion of such cases was larger for Anxiety 8a than for Life Satisfaction 8b. Additionally, the variance of RP posterior means, corresponding to the increase (equation (11)), was much smaller in Anxiety 8a than in Life Satisfaction 8b.

SHORTENED SF Analysis

The averages of 1- and 2-item shortening (reference points) were 11.87% and 25.29%. The percent increase for Anxiety 8a (2.55%) was less than a quarter of the mean increase for 1-item shortening, while the percent increase for Life Satisfaction 8b (39.56%) was more than 1.5 times the average percent increase for 2-item shortening.

Conclusions and Discussions

We developed a method that accurately measures the increase by incorporating the mathematical relationship that SS posterior is the mixture of RP posteriors. Using this method, we demonstrated that the increase can be either inconsequentially small or large, findings not observed in previous studies. Because the current guidance might mislead users when SFs exhibit such increases, we recommend including the increase of each SF in the guidance so that users can make informed decisions when choosing a scoring method.

Distribution of the Increases

The next step would involve computing the increases for more than 200 PROMIS SFs using our method to obtain a distribution of these increases, including the proportions of SFs with inconsequential and large increases. This would reveal the average increase as well as the proportions of those extreme increases in SFs within PROMIS. That knowledge would contribute to our understanding of SS scoring in PROMIS and help in writing appropriate scoring guidance.

Cause of the Increase

Understanding the cause of the increase is important, as it may provide insights into how to mitigate it. Our preliminary analysis indicated that a larger variance in slope parameters was associated with a larger increase. When constructing SFs, higher-slope items are generally selected for their superior discrimination, while lower-slope items are sometimes included for substantive reasons. This suggests that incorporating lower-slope items not only reduces overall discrimination but also contributes to the increase. While selecting lower-slope items may decrease the variance in slopes and potentially reduce the increase, it would come at the cost of lower discrimination, which we do not recommend.

The Increase in Target Population

In this manuscript, we specified SS WEIGHTS to reflect the target population of each SF. However, a user could obtain a more accurate expected increase by specifying SS WEIGHTS to reflect her study. For example, suppose she uses Anxiety 8a for screening anxiety and expects about 30% of samples to have the minimum score, at which the SS-level increase is 0. The larger weight of 30% on 0 and smaller weights on others would even lower the SF-level increase.

The Increase in SD

Using the formula that converts the increase in variance to one in SD, the increases in standard deviation in Anxiety 8a and Life Satisfaction 8b are 1.27% and 18.13%, respectively (Appendix B). Although the variance should be used in the computation stage to reflect the mathematical relationship between SS and RP posterior variances as found in this manuscript, the standard deviation might be more appropriate for improving interpretation in the guidance.

Supplemental Material

Supplemental Material - Increase of Uncertainty in Summed-Score-Based Scoring in Non-Rasch IRT

Supplemental Material for Increase of Uncertainty in Summed-Score-Based Scoring in Non-Rasch IRT by Eisuke Segawa in Applied Psychological Measurement.

Funding: The author received no financial support for the research, authorship, and/or publication of this article.

The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Supplemental Material: Supplemental material for this article is available online.

ORCID iD

Eisuke Segawa https://orcid.org/0000-0003-1103-9448

References

  1. Ayala R. J. D. (2009). The theory and practice of item response theory. The Guilford Press. [Google Scholar]
  2. Cai L. (2015). Lord–Wingersky algorithm version 2.0 for hierarchical item factor models with applications in test scoring, scale alignment, and model fit testing. Psychometrika, 80(2), 535–559. 10.1007/s11336-014-9411-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Cella D., Riley W., Stone A., Rothrock N., Reeve B., Yount S., Amtmann D., Bode R., Buysse D., Choi S., Cook K., Devellis R., DeWalt D., Fries J. F., Gershon R., Hahn E. A., Lai J. S., Pilkonis P., Revicki D., PROMIS Cooperative Group . (2010). The patient-reported outcomes measurement information system (PROMIS) developed and tested its first wave of adult self-reported health outcome item banks: 2005–2008. Journal of Clinical Epidemiology, 63(11), 1179–1194. 10.1016/j.jclinepi.2010.04.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Cella D., Yount S., Rothrock N., Gershon R., Cook K., Reeve B., Ader D., Fries J. F., Bruce B., Rose M., PROMIS Cooperative Group . (2007). The patient-reported outcomes measurement information system (PROMIS): Progress of an NIH roadmap cooperative group during its first two years. Medical Care, 45(5 Suppl 1), S3–S11. 10.1097/01.mlr.0000258615.42478.55 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Chalmers R. P. (2012). Mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. 10.18637/jss.v048.i06 [DOI] [Google Scholar]
  6. Forrest C. B., Devine J., Bevans K. B., Becker B. D., Carle A. C., Teneralli R., Moon J. H., Tucker C. A., Ravens-Sieberer U. (2018). Development and psychometric evaluation of the PROMIS pediatric life satisfaction item banks, child-report and parent-proxy editions. Quality of Life Research: An International Journal of Quality of Life Aspects of Treatment, Care and Rehabilitation, 27(1), 217–234. 10.1007/s11136-017-1681-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Frühwirth-Schnatter S. (2006). Finite mixture and markov switching models. Springer. [Google Scholar]
  8. HealthMeasures . (2024). Northwetern university. https://www.healthmeasures.net/index.php [Google Scholar]
  9. Lord F. M., Wingersky M. S. (1984). “Comparison of IRT true-score and equipercentile observed-score” equatings. Applied Psychological Measurement, 8(4), 453–461. 10.1177/014662168400800409 [DOI] [Google Scholar]
  10. Pilkonis P. A., Choi S. W., Reise S. P., Stover A. M., Riley W. T., Cella D., PROMIS Cooperative Group . (2011). Item banks for measuring emotional distress from the patient-reported outcomes measurement information system (PROMIS): Depression, anxiety, and anger. Assessment, 18(3), 263–283. 10.1177/1073191111411667 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. R Core Team . (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/ [Google Scholar]
  12. Samejima F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika, 34(S1), 1-97. [Google Scholar]
  13. Segawa E., Benjamin S., Cella D. (2020). A comparison of computer adaptive tests (CATs) and short forms in terms of accuracy and number of items administrated using PROMIS profile. Quality of Life Research, 29(1), 213–221. 10.1007/s11136-019-02312-8 [DOI] [PubMed] [Google Scholar]
  14. Steinberg L., Thissen D. (2013). Item response theory. In Comer J. S., Kendall P. C. (Eds.), The oxford handbook of research strategies for clinical psychology. Oxford University Press. [Google Scholar]
  15. Thissen D., Maria O. (2001). Item response theory for items scored in two categories. In Test scoring (pp. 85–152). Routledge. [Google Scholar]
  16. Thissen D., Nelson L., Rosa K., McLeod L. D. (2001). Item response theory for items scored in more than two categories. In Test scoring (pp. 153–198). Routledge. [Google Scholar]
  17. Thissen D., Pommerich M., Billeaud K., Williams V. S. L. (1995). Item response theory for scores on tests including polytomous items with ordered responses. Applied Psychological Measurement, 19(1), 39–49. 10.1177/014662169501900105 [DOI] [Google Scholar]
  18. Wickham H., Averick M., Bryan J., Chang W., McGowan L. D. ’A., Francois R., Garrett G., Hayes A., Henry L., Hester J., Kuhn M., Pedersen T., Miller E., Bache S., Müller K., Ooms J., Robinson D., Seidel D., Spinu V., Yutani H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. 10.21105/joss.01686 [DOI] [Google Scholar]
  19. Zhu H. (2024). Create Awesome LaTeX Table with knitr:: kable and kableExtra.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material - Increase of Uncertainty in Summed-Score-Based Scoring in Non-Rasch IRT

Supplemental Material for Increase of Uncertainty in Summed-Score-Based Scoring in Non-Rasch IRT by Eisuke Segawa in Applied Psychological Measurement.


Articles from Applied Psychological Measurement are provided here courtesy of SAGE Publications

RESOURCES