Abstract
Objective:
The RDC/TMD contains two screening instruments for depression and somatization; the items were adopted in toto from the full SCL90. The present study sought to test whether extraction of two subscales (depression, somatization) from a well-known and widely used validated instrument (SCL-90) affected subscale reliability and validity.
Methods:
The full SCL90R and a modified version containing only the depression and somatization scales were administered in counterbalanced order to 103 subjects. As another test of context, a subset of participants completed the modified and full versions as part of a larger battery of instruments relevant to facial pain. Statistics included internal reliability for item analysis and intra-class correlation (ICC) and Lin’s Concordance Correlation Coefficient (CCC) for total scale score reliability.
Results:
Internal reliability was 0.95 for depression and 0.87 for somatization, independent of test form. Total scale scores were reliable across test versions, with both ICC and CCC approximately 0.95 for depression and 0.91 for somatization. Permutation tests using the CCC indicated a mild influence on the somatization score but not the depression score due to order effects, but these effects were not significant when considering the 95% CIs based on resampling methods.
Conclusion:
Whether items from other subscales are present or not does not affect the internal reliability or parallel forms reliability of the total scores from either depression or somatization. Context of administration, via order of forms completion, does not alter total score or reliability of depressive items but may alter total scores for somatization.
Introduction
Psychological self-report instruments are used extensively in medical and dental research. Increasingly complex study designs often place high burdens on subject participation and one method to reduce such burdens is to tailor the self-report assessments by extracting selected subscales from a validated parent instrument. Such a strategy was used in developing the Research Diagnostic Criteria for Temporomandibular Disorders (RDC/TMD). The RDC/TMD (translated into over 20 languages) is the most widely used research tool in the clinical setting for the diagnosis of TMD and for the assessment of psychosocial distress in a TMD population 1. TMDs are a group of musculoskeletal pain conditions associated with the muscles of mastication and/or the temporomandibular joint2, and they affect up to 18% of the US population3.
As with all chronic pain conditions, psychosocial distress and mental illness are quite common in TMD clinic populations.4–8 The RDC/TMD uses a dual axis diagnostic and classification system that includes one axis to record clinical physical findings and a second axis to record behavioral, psychological, and psychosocial status. The physical axis provides clinical researchers with a standardized system that can be evaluated for its use in examining, diagnosing and classifying the most commonly appearing subtypes of temporomandibular disorders. The biobehavioral axis (Axis II) was intentionally designed as a brief screening tool to assess for two pain-relevant psychological constructs (depression and somatization) via 32 items extracted from the SCL-909 and to evaluate pain-related interference via the Graded Chronic Pain Index10. Depression and somatization, as two scales comprising the SCL-90, were specifically selected for screening in this pain population because of the very strong theoretical and empirically demonstrated relationship of those constructs to the experience and progression of chronic pain. Importantly, the extraction of the two subscales from the SCL-90 for the RDC/TMD was accompanied by the administration of the isolated subscales in a random population design, developing independent norms for use of those two subscales in isolation of the parent instrument1 and shown to exhibit in their extracted form solid psychometric values11. Consequently, the use of extracted sub-scales in the context of the RDC/TMD is psychometrically justified, but the larger question is whether the scale values and their interpretation can be generalized to information as obtained by the parent instrument (the SCL-90 or SCL-90R).
The classical psychometric literature 12 13, however, has considered subscale extraction for independent application inappropriate and, moreover, formal guidelines for instrument development explicitly indicate that if a subscale is extracted, absence of alteration in the scores needs to be demonstrated 14. While the extraction of selected items from a parent instrument is associated with two established “sins” – the assumption that the reliability and validity of the parent items automatically applies to the extracted items, and the belief that less validity evidence is consequently needed 13, the claim of potential problems associated with the extraction of entire subscales is accompanied by little to no evidence. While the classical psychometric canon has cautioned against extraction of sub-scales from parent instruments, assumptions within both classical test theory (CTT) as well as item response theory (IRT) regarding item invariance would suggest that the performance of individual items would remain consistent regardless of context 15.
In a review of the literature on short form use and methodology, Smith and colleagues suggested that the literature has been characterized by an overly optimistic view that validity will transfer from the parent form to the isolated sub-scales 13. They suggest that methodologic and psychometric principles should be equally applied to the isolated subscales in order to develop valid clinical assessment tools. Within the RDC/TMD, the depression and somatization scales have been normed on large samples1, used in a wide range of studies in the US and internationally16–20, and have demonstrated reliability and validity with other similar measures11. However, the relationship of these two scales to the original scales—that is, as used in the RDC/TMD and as measured in the SCL-90—has never been assessed. This question has direct implications for the equivalence of the short version containing 32 items as used in the RDC/TMD and the long form of the SCL-90, and, of course, this question has more general implications for other instruments in wide-spread use in the same manner. The aim of the present study was to test whether extraction of the two subscales in the RDC/TMD affected the subscale score reliability and whether scores from the RDC/TMD subscales are comparable to the same scales when the whole SCL90/R is administered.
Methods
Subjects.
103 subjects between 18 and 65 years of age and proficient with English language usage were recruited sequentially from two main sources, a private facial pain practice (n=51) and dental school patients identified as having special social and/or financial needs (n=52). The latter group was known to exhibit life stress that interferes with their ability to fully participate in their dental treatment, and among that group, subjects were either patients in a specialty pain teaching clinic (n=15) or were general dental school patients (n=37). We sampled from the indicated populations in order to minimize the number of low responses across the two testing situations since that would upwardly bias any agreement. Of subjects approached for the study, there were 2 refusals, and 114 subjects entered the study; 11 subjects returned incomplete data, and the final sample of 103 subjects comprised 27 males (mean age 44.7, SD 12.3) and 76 females (mean age 41.5, SD 12.3). The preponderance of females is consistent with the gender distribution in each recruitment source. Subjects received monetary compensation of $10 for completing the study. The study was approved by the Health Sciences IRB, and informed consent was obtained from each subject.
Procedures.
The 32 RDC/TMD items assessing depression and somatization as extracted from the original SCL-90 comprised the “modified form” of the targeted instrument. The full SCL-90 comprised the “full form”. Item sequence in the modified form was as published in the original RDC/TMD, which followed, for the most part, the ordering of the corresponding items in the SCL-90; the item sequence in the full form was exactly as published in the SCL-90. In using “depression” and “somatization” as the two target constructs, two assessment domains were addressed in the present study: a set of items that relates to more psychological states (depression symptoms, per DSM IV) and a set of items that relates to more physical symptom states (specifically, non-functional symptoms) such as would easily be found in a medical symptom checklist.
Each subject received both forms, counter-balancing order of administration across subjects as they entered the study. The counter-balancing of the modified and full forms of the target instrument created a context effect, in that those subjects randomly assigned to the order of completing the full form first carried context effects of the other constructs that comprise the SCL-90 to the modified form, completed second. The subjects assigned to the reverse sequence carried to the second form administration a potentially stronger bias of depression and somatization.
A second factor was presence vs absence of other instruments (e.g., disability, pain symptoms, limitation, stress experience items, demographics) administered at the same time as the target instruments. Because this study was conducted in a clinical setting, some subjects were recruited for this study separate from their clinical process and some subjects were recruited as part of their standard clinical evaluation. While assignment to counter-balanced order was random per entry into the study, subjects who completed other forms with the target instruments did so on a quasi-experimental basis21.
Subjects from the respective settings were recruited, and the first study instrument was administered either in person or at home (depending on clinic flow procedures); the second form was always completed at home. All forms were self-administered. Forms from home were mailed in, and postal date was compared against declared date of completion in order to verify time interval between administrations. The requested time separation between administrations was 2 days, with an allowable window of 1-9 days. The ideal time interval for assessing parallel forms of an instrument is to administer on the same day; however, endorsement by recall from the prior administration would potentially confound the study goals, and a long interval between administrations could result in low agreement due to the person’s mood or bodily symptom states changing. Two days, as the target interval, was selected based on clinical experience that the constructs under examination do not typically change over that short interval.
Scoring.
Consistent with recommended practice for the SCL-9022, a summary score was created for each scale within each instrument by computing the simple sum of the endorsed ordinal rank for each item. Following published scoring rules for the depression and somatization scales within the RDC/TMD assessment protocol, the depression score was based on 20 items (derived from the 13 original SCL-90 items for depression and the 7 items for vegetative symptoms) and the somatization score was based on 12 items. Scoring of the respective constructs in the long form was computed based on the same items. Per RDC/TMD guidelines, the sums were adjusted for missing values, as long as there were at least 2/3 valid responses within a construct; there were 94% complete responses for the depression scales with 6% of subjects with up to 2 missing items, and there was 96% complete responses for the somatization scales with 4% with 1 missing item. Then, the summed score was rescored on a 0-4 metric based on the number of items with valid responses.
Statistical Analyses.
Internal reliability, via Cronbach’s alpha, was computed as an index of individual item performance. Raw scores for the two versions (modified, full) of each scale were compared using % difference, per the other two factors (counter-balanced order; isolation of instrument administration), in order to provide a descriptive summary of how the scales performed. In keeping with general methods of presenting test reliability, the modified form and the full form were compared using several approaches: Pearson correlation, ICC (fixed raters)23, and an ICC computed from a full factorial model which included coding for the other two experimental factors present in the data collection methods. However, while the ICC is an often used statistic for instrument reliability, it is not without problems or critique24–28. The most notable problem is the large impact that sampling bias has on restriction of range, leading to a biased and underestimated statistic. Lin developed the Concordance Correlation Coefficient (CCC) in order to resolve these criticisms29, and consequently the CCC is used for the primary statistic of reliability between the modified version and the full version for each construct. Note that a CCC valued 1.0 denotes perfect agreement, and a value of 0.0 denotes no agreement.
To evaluate the extent that the modified and full forms of depression and somatization agree, the CCC was computed for different subsets of the data. Bootstrap re-sampling methods were used to obtain 95% confidence intervals. Specifically, re-sampling with replacement of the data was simultaneously done from each group; for each re-sample, the CCC was calculated. This procedure was repeated 10,000 times. A confidence interval for the 2.5th and 97.5th percentiles from this simulated distribution was then obtained. In order to statistically evaluate differences observed between CCC’s corresponding to subsets of interest, permutation testing methods were used. The data were permuted, ignoring subset assignment, after which the difference in CCC between the two groups was calculated. This was done 10,000 times in order to obtain the required null distribution from which the 2-sided p-value was obtained.
In addition to the CCC which incorporates magnitude of values, equivalence of summary scores was also assessed with factorial ANOVA using partial sums of squares in order to assess the factors of counter-balanced order and isolation of instrument administration on difference scores derived from the full form and modified form. In addition, as a secondary analysis in order to assess how the presence of other instruments affects respondent behavior, another set of ANOVAs were computed testing the effects of these factors on each of the available raw scores from the target instruments. From the associated factorial cell means, differences between contrasts of interest were computed and converted to percentages in order to estimate any practical impact based on the various factors implemented in this study. Stata 8.0 and SAS v9 were used for statistical analyses.
Results
While a 2-day interval between administrations was deemed optimal, subjects completed the second instrument from 1-9 days after the first (Figure 1). Fifty-four percent of the sample completed the second form within two days of the first form. In order to assess whether a very short interval vs a long interval between time1 and time2 administrations had any impact on responding to the second instrument, we plotted the difference scores of the summed score responses from time2 to time1 as a dot-plot in order to determine whether there was a trend over the time interval between adminstrations25. As the dot-plots in Figure 2 demonstrate, the short interval of 1-2 days resulted in the same absence of effect as did the longer periods, for both depression and somatization subscales. The data exhibited descriptive values sufficient for the present study: the mean value of depression was 1.1 (SD 0.91; min 0, max 3.5) and the mean of somatization was 1.04 (SD 0.73, min 0.05, max 3.05), and the score means for each measure did not differ across the subject groups based on recruitment source (p>0.14).
Figure 1.

Histogram of number of days between self-administration of first study form and second study form. See Figure 2 for assessment of implications of this difference.
Figure 2.

Scatter plot and dot plot of raw data. Scatter plot displays total scale score (adjusted for missing) for each of modified and full form versions of depression and of somatization. Dot-plot displays the difference between the score obtained from instrument administered at time-2 and score from time-1 administration, according to the number of days between administrations. The Lowess regression line demonstrates absence of appreciable effect over time upon the difference in total score.
Internal reliability, via Cronbach’s alpha, was computed for only non-missing data (Table 1). Internal reliability was excellent for both constructs of depression and somatization, with that of depression slightly higher than that of somatization. There was no difference in overall internal reliability according to whether the modified or full form version was used, for either the depression or somatization subscales.
Table 1. Internal Reliability and descriptive statistics.
Cronbach’s alpha estimates for each construct, according to whether full instrument (full) or subset of items (modified). The mean scores, comparing full vs modified, did not differ (paired t-test, p>0.05) within each of depression or somatization.
| Overall alpha | Mean | SD | Range | |
|---|---|---|---|---|
| Depression | ||||
| — Full | 0.953 | 1.08 | 0.90 | 0 – 3.5 |
| — Modified | 0.948 | 1.13 | 0.87 | 0 – 3.5 |
| Somatization – All items | ||||
| — Full | 0.868 | 1.00 | 0.74 | 0.1 – 3.0 |
| — Modified | 0.868 | 1.08 | 0.76 | 0 – 3.1 |
Standard reliability statistics are shown in Table 2, indicating highly comparable responses between Pearson and ICC statistics; given the adherence of the raw data to the line of unity as shown in the scatter plots in Figure 2, the equivalence of the statistics is not surprising. While the items for depression exhibit a higher level of reliability between modified and full versions of the instrument, the reliability statistics for both depression and somatization are excellent.
Table 2. Summary Reliability Statistics.
Comparison of total scale scores from modified vs full forms of the respective SCL-90 subscales.
| Depression | Somatization | |
|---|---|---|
| Pearson Correlation | 0.961 | 0.908 |
| Simple ICC * (fixed raters) | 0.960 | 0.907 |
| Complex ICC ** (fixed raters) | 0.959 | 0.905 |
Based on 1-way ANOVA, comparing modified to full instrument total scores.
Based on full ANOVA factorial model: modified vs full, form sequence from counter-balancing, and whether administered with Other Instruments
The CCC and 95% confidence interval for each construct, according to each of the different experimental factors in this study, are shown in Table 3; for these analyses 6 subjects were dropped because of missing data related to whether they completed study forms with other forms or not. The experimental factors of counterbalancing and modified vs full instruments make little difference in the reliability of scores for either of depression or somatization; in contrast, whether one of the target instruments (modified or full) was completed in isolation of other instruments or not results in a substantial shift in the CCC for somatization but not depression. In order to assess whether that shift in CCC reliability was significant, the difference between the respective CCC statistics was computed and tested against a distribution created via permutation tests. As shown in Table 4, none of these differences were significant, demonstrating that neither of the two observed experimental factors (whether modified vs full instrument was administered first; completion of the instrument in isolation or with other instruments) had any appreciable effect on the scale reliability.
Table 3. Estimated CCC for each group .
Isolques 0 = no, and 1=yes for whether the form was administered in isolation of other instruments (cf., with other instruments). Formseq 1 = modified → full, while Formseq 2 = full → modified. Sample size for total sample in this analysis (n=97) differs from the n=103; see text for explanation.
| Pair of Comparing Variable |
Isolques | Formseq | Sample Size |
Estimated CCC | 95% Bootstrap CI of CCC (simulation=10,000) |
|---|---|---|---|---|---|
| DEP | Overall | 97 | .958 | [ .932, .975] | |
| 0 | · | 24 | .959 | [ .867, .984] | |
| 1 | · | 73 | .956 | [ .931, .972] | |
| · | 1 | 47 | .939 | [ .886, .969] | |
| · | 2 | 50 | .978 | [ .956, .990] | |
| 0 | 1 | 12 | .926 | [ .520, .974] | |
| 0 | 2 | 12 | .987 | [ .849, .974] | |
| 1 | 1 | 35 | .940 | [ .872, .974] | |
| 1 | 2 | 38 | .974 | [.943, .990] | |
| SOM | Overall | 97 | .894 | [.837, .928] | |
| 0 | · | 24 | .696 | [.458, .816] | |
| 1 | · | 73 | .920 | [.875, .950] | |
| · | 1 | 47 | .777 | [.648, .871] | |
| · | 2 | 50 | .955 | [.875, .950] | |
| 0 | 1 | 12 | .597 | [.189, .770] | |
| 0 | 2 | 12 | .797 | [.423, .903] | |
| 1 | 1 | 35 | .812 | [.672, .922] | |
| 1 | 2 | 38 | .970 | [.947, .984] | |
Table 4. Estimated difference in CCC between different groups.
See Table 3 for explanation of codes.
| Pair of Comparing Variable | Comparing | Sample size | Estimated CCC(1)-CCC(2) | P-value |
|---|---|---|---|---|
| DEP | isolques=0 isolques=1 |
N1=24 N2=73 |
.00319 | .985 |
| formseq=1 formseq=2 |
N1=47 N2=50 |
−.0387 | .791 | |
| SOM | isolques=0 isolques=1 |
N1=24 N2=73 |
−.224 | .118 |
| formseq=1 formseq=2 |
N1=47 N2=50 |
.−.178 | .152 |
Scale scores for each construct are shown in Table 1. Using factorial ANOVA for the difference score between full form and modified form for depression, there were no significant main or interaction effects (p>0.14); the same was true for the somatization instruments (p>0.17). These results indicated that scale scores for the full form and the modified form were not different from one another, for each of depression and somatization, within the two factors assessed in this study. In contrast, administration of other instruments simultaneous with the target instrument significantly decreased scale scores, regardless of forms sequence, for the full form (p=0.045) and marginally for the modified form (p=0.053) for depression; there was no impact by administration of other instruments on either the full form score (p=0.35) or the modified form score (p=0.35) of somatization. Inspection of raw means, partitioned by the two study factors, disclosed that individuals consistently exhibited a pattern of endorsing a lower level of depression symptoms when the target instrument was administered with other instruments vs when administered alone. Importantly, this was equally true for each of the full and modified test forms. When the mean total values were re-expressed as a percentage difference of the factorial mean values, the simple difference in scale scores between the full instrument and the modified instrument for each of depression and somatization was 5% and 7%, respectively; in contrast, the presence of other instruments resulted in differences of up to 35% in the scale scores for depression instruments.
Discussion
This study was undertaken in order to answer two questions. The more specific and answerable question is whether the two subscales of depression and somatization, as used in the RDC/TMD, can be extracted from the SCL-90 and retain their validity and reliability with respect to retaining the same interpretation of the construct. The importance of this specific question lies in the daily clinical use of administering these items world-wide in TMD assessment. The second question is one of psychometric principle, and asks if the general caution against subscale extraction is indeed necessarily warranted. From a validity perspective, these data indicate that item extraction from the parent instrument was successful. Scores are comparable regardless of using the full instrument or the subset of scales and the resulting instrument is shorter and more appropriately tailored for the specific research task. This finding, by itself, suggests that the use of the SCL-90 subscales of depression and somatization in the RDC/TMD protocol is indeed valid not only in terms of the separate validity data published for the RDC/TMD protocol, but now also for comparability of scores obtained using the RDC/TMD protocol to scores obtained in settings where the full SCL-90 is used.
Counter-balancing in this study resulted in 100% of the subjects completing the same items again, with 50% of the subjects doing so a second time with only the 2 scales comprising the modified form, and with 50% of the subjects doing so a second time with the other 58 items from the SCL-90 interwoven into the study’s primary target items. Counter-balancing thus also forms an experimental variable of tightly controlled context, and this more local context did not appreciably alter subject responses. Similarly, reduction in the number of items for the CES-D, a tool comparable to the SCL-90 for assessing depression11, did not result in any appreciable alteration in its core psychometric properties30. In contrast, studies examining serial order of items report consistent changes in response patterning due to changing item order (as would occur in the present study in the full instrument form); however, the studies examining serial order effects have been limited to personality assessment 31. Overall, these findings suggest that changes in internal instrument structure, depending on the underlying construct, can occur without affecting scoring properties.
A different type of context effect occurred in this study when other instruments were administered at the same time as our target instruments. In the assessment of personality with self-report instruments, it is hypothesized that a self-reflective focus 32 and consequent engagement of the self 33 is created by the context of the instrument, and that that results in the individual endorsing more rather than fewer characteristics about themselves due to better access to memory. The present data suggest the opposite pattern for depression symptoms, at least in a dental setting, in that other instruments administered with the target instrument resulted in the individuals reporting less symptomatology compared to when they completed the test instrument alone. It is tempting to interpret this as due to time effects: a longer instrument results in the individual allocating perhaps less time pondering individual items and hence underreport relative to when they complete the target instrument alone. However, that the somatization data did not yield such a pattern suggests that at least one other factor may be operating. Since the bulk of the other items comprising the other instruments were focused on pain and functioning, perhaps the individuals reframed their depressive symptoms as part of the pain disorder (hence, under-reporting the depression); and the somatization symptoms (i.e., physical body symptoms) reporting did not decline because they are congruent with the pain disorder. Other evidence suggests that the content of the early items within an instrument appears to clarify the meaning of later items in the instrument with the consequence of improved overall reliability33. While these are clearly important considerations for future research addressing self-report based information, for the present study the important conclusion is that these other influences affected not only the modified form but to the same extent the full form..
Overall, the present findings suggest that the general caution against use of extracted subscales must be considered in context: different kinds of items can be expected to either be sensitive to, or not sensitive to, context effects, and while that context may be embedded within the administered target items, that context is certainly also created by the proximity of other instruments in addition to the items comprising a parent instrument, and, finally, that context affects items differentially. In the case of the present study, items assessing bodily symptoms were more sensitive to the items in other instruments assessing pain-relevant symptoms, while pain-relevant mood items (i.e., depression) were not sensitive to that context.
The context of the present study itself should be considered as yet one more layer in this investigation. Over 50% of the subjects were currently being either evaluated or treated in a dental (not psychological) facility for a chronic pain condition, and the remainder of the subjects, while being evaluated by a social work unit, were available for recruitment in this setting due to their primary complaint of dental problems. Hence, it would be expected that this subject sample had a high likelihood of substantial priming of physical (vs mental health) symptoms and beliefs.
The present study has several limitations. Foremost is that the study used only two subscales from one multi-dimensional instrument; this hardly answers the larger question of whether any subscale from all multi-dimensional instruments can be extracted. Indeed, if extracted from a parent instrument, subscales should be validated independently without presumption that the psychometric properties of the parent carry over to the extracted subscales 13. A second limitation is that the administration of other instruments was performed in only a subset of subjects and although that sample was adequate in size for the boot-strap resampling procedure potential bias in sampling of those subjects (since they came from a different clinical population) cannot be ruled out. A third limitation is that this study only examined the impact of other self report instruments and items on performance of subscales. Future research in the clinical setting should also examine the impact of clinical examinations on performance of instruments. Finally, this study was conducted in a dental school, among patients who despite an approximate 50% rate of diagnosable mental disorders were nevertheless seeking somatic help for somatically oriented problems (at least, based on the nature of the chief complaints among these two populations); subjects recruited from other areas may behave differently with respect to the tension between constructs such as depression vs somatization.
As measured with SCL-90 items using two questionnaire versions, depression was noted to exhibit extremely robust reliability and to not be influenced by adjacent items or other instruments. Somatization exhibited very good reliability, but was also more influenced by context of immediately adjacent items and other instruments. Consequently, extraction of subscales from other instruments should be empirically validated for retaining the expected reliability. In sum, current wisdom says that sub-scales cannot be extracted, yet we turn a blind eye to other context influences potentially affecting response behaviors. Because of increasing pressures to administer shorter tests (either by reducing the number of items, or by administering only the necessary subscales of a multidimensional instrument), the question of whether subscales can be extracted from a parent instrument is even more relevant clinically in current settings in addition to its importance for psychometric theory. These data suggest that extraction and administration of a sub-scale need not necessarily compromise the reliability and validity of that scale.
Acknowledgments
The authors thank Dr. Yoly Gonzalez, Dr. Lance Ortman, and the School of Dental Medicine staff at the University at Buffalo for their help during the recruitment phase. This research was presented at the International Association for Dental Research, Göteborg, Sweden, July 2003. Research supported by NIH-DE-13331 from NIDCR/NIH, Bethesda, MD, US. The RDC/TMD Validation Study Group is comprised of: University of Minnesota: Eric Schiffman (Study PI), Mansur Ahmad, Gary Anderson, Quintin Anderson, Pat Carlson, Mary Haugan, Amanda Jackson, John Look, Wei Pan; University at Buffalo: Richard Ohrbach (PI), Leslie Garfinkel, Yoly Gonzalez, Krishnan Kartha, Sharon Michalovic, Betsy Seyler-Piccolo, and Theresa Speers; University of Washington: Edmond Truelove (PI), Lars Hollender, Kimberly Huggins, Kathy Scott, and Earl Sommers.
Contributor Information
Richard Ohrbach, University at Buffalo.
Jeffrey Sherman, University of Washington.
Carla Beneduce, University at Buffalo.
Kimberly Zittel-Palamara, University at Buffalo.
Youngju Pak, University at Buffalo.
References
- 1.Dworkin SF, LeResche L. Research Diagnostic Criteria for Temporomandibular Disorders: Review, Criteria, Examinations and Specifications, Critique. J Craniomandib Disord Facial Oral Pain 1992;6:301–355. [PubMed] [Google Scholar]
- 2.McNeill C, Mohl ND, Rugh JD et al. Temporomandibular disorders: diagnosis, management, education, and research. JADA 1990;120:253–263. [DOI] [PubMed] [Google Scholar]
- 3.Von Korff M, Dworkin SF, LeResche L et al. An epidemiologic comparison of pain complaints. Pain 1988;32:173–183. [DOI] [PubMed] [Google Scholar]
- 4.Kight M, Gatchel RJ, Wesley L. Temporomandibular disorders: evidence for significant overlap with psychopathology. Health Psychol 1999;18:177–182. [DOI] [PubMed] [Google Scholar]
- 5.Ohrbach R, LeResche L, Dworkin SF. Longitudinal changes in TMD: Influence of Baseline Findings and Treatment-Seeking. Journal of Dental Research 76, 389. 1997. [Google Scholar]
- 6.Dworkin SF. Behavioral, emotional, and social aspects of orofacial pain. In: Stohler CS, Carlson DS, eds. Biological and Psychological Aspects of Orofacial Pain. Ann Arbor, MI: Center for Human Growth and Development,1994:93–112. [Google Scholar]
- 7.Wilson L, Dworkin SF, Whitney C et al. Somatization and pain disperson in chronic temporomandibular disorder pain. Pain 1994;57:55–61. [DOI] [PubMed] [Google Scholar]
- 8.Turner JA, Dworkin SF. Screening for psychosocial risk factors in patients with chronic orofacial pain: recent advances. JADA 2004;135:1119–1125. [DOI] [PubMed] [Google Scholar]
- 9.Derogatis LR, Lipman RS, Covi L. SCL-90: an outpatient psychiatric rating scale--preliminary report. Psychopharmacology 1973;9:13–28. [PubMed] [Google Scholar]
- 10.Von Korff M, Ormel J, Keefe FJ et al. Grading the severity of chronic pain. Pain 1992;50:133–149. [DOI] [PubMed] [Google Scholar]
- 11.Dworkin SF, Sherman JJ, Mancl L et al. Reliability, validity, and clinical utility of RDC/TMD Axis II scales: Depression, non-specific physical symptoms, and graded chronic pain. J Orofacial Pain 2002;16:207–220. [PubMed] [Google Scholar]
- 12.Anastasi A Psychological Testing. New York: Macmillan Publishing Company, 1988. [Google Scholar]
- 13.Smith GT, McCarthy DM, Anderson KG. On the sins of short-form development. Psychological Assessment 2000;12:102–111. [DOI] [PubMed] [Google Scholar]
- 14.American Educational Research Association. Standards for Educational and Psychological Testing. Washington, D.C.: Author, 1999. [Google Scholar]
- 15.Embretson SE, Reise SP. Item Response Theory for Psychologists. Mahwah,NJ: Lawrence Erlbaum, 2000. [Google Scholar]
- 16.List T, Dworkin SF. Comparing TMD diagnoses and clinical findings at Swedish and U.S. TMD centers using Research Diagnostic Criteria for Temporomandibular Disorders. J Orofacial Pain 1996;10:240–253. [PubMed] [Google Scholar]
- 17.Lobbezoo F, van Selms MKA, John MT et al. Use of the Research Diagnostic Criteria for Temporomandibular Disorders for multinational research. Translation efforts and reliability assessments in The Netherlands. J Orofacial Pain 2004;19:301–308. [PubMed] [Google Scholar]
- 18.Wahlund K, List T, Dworkin SF. Temporomandibular disorders in children and adolescents: Reliability of a questionnaire, clinical examination, and diagnosis. J Orofacial Pain 1998;12:42–52. [PubMed] [Google Scholar]
- 19.Yap AU, Dworkin SF, Chua EK et al. Prevalence of temporomandibular disorder subtypes, psychologic distress, and psychosocial dysfunction in Asian patients. J Orofacial Pain 2003;17:21–28. [PubMed] [Google Scholar]
- 20.John MT, Hirsch C, Reiber T et al. Translating the Research Diagnostic Criteria for Temporomandibular Disorders into German: Evaluation of content and process. J Orofacial Pain 2006;20:43–52. [PubMed] [Google Scholar]
- 21.Campbell DT, Stanley JC. Experimental and Quasi-Experimental Designs for Research. Boston: Houghton Mifflin Company, 1963. [Google Scholar]
- 22.Derogatis LR. SCL-90-R: Administration, Scoring and Procedures Manual-II, for the Revised Version. Towson, MD: Clinical Psychometric Research, 1983. [Google Scholar]
- 23.Shrout PE, Fleiss JL. Intraclass correlations: uses in assessing rater reliability. Psychol Bull 1979;86:420–428. [DOI] [PubMed] [Google Scholar]
- 24.Bartko JJ, Carpenter WT Jr. On the methods and theory of reliability. J Nerv Ment Dis 1976;163:307–317. [DOI] [PubMed] [Google Scholar]
- 25.Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet 1986;1:307–310. [PubMed] [Google Scholar]
- 26.-----A note on the use of the intraclass correlation coefficient in the evaluation of agreement between two methods of measurement. Comput Biol Med 1990;20:337–340. [DOI] [PubMed] [Google Scholar]
- 27.Chinn S The assessment of methods of measurement. Stat Med 1990;9:351–362. [DOI] [PubMed] [Google Scholar]
- 28.Ludbrook J Comparing methods of measurement. Clinical and Experimental Pharmacology and Physiology 1997;24:193–203. [DOI] [PubMed] [Google Scholar]
- 29.Lin LI. A concordance correlation coefficient to evaluate reproducibility. Biometrics 1989;45:255–268. [PubMed] [Google Scholar]
- 30.Cole JC, Rabin AS, Smith TL et al. Development validation of a Rasch-derived CES-D short form. Psychological Assessment 2004;16:360–372. [DOI] [PubMed] [Google Scholar]
- 31.Steinberg L Context and serial-order effects in personalty measurement: limits on the generality of measuring changes the measure. J Person Soc Psych 1994;66:341–349. [Google Scholar]
- 32.Hamilton JC, Shuminsky TR. Self-awareness mediates the relationship between serial position and item reliability. J Person Soc Psych 1990;59:1301–1307. [Google Scholar]
- 33.Knowles ES, Byers B. Reliability shifts in measurement reactivity: driven by content engagement or self-engagement? J Person Soc Psych 1996;70:1080–1090. [DOI] [PubMed] [Google Scholar]
