Abstract
In the controversy over civil commitment procedures, reliability as well as validity of clinicians’ assessments have been challenged. In this study, the generalizability of TRIAD, an observational assessment device, was tested. Patient load strongly influenced the degree to which TRIAD predicted case disposition and clinician global ratings of dangerousness and grave disability. Given comparable patient–clinician ratios, TRIAD predicted 81% to 86% of case dispositions. Agreement between clinician global assessments and TRIAD ratings was high to moderate. Clinicians apparently agree on sets of indicators which can be consistently weighted, but the application of the standard described by TRIAD may be jeopardized by increasing patient loads.
The civil commitment process has been criticized on two grounds. First, the process is said to be arbitrary in that clinicians do not apply legal standards in a consistent way.1 That is, clinicians’ decisions are unreliable and therefore inequitable. Second, the civil commitment laws seemingly require that psychiatric personnel make predictions which no one is qualified to make. That is, clinicians’ predictions are invalid.2
Logically, consistency or reliability precedes validity. You cannot know whether you are measuring true dangerousness or disability (validity) without knowing first that you are measuring the same thing each time (reliability). The research concept of reliability, or measuring the same thing each time, translates, in the commitment process, into the practical issue of the consistent application of legal standards—that is, equity in the evaluation of patients. It is to the issue of consistency or reliability that we have addressed ourselves. In so doing we also have addressed the question of equity.
A major function of the psychiatric emergency service (PES) is the implementation of the civil commitment laws in the process of involuntarily hospitalizing patients. The courts and legislatures have established criteria for emergency involuntary hospitalization, usually similar to the California criteria, danger to self, danger to others, and grave disability.3 However, the states have left the substantive interpretation of civil commitment criteria to mental health professionals, assuming that, in the absence of evidence for predictive accuracy, there are at least professional standards which can be consistently applied. In view of this assumption, it is surprising to find that, of the several previously published studies which have examined clinical reasons for admission decisions,4 none has attempted to describe the clinical application of any of the legal or statutory criteria.
The present authors assumed that demonstrating consistency in the application of the legal criteria would require an adequate description of the bases of the clinical evaluation.5 We developed an index entitled “Three Ratings of Involuntary Admissibility (TRIAD)” to reflect the way clinicians in psychiatric emergency rooms interpret and apply the concepts “danger to self,” “danger to others,” and “grave disability” due to mental disorder. TRIAD was developed through an iterative process resulting in the identification and ranking of patterns of behavior and circumstance more or less likely to lead to the determination that a patient is involuntarily admissible by the standards of California’s Lanterman–Petris–Short (LPS) Act.6
We theorized that through professional training and experience, clinicians are sensitized to clusters or patterns of behavior and circumstance that are believed to be associated with danger to self, danger to others, and grave disability, and that clinicians internalize scales by which they weigh or rank these patterns. (On each of these scales several quite different-looking patterns are given each weight or rank.) Clinicians, we proposed, would react to some patterns as unambiguously dangerous or not dangerous, and they would consistently respond to these patterns with decisions that a person was admissible or not admissible under involuntary criteria. Admission decisions would therefore be highly consistent in cases involving unambiguous patterns. Other patterns would be experienced as more ambiguous, and this ambiguity would lead to a greater variation in the outcome of the evaluation process.
In 1981, we observed 89 cases at two urban county hospital psychiatric emergency rooms in the San Francisco Bay Area, leading to the following two conclusions.7 First, we had developed an index of indicators used by clinicians with interrater reliability coefficients of Pearson’s r equal to 0.89, Danger to Self score; 0.94, Danger to Others score; 0.77, Grave Disability score; 0.89, total TRIAD score. Second, TRIAD appeared to be a valid index of those aspects of a case to which clinicians respond, and the weights they assign various constellations of those aspects, in evaluating dangerousness and disability. This tentative conclusion was drawn from the fact that TRIAD scores correctly predicted dispositions in 82% of 89 cases observed in these two psychiatric emergency rooms. The results of this study further indicated that clinicians had employed shared constructs of danger to self, danger to others, and grave disability, and that these constructs could be reliably applied in actual cases.
In 1983 we undertook a new investigation, the results of which are reported below, to determine the generalizability (external validity) of the evaluative criteria embodied in TRIAD across time and across settings.
Method
One hundred one cases were observed at two psychiatric emergency rooms. Fifty-one of these case observations were made at one of the original hospital settings in which the 1981 study was conducted. Fifty observations were made at a new emergency room—one which was different from the original rooms in several ways. Unlike the original settings, which were separate psychiatric emergency facilities within the public general hospitals of large cities, the new setting was an emergency psychiatric service operating within the medical emergency room of a public general hospital in a suburban community.
Each case observation involved the independent assessment of a case by a researcher who accompanied the clinician and the patient through the psychiatric emergency evaluation of that case. The clinician simply went about his business in the course of the evaluation, keeping the research observer informed as to any information he had received by phone or in writing. Scores on the TRIAD index were computed independently by the observer on the basis of the information gathered up to the point at which the clinician made a disposition decision on the case. The clinician at no time was aware of the contents of the TRIAD form. The major criterion for determining the accuracy of the TRIAD assessment as a reflection of the clinician’s assessment was the agreement between TRIAD Severity scores and the clinician’s disposition decision. The authors hypothesized that the higher the Severity score on TRIAD, the greater the probability of a decision to hold a patient involuntarily.
A second criterion, introduced in the 1983 study, was the agreement between TRIAD scores and clinician global ratings of patients. At the time the researcher scored TRIAD, the clinician was asked to provide an independent rating of the patient on global scales of Danger to Self, Danger to Others, and Grave Disability, with ranges and points equivalent to the possible scores on the structured TRIAD instrument.
The TRIAD instrument
TRIAD consists of three checklists with a total of 84 numbered items which can be combined to yield 146 patterns of behavior and circumstance relevant to the clinical prediction of violence and suicide and the assessment of grave disability. On each of the scales, several patterns are assigned the highest score, several are assigned the next highest score, and so on. No pattern combines more than 9 items, and most involve 2, 3, or 4 items.
For example, “threatened to harm another” is one item which, by itself, scores at a moderate level (2), on the Danger to Others scale. However, such a threat may yield a higher score (4) if it occurs with three other particular items. The first additional item has to do with provocation or lack thereof. The others involve indications of a concrete plan and/or weapon, and/or being in a volatile or unpredictable or enraged state, and/or having a history of assault.
According to the assumptions embodied in TRIAD, if such a presenting picture is accompanied by mental disorder, the evaluating clinician will determine the patient is clearly admissible by the LPS standards. In order to prevent hospitalization, the clinician may attempt to bring about some change in the picture through crisis intervention or medication in the emergency room, but if these efforts fail, admission will follow. If the efforts succeed, the Danger to Others score will be lower than it would otherwise have been. Other patterns seem equally clear, but some are more ambiguous and yield intermediate scores. Scoring on TRIAD is completed at the time of disposition by finding the standard pattern that includes the checked items and yields the highest score.
As discussed above, the primary objective of the research effort was to document the degree of consistency, or the reliability, of the clinicians’ assessments of dangerousness and grave disability across patients, settings, and time. Based on the observation that clinicians in the midst of the emergency evaluation frequently ask themselves, “Is this person holdable?”, we posited, in our development of TRIAD, the construct “admissibility.”8 We hypothesized that clinicians find a patient more or less “admissible” or “holdable” on the basis of danger to self, danger to others, or grave disability. The question “How admissible or holdable is a patient?” is roughly equivalent to the questions “How dangerous is this person to himself? How dangerous is he to others? How disabled is he?” Thus, TRIAD is a measure of the extent to which the behavior and circumstances of the patient are commensurate with clinicians’ concept of admissibility—that is, their concepts of dangerousness and grave disability.
TRIAD Severity level is the overall severity of the patient’s presentation on the three legal criteria. Severity is the degree to which the patient’s presentation on any one criterion or across criteria9 corresponds to the hypothesized construct “involuntary admissibility.” Severity scores take into account individual scale scores and also the sum of all scale scores.
Results
In general, the results of the study indicated that, indeed, the TRIAD instrument could reliably predict case disposition across the 18-month stretch of time and across two quite different psychiatric emergency settings. However, in the second round of evaluations (1983), TRIAD proved to be a less consistent predictor of disposition. This difference resulted entirely from a decline in the accuracy of TRIAD predictions within Facility #1 between Time 1 (1981) and Time 2 (1983). As will be detailed below, the decline in predictive power within the setting and across time is explained by changes in the PES environment.
TRIAD performance across time
During May and June 1983, 51 cases were observed and rated at Facility #1, one of the two sites of the 1981 study. In July, August, and September 1983, 50 psychiatric emergency cases were studied at Facility #2. Of the 101 cases observed in 1983, severity of presentation, as measured by TRIAD, correctly predicted disposition of 78% (gamma = 0.80), as compared with 82% of 89 cases at two facilities in 1981 (gamma = 0.86).
As clinicians are legally required to apply the criteria of dangerousness and grave disability only in cases of involuntary admission, TRIAD may be expected to be a better predictor of involuntary than voluntary admissions. When patients retained voluntarily (n = 8 in 1983 and n = 4 in 1981) were eliminated from the sample, TRIAD correctly predicted disposition in 80% of 93 cases in 1983 (gamma = 0.82), as compared with 81% of 85 cases at two facilities in 1981 (gamma = 0.85). In distinguishing patients released from patients retained involuntarily, therefore, TRIAD performed with almost equal accuracy in both years.
Prediction of disposition at Facility #1
To assess the generalizability of TRIAD across time, we compared TRIAD’S rates of success in predicting disposition at Facility #1 in 1981 and 1983. The results are summarized in tables 1 and 2.
TABLE 1.
Disposition of cases by Severity level: Facility #1,1981 (n = 31)
| Severity level | Released | Retained voluntarily | Retained involuntarily | Total |
|---|---|---|---|---|
| Level 1 (DSS, DOS, GDS = 0 or 1; total ≤ 3) |
3 (75%) | 0 | 1 (25%) | 4 (100%) |
| Level 2 (DSS, DOS, GDS = 2; total = 2) |
5 (83.3%) | 0 | 1 (16.7%) | 6 (100%) |
| Level 3 (DSS, DOS, GDS = 2; total = 3) |
1 (50%) | 0 | 1 (50%) | 2 (100%) |
| Level 4 (DSS, DOS, GDS = 3 or 4 or total ≥ 4) |
3 (15.8%) | 0 | 16 (84.2%) | 19 (100%) |
NOTE: In this and all of the following tables, DSS is “Danger to Self score”; DOS is “Danger to Others score”; and GDS is “Grave Disability score.” For summary purposes, those case dispositions considered “correctly predicted” include cases in which patients scoring at Severity Level 4 were retained and patients scoring at Severity Levels 1 through 3 were released.
TABLE 2.
Disposition of cases by Severity level: Facility #1, 1983 (n = 51)
| Severity level | Released | Retained voluntarily | Retained involuntarily | Total |
|---|---|---|---|---|
| Level 1 (DSS, DOS, GDS = 0 or 1; total ≤ 3) |
7 (58.3%) | 0 | 5 (41.7%) | 12 (100%) |
| Level 2 (DSS, DOS, GDS = 2; total = 2) |
3 (50%) | 1 (16.7%) | 2 (33.3%) | 6 (100%) |
| Level 3 (DSS, DOS, GDS = 2; total = 3) |
0 | 0 | 1 (100%) | 1 (100%) |
| Level 4 (DSS, DOS, GDS = 3 or 4 or total ≥ 4) |
6 (18.8%) | 2 (6.3%) | 24 (75%) | 32 (100%) |
TRIAD correctly predicted disposition of 70% of cases observed in Facility #1 in 1983 (conditional gamma = 0.62), as compared with 81% in 1981 (conditional gamma = 0.81).10 Eliminating patients retained voluntarily in 1983 (none of our sample were voluntarily, retained there in 1981), the 1983 proportion of correct predictions was 71% (gamma = 0.61). While the relationship between TRIAD Severity level and disposition continued to be impressive, the 11% drop in the accuracy of TRIAD for predicting disposition in the setting is noteworthy. The reasons for this change will be explored below.
As mentioned earlier, we expected that those patients scoring at the lowest level of Severity (Level 1) would be released after the evaluation. Further, we expected that disposition of cases at the middle levels (Levels 2 and 3) might be influenced by other contingencies and that those patients scoring at the highest level would be retained. Although disposition is admittedly an imperfect criterion for the concept “admissibility,” it is one of the only two available criteria, and deviation from these expectations raises questions about the reliability (and validity) of TRIAD. In 1983, there were more such discrepancies than in 1981. The decreased accuracy of TRIAD in 1983 reflects increases in the number of false positives (high-scorers released) and false negatives (low-scorers retained), as well as more frequent retention of patients in the moderate range of TRIAD Severity (Levels 2 and 3).
Patient characteristics and discrepant outcomes
Discrepant cases at Facility #1 were few in 1981; false positives were likely to have high scores on Danger to Self, to have nonpsychotic diagnoses, and to be (slightly) disproportionately white, while false negatives were involuntary at entry, psychotic, and female. In 1983 the same trends were observed, with the addition that high-scorers released were usually voluntary at entry. However, a more compelling explanation for the different decisions on similar cases over the two study periods arises from a look at changes in the environment of the Facility #1 PES.
Differences in Facility #1 PES environment
Housed separately from the medical emergency room, the psychiatric emergency service of Facility #1 is a medically oriented, locked unit in an urban county general hospital. During May and June 1983, the period of our second round of observations, 1,433 patients entered the PES (an average of 716 per month). The average hourly census of the PES was approximately 9 patients, with 2 psychiatrists and 1 social worker on duty to evaluate patients during the day shift and 1 psychiatrist on duty at night. An evaluating clinician was assigned sole responsibility for case disposition. Clinicians were supported by a nursing staff of 5 per shift. (Psychiatrists were also responsible for 106 consultations on medical wards during the 2-month period.)
The above description of Facility #1 in 1983 applies in all particulars to 1981 except for the patient load. In November and December 1981, the first period of study there, the service saw an average of 670 patients per month as compared with 716 per month in the 1983 study period. The modal number of patients checking into the PES on days of case observation for this study increased from 20 in 1981 to 24 in 1983, and the mean from 23.8 to 25.6. According to PES statistics, the average hourly census on the unit was 5.2 patients in 1981 as compared with 8.9 in 1983. As the level of staffing remained the same, the ratio of patients to clinicians increased.
Effect of patient load
That this increase in patient load at Facility #1 may account for the drop in percentage of dispositions correctly predicted by TRIAD is confirmed by controlling for the number of patients in the PES at the time of the case observation. A measure of patient load most likely to affect the decision-making environment is the total number of patients to whom staff must attend in the PES at the same time the study subject is being evaluated. The number of patients in PES jurisdiction at the time a case was observed ranged from 4 to 17. (In 1 case, the number was not obtained.) In 21 cases, the range was 4 to 10 (see table 3). TRIAD correctly predicted disposition for 81% of those 21 cases (conditional gamma = 0.92). This is the same percentage of correct predictions obtained in 1981, when Facility #1 was in general not as busy as in 1983.
TABLE 3.
Disposition of cases by TRIAD Severity level with 4–10 patients in PES at evaluation: Facility #1, 1983
| Severity level | Released | Retained | Total |
|---|---|---|---|
| Level 1 | 5 (100%) | 0 | 5 (100%) |
| Level 2 | 1 (50%) | 1 (50%) | 2 (100%) |
| Level 3 | 0 | 0 | 0 |
| Level 4 | 3 (21%) | 11 (79%) | 14 (100%) |
| Total | 9 (43%) | 12 (57%) | 21 (100%) |
For 29 cases in which 11 to 17 patients were in PES jurisdiction at the time of the observation (see table 4), TRIAD predicted only 66% of dispositions (conditional gamma = 0.36). Although no data were collected in 1981 on the number of patients in the PES at the time of the observation, it is reasonable to believe, given the lower average number of patients entering the day of observation and given the lower average hourly census, that the number was closer to the range in which TRIAD performed more successfully in 1983. The relationship between TRIAD Severity and disposition remains very strong and, when controlling for the number of patients in the PES, even improves as compared with 1981—that is, gamma = 0.81 (n = 31) in 1981 as compared with gamma = 0.92 (n = 21) in 1983.
TABLE 4.
Disposition of cases by TRIAD Severity level with 11–17 patients in PES at evaluation: Facility #1,1983
| Severity level | Released | Retained | Total |
| Level 1 | 2 (29%) | 5 (71%) | 7 (100%) |
| Level 2 | 2 (66.6%) | 1 (33.3%) | 3 (100%) |
| Level 3 | 0 | 1 (100%) | 1 (100%) |
| Level 4 | 3 (17%) | 15 (83%) | 18 (100%) |
| Total | 7 (24%) | 22 (76%) | 29 (100%) |
When the number of patients in the PES at Facility #1 was 10 or fewer, the number of false negatives fell to zero. However, the proportion of false positives rose slightly, indicating that discrepant releases are at least as likely to take place with fewer patients to be seen.
What clinical considerations might lead to retention of low-scoring patients when patient load is high? With 11 to 17 patients present, all psychotic patients were retained, whereas nonpsychotic patients had a 50% chance of release (gamma = 1.0). With fewer patients in the emergency room, the relationship of psychosis to disposition (gamma = 0.72) was less strong than the relationship of TRIAD Severity to disposition (gamma = 0.92). Moreover, the covariance of Severity and psychosis was greater when there were fewer patients (gamma = 0.2) than when more patients were present (gamma = 0.05) and overall (gamma = 0.02). All of the low-scoring patients retained had psychotic diagnoses.
Additional information is needed to account for the differential retention of low-scoring psychotic patients when the service is busiest. Most (57%) of the patients retained at Facility #1 were held for further observation in the PES rather than immediate transfer to an inpatient unit. Apparently, therefore, when cases back up in the PES, clinicians retain those patients who are more difficult to evaluate until such time as they can devote more attention to them. This explanation is supported by the finding that patients with low scores who were retained had not only psychotic diagnoses, but also relatively high levels of symptomatology.
While one interpretation of this phenomenon is that clinicians are applying a “need for treatment” standard rather than an acceptable California legal standard, another is that they simply find these psychotic, multisymptom patients impossible to evaluate rapidly and put off until later asking questions that would either yield higher TRIAD scores or convince them that the patients are not admissible. Indeed, of the 5 false negatives, 4 were retained in PES for further evaluation, and the 1 referred to an inpatient unit specifically requested admission. However, for those patients unwillingly retained, the difference between formal admission for observation and detention in PES may be very abstract indeed.
In sum, it appears that the drop in percentage of correct predictions by TRIAD at Facility #1 across time may be attributed primarily to the influence of patient load on the decision-making process. This conclusion is bolstered by the difference in the level of agreement between TRIAD scores and clinician global ratings (CGRs) when controlling for patient load (see tables 5 and 6). CGRs were obtained in 49 of the 51 cases observed at Facility #1.
TABLE 5.
Correlation coefficients for TRIAD scores and clinician global ratings of cases with 4–10 patients in PES at evaluation: Facility #1, 1983 (n = 20)
| TRIAD total | TRIAD DSS | TRIAD DOS | TRIAD GDS | |
|---|---|---|---|---|
| CGR total | 0.7 (p ≤ 0.0001) | |||
| CGR DSS | 0.74 (p ≤ 0.0001) | * | * | |
| CGR DOS | * | 0.63 (p ≤ 0.0001) | * | |
| CGR GDS | * | * | 0.79 (p ≤ 0.0001) |
Nonsignificant correlations ranging from −0.18 to 0.25.
TABLE 6.
Correlation coefficients for TRIAD scores and clinician global ratings of cases with 11–17 patients in PES at evaluation: Facility #1, 1983 (n = 29)
| TRIAD total | TRIAD DSS | TRIAD DOS | TRIAD GDS | |
|---|---|---|---|---|
| CGR total | 0.27 (n.s.) | |||
| CGR DSS | 0.48 (p ≤ 0.01) | * | † | |
| CGR DOS | * | 0.49 (p ≤ 0.01) | ‡ | |
| CGR GDS | * | * | 0.41 (p ≤ 0.05) |
Nonsignificant correlations ranging from −0.24 to 0.05.
Significant negative correlation.
Unexpected significant correlation.
When 11 to 17 patients were present in the PES (see table 6), the correlation between the TRIAD total and the CGR was nonsignificant (r = 0.27), but the correlation rose to r = 0.7 (p ≤ 0.0001) with 4 to 10 patients present (see table 5). With fewer patients in the PES, the correlation between TRIAD and the clinician measure of Danger to Self rose from r = 0.48 (p ≤ 0.01) to r = 0.74 (p ≤ 0.0001), between the two measures of Danger to Others from r = 0.49 to r = 0.63, and between the two measures of Grave Disability from r = 0.41 to r = 0.79. The differences in strength of agreement of the total scores and the Grave Disability scores over the two conditions are significant at the level of p ≤ 0.05. It would appear, therefore, that with fewer distractions and/or less pressure for time, clinicians base their ratings of patients on roughly those factors reflected in the TRIAD scores. When distractions and/or time pressure increase, the standard described by TRIAD is less likely to be applied.
TRIAD performance across settings
To assess the generalizability of TRIAD across settings, we compared TRIAD scores with clinician ratings and dispositions of cases observed at Facility #2 in 1983 with those obtained at Facility #1. In this and the following sections, we will report the results of this comparison. First, however, we will describe Facility #2.
TRIAD scores and CGRs in two emergency rooms, 1983
Housed in the medical emergency room of a suburban general hospital, the PES of Facility #2 was staffed by teams composed of 2 licensed psychiatric technicians (LPTs) and a physician. Each patient was seen by an LPT, who did the initial work-up and turned the case over to a physician with a report and recommendation for disposition. The patient was then evaluated by the physician as well. Frequently, the physician and LPT were actually in communication about the case from start to finish. Nursing care was provided by medical emergency room staff.
In the preceding section we noted that agreement between clinicians’ ratings and TRIAD scores at Facility #1, given a ratio of 3 evaluating clinicians and 4 to 10 patients, ranged from r = 0.63, Danger to Others, to r = 0.79, Grave Disability score. The agreement on total scores was r = 0.7 and on Danger to Self was r = 0.74. At Facility #2, independent CGRs were obtained from both the LPT (see table 7) and the physician (M.D.) (see table 8) on each case. The two sets of clinicians agreed with the TRIAD total score at the same level (r = 0.59), and the strength of agreement with TRIAD by each set of clinicians was similar on two of the scales: r = 0.67 and r = 0.7 on Danger to Others, and r = 0.7 and r = 0.71 on Grave Disability.
TABLE 7.
Correlation coefficients for TRIAD scores and psychiatric technician global ratings of cases: Facility #2, 1983 (n = 48)
| TRIAD total | TRIAD DSS | TRIAD DOS | TRIAD GDS | |
|---|---|---|---|---|
| Clinician (LPT) GR total |
0.59 (p ≤ 0.0001) | |||
| Clinician (LPT) GR DSS |
0.57 (p ≤ 0.0001) | * | * | |
| Clinician (LPT) GR DOS |
0.7 (p ≤ 0.0001) | † | ||
| Clinician (LPT) GR GDS |
* | * | 0.71 (p ≤ 0.0001) |
Nonsignificant correlations ranging from 0.04 to 0.19.
Significant moderate correlation (r = 0.37, p ≤ 0.01).
TABLE 8.
Correlation coefficients for TRIAD scores and physician global ratings of cases: Facility #2, 1983 (n = 50)
| TRIAD total | TRIAD DSS | TRIAD DOS | TRIAD GDS | |
|---|---|---|---|---|
| Clinician (M.D.) GR total |
0.59 (p ≤ 0.0001) | |||
| Clinician (M.D.) GR DSS |
0.47 (p ≤ 0.0001) | * | * | |
| Clinician (M.D.) GR DOS |
0.67 (p ≤ 0.0001) | * | ||
| Clinician (M.D.) GR GDS |
* | * | 0.7 (p ≤ 0.0001) |
Nonsignificant correlations ranging from 0.003 to 0.198.
However, on Danger to Self, physicians’ agreement with TRIAD was moderate (r = 0.47), while LPTs agreed slightly more strongly (r = 0.57). This variation in strength of agreement and the lower level of agreement on Danger to Self at Facility #2 as compared with Facility #1 is interesting in light of the fact that at Facility #2, the Danger to Self scale was highly predictive of disposition. While the clinicians at Facility #2 were inclined to give high ratings on Danger to Self less often than TRIAD, 100% of the patients at that facility who scored high on TRIAD’S Danger to Self scale were retained.
In sum, the validity of TRIAD as a reflection of clinicians’ assessments was supported in the two facilities by agreement with CGRs ranging from r = 0.59 to r = 0.7 on the total score, from r = 0.47 to r = 0.74 on the Danger to Self score, from r = 0.63 to r = 0.7 on the Danger to Others, and from r = 0.7 to r = 0.79 on Grave Disability.
TRIAD scores and disposition in two emergency rooms, 1983
Having compared agreement between CGRs and TRIAD in the two facilities, we will now compare disposition decisions and TRIAD scores of cases observed at Facility #2 in 1983 with the 1983 results at Facility #1.
TRIAD correctly predicted 86% of 50 disposition decisions (conditional gamma = 0.92) in the PES of Facility #2 in the period July through September 1983. Table 9 reflects the disposition of cases at each TRIAD Severity level. If cases of voluntary admission are excluded from the analysis, 91% of dispositions are correctly predicted by TRIAD, and the relative reduction of error attributable to TRIAD approaches 95% (gamma = 0.95). Only 3 cases at Facility #2 yielded dispositions clearly discrepant from those predicted by TRIAD—that is, Severity Level 1 patients retained and Severity Level 4 patients released.
TABLE 9.
Disposition of cases by Severity level: Facility #2, 1983 (n = 50)
| Severity level | Released | Retained voluntarily | Retained involuntarily | Total |
|---|---|---|---|---|
| Level 1 (DSS, DOS, GDS = 0 or 1; total ≤ 3) |
10 (83.3%) | 1 (8.3%) | 1 (8.3%) | 12 (100%) |
| Level 2 (DSS, DOS, GDS = 2; total = 2) |
1 (25%) | 1 (25%) | 2 (50%) | 4 (100%) |
| Level 3 (DSS, DOS, GDS = 2; total = 3) |
1 (100%) | 0 | 0 | 1 (100%) |
| Level 4 (DSS, DOS, GDS = 3 or 4 or total ≥ 4) |
2 (6.1%) | 3 (9.1%) | 28 (84.8%) | 33 (100%) |
Thus, TRIAD was a more successful predictor of disposition at Facility #2 than at Facility #1, where only 70% of 51 case dispositions (gamma = 0.62) were correctly predicted in 1983. A major difference between the two settings was patient load. As discussed above, patient load also represented a major difference between Facility #1 in 1983 and Facility #1 in 1981, when TRIAD Severity level predicted case disposition in 81% of 31 cases (conditional gamma = 0.81).
At Facility #2, the number of patients in the PES jurisdiction at the time of evaluation of study subjects ranged from 1 to 4. At Facility #1, the number ranged from 4 to 17. Despite this difference in patient load, the level of staffing at the two facilities was comparable with regard to numbers and experience of evaluating clinicians.
At Facility #1, a total of 10 psychiatrists and 1 social worker were observed, with PES experience ranging from 2 to 13 years. They worked in shifts of 1 psychiatrist in the evenings and 2 physicians and 1 social worker in the daytime. The length of experience of an evaluator in a particular case averaged 4.7 years. At Facility #2, 9 LPTs and 13 physicians were observed, working in shifts of 2 LPTs and 1 physician. In all observed cases, LPTs were experienced in psychiatric emergency settings (1 to 7 years), while in many cases (n = 24), the assigned physician was a new medical or, more often, psychiatric resident with no more than a few weeks of PES experience. The overall average length of experience of the LPT in a particular case was 4.8 years and of the physician, 3 years.
At Facility #1, the proportion of dispositions correctly predicted by TRIAD Severity rose from 70% of 50 cases overall to 81% of 21 cases (conditional gamma = 0.92) in which the patient load ranged from 4 to 10. Thus, given a comparable patient load, the percentage of correct predictions at Facility #1 in 1983 was equivalent to the percentage of correct predictions in 1981, and the strength of the relationship between TRIAD Severity and disposition at Facility #1 in 1983 was equivalent to that at Facility #2 in 1983 (gamma = 0.92).
As the greatest number of patients in the PES at Facility #2 was equivalent to the lowest number at Facility #1, it is not possible to compare TRIAD performance given the same number of patients at the two facilities. However, given the improvement in TRIAD performance at Facility #1 when patient load dropped, it is not unreasonable to speculate that if the number of patients at Facility #1 had dropped to Facility #2’s range (1 to 4), the rate of correct predictions of disposition based on TRIAD Severity would have further improved, possibly to the level of 86% attained at Facility #2.
From the above it is reasonable to conclude that the lower the patient–clinician ratio in PES jurisdiction, the more reliable is disposition as a criterion for the validity of TRIAD as a measure of the clinician’s construct “admissibility.” Given 3 evaluating clinicians and 1 to 10 patients, TRIAD Severity level predicted disposition in 81% to 86% of cases in both urban and suburban settings and in the same setting across time.
Given the conditions stated above, most dispositions that are not clearly predictable from TRIAD Severity involve high-scorers who are released or moderate-scorers who are retained. As moderate scores are assigned to those TRIAD patterns thought to represent the greatest ambiguity in relation to legal criteria, the 50–50 split in the number of patients released and held involuntarily is not unexpected. When additional moderate-scoring cases have been observed, it will be possible to say with some certainty whether the split results entirely from the ambiguity of the presenting patterns or whether some patterns are scored incorrectly, (vis-à-vis weights attributed to them by clinicians) on TRIAD.
The discrepant cases of greatest relevance to the validity of TRIAD are those in which high-scorers are released (false positives), because they continue to occur when patient load is low. However, many false positive cases involve patients with low levels of psychopathology, which apparently lead clinicians to doubt that they meet the second legal requirement for admission—that their dangerousness or disability be due to mental disorder. In most of these cases there is also a less restrictive setting available for treatment. Thus, release in some false positive cases may be consistent with legal requirements and social policy.
Caveat
The purpose of the study was to determine whether PES clinicians are employing a shared professional standard in their application of the legal criteria for involuntary admission. Our findings indicate that there is such a standard and that TRIAD roughly describes it. This provides some assurance that it is possible for patients in a psychiatric emergency room to be treated equitably in the matter of civil commitment. However, caution should be used in interpreting these results. The study focuses on only one of many processes for which the emergency room clinician is responsible—that is, the assessment of the patient’s dangerousness and disability, given the data which come to light.
Another process for which the clinician is responsible is the data gathering itself. The thoroughness and skill with which the data are elicited were not addressed by the study, although TRIAD does provide a necessary framework for such an evaluation. Given similar data as to the patient’s behavior and circumstances, clinicians apparently will make similar judgments. However, clinicians may vary in the skill with which they elicit data (or settings may vary in the degree to which they support thorough evaluation) to such an extent that cases which are in fact similar may be judged differently because relevant information never comes to light.
Other processes for which the clinician is responsible include accurate diagnosis, emergency treatment (whether psychological, social, or medical) that may reduce the need for hospitalization, and the evaluation of the patient’s response to emergency treatment. Variations in the adequacy of these processes, which may also interfere with equitable treatment of patients, were beyond the scope of this study.
Conclusion
The evaluation of dangerousness and disability due to mental disorder, part of the emergency civil commitment procedure, remains a controversial activity, questioned in the courts and even among the mental health professionals who participate in it. The most difficult challenge has been to the ability of those doing the evaluation to predict future dangerousness. Even though there is some evidence that this can be done for the short term,11 there is widespread skepticism about this process. Aside from skepticism about the validity of clinicians’ predictions of dangerous behavior or assessments of disability, there has been skepticism about the reliability of clinicians’ judgments—that is, the extent to which they apply the legal criteria equitably. Little effort has been made previously to document the reliability of clinical assessments of dangerousness and disability.
The results of the research reported here go a long way toward documenting agreement among San Francisco Bay Area clinicians on sets of indicators which can be consistently weighed in the evaluation of dangerousness and disability. However, intensified pressures on psychiatric emergency services may compromise the application of this standard and contribute in a measurable way to inequities in the evaluation of patients. The data reported here demonstrate that there is a limit to the number of patients to which even experienced and highly trained clinicians can respond in an equitable manner. Studies of additional emergency rooms may reveal other factors that impinge upon the equitable application of the legal criteria.
Acknowledgments
Research for this article was supported by NIMH grant MH37310-02 and by the University of California, Berkeley Campus Committee on Research.
Notes
- 1.Morse SJ, “A Preference for Liberty: The Case Against Involuntary Commitment of the Mentally Disordered,” C.L.R 70 (1982): 54–106; [Google Scholar]; Ennis BJ and Litwack TR, “Psychiatry and the Presumption of Expertise: Flipping Coins in the Courtroom,” C.L.R 62 (1974): 693–752; [Google Scholar]; Scheff TJ, “Two Studies of the Societal Reaction,” in Being Mentally III: A Sociological Theory, 2d ed. (New York: Aldine Publishing Company, 1984), ch. 6, pp. 90–113. [Google Scholar]
- 2.Morse, supra note 1; Ennis and Litwack, supra note 1.
- 3.Schwitzgebel RK, “Survey of Civil Commitment Statutes,” in Civil Commitment and Social Policy: On Evaluation of the Massachusetts Mental Health Reform Act of 1970, ed. McGarry AL, Schwitzgebel RK, Lipsitt PD, and Lelos D (Rockville, Md: Center for Studies of Crime and Delinquency, 1981), pp. 47–83. [Google Scholar]
- 4.Allen RH, Weinman M, Lorimor R, and Claghorn JL, “A Multi-tiered System for the Least Restrictive Setting,” Am. J. Psychiatry 137 (1980): 968–71; [DOI] [PubMed] [Google Scholar]; Baxter S, Chodoroff B, and Underhill R, “Psychiatric Emergencies: Dispositional Determinants and the Validity of the Decision to Admit,” Am. J. Psychiatry 124 (1968): 1542–46; [DOI] [PubMed] [Google Scholar]; Browning CH, Tyson RL, and Miller SI, “A Study of Psychiatric Emergencies: Part II. Suicide,” Psychiatry & Med. 1 (1970): 359–66; [DOI] [PubMed] [Google Scholar]; Feigelson EB, Davis EB, MacKinnon R, Shands HC, and Schwartz CC, “The Decision to Hospitalize,” Am. J. Psychiatry 135 (1978): 354–57; [DOI] [PubMed] [Google Scholar]; Hanson GD and Babigian HM, “Reasons for Hospitalization From a Psychiatric Emergency Service,” Psychiatric Q. 3 (1974): 336–51; [DOI] [PubMed] [Google Scholar]; Kirstein L, Prusoff B, Weismann M, and Dressier DM, “Utilization Review of Treatment for Suicide Attempts,” Am. J. Psychiatry 132 (1975): 22–27; [DOI] [PubMed] [Google Scholar]; Mendel WM and Rapport S, “Determinants of the Decision for Psychiatric Hospitalization,” Arch. Gen. Psychiatry 20 (1969): 321–28; [DOI] [PubMed] [Google Scholar]; Meyerson AT, Moss JZ, Belville R, and Smith H, “Influence of Experience on Major Clinical Decisions: Training Implications,” Arch. Gen. Psychiatry 36 (1979): 423–27; [DOI] [PubMed] [Google Scholar]; Rose SO, Hawkins J, and Apodaca L, “Decision to Admit,” Arch. Gen. Psychiatry 34 (1977): 418–21; [DOI] [PubMed] [Google Scholar]; Schwartz MD and Errera P, “Psychiatric Care in a General Hospital Emergency Room: II. Diagnostic Features,” Arch. Gen. Psychiatry 9 (1963): 113–21; [DOI] [PubMed] [Google Scholar]; Streiner DL, Goodman JT, and Woodward CA, “Correlates of the Hospitalization Decision: A Replicative Study,” Can. J. Pub. Health 66 (1975): 411–15; [PubMed] [Google Scholar]; Tischler B, “Decision-Making Processes in the Emergency Room,” Arch. Gen. Psychiatry 14 (1966): 69–78; [DOI] [PubMed] [Google Scholar]; Tyson RL, Miller SI, and Browning CH, “A Study of Psychiatric Emergencies: Part I. Demographic Data,” Psychiatry in Med. 1 (1970): 349–57; [DOI] [PubMed] [Google Scholar]; Wood EC, Rakusin JM, and Morse E, “Resident Psychiatrist in the Admitting Office,” Arch. Gen. Psychiatry 13 (1965): 54–61. [DOI] [PubMed] [Google Scholar]
- 5.Segal SP, Watson MA, and Nelson LS, “Indexing Civil Commitment in Psychiatric Emergency Rooms,” Annals (American Academy of Political and Social Sciences) 484 (1986): 56–69. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Cal. Welf. & Inst. Code § 5000.
- 7.Segal et al., supra note 5.
- 8. Id.
- 9.Our observation led us to believe that when a patient comes into the emergency room the clinician focuses his assessment on the area suggested by the patient’s major presenting behavioral problem. For example, a suicide threat will lead to an assessment of danger to self rather than disability or danger to others. These areas will be explored secondarily, as a result of information that comes to light in the assessment of danger to self. If the patient does not present a strong picture of admissibility on any one criterion, the overall picture (moderate and/or low-level presentation on more than one criterion) becomes salient for the disposition. In our analysis, therefore, we attended not only to the patient’s presentation on individual criteria, but also to the overall presentation.
- 10.Cases in which dispositions are said to be “correctly predicted” were cases in which patients scoring at Severity Level 4 were retained and patients scoring at Severity Levels 1 through 3 were released. We consider this approach satisfactory for summarizing our findings, although it does not quite reflect the complexity of our expectations for Severity Levels 2 and 3.
- 11.Monahan J, “Prediction Research and the Emergency Commitment of Dangerous Mentally Ill Persons: A Reconsideration,” Am. J. Psychiatry 135 (1978): 198–201; [DOI] [PubMed] [Google Scholar]; Skodol AE and Karasu TB, “Toward Hospitalization Criteria for Violent Patients,” Comprehensive Psychiatry 21 (1980): 162–66. [DOI] [PubMed] [Google Scholar]
