Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 29.
Published in final edited form as: Rehabil Psychol. 2024 Feb 15;69(4):326–334. doi: 10.1037/rep0000550

Initial development and psychometric properties of the Therapist Quality Scale

MA Day 1,2, LC Ward 1, DM Ehde 2, ME Mendoza 2, KM Phillips Reindel 3, BE Thorn 4, I Bindicsova 1, MP Jensen 2
PMCID: PMC13124244  NIHMSID: NIHMS2054795  PMID: 38358711

Abstract

Purpose/Objective.

This study sought to develop and evaluate the psychometric properties of a brief measure of therapist quality that is applicable for use across different types of psychosocial chronic pain treatments: the Therapist Quality Scale (TQS).

Research Method/Design.

The initial pool of 14 items was adapted from existing measures, with items selected that are relevant across interventions tested in a parent trial comparing mindfulness meditation, cognitive therapy, and behavioral activation for chronic back pain (N=302 participants) from which data for this study were obtained. A random selection of 25% of video-recorded sessions from each cohort was coded for therapist quality (2 randomly selected sessions per group), with 66 sessions included in the final analyses (n=33 completed pairs). Items were coded on a 7-point Likert-type scale. Exploratory Factor Analysis (EFA) and reliability estimates were generated.

Results.

Correlations between the 14 initial items rated for Session 1 and 2 were generally significant, suggesting reliability in ratings. EFA showed a single factor solution that provided a parsimonious explanation of the correlational structure for both sessions. Eight items with factor loadings of ≥.50 in both sessions were selected to form the TQS. Reliability analyses demonstrated that all items contributed to scale reliability, and internal consistency reliabilities were good (αs≥.87).

Conclusions/Implications.

The TQS provides a brief, reliable measure that is applicable for use across different types of treatments to rate the quality of therapist’s delivery. The items assess quality in delivering specific techniques, maintaining session structure, and in developing and maintaining therapeutic rapport.

Keywords: Therapist Quality Scale, fidelity, clinical trials, psychosocial treatment, chronic pain

Introduction

Current guidelines for chronic pain recommend non-pharmacological interventions as the first-line treatment approach (Dowell et al., 2016). Unlike medications, most non-pharmacological interventions involve a therapist (e.g., physiotherapist, psychologist, counselor, etc.) who works with and teaches individuals with chronic pain specific and, ideally, evidence-based pain management skills. While a critical component of evidence-based rehabilitation psychology practice is to deliver the treatment(s) shown to have efficacy, it is also important for the therapist to deliver the treatment in a competent manner (Sackett et al., 1996).

Treatment fidelity, in the context of clinical trials, includes the monitoring of therapist adherence to delivering core features of the treatment, as well as therapist competency or quality in the skillful delivery of the treatment (Proctor et al., 2011). Treatment fidelity ratings have been recognised as an important moderator of psychological treatment outcome for over 40 years (e.g., Muse & McManus, 2013; Young & Beck, 1980). Within the broader psychotherapy literature, research has shown that higher treatment fidelity consistently predicts more beneficial outcomes (e.g., Hogue et al., 2008; Schoenwald et al., 2008; Stirman et al., 2013; Strunk et al., 2010). Moreover, a meta-analysis concluded that therapist effects have a modest role in accounting for outcomes of clinical trials (i.e., d ~ 0.35) and large effects on naturalistic therapy outcomes (i.e., d ~ 0.55) (Baldwin & Imel, 2013). Recognising this body of evidence, the extension of the CONSORT statement to reporting randomized controlled trials (RCTs) of non-pharmacological interventions made it a requirement to include the provision of a standardized manual detailing the structure and content of therapy, as well as to detail how provider adherence to the manual was assessed (Boutron et al., 2008). The American Psychological Association’s 2018 Journal Article Reporting Standards (JARS) for Quantitative Research in Psychology extended this even further, and also requires that authors describe the methods for, and results of, assessing therapists’ adherence and competence/quality in delivering an intervention per protocol for clinical trials (Appelbaum et al., 2018).

A few measures have been developed and used to assess elements of treatment fidelity in clinical trials. One scale developed by Yates and colleagues rates the standard of clinical trials evaluating psychological pain interventions based on: (1) the existence of a treatment manual, (2) the reporting of adherence to that manual, and (3) the specification that the therapists were trained to provide the treatment based on the manual (Yates et al., 2005). However, while rigorous clinical trials increasingly report on adherence, therapist quality has rarely been directly evaluated in chronic pain research (Morley, 2011; Vlaeyen, 2005).

The dearth of measures evaluating, and trials reporting, therapist quality is likely attributable to both the complexity of the construct as well as to how cumbersome the assessment process is. It entails first determining the most important therapist quality domains, developing items to assess those domains, and then training coders to rate therapist quality using those items across treatments. The process is further complicated by the fact that non-pharmacological interventions include both unique components (e.g., teaching cognitive restructuring skills in many, but not all, cognitive-behavioral therapy interventions) and shared components (e.g., time management of treatment sessions, therapist empathy).

Some measures have been developed to assess the quality of the treatment session delivery for specific interventions, primarily cognitive behavioral, such as the Cognitive Therapy Rating Scale (Muse & McManus, 2013; Vallis et al., 1986; Young & Beck, 1980), and the Cognitive Therapy Adherence and Competence Scale (Barber, 2003). However, a significant limitation of these measures is that they are specific to just one intervention. As a result, they do not allow for the assessment of quality in the delivery of more than one treatment. Therefore, they cannot be used to compare the quality of treatment delivery between studies or evaluate the role of therapist quality in outcome across different treatment conditions in the same study.

One strategy for addressing this issue is to develop and use proxy measures for therapist quality, such as amount of therapist training and/or supervision provided, or the number of years of experience of the therapist(s). However, such proxy measures have not been found to be associated with treatment outcome (Resnik & Hart, 2003). Thus, in order to understand the role of therapist quality on outcome, there continues to be a need for a reliable and valid measure that can be used to assess quality across more than one psychosocial treatment.

Given these considerations, the aim of this work was to begin the process of developing a measure that could be used to assess therapist quality (i.e., the quality with which the clinician engaged in delivering different treatment procedures) and to evaluate the initial psychometric properties of this measure. The items needed to assess treatment adherence are fairly straightforward; they usually involve identifying the treatment components described in a treatment manual, and then asking raters to determine whether or not that treatment component was provided. As noted previously, however, assessing therapist quality is much more complex; developing such a measure is the primary focus of the current study.

To develop what we view as the first iteration of a measure, we began by adapting previously developed treatment fidelity rating forms (Barber, 2003; Day, 2017) to make them applicable to assess therapist quality across three theoretically unique chronic pain treatments that were evaluated in a recent RCT (Day et al., 2020): cognitive therapy (CT), behavioral activation (BA), and mindfulness meditation (MM). The rating forms created contained three distinct sections with associated quality items measuring: (1) quality in teaching or encouraging specific pain management techniques, (2) quality in providing structure to the treatment sessions, and (3) quality with respect to developing a therapeutic relationship (e.g., common factors such as being empathetic, being genuine, etc.). Items were written such that the same therapist quality domain could be assessed across all three treatments investigated in the trial from which the data for the current study were generated (i.e., CT, BA, MM). In addition, detailed coding manuals were developed to make the criteria for each rating as clear as possible (i.e., based on observable therapist behaviors; see online supplementary material). Finally, the individuals who rated quality were trained to a high standard of inter-rater-reliability with experts. This study reports the initial factor structure and reliability of this measure.

Methods

Study Design

The data for this measurement study came from a 3-group parallel (1:1:1), single-blind RCT (Day et al., 2020). The sample was 302 individuals with chronic pain, with low back pain experienced as a primary or secondary pain condition for at least 6-months duration. Data collection took place between October, 2018 and March, 2022. Participants were randomly assigned to eight, 1.5-hour, group-delivered Zoom videoconference sessions of CT, BA or MM with thirteen cohorts in total; all sessions were video-recorded. A random selection of 25% of sessions from each cohort were selected for coding for quality rating purposes (2 randomly selected sessions per group), for a total of 78 sessions coded. This research was approved by the Human Participants Review board at the University of Washington and the study sponsor. Informed consent was obtained from all participants. The parent trial was pre-registered on clinicaltrials.gov (Identifier: NCT03687762). See the trial protocol paper for details on the inclusion/exclusion criteria and other trial procedures (Day et al., 2020).

Item pool

The items analysed in this research are originally from the Mindfulness-Based CT-Adherence Appropriateness and Quality Scale (MBCT-AAQS; Day, 2017), which was originally adapted from the CT Adherence and Competence Scale (CTACS; Barber, 2003) with input from the MBCT-Adherence Scale (Segal, Teasdale, Williams, and Gemar, 2002). The rationale for using the items from the MBCT-AAQS and not the more widely used CT Scale (Young & Beck, 1980) was based on Whisman’s critique of the CT Scale, which highlighted how multiple concepts are addressed by one item in the CT Scale (Whisman, 1993); the items from the MBCT-AAQS generally assess single concepts. The MBCT-AAQS contains 29 items grouped into three rationally defined subscales (i.e., not developed via statistically defined factor loadings, as is also the case for the CTACS): (1) specific (mindfulness-based cognitive therapy) techniques; (2) therapy session structure; and (3) development of a collaborative therapeutic relationship. The item responses are based on a 7-point Likert-type scale ranging from 0 (“Poor”) to 6 (“Excellent) for quality. The anchor descriptors for the quality items vary across each item. For any items reflecting an intervention or process that did not occur in that session, a “not applicable” rating for quality is allowed.

Two items from the “specific techniques” subscale of the MBCT-AAQS were retained and used in the initial item pool used in this study: (1) guided discovery; and (2) commitment to practice. Four items from the “session structure” scale of the MBCT-AAQS were selected and pertain to: (1) agenda setting/orientation to session; (2) elicitation of participant feedback and questions throughout the session (modified for “group participant...” herein); (3) maintaining session focus/structure; and (4) session summary. One new item was created in the current research and included in the initial item pool: (5) time management throughout the session. In the MBCT-AAQS, time management was subsumed within the focus/structure item, but in the CT Scale “time management” is assessed using a separate item. We adopted this latter approach for the initial item pool used here. The “development of a collaborative relationship” items from the MBCT-AAQS that were included in the initial item pool were: (1) socialization to the treatment model; (2) warmth/genuineness/congruence; (3) acceptance/respect; (4) attentiveness; (5) accurate empathy; (6) collaboration; and (7) encouragement of group cohesion.1

Because the overwhelming majority of the items for the scale examined here are from the MBCT-AAQS, while at the same time those items were selected to be applicable to a variety of chronic pain treatments (i.e., not just MBCT) and that the focus is on assessing quality only (i.e., adherence is rated per specific manuals with different techniques and appropriateness does not apply across all items – see above footnote), we called the scale that is being validated here the Therapist Quality Scale (TQS). As “quality” is a subjective evaluation, in this research we developed detailed coding manuals for each condition and extended descriptors for the anchor descriptions (see Supplementary material for the TQS Rating Form and TQS Coding Manual).

Therapists and session raters

The five clinical therapists in this trail all had PhD degrees in clinical psychology, had experience in working with individuals with chronic pain, were trained and supervised to deliver all three treatments, and were also trained to code the therapy sessions for quality ratings using the TQS. They did not rate their own sessions. In total, five therapists conducted the treatment sessions, although a single therapist provided therapy to nearly half of the participants for which quality ratings were obtained, delivering 18 of 39 coded sessions as this therapist provided more therapy sessions in the trial than the other therapists. The other therapists delivered 7, 6, 5, or 3 sessions. Four of the therapists coded the sessions (excluding their own). The duration for completing coding for each session was generally the same as the duration of the sessions in the trial (1.5-hours), with ratings of items completed while reviewing the session.

Prior to commencing coding, each therapist was trained by Dr. Day with 6-hours of initial instruction on how to use the TQS rating forms for the three treatments. They then coded two therapy sessions for each condition (i.e., six sessions in all) for training purposes; these same sessions were also coded by experts in that treatment condition (Ehde for CT; Jensen for BA; Day for MM). The sessions used for coder training were not included in the randomized selection of sessions coded for fidelity reporting purposes. After this, Dr. Day calculated two-way mixed effects intraclass correlation coefficients (ICC) using absolute agreement definition for expert/rater coding reliability for each condition. If the ICC was ≥.90 for one or all three treatment conditions, then the coders commenced rating the sessions for those conditions to be included in the final quality ratings. However, if the ICC was <.90 for any condition, Dr. Day provided an additional 1-hour training session during which she provided feedback on areas where the coder’s ratings diverged from the expert’s ratings and discussed the rationale for ratings given. The raters then coded two additional therapy sessions for the condition(s) needed; these sessions were also coded by the experts. After this second round of coding/training, the ICC reliability criteria was met by all therapists/raters. The four raters then conducted ratings for 12, 10, 8 or 3 (33 of 39) pairs of therapy sessions. Data from the six pairs of sessions with different raters for the two sessions were excluded from the present analyses; so, analyses were conducted on TQS ratings from 66 sessions total, or 33 pairs of rated sessions (i.e., where the same coder rated both sessions in a given cohort which was required for the planned reliability analyses).

Data analysis

To examine the correlational structure of TQS item scores, we conducted exploratory factor analyses (EFAs) on intra-session item correlations from the first and second sessions (n = 33 completed pairs) of quality ratings for the items that are shared across all three treatments. The number of factors was determined by Mplus significance testing, and scales were constructed only when factors from the two sessions were significantly (p < .05) congruent (according to factor loading correlations). Items were retained for a scale if correlations were ≥ .50 and statistically significant (p < .05) on corresponding factors from both sessions.

Mplus (version 8.6) was used for the EFA (Muthén & Muthén, 2017), and SPSS (version 27) was used for all other analyses, including alpha coefficients and descriptive statistics. Models were estimated using robust maximum likelihood (MLR), which adjusts standard errors for violations of normality. Estimation with MLR uses all of the data so that cases were not lost because of missing values.

Results

Table 1 shows the means and SDs of the 14 TQS ratings for the two sessions and also provides correlations between ratings from the two sessions and the results of paired-difference t-tests. As indicated by the sample sizes, there were very few missing values. Mean ratings were above 4.00 (i.e., indicating very high quality), except in one instance, and SDs ranged from 0.55 to 1.17.

Table 1.

Means, SDs, and Correlations of Ratings from Sessions 1 and 2 with Paired-Difference t-tests.


Session 1 Session 2

Rating N M SD M SD r p t p

Agenda setting 32 4.34 0.65 4.22 0.91 .36 .043 0.78 .442
Guided discovery 33 3.82 0.77 4.00 0.79 .57 .001 −1.44 .160
Group participant feedback 33 4.27 0.67 4.45 0.79 .46 .007 −1.36 .184
Focus/structure 33 4.45 0.67 4.27 0.67 .27 .125 1.29 .206
Session summary 33 4.09 1.07 4.18 0.98 .22 .216 −0.41 .687
Time management 33 4.39 1.17 4.39 1.00 .24 .183 0.00 1.000
Socialization to treatment model 33 4.21 0.74 4.24 0.87 .55 .001 −0.23 .823
Warmth/genuineness/congruence 33 4.39 0.56 4.64 0.65 .24 .188 −1.85 .073
Acceptance/respect 33 4.48 0.67 4.45 0.67 .37 .034 −0.47 .645
Attentiveness 33 4.36 0.70 4.36 0.55 .30 .094 0.00 1.000
Accurate empathy 33 4.33 0.98 4.39 0.66 .59 <.001 −0.53 .601
Collaboration 33 4.27 0.98 4.48 0.83 .45 .009 −1.27 .214
Encouragement of group cohesion 30 4.20 0.96 4.40 0.93 .45 .013 −1.10 .281
Commitment to practice 33 3.79 0.78 4.15 0.71 .28 .109 −2.33 .026

Correlations between corresponding TQS ratings from the two sessions were statistically significant (p < .05) for 8 of 14 comparisons, but the coefficients were generally low (as low as .22); none exceeded .59. The paired-difference t-tests in Table 1 revealed only one statistically significant (p < .05) mean difference between the two ratings of the item related to eliciting from participants a “commitment to practice.”

EFAs of the 14 TQS ratings were conducted separately for two rated sessions, and standardized factor loadings for the two sessions are shown in Table 2. A single factor solution was found to provide a reasonably parsimonious explanation of the correlational structure for both sessions. For the first session rated (Session 1), the single factor accounted for 32% of the total variance in the quality ratings, and an additional factor did not produce a significant increment in covariance explained (p = .123). For the second session rated (Session 2), a single factor accounted for 35% of the total variance, and a two-factor solution failed to converge. Factor loadings for Session 1 ranged from .00 to .78 and were statistically significant (p < .05) for 10 of 14 ratings. Loadings for Session 2 ratings ranged from .23 to .89, and 11 of 14 ratings were statistically significant. The correlation of r = .61 between loadings indicated a degree of congruence in factor patterns from the two sessions.

Table 2.

Standardized loadings from Factor Analyses with Robust Maximum Likelihood Estimation and Correlations of Ratings from Sessions 1 and 2.

Item Session 1 Correlation Sessions 1 & 2

Guided discovery1 .74* .57*
Commitment to practice .15 .28
Agenda setting .26* .36*
Group participant feedback1 .66* .46*
Focus/structure .46* .27
Session summary .43* .22
Time management .24* .24
Socialization to treatment model .54* .55*
Warmth/genuineness/congruence1 .71* .24
Acceptance/respect1 .73* .37*
Attentiveness1 .61* .30
Accurate empathy1 .66* .59*
Collaboration1 .64* .45*
Encouragement of group cohesion1 .72* .45*

Note. Correlations are from Table 1.

1

Item selected for final Therapy Quality Scale.

*

p < .05 for factor loading or correlation.

Based on the results of the two EFAs, the eight items with factor loadings of .50 or greater in both sessions were selected to form a global measure of therapist quality that was computed as the average of the eight ratings, with at least one item included from each of the theorized three quality domains. Mean scores for the TQS for Session 1 (M = 4.27, SD = 0.56) and Session 2 (M = 4.41, SD = 0.53) did not differ significantly, t(32) = −1.65, p = .108. The overall TQS scores were significantly correlated (r = .59, p < .001). Reliability analyses demonstrated that all items contributed to scale reliability in the data from each session, and internal consistency reliabilities were good for the ratings of both Session 1 (α = .88) and Session 2 (α = .87). After correction for attenuation due to unreliability, TQS scores from the two sessions had a correlation of r = .67. Thus, Session 1 and Session 2 quality ratings shared about 45% of their reliable variances.

Analyses of variance (ANOVAs) were conducted with repeated measures on Session (1 versus 2) to evaluate possible effects of treatment condition, therapist, or rater on the scale scores. An ANOVA with treatment condition as a between-subject comparison revealed no significant differences in scale scores, F(2, 30) = 0.45, p = .643, and session ratings were not significantly different, F(1, 30) = 2.56, p = .120. The Treatment Condition X Session interaction was also not statistically significant, F(2, 30) = 0.16, p = .850. A second repeated-measures ANOVA resulted in a non-significant effect for therapist, F(4, 28) = 0.48, p = .754, and the Therapist X Session interaction was not statistically significant, F(4, 28) = 1.37, p = .268. Finally, TQS scores did not differ significantly as a function of rater, F(3, 29) = 0.25, p = .864, and the Rater X Session interaction was also not statistically significant, F(3, 29) = 0.48, p = .697.

Discussion

In evidence-based rehabilitation psychology practice, a recognised critical element is that the therapist be able to deliver a given empirically supported treatment with competence/quality (Sackett et al., 1996). However, a significant gap in the field in general, and in the study of chronic pain treatment in particular, has been that in generating evidence to evaluate treatment efficacy, the quality of the therapeutic interactions during the treatment sessions has been assumed, yet rarely empirically confirmed. As a result, in most published clinical trials, it is unclear if the treatments were delivered as intended and with adequately high quality. If any of these treatments was not provided with high quality, it is possible that the effect sizes of so-called empirically supported treatments may be underestimated.

The brief, global scale of therapist quality developed in this study – the TQS – was intended to be an initial step towards equipping researchers with a practical measure that can be readily adopted to address this challenge in future comparative trials. Existing fidelity measures in the broad psychotherapy literature have encompassed three domains of therapist quality (Barber, 2003; Day, 2017), and our initial item pool contained items intended to assess each of these three domains: (1) delivering specific psychological techniques, (2) maintaining session structure, and (3) establishing and maintaining a collaborative relationship via harnessing common therapeutic factors. Although the third domain was the aspect of therapist quality most represented in the final global measure, the TQS does successfully contain items assessing each of these three domains.

Only one of the two hypothesized “specific technique” items included in the initial pool demonstrated adequate reliability in the current analyses; this item assesses therapist quality in engaging participants/clients in a guided discovery. Guided discovery here involves the therapist using a balance of open-ended questions, reflective, confrontive, and interpretive responses to guide clients to explore and deepen their understanding of important issues (e.g., their understanding of the accuracy of exaggerated thoughts in cognitive therapy, for example) and to link this understanding to the theory underlying the intervention being delivered. The results showed that ratings of therapists using this technique were only marginally reliable across sessions. However, such guided discovery is an important feature of Socratic Dialogue, which has been described as a cornerstone technique in cognitive-behavioral therapies, and a core competency for therapists (Clark & Egan, 2015). Within the depression literature, ratings of therapist use of Socratic questioning has been shown to predict next-session symptom improvement, even while controlling for therapeutic alliance (Braun et al., 2015). This guided discovery item having loaded on the TQS is consistent with this prior research supporting its importance. That it was only marginally reliable across the sessions, however, suggests that possibly therapists are inconsistent with being able to facilitate this guided discovery and might be more effective some days vs. other days, or that it may be related to factors other than therapist skill, such as client insight levels or group size or dynamics. Another possibility is that the item assessing this domain is difficult to code reliably. More research, perhaps with additional items developed to assess this therapist skill, is needed to help determine if it is possible to assess this domain more reliably.

The other item we had hypothesized to tap the “specific technique” domain pertained to therapists motivating and eliciting from clients a commitment to between-session skills practice. However, the ratings for this item were not reliable across sessions, and it did not load onto the TQS. As we retrospectively considered the possible reasons for this finding, we examined the treatment manuals carefully, and determined that they did not explicitly emphasize this element to a great depth; rather, we had assumed that the therapists would engage in this therapist behavior during homework assignment and review. This may be an example of where therapists have “adhered” to the manual and noted the need to commit to daily skills practice, and yet they may not have reliably elaborated upon this to enhance motivation and commitment. An alternative possible explanation is that because treatment was group delivered in the parent RCT, this may have made it more challenging for therapists to elicit commitment to skills practice from each individual participant. However, given that client’s use of “commitment language” in studies of Motivational Interviewing (MI) in the broad psychotherapy literature has been consistently linked to subsequent behavioral change (i.e., in this instance, skills practice; Miller & Rose, 2009), future trials would need to ensure that this therapy skill is not only taught, but that the manuals provide more explicit guidance regarding when and how to use this skill, as well as how often to use this skill.

Of the five items included in the initial item pool hypothesized to assess therapist quality in maintaining “session structure,” the results showed only one was reliable and loaded onto the final scale. This item assesses therapist quality in eliciting feedback and questions from clients throughout the session, making appropriate adjustments to the session on the basis of this feedback, and responding to questions in a nuanced, supportive manner. Interestingly, this is the only item that taps the importance of interactive reciprocity between therapists and clients, while maintaining flow/structure in the therapeutic process throughout a session. Each of the other four items that did not load onto the final global scale pertained mostly to therapist behavior occurring in relative isolation of client interaction (i.e., agenda setting, time management, maintaining focus and providing a session summary). Possibly, these latter session structure components of therapy might be best assessed for therapist adherence alone in clinical trials given their non-interactive nature.

The third domain that the initial item pool was hypothesized to correspond to was the “development of a collaborative relationship.” Given this domain encapsulates common factors of therapy, it was relatively intuitive that this lends itself well to developing a measure that is applicable to assess therapist quality across a range of treatments, as was the objective of this research. Thus, the largest number of initial items were hypothesized to map on to this domain. Of the seven items in the initial pool, six loaded onto the final TQS, including therapist quality in demonstrating warmth/genuineness/congruence, acceptance/respect, attentiveness, accurate empathy, collaboration, and active encouragement of group cohesion. The only item that did not load onto the final global scale pertained to therapist’s socializing clients to the treatment model. A possible reason for this may be that this item represents therapists engaging in the more psychoeducation aspects of the therapy sessions, which may require less technical skill to deliver with quality. This again highlights that therapeutic processes that entail reciprocal interaction may be a distinguishing aspect of therapist quality, including in relation to these common factors. That six items pertaining to common factors loaded onto the final TQS can be considered a strength of the measure, as it has been suggested that these factors should not be considered in isolation because they likely operate in a synergistic manner with specific techniques to contribute to therapeutic effectiveness (Appelbaum, 1978; Weinberger, 1995).

The importance of common factors in therapy has been underscored by both debate regarding “primacy” and also a plethora of empirical evidence that undeniably supports their role in influencing outcome. Carl Rogers argued common factors are the sine qua non of psychotherapy and are both necessary and sufficient for therapeutic progress (Rogers, 1957). Others have argued that while necessary, they may not be sufficient and that specific techniques in an action-oriented approach are required (e.g., Wampold, 2015). Either way, we would agree with John Mill, who proposed that such common factors may be primary in the therapeutic process because they usually precede anything else (Mill, 1881). Beyond the primacy debate however, research has consistently demonstrated the importance of common factors in accounting for positive outcomes (Bennett, 2011; Day et al., 2016; Knuuttila, 2012; Lambert, 2002). We found in secondary analyses of a prior trial that tested both theoretically-derived mechanisms (e.g., pain catastrophizing, pain control beliefs, mindfulness) as well as common factors that working alliance was the only mechanism significantly associated with improved pain intensity during cognitive therapy, mindfulness meditation and mindfulness-based cognitive therapy (Day et al., 2020). However, this emerging body of research in the field of chronic pain that examines such common factors has tended to focus on client self-reported variables, such as client ratings of working alliance using the Working Alliance Inventory (Hatcher & Gillaspy, 2006). The global TQS measure developed in this research may extend this area of study by taking into consideration directly observed therapist quality indicators.

Limitations

In this trial, therapists were selected who were PhD-level psychologists with experience working with chronic pain, who also received further intensive training and supervision. As such, all were rated as delivering the treatments with a very high level of quality, which limited the variability in the quality scores. While this is desirable in a clinical trial (i.e., as it confirms that treatment was delivered with competence and results will be a true representation of the respective treatment effects), it precluded investigation of validity of the measure; for example, to determine how quality ratings might be associated with treatment outcome. Such validity analyses would be possible in a study that sought to examine the role of therapist quality in outcome rather than to evaluate efficacy, specifically, by training therapists to vary quality (i.e., purposely provide treatment with more or less quality, depending on the condition). Alternatively, one could include therapists in the trial who might be expected to vary naturally in quality (e.g., different professional backgrounds, levels of training etc.). As noted earlier, a further limitation was that although this was a large RCT with >300 participants, treatment was delivered as group therapy. As a result, there were fewer sessions than there were study participants. This resulted in relatively few pairs of session ratings – just 33 pairs. This limited the power for conducting the planned factor analyses. It also limits the findings to therapists’ quality for group therapy. Future research using the TQS would ideally allow for a larger number of ratings and, additionally, for evaluating its psychometrics in individually delivered (1:1) therapy.

Despite the study’s limitations, the findings indicate that the TQS is a reliable measure that can be used to assess therapist quality in harnessing both specific and common therapeutic factors across different types of chronic pain treatments. This increases the opportunity for researchers to be able to study therapist quality between studies and allows for the monitoring of therapist quality during a clinical trial that compares different active chronic pain treatments. We do not consider therapist quality to be a stable trait, but it can vary, and also with supervision and feedback, it can be improved. The TQS ratings could be used to inform rehabilitation psychologist’s training (i.e., for the potential need for up-skilling certain aspects of delivery by taking a MI training workshop, for example) and also for such supervision and feedback purposes. Such a data-driven approach to training and supervision could therefore advance therapist skills in delivering specific psychotherapy techniques for chronic pain management, in maintaining session structure, and in developing and maintaining therapeutic rapport. Given the importance of therapist quality for outcome, being able to monitor quality (and correct as needed) could increase the rigor of clinical trials. Further, it would afford the capacity to more precisely determine the effect sizes of psychological treatments, when they are delivered both as intended, and with quality across sessions.

Supplementary Material

supplemental material

Figure 1.

Figure 1.

Results from application of the latent variable multiple regression model demonstrating that Quality ratings did not vary as a function of rater, therapist, or treatment condition.

Impact.

  • The quality in which various psychosocial treatments are delivered has been shown to predict outcome improvement in the broader psychotherapy literature; however, therapist quality is rarely assessed in clinical trials.

  • The Therapist Quality Scale (TQS) is the first validated measure with demonstrated reliability that can be used to assess therapist quality in the delivery of different types of psychosocial chronic pain treatments, thereby providing a tool to assess the role of quality in accounting for outcome change in comparative trials and naturalistic settings.

  • The TQS could also be used for therapist training and supervision purposes, to advance therapist skills in delivering specific psychotherapy techniques for chronic pain management, in maintaining session structure, and in developing and maintaining therapeutic rapport.

Acknowledgments

Research reported in this manuscript was supported by the National Center for Complementary and Integrative Health of the National Institutes of Health under the Award Number 1 R01 AT008559–01A1. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. The authors have no conflicts of interest to report.

Footnotes

1

An additional adaptation was made to this scale in the parent trial, in that only quality ratings were made for these items and not adherence, as “adhering” to empathy for example, is not logically assessable.

References

  1. Appelbaum M, Cooper H, Kline RB, Mayo-Wilson E, Nezu AM, Rao SM (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1):3–25. Erratum in: American Psychologist 2018 Oct;73(7):947. [DOI] [PubMed] [Google Scholar]
  2. Appelbaum SA (1978). Pathways to change in psychoanalytic therapy. Bulletin of the Menninger Clinic, 42, 239–251. [PubMed] [Google Scholar]
  3. Baldwin SA, & Imel ZE (2013). Therapist effects: finding and methods. In Lambert MJ (Ed.), Bergin and Garfield’s handbook of psychotherapy and behavior change (pp. 258–297). Wiley. [Google Scholar]
  4. Barber JP, Liese BS, Abrams MJ (2003). Development of the cognitive therapy adherence and competence scale. Psychotherapy Research, 13(2), 205–221. [Google Scholar]
  5. Bennett JK, Fuertes JN, Kietel M, Phillips R (2011). The role of patient attachment and working alliance on patient adherence, satisfaction, and health related quality of life in lupus treatment. Patient Education and Training, 85, 53–59. [DOI] [PubMed] [Google Scholar]
  6. Boutron I, Moher D, Altman DG, Schulz KF, & Ravaud P (2008). Extending the CONSORT Statement to randomized trials of nonpharmacologic treatment: Explanation and elaboration. Annals of Internal Medicine, 148(4), 295–309. [DOI] [PubMed] [Google Scholar]
  7. Braun JD, Strunk DR, Sasso KE, & Cooper AA (2015). Therapist use of Socratic questioning predicts session-to-session symptom change in cognitive therapy for depression. Behaviour Research and Therapy, 70, 32–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Clark GI, & Egan SJ (2015). The Socratic Method in cognitive behavioural therapy: A narrative review. Cognitive Therapy & Research, 39, 863–879. [Google Scholar]
  9. Day MA (2017). Mindfulness-Based Cognitive Therapy for Chronic Pain: A Clinical Manual and Guide. Wiley. [Google Scholar]
  10. Day MA, Ehde DM, Burns J, Ward LC, Friedly JL, Thorn BE, Ciol MA, Mendoza E, Chan JF, Battalio SL, Borckardt J, & Jensen MP (2020). A randomized trial to examine the mechanisms of cognitive, behavioral and mindfulness-based psychosocial treatments for chronic pain: Study protocol. Contemporary Clinical Trials, 93. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Day MA, Halpin J, & Thorn BE (2016). An empirical examination of the role of common factors of therapy during a mindfulness-based cognitive therapy intervention for headache pain. Clinical Journal of Pain, 32(5), 420–427. [DOI] [PubMed] [Google Scholar]
  12. Day MA, Ward LC, Thorn BE, Burns J, Ehde DM, Barnier AJ, Mattingley JB, Jensen MP (2020). Mechanisms of mindfulness meditation, cognitive therapy, and mindfulness-based cognitive therapy for chronic low back pain. Clinical Journal of Pain, 36(10), 740–749. [DOI] [PubMed] [Google Scholar]
  13. Dowell D, Haegerich TM, & Chou R (2016). CDC guideline for prescribing opioids for chronic pain—United States, 2016. JAMA, 315(15), 1624–1645. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Hatcher RL, & Gillaspy A (2006). Development and validation of a revised short version of the Working Alliance Inventory. Psychotherapy Research, 16, 12–25. [Google Scholar]
  15. Hogue A, Henderson CE, Dauber S, Barajas PC, Fried A, & Liddle HA (2008). Treatment adherence, competence, and outcome in individual and family therapy for adolescent behavior problems. Journal of Consulting and Clinical Psychology, 76, 544–555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Knuuttila V, Kuusisto K, Saarnioi P, Numm T (2012). Early working alliance in outpatient substance abuse treatment: Predicting substance use frequency and client satisfaction. Clinical Psychologist, 16(3), 123–135. [Google Scholar]
  17. Lambert MJ, Barley DE (2002). Research summary on the therapeutic relationship and psychotherapy outcome. In Norcross JC (Ed.), Psychotherapy relationships that work: Therapist contributions and responsiveness to patients (pp. 17–36). Oxford University Press. [Google Scholar]
  18. Mill JS (1881). A system of logic (8th Ed.). Harper. [Google Scholar]
  19. Miller WR, & Rose GS (2009). Toward a theory of motivational interviewing. American Psychologist, 64(6), 527–537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Morley S (2011). Efficacy and effectiveness of cognitive behaviour therapy for chronic pain: Progress and some challenges. Pain, 152(3 Suppl), S99–106. [DOI] [PubMed] [Google Scholar]
  21. Muse K, & McManus F (2013). A systematic review of methods for assessing competence in cognitive–behavioural therapy. Clinical Psychology Review, 33(3), 484–499. [DOI] [PubMed] [Google Scholar]
  22. Muthén LK, & Muthén BO (2017). Mplus users guide (Version 8). Muthén & Muthén. [Google Scholar]
  23. Proctor E, Silmere H, Raghavan R, Hovmand P, Aarons G, Bunger A, Griffey R, & Hensley M (2011). Outcomes for implementation research: Conceptual distinctions, measurement challenges and research agenda. Administration and Policy in Mental Health and Mental Health Services Research, 38, 65–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Resnik L, & Hart DL (2003). Using clinical outcomes to identify expert physical therapists. Phys Ther, 83, 990–1002. [PubMed] [Google Scholar]
  25. Rogers CR (1957). The necessary and sufficient conditions of therapeutic personality change. J Consult Psychol, 21(2), 95–103. [DOI] [PubMed] [Google Scholar]
  26. Sackett DL, Rosenberg WM, Gray JA, Haynes RB, & Richardson WS (1996). Evidence based medicine: what it is and what it isn’t. British Medical Journal, 312(7023), 71–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Schoenwald SK, Carter RE, Chapman JE, & Sheidow AJ (2008). Therapist adherence and organizational effects on change in youth behavior problems one year after multisystemic therapy. Administration and Policy in Mental Health and Mental Health Services Research 35, 379–394. [DOI] [PubMed] [Google Scholar]
  28. Segal ZV, Teasdale JD, Williams JM, Gemer MC (2002). The mindfulness-based cognitive therapy adherence scale: inter-rater reliability, adherence to protocol and treatment distinctiveness. Clinical Psychology & Psychotherapy, 9, 131–138. [Google Scholar]
  29. Stirman SW, Miller CJ, Toder K, & Calloway A (2013). Development of a framework and coding system for modifications and adaptations of evidence-based interventions. Implementation Science, 8, 1–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Strunk DR, Brotman MA, & DeRubeis RJ (2010). The process of change in cognitive therapy for depression: Predictors of early inter-session symptom gains. Behaviour Research and Therapy, 48, 599–606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Vallis TM, Shaw BF, & Dobson KS (1986). The cognitive therapy scale: Psychometric properties. Journal of Consulting and Clinical Psychology, 54, 381–385. [DOI] [PubMed] [Google Scholar]
  32. Vlaeyen JW, Morley S (2005). Cognitive-behavioral treatments for chronic pain: What works for whom? Clinical Journal of Pain, 21(1), 1–8. [DOI] [PubMed] [Google Scholar]
  33. Wampold BE (2015). How Important are the Common Factors in Psychotherapy? An Update. World Psychiatry, 14(3), 270–277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Weinberger J (1995). Common factors aren’t so common: The common factors dilemma. Clinical Psychology, 2(1), 45–69. [Google Scholar]
  35. Whisman MA (1993). Mediators and moderators of change in cognitive therapy of depression. Psychological Bulletin, 114, 248–265. [DOI] [PubMed] [Google Scholar]
  36. Yates SL, Morley S, Eccleston C, & de CWAC (2005). A scale for rating the quality of psychological trials for pain. Pain, 117(3), 314–325. [DOI] [PubMed] [Google Scholar]
  37. Young J, & Beck A (1980). Cognitive therapy scale: Rating manual. Philadelphia: Center for Cognitive Therapy. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

supplemental material

RESOURCES