Abstract
While it is known that therapists vary in effectiveness, it is unclear what therapist-level characteristics predict this variation. We conducted a large-scale, preregistered study (n = 97 therapists from the United States and Canada, n = 6,152 patients) examining a multimodal set of 38 therapist-level predictors that have been empirically or theoretically linked with patient outcomes. We examined associations with pre-post change and rate of change in psychological distress, and likelihood of attending >1 treatment session. We largely did not find associations between therapist-level characteristics and patient outcomes. Most predictors failed to replicate across sensitivity analyses and/or were non-significant following p-value correction. The most robust evidence suggested that interpersonal capacities assessed via a performance task are associated with likelihood of attending >1 treatment session. A key limitation of the study is small therapist effects which may have reduced statistical power. Empirically, it remains uncertain what qualities characterize highly effective therapists.
Keywords: psychotherapy, expertise, therapist effects, multimodal assessment, behavioral measures
The idea that some therapists are better than others dates back to the 1930s (e.g., Rosenzweig, 1936). Therapist differences in outcome have been studied empirically for over 40 years (Luborsky et al., 1985), with a rapid increase in work in this area in the past 20 years (Wampold & Owen, 2021). It is now well established that therapists indeed differ systematically in their patients’ outcomes – that is, some therapists are consistently more effective than other therapists (Wampold & Owen, 2021). A meta-analysis of 20 studies that explicitly focused on the therapist effect (Johns et al., 2019) estimated that 5% of variance in treatment outcome is attributable to therapists (i.e., therapist-level intraclass correlation [ICC] = .050). This estimate is identical to that provided by an earlier meta-analysis of 46 studies that allowed the inclusion of studies that reported therapist-level ICCs, but were not specifically designed to study the therapist effect (Baldwin & Imel, 2013). While 5% would appear to be a relatively modest proportion of variance explained, it is similar in magnitude or larger than that associated with other primary contributors to treatment outcome (e.g., differences between treatments, therapists’ adherence and competence, therapeutic alliance; Power et al., 2022; Wampold & Owen, 2021).
Whether or not therapists can improve their outcomes is an area of active debate (Wampold & Owen, 2021). Some have argued that the context of therapy is a difficult environment in which to learn, due to its unpredictable nature and therapists’ lack of access to feedback necessary for improving performance (Tracey et al., 2014). Thus, differences between therapists may be due more to innate qualities (e.g., interpersonal skill; Anderson et al., 2009) than training or experience (Goldberg, Rousmaniere et al., 2016; Minami et al., 2009). Drawing from the science of expertise (e.g., Ericsson & Charness, 1994), there has been increasing emphasis on the use of deliberate practice (e.g., rehearsal of specific therapeutic skills paired with outcome monitoring and feedback; Miller et al., 2013; Rousmaniere, 2017) as a means for improving therapists’ outcomes. Although empirical data are limited, there is evidence supporting both the notion that therapists do not typically improve (Germer et al., 2022; Goldberg, Rousmaniere, et al., 2016) and that they can improve in an adequately enriched learning context (e.g., Goldberg, Babins-Wagner, et al., 2016). Regardless of whether therapists’ skills are inborn or trainable, there is general agreement that therapists vary in effectiveness.
Given therapists differ in their patients’ outcomes, a key question naturally follows: what are the qualities of effective therapists? The answer to this question may lie in one (or both) of two types of qualities: (a) therapist actions during psychotherapy with a given patient (e.g., therapists’ use of specific skills such as Socratic questioning or empathic reflections; Hill & Norcross, 2023), which can be measured as in-session behavior, and (b) therapist characteristics (e.g., training, relational skills; Heinonen & Nissen-Lie, 2019), measured outside of psychotherapy. Therapists’ in-session behavior is certainly worthy of investigation. However, in-session behavior is a somewhat limited explanation for therapist effects given that therapists actions within psychotherapy are dependent to a large degree on the particular patient being treated (Boswell et al., 2013; Wampold & Imel, 2015). Assessing what the therapist brings to psychotherapy, including various characteristics (e.g., age, experience, personality) as well as skills (e.g., ability to express empathy with difficult patients) has self-evident implications for training, professional development, and personnel selection. Here we aim examine the association between therapist qualities assessed outside of therapy and treatment outcomes.
The most recent systematic review of the link between therapist characteristics and treatment outcomes (Heinonen & Nissen-Lie, 2019) identified 31 studies measuring professional characteristics (e.g., subjective efficacy, professional self-doubt), socio-emotional characteristics such as performance-based measures of interpersonal skills (e.g., Facilitative Interpersonal Skills [FIS] task; Anderson et al., 2009), multicultural counseling competencies, and personal characteristics (e.g., attachment style). Therapists’ self-report of some constructs (e.g., professional self-doubt, attachment style, time spent engaging in deliberate practice) have shown links with outcomes. However, these constructs have largely been examined in only one or two studies that included relatively small numbers of patients and/or therapists (Heinonen & Nissen-Lie, 2019). Consistent with other reviews (e.g., Lingiardi et al., 2018; Wampold & Owen, 2021), the strongest evidence in Heinonen and Nissen-Lie’s (2019) review emerged for therapists’ interpersonal skills assessed via performance tasks such as the FIS. This task involves therapists recording responses to vignettes of patients (played by actors) displaying interpersonally challenging behavior in psychotherapy. Responses are then coded by raters.
It is difficult to draw firm conclusions from this literature. Dozens of candidate therapist characteristics have been examined across studies, with most studies reporting only a small number of potential characteristics (Heinonen & Nissen-Lie, 2019). Further, some theoretically important therapist characteristics (e.g., multicultural orientation; Davis et al., 2018) have not been thoroughly investigated. Thus, there is a need for large-scale evaluations of a range of therapist characteristics. Given evidence that non-self-report measures of therapist characteristics have often been most consistently predictive of patient outcomes (e.g., FIS skills), multimodal measurement may be particularly worthwhile. By multimodal evaluation, we are referring to assessment of therapist characteristics using multiple modes of data collection (i.e., self-report, performance tasks such as the FIS, behavioral measures).
Current Study
We aimed to evaluate a diverse set of therapist characteristics that have been empirically or are theoretically linked with patient outcomes (Heinonen & Nissen-Lie, 2019). We focused on four categories of predictors: interpersonal skills and related domains, multicultural orientation and related domains, professional characteristics (e.g., use of deliberate practice, years of experience), and attached-related measures. To evaluate these constructs, we gathered multimodal data through self-report, performance tasks, and computerized behavioral measures. Given a lack of certainty regarding characteristics most likely to predict patient outcomes, we intentionally gathered a large set of candidate variables. We conducted a preregistered set of analyses in a sample of therapists and patients drawn from clinics in the United States and Canada. We report results with and without false-discovery rate (FDR; Benjamini & Hochberg; 1995) p-value correction.
Transparency and Openness
Preregistration
Study hypotheses and analytic plan were preregistered. This included preregistrations focused on interpersonal skills and related constructs (https://osf.io/pfkdj/overview), professional training and related constructs (https://osf.io/ecv2z/overview), multicultural orientation and related constructs (https://osf.io/wp6kv/overview), and hypothesized mediation of therapist effects via therapist attachment style (https://osf.io/7xn48/overview). Preregistrations were submitted when a portion of the therapist data was collected, but prior to any analyses. Linked patient data had not been received at the time of preregistration. We made ten deviations from our preregistration which are described in Method.
Data, Materials, Code, and Online Resources
Due to IRB restrictions, we are not able to share study data. However, analysis scripts are publicly available (https://osf.io/am4x9/files/osfstorage). All study materials except the FIS and Multicultural Orientation Task are publicly or commercially available (e.g., through Inquisit).
Reporting
We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.
Ethical Approval
This research was approved by the University of Wisconsin – Madison Institutional Review Board ([IRB], 2019–0091).
Method
Participants and Procedure
Therapists were recruited from clinics in the United States and Canada that use routine outcome monitoring systems (e.g., Outcome Questionnaire – 45 [OQ-45; Lambert et al., 2004], Behavioral Health Monitor [BHM; Kopta & Lowry, 2002]). Study staff sent email invitations to directors of clinics known to be using the OQ-45 or BHM-20. Invitations were forwarded to clinical staff (licensed staff and clinical trainees) who were invited to complete an approximately 2-hour assessment battery. All assessments were delivered online. Questionnaire measures were delivered through Qualtrics. Performance tasks (FIS and Multicultural Orientation [MCO] task; Stewart et al., 2024) were completed using the Skillsetter platform (www.Skillsetter.com). Computerized behavioral measures were completed using the Inquisit platform (www.millisecond.com/) and a web-app built by our team (for the Empathic Accuracy Task). Participants were paid $200 for completing the assessment battery. Patient data provided by participating clinics were linked with therapist data. Therapists completed assessments between November 2019 and January 2021. Patients’ first session occurred between November 2010 and October 2024. However, data used in all but one sensitivity analysis were restricted to patients whose first session occurred on or after the date the therapist assessment occurred.
We included an estimated power table in our preregistration and aimed to collect data from between 100 and 200 therapists anticipating that each would have 20 patients. Using variance estimates from a similar naturalistic psychotherapy data set (Goldberg et al., 2023), we estimated that 100 therapists would provide approximately 80% power to detect a difference in patient outcomes associated with a given therapist-level predictor in standard deviation units (i.e., Cohen’s d) of d ≥ 0.40. We estimated a sample of 200 therapists would provide ≥ 80% power to detect a difference of d ≥ 0.30. We anticipated that we would likely be underpowered to detect small effects (ds ≤ 0.20).
Therapists
A total of 97 therapists provided data for one or more predictors and could be matched with patient data. Therapists were recruited from 13 clinics in the United States (n = 10) and Canada (n = 3). Most therapist (71.1%) were recruited from a large community mental health agency in Canada. The remaining therapists were recruited from university counseling centers, with 10 based in the United States and two based in Canada. Therapists were on average 35.77 years old (standard deviation [SD] = 9.99); 81.4% identified as women, 15.4% as men (one identifying as a transman), 3.1% did not report a gender; 69.1% identified as non-Latinx White, 6.2% as Black, 20.6% as Asian, 1.0% as Latinx, 1.0% as “other” race/ethnicity, and 2.1% did not report a race/ethnicity. Therapists had an average of 63.42 (SD = 79.62, median = 27) patients in the data set.
Patients
A total of 6,152 patients were seen by the 97 therapists who provided therapist-level data. Patients attended individual psychotherapy for a total of 28,843 sessions. Data were analyzed from the first episode of care with gaps >120 days used to demarcate distinct episodes (Goldberg et al., 2016). Patients’ initial session occurred on average 1.29 years (SD = 1.17) after the therapists completed the assessment battery.
Most patients (86.8%) were from the community mental health agency in Canada, with the remaining patients coming from university counseling centers in the United States (10.3%) and Canada (3.0%). Patients were on average 35.05 years old (SD = 12.63); 61.2% identified as women, 33.3% as men, 5.5% did not report a gender; 53.9% identified as non-Latinx White, 4.1% as “other” race/ethnicity, 9.0% as Asian, 2.3% as Black, 2.4% as Native American / Aboriginal / First Nations / Alaskan Native, 2.2% as Latinx, 1.3% as Middle Eastern, 0.1% as multiracial, and 24.8% did not report a race/ethnicity. Detailed clinical information was not available at the patient level. However, common reasons for seeking treatment at the Canadian community mental health agency included family/marital problems, depression, anxiety, stress, and eating disorders (Rousmaniere et al., 2016) and common reasons for seeking treatment at college counseling centers include depression, anxiety, stress, as well as family and academic concerns (Xiao et al., 2017).
Measures
Patient Outcome Measures
Two different outcome measures were used across the 13 clinics. The community mental health agency used the OQ-45 (Lambert et al., 2004). This is a 45-item self-report measure designed to index psychological distress in the context of psychotherapy. Scores ≥63 are considered clinically elevated (Lambert et al., 2004). Internal consistency for the total score was high (α = .94). The university counseling centers used the BHM-20 (Kopta & Lowry, 2002). This is a 20-item self-report measure that was also designed to index psychological distress in the context of psychotherapy. Scores ≤2.78 on the total score (i.e., General Mental Health scale) are considered clinically elevated (Kopta et al., 2015). Internal consistency for the total score was high (α = .91).
Therapist-Level Characteristics
Therapists completed an assessment battery that included self-report questionnaires, performance tasks, and computerized behavioral tasks. Tasks were grouped in four primary content domains: FIS-related measures, multicultural-related measures, professional characteristics, and attachment-related measures (Table 1). Given the FIS has shown the most consistent associations with patient outcomes, many measures were selected to assess constructs broadly related to the eight FIS domains (described below). Given the growing literature highlighting the importance of multicultural capacities in therapy, we also included the MCO task (described below) and related measures. We assessed several professional characteristics that have been empirically or are theoretically linked with patient outcomes (Heinonen & Nissen-Lie, 2019). Finally, we assessed attachment style and a related construct (childhood trauma). As noted above, we intentionally gathered a large set of candidate characteristics based on the view that a more comprehensive (rather than narrow) evaluation would be most valuable at the current stage of work in this area.
Table 1.
Study Measures Organized by Domain
| Measure | Domain | Modality |
|---|---|---|
| FIS total | FIS overall | Performance task |
| FIS verbal fluency | FIS verbal fluency | Performance task |
| Word fluency test | FIS verbal fluency | Computerized behavioral measure |
| FIS hope and positive expectations | FIS hope and positive expectations | Performance task |
| Depression, Anxiety, and Stress Scale | FIS hope and positive expectations | Self-report |
| Psychological Well-being | FIS hope and positive expectations | Self-report |
| FIS persuasiveness | FIS persuasiveness | Performance task |
| Ten Item Personality Inventory extraversion | FIS persuasiveness | Self-report |
| FIS emotional expression | FIS emotional expression | Performance task |
| Emotional Stroop task | FIS emotional expression | Computerized behavioral measure |
| FIS warmth, acceptance, and understanding | FIS warmth, acceptance, and understanding | Performance task |
| Ten Item Personality Inventory agreeableness | FIS warmth, acceptance, and understanding | Self-report |
| Five Facet Mindfulness Questionnaire Short Form total score | FIS warmth, acceptance, and understanding | Self-report |
| Self-Compassion Scale Short Form total score | FIS warmth, acceptance, and understanding | Self-report |
| Dispositional Contempt | FIS warmth, acceptance, and understanding | Self-report |
| FIS empathy | FIS empathy | Performance task |
| Interpersonal Reactivity Index | FIS empathy | Self-report |
| Empathic accuracy task | FIS empathy | Computerized behavioral measure |
| FIS alliance bond capacity | FIS alliance bond capacity | Performance task |
| FIS alliance rupture and repair | FIS alliance rupture and repair | Performance task |
| NIH Toolbox Loneliness | FIS alliance bond capacity / rupture and repair | Self-report |
| NIH Toolbox Emotional Support | FIS alliance bond capacity / rupture and repair | Self-report |
| Multicultural Orientation task overall | Multicultural | Performance task |
| Multicultural Orientation task comfort | Multicultural | Performance task |
| Multicultural Orientation task humility | Multicultural | Performance task |
| Multicultural Orientation task opportunity | Multicultural | Performance task |
| Racial Implicit Association task | Multicultural | Computerized behavioral measure |
| Multicultural Knowledge, Awareness, and Skills Scale total score | Multicultural | Self-report |
| Use of feminist and/or multicultural theory | Multicultural | Self-report |
| Deliberate practice hours | Professional characteristics | Self-report |
| Difficulties in Therapeutic Practice professional self-doubt | Professional characteristics | Self-report |
| Years providing psychotherapy | Professional characteristics | Self-report |
| Perceived efficacy | Professional characteristics | Self-report |
| Years of personal therapy | Professional characteristics | Self-report |
| Multitheoretical List of Interventions therapeutic technique diversity | Professional characteristics | Self-report |
| Experiences in Close Relationships anxious | Attachment | Self-report |
| Experiences in Close Relationships avoidant | Attachment | Self-report |
| Childhood Trauma Questionnaire total | Attachment | Self-report |
Note. FIS = Facilitative Interpersonal Skills; NIH = National Institutes of Health.
Performance Tasks.
Facilitative Interpersonal Skills task (FIS; Anderson et al., 2009).
The FIS task involves responding to brief video vignettes of actors depicting interpersonally challenging moments in psychotherapy. To reduce participant burden, the current study used four vignettes. Participants are asked to record themselves responding as if they were the patients’ therapist. Recorded responses are rated on eight standardized FIS items (Anderson et al., 2009): verbal fluency; hope and positive expectations; persuasiveness; emotional expression; warmth, acceptance, and understanding; empathy; alliance-bond capacity; and alliance rupture-repair responsiveness. Ratings range from 1 (skill deficit) to 5 (optimal presence of skill). In the current study, three raters who were graduate students in counseling psychology received 17 hours of training in FIS ratings over the course of 10 training sessions by licensed psychologists who were former students of the FIS developer. Initial training occurred using non-study videos until the raters achieved adequate inter-rater reliability (ICC2,3 ≥ .60; Cicchetti, 1994). Each video was rated by all three raters. Inter-rater reliability across all items was good (ICC = .67), although it ranged from fair to excellent across FIS domains (ICCs = .75, .58, .62, .66, .70, .68, .68, .60, for verbal fluency; hope and positive expectations; persuasiveness; emotional expression; warmth, acceptance, and understanding; empathy; alliance-bond capacity; and alliance rupture-repair responsiveness). Item scores were computed by averaging across the three raters. The FIS total score was computed by averaging across the eight domains (α = .94).
Multicultural Orientation Task (MCO; Stewart et al., 2024).
The MCO also involves responding to brief video vignettes. The current study used four vignettes. Each vignette depicts a patient sharing content with culturally relevant elements (e.g., a Black man experiencing discrimination at work, a gay man struggling with dating). Raters evaluate the degree to which the therapists’ response demonstrated cultural comfort (i.e., comfort having a conversation about culture, three items), cultural humility (i.e., non-superior stance toward patients, three items), and cultural opportunity (i.e., seizing the moment to ask about cultural identities, one item). A final item assesses the degree to which a therapist’s response was generally therapeutic. Items are rated from 1 (poor performance) to 5 (high performance). A second set of three raters who were graduate students in counseling psychology received 10 hours of training in MCO ratings over the course of eight training sessions by one of the MCO developers. Training occurred using non-study videos until the raters achieved adequate inter-rater reliability (ICC2,3 ≥ .60; Cicchetti, 1994). Each video was rated by all three raters. Inter-rater reliability across all items was excellent (ICCs = .81), although it ranged from fair to excellent across MCO domains (ICCs = .55, .62, .91, and .64, for comfort, humility, opportunity, and generally therapeutic, respectively). Item scores were computed by averaging across the three raters. Cultural comfort and cultural opportunity items were averaged for each scale (αs = .91 and .89, respectively).
Computerized Behavioral Measures and Self-Report Questionnaires.
A description of these measures is included in Supplemental Materials Table 1.
Data Analysis
We made ten deviations from the preregistered plan. First, we restricted our primarily analyses to patients whose course of therapy began on or after the day therapists completed their assessments. This was done based on the notion that these characteristics may not be stable over time and that it may be problematic to predict patient outcomes that occurred prior to this assessment. Second, we report results with an FDR-correction applied. Third, we do not report analyses examining FIS as a mediator between attachment security and patient outcomes as this was outside our focus on therapist-level predictors of patient outcomes. Fourth, we do not report analyses examining a face rating task as a measure of racial bias, therapist grade point average, or therapist graduate record exam scores. The face rating task could not be used due to a technical error in which two versions of the task were used over the course of the study. Grade point average and graduate record exam scores were reported in inconsistent ways and infrequently such that they were deemed unusable. Fifth, we included a dummy-coded variable indicating whether data was drawn from a clinic using the BHM-20 (or the OQ-45) based on exploratory data analyses demonstrating that changes on the OQ-45 were substantially larger than changes on the BHM-20 on average (ds = −0.54 and −0.36, for OQ-45 and BHM-20, respectively, p < .001). Sixth, we added three sensitivity analyses that examined results restricted to the OQ-45 and BHM-20 and a model covarying the time between a patient’s first session and when the therapist completed the assessment battery. Seventh, to help with model convergence, we z-scored all therapist-level predictor variables. Eighth, we modeled clinic as a random intercept. Ninth, we modified the random effects for models that would not converge. Tenth, we conducted a set of analyses using all patient data, regardless of whether a patient’s first session occurred prior to their therapist’s assessment. This sensitivity analysis included 109 therapists who saw 15,550 patients.
In total, we examined 38 therapist-level predictors (Table 1) of three patient-level outcomes: pre-post change in psychological distress (on OQ-45 or BHM-20), rate of change in psychological distress, and early termination (i.e., whether patients completed >1 treatment session)1. Separate sets of models were conducted for each predictor. Data were analyzed using the ‘lme4 package (Bates et al., 2015) in R (R Core Team, 2024). Maximum likelihood estimation which is robust to data missing at random (Graham, 2009) was used across models.
For pre-post change, we used two-level multilevel models (i.e., patients nested within therapists) with therapist-level characteristics entered as Level 2 predictors:
| (Equation 1) |
where Yij reflects the pre-post changes in psychological distress on the OQ-45 or BHM-20 in standardized Cohen’s d units for patient i seen by therapist j. This value was calculated by subtracting psychological distress at the first session from psychological distress at the final session and dividing the difference score by the first session SD. In order to allow inclusion in the same model, scores on the BHM-20 (where higher scores reflect lower levels of psychological distress; Kopta et al., 2002) were reversed prior to calculating d. Thus, smaller (e.g., more negative) ds reflect greater reductions in psychological distress. The fixed intercept (β00) reflects the overall mean change in psychological distress and the fixed slope for the BHM-20 (β01) reflects the overall effect of whether the BHM-20 was used. The fixed slope for the therapist-level predictors (β02) reflects the relationship between each of the 38 predictors with pre-post change in psychological distress. Negative coefficients for β02 indicate that higher levels of the predictor (e.g., FIS total score) was associated with larger (i.e., more negative) pre-post reductions in psychological distress. Random effects (shown within brackets) included a random intercept at the therapist level (U0j) and an error term (eij). Including a random intercept at the clinic level, which would have produced a three-level model (i.e., patients nested within therapists nested within clinics), resulted in a singular fit due to insufficient variance at the clinic level. Thus, a two-level model was used. Consistent with our preregistration, the pre-post model and all other models examining the Multitheoretical List of Therapeutic Interventions–30 (MULTI-30; Solomonov et al., 2019) included a linear and quadratic version of the entropy score. We examined interactions with the quadratic version. An empty model (i.e., Equation 1 with the therapist characteristic predictor [i.e., β02] omitted) was used to calculate therapist-level ICCs. Specifically, we divided the variance in pre-post change at the therapist level by the total variance.
For rate of change, we used four-level multilevel models (i.e., sessions nested within patients nested within therapists nested within clinics). To determine how best to model rate of change, we compared models with linear only; linear and quadratic; linear, quadratic, and cubic; and log-linear session terms. The model with log-linear session number showed superior fit (Bayesian Information Criterion [BIC] = 54114.25, 53573.07, 52862.28, 52307.04; for linear only; linear and quadratic; linear, quadratic, and cubic; and log-linear respectively). Adding a random slope for the log-linear session number coefficient to the log-linear model further improved the fit (BIC = 50559.04, χ2 [4] = 1786.60, p < .001). Therapist-level characteristics were entered as Level 3 predictors in interaction with the log session term:
| (Equation 2) |
where Yijkl reflects level of psychological distress on the OQ-45 or BHM-20 at session i for patient j seen by therapist k in clinic l. In order to allow inclusion in the same model, scores were standardized (i.e., z-scored) by subtracting the grand mean score (on OQ-45 or BHM-20) at the first session across the full sample from a given patient’s score and dividing by the sample SD at the first session. As for the pre-post models, BHM-20 scores were reverse scored so that lower scores reflect lower psychological distress. Psychological distress scores were predicted by a fixed intercept (β0000), a fixed slope for the BHM-20 (β0001), a fixed slope for a natural log-transformed session number (β1000), and fixed slope for the therapist-level predictors (β0010). The fixed slope for the interaction between the log session number term and the therapist-level predictors (β0020) is the coefficient of interest for evaluating whether therapist-level predictors moderate rate of change in psychological distress. The model included random intercepts at the patient level (U0jkl), the therapist level (U00kl), the clinic level (U000l), and an error term (eijkl). The model also included a random slope that allowed the relationship between log session number and psychological distress to vary across patients (U1jkl) and therapists (U01kl).
Multilevel logistic regression was used to examine therapist-level predictors of whether a given patient attended >1 treatment session. This model was equivalent to Equation 1, except a dichotomous variable (coded as 1 when a given patient attended >1 treatment session and 0 when a given patient did not attend >1 treatment session) was entered as the dependent variable and a random intercept was included at the clinic level (i.e., producing a three-level model with patients nested within therapists nested within clinics). A logit link function was used to model the dichotomous outcome. Positive coefficients for the therapist-level predictors reflect a greater likelihood that a therapists’ patients on average were more likely to attend >1 treatment session. As for the pre-post change models, an empty model (i.e., Equation 1 but with the therapist characteristic predictor [i.e., β01] omitted and a clinic random intercept included) was used to calculate therapist-level ICCs. Specifically, we divided the variance in likelihood of attending >1 treatment session at the therapist and clinic levels by the sum of total variance combined with π2 / 3 (Rabe-Hesketh & Skrondal, 2012).
Sensitivity Analyses
Models were rerun in eight sensitivity analyses. The decision to conduct a large set of sensitivity analyses was motivated by multiverse analysis within psychology that highlights the increased transparency these analyses can provide (Steegen et al., 2016). Our four preregistered sensitivity analyses included models with outliers on therapist-level predictors excluded (i.e., values ≥3 SDs from sample mean), restricting to therapist with ≥5 patients, restricting to patients with ≥3 sessions, and restricting to patients with clinically elevated symptoms at baseline2. In addition, we ran separate models for the OQ-45 and BHM-20, a model where we covaried the time between the date of each patient’s first session and when the therapist completed the assessment battery, and a model that used all patient data regardless of whether the assessment battery occurred later than a patient’s first session. For pre-post change and early termination prediction sensitivity analyses restricted to the OQ-45 or BHM-20, we used the models in Equation 1 and 2, but with the BHM-20 term omitted. The clinic random effect was omitted for the OQ-45 models as only one clinic used the OQ-45. For the rate of change model, we followed the same model building steps to identify the best fitting model for the sensitivity analyses restricted to the OQ-45 or BHM-20. In all cases, models with log-linear session number and random slopes fit best. In instances where models did not converge (e.g., due to singular fit), random intercepts at the clinic level were removed.
Results
Patients attended an average of 4.69 sessions (SD = 4.45). Average pre-post change in psychological distress was d = −0.50 (SD = 0.77). Most patients (77.6%) had >1 sessions. Most patients (75.7%) had clinically elevated psychological distress at baseline. Descriptive statistics for therapist-level variables are reported in Supplemental Materials Table 2. Patient demographic characteristics across analytic samples are reported in Supplemental Materials Table 3. The therapist-level ICC for pre-post change in psychological distress was 1.63%. The therapist-level ICC when restricting to patients with >1 session was 1.56%. Therapist-level ICCs for the OQ-45 and BHM-20 were 1.59% and 5.27%, respectively. The therapist-level ICC for likelihood of attending >1 treatment session was 5.74% (ICCs = 4.96% and 18.5% for data drawn from clinics using the OQ-45 and BHM-20, respectively).
Pre-Post Change
Results of the pre-post change model are reported in Table 2 and Supplemental Materials Tables 4 to 11. In the primary model, only higher FIS verbal fluency was associated with larger pre-post reductions in distress (b = −0.0380). A similar effect was observed in four sensitivity analyses (outliers removed, therapists with ≥5 patients, clinically elevated at baseline, covarying time to therapist assessment). Across the eight sensitivity analyses, there were eight additional instances of significant predictors. Higher therapist-level negative affect, higher FIS emotional expression, higher FIS empathy, more use of deliberate practice, and higher professional self-doubt were associated with smaller pre-post reductions in distress. More years of providing therapy and higher self-rated effectiveness were associated with larger pre-post reductions in distress. None of the associations in the pre-post change models survived p-value correction.
Table 2.
Pre-Post Model Results
| Predictor | Level 1 n | Level 2 n | Level 3 n | Estimate | p | p FDR |
|---|---|---|---|---|---|---|
| FIS total | 4,728 | 84 | NA | −0.0025 | 0.884 | 0.953 |
| FIS verbal fluency | 4,728 | 84 | NA | −0.0380 | 0.037 | 0.953 |
| Word fluency test | 4,122 | 69 | NA | 0.0044 | 0.828 | 0.953 |
| FIS hope and positive expectations | 4,728 | 84 | NA | −0.0005 | 0.976 | 0.976 |
| Depression, Anxiety, and Stress Scale | 4,964 | 88 | NA | 0.0290 | 0.077 | 0.953 |
| Psychological Well-being | 4,964 | 88 | NA | −0.0212 | 0.214 | 0.953 |
| FIS persuasiveness | 4,728 | 84 | NA | −0.0148 | 0.409 | 0.953 |
| Ten Item Personality Inventory extraversion | 4,964 | 88 | NA | −0.0134 | 0.451 | 0.953 |
| FIS emotional expression | 4,728 | 84 | NA | 0.0159 | 0.380 | 0.953 |
| Emotional Stroop task | 3,941 | 69 | NA | 0.0069 | 0.728 | 0.953 |
| FIS warmth, acceptance, and understanding | 4,728 | 84 | NA | 0.0112 | 0.525 | 0.953 |
| Ten Item Personality Inventory agreeableness | 4,964 | 88 | NA | −0.0035 | 0.837 | 0.953 |
| Five Facet Mindfulness Questionnaire Short Form total score | 4,964 | 88 | NA | 0.0038 | 0.826 | 0.953 |
| Self-Compassion Scale Short Form total score | 4,964 | 88 | NA | 0.0107 | 0.526 | 0.953 |
| Dispositional Contempt | 4,964 | 88 | NA | 0.0213 | 0.157 | 0.953 |
| FIS empathy | 4,728 | 84 | NA | 0.0097 | 0.581 | 0.953 |
| Interpersonal Reactivity Index | 4,964 | 88 | NA | 0.0050 | 0.772 | 0.953 |
| Empathic accuracy task | 4,566 | 82 | NA | −0.0019 | 0.912 | 0.953 |
| FIS alliance bond capacity | 4,728 | 84 | NA | 0.0026 | 0.870 | 0.953 |
| FIS alliance rupture and repair | 4,728 | 84 | NA | 0.0021 | 0.903 | 0.953 |
| NIH Toolbox Loneliness | 4,964 | 88 | NA | 0.0177 | 0.283 | 0.953 |
| NIH Toolbox Emotional Support | 4,964 | 88 | NA | 0.0016 | 0.928 | 0.953 |
| Multicultural Orientation task overall | 4,826 | 85 | NA | 0.0138 | 0.462 | 0.953 |
| Multicultural Orientation task comfort | 4,826 | 85 | NA | 0.0052 | 0.790 | 0.953 |
| Multicultural Orientation task humility | 4,826 | 85 | NA | 0.0121 | 0.527 | 0.953 |
| Multicultural Orientation task opportunity | 4,826 | 85 | NA | 0.0029 | 0.862 | 0.953 |
| Racial Implicit Association task | 4,122 | 69 | NA | −0.0091 | 0.617 | 0.953 |
| Multicultural Knowledge, Awareness, and Skills Scale total score | 4,964 | 88 | NA | −0.0156 | 0.347 | 0.953 |
| Use of feminist and/or multicultural theory | 5,181 | 96 | NA | 0.0124 | 0.473 | 0.953 |
| Deliberate practice hours | 4,773 | 82 | NA | 0.0247 | 0.134 | 0.953 |
| Difficulties in Therapeutic Practice professional self-doubt | 4,964 | 88 | NA | −0.0022 | 0.900 | 0.953 |
| Years providing psychotherapy | 5,171 | 95 | NA | −0.0142 | 0.565 | 0.953 |
| Perceived efficacy | 5,171 | 95 | NA | −0.0117 | 0.512 | 0.953 |
| Years of personal therapy | 5,047 | 93 | NA | 0.0103 | 0.516 | 0.953 |
| Multitheoretical List of Interventions therapeutic technique diversity | 4,964 | 88 | NA | −0.006 | 0.660 | 0.953 |
| Experiences in Close Relationships anxious | 4,964 | 88 | NA | 0.0068 | 0.691 | 0.953 |
| Experiences in Close Relationships avoidant | 4,964 | 88 | NA | 0.0041 | 0.812 | 0.953 |
| Childhood Trauma Questionnaire total | 4,704 | 80 | NA | 0.0307 | 0.078 | 0.953 |
Note. FIS = Facilitative Interpersonal Skills; NIH = National Institutes of Health p = p-value; pFDR = false-discovery rate corrected p-values; Estimate = fixed effect of each predictor variable from multilevel model predicting pre-post changes in psychological distress; Level 1 = patient level; Level 2 = therapist level; Level 3 = clinic level; NA = random intercept omitted due to model non-convergence. Negative estimates reflect associations with larger pre-post reductions in psychological distress.
Rate of Change
Results of the rate of change model are reported in Table 3 and Supplemental Materials Tables 12 to 19. In the primary model, negative affect and three FIS dimensions (emotional expression, empathy, alliance rupture-repair responsiveness) were the only therapist-level predictors of rate of change in psychological distress, with higher scores on each being associated with a slower (less steep) decrease in psychological distress over time (interaction bs = 0.028 to 0.039, ps = .011 and .041). These effects did not survive p-value correction. Negative affect and various FIS scores were significant predictors in some (but not all) sensitivity analyses, always in the same direction as in the primary model (i.e., higher scores, slower decrease in psychological distress). Higher MCO scores (four models), loneliness (two models), verbal fluency, emotional Stroop interference, self-compassion, self-report empathy, agreeableness, and use of a feminist and/or multicultural theoretical orientation were associated with a slower decrease in psychological distress over time. Higher racial implicit bias was associated with a faster decrease in distress over time in one model. In total, 17 variables emerged as significant predictors when restricted to the BHM-20 data set including the FIS total score and seven FIS subscales, verbal fluency, emotional Stroop interference, self-compassion, self-reported empathy, MCO generally therapeutic response score and two MCO subscales, racial implicit bias, and use of a feminist and/or multicultural theoretical orientation. Fifteen of these effects survived p-value correction. However, these effects were based on a small number of therapists (ns ≤28) and most did not emerge as significant predictors in the primary model or the other sensitivity analyses, and thus should be interpreted very cautiously. No other effects survived p-value correction.
Table 3.
Rate of Change Model Results
| Predictor | Level 1 n | Level 2 n | Level 3 n | Level 4 n | Estimate | p | p FDR |
|---|---|---|---|---|---|---|---|
| FIS total | 23,719 | 5,426 | 84 | 12 | 0.0287 | 0.052 | 0.363 |
| FIS verbal fluency | 23,719 | 5,426 | 84 | 12 | 0.0040 | 0.797 | 0.864 |
| Word fluency test | 21,968 | 4,763 | 69 | 8 | 0.0060 | 0.719 | 0.854 |
| FIS hope and positive expectations | 23,719 | 5,426 | 84 | 12 | 0.0225 | 0.123 | 0.389 |
| Depression, Anxiety, and Stress Scale | 24,978 | 5,712 | 88 | 12 | 0.0279 | 0.041 | 0.363 |
| Psychological Well-being | 24,978 | 5,712 | 88 | 12 | −0.0157 | 0.270 | 0.489 |
| FIS persuasiveness | 23,719 | 5,426 | 84 | 12 | 0.0269 | 0.072 | 0.363 |
| Ten Item Personality Inventory extraversion | 24,978 | 5,712 | 88 | 12 | −0.012 | 0.403 | 0.589 |
| FIS emotional expression | 23,719 | 5,426 | 84 | 12 | 0.0387 | 0.013 | 0.247 |
| Emotional Stroop task | 21,044 | 4,605 | 69 | 8 | 0.0194 | 0.221 | 0.489 |
| FIS warmth, acceptance, and understanding | 23,719 | 5,426 | 84 | 12 | 0.0172 | 0.243 | 0.489 |
| Ten Item Personality Inventory agreeableness | 24,978 | 5,712 | 88 | 12 | 0.0154 | 0.258 | 0.489 |
| Five Facet Mindfulness Questionnaire Short Form total score | 24,978 | 5,712 | 88 | 12 | 0.0042 | 0.762 | 0.864 |
| Self-Compassion Scale Short Form total score | 24,978 | 5,712 | 88 | 12 | 0.0156 | 0.253 | 0.489 |
| Dispositional Contempt | 24,978 | 5,712 | 88 | 12 | 0.0119 | 0.346 | 0.564 |
| FIS empathy | 23,719 | 5,426 | 84 | 12 | 0.0375 | 0.011 | 0.247 |
| Interpersonal Reactivity Index | 24,978 | 5,712 | 88 | 12 | 0.0242 | 0.074 | 0.363 |
| Empathic accuracy task | 22,922 | 5,235 | 82 | 13 | 0.0018 | 0.903 | 0.903 |
| FIS alliance bond capacity | 23,719 | 5,426 | 84 | 12 | 0.0206 | 0.137 | 0.400 |
| FIS alliance rupture and repair | 23,719 | 5,426 | 84 | 12 | 0.0316 | 0.029 | 0.363 |
| NIH Toolbox Loneliness | 24,978 | 5,712 | 88 | 12 | 0.0172 | 0.208 | 0.489 |
| NIH Toolbox Emotional Support | 24,978 | 5,712 | 88 | 12 | −0.0037 | 0.798 | 0.864 |
| Multicultural Orientation task overall | 24,115 | 5,541 | 85 | 12 | 0.0277 | 0.078 | 0.363 |
| Multicultural Orientation task comfort | 24,115 | 5,541 | 85 | 12 | 0.0260 | 0.116 | 0.389 |
| Multicultural Orientation task humility | 24,115 | 5,541 | 85 | 12 | 0.0243 | 0.117 | 0.389 |
| Multicultural Orientation task opportunity | 24,115 | 5,541 | 85 | 12 | 0.0229 | 0.086 | 0.363 |
| Racial Implicit Association task | 21,968 | 4,763 | 69 | 8 | −0.0131 | 0.387 | 0.588 |
| Multicultural Knowledge, Awareness, and Skills Scale total score | 24,978 | 5,712 | 88 | 12 | 0.0058 | 0.685 | 0.840 |
| Use of feminist and/or multicultural theory | 25,957 | 5,963 | 96 | 13 | 0.0134 | 0.326 | 0.563 |
| Deliberate practice hours | 24,759 | 5,543 | 82 | 12 | −0.0127 | 0.356 | 0.564 |
| Difficulties in Therapeutic Practice professional self-doubt | 24,978 | 5,712 | 88 | 12 | −0.0090 | 0.533 | 0.675 |
| Years providing psychotherapy | 25,904 | 5,948 | 95 | 13 | −0.0045 | 0.819 | 0.864 |
| Perceived efficacy | 25,904 | 5,948 | 95 | 13 | −0.0028 | 0.841 | 0.864 |
| Years of personal therapy | 25,386 | 5,807 | 93 | 13 | 0.0178 | 0.149 | 0.404 |
| Multitheoretical List of Interventions therapeutic technique diversity | 24,978 | 5,712 | 88 | 12 | −0.0157 | 0.217 | 0.489 |
| Experiences in Close Relationships anxious | 24,978 | 5,712 | 88 | 12 | 0.0109 | 0.436 | 0.614 |
| Experiences in Close Relationships avoidant | 24,978 | 5,712 | 88 | 12 | 0.0092 | 0.518 | 0.675 |
| Childhood Trauma Questionnaire total | 23,547 | 5,431 | 80 | 11 | 0.0104 | 0.456 | 0.619 |
Note. FIS = Facilitative Interpersonal Skills; NIH = National Institutes of Health p = p-value; pFDR = false-discovery rate corrected p-values; Estimate = interaction between each predictor variable and log session number from multilevel models predicting psychological distress; Level 1 = session level; Level 2 = patient level; Level 3 = therapist level; Level 4 = clinic level. Negative estimates reflect associations with faster reductions in psychological distress over time.
Early Termination
Results of the early termination prediction model are reported in Table 4 and Supplemental Materials Tables 20 to 26. In the primary model, FIS verbal fluency, FIS persuasiveness, and mindfulness were all associated with greater likelihood of attending >1 treatment session. Deliberate practice was associated with lower likelihood of attending >1 treatment session. None of these effects survived p-value correction in the primary model.
Table 4.
Results of Multilevel Logistic Regression Models Predicting Attendance at >1 Treatment Session
| Predictor | Level 1 n | Level 2 n | Level 3 n | Estimate | p | p FDR |
|---|---|---|---|---|---|---|
| FIS total | 5,584 | 85 | 12 | 0.1124 | 0.078 | 0.494 |
| FIS verbal fluency | 5,584 | 85 | 12 | 0.1886 | 0.003 | 0.076 |
| Word fluency test | 4,905 | 70 | 8 | −0.0212 | 0.766 | 0.845 |
| FIS hope and positive expectations | 5,584 | 85 | 12 | 0.0857 | 0.185 | 0.684 |
| Depression, Anxiety, and Stress Scale | 5,878 | 89 | 12 | −0.0281 | 0.651 | 0.825 |
| Psychological Well-being | 5,878 | 89 | 12 | 0.0183 | 0.774 | 0.845 |
| FIS persuasiveness | 5,584 | 85 | 12 | 0.1785 | 0.004 | 0.076 |
| Ten Item Personality Inventory extraversion | 5,878 | 89 | 12 | 0.0132 | 0.839 | 0.883 |
| FIS emotional expression | 5,584 | 85 | 12 | 0.0746 | 0.273 | 0.684 |
| Emotional Stroop task | 4,768 | 70 | 8 | 0.0937 | 0.210 | 0.684 |
| FIS warmth, acceptance, and understanding | 5,584 | 85 | 12 | 0.0522 | 0.425 | 0.734 |
| Ten Item Personality Inventory agreeableness | 5,878 | 89 | 12 | 0.0719 | 0.227 | 0.684 |
| Five Facet Mindfulness Questionnaire Short Form total score | 5,878 | 89 | 12 | 0.1315 | 0.022 | 0.209 |
| Self-Compassion Scale Short Form total score | 5,878 | 89 | 12 | 0.0424 | 0.486 | 0.769 |
| Dispositional Contempt | 5,878 | 89 | 12 | 0.0161 | 0.775 | 0.845 |
| FIS empathy | 5,584 | 85 | 12 | 0.0731 | 0.267 | 0.684 |
| Interpersonal Reactivity Index | 5,878 | 89 | 12 | −0.0111 | 0.860 | 0.883 |
| Empathic accuracy task | 5,393 | 83 | 13 | −0.0376 | 0.570 | 0.809 |
| FIS alliance bond capacity | 5,584 | 85 | 12 | 0.0286 | 0.640 | 0.825 |
| FIS alliance rupture and repair | 5,584 | 85 | 12 | 0.0643 | 0.311 | 0.684 |
| NIH Toolbox Loneliness | 5,878 | 89 | 12 | 0.1134 | 0.064 | 0.486 |
| NIH Toolbox Emotional Support | 5,878 | 89 | 12 | −0.1049 | 0.105 | 0.570 |
| Multicultural Orientation task overall | 5,701 | 86 | 12 | 0.0343 | 0.627 | 0.825 |
| Multicultural Orientation task comfort | 5,701 | 86 | 12 | 0.0862 | 0.226 | 0.684 |
| Multicultural Orientation task humility | 5,701 | 86 | 12 | 0.0207 | 0.778 | 0.845 |
| Multicultural Orientation task opportunity | 5,701 | 86 | 12 | 0.0621 | 0.312 | 0.684 |
| Racial Implicit Association task | 4,905 | 70 | 8 | 0.0651 | 0.344 | 0.684 |
| Multicultural Knowledge, Awareness, and Skills Scale total score | 5,878 | 89 | 12 | 0.0336 | 0.575 | 0.809 |
| Use of feminist and/or multicultural theory | 6,152 | 97 | 13 | −0.0062 | 0.930 | 0.930 |
| Deliberate practice hours | 5,723 | 83 | 12 | −0.1541 | 0.006 | 0.076 |
| Difficulties in Therapeutic Practice professional self-doubt | 5,878 | 89 | 12 | −0.0601 | 0.360 | 0.684 |
| Years providing psychotherapy | 6,137 | 96 | 13 | −0.0696 | 0.444 | 0.734 |
| Perceived efficacy | 6,137 | 96 | 13 | −0.0250 | 0.736 | 0.845 |
| Years of personal therapy | 5,994 | 94 | 13 | 0.0617 | 0.356 | 0.684 |
| Multitheoretical List of Interventions therapeutic technique diversity | 5,878 | 89 | 12 | −0.0306 | 0.563 | 0.809 |
| Experiences in Close Relationships anxious | 5,878 | 89 | 12 | 0.0492 | 0.425 | 0.734 |
| Experiences in Close Relationships avoidant | 5,878 | 89 | 12 | 0.0748 | 0.281 | 0.684 |
| Childhood Trauma Questionnaire total | 5,592 | 81 | 11 | −0.0676 | 0.293 | 0.684 |
Note. FIS = Facilitative Interpersonal Skills; NIH = National Institutes of Health p = p-value; pFDR = false-discovery rate corrected p-values; Estimate = fixed effect of each predictor variable from multilevel model predicting likelihood of attending >1 treatment session (in log units); Level 1 = patient level; Level 2 = therapist level; Level 3 = clinic level. Positive estimates reflect associations with greater likelihood of attending >1 treatment session.
All four variables that were significant predictors in the primary model (FIS verbal fluency, FIS persuasiveness, mindfulness, deliberate practice) showed associations in the same direction across multiple sensitivity analyses. The FIS total score was also associated with greater likelihood of attending >1 treatment session in two sensitivity analyses. Years of experience, self-rated effectiveness, professional self-doubt, and therapeutic technique diversity were associated with lower likelihood of attending >1 treatment session each in one model. Self-rated effectiveness showed the opposite effect in one model. Use of a feminist and/or multicultural theoretical orientation, MCO generally therapeutic response score, MCO opportunity, and MCO humility were associated with greater likelihood of attending >1 treatment session. Across sensitivity analyses, nine effects survived p-value correction, eight of which reflected the association between FIS verbal fluency or FIS persuasiveness with greater likelihood of attending >1 treatment session.
Discussion
We conducted the first relatively large-scale, multimodal, preregistered assessment of the relationship between therapist-level characteristics and variation in patient outcomes (i.e., therapist effects). We examined 38 predictor variables across three types of outcomes (pre-post change, rate of change, early termination). We evaluated the stability of results across eight sensitivity analyses. Taken together, we found limited evidence suggesting that therapist characteristics impact treatment outcome. Despite selecting a large set of candidate predictors that have been empirically or theoretically linked with outcomes, results were largely null. The few statistically significant findings often failed to replicate across sensitivity analyses and/or failed to survive p-value correction. The most consistent evidence suggested that higher FIS verbal fluency and FIS persuasiveness are associated with greater likelihood of patients attending >1 session of therapy. This finding suggests that therapist interpersonal capacities assessed through brief responses to patient vignettes may contain signals that patients use when determining whether to return for a second therapy session.
Aside from demonstrating linkages between FIS verbal fluency and persuasiveness with early termination (which, to our knowledge, has not been reported before so cannot really be considered a replication), our results are largely disappointing for those hoping to identify reliable therapist-level predictors of treatment outcome. However, before drawing a firm conclusion that we simply cannot predict therapist effects, it is important to acknowledge methodological factors that may have influenced the current findings. The most obvious limitation is the likely low statistical power. Notably, our recruited sample size (n = 97 therapists, n = 6,152 patients) is larger than most even highly cited associations in the literature (e.g., Anderson et al., 2009; Chow et al., 2015; Nissen-Lie et al., 2013). However, the therapist effect in the current sample (~1–2%) was considerably smaller than that reported in both meta-analyses as well as large-scale studies using individual patient data (i.e., ICCs ~ 5%; Johns et al., 2019; Schiefele et al., 2017), as well as smaller than recent studies that have demonstrated linkages between therapists’ interpersonal skills and outcomes (e.g., Schwartz et al., 2025). It is not possible to determine precisely why the therapist effect was smaller in the current study. One potential contributor is the relatively large number of therapists and patients per therapist, both of which tend to be associated with smaller therapist effects (Schiefele et al., 2016). In addition, the Canadian site implemented routine outcome monitoring and encouraged the use of deliberate practice; this may improve patient outcomes generally and reduce variability between therapists (de Jong et al., 2021; Goldberg, Babins-Wagner, et al., 2016). Our final therapist sample was also lower than our target sample size (100–200) and technical issues with some measures (especially the computerized behavioral measures) further reduced sample size and statistical power for models involving those predictors. Moreover, effect sizes observed for most models in the current study were modest. For example, in the primary pre-post model, the largest effect was b = −0.0380 (for FIS verbal fluency), which indicates that a one SD increase in a therapist’s FIS verbal fluency score was associated with a 0.038 lower pre-post d for their patients. Similarly, in the early termination model, the largest effect was b = 0.1886 (also for FIS verbal fluency), which indicates that a one SD increase in a therapist’s FIS verbal fluency score was associated with a 21% increase in the likelihood of their patients attending more than one session (i.e., e0.1886 = 1.208). Both of these effect sizes are below commonly used benchmarks for small magnitude effects (Chen et al., 2010; Cohen, 1992).
We hope this study does not discourage efforts to identify therapist-level characteristics that are linked with patient outcomes. If anything, we hope this study highlights the importance of attending to methodological rigor in our efforts in this area, by preregistering our hypotheses prior to data analysis and by recruiting ideally much larger samples of therapists and patients. The many significant results that emerged when examining data from a relatively small sample of therapists (ns ≤28 for the BHM-20 data set) that did not emerge in other models highlights the danger of drawing conclusions from small numbers of therapists. Unfortunately, likely due to logistical challenges with collecting linked therapist- and patient-level data, much of the existing literature on predictors of therapist differences is based on small numbers of therapists and patients (Heinonen & Nissen-Lie, 2019; Wampold & Owen, 2021, although see Schwartz et al., 2025). Had we selectively reported our results, we could have reported that therapists with higher FIS verbal fluency, higher self-rated effectiveness, more years of providing therapy, and lower negative affect have patients whose symptoms reduce more rapidly. However, this would have ignored the many unexpected associations that emerged in this sensitivity analysis. We just as easily could have concluded that those with higher FIS emotional expression and FIS empathy, higher professional self-doubt, and more use of deliberate practice have patients whose symptoms reduce more slowly.
Research on therapist-level predictors of outcomes that is truly definitive may well require thousands of therapists, especially in contexts where therapist differences are small. This may be feasible in health systems that routinely monitor outcomes (e.g., National Health Service in the United Kingdom). Given that therapist interpersonal skills have been most consistently associated with patient outcomes (Heinonen & Nissen-Lie, 2019), it may be crucial to include non-self-report measures of these capacities. While conducting thousands of FIS assessments would be extraordinarily expensive using human ratings, advances in artificial intelligence and machine learning suggest that automated scoring may soon be ready for use (Goldberg et al., 2021; Goldberg et al., in press). If coupled with text-based administration (Zech et al., 2023), a scalable FIS may not be too far in the future.
In addition to potentially modest statistical power, several other limitations are worth mentioning. First, though drawn from two countries and multiple clinics, our sample of therapists and patients were relatively homogenous in terms of race/ethnicity which limits generalizability to more racially/ethnically diverse populations. Second, there surely is variation in the degree to which the therapist characteristics we assessed are stable qualities. Some characteristics (e.g., negative affect, mindfulness) very likely change over time, even rapidly so. Other characteristics (e.g., FIS skills) may change systematically in the context of training (Anderson et al., 2020). We examined associations between these characteristics assessed at one time point and patient outcomes collected over years. This may have further reduced the likelihood that we were able to detect relevant signals. Although results were similar when covarying time between therapist assessment and a patient’s first session and when expanding the window of acceptable patient data to include sessions that occurred prior to the therapists’ assessment, ideally future studies will include multiple assessments of therapist characteristics to allow examination of their association with patient outcomes at times more proximal to when treatment actually occurred. Gathering such measures during therapist training and beyond may help clarify the degree which therapist expertise develops over the course of training and clinical practice (or not; Tracey et al., 2014). Third, although we examined several types of outcomes (i.e., pre-post change, rate of change, early termination), there were additional ways of defining outcome that may have been considered (e.g., reliable change index; defining early termination as attending only two sessions or not attending the last scheduled appointment; Jacobson & Truax, 1991; Smith & Greenberg, 2012) which may have produced different results. Fourth, we focused only on therapist predictors of outcomes and did not model patient factors or the interaction between therapist and patient factors. This may be an especially important limitation to address in future studies using large samples and analytic approaches particularly well suited to examining interactions (e.g., random forest machine learning methods; Denisko & Hoffman, 2018). Fifth, we examined only a subset of the universe of potentially relevant therapist characteristics. Further, some of the included characteristics and the measures we used to assess these constructs may be impacted by confounding factors that may obscure a meaningful relationship between these constructs and patient outcomes (e.g., acting ability may impact FIS performance but may or may not impact behavior in therapy). Finally, it would be valuable for future studies to employ statistical methods particularly well suited to handling multiple correlated predictors, such as many machine learning approaches. We considered employing machine learning in the current study. However, we opted to conduct individual multilevel models to mirror prior work in this area, to clearly characterize any associations with the individual characteristics examined, and to maintain consistency with our preregistered analytic plan. Moreover, the sample size of therapists (n = 97) is woefully below the recommended sample size for machine learning (e.g., n = 500–1000; Zantvoort et al., 2024). Nonetheless, we are hopeful that future studies employing machine learning on adequately large (and ideally adequately rich, e.g., using session recordings; Kuo et al., 2024) datasets will greatly enhance our understanding of therapist contributors to patient outcome. Evaluating whether models trained using one data source (e.g., from a large Canadian mental health agency) generalize to another data source (e.g., center in the U.S.) could be a valuable future direction, given evidence that generalizability is often poor (Chekroud et al., 2024).
In conclusion, it remains an open question as to what characteristics make one therapist more effective than another. Therapist interpersonal skills (e.g., assessed via FIS) have been suggested as the most consistent predictor of patient outcomes, yet we largely failed to replicate these associations or the various other associations that have been reported in the literature. Some interpersonal capacities (i.e., verbal fluency, persuasiveness) do appear to be linked with likelihood of attending more than one therapy session, but these capacities are largely not associated with larger or faster reductions in symptoms over time. We hope this study encourages further large-scale, preregistered, multimodal assessment of therapist characteristics that can be used to ultimately improve treatment for the benefit of patients.
Supplementary Material
Acknowledgements
The authors are grateful to William T. Hoyt for early discussions about this project.
Funding
SBG was partially supported by the National Center for Complementary and Integrative Health (K23AT010879) and the National Institute of Mental Health (R01MH139512). Nili Solomonov was supported by a grant from the National Institute of Mental Health (K23MH123864). This project was funded in part by the 50th Anniversary Grant from the American Psychological Association Division 29 (Psychotherapy).
Footnotes
Conflicts of Interest
SMK is the CEO of CelestHealth which operates the Behavioral Health Monitor (BHM) measure. The remaining authors have no conflicts to disclose.
Of note, a wide variety of operationalizations of “early termination” have appeared in the psychotherapy literature (Smith & Greenberg, 2012). We used attending only a single session, one of the four operationalizations studied by Hatchett and Park (2003), as a definition of early termination that could be applied reliably across samples. However, results may have differed had another method been used to define early termination.
At the request of an anonymous reviewer, we examined the association between attending ≥3 sessions and reporting clinically elevated symptoms at baseline. Rates of clinically elevated symptoms were similar across the groups (72.5% versus 77.7%, for those attending <3 sessions and those attending ≥3 sessions, respectively). However, these characteristics were very modestly correlated (r = .06, p < .001).
References
- Anderson T, Ogles BM, Patterson CL, Lambert MJ, & Vermeersch DA (2009). Therapist effects: Facilitative interpersonal skills as a predictor of therapist success. Journal of Clinical Psychology, 65(7), 755–768. doi: 10.1002/jclp.20583 [DOI] [PubMed] [Google Scholar]
- Anderson T, Perlman MR, McCarrick SM, & McClintock AS (2020). Modeling therapist responses with structured practice enhances facilitative interpersonal skills. Journal of Clinical Psychology, 76(4), 659–675. [DOI] [PubMed] [Google Scholar]
- Baldwin SA, & Imel ZE (2013). Therapist effects: Findings and methods. In Lambert MJ (Ed.), Bergin and Garfield’s Handbook of Psychotherapy and Behavior Change (6th ed.) (p. 258–297). Hoboken, NJ: Wiley & Sons. [Google Scholar]
- Bates D, Mächler M, Bolker B, & Walker S (2015). Fitting Linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. doi: 10.18637/jss.v067.i01 [DOI] [Google Scholar]
- Benjamini Y, & Hochberg Y (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, 57(1), 289–300. [Google Scholar]
- Boswell JF, Gallagher MW, Sauer-Zavala SE, Bullis J, Gorman JM, Shear MK, Woods S, & Barlow DH (2013). Patient characteristics and variability in adherence and competence in cognitive-behavioral therapy for panic disorder. Journal of Consulting and Clinical Psychology, 81(3), 443–454. doi: 10.1037/a0031437 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chekroud AM, Hawrilenko M, Loho H, Bondar J, Gueorguieva R, Hasan A, Kambeitz J, Corlett PR, Koutsouleris N, Krumholz HM, Krystal JH, & Paulus M (2024). Illusory generalizability of clinical prediction models. Science, 383(6679), 164–167. [DOI] [PubMed] [Google Scholar]
- Chen H, Cohen P, & Chen S (2010). How big is a big odds ratio? Interpreting the magnitude of odds ratios in epidemiological studies. Communications in Statistics-Simulation and Computation, 39, 860–864. doi: 10.1080/03610911003650383 [DOI] [Google Scholar]
- Cicchetti D (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychological Assessment, 6(4), 284–290. [Google Scholar]
- Cohen J (1992). A power primer. Psychological Bulletin, 112(1), 155–159. [DOI] [PubMed] [Google Scholar]
- Davis DE, DeBlaere C, Owen J, Hook JN, Rivera DP, Choe E, … & Placeres V (2018). The multicultural orientation framework: A narrative review. Psychotherapy, 55(1), 89–100. [DOI] [PubMed] [Google Scholar]
- de Jong K, Conijn JM, Gallagher RA, Reshetnikova AS, Heij M, & Lutz MC (2021). Using progress feedback to improve outcomes and reduce drop-out, treatment duration, and deterioration: A multilevel meta-analysis. Clinical Psychology Review, 85, 102002. [DOI] [PubMed] [Google Scholar]
- Denisko D, & Hoffman MM (2018). Classification and interaction in random forests. Proceedings of the National Academy of Sciences, 115(8), 1690–1692. [Google Scholar]
- Ericsson KA, & Charness N (1994). Expert performance: Its structure and acquisition. American Psychologist, 49(8), 725–747. [Google Scholar]
- Germer S, Weyrich V, Bräscher A-K, Mütze K, & Witthöft M (2022). Does practice really make perfect? A longitudinal analysis of the relationship between therapist experience and therapy outcome: A replication of Goldberg, Rousmaniere, et al. (2016). Journal of Counseling Psychology, 69(5), 745–754. [DOI] [PubMed] [Google Scholar]
- Goldberg SB, Babins-Wagner R, Imel ZE, Caperton DD, Weitzman LM, & Wampold BE (2023). Threat alert: The effect of outliers on the alliance–outcome correlation. Journal of Counseling Psychology, 70(1), 81–89. doi: 10.1037/cou0000638 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldberg SB, Flemotomos N, Martinez VR, Tanana M, Kuo P, Pace BT, Villatte JL, Georgiou PG, Van Epps J, Imel ZE, & Atkins DC (2020). Machine learning and natural language processing in psychotherapy research: Alliance as example use case. Journal of Counseling Psychology, 67(4), 438–448. doi: 10.1037/cou0000382 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldberg SB, Rousmaniere T, Miller SD, Whipple J, Nielsen SL, Hoyt WT, & Wampold BE (2016). Do psychotherapists improve with time and experience? A longitudinal analysis of outcomes in a clinical setting. Journal of Counseling Psychology, 63(1), 1–11. doi: 10.1037/cou0000131 [DOI] [PubMed] [Google Scholar]
- Goldberg SB, Tanana M, Imel ZE, Atkins DC, Hill CE, & Anderson T (2021). Can a computer detect interpersonal skills? Using machine learning to scale up the Facilitative Interpersonal Skills task. Psychotherapy Research, 31(3), 281–288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Goldberg SB, Tanana M, Stewart SH, Williams CY, Atkins DC, Imel ZE, & Owen J (in press). Automating the assessment of multicultural orientation through machine learning and natural language processing. Psychotherapy. [Google Scholar]
- Graham JW (2009). Missing data analysis: Making it work in the real world. Annual Review of Psychology, 60, 549–576. [Google Scholar]
- Hatchett GT, & Park HL (2003). Comparison of four operational definitions of premature termination. Psychotherapy: Theory, Research, Practice, Training, 40(3), 226–231. [Google Scholar]
- Heinonen E, & Nissen-Lie HA (2020). The professional and personal characteristics of effective psychotherapists: A systematic review. Psychotherapy Research, 30(4), 417–432. doi: 10.1080/10503307.2019.1620366 [DOI] [PubMed] [Google Scholar]
- Hill CE, & Norcross JC (Eds.). (2023). Psychotherapy skills and methods that work. Oxford University Press. [Google Scholar]
- Jacobson N & Truax P (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. [DOI] [PubMed] [Google Scholar]
- Johns RG, Barkham M, Kellett S, & Saxon D (2019). A systematic review of therapist effects: A critical narrative update and refinement to Baldwin and Imel’s (2013) review. Clinical Psychology Review, 67, 78–93. doi: 10.1016/j.cpr.2018.08.004 [DOI] [PubMed] [Google Scholar]
- Kopta SM, & Lowry JL (2002). Psychometric evaluation of the Behavioral Health Questionnaire-20: A brief, comprehensive instrument for assessing mental health. Psychotherapy Research, 12, 413–426. [Google Scholar]
- Kopta M, Owen J, & Budge S (2015). Measuring psychotherapy outcomes with the Behavioral Health Measure-20: Efficient and comprehensive. Psychotherapy, 52(4), 442–448. doi: 10.1037/pst0000035 [DOI] [PubMed] [Google Scholar]
- Kuo PB, Tanana MJ, Goldberg SB, Caperton DD, Narayanan S, Atkins DC, & Imel ZE (2024). Machine-Learning-Based Prediction of Client Distress From Session Recordings. Clinical Psychological Science, 12(3), 435–446. doi: 10.1177/216770262311726 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kvarven A, Strømland E, & Johannesson M (2020). Comparing meta-analyses and preregistered multiple-laboratory replication projects. Nature Human Behaviour, 4(4), 423–434. [Google Scholar]
- Lambert MJ, Morton JJ, Hatfield D, Harmon C, Hamilton S, Reid RC,… Burlingame GB (2004). Administration and scoring manual for the Outcome Questionnaire-45. Orem, UT: American Professional Credentialing Services. [Google Scholar]
- Lingiardi V, Muzi L, Tanzilli A, & Carone N (2018). Do therapists’ subjective variables impact on psychodynamic psychotherapy outcomes? A systematic literature review. Clinical Psychology & Psychotherapy, 25(1), 85–101. doi: 10.1002/cpp.2131 [DOI] [PubMed] [Google Scholar]
- Luborsky L, McLellan AT, Woody GE, O’Brien CP, & Auerbach A (1985). Therapist success and its determinants. Archives of General Psychiatry, 42(6), 602–611. [DOI] [PubMed] [Google Scholar]
- Miller SD, Hubble MA, Chow DL, & Seidel JA (2013). The outcome of psychotherapy: Yesterday, today, and tomorrow. Psychotherapy, 50(1), 88–97. doi: 10.1037/a0031097 [DOI] [PubMed] [Google Scholar]
- Minami T, Davies DR, Tierney SC, Bettman JE, McAward SM, Averill LA,… Wampold BE (2009). Preliminary evidence on the effectiveness of psychological treatments delivered at a university counseling center. Journal of Counseling Psychology, 56(2), 309–320. doi: 10.1037/a0015398 [DOI] [Google Scholar]
- Power N, Noble LA, Simmonds-Buckley M, Kellett S, Stockton C, Firth N, & Delgadillo J (2022). Associations between treatment adherence–competence–integrity (ACI) and adult psychotherapy outcomes: A systematic review and meta-analysis. Journal of Consulting and Clinical Psychology, 90(5), 427–445. [DOI] [PubMed] [Google Scholar]
- R Core Team (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/ [Google Scholar]
- Rabe-Hesketh S, & Skrondal A (2012). Multilevel and longitudinal modeling using Stata, Volume II: Categorical responses, counts, and survival. Stata Press. [Google Scholar]
- Rosenzweig S (1936). Some implicit common factors in diverse methods in psychotherapy. American Journal of Orthopsychiatry, 6, 412–415. [Google Scholar]
- Rousmaniere T (2017). Deliberate practice for psychotherapists: A guide to improving clinical effectiveness. New York: Routledge. [Google Scholar]
- Rousmaniere TG, Swift JK, Babins-Wagner R, Whipple JL, & Berzins S (2016). Supervisor variance in psychotherapy outcome in routine practice. Psychotherapy Research, 26(2), 1–10. doi: 10.1080/10503307.2014.963730 [DOI] [PubMed] [Google Scholar]
- Schiefele AK, Lutz W, Barkham M, Rubel J, Böhnke J, Delgadillo J, … & Lambert MJ (2017). Reliability of therapist effects in practice-based psychotherapy research: A guide for the planning of future studies. Administration and Policy in Mental Health and Mental Health Services Research, 44, 598–613. [DOI] [PubMed] [Google Scholar]
- Schwartz B, Hehlmann MI, Deisenhofer AK, Rubel JA, Fischer L, Lutz W, & Schöttke H (2025). Elucidating Therapist Differences: Therapists’ Interpersonal Skills and Their Effect on Treatment Outcome. Behaviour Research and Therapy, 104689, 1–11. [Google Scholar]
- Simmons JP, Nelson LD, & Simonsohn U (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. doi: 10.1177/0956797611417632 [DOI] [PubMed] [Google Scholar]
- Solomonov N, McCarthy KS, Gorman BS, & Barber JP (2019). The Multitheoretical List of Therapeutic Interventions–30 items (MULTI-30). Psychotherapy Research, 29(5), 565–580. doi: 10.1080/10503307.2017.1422216 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Steegen S, Tuerlinckx F, Gelman A, & Vanpaemel W (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712. [DOI] [PubMed] [Google Scholar]
- Stewart SH, Drinane JM, Owen J, & Dumas D (2024). Psychotherapy training via a task-based assessment of the multicultural orientation framework: A pilot study. Counselling and Psychotherapy Research, 24(1), 190–198. [Google Scholar]
- Swift JK, & Greenberg RP (2012). Premature discontinuation in adult psychotherapy: A meta-analysis. Journal of Consulting and Clinical Psychology, 80(4), 547–559. doi: 10.1037/a0028226 [DOI] [PubMed] [Google Scholar]
- Tracey TJG, Wampold BE, Lichtenberg JW, & Goodyear RK (2014). Expertise in psychotherapy: An elusive goal? American Psychologist, 69(3), 218–229. doi: 10.1037/a0035099 [DOI] [PubMed] [Google Scholar]
- Wampold B, & Imel ZE (2015). The great psychotherapy debate: The evidence for what makes psychotherapy work (2nd ed.). New York: Routledge. [Google Scholar]
- Wampold BE, & Owen J (2021). Therapist effects: History, methods, magnitude and characteristics of effective therapists (pp. 297–326). In Barkham M, Lutz W, and Castonguay LG (Eds.), Bergin and Garfield’s Handbook of Psychotherapy (7th ed.). John Wiley & Sons, Inc. [Google Scholar]
- Xiao H, Carney DM, Youn SJ, Janis RA, Castonguay LG, Hayes JA, & Locke BD (2017). Are we in crisis? National mental health and treatment trends in college counseling centers. Psychological Services, 14(4), 407–415. [DOI] [PubMed] [Google Scholar]
- Zantvoort K, Nacke B, Görlich D, Hornstein S, Jacobi C, & Funk B (2024). Estimation of minimal data sets sizes for machine learning predictions in digital mental health interventions. npj Digital Medicine, 7, 361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zech J, Foley VK, Hull TD, & Anderson T (2023). Assessing the quality of digital patient-therapist communication: the development and validation of a text-based facilitative interpersonal skills task. Psychotherapy Research, 33(6), 743–756. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
