ABSTRACT
Background and Purpose
Therapeutic alliance (TA)—the collaborative therapist–client partnership—is pivotal to outcomes, yet TA tools tailored to telerehabilitation are limited. We developed and preliminarily validated a therapist‐reported measure for physiotherapists (PTs), occupational therapists (OTs), and speech and language therapists (SLTs): the Working Alliance Inventory–Telerehabilitation (WAI–Tele‐Reha‐G).
Methods
Items drew on expert interviews, adaptation guidelines, and the German WAI–Short Revised (therapist version) and were refined via cognitive interviews (n = 9) and pilot testing (n = 34). Validation involved German‐speaking PTs, OTs, and SLTs (test: n = 128; retest: n = 84). Rasch partial credit modelling assessed model fit, unidimensionality, local independence, threshold ordering, and differential item functioning (DIF). Reliability (Cronbach's alpha, Person Separation Index [PSI]), test–retest reliability (Lin's concordance correlation coefficient [CCC]), convergent validity (with the Unified Theory of Acceptance and Use of Technology‐2 Questionnaire [UTAUT‐2], Technology Commitment–Short Scale, 5‐item Revised Vocational Self‐Efficacy Scale), known‐groups validity, and targeting were evaluated.
Results
Goal, task and bond scales generally fit the Rasch model and were unidimensional with ordered thresholds; no DIF was detected in this sample. Minor local dependence in the task scale was handled via a superitem during Rasch estimation; scoring remained at the item level. Internal consistency was good (alpha = 0.80–0.81) and PSI modest (0.72–0.74). Test–retest reliability was excellent (CCC = 0.81–0.93). Correlations among goal, task, and bond were moderate‐to‐strong; correlations with technology acceptance/commitment and vocational self‐efficacy were low‐to‐moderate, as hypothesised. Ceiling effects were < 15%, with targeting skewed towards higher TA levels.
Discussion
The WAI–Tele‐Reha‐G shows preliminary evidence of satisfactory psychometric properties for a therapist‐reported assessment of TA in telerehabilitation. Interpretation should consider the modest sample, therapist‐only perspective, targeting skew, and Rasch adjustments. Future work should add more challenging items to improve discrimination at high TA levels and develop a patient‐reported version to capture the dyadic nature of TA.
Keywords: telerehabilitation, therapeutic alliance, validation studies as topic
1. Introduction
Telerehabilitation delivers remote rehabilitation via technology (Shem et al. 2022) to enhance accessibility and knowledge exchange among professionals, providers, and patients, and is effective for people with motor, neurological, and musculoskeletal conditions (Shem et al. 2022). Both synchronous (e.g., video, chat) and asynchronous methods (e.g., image/sensor data sharing) are used by physicians, psychologists, nurses, physiotherapists (PTs), occupational therapists (OTs), and speech and language therapists (SLTs) (Shem et al. 2022).
Therapeutic alliance (TA), the collaborative partnership between therapist and client (Bachelor and Horvath 1999), originates in psychotherapy (Hentschel 2005), with several conceptualisations (Bordin 1979; Fluckiger et al. 2018). Bordin's model highlights agreement on goals and tasks and the quality of the bond (Bordin 1979). Strong TA relates to positive outcomes in psychotherapy (Fluckiger et al. 2018) and rehabilitation (A. M. Hall et al. 2010). In our prior qualitative work with PTs, OTs, and SLTs, core TA themes in telerehabilitation included communication, trust, respect, and agreement on goals and tasks, alongside caregiver support for safety and connectivity, caregivers as co‐therapists, and strategies to enhance autonomy and motivation (Seebacher et al. 2024). These findings suggest that TA in telerehabilitation may involve relational and contextual elements that differ from traditional face‐to‐face rehabilitation settings. For example, remote care relies on technology‐mediated communication, management of technical disruptions, and, in some cases, active caregiver involvement in therapy delivery. In addition, the limited opportunities for physical presence and bodily touch may influence the development and maintenance of TA compared with face‐to‐face rehabilitation. Given TA's relevance for adherence and outcomes, appropriate TA measures are needed (Paap et al. 2019). A review identified four physiotherapy TA instruments and recommended the Working Alliance Inventory (WAI; Horvath and Greenberg 1989) as the most suitable (Gutiérrez‐Sánchez et al. 2021). The Physiotherapy Therapeutic Relationship Measure (P‐TREM) was recently developed for musculoskeletal outpatients with content validity assessed (McCabe et al. 2022), but telerehabilitation‐specific tools are lacking. We therefore adapted the German WAI–Short Revised therapist version (WAI‐SRT) for PTs, OTs, and SLTs in telerehabilitation and validated the WAI‐Tele‐Reha‐G (Hatcher and Gillaspy 2006; Horvath and Greenberg 1989). While the German WAI‐SRT was used due to the team's base in Austria and Germany, contextual (rehabilitation vs. telerehabilitation) and methodological considerations are broadly applicable.
2. Methods
2.1. Study Design
We conducted a sequential mixed‐methods study comprising qualitative questionnaire development followed by quantitative psychometric validation in German‐speaking PTs, OTs, and SLTs. Reporting follows COnsensus‐based Standards for the selection of health Measurement Instruments (COSMIN) guidelines (Supporting Information S7).
2.2. Procedure
The study proceeded in two phases (Figure 1). Phase 1 (15 May–31 October 2020) focused on instrument adaptation through expert interviews, cognitive interviewing and pilot testing, including a readability assessment. Phase 2 (13 March 2023–31 January 2024) examined psychometric properties. An expert panel (four PTs, one psychometrician, one neuroscientist; four PhD, two MSc degrees) oversaw instrument development and analysis.
FIGURE 1.

Flow diagram of study procedures. OTs, occupational therapists; PTs, physiotherapists; SLTs, speech and language therapists; T1, test; T2, retest; WAI‐Tele‐Reha‐G, Working‐Alliance‐Inventory‐Telerehabilitation‐German.
2.3. Participants
As the aim was to adapt and validate the therapist version of the WAI‐SRT for telerehabilitation, we included therapists only. For cognitive interviews, we used purposive quota sampling (Table 1) to recruit German‐speaking PTs, OTs and SLTs with telerehabilitation experience, plus therapists with additional psychology training (target n = 8–12 experts, defined by specialised professional knowledge (Gläser and Laudel 2010)).
TABLE 1.
Baseline characteristics of participants in questionnaire development and validation.
| Characteristic | Cognitive interview sample (n = 9) | Pilot test sample (n = 34) | Validation sample (n = 128) |
|---|---|---|---|
| Age (y) a | — | 34.9 (10.3) | 40.5 (11.0) |
| Sex, female b | 7 (77.8) | 27 (79.4) | 94 (73.4) |
| Professional experience (y) a | — | 10.5 (10.4) | 17.4 (14.4) |
| Profession b | |||
| PT c | 3 (33.3) | 24 (70.6) | 66 (51.6) |
| OT c | 2 (22.2) | 4 (11.8) | 40 (31.3) |
| SLT c | 2 (22.2) | 6 (17.6) | 22 (17.1) |
| PT & P | 2 (22.2) | 0 (0.0) | 0 (0.0) |
| Highest degree b | |||
| Master/doctorate | 9 (100.0) | 7 (20.6) | 47 (36.7) |
| Bachelor/diploma | 0 (0.0) | 27 (79.4) | 81 (63.3) |
| Professional setting b | |||
| Clinical | 3 (33.3) | 33 (97.1) | 68 (53.1) |
| Teaching/research | 4 (44.5) | 0 (0.0) | 4 (3.2) |
| Mixed | 2 (22.2) | 1 (2.9) | 56 (43.7) |
| Country of work b | |||
| Austria | 4 (44.5) | 13 (38.2) | 68 (53.1) |
| Germany | 3 (33.3) | 12 (35.3) | 32 (25.0) |
| Switzerland | 2 (22.2) | 9 (26.5) | 23 (18.0) |
| Austria and Germany | 0 (0.0) | 0 (0.0) | 5 (3.9) |
Abbreviations: M, mean; N, number; OT, occupational therapist; P, psychologist; PT, physiotherapist; SD, standard deviation; SLT, speech and language therapist; y: years.
Mean (standard deviation).
Frequencies.
These therapists had prior experience in telerehabilitation.
Pilot testing involved German‐speaking therapists with telerehabilitation experience (n = 34). For validation, convenience and snowball sampling targeted PTs, OTs, and SLTs practising telerehabilitation in Austria, Germany, and Switzerland. We aimed for n = 500 to maximise stability of Rasch calibrations and permit robust DIF testing, with a planned second sample if amendments were required (Linacre 1994). Recruitment used professional associations, online newsletters/ads, and outreach to hospitals, rehabilitation centres, health science universities, and freelance clinics.
2.4. Ethics Approval Statement and Study Registration
Development phases did not require ethics approval because only experts participated; written consent was obtained before cognitive interviews and pilot testing. Validation received approval from the Ethics Committee of the Danube Private University (DPU), Krems, Austria (date 4 February 2023; approval number GZ: DPU‐EK/026). The study complied with the EU GDPR (DSGVO 2016/679), Austrian (DSG 2019) and German (BDSG 2017) data protection laws, and the Declaration of Helsinki (2013).
WAI‐Tele‐Reha‐G development was prospectively registered (ISRCTN10132326, 02.05.2020); validation was prospectively registered in the DRKS (DRKS00030941, 06.02.2023).
2.5. Outcome Measures
Demographic and professional data included age, sex, profession, highest degree, country, professional years of experience, setting, and field. Convergent validity comparators were selected as adjacent constructs relevant to technology‐mediated care: technology acceptance/intention to use (Unified Theory of Acceptance and Use of Technology‐2 [UTAUT‐2] questionnaire), technology‐related commitment/confidence (Technology Commitment (TB) short scale), and vocational self‐efficacy (5‐item Revised Vocational Self‐Efficacy Scale [BSW‐5‐Rev]). We hypothesised low‐to‐moderate positive correlations with WAI–Tele‐Reha‐G scores. We therefore administered:
UTAUT‐2 questionnaire (German version) (Harborth and Pape 2018) adapted from Pokémon Go (Venkatesh et al. 2012) replacing ‘playing Pokémon Go’ with ‘using telerehabilitation’. It assesses habit, performance expectancy, effort expectancy, social influence, facilitating conditions, hedonic motivation, and behavioural intention (7‐point Likert; higher scores indicate greater willingness to adopt telerehabilitation; subscales only) (Harborth and Pape 2018). The German UTAUT‐2 questionnaire showed good to excellent internal consistency across subscales (Cronbach's α = 0.733–0.951), with evidence supporting convergent and discriminant validity as well as an appropriate factor structure (Harborth and Pape 2018).
TB short scale (Neyer et al. 2012): technology acceptance, technology competence confidence and technology control confidence (three 4‐item facets; 5‐point Likert; higher scores indicate greater willingness). The TB short scale demonstrated good internal consistency (α = 0.74–0.84) and good structural validity based on confirmatory factor analysis (Neyer et al. 2012).
5‐item Revised Vocational Self‐Efficacy Scale (BSW‐5‐Rev): vocational self‐efficacy (4‐point Likert; higher scores indicate greater vocational self‐efficacy; Knispel et al. 2021). BSW‐5‐Rev was used to examine professional self‐efficacy as relevant to TA. The BSW‐5‐Rev showed acceptable internal consistency (α = 0.73) and evidence of construct validity through expected associations with related psychological constructs and self‐efficacy measures (Knispel et al. 2021).
2.6. WAI‐Tele‐Reha‐G Development, Content Validity, and Readability Evaluation
Guided by prior qualitative work with 12 telerehabilitation experts (Seebacher et al. 2024) and adaptation guidelines (Tsang et al. 2017), three authors (C.G., G.D., B.S.) adapted the validated German WAI‐SRT to telerehabilitation. The original WAI‐SRT covers goal (items 1, 4, 6, 11), task (2, 8, 10, 12), and bond (3, 5, 7, 9) with 12 items rated from 1 (‘seldom’) to 5 (‘always’) (Hatcher and Gillaspy 2006; Horvath and Greenberg 1989); higher scores indicate better TA. Permission to use the WAI‐SRT was obtained (Horvath and Greenberg 1989). Iterative adaptation removed redundancy, introduced telerehabilitation‐specific content (e.g., technology, exercise safety, caregiver involvement, autonomy), and refined wording by consensus to yield version 0.1 (v0.1).
Cognitive interviews were conducted with 9 experts using think‐aloud and verbal probing (see Table S1 for an interview guide) (Knafl et al. 2007) to assess relevance, comprehensiveness and clarity of instructions, items and response options; this informed version 0.2.
Pilot testing of v0.2 with 34 telerehabilitation therapists included an online survey (SurveyMonkey Inc., San Mateo, CA, USA) and open‐ended feedback confirmed clarity; minor wording refinements (adaptation 3) produced the final WAI–Tele‐Reha‐G.
Readability was evaluated using the LIX (German, ‘Lesbarkeitsindex’; readability index) by Björnson (Lenhard and Lenhard 2014‐2022):
| (1) |
where > 60 = very difficult text, 50–60 = difficult, 40–50 = moderately difficult, and < 40 = easy to very easy (Lenhard and Lenhard 2014–2022). Readability was improved by simplifying wording (e.g., replacing long/technical terms where possible).
2.7. Data Collection for the WAI‐Tele‐Reha‐G Validation
Validation data were collected online (LamaPoll, Lamano GmbH & Co. KG, Berlin, Germany) and piloted before study launch. All measures were administered at baseline (T1); the WAI–Tele‐Reha‐G was repeated after 2–3 weeks (T2) to assess test–retest reliability.
2.8. Data Analysis
2.8.1. Internal Construct Validity
Rasch analysis was conducted in RUMM2030 (www.rummlab.com.au/) using the partial credit model (PCM) for polytomous items, as supported by significant likelihood ratio tests (Masters 1982). We examined threshold ordering, overall and item fit, unidimensionality, local independence (Rasch 1980), and differential item functioning (DIF) by sex, age (tertiles), profession, education, country and setting (Andrich 1988; Holland and Wainer 1993). Unidimensionality was assessed via t‐tests comparing person estimates from positive versus negative loadings on the first principal component of residuals (Tennant and Conaghan 2007). We expected non‐significant chi‐square statistics (Bonferroni‐adjusted, baseline p = 0.05) for item‐person and individual fits (Mills et al. 2016), standardised item fit residuals within ± 2.5 (99% CI), and residual mean approximates 0 and SD approximates 1 (Supporting Information S1; Mills et al. 2016). Local dependence (LD) was flagged by item residual correlations > 0.2 above the mean (Christensen et al. 2017). When LD occurred between conceptually overlapping items, we formed a testlet (superitem) by summing the dependent items for Rasch estimation to absorb local covariance while retaining content coverage (Supporting Information S1) (Jiao et al. 2013; Li et al. 2006). This adjustment was restricted to Rasch modelling; user‐facing scoring remained at the item level. DIF was evaluated using the analysis of variance of standardised residuals with Bonferroni adjustment.
2.8.2. Convergent and Known‐Groups Validity
We computed Spearman's rank correlations (ρ) with 95% CIs among WAI‐Tele‐Reha‐G scales and with UTAUT‐2, TB, and BSW‐5‐Rev. Based on prior TA research in physiotherapy (Linares‐Fernández et al. 2021) and rehabilitation (Paap et al. 2019), we expected moderate to strong positive associations among goal, task, and bond (ρ = 0.5–0.90) and low to moderate positive correlations with UTAUT‐2, TB, and BSW‐5‐Rev (ρ = 0.3–0.70) (Hinkle et al. 2003). Known‐groups validity was explored using Mann–Whitney U tests by sex, with Hodges–Lehmann median differences (95% CI). No sex differences were expected. Descriptive statistics (frequencies, mean [SD], median [IQR]) and validity analyses were conducted in IBM SPSS 29.0 (Armonk, NY: IBM Corp.). Two‐sided p < 0.05 was considered significant.
2.8.3. Reliability and Targeting
Internal consistency was estimated with Cronbach's alpha and the Person Separation Index (PSI), interpreted similarly with a minimum acceptable 0.7 (Fisher 1992). Test–retest reliability was assessed using Lin's concordance correlation coefficient (CCC) with 95% CIs utilising MedCalc (MedCalc Software Ltd, Ostend, Belgium; Lin 1989). Median (range) WAI‐Tele‐Reha‐G scores at test/retest were calculated. We computed the standard error of measurement (SEM) and minimum detectable change (MDC, 95% CI) and examined floor/ceiling effects (Supporting Information S1). Targeting was evaluated by inspecting person–item threshold distributions to assess alignment of item difficulty with respondent ability along the TA continuum (Pomeroy et al. 2020).
3. Results
3.1. Phase 1: WAI‐Tele‐Reha‐G Development
3.1.1. Adaptation 1
Expert interview results are reported elsewhere (Seebacher et al. 2024). Phase 1 data collection proceeded smoothly. In adapting the German WAI‐SRT for telerehabilitation, original items 1–3 and 5–7 were retained or slightly revised to emphasise respect and appreciation, focusing on key TA aspects relevant to both psychology and telerehabilitation. Additions included phrases related to telerehabilitation and new items addressing technology, exercise safety, caregivers as co‐therapists, and patient autonomy. Items 8–10 were removed due to redundancy or low relevance, and item 4 was revised to focus on the ‘goal’ rather than ‘task’. Modified and new items are shown in Tables S2 and S3, forming WAI‐Tele‐Reha‐G v0.1 (Supporting Information S2). All adaptations from versions 0.1 to 0.3 are reported in Tables [Link], [Link], [Link] and Supporting Information [Link], [Link], [Link].
3.1.2. Adaptation 2
Baseline characteristics of the cognitive interview sample (n = 9) are presented in Table 1. Experts rated WAI‐Tele‐Reha‐G v1.0 as clear with an appropriate item count. Response options were revised to always, often, sometimes, rarely, never. To simplify the layout, second‐page instructions were removed and wording such as ‘in telerehabilitation’ was standardised. A paragraph on relatives and a sentence clarifying the therapist's perspective were added (Supporting Information S3), resulting in v0.2.
3.1.3. Adaptation 3
Pilot test baseline characteristics are presented in Table 1. Feedback from 34 telerehabilitation therapists supported the clarity of response options and layout for v0.2 (Supporting Information S4). Wording changes included replacing ‘shows me’ with ‘makes me understand’ in items 7, 9, and 12, standardising ‘within telerehabilitation,’ and simplifying the relative‐related phrase; minor linguistic adjustments to instructions did not alter comprehension. See Supporting Information S5 for v0.3.
3.1.4. Readability
To enhance readability (LIX), key terms were adapted. Long words such as ‘telerehabilitation’, ‘within’, ‘regularly’, and ‘verbal’ increased the LIX; thus, ‘telerehabilitation’ was shortened to ‘TR’ and ‘within’ to ‘in.’ For 17 items (258 words; 17 sentences; 15.18 words/sentence; 58 long words), LIX = 37.68, indicating easy reading comparable to children's/young adult literature (Lenhard and Lenhard 2014‐2022).
3.2. Phase 2: WAI‐Tele‐Reha‐G Validation
Despite contacting around 90 hospitals/rehabilitation centres, 10 universities, and 240 clinics, and > 2300 survey starts, 128 therapists completed T1 and 84 completed T2. Baseline characteristics of the validation sample are presented in Table 1.
3.2.1. Internal Construct Validity
Rasch PCM provided preliminary support for the internal construct validity of the WAI–Tele‐Reha‐G subscales, with important caveats. The goal scale showed acceptable overall fit to the Rasch model with evidence of unidimensionality, local independence, and ordered thresholds. The bond scale similarly fits the model with unidimensionality, local independence, and ordered thresholds. For the task scale, initial analyses indicated minor misfit with local dependence and disordered thresholds for one item. The strongest residual correlation was between items 7 and 12, both addressing therapist instructional guidance and support in telerehabilitation. On conceptual and statistical grounds, these items were combined into a testlet (superitem) for Rasch estimation, which resolved LD, yielded ordered thresholds and satisfactory overall fit. This adjustment was confined to Rasch modelling; user‐facing scoring remained at the item level.
No evidence of DIF by sex, age (tertiles), profession, education, country or setting was detected in this sample; however, subgroup sizes were modest and these invariance findings should be considered provisional. Given the achieved sample and upper‐end targeting, item location standard errors (SEs) were larger than anticipated under the planned N = 500, ranging from approximately 0.14 to 0.30 logits across items (Table S5). By contrast, with larger samples and good targeting, item SEs in rating‐scale Rasch models are often around 0.06–0.10 logits, reflecting the 1/√N relationship between SE and sample size (Bond and Fox 2007; Hagell and Westergren 2016; Linacre 1994). Detailed item/subtest fit is shown in Table S5, and Table 2 summarises model fit against ideal values.
TABLE 2.
Rasch model fit and reliability of the WAI‐Tele‐Reha‐G.
| Analysis WAI‐Tele‐Reha‐G | Item residuals | Person residuals | Chi‐square test a | PSI b | Alpha | Unidimensionality c | Floor effects d | Ceiling effects d | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | SD | Mean | SD | Value (df) | p | % tests > 5% (95% CI) | |||||
| Goal scale | |||||||||||
| 5‐item scale | 0.14 | 1.04 | −0.33 | 0.99 | 10.09 (10) | 0.432 | 0.72 | 0.80 | 0.8 (3.0–4.6) | 0.8 | 13.3 |
| Task scale | |||||||||||
| 6‐item scale | 0.30 | 1.21 | −0.30 | 0.96 | 21.25 (12) | 0.046 | 0.74 | 0.81 | 8.6 (4.8–12.4) | 0.0 | 8.6 |
| Task component e | 0.05 | 1.53 | −0.48 | 1.11 | 26.10 (24) | 0.102 | 0.73 | 0.75 | 4.7 (2.6–6.9) | 0.0 | 8.6 |
| Bond scale | |||||||||||
| 6‐item scale | −0.33 | 1.09 | −0.31 | 0.56 | 18.08 (12) | 0.113 | 0.71 | 0.82 | 4.7 (0.9–8.5) | 0.0 | 12.5 |
| Ideal values | 0.00 | 1.00 | 0.00 | 1.00 | — | > 0.01 f | ≥ 0.70 | > 0.70 | LCI < 5.0 | < 15 | < 15 |
Abbreviations: Alpha, Cronbach’s alpha; CI, confidence interval; df, degrees of freedom; LCI, lower bound of the 95% CI; PSI, person separation index; SD, standard deviation.
The Chi‐square test is used for the item‐trait interaction.
Based on independent samples t‐tests, comparing person residuals loading positively and negatively on the first principal component (95% CI).
Values represent percentages.
Combination of one superitem (items 7 and 12) and original items 4, 8, 9, and 14.
Bonferroni‐adjusted based on the number of items; with 5 items, p > 0.01, with 6 items, p > 0.008.
3.2.2. Convergent and Known‐Groups Validity
Table 3 shows moderate to strong positive correlations among the WAI‐Tele‐Reha‐G goal, task, and bond scales (ρ = 0.67–0.86), as expected for related facets of TA. As hypothesised, low‐to‐moderate positive correlations were observed with technology acceptance/intention to use (UTAUT2), technology‐related confidence (Technology Commitment) and vocational self‐efficacy (BSW‐5‐Rev). Known‐groups testing revealed no significant sex differences in TA (Tables S6 and S7).
TABLE 3.
Convergent validity and internal consistency of WAI‐Tele‐Reha‐G and comparator scales were assessed for the study sample.
| Other scales | Alpha | Spearman's rank correlation coefficients a | ||
|---|---|---|---|---|
| Goal scale | Task scale | Bond scale | ||
| Goal scale | 0.803 | — | 0.799*** | 0.765*** |
| Task scale | 0.809 | 0.799*** | — | 0.757*** |
| Bond scale | 0.800 | 0.765*** | 0.757*** | — |
| Unified theory of acceptance and use of technology‐2 questionnaire | ||||
| Habit | 0.776 | 0.315ns | 0.255ns | 0.149ns |
| Performance expectancy | 0.885 | 0.426*** | 0.427*** | 0.370*** |
| Effort expectancy | 0.849 | 0.404*** | 0.420*** | 0.336** |
| Social influence | 0.782 | 0.262* | 0.299** | 0.137ns |
| Hedonic motivation | 0.693 | 0.222ns | 0.280* | 0.208ns |
| Facilitating conditions | 0.760 | 0.372** | 0.335** | 0.359** |
| Behavioural intention | 0.977 | 0.380*** | 0.362** | 0.321** |
| Short scale for measuring technology commitment | ||||
| Technology acceptance | 0.877 | 0.201ns | 0.273* | 0.239ns |
| Technology competence confidence | 0.886 | 0.306* | 0.347** | 0.307* |
| Technology control confidence | 0.809 | 0.415*** | 0.457*** | 0.376*** |
| BSW‐5‐Rev | 0.719 | 0.512*** | 0.470*** | 0.416*** |
Abbreviations: Alpha, Cronbach's alpha; BSW‐5‐Rev, 5‐item Revised Vocational Self‐Efficacy Scale.
*** Correlation is significant at the < 0.0001 level; ** at < 0.01 level; * at < 0.05 level; ns indicates not significant (2‐tailed, with p‐values corrected for 14 comparisons).
3.2.3. Reliability and Targeting
Internal consistency was good (Cronbach's α = 0.80–0.81), and the PSI indicated modest person separation (0.72–0.74) (Table 2). Measurement error indices were acceptable, and test–retest reliability over 2–3 weeks was excellent (Lin's CCC = 0.81–0.93) (Table 4). Standard error of measurement and minimum detectable change values are reported in Table 4.
TABLE 4.
Reliability of the WAI‐Tele‐Reha‐G.
| WAI‐Tele‐Reha‐G | Goal scale | Task scale | Bond scale |
|---|---|---|---|
| Measurement error | |||
| SEM | 0.76 | 0.65 | 0.93 |
| MDC points/max score | 2.87/20 | 2.96/24 | 2.94/24 |
| MDC percentage | 14.23 | 12.34 | 12.25 |
| Test‐retest reliability | |||
| CCC per Lin (95% CI) | 0.87 (0.81–0.91) | 0.88 (0.83–0.92) | 0.89 (0.85–0.93) |
| Pearson ρ a | 0.88 | 0.89 | 0.91 |
| Factor Cb b | 0.99 | 0.99 | 0.99 |
| Values for T1 c | 21 (5–25) | 25 (6–30) | 28 (6–30) |
| Values for T2 c | 20 (5–25) | 24 (6–30) | 28 (6–30) |
Abbreviations: CI, confidence interval; MDC, minimum detectable change; SEM, standard error of measurement; T1, test 1; T2, test 2, retest (2–3 weeks after T1).
Pearson ρ: Precision measure for the concordance correlation coefficient (CCC) per Lin.
Bias correction factor Cb: CCC accuracy per Lin.
Median (minimum‐maximum).
Floor effects were absent or negligible; ceiling effects were below 15% across subscales (Table 2). However, person–item threshold distribution maps showed a clear upper‐end skew: mean person locations were above the scale average for each subscale, respondents clustered at higher TA levels, and there were relatively few high‐difficulty thresholds (Figure 2). These patterns indicate reduced discrimination at the upper end of the continuum and likely contributed to PSI being lower than Cronbach's alpha. In contexts where TA is generally rated highly, sensitivity to between‐therapist differences and to small improvements among therapists already reporting strong TA may therefore be limited. To aid interpretation, a transformation table is provided to convert raw scores to interval‐scaled estimates (complete datasets required; Table S8), and minimum detectable change (MDC95) values are reported (Table 4). The WAI–Tele‐Reha‐G is provided in Supporting Information S6.
FIGURE 2.

Targeting of the WAI‐Tele‐Reha‐G Goal (A), task (B), and bond (C) scales. Person–item threshold distribution maps showing the distribution of therapists' therapeutic alliance (TA) levels (upper histograms) relative to item threshold locations (lower panels) along the latent therapeutic alliance continuum (logits). Closer alignment between person and item distributions reflects better targeting of the scale to the sample, whereas clustering of respondents at higher levels may indicate reduced discrimination at high therapeutic alliance levels.
4. Discussion
We adapted the German WAI‐SRT for use in telerehabilitation across PT, OT, and SLT and undertook an initial psychometric evaluation of the WAI–Tele‐Reha‐G. Following iterative content development and minor Rasch adjustments, the instrument showed preliminary evidence of internal construct validity, acceptable reliability indices, and the expected pattern of correlations with theoretically adjacent constructs. Given the modest sample and therapist‐only perspective, these findings should be interpreted as early‐stage.
The findings provided preliminary support for internal construct validity. The goal, task and bond scales met key Rasch expectations for overall fit, unidimensionality, LD (after adjustment), ordered thresholds and lack of DIF by the examined groups in this sample. The task scale initially displayed LD and disordered thresholds for one item. The strongest residual correlation occurred between two items that both assess therapist instructional guidance and support within telerehabilitation. On conceptual and statistical grounds (Marais and Andrich 2008), we formed a testlet (superitem) by summing these two items for Rasch estimation, which resolved LD and yielded ordered thresholds and satisfactory model fit. This adjustment was confined to Rasch modelling; practical scoring remains at the item level. While testlet formation is an accepted pragmatic approach (Wang and Wilson 2005), future work should examine whether rewording or replacing one of the overlapping items can remove the need for a testlet and further strengthen the scale.
Convergent validity was consistent with a priori expectations. Correlations among the WAI–Tele‐Reha‐G subscales were moderate to strong, and associations with technology acceptance (UTAUT2), technology commitment and vocational self‐efficacy were low to moderate. These comparators were selected as adjacent constructs expected to influence alliance in technology‐mediated care—e.g., communication fluency under technical constraints and confidence in collaboratively negotiating tasks and goals remotely—without being isomorphic with TA. The observed pattern supports conceptual relatedness without redundancy. Future studies should incorporate more direct comparators (e.g., therapist‐reported TA in face‐to‐face contexts and patient‐reported TA).
Reliability indices were satisfactory overall. Internal consistency was good (Cronbach's alpha 0.80–0.81), and test–retest reliability was excellent (concordance correlation coefficient 0.81–0.93). Person Separation Index (0.72–0.74) was modest, which likely reflects targeting (Andrich 1982). Person–item maps indicated clustering of respondents at higher levels of TA and relatively fewer difficult items. Although ceiling effects were below the commonly used 15% threshold, this upper‐end skew implies reduced discrimination among therapists who already report very strong alliances and may attenuate sensitivity to small improvements in that range. Until more challenging items are added or response anchors refined, we recommend cautious interpretation of high‐end scores and the use of interval‐scaled estimates and minimum detectable change values when monitoring change over time.
Several physiotherapy alliance instruments—predominantly patient‐reported—have been developed using traditional psychometric methods, with mixed results (Gutiérrez‐Sánchez et al. 2021). For example, a Rasch analysis of the Working Alliance Theory of Change Inventory supported unidimensionality after removing seven items but reported ceiling effects, indicating scope for further refinement (A. Hall et al. 2012). The Physiotherapy Therapeutic Relationship Measure offers a theory‐based framework, although its psychometric properties remain to be evaluated (McCabe et al. 2022). Our work extends this literature by providing a therapist‐reported, telerehabilitation‐specific instrument evaluated with Rasch modelling, while acknowledging that additional development is required to improve high‐end discrimination and to incorporate patient perspectives.
Several limitations warrant a cautious interpretation. Despite extensive outreach, the achieved sample (N = 128; retest N = 84) was below the planned n = 500 intended to maximise stability of Rasch calibrations and permit robust DIF testing. Under‐recruitment relative to the n = 500 target likely widened item and person standard errors, attenuated person separation, and reduced power for subgroup and DIF analyses; therefore, invariance and subgroup findings should be considered provisional (Bond and Fox 2007). In addition, this study validated only the therapist version. Because TA is inherently dyadic, therapist ratings may not fully capture patients' perceptions and experiences in telerehabilitation, limiting ecological validity. Developing and validating a complementary patient‐reported version is a priority. Taken together, these findings suggest that the WAI–Tele‐Reha‐G has promise as a therapist‐reported instrument for formative applications in telerehabilitation—such as reflective practice, training and group‐level service evaluation. At this stage, we advise caution when using the measure to differentiate very high levels of alliance or to inform individual patient‐level decisions.
Future work should (1) recruit larger and more diverse samples to stabilise item calibrations and enable robust DIF and subgroup analyses across professions, countries and settings; (2) refine the task items to reduce local dependence and re‐evaluate the need for a testlet; (3) add more challenging items and/or optimise response anchors to improve discrimination at the upper end of the continuum; (4) develop and validate a patient‐reported version to enable dyadic assessment and stronger convergent validity testing; and (5) pursue cross‐linguistic calibration and pooling to support international comparability. In parallel, we will support translation, cultural adaptation and cross‐linguistic Rasch analyses through data pooling to enhance international comparability.
5. Implications on Physiotherapy Practice
The WAI–Tele‐Reha‐G is a therapist‐reported instrument that, based on preliminary validation, may be useful for reflective practice, supervision and training, and group‐level service evaluation in telerehabilitation across PT, OT, and SLT. Clinicians can use the goal, task and bond subscales to identify strengths and areas for development in communication, collaborative goal setting and task guidance when working remotely. Given under‐recruitment in the validation sample, therapist‐only data, and some skew towards higher scores, we recommend cautious interpretation—particularly when differentiating very high levels of TA or making individual patient‐level decisions. Where monitoring over time is intended, interval‐scaled scores (via the transformation table) and minimum detectable change values should be used. At this stage, the measure is best applied for formative purposes (e.g., quality improvement, team feedback, educational curricula) rather than as a stand‐alone clinical decision aid. Further work, including larger and more diverse samples, refinement to improve discrimination at the upper end of the scale, and development of a patient‐reported version to enable dyadic assessment, is needed before broader clinical implementation.
Funding
Parts of this work were supported by VASCage GmbH, Innsbruck, Austria (first questionnaire adaptation in Study Phase 1).
Ethics Statement
Development phases did not require ethics approval because only experts participated; written consent was obtained before cognitive interviews and pilot testing. Validation received approval from the Ethics Committee of the Danube Private University (DPU), Krems, Austria (date 4 February 2023; approval number GZ: DPU‐EK/026). The study complied with the EU GDPR (DSGVO 2016/679), Austrian (DSG 2019) and German (BDSG 2017) data protection laws, and the Declaration of Helsinki (2013). WAI‐Tele‐Reha‐G development was prospectively registered (ISRCTN10132326, 02.05.2020) and validation was prospectively registered in the DRKS (DRKS00030941, 06.02.2023).
Consent
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Supporting Information S1
Supporting Information S2
Supporting Information S3
Supporting Information S4
Supporting Information S5
Supporting Information S6
Supporting Information S7
Table S1: Cognitive interview guide.
Table S2: Initial modifications to the English original WAI‐SRT for telerehabilitation adaptation (Adaptation 1A).
Table S3: Newly formulated items with justification (Adaptation 1B).
Table S4: Modifications implemented based on expert feedback, along with justifications for the changes (Adaptation 2).
Table S5: Fit of individual WAI‐Tele‐Reha‐G items and subtests to the Rasch model.
Table S6: Convergent validity and internal consistency of WAI‐Tele‐Reha‐G and comparator scales assessed for the study sample.
Table S7: Known‐groups validity of the WAI‐Tele‐Reha‐G.
Table S8: Conversion of raw scores to interval‐scale latent estimates for the WAI‐Tele‐Reha‐G.
Acknowledgements
We sincerely thank all the professional associations that shared the survey links for therapists. We also appreciate the experts and therapists for their significant contributions to this study. Open Access funding provided by Medizinische Universitat Innsbruck.
Hoffmann, Laura , Geimer Carole, Diermayr Gudrun, Reindl Markus, Horton Mike C., and Seebacher Barbara. 2026. “Development and Validation of the WAI‐Tele‐Reha‐G: A Measure of Therapeutic Alliance in Telerehabilitation,” Physiotherapy Research International: e70309. 10.1002/pri.70309.
Laura Hoffmann and Carole Geimer contributed equally to this work.
Data Availability Statement
The German dataset is available for pooling to aid in translating and adapting the WAI‐Tele‐Reha‐G for Rasch analysis upon contacting the corresponding author and signing a data sharing contract, enabling assessment of language version equivalence.
References
- Andrich, D. 1982. “An Index of Person Separation in Latent Trait Theory, the Traditional KR20 Index, and the Guttman Scale Response Pattern.” Education Research & Perspectives 9: 95–104. [Google Scholar]
- Andrich, D. 1988. Rasch Models for Measurement Series: Quantitative Applications in the Social Sciences No. 68. Sage. [Google Scholar]
- Bachelor, A. , and Horvath A.. 1999. “The Therapeutic Relationship.” In The Heart and Soul of Change: What Works in Therapy, edited by Hubble M. A., Duncan B., and Miller S. D., 299–307. American Psychological Association. [Google Scholar]
- Bond, T. G. , and Fox C. M.. 2007. Applying the Rasch Model: Fundamental Measurement in the Human Sciences. 2nd ed. Lawrence Erlbaum Associates. [Google Scholar]
- Bordin, E. 1979. “The Generalizability of the Psychoanalytic Concept of the Working Alliance.” Psychotherapy 16, no. 3: 252–260. 10.1037/h0085885. [DOI] [Google Scholar]
- Christensen, K. B. , Makransky G., and Horton M.. 2017. “Critical Values for Yen’s Q3: Identification of Local Dependence in the Rasch Model Using Residual Correlations.” Applied Psychological Measurement 41, no. 3: 178–194. 10.1177/0146621616677520. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fisher, W. J. 1992. “Reliability, Separation, Strata Statistics.” Rasch Measurement Transactions 6, no. 3: 238. [Google Scholar]
- Fluckiger, C. , Del Re A. C., Wampold B. E., and Horvath A. O.. 2018. “The Alliance in Adult Psychotherapy: A Meta‐Analytic Synthesis.” Psychotherapy 55, no. 4: 316–340. 10.1037/pst0000172. [DOI] [PubMed] [Google Scholar]
- Gläser, J. , and Laudel G.. 2010. Experteninterviews und Qualitative Inhaltsanalyse als Instrumente Rekonstruierender Untersuchungen. 4 ed. VS Verlag. [Google Scholar]
- Gutiérrez‐Sánchez, D. , Pérez‐Cruzado D., and Cuesta‐Vargas A. I.. 2021. “Systematic Review of Therapeutic Alliance Measurement Instruments in Physiotherapy.” Physiotherapie Canada 73, no. 3: 212–217. 10.3138/ptc-2019-0077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hagell, P. , and Westergren A.. 2016. “Sample Size and Statistical Conclusions From Tests of Fit to the Rasch Model According to the Rasch Unidimensional Measurement Model (Rumm) Program in Health Outcome Measurement.” Journal of Applied Measurement 17, no. 4: 416–431. PMID: 28009589. [PubMed] [Google Scholar]
- Hall, A. , Ferreira M., Clemson L., Ferreira P., Latimer J., and Maher C.. 2012. “Assessment of the Therapeutic Alliance in Physical Rehabilitation: A RASCH Analysis.” Disability & Rehabilitation 34, no. 3: 257–266. 10.3109/09638288.2011.606344. [DOI] [PubMed] [Google Scholar]
- Hall, A. M. , Ferreira P. H., Maher C. G., Latimer J., and Ferreira M. L.. 2010. “The Influence of the Therapist‐Patient Relationship on Treatment Outcome in Physical Rehabilitation: A Systematic Review.” Physical Therapy 90, no. 8: 1099–1110. 10.2522/ptj.20090245. [DOI] [PubMed] [Google Scholar]
- Harborth, D. , and Pape S.. 2018. German Translation of the Unified Theory of Acceptance and Use of Technology 2 (UTAUT2) Questionnaire.
- Hatcher, R. L. , and Gillaspy J. A.. 2006. “Development and Validation of a Revised Short Version of the Working Alliance Inventory.” Psychotherapy Research 16, no. 1: 12–25. 10.1080/10503300500352500. [DOI] [Google Scholar]
- Hentschel, U. 2005. “Die Therapeutische Allianz.” Psychotherapeut 50, no. 5: 305–317. 10.1007/s00278-005-0440-3. [DOI] [Google Scholar]
- Hinkle, D. E. , Wiersma W., and Jurs S. G.. 2003. Applied Statistics for the Behavioral Sciences. 5th ed. ed. Houghton Mifflin. [Google Scholar]
- Holland, P. W. , and Wainer H.. 1993. Differential Item Functioning. Lawrence Erlbaum. [Google Scholar]
- Horvath, A. O. , and Greenberg L. S.. 1989. “Development and Validation of the Working Alliance Inventory.” Journal of Counseling Psychology 36, no. 2: 223–233. 10.1037/0022-0167.36.2.223. [DOI] [Google Scholar]
- Jiao, H. , Wang S., and He W.. 2013. “Estimation Methods for One‐Parameter Testlet Models.” Journal of Educational Measurement 50, no. 2: 186–203. 10.1111/jedm.12010. [DOI] [Google Scholar]
- Knafl, K. , Deatrick J., Gallo A., et al. 2007. “The Analysis and Interpretation of Cognitive Interviews for Instrument Development.” Research in Nursing & Health 30, no. 2: 224–234. 10.1002/nur.20195. [DOI] [PubMed] [Google Scholar]
- Knispel, J. , Wittneben L., Slavchova V., and Arling V.. 2021. “Skala Zur Messung der Beruflichen Selbstwirksamkeitserwartung (BSW‐5‐Rev).” In Zusammenstellung Sozialwissenschaftlicher Items und Skalen (ZIS). 10.6102/zis303. [DOI] [Google Scholar]
- Lenhard, W. , and Lenhard A.. 2014–2022. “Berechnung des Lesbarkeitsindex LIX Nach Björnson.” http://www.psychometrica.de/lix.html.
- Li, Y. , Bolt D. M., and Fu J.. 2006. “A Comparison of Alternative Models for Testlets.” Applied Psychological Measurement 30, no. 1: 3–21. 10.1177/0146621605275414. [DOI] [Google Scholar]
- Lin, L. I.‐K. 1989. “A Concordance Correlation Coefficient to Evaluate Reproducibility.” Biometrics 45, no. 1: 255–268. 10.2307/2532051. [DOI] [PubMed] [Google Scholar]
- Linacre, J. M. 1994. “Sample Size and Item Calibration Stability.” Rasch Measurement Transactions 7, no. 4: 328. [Google Scholar]
- Linares‐Fernández, M. T. , La Touche R., and Pardo‐Montero J.. 2021. “Development and Validation of the Therapeutic Alliance in Physiotherapy Questionnaire for Patients With Chronic Musculoskeletal Pain.” Patient Education and Counseling 104, no. 3: 524–531. 10.1016/j.pec.2020.09.024. [DOI] [PubMed] [Google Scholar]
- Marais, I. , and Andrich D.. 2008. “Formalizing Dimension and Response Violations of Local Independence in the Unidimensional Rasch Model.” Journal of Applied Measurement 9, no. 3: 200–215. [PubMed] [Google Scholar]
- Masters, G. N. 1982. “A Rasch Model for Partial Credit Scoring.” Psychometrika 47, no. 2: 149–174. 10.1007/BF02296272. [DOI] [Google Scholar]
- McCabe, E. , Miciak M., Roduta Roberts M., et al. 2022. “Development of the Physiotherapy Therapeutic Relationship Measure.” European Journal of Physiotherapy 24, no. 5: 287–296. 10.1080/21679169.2020.1868572. [DOI] [Google Scholar]
- Mills, R. J. , Tennant A., and Young C. A.. 2016. “The Neurological Sleep Index: A Suite of New Sleep Scales for Multiple Sclerosis.” Multiple Sclerosis Journal—Experimental, Translational and Clinical 2: 2055217316642263. 10.1177/2055217316642263. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Neyer, F. J. , Felber J., and Gebhardt C.. 2012. “Entwicklung und validierung einer kurzskala zur erfassung von technikbereitschaft. [Development and validation of a brief measure of technology commitment.].” Diagnostica 58, no. 2: 87–99. 10.1026/0012-1924/a000067. [DOI] [Google Scholar]
- Paap, D. , Schrier E., and P.U D.. 2019. “Development and Validation of the Working Alliance Inventory Dutch Version for Use in Rehabilitation Setting.” Physiotherapy Theory and Practice 35, no. 12: 1292–1303. 10.1080/09593985.2018.1471112. [DOI] [PubMed] [Google Scholar]
- Pomeroy, I. M. , Tennant A., Mills R. J., Young C. A., and Group T. O. S.. 2020. “The WHOQOL‐BREF: A Modern Psychometric Evaluation of Its Internal Construct Validity in People With Multiple Sclerosis.” Quality of Life Research 29, no. 7: 1961–1972. 10.1007/s11136-020-02463-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rasch, G. 1980. Probabilistic Models for Some Intelligence and Attainment Tests. University of Chicago Press. [Google Scholar]
- Seebacher, B. , Geimer C., Neu J., Schwarz M., and Diermayr G.. 2024. “Identifying Central Elements of the Therapeutic Alliance in the Setting of Telerehabilitation: A Qualitative Study.” PLoS One 19, no. 3: e0299909. 10.1371/journal.pone.0299909. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Shem, K. , Irgens I., and Alexander M.. 2022. “Chapter 2—Getting Started: Mechanisms of Telerehabilitation.” In Telerehabilitation, edited by Alexander M., 5–20. Elsevier. [Google Scholar]
- Tennant, A. , and Conaghan P. G.. 2007. “The Rasch Measurement Model in Rheumatology: What Is It and Why Use It? When Should It Be Applied, and What Should One Look for in a Rasch Paper?” Arthritis & Rheumatism 57, no. 8: 1358–1362. 10.1002/art.23108. [DOI] [PubMed] [Google Scholar]
- Tsang, S. , Royse C. F., and Terkawi A. S.. 2017. “Guidelines for Developing, Translating, and Validating a Questionnaire in Perioperative and Pain Medicine.” Supplement, Saudi Journal of Anaesthesia 11, no. S1: S80–S89. 10.4103/sja.SJA_203_17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Venkatesh, V. , Thong J. Y. L., and Xu X.. 2012. “Consumer Acceptance and Use of Information Technology: Extending the Unified Theory of Acceptance and Use of Technology.” MIS Quarterly 36, no. 1: 157–178. 10.2307/41410412. [DOI] [Google Scholar]
- Wang, W.‐C. , and Wilson M.. 2005. “The Rasch Testlet Model.” Applied Psychological Measurement 29, no. 2: 126–149. 10.1177/0146621604271053. [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supporting Information S1
Supporting Information S2
Supporting Information S3
Supporting Information S4
Supporting Information S5
Supporting Information S6
Supporting Information S7
Table S1: Cognitive interview guide.
Table S2: Initial modifications to the English original WAI‐SRT for telerehabilitation adaptation (Adaptation 1A).
Table S3: Newly formulated items with justification (Adaptation 1B).
Table S4: Modifications implemented based on expert feedback, along with justifications for the changes (Adaptation 2).
Table S5: Fit of individual WAI‐Tele‐Reha‐G items and subtests to the Rasch model.
Table S6: Convergent validity and internal consistency of WAI‐Tele‐Reha‐G and comparator scales assessed for the study sample.
Table S7: Known‐groups validity of the WAI‐Tele‐Reha‐G.
Table S8: Conversion of raw scores to interval‐scale latent estimates for the WAI‐Tele‐Reha‐G.
Data Availability Statement
The German dataset is available for pooling to aid in translating and adapting the WAI‐Tele‐Reha‐G for Rasch analysis upon contacting the corresponding author and signing a data sharing contract, enabling assessment of language version equivalence.
