Skip to main content
International Journal of Sports Physical Therapy logoLink to International Journal of Sports Physical Therapy
. 2020 Aug;15(4):487–500.

THE RELIABILITY OF CLINICAL BALANCE TESTS UNDER SINGLE-TASK AND DUAL-TASK TESTING PARADIGMS IN UNINJURED ACTIVE YOUTH AND YOUNG ADULTS

Thaer S Manaseer 1,2,3, Jackie L Whittaker 1,3,4,5, Codi Isaac 3,6, Kathryn Schneider 7,8,9,10,11, Mary Roduta Roberts 1 2, Douglas P Gross 1,3
PMCID: PMC7735688  PMID: 33354382

Abstract

Background:

Previous researchers have suggested that balance control deficits are detected more accurately with dual-task testing than single-task testing. However, it is necessary to examine the clinimetric properties of dual-task testing before employing it in clinical and research settings.

Objective:

To examine and compare the relative and absolute reliability of the Balance Error Scoring System (BESS), Tandem Gait Test (TGT), and Clinical Reaction Time (CRT) under single and dual-task conditions in uninjured active youth and young adults.

Study Design: Single-group, repeated-measures study.

Methods:

Twenty-three individuals [9 female; median age 17 years] completed three trials of the BESS, TGT, and CRT under single and dual-task testing conditions during testing session one. Two raters assessed participants to assess inter-rater reliability. Either later on the same day or the following day, the protocol was repeated by one rater to assess intra-rater reliability. The average of three trials was used to calculate intra-rater (between-session) and inter-rater (within-session) intraclass correlation coefficient (ICC), standard error of measurement (SEM), minimal detectable change (MDC), and Cohen's Kappa coefficient for tests as appropriate under both conditions. Bland-Altman plots (mean difference and 95% limits of agreement) were used to assess for a systematic error associated with a learning effect.

Results:

Only one participant attended the second session on the following day, while 22 participants (95%) attended the second session within four hours after testing session one. Under single-task testing, estimated ICCs, SEMs, MDCs, and Kappa coefficients ranged from 0.24 to 0.99, 0.3 to 23, 0.8 to 64, and 0.03 to 0.64, respectively. Under dual-task testing, estimated ICCs, SEMs, MDCs, and Kappa coefficients ranged from 0.70 to 0.99, 0.4 to 17, 1.1 to 47, and 0.39 to 0.83, respectively. A learning effect was identified for all tests under all conditions.

Conclusion:

The BESS is the only clinical test that demonstrated acceptable reliability for clinical use under single-task testing conditions. The BESS, TGT, and CRT all demonstrated acceptable reliability for clinical use under dual-task testing conditions. A practice session should be used to reduce the possible learning effect seen. Further studies examining sources of the systematic error observed are needed.

Level of Evidence:

2b.

Keywords: Adolescent, gait, psychometrics, reaction time, sport, young adult, movement system.

INTRODUCTION

Non-instrumented assessment of balance control is a common practice in clinical and on-field sports medicine, rehabilitation, and training settings.1For example, the Balance Error Scoring System (BESS),2Tandem Gait Test (TGT),3and Clinical Reaction Time (CRT)4are frequently used in baseline pre-season testing and after concussion to examine static balance, dynamic balance, and reaction time, respectively.5

The clinimetric properties of the BESS, TGT, and CRT have been examined to varying degrees. The BESS involves three stances, double, single and tandem. Each stance takes 20 seconds to complete and is performed on both a firm and unstable surface. The intra- and inter-rater reliability of the BESS has been previously examined in uninjured children, youth, and adult athletes with observed intra-class correlation coefficient (ICC) ranging from 0.57 to 0.98.2,6-9The BESS has also demonstrated varying degrees of criterion-related validity and concurrent validity in comparison to kinematic measures of postural sway in uninjured adult male (r=0.3-0.79, p<0.01)10and adolescent (r = 0.54, p = 0.001)11athletes. However, Quatman-Yates et al12suggested that the BESS may be limited for producing accurate assessments of postural control abilities in young athletes with concussion.

The TGT involves walking in a forward direction as accurately and quickly as possible down and back along a 38mm-wide three-meter line, with an alternate foot heel-to-toe gait. Despite the wide usage of the TGT for dynamic balance assessment, it is difficult to synthesize the test's clinimetric properties as there is currently no standardized testing protocol. The intra- and inter-rater reliability of the TGT protocol described by Koyama et al.13has been previously examined in uninjured adults with ICCs ranging from 0.70 to 0.95. The TGT has also demonstrated evidence of concurrent-validity (r>0.67, p<0.01) with Timed Up and Go test scores in uninjured adults.13

Finally, the CRT involves catching a falling numbered-rod as quickly as possible. The drop distance is then converted to speed. The intra- and inter-rater reliability of the CRT has been previously examined in uninjured athletes with ICCs ranging between 0.74 and 0.76.14The CRT has also demonstrated evidence of criterion-related validity with computerized reaction times (r=0.54, p-value not provided) in uninjured athletes..14

When a more accurate assessment of balance is needed, as in research settings, instrumented assessment of balance control using laboratory measures have been employed.15Commonly used testing paradigms across studies involving instrumented assessments included single- and dual-task testing paradigms.16-20For single-task testing, examined individuals are asked to control their balance without performing a concurrent cognitive task. For dual-task testing, examined individuals are asked to control their balance while performing a concurrent cognitive task. The authors of these studies demonstrated that adding a concurrent cognitive task to a balance task can provide important information about balance control impairments that may not be identified with single-task testing.16-20Further, a recent systematic review21reported that dual-task testing identified balance impairments later in the recovery period following concussion than single-task testing. Although instrumented assessments are not clinically feasible given the time, cost and need for specialized equipment and technicians, the translation of the dual-task paradigm to clinical balance testing through the addition of a cognitive task to tests such as the BESS,22TGT,23and CRT14may improve the robustness of clinical balance assessment.

Before employing the dual-task BESS, TGT, and CRT in clinical or clinical research settings, it is necessary to examine their clinimetric properties (i.e., reliability and validity).24The primary objective of this study was to examine the relative and absolute reliability of the dual-task BESS, TGT and CRT in a sample of uninjured active youth and young adults. It was hypothesized that dual-task testing would demonstrate acceptable reliability for clinical use. A secondary objective of this study was to compare the reliability of the BESS, TGT, and CRT under single-task versus dual-task testing.

METHODS

Design

This was a single-group, repeated-measures study examining the relative and absolute intra-rater (between-sessions) and inter-rater (within-session) reliability of three clinical tests of balance control under single and dual-task conditions. Relative reliability is the degree to which tested individuals maintain their position in a sample with repeated measurements. Absolute reliability, on the other hand, is the degree to which scores on repeated measurements vary for tested individuals.25Ethics approval (No: Pro00077091) was acquired from the University of Alberta Health Research Ethics Board, and informed consent and/or assent was obtained from all participants prior to testing as appropriate.

Participants

Participants included a convenience sample of uninjured active individuals who were 13 – 24 years old. ‘Active’ was operationalized as Cincinnati Sports Activity Scale level one or two.26Participants were recruited from local sport organizations and through advertisements, social media, and word of mouth. Participants were excluded if they were not active in recreational or competitive sport; had suffered a concussion within the prior 12 months; reported a lower extremity injury that resulted in time lost from recreational/sport activities for greater than one week within the prior three months, had an inner ear or sinus infection over the week prior to testing, had an uncorrectable (i.e., neither with vision glasses nor contacts) vision condition at time of testing, had a history of cognitive deficits including concentration abnormalities, history of attention deficit hyperactivity disorder; or were non-English speakers. Sample size was estimated based on guidelines provided by Walter et al.27Based on previously reported reliability of the BESS (ICC = 0.87),6twenty-one participants were needed for reliability analysis using two repetitions to achieve a power of 80% with alpha of 0.05 for clinically acceptable reliability (ICC = 0.6).28

Procedures

All data were collected at a private physiotherapy clinic over two testing sessions. Two physical therapists (TM, CI) rated individual participant's performance of three trials of the BESS, TGT and CRT under single and dual-task conditions. Each of the two physical therapists had more than five years of experience administering the BESS, TGT, and CRT in clinical settings. They met before data collection to review and discuss the test instructions and scoring procedures. At testing session one, participants were asked to complete a study questionnaire that gathered information about demographics and medical history. Participants were then familiarized with testing procedures before data collection started. Next, the two physical therapists independently rated and recorded individual participant's performance simultaneously to evaluate the inter-rater reliability of the BESS, TGT and CRT under single and dual-task conditions. One physical therapist (TM) provided tests’ instructions for all participants. Participants performed the BESS, TGT, and then CRT under both the single- and dual-task testing conditions, with the single-task testing condition performed first. Participants were given a one-minute rest between trials to minimize fatigue. Consistent with a guideline for reliability research design,29one of the physical therapists (TM) repeated testing of all participants for all tasks under all conditions either later on the same day or the following day to evaluate the intra-rater reliability. The BESS, TGT, and CRT were administered in the same order as in session one. The two physical therapists followed a specific script for the BESS, TGT, and CRT in order to standardize test instruction between raters and testing sessions. Participants completed the tests with their shoes on. No feedback regarding testing outcomes was given to participants or shared between physical therapists during or after testing.

Outcome Measures

Demographics and medical history. A questionnaire adapted from the Sports Concussion Assessment Tool–5th edition (Appendix 1) was used to collect information on participants’ demographics (i.e., sex, age, and the primary played sport) and medical history (i.e., history of previous concussions and current medications).30

The Balance Error Scoring System (Figure 1). The BESS is used to evaluate static balance ability and involves three stances, double, single and tandem. Each stance takes 20 seconds to complete and is performed on both a firm and unstable surface. A stance in the BESS is scored based on the number of errors a participant commits, with one point given for each error. Possible errors include lifting the hands off the iliac crests, opening the eyes, stepping, stumbling, falling, remaining out of position for more than five seconds, moving the hip into more than 30 degrees of flexion or abduction, or lifting the forefoot or heel. A maximum of 10 points per stance is allowed. If a participant is unable to maintain a stance for five seconds, a maximum score of 10 was given for that stance. The total score of the BESS ranges from 0 to 60, and is calculated as the sum of the error points given for each of the six stances.10For the dual-task condition, participants were asked to subtract by seven from a randomly assigned number while performing the BESS. This cognitive task is frequently used in dual-task balance assessment.16

Figure 1.

Figure 1.

CThe Balance Error Scoring System. Top row, firm surface condition. Bottom row, soft surface condition. Left column, parallel stance. Middle column, single-leg stance. Right column, tandem stance.

The Tandem Gait Test (Figure 2). The TGT is used to evaluate dynamic balance control and involves walking in a forward direction as fast and accurately as possible down and back along a 38mm-wide three-meter line using an alternate foot heel-to-toe gait. During the TGT, the administrator notes whether the evaluee steps off the line, separates his/her heel and toe, or touches the examiner or an object for support.30In the current study, the time in seconds required for participants to complete the test (i.e., TGT-Time), as well as the participant's ability to successfully complete the test (i.e., TGT-Error, pass/fail) were collected. For the dual-task condition, participants were asked to spell-out a five-letter word backward while performing the TGT.16

Figure 2.

Figure 2.

The Tandem Gait Test. (a) Starting point. (b) Heel-to-toe walking. (c) Turning. (d) heel-to-toe walking back to the starting point.

The Clinical Reaction Time (Figure 3). The CRT is used to evaluate reaction time and requires a participant to sit on a chair with the dominant hand resting opened on a flat, horizontal table. During the CRT, the examiner vertically suspends a rigid 80cm cylinder coated in high-friction tape, marked in ½ cm increments, and affixed to a weighted disk at one end. At predetermined, random time intervals ranging from 4 to 15 seconds, the examiner releases the apparatus and the participant catches it as quickly as possible. The distance the apparatus falls in centimeters is recorded by measuring from the top of the disk to the most superior aspect of the participant's hand. This distance is then converted to clinical reaction time, in milliseconds, using the formula for a free body falling under the influence of gravity (d?=?½ gt²; where d = distance, g = 9.8 m/s², and t = time.31For the dual-task condition, participants were asked to verbally spell a five-letter word backward while waiting for the testing apparatus to fall.16

Figure 3.

Figure 3.

The Clinical Reaction Time Test. (a) Demonstration of the starting athlete and tester positioning. (b) Demonstration of the post-drop athlete and tester positioning.

Analysis

Appropriate descriptive statistics were used to summarize all outcomes. For balance tests with continuous outcomes including the BESS, TGT-Time, and CRT, ICC2,1 with 95% confidence intervals (CI) were calculated based on trial one and the average of three trials to estimate relative intra-rater and inter-rater reliability.29ICC estimates were interpreted as acceptable if they were ≥0.60.28Standard error of measurement (SEM),24and minimal detectable changes at the 95% confidence level (MDC95) based on trial one and average of three trials were calculated to estimate absolute intra-rater and inter-rater reliability.32SEM was calculated as SEM = pooled Standard Deviation x (√1 - ICC). MDC95 was calculated as MDC95 = 1.96 x SEM x √2. Bland-Altman plots (i.e., mean difference and 95% limits of agreement) were used to assess for systematic bias between the first and third trial at session one, and between sessions using the average of three trials at session one minus the average of three trials at session two of all tests and conditions using data from rater one (TM).33For TGT-Error, Cohen's Kappa coefficients (κ; 95% CI) for three trials were calculated to estimate intra-rater and inter-rater agreement.34All analyses were performed using IBM SPSS 25 for Windows (Armonk, New York).

RESULTS

Participants

Of the 40 individuals who expressed interest in participating in the study, four did not meet the inclusion criteria (history of concussion within the year prior to testing), three declined to participate (time constraints), and nine did not respond to communications leaving a study sample of 24 participants. One participant withdrew after providing consent, but prior to testing for undisclosed personal reasons. The recruited sample included 23 participants. The median age of participants was 17 years (ranging from 13 to 24), and 39% (n=9) were female. The majority (65.2%) of participants played hockey, ringette (i.e., a game of Canadian origin for women and girls that is played on ice with two teams of six players on skates whose object is to drive a rubber or plastic ring into the opponents’ goal with a straight stick), or soccer. Nine of the participants (39%) had suffered a concussion greater than one year prior to testing. One participant (4.3%) reported current use of antibiotics for acne.

All participants (n=23) completed two sessions of testing. Only one participant attended the second session on the following day, while 22 participants (95%) attended the second session within four hours after testing session one. Summary statistics for participants’ performance on the BESS, TGT-Time, TGT-Error, and CRT are summarized by session, rater, and task in Table 1.

Table 1.

Descriptive Statistics for the BESS, TGT, and CRT by Session, Rater, and Task (n=23).

Condition Trial Session 1, Rater 1 Session 1, Rater 2 Session 2, Rater 1
BESS errors TGTT second TGTE % pass CRT ms BESS errors TGTT second TGTE % pass CRT ms BESS errors TGTT second TGTE % pass CRT ms
ST Trial 1 16 (24) 19 (25) 69 260 (90) 18 (26) 20 (24) 43 270 (150) 15 (30) 16 (15) 73 280 (140)
Average of 3 trials 14 (26) 19 (14) 73 250 (70) 16 (29) 19 (14) 49 260 (110) 15 (30) 16 (10) 76 250 (90)
DT Trial 1 14 (34) 20 (25) 78 240 (110) 14 (31) 19 (24) 82 270 (200) 15 (30) 19 (27) 78 280 (170)
Average of 3 trials 13 (30) 20 (23) 84 250 (110) 14 (27) 21 (23) 79 260 (140) 14 (30) 18 (23) 77 240 (100)

Note. BESS: Values are presented as median (range; defined as largest value minus smallest value) or percentage. BESS: Balance Error Scoring System, CRT: clinical reaction time, DT: dual-task, ms: milliseconds, ST: single-task, TGTE: pass/fail in the Tandem Gait Test, TGTT: time required to complete the Tandem Gait Test.

Relative Reliability

Table 2 summarizes intra- and inter-rater ICC's (95% CI) estimates for the BESS, TGT-Time, and CRT under single- and dual-task conditions calculated using trial one only, while Table 3 presents these estimates calculated using an average of all three trials. Inter-rater reliability ICC for the CRT (single-task) based on trial one could not be estimated due to an absence of variance in scores between raters (i.e., Negative ICC values obtained).35All of the ICCs calculated based on the average of three trials were > 0.60 for dual-task testing across tasks and conditions. Fifty-percent of the ICCs calculated based on the average of three trials were > 0.60 for single-task testing across tasks and conditions. Table 4 presents intra- and inter-rater Cohen's κ estimates for the TGT-Error under both single- and dual-task conditions. In general, the dual-task condition provided higher Cohen's κ estimates compared to single-task conditions.

Table 2.

Intra- and Inter-rater Reliability Estimates for the BESS, TGT, and CRT by Session, Rater, and Task (n=23).

Estimates Based on Trial 1
Intra-rater Inter-rater
ICC (2,1) (95% CI) SEM MDC ICC (2,1) (95% CI) SEM MDC
BESS-ST errors 0.72 (0.45,0.87) 2.4 6.6 0.89 (0.76,0.95) 1.8 4.9
BESS-DT errors 0.62 (0.31,0.82) 2.8 7.8 0.95 (0.89,0.98) 1.7 4.7
TGTT-ST second 0.20 (0,0.53) 3.8 10.5 0.98 (0.97,0.99) 0.7 1.9
TGTT-DT second 0.81 (0.61,0.91) 2.4 6.6 0.94 (0.87,0.97) 1.4 3.8
CRT-ST millisecond 0.10 (0,0.44) 28 77 - - -
CRT-DT millisecond 0.31 (0.01,0.62) 26 72 0.40 (0.00,0.70) 32 88

Note. (-): Values that could not be calculated due to low examiner variance, BESS: Balance Error Scoring System, CI: confidence interval, CRT: clinical reaction time, DT: dual-task, ICC: intra-class correlation coefficient, MDC: minimal detectable change, SEM: standard error of measurement, ST: single-task, TGTT: time required to complete the Tandem Gait Test.

Table 3.

: Intra- and Inter-rater Reliability Estimates for the BESS, TGT, and CRT by Session, Rater, and Task (n=23).

Estimates Based on the Average of Three Trials
Intra-rater Inter-rater
ICC (2,1) (95% CI) SEM MDC ICC (2,1) (95% CI) SEM MDC
BESS- ST errors 0.94 (0.87,0.97) 1.5 4.3 0.96 (0.88,0.98) 1.2 3.5
BESS- DT errors 0.94 (0.86,0.97) 1.4 3.8 0.98 (0.97,0.99) 1.1 3
TGTT- ST second 0.54 (0,0.79) 2 5.5 0.99 (0.99,0.99) 0.3 0.8
TGTT- DT second 0.94 (0.83,0.98) 1.2 3.3 0.99 (0.98,0.99) 0.4 1.1
CRT- ST millisecond 0.59 (0.04,0.82) 13 36 0.24 (0,0.68) 23 64
CRT- DT millisecond 0.90 (0.77,0.96) 11 30 0.70 (0.29,0.87) 17 47

Note. BESS: Balance Error Scoring System, CI: confidence interval, CRT: clinical reaction time, DT: dual-task, ICC: intra-class correlation coefficient, MDC: minimal detectable change, SEM: standard error of measurement, ST: single-task, TGTT: time required to complete the Tandem Gait Test.

Table 4.

Kappa Statistic (κ) for the Pass/Fail Task in the Tandem Gait Test (n=23).

Intra-rater Inter-rater
Trial 1 Trial 2 Trial 3 Trial 1 Trial 2 Trial 3
κ (95% CI) κ (95% CI) κ (95% CI) κ (95% CI) κ (95% CI) κ (95% CI)
Single-Task 0.03 (0,0.52) 0.23 (0,0.80) 0.64 * (0.1,1) 0.17 (0,0.52) 0.28 (0,0.68) 0.37 * (0.09,0.78)
Dual-Task 0.23 (0,0.80) 0.32 (0,0.91) 0.40 * (0.01,0.96) 0.58 * (0.01,1.00) 0.59 * (0.01,1.00) 0.83 * (0.13,1.00)

Note. *p<0.05, CI: confidence interval.

Absolute Reliability

Table 2 summarizes intra- and inter-rater SEM and MDC estimates for all tests under single- and dual-task conditions calculated using trial one only, while Table 3 presents these estimates calculated using an average of all three trials. Overall, administering the BESS, TGT-Time, and CRT three times and averaging the three trials provided lower SEMs and MDCs under both single- and dual-task testing. Based on the average of three trials, dual-task testing provided lower SEMs and MDCs for the BESS, TGT-Time, and CRT compared to single-task testing.

Table 5 summarizes the mean difference (SD) and 95% limits of agreement associated with Bland-Altman plots for between trials and between sessions of all tests and conditions. There was a positive shift in the difference scores related to single- and dual-task BESS, TGT-Time, and CRT between trials and between sessions (Figure 4 presents an example). Appendix 2 shows all of the Bland-Altman plots for between trials and between sessions of all tests and conditions. The positive shift observed remained after stratifying the analysis by sex and age.

Table 5.

Summary Data Associated with Bland-Altman Plots.

Test Mean Difference (SD) T1 to T3 95% LOA T1 to T3 Mean Difference (SD) S1 to S2 95% LOA S1 to S2
BESS-ST errors 2.5 (6.3) −9.8 to 14.8 0.2 (2.2) −4.2 to 4.5
BESS-DT error 2.1 (7.1) −11.7 to 15.9 0.1 (3.5) −6.8 to 6.9
TGTT-ST second 2.8 (3.9) −4.8 to 10.3 1.8 (3.2) −4.5 to 8.3
TGTT-DT second 1.3 (2.3) −3.3 to 5.9 1.2 (1.9) −2.7 to 5.1
CRT-ST milliseconds 3.0 (29.6) −54.9 to 61.0 4.3 (23) −40.9 to 49.6
CRT-DT milliseconds 6.5 (26.0) −44.5 to 57.5 0.8 (17) −33.0 to 34.7

Note. BESS: Balance Error Scoring System, CRT: clinical reaction time, DT: dual-task, LOA: limits of agreement, S: session, SD: standard deviation, ST: single-task, TGTT: time required to complete the Tandem Gait Test, T: trial.

Figure 4.

Figure 4.

Bland-Altman plot of the difference in the single-task Tandem Gait Test (seconds) between sessions one and two against the mean difference. The solid horizontal line represent the mean difference. The dashed horizontal lines represent the 95% upper and lower limits of agreement.

DISCUSSION

This novel research demonstrates that averaging observations across three trials produces clinically acceptable relative and absolute intra-rater and inter-rater reliability for the BESS, TGT-Time, and CRT in uninjured active youth and young adults. This was true in in both dual- and most single-task conditions. Given this and the potential learning effect observed, it is recommended that a practice session be administered prior to performing the BESS, TGT and CRT regardless of dual- or single-task conditions. Further, dual-task testing demonstrated higher relative and absolute reliability compared to single-task testing.

As there is a paucity of evidence about the reliability of dual-task clinical balance testing direct comparisons to previous studies are limited. Ross et al22reported higher estimates of inter-session reliability (ICC = 0.81, SEM = 1.87) based on one trial of the dual-task BESS in a sample of uninjured college students. On the other hand, a comparable estimate of inter-rater reliability (ICC = 0.87) based on repeated administrations of the dual-task CRT has been reported in a sample of uninjured athletes.14To the authors’ knowledge, there are no previous studies examining the reliability of TGT under dual-task condition.

There are substantially more studies examining the reliability of single-task clinical balance testing compared to dual-task testing.2,7,9,13,14Compared to the current study, Finnoff et al7reported a higher inter-rater reliability estimate (MDC=9.4) based on one trial of the single-task BESS. It is difficult to hypothesize the reasons for this difference as the authors did not report participant demographics, or a confidence interval for the MDC estimate. Similarly, Schneiders et al36reported higher intra-rater reliability estimates based on one trial (ICC=0.54) and the average of three trials (ICC=0.70) of the single-task TGT-Time in a sample of uninjured individuals (mean age=22.2 ± 3.8 years). Possible explanations for this difference may be the younger sample and the systematic error in the TGT-Time identified in the current study, which reduced the stability of the TGT-Time between testing sessions. One explanation for the systematic error in the TGT-Time for between-sessions measurements is a learning effect. Specifically, participants tended to walk faster in session two compared to session one (1.8 seconds in maximum mean difference; Table 5). Further studies comparing the systematic error of the single-task TGT-Time in adolescent versus adults are needed. Finally, Eckner et al14,37reported higher intra-rater (ICC=0.76) and inter-rater (ICC=0.74) reliability estimates based on repeated administration of the single-task CRT in samples of uninjured individuals (age range 8-30 years). The difference may be attributed to the limited number of repetitions averaged in the current study (three trials) compared to the later (eight repetitions), which may have contributed to greater between-rater variability. To the authors’ knowledge, there are no previous studies examining the reliability of single-task TGT-Error.

The results from the current study suggest that repeated administration of the dual-task BESS, TGT-Time, and CRT are associated with a learning effect. Specifically, participants tended to commit fewer errors on the BESS, walk faster on the TGT, and react faster on the CRT in trial three compared to trial one, and in session two compared to session one (Table 5). Further studies are needed to determine the effect of different sources of variance within these tests, potentially informed by generalizability theory. Examples of the sources of variance include age, sex, number of trials, fatigue, and footwear.38-41The results, on the other hand, further supports previous investigations suggesting that repeated administration produces a learning effect with the single-task BESS,38TGT-Time,13and CRT (mainly with three trials).42For instance, participants’ performance on the BESS, TGT-Time, and CRT enhanced in trial three compared to trial one, and in session two compared to session one (see Table 5).

In a previous systematic review on dual-task assessments for use in concussion management,16Register-Mihalik and colleagues observed that dual-task testing in some cases was more reliable than single-task testing. While the findings of the current study are consistent with this observation, the exact reason for this is not clear. The authors of the current study speculate this may represent higher measurement consistency under dual-task conditions, but could also be attributable to higher between-participant variability under dual-task compared to single-task condition (Table 1).

Clinical Recommendations

Adding a cognitive task to the BESS, TGT, and CRT enhanced the tests’ reliability without compromising ease of administration or time in the clinic. Based on data analyses, the dual-task BESS, TGT-Time, and CRT demonstrated acceptable reliability for clinical use. Clinicians may consider interpreting the mean score from three administrations of dual-task clinical balance testing to obtain the most reliable scores. The findings of the current study suggest that there is a 95% probability that a change of four points on the BESS, 3.3 seconds on the TGT-Time, and 30 milliseconds on the CRT is due to a change in postural stability rather than variability in trial scores by a single rater (based on the average score of three repetitions by the same rater). Further, there is a 95% probability that a change of four points on the BESS, 1.1 seconds on the TGT-Time, and 47 milliseconds on the CRT is due to a change in postural stability rather than variability in scores between raters (based on the average score of three repetitions by the same rater). For the dual-task TGT-Error, clinicians may consider interpreting the score from trial three to obtain the most reliable intra- and inter-rater scores. Finally, multiple test exposures may lead to improved balance due to a learning effect. Clinicians should implement a practice session before final testing to eliminate the learning effect seen.

Strengths and Limitations

The main strength of the current study is that both intra-rater and inter-rater reliability was evaluated using both relative (i.e., ICC) and absolute (i.e., SEM and MDC) reliability methodologies. That being said, this study has some limitations. Specifically, the order of tasks and testing conditions administration was not random, which added standardization but may have contributed to the observed learning effect. Thus, it is recommend that clinicians implement a practice session before final testing to reduce the potential learning effect observed. Further, the current study involved only uninjured participants, which limits the generalizability of the findings to injured youth and young adults. As the majority (61%) of the recruited sample were males, the generalizability of the findings to females may be limited. Similarly, the generalizability of the findings to different sports may be limited due to the fact that the majority (65.2%) of the recruited sample played hockey, ringette, or soccer. Individuals who were recruited may be at different levels of maturation, which may contribute to the variability in the results. Finally, the precision of the single-task TGT-Time, single- and dual-task TGT-Error, and dual-task CRT was low, which may be attributable to test instability and/or the small and homogenous sample recruited.

The current study is a first step toward establishing a line of research evaluating the clinimetric properties of clinical dual-task testing paradigms for identifying balance control deficits in neurologically impaired adolescents and young adults. Future studies examining the reliability, validity, and responsiveness of the dual-task BESS, TGT, and CRT are needed in larger and more representative samples of neurologically impaired adolescents and young adults.

CONCLUSION

Assessment of balance control is a common practice in clinical and training settings. Studies involving instrumented assessments have demonstrated that adding a concurrent cognitive task to a balance task can provide important information about balance control impairments that may not be identified with single-task testing. As the instrumented assessment of balance control is not always clinically feasible, it has been suggested that adding a cognitive task to the commonly used BESS, TGT, and CRT may improve the robustness of clinical balance assessment. The findings of the current study suggest that administering the dual-task BESS, TGT, and CRT three times and averaging the three trials provided acceptable reliability for clinical use. Further, dual-task testing demonstrated higher relative and absolute reliability compared to single-task testing.

Appendix I. Participants Background.

graphic file with name ijspt-15-487-A001.jpg

Appendix 2. Bland-Altman plots of the difference in balance tests between trials and between sessions against the mean difference. The solid horizontal lines represent the mean difference. The dashed horizontal lines represent the 95% upper and lower limits of agreement. a) Trial 1 vs. Trial 3 BESS, single task, b) Trial 1 vs. Trial 3 BESS, dual task, c) Trial 1 vs. Trial 3 Tandem Gait test, single task, d) Trial 1 vs. Trial 3 Tandem Gait test, dual task, e) Trial 1 vs. Trial 3 Clinical Reaction Time, single task, f) Trial 1 vs. Trial 3 Clinical Reaction Time, dual task, g) Session 1 vs. Session 2 BESS, single task, h) Session 1 vs. Session 2 BESS, dual task, i) Session 1 vs. Session Tandem Gait test, single task, j) Session 1 vs. Session Tandem Gait test, dual task, k) Session 1 vs. Session 2 Clinical Reaction Time, single task, l) Session 1 vs. Session 2 Clinical Reaction Time, dual task

graphic file with name ijspt-15-487-A002.jpg

References

  • 1.Huxham FE Goldie PA Patla AE. Theoretical considerations in balance assessment. Aust J Physiother. 2001;47(2):89-100. [DOI] [PubMed] [Google Scholar]
  • 2.Bell DR Guskiewicz KM Clark MA Padua DA. Systematic review of the balance error scoring system. Sports Health. 2011;3(3):287-295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Margolesky J Singer C. How tandem gait stumbled into the neurological exam: a review. Neurol Sci. 2018;39(1):23-29. [DOI] [PubMed] [Google Scholar]
  • 4.National Research Council, & Committee on sports-related concussions in youth. Sports-related concussions in youth: improving the science, changing the culture. National Academies Press. Washington. 2014:appendix C. [PubMed]
  • 5.McCrory P Meeuwisse W Dvorak J, et al. Consensus statement on concussion in sport-the 5(th) international conference on concussion in sport held in Berlin, October 2016. Br J Sports Med. 2017;51(11):838-847. [DOI] [PubMed] [Google Scholar]
  • 6.Murray N Salvatore A Powell D Reed-Jones R. Reliability and validity evidence of multiple balance assessments in athletes with a concussion. J Athl Train. 2014;49(4):540-549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Finnoff JT Peterson VJ Hollman JH Smith J. Intrarater and interrater reliability of the Balance Error Scoring System (BESS). Phys Med Rehabil. 2009;1(1):50-54. [DOI] [PubMed] [Google Scholar]
  • 8.Cushman D Hendrick J Teramoto M Fogg B, et al. Reliability of the balance error scoring system in a population with protracted recovery from mild traumatic brain injury. Brain Inj. 2018;32(5):569-574. [DOI] [PubMed] [Google Scholar]
  • 9.Hansen C Cushman D Chen W Bounsanga J, et al. Reliability testing of the balance error scoring system in children between the ages of 5 and 14. Clin J Sport Med. 2017;27(1):64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Riemann BL Guskiewicz KM Shields EW. Relationship between clinical and forceplate measures of postural stability. J Sport Rehabil. 1999;8(2):71-82. [Google Scholar]
  • 11.Alsalaheen BA Haines J Yorke A Stockdale K Broglio SP. Reliability and concurrent validity of instrumented balance error scoring system using a portable force plate system. Phys Sportsmed. 2015;43(3):221-226. [DOI] [PubMed] [Google Scholar]
  • 12.Quatman-Yates C Hugentobler J Ammon R Mwase N Kurowski B Myer GD. The utility of the balance error scoring system for mild brain injury assessments in children and adolescents. Phys Sportsmed. 2014;42(3):32-38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Koyama S Tanabe S Itoh N Saitoh E, et al. Intra-and inter-rater reliability and validity of the tandem gait test for the assessment of dynamic gait balance. Eur J Physiother. 2018;20(3):135-140. [Google Scholar]
  • 14.Eckner JT Richardson JK Kim H Joshi MS et al. Reliability and criterion validity of a novel clinical test of simple and complex reaction time in athletes. Percept Mot Skills. 2015;120(3):841-859. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Muro-de-la-Herran A Garcia-Zapirain B Mendez-Zorrilla A. Gait analysis methods: an overview of wearable and non-wearable systems, highlighting clinical applications. Sensors. 2014;14(2):3362-3394. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Register-Mihalik JK Littleton AC Guskiewicz KM. Are divided attention tasks useful in the assessment and management of sport-related concussion? Neuropsychol Rev. 2013;23(4):300-313. [DOI] [PubMed] [Google Scholar]
  • 17.Woollacott M Shumway-Cook A. Attention and the control of posture and gait: a review of an emerging area of research. Gait Posture. 2002;16(1):1-14. [DOI] [PubMed] [Google Scholar]
  • 18.Fraizer EV Mitra S. Methodological and interpretive issues in posture-cognition dual-tasking in upright stance. Gait Posture. 2008;27(2):271-279. [DOI] [PubMed] [Google Scholar]
  • 19.Yogev-Seligmann G Hausdorff JM Giladi N. The role of executive function and attention in gait. Mov Disord. 2008;23(3):329-342. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Al-Yahya E Dawes H Smith L Dennis A, et al. Cognitive motor interference while walking: a systematic review and meta-analysis. Neurosci Biobehav Rev. 2011;35(3):715-728. [DOI] [PubMed] [Google Scholar]
  • 21.Manaseer TS Gross DP Dennett L Schneider K, et al. Gait Deviations Associated With Concussion: A Systematic Review. Clin J Sport Med. 2017;10.1097/JSM.0000000000000537. [DOI] [PubMed] [Google Scholar]
  • 22.Ross LM Register-Mihalik JK Mihalik JP McCulloch KL, et al. Effects of a single-task versus a dual-task paradigm on cognition and balance in healthy subjects. J Sport Rehabil. 2011;20(3):296-310. [DOI] [PubMed] [Google Scholar]
  • 23.Howell DR Osternig LR Chou LS. Single-task and dual-task tandem gait test performance after concussion. J Sci Med Sport. 2017;20(7):622-626. [DOI] [PubMed] [Google Scholar]
  • 24.Portney LG Watkins MP. Foundations of clinical research: applications to practice. Vol 2: Prentice Hall; Upper Saddle River, NJ; 2000. [Google Scholar]
  • 25.Atkinson G Nevill AM. Statistical methods for assessing measurement error (reliability) in variables relevant to sports medicine. Sports Med. 1998;26(4):217-238. [DOI] [PubMed] [Google Scholar]
  • 26.Barber-Westin SD Noyes FR. McCloskey JW. Rigorous statistical reliability, validity, and responsiveness testing of the Cincinnati knee rating system in 350 subjects with uninjured, injured, or anterior cruciate ligament-reconstructed knees. Am J Sports Med. 1999;27(4): 402-416. [DOI] [PubMed] [Google Scholar]
  • 27.Walter SD Eliasziw M Donner A. Sample size and optimal designs for reliability studies. Stat Med. 1998;17(1):101-110. [DOI] [PubMed] [Google Scholar]
  • 28.Anastasi A. Psychological Testing. 6th ed. New York, NY: Macmillan; 1998. [Google Scholar]
  • 29.Koo TK Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155-163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Echemendia RJ Meeuwisse W McCrory P Davis GA, .et al. The Sport Concussion Assessment Tool 5th Edition (SCAT5): Background and rationale. Br J Sports Med. 2017;51(11):848-850. [DOI] [PubMed] [Google Scholar]
  • 31.Eckner JT Whitacre RD Kirsch NL Richardson JK. Evaluating a clinical measure of reaction time: an observational study. Percept Mot Skills. 2009;108(3):717-720. [DOI] [PubMed] [Google Scholar]
  • 32.Roebroeck ME Harlaar J Lankhorst GJ. The application of generalizability theory to reliability assessment: an illustration using isometric force measurements. Phys Ther. 1993;73(6):386-395. [DOI] [PubMed] [Google Scholar]
  • 33.Bland JM Altman D. Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet. 1986;327(8476):307-310. [PubMed] [Google Scholar]
  • 34.McHugh ML. Interrater reliability: the kappa statistic. Biochem Med. 2012;22(3):276-282. [PMC free article] [PubMed] [Google Scholar]
  • 35.Loosveldt G. Face-to-face interviews. International handbook of survey methodology. 2008:201-220. [Google Scholar]
  • 36.Schneiders AG Sullivan SJ Gray AR Hammond-Tooke GD, et al. Normative values for three clinical measures of motor performance used in the neurological assessment of sports concussion. J Sci Med Sport. 2010;13(2):196-201. [DOI] [PubMed] [Google Scholar]
  • 37.Eckner JT Kutcher JS Richardson JK. Between-seasons test-retest reliability of clinically measured reaction time in National Collegiate Athletic Association Division I athletes. J Athl Train. 2011;46(4):409-14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Valovich TC Perrin DH Gansneder BM. Repeat administration elicits a practice effect with the balance error scoring system but not with the standardized assessment of concussion in high school athletes. J Athl Train. 2003;38(1):51. [PMC free article] [PubMed] [Google Scholar]
  • 39.Wilkins JC Valovich McLeod TC Perrin DH Gansneder BM. Performance on the Balance Error Scoring System Decreases After Fatigue. J Athl Train. 2004;39(2):156-161. [PMC free article] [PubMed] [Google Scholar]
  • 40.Schneiders AG Sullivan SJ Kvarnstrom J Olsson M, et al. The effect of footwear and sports-surface on dynamic neurological screening for sport-related concussion. J Sci Med Sport. 2010;13(4):382-386. [DOI] [PubMed] [Google Scholar]
  • 41.Broglio SP Zhu W Sopiarz K Park Y. Generalizability theory analysis of balance error scoring system reliability in healthy young adults. J Athl Train. 2009;44(5):497-502. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Nguyen CN Clements RN Porter LA Clements NE, et al. Examining practice and learning effects with serial administration of the clinical reaction time test in healthy young athletes. J Sport Rehabil. 2018:1-22. [DOI] [PubMed] [Google Scholar]

Articles from International Journal of Sports Physical Therapy are provided here courtesy of North American Sports Medicine Institute

RESOURCES