Skip to main content
HHS Author Manuscripts logoLink to HHS Author Manuscripts
. Author manuscript; available in PMC: 2023 Nov 2.
Published in final edited form as: Assist Technol. 2021 May 3;34(6):674–683. doi: 10.1080/10400435.2021.1910375

Exploration of multimodal alternative access for individuals with severe motor impairments: Proof of concept

Kelsey Mandak 1, Janice Light 1, Savanna Brittlebank-Douglas 1
PMCID: PMC9136588  NIHMSID: NIHMS1698456  PMID: 33780326

Abstract

Many individuals with complex communication needs and severe motor impairments are unable to control technologies through conventional means and require alternative access techniques to achieve accurate and efficient access. With current alternative access techniques, individuals with severe motor impairments are limited in that they can only use one access technique at a time. The purpose of this project was to test proof of concept of a new multimodal access technique which integrated eye gaze and single switch scanning selection techniques. The aims were to investigate the learning patterns of two adults with severe cerebral palsy when first introduced to the multimodal access technique and then to compare the accuracy and efficiency of multimodal to single-modality access when selecting targets on an AAC visual scene display. The participants learned to use the multimodal access technique; they demonstrated improvements in their accuracy of selection across sessions and attained at least 80% accuracy within 3–15 training sessions. Both participants achieved similar accuracy with multimodal access compared to single-modality, but took longer to select targets with multimodal access compared to single-modality. The potential utility of multimodal access is explored as well as important avenues for future research.

Keywords: augmentative and alternative communication, older adults, computer access


Despite the significant advances in augmentative and alternative communication (AAC) access techniques, there are many individuals with complex communication needs (CCN) and severe motor impairments who find it challenging to access assistive technologies accurately and efficiently using conventional techniques (e.g., keyboard, mouse; Beukelman & Light, 2020). For individuals with severe speech and motor impairments, alternative access techniques are often required (Fager, Beukelman, Fried-Oken, Jakobs, & Baker, 2012). Alternative access techniques include a wide range of options such as alternative keyboards and mouse options, switch input devices, head-tracking, and eye-tracking technologies. These alternative access techniques include various control sites (e.g., eyes, knee, head, etc.) and a number of interfaces (e.g., onscreen keyboard, joystick, switches, etc.) which can be used to directly or indirectly select items. For example, an individual with severe motor impairments may use eye gaze to control AAC technology via direct selection. Another individual who is unable to use precise eye gaze may use single switch scanning to control assistive technology (i.e., indirect selection).

Although alternative access techniques benefit individuals with CCN and motor impairments, many challenges still remain. With current access techniques, individuals are limited in that they can only use one access modality at a time (e.g., only eye gaze or only switch-access). Although single-modality access may suffice for some, it can lead to extreme fatigue, over-use injuries, inaccurate control, and inefficiency for others (Fager, Bardach, Russell, & Higginbotham, 2012). For example, an individual who uses eye gaze may communicate effectively in the morning, but may fatigue throughout the day due to the repetitive ocular movements. Another individual who uses single switch scanning may become cognitively fatigued quickly as scanning requires extended time to generate a spontaneous message. Other individuals may be able to utilize an access technique to some extent, but not with sufficient precision to make it a viable means to control assistive technology. For example, some individuals with significant motor impairments lack the precise oculomotor control required to use eye tracking access; as a result, they must default to single switch scanning, a much slower and more cognitively demanding technique. Still others with severe motor impairments may demonstrate the capability to use a specific access technique when properly positioned in a wheelchair with customized supports, but may be unable to use the same access technique with the same level of precision in other positions throughout the day (e.g., in bed; Fager et al., 2019).

There is an urgent need to identify more effective and efficient ways for individuals with severe motor impairments to access assistive technologies. An alternative to single-modality access is multimodal access. Multimodal access combines two access modalities (e.g., eye gaze and switch access) in order to reduce motor demands placed on the individual (see Fager, 2018 for a review). This study investigated a new multimodal access technology developed through the Rehabilitation Engineering Research Center on Augmentative and Alterative Communication (RERC on AAC) which integrates eye gaze and single switch scanning techniques to operate AAC technology (Jakobs, Jakobs, Fager, Beukelman, & Light, 2014). Specifically, in this study, the multimodal access technique was investigated as individuals with severe motor impairments who were not literate made selections on visual scene displays (VSDs). VSDs are a type of AAC display that uses photographs of meaningful activities with programmed hotspots. Hotspots are pre-programed areas of a screen that, when selected, produce a pre-recorded spoken word or phrase to communicate a word or message. Research shows that VSDs can be used effectively to support the communication of various populations, including beginning communicators with developmental disabilities (e.g., autism spectrum disorder, cerebral palsy, Down syndrome) and adults with acquired conditions (e.g., aphasia, TBI, etc.) (Beukelman & Light, 2020). For these individuals, the use of VSDs offers several potential advantages including a) the inclusion of the actual social contexts in which language is learned (e.g., photographs of meaningful events) and b) the preservation of the visual and functional relationships that are present in real life. These attributes of VSDs contrast with those of grid AAC displays which present vocabulary concepts outside of the contexts in which they are learned in a decontextualized manner.

The multimodal access technique allows individuals to use eye gaze to highlight an approximate area of the VSD (e.g., upper right corner). Once the area is highlighted, the individual uses single switch scanning to scan through the hotspots closest to the highlighted area (e.g., any hotspots in the upper right corner). This multimodal access technique reduces the burden for precise oculomotor control, capitalizing on eye gaze to select the general location and thereby limit the array size of potential targets. Since the individual only needs to scan through the hotspots in the highlighted area, the time demands that are typical of scanning are reduced. However, multimodal access requires individuals to coordinate the use of two access techniques, potentially increasing the learning demands.

Preliminary research has explored the use of this type of multimodal access to an onscreen keyboard by a literate individual with an acquired disability resulting in severe motor impairments; this research has demonstrated promising results (Fager et al., 2019). Specifically, a single-subject alternating treatment study compared the performance of a participant when using eye-tracking only or multimodal access to spell randomly selected sentences using an onscreen keyboard (Fager et al., 2019). During 10 alternating sessions (5 eye-tracking, 5 multimodal), data were collected on accuracy of first attempts to target letters per sentence and number of errors per sentence. Across the 10 sessions, the participant demonstrated increased accuracy with the multimodal technique (91%) versus using eye-tracking only (63%). She additionally demonstrated less errors with the multimodal technique (average of 2 errors per sentence) compared to using eye-tracking only (average of 8 errors per sentence).

Some research has additionally explored the use of multimodal systems for purposes other than AAC alternative access and has demonstrated other benefits such as ease of use and reduced strain. Biswas and Langdon (2011) compared the use of eye-tracking alone to a multimodal eye-tracking and switch scanning system during a computer navigational task for individuals without disabilities. The participants rated the multimodal eye-tracking and scanning access method as less strenuous than eye-tracking alone. With the multimodal technique, users are able to switch back and forth between eye gaze and scanning, which allows the eye muscles to rest (Biswas & Langdon, 2011).

Evidence is emerging for other multimodal access combinations as well, including tongue motion, head tracking, speech recognition, and pen input. For example, Oviatt (2000) demonstrated a 19 – 35% reduction in the total error rate for a speech recognition system when a multimodal gesture-based (i.e., speech + pen input) system was used. Sahadat, Alreja, and Ghovanloo (2018) evaluated a wearable headset that utilized simultaneious multimodal input to provide computer access, including: 1) head tracking to control the mouse point, 2) tongue motion to provide the mouse clicks, and 3) speech recognition to provide the typing. Although the prototype technology was developed for individuals with severe motor impairments, the study evaluated the device with 15 individuals without disabilities while composing and sending an email. Over the course of four trials, participants improved in their task completion rates and accuracy, though slower and less accurate than when using a keyboard-and-mouse combination (i.e., the “gold standard”). Although these results demonstrate the exploration and potential of multimodal technology, most participants in the available studies were individuals without disabilities. Future research and technology development must include individuals with disabilities in order to understand the true potential, benefits, and challenges of multimodal access technologies.

The purpose of this study was to test proof of concept when using the prototype multimodal access technique for non-literate individuals with severe motor impairments when accessing AAC VSD technologies. Although Fager et al. (2019) have demonstrated the potential of multimodal technology when using an on-screen keyboard, multimodal access has not been investigated as a means to access AAC VSDs. Multimodal access to VSDs may differ from keyboards as there may be fewer targets and therefore less demands for precision with the former type of AAC display. With fewer targets in VSDs, it is also likely that the spacing between the representations (i.e., hotspots) is greater, as well as the size of the representations. For many individuals with CCN and severe motor impairments, the characteristics of VSDs may also be advantageous from an access standpoint (i.e., require less precision). Specifically, the aims of this project were to: (1) determine the learning patterns of individuals when first introduced to the new multimodal access technique, and (2) compare the accuracy and efficiency of multimodal access (i.e., integrated eye gaze and switch scanning) to single modality access (i.e., switch scanning alone and eye gaze alone) when selecting target hotspots on VSDs.

Methods

Participants

Individuals were eligible for inclusion in this study if they a) were between 21 and 60 years old; b) demonstrated a severe motor impairment that required alternative access to control assistive technology; c) demonstrated symbolic recognition, that is, able to recognize photographs or line drawings; d) were not functionally literate; e) were able to follow 2–3 step directions (e.g., “Find the fish tank. When you’ve found it, press the button.”); and f) had adequate vision (with or without correction) to locate all items on the AAC display and adequate hearing at conversational levels to respond to verbal instructions.

The two participants in this study, Mark (42-years old) and Dillon (52-years old; both pseudonyms), were both recruited from an adult training facility for adults with multiple disabilities. They were both diagnosed with severe spastic cerebral palsy and severe dysarthria. Both were non-ambulatory and used power-wheelchairs for mobility. The upper extremity control of both men was impaired bilaterally, resulting in the inability to effectively use their hands or fingers to directly select from an AAC display. See Table 1 for a description of each participant and his alternative access techniques.

Table 1.

Description of Participants and their Alternative Access Techniques

Participant Age/Diagnosis Typical access technique Investigated Access Techniques
Multimodal Eye gaze Scanning
Mark 42/Cerebral Palsy Eye gaze with switch selectiona Eye gaze with switch scanninga Eye gaze with switch selectiona Switch scanninga
Dillon 52/Cerebral Palsy Eye gaze with dwell selection Eye gaze with switch scanningb Eye gaze with dwell selection Switch scanningb
a

Mark used a jelly-bean switch located on the inside of his left knee. He activated the switch via knee adduction.

b

Dillon used a jelly-bean switch mounted on the underside of a large plexi-glass tray attached to his wheelchair. He activated the switch by slightly lifting his left leg.

At the time of the study, Mark was using eye gaze with switch selection to control his AAC device, a Tobii Dynavox i-15+ speech-generating device1 with Compass software2 and a 25-item grid. Mark used eye gaze to locate a target on his device and then used knee adduction to activate a jelly-bean switch that was located on the inside of his left knee to select the item. At the time of the study, Mark had been using his device with eye gaze access for three years. Prior to obtaining his current system, Mark relied on single switch scanning (via knee adduction) to control his previous AAC technology.

At the time of the study, Dillon was also using eye gaze (with dwell selection) to access his device, a Tobii Dynavox i-15+ speech-generating device1 with Snap Core First4. With his device, he was using page sets with 6 – 16 symbols per page. Dillon had been using this device for approximately 3 months prior to the study. Previously, Dillon mainly communicated through use of facial expressions, eye gaze to items or people, or shaking his head for yes/no responses. He intermittently used single-switch scanning and a free iPad app (i.e., SoundingBoard5). In this study, Dillon lifted his left leg to activate a jelly-bean switch, which was placed on the underside of a large plexiglass tray attached to his wheelchair.

Materials

The AAC technology with multimodal access used in this study was a prototype developed through the RERC on AAC by Jakobs and colleagues (2014). The system consisted of the following components: (a) a 12-inch Microsoft Surface tablet, (b) a Tobii-Dynavox infrared eye tracker (PCeye mini6) mounted to the bottom of the tablet, (c) a USB switch interface7, and (d) a jelly bean switch3.

The prototype technology was custom software which allowed the participants to switch between eye gaze and switch scanning within an application called EasyVSD8. Within EasyVSD, two VSDs were programmed for each participant; these VSDs were personalized photographs of each participant in his natural environment with hotspots for familiar people, objects, and/or activities. Each VSD was then programmed with 10 hotspots, consisting of 6 target items and 4 non-target items. To ensure a relatively balanced distribution of hotspots throughout the VSD, each VSD was visually divided into 6 separate sections (i.e., upper left, middle, and right; lower left, middle, and right) and the 6 targets were located in at least 5 of the 6 sections of the VSD. See Figure 1 for an example of a VSD with hotspots made visible for the reader to show the distribution.

Figure 1.

Figure 1.

Illustration of the Multimodal Access Technique when Completing the Target Selection Task in the VSD App

Training: Learning Multimodal Access

Before the comparison of the access modalities took place, the participants were trained to use multimodal access. In order to explore the learning demands of the novel access method, accuracy of target selection and time-to-target were measured for each participant across multiple sessions. Before data collection began, the researcher introduced a practice VSD and the multimodal access technique to each participant. First, the researcher introduced the task by stating, “Today we’re going to play a game, to see how fast you can find things on the screen. Sarah is going to show us how to play the game. She is going to listen to what I say, find it on the screen with her eyes, press her button when she thinks she is close, and then press it again when she finds it. Let’s try it out.” Following the introduction, the researcher provided one verbal instruction at a time (e.g., “Find the fish tank”) and a third-party model (a therapist from the adult program) completed the task in view of the participant, modeling the use of the multimodal access technique while also providing an verbal description of her actions (e.g., “I am looking at the fish tank”). The researcher and therapist completed these steps for three different hotspots. For each hotspot, the researcher modeled a different situation (i.e., selecting the target on the first scan through, selecting the target on the second scan through, and making an error and correcting it). During all models, the researcher or therapist provided a verbal description of her actions (e.g., “Sarah chose the plant instead of the fish tank. She has to make sure the circle is around the fish tank before pressing the button”).

Once the task had been modeled, the participant then practiced the task with prompting and feedback from the researcher (e.g., “Find the fish tank. Find it on the screen with your eyes.”; “When you’ve found the fish tank, press your switch.”; “Press your switch when the scan circle is around the fish tank”). Positive verbal feedback was provided after the participants completed each step. If the participant did not start the scan or did not try to select the target once the scan was started, the researcher provided the following prompts: “Look right at fish tank and press the button when you think you’re close”; “Remember, you want to hit your button again when the circle is around the fish tank.” Once the participants accurately selected the target two times consecutively with prompting support, they then independently practiced multimodal access to select two additional targets, this time without prompting from the researcher (i.e., the only prompt they received was, “Find the fish tank”). Two consecutive and independently made accurate target selections were required in order to begin the data collection procedures for the training.

Once the participants demonstrated that they understood the task requirements of multimodal access, they were introduced to two new personalized VSDs and began a series of training sessions. Training sessions took approximately 10 to 15 minutes and were scheduled 1–2 times a week. During each training session, the participants were instructed to select targets across various screen locations on their personalized VSDs. Each participant was prompted to find a total of 12 targets (i.e., 6 on each VSD) using the multimodal access technique (e.g., “Find the birdhouse”). During each of the trials, data were collected on the participant’s accuracy of selection and time to reach target. The training sessions continued for a minimum of 10 sessions until each participant demonstrated at least 80% accuracy.

The training sessions also provided an opportunity for the researchers to assess the functionality of the multimodal access prototype and make iterative changes to enhance performance. As originally designed, the multimodal access technique did not provide any visual feedback to the participants as they looked at the screen. After 9 training sessions, the technology was adapted to include a new feature on the screen (the eye-circle) that moved with a participant’s point of gaze and provided visual feedback for the participant. The circle coincided with the radius of the scan area (i.e., any hotspots within the circle were included in the scan). Both participants preferred to receive this visual feedback so this feature was utilized in all subsequent sessions.

Comparison of Access Modalities

Once the participants had completed a minimum of 10 training sessions and had attained at least 80% accuracy with the multimodal access technique during training, the comparison of access modalities began by using an alternating treatment single-case design (Kazdin, 2011). The use of this single-case design was chosen due to the heterogenous nature of the participants (i.e., significant differences in their abilities and characteristics) and the primary research aim of comparing the three access modalities (Kratochwill et al., 2010). In accordance with this design, the participants engaged in five individual sessions to compare performance across the multimodal and single-modality access techniques. During each session, the participants completed 12 trials in three alternating treatment conditions (i.e., eye gaze alone, scanning alone, and multimodal access; Kazdin, 2011). In the eye gaze alone condition, the participants were required to look directly at the desired target in order to make an accurate selection. In the scanning alone condition, the participants used automatic scanning in what can best be described as a circular pattern (i.e., the scan started with the hotspot that was most left on the screen and continued with the hotspot in closest proximity in a circular pattern). The sequence of the access techniques was counterbalanced in each session to control for order effects. In each session, the participants completed 12 target selection trials with each of the three access techniques, for a total of 36 selections in each session. The 12 target selection trials were completed using the two personalized VSDs that were introduced during the training (i.e., 6 targets per VSD). The two personalized VSDs were used for each participant during each individual session.

Measures

The dependent variables in the training and comparison sessions were percent accuracy of target selection out of 12 trials and time-to-target (in seconds). To be considered correct, the participants needed to select the target hotspot on the VSD in order to produce the spoken output from the tablet. A selection was considered incorrect if the participants chose the incorrect hotspot or if the participant initially made an error and then self-corrected. Time-to-target was measured as the time from the completion of the researcher’s verbal prompt (e.g., “Find the fish tank”) to the onset of the speech output of the target word spoken from the tablet upon selection (e.g. “Fish tank”). Time-to-target was only measured for accurate target selections.

All data from all sessions were collected using video and screen recording. Following each session, a graduate research assistant viewed each video and recorded the accuracy and time-to-target. To ensure reliability of data collection, the first author recorded the accuracy of selections and time-to-target for six of the training sessions for each participant. Her data were then compared with the research assistant’s in order to determine inter-rater agreement of data. Inter-rater agreement for accuracy and time-to-target was calculated by taking the number of agreements divided by the total number of items multiplied by 100. In order for time-to-target measures to be considered in agreement, the measured time of one rater had to be within 1 s of the other rater’s. Inter-rater agreement for accuracy of target selection was 100% for both participants. Agreement for time-to-target was 93% for Mark and 95% for Dillon.

Data Analysis

For the training sessions, data on the dependent variables (i.e., accuracy and time-to-target) were collected and summarized for consecutive sessions until the participants completed a minimum of 10 sessions and attained at least 80% accuracy; these data were graphed separately for each variable for each participant. Time-to-target was graphed as mean time in seconds across the correct trials. Changes in the level and trend of each dependent variable were analyzed visually for each participant across sessions to determine the demands of learning to use the multimodal access technique following initial introduction.

For the comparison sessions, in order to compare performance across the different access techniques (i.e., multimodal access compared to single-modality access, either eye gaze alone or single switch scanning alone), the alternating treatment data were graphed separately for each participant for each dependent variable. The graphs included performance data with each access technique for each of the five sessions. These data were then visually analyzed in accordance with single-subject standards (Kratochwill et al., 2010). Specifically, the level stability (i.e., variability) and change for each variable were compared for each participant across the three conditions (i.e., multimodal, eye gaze alone, scanning alone).

Results

Learning the Multimodal Access Technique

Mark completed a total of 10 training sessions with the multimodal access technique (the first 9 sessions without visual feedback and the final session with the eye-circle marking the location of his gaze on the screen). As shown in Figure 2, Mark learned the multimodal access technique quickly as evidenced by his accelerating trend; he attained at least 80% accuracy (i.e., criterion) by Session #3 and maintained this level of accuracy across all but one of the subsequent sessions (i.e., 67% in Session 7). His mean time-to-target showed improvement (i.e., decelerating trend) from 26.8 s in his initial session to means of 12.6 s, 19.9 s, and 14.6 s in sessions 7 through 9. With the design change in session 10 to incorporate the eye-circle, Mark continued to demonstrate accurate performance (i.e., 92% accuracy) of target selection and he achieved his lowest mean of 8.3 s.

Figure 2.

Figure 2.

Percent Accuracy and Mean Latency to Target Selection for each Participant during the Learning Sessions in Study #1 when Initially Introduced to Multimodal Access

Dillon completed a total of 17 training sessions, including nine without visual feedback as to his gaze location on the AAC display and the final eight sessions with the eye-circle feature. As shown in Figure 2, in the first nine sessions (i.e., without visual feedback), Dillon’s performance was highly variable with accuracy ranging from 17–75% (M = 54.6%) and mean time-to-target ranging from 18.9 to 45.6 s across sessions. Once the eye-circle feature was introduced to provide visual feedback on his gaze location, Dillon demonstrated an accelerating trend with his accuracy of target selection improving from 50% in session 10 to 83.3% in session 17 (range: 50–83.3%, M = 63.5%). His mean time-to-target improved overall from 43.7 s in session 10 to 29.4 s in session 17, although this measure continued to show variability across sessions (range: 19.1–51.8 s).

Comparison of Multimodal Access to Unimodality Access

Once training was complete, Mark and Dillon participated in five comparison sessions. The percent accuracy and mean time-to-target for each participant during the five comparison sessions can be found in Figure 3. Overall, Mark was marginally more accurate with multimodal access (M = 96.7%, SD = 4.4%), followed by eye gaze (M = 93.3%, SD = 7.1%), and scanning (M = 80.0%, SD = 13.9%). He demonstrated more variability in his accuracy with single switch scanning than the other two access techniques. For Mark, eye gaze resulted in the fastest time-to-target (M = 3.6 s, SD = 0.8 s), followed by scanning (M = 11.1 s, SD = 1.0 s), and multimodal (M = 11.1 s, SD = 3.6 s).

Figure 3.

Figure 3.

Comparison of Percent Accuracy and Mean Latency to Target Selection for each Participant across the Three Access Techniques during Study #2 (Alternating Treatment Design)

Dillon’s accuracy was similar for multimodal access (M = 78.3%, SD = 12.7%), and scanning alone (M = 76.7%, SD = 6.9%); he was less accurate with eye gaze alone (M = 60.1%, SD = 10.8%). Dillon’s time-to-target was fastest and least variable (M = 24.7 s, SD = 4.5 s) with scanning alone, followed by eye gaze alone (M = 34.5 s, SD = 7.8 s), and multimodal access (M = 39.9 s, SD = 13.9 s) which both showed greater variability.

Discussion

To-date, this is the first study to investigate multimodal access as a means for individuals with severe motor impairments who are nonliterate to control VSD AAC technology. The multimodal technique leveraged eye gaze to identify the general location of the target hotspot within the VSD and then used single switch scanning to precisely select the target hotspot within that general area. Despite the theoretical benefits of the novel access technique, including increased efficiency and reduced visual-motor demands, it can be challenging for individuals to learn to coordinate multiple modalities and complete a sequence of steps to select a target to communicate. It is therefore critical to gain a better understanding of the multimodal access method and how these learning demands may impact performance.

Learning to Use Multimodal Access

Results of the training sessions demonstrated that both individuals with severe cerebral palsy were able to learn to use the multimodal access technique. Mark required a relatively short period of time, requiring only three sessions to attain at least 80% accuracy (a total of approximately 40 minutes of practice). Dillon required a total of 15 sessions (a total of approximately 3 hours) to reach criterion, and was still quite variable in his performance during the final training sessions. Despite the varying learning patterns of the two participants, they were both able to learn to manage the visual-motor demands of the access tasks and the cognitive demands of sequencing these tasks correctly to select the targets. These results are in keeping with previous research into the learning demands of multimodal access; Fager et al. (2019) reported preliminary data showing that an adult with a brainstem stroke also learned to use multimodal access (combination of eye gaze and single switch scanning) to reliably access a keyboard in a short amount of time.

The visual feedback to eye gaze location on the screen, provided by the eye-circle, seemed to positively impact the time-to-target for Mark and the accuracy of performance for Dillon. Impairment of the oculomotor system is common in individuals with cerebral palsy and can result in abnormalities in fixations, smooth pursuits, and ocular movements (Fazzi et al., 2012; Park et al., 2016). The explicit feedback provided by the eye-circle may have supported both participants in monitoring their oculomotor control, thus improving performance with the multimodal access technique.

Mark demonstrated a steady increase in his accuracy and a steady decrease in his time to select the target across the 10 training sessions. Dillon’s accuracy showed a relatively steady increase over the final eight sessions once the eye-circle was introduced; however, his time-to-target was more variable, perhaps reflecting a need for more practice to build greater fluency in his performance. Treviranus and Roberts (2003) noted that typically instruction in new access techniques is terminated once the individual demonstrates an acceptable level of accuracy; they argue that instruction should continue over a sustained period of time so that the individual can build fluency (automaticity) in their performance. Although the participants both learned to use the access technique with greater than 80% accuracy, it is doubtful that they had sufficient practice with the new technique to develop automaticity and maximize efficiency.

Furthermore, the study utilized the seating, positioning, single switches, and scanning techniques used by the participants prior to the study. The participants, especially Dillon, may have benefitted from adjustments to their seating and positioning or from incorporation of a different type of cursor control technique for scanning (e.g., step scanning) rather than automatic scanning as used in this study (Beukelman & Light, 2020). Although Dillon demonstrated the fasted time-to-target with automatic scanning compared to the other access methods, it should be noted that he still required almost 25 seconds to reach the target. It is possible that another scanning method, such as step scanning, could reduce his time-to-target. Future research is required to investigate the effects of multimodal access incorporating different cursor control techniques with individuals with severe motor impairments.

Comparison of Multimodal versus Single-modality Access

Although the participants demonstrated a slightly higher mean accuracy level with the multimodal technique, there was high variability and overlap among the participants’ multimodal and single-modality accuracy values. Based on the available data, no clear advantages nor disadvantages of multimodal access can be concluded in terms of accuracy of target selection.

Regarding time-to-target, neither participant demonstrated the fastest time using the multimodal access technique. Mark was consistently fastest selecting target hotspots using eye gaze alone. This result is not surprising since Mark demonstrated sufficient oculomotor control to use eye gaze reliably (i.e., greater than 80% accuracy with eye gaze alone) and he had been using eye gaze to access his AAC technology over the past 3 years, providing sufficient time to develop automaticity of performance. Although multimodal access reduces the oculomotor demands of eye gaze alone, it adds additional steps to select target hotspots, increasing the time required to make a selection. Thus, the multimodal combination of eye gaze and scanning did not provide much benefit to Mark when completing the target selection task.

However, it is important to consider that Mark’s actual rate of communication when using the access modalities was never evaluated. The demands of the target selection task were very different from those of a true communicative interaction. For example, there was no need for participants to correct errors during the target selection task. On the contrary, when Mark is participating in actual interactions, corrections may be required to circumvent communication breakdowns and these corrections would take additional time. When using his actual device which had a more complex symbol-based grid display (e.g., increased number of representations), Mark would frequently make errors with his eye gaze selection and quickly correct his errors. Thus, in clinical practice, decisions regarding access must consider the interaction between accuracy and time-to-target to determine the trade-off. In tasks that prioritize accuracy over rate, or those that require higher precision of access, the option of multimodal access would potentially be beneficial even for someone like Mark.

Overall, Dillon selected the targets fastest using scanning, followed by eye gaze, and then multimodal access. However, his data suggest that he had difficulty using all three access techniques efficiently and often required longer than 20 s to select a target. His time-to-target with eye gaze alone and with scanning alone was more stable, but his time-to-target with multimodal access remained highly variable across trials and sessions suggesting that he would benefit from ongoing practice to achieve greater automaticity. As discussed earlier, individuals with cerebral palsy often have challenges with oculomotor control, which could potentially explain Dillon’s difficulty when using eye gaze access alone. He demonstrated mean accuracy of less than 70% using eye gaze alone (range 42 to 67% across sessions), suggesting that he would need to spend significant time correcting errors during communication interactions with this access technique. Dillon was also relatively new at using the eye gaze access method with his current communication device (i.e., just over 3 months). It is not uncommon for individuals to require months or over a year to learn the necessary operational skills and eye gaze control to access their communication system effectively (Donegan, 2012; Donegan & Oosthuizen, 2006).

Dillon demonstrated similar accuracy with multimodal access and scanning, but overall his time-to-target was faster with scanning. This result is surprising since, in theory, the multimodal access technique reduced the size of the scanning array and thereby should have been more efficient. The smaller scanning array did not appear to offer much benefit to Dillon since he was still faced with the timing demands of automatic scanning. Regardless of the smaller array, Dillon continued to have difficulties with accurately timing his switch activations. Also, multimodal access would have theoretically been more beneficial for Dillon if his oculomotor control was sufficient to consistently select the general location of the target on the VSD. The consistent improvements in his accuracy of performance in sessions 10 to 17 (once the eye-circle was provided) suggest that his oculomotor control may be sufficient to support multimodal access; however, he may require additional practice to attain more efficiency and automaticity. As stated above, Dillon was a relatively new eye gaze user and would likely benefit from more experience with the access method. At the time of the study, the coordination of his new access method (i.e., eye gaze) together with scanning may have been too demanding from a learning perspective.

To date, there has been only minimal research to investigate the effects of multimodal access. However, a comparison of the results of this study to the preliminary results reported by Fager and colleagues (2019) suggest that multimodal access may offer greater advantage to literate adults with acquired disabilities who are using a keyboard to spell messages than it did for the individuals with cerebral palsy using photo VSDs to communicate messages in this study. Fager and colleagues reported that an individual with a brainstem stroke was significantly more accurate on first attempts to select target letters from an onscreen keyboard when using multimodal access (91% accuracy) compared to eye gaze alone (63% accuracy) and she made significantly fewer errors overall when using multimodal access (2 errors per sentence) compared to eye gaze alone (8 errors per sentence). There may be numerous factors to explain the greater benefits in the Fager et al. (2019) study compared to the present study, including factors related to the participants and their intrinsic characteristics (e.g., motor, sensory perceptual, cognitive, linguistic, and literacy skills) as well as factors related to the AAC user interface display and the nature of the task. Spelling messages using an onscreen keyboard requires greater precision of access than communicating by selecting hotspots on a VSD. Most VSDs include fewer than 10 targets (i.e., a smaller selection set) and these targets are often relatively large (e.g., a person, a shared activity). With a smaller selection set, it is also likely that spacing between the targets and size of the targets are both greater than when using a keyboard display. Due to these physical characteristics, less precision is often required of individuals when accessing VSDs. Thus, multimodal access may prove to be more beneficial when completing tasks that require greater precision of selection, such as when using a keyboard display.

Although multimodal access might not be advantageous for some individuals compared to single-modality access under optimal conditions, it may serve as an important alternative or back-up access technique in certain situations (e.g., when positioning is not optimal for access, if there are calibration issues when using eye gaze, if there is significant demand for accurate communication; Fager et al., 2019). When faced with such barriers or challenges, a multimodal access technique may be appropriate. For example, Mark, the participant in the present study, used the multimodal access technique in a subsequent research investigation to complete a task in which accuracy was necessary.

Limitations and Future Research Directions

Although the results of this proof of concept study have furthered our understanding of multimodal access, there are several limitations and future research directions to consider. First, this study only included two individuals with severe cerebral palsy. In addition, this study only considered one potential combination of access techniques – eye gaze and scanning. Furthermore, the study only considered performance within structured trials and within those trials, only measured time-to-target for accurate target selections. Without measuring time-to-target for all trials, it is plausible that a faster time-to-target could be largely impacted by lower accuracy values for some participants. Future investigation is necessary with more individuals who have severe motor impairments to better understand when multimodal access may be effective and to identify additional combinations of access techniques that could be beneficial. Future research is also required to investigate the conditions under which multimodal access is most effective. Multimodal access may be more beneficial when completing tasks that require accuracy and using displays that require precision. This is the first study to consider the impact of multimodal access with individuals who are not literate, in this case using photo VSDs as AAC displays. Future investigations should explore the potential benefits with other types of AAC displays for individuals who are nonliterate, such as symbol-based grid displays.

Conclusion

“Despite the advances in standard and assistive technologies, individuals with the most severe motor impairments are experiencing a growing digital divide as their needs continue to be unmet” (Fager et al., 2019; p. 10). Now more than ever, the research and development of emerging access technologies is critical in order to meet the needs of these individuals. In this preliminary study, two adults who had severe motor impairments and were nonliterate, successfully learned how to use a multimodal access technique (combining eye gaze and single switch scanning) to control AAC VSD technology. Results suggested the potential benefits of multimodal access, hypothetically in communication situations in which accuracy and precision are vital. Additional investigations into multimodal access techniques are warranted in order to fully discover their potential and to ensure accurate and efficient access options for all individuals with complex communication needs who require access technologies.

Acknowledgements

Augmentative and Alternative Communication (The RERC on AAC), funded by Grant #90RE5017 from the National Institute on Disability, Independent Living, and Rehabilitation (NIDILRR) within the Administration for Community Living (ACL) of the U.S. Department of Health and Human Services (HHS). The contents of this paper do not necessarily represent the policy of NIDILRR, ACL, or HHS, and you should not assume endorsement by the Federal Government.

This project was supported by funding from the Rehabilitation Engineering Research Center on Augmentative and Alternative Communication (The RERC on AAC), funded by Grant #90RE5017 from the National Institute on Disability, Independent Living, and Rehabilitation (NIDILRR) within the Administration for Community Living (ACL) of the U.S. Department of Health and Human Services (HHS).

Footnotes

1.

The I-15+ is a speech-generating device by Tobii Dynavox (https://www.tobiidynavox.com)

2.

Compass is an AAC software by Tobii Dynavox (https://www.tobiidynavox.com)

3.

A jelly-bean switch is a mechanical, wired switch that is activated by pressure; developed by AbleNet Inc. (https://www.ablenetinc.com)

4.

Snap Core First is an AAC software by Tobii Dynavox (https://www.tobiidynavox.com)

5.

Sounding Board is a free communication app developed by AbleNet Inc. (https://www.ablenetinc.com)

6.

PCeye Mini is an eye-tracker by Tobii Dynavox (https://www.tobiidynavox.com)

7.

The USB switch interface used in this study was the Swifty by Origin Instruments (https://orin.com).

8.

EasyVSD is an app developed by InvoTek, Inc. (http://www.invotek.org), under the Rehabilitation Engineering Research Center on Augmentative and Alternative Communication.

References

  1. Beukelman D, & Light J (2020). Augmentative and alternative communication: Supporting children & adults with complex communication needs. Paul H. Brookes Publishing Co.: Baltimore, MD. [Google Scholar]
  2. Biwas P, & Langdon P (2011). A new input system for disabled users involving eye gaze tracker and scanning interface. Journal of Assistive Technologies, 5(2), 58–66. 10.1108/17549451111149269 [DOI] [Google Scholar]
  3. Biwas P, & Langdon P (2015). Multimodal intelligent eye-gaze tracking system. International Journal of Human-Computer Interaction, 31(4), 277–294. doi: 10.1080/10447318.2014.1001301 [DOI] [Google Scholar]
  4. Donegan M, & Oosthuizen L (2006). The “KEE”concept for eye-control and complex disabilities: Knowledge-based, end user-focused and evolutionary. Proceedings of the 2nd Conference on Communication by Gaze Interaction (COGAIN). Retrieved from http://wiki.cogain.org/index.php/Downloads [Google Scholar]
  5. Donegan M (2012). Participatory design: The story of Jayne and other complex cases. In Majaranta P, Aoki H, Donegan M, Hansen DW, Hansen JP, Hyrskykari A, & Räihä K (Eds.), Gaze interaction and applications of eye tracking: Advances in assistive technologies (pp. 55–61). Hershey, PA: Medical Information Science Reference. [Google Scholar]
  6. Fager S (2018). Alternative access for adults who rely on augmentative and alternative communication. Perspectives of the ASHA Special Interest Groups, 3(12), 6–12. 10.1044/persp3.SIG12.6 [DOI] [Google Scholar]
  7. Fager S, Bardach L, Russell S, & Higginbotham J (2012). Access to augmentative and alternative communication: New technologies and clinical decision-making. Journal of Pediatric Rehabilitation Medicine, 5(1), 53–61. 10.3233/PRM-2012-0196 [DOI] [PubMed] [Google Scholar]
  8. Fager S, Beukelman DR, Fried-Oken M, Jakobs T, & Baker J (2012). Access interface strategies. Assistive Technology, 24(1), 25–33. 10.1080/10400435.2011.648712 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Fager S, Fried-Oken M, Jakobs T, & Beukelman DR (2019). New and emerging access technologies for adults with complex communication needs and severe motor impairments: State of the science. Augmentative and Alternative Communication, 35(1), 13–25. 10.1080/07434618.2018.1556730 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Fazzi E, Signorini SG, La Piana R, Bertone C, Misefari W, Galli J, … & Bianchi PE (2012). Neuro-ophthalmological disorders in cerebral palsy: Ophthalmological, oculomotor, and visual aspects. Developmental Medicine & Child Neurology, 54(8), 730–736. 10.1111/j.1469-8749.2012.04324.x [DOI] [PubMed] [Google Scholar]
  11. Jakobs T, Jakobs E, Fager S, Beukelman D, & Light J (2014). Developing multimodal technologies to improve access. Retrieved from, https://rerc-aac.psu.edu/development/d1-developing-multimodal-technologies-to-improve-access/
  12. Kazdin AE (2011). Single-case research designs: Methods for clinical and applied settings (2nd ed.). New York, NY: Oxford University Press. [Google Scholar]
  13. Kratochwill TR, Hitchcock J, Horner RH, Levin JR, Odom SL, Rindskopf DM & Shadish WR (2010). Single-case designs technical documentation. Retrieved from What Works Clearinghouse website: http://ies.ed.gov/ncee/wwc/pdf/wwc_scd.pdf.
  14. Light J, & McNaughton D (2013). Putting people first: Rethinking the role of technology in augmentative and alternative communication. Augmentative and Alternative Communication, 29(4), 299–309. 10.3109/07434618.2013.848935 [DOI] [PubMed] [Google Scholar]
  15. Oviatt S (2000). Taming recognition errors with a multimodal interface. Communication of the ACM, 43(9), 45–51. 10.1145/348941.348979 [DOI] [Google Scholar]
  16. Park MJ, Yoo YJ, Chung CY, & Hwang JM (2016). Ocular findings in patients with spastic type cerebral palsy. BMC ophthalmology, 16(1), 195. 10.1186/s12886-016-0367-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Sahadat MN, Alreja A, & Ghovanloo M (2018). Simultaneous multimodal PC access for people with disabilities by integrating head tracking, speech recognition, and tongue motion. IEEE Transactions on Biomedical Circuits and Systems, 12(1), 192–201. 10.1109/TBCAS.2017.2771235 [DOI] [PubMed] [Google Scholar]
  18. Treviranus J, & Roberts V (2003). Supporting competent motor control of AAC systems. In Light J, Beukelman D, & Reichle J (Eds.), Communicative competence for individuals who use AAC (pp. 107–145). Baltimore: Paul H. Brookes. [Google Scholar]

RESOURCES