Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2021 Jul 13.
Published in final edited form as: J Second Lang Pronunciation. 2018 May 31;4(1):129–153. doi: 10.1075/jslp.00006.bli

Computer-Assisted Visual Articulation Feedback in L2 Pronunciation Instruction: A Review*

Heather Bliss 1,2, Jennifer Abel 2, Bryan Gick 2,3
PMCID: PMC8276941  NIHMSID: NIHMS1009630  PMID: 34262851

Abstract

Language learning is a multimodal endeavor; to improve their pronunciation in a new language, learners access not only auditory information about speech sounds and patterns, but also visual information about articulatory movements and processes. With the development of new technologies in computer-assisted pronunciation training (CAPT) come new possibilities for delivering feedback in both auditory and visual modalities. The present paper surveys the literature on computer-assisted visual articulation feedback, including direct feedback that provides visual models of articulation and indirect feedback that uses visualized acoustic information as a means to inform articulation instruction. Our focus is explicitly on segmental features rather than suprasegmental ones, with visual feedback conceived of as providing visualizations of articulatory configurations, movements, and processes. In addition to discussing types of visual articulation feedback, we also consider the criteria for effective delivery of feedback, and methods of evaluation.

Keywords: Multimodality, visual feedback, articulation, CAPT, segmental features

1. Introduction

It has long been observed that language learners can access not only auditory but also visual information to acquire the speech sounds and patterns in a new language (e.g., Catford & Pisoni, 1970; Navarra & Soto-Faraco, 2007). Accordingly, language teachers have incorporated both modalities into their instructional methods, and in particular there is a history of using visual stimuli to provide feedback to learners on their pronunciation. Much of the focus of this work has been on the development and implementation of visual feedback methods for teaching suprasegmental features (such as tone, stress, rhythm, and intonation), but visual information has also been used to provide feedback on segmental features (such as place or manner of articulation). While there are a number of comprehensive reviews of visual feedback on suprasegmental features (Chun, 1989, 1998, 2002, 2013; Hardison, 2004; Hincks, 2015; Levis & Pickering, 2004), to our knowledge there are no comparable reviews of visual feedback on segmental features. This paper is an attempt to fill that gap. Because segmental contrasts are achieved through manipulating the articulatory apparatus, visual feedback on segmental features can be conceived of as providing visualizations of articulatory configurations, movements, and processes. Expanding on the brief overview that Kartushina, Hervais-Adelman, Frauenfelder, & Golestani (2015, pp. 818–20) provide as background to their own study, the present paper’s focus is on computer-assisted visual articulation feedback, i.e., feedback that provides visual information about learners’ articulations of individual speech sounds, as compared to the articulations of native speakers, as a way to improve learners’ own articulations.

We begin with a survey of methods that have been used in studies that have reported on the use of visual articulation feedback explicitly, setting aside studies on feedback delivered via other modalities (e.g., auditory) or on other aspects of pronunciation (e.g., suprasegmental features). Our intent is to provide a narrative review, with a chronology of the development of visual articulation feedback methods; our criteria for inclusion are consequently less stringent than those required by a systematic review or meta-analysis.

Computer-assisted visual articulation feedback can be direct, displaying articulatory mechanisms and processes, or indirect, displaying acoustic information that informs articulation. We consider both. For visual articulation feedback to be successful, it needs to be easily interpreted by learners, and in order to determine success, studies need to control for other variables that can affect pronunciation and to report clearly on the methods and results. The present review comments on these factors as they pertain to the studies considered. Section 2 provides some context for the role of computer-assisted visual articulation feedback in the field of L2 pronunciation instruction. Section 3 gives an overview of the various types of computer-assisted visual articulation feedback, including indirect feedback, direct feedback and simulation approaches. Section 4 considers the criteria that must be met for computer-assisted visual articulation feedback to be successful in a pronunciation instruction program. Section 5 discusses methods of evaluation and reporting in the studies surveyed, and Section 6 presents conclusions and suggests areas for further study.

2. Situating visual articulation feedback in the context of L2 learning

Corrective feedback is the primary source of negative evidence for language learners; while the well-formed speech they are exposed to as input comprises positive evidence, corrective feedback is usually given in response to what are deemed incorrect productions, providing negative evidence to support the target grammar. Although it has been suggested by some researchers that L2 learning does not benefit from negative evidence provided by corrective feedback (Schwartz, 1993; Truscott, 2007), Lee et al. (2015) state that “hundreds of primary studies and 18 meta-analyses of feedback research … have shown positive effects for feedback” (p. 360); they go on to note that this body of research has, for the most part, exclusively focused on the role of feedback in correcting lexical and morphosyntactic errors, but that in their own meta-analysis of 86 studies of pronunciation instruction, feedback was also shown to play a positive role, suggesting that feedback is beneficial across the board in L2 learning.

While feedback is naturally an essential part of learning, developments in CAPT (Computer- Assisted Pronunciation Training) have enabled more technologically informed approaches to provide feedback in pronunciation instruction. If implemented successfully, CAPT systems have the potential advantage of reducing the burden of pronunciation teaching for instructors by providing opportunities for self-paced autonomous practice (Neri, Cucchiarini, Strik, & Boves 2002). Lee et al. (2015) note that feedback and technology often go hand-in-hand; CAPT systems in part developed as a way to address the problem of how to provide individualized feedback to learners. Interestingly, however, while Lee et al. observe that feedback in general has a positive effect on learning, they find the opposite for technology, noting that studies with technology-enhanced pronunciation instruction produce smaller effects than those relying exclusively on human-delivered instruction. They suggest that, while CAPT and other technological approaches have great potential, “there is clearly a need for research seeking to improve technology-enhanced instructional materials” (p. 361). Similarly, Neri et al. (2002) note that many CAPT systems fall short of their potential, and they develop a series of recommendations for developing CAPT systems that can better meet pedagogical requirements.

One way in which technology can be employed effectively in pronunciation instruction is by providing information not normally available to the learner, or by providing pronunciation information through different or multiple sensory modalities, offering learners the potential for a more multimodal or multisensory experience. CAPT systems can readily incorporate both auditory and visual cues, drawing on innovative methods for visualizing the speech signal. The idea of using computer-assisted visual feedback in language teaching has been investigated for over half a century (Abberton & Fourcin, 1975; Anderson, 1960; de Bot, 1980; Léon & Martin, 1972; Vardanian, 1964). As Lambacher (1999) observes, many of the efforts in this area have focused on teaching suprasegmental elements of speech. Nevertheless, there is now a growing body of literature addressing the use of computer-assisted visual feedback in improving L2 articulation.

3. Types of computer-assisted visual articulation feedback

This section surveys the various types of computer-assisted visual articulation feedback that have been described in the literature. Following Kartushina et al. (2015), we draw a distinction between indirect and direct types of visual articulation feedback. Feedback that uses visualized acoustic information, such as spectrographic, formant, or vocal tract resonance displays, falls into the category of indirect feedback, as it provides visual information about the articulation of L2 sounds derived from acoustic analyses. On the other hand, feedback that relies on visualization of articulation, such as that provided by ultrasound or intra-oral techniques, is categorized as direct feedback, as it “provides the participant with an immediate and dynamic view of the position and movements of their articulators during production” (Kartushina et al., 2015, p. 819). A third approach uses computer simulations of human articulation to provide information to learners. All three methods are described with references to illustrative examples in detail below which will be followed by in-depth critical assessment in Sections 4 and 5.

3.1. Indirect feedback using visualized acoustic information

Speech employs movements of the larynx, jaw, tongue, lips, etc. to produce audible variations in air pressure and to change the size and shape of the vocal tract (as a resonating chamber). These variations can be represented by visible displays of pressure waveforms and spectral information that trace the frequency and intensity of the acoustic signal over time. As waveforms and spectrograms indirectly encode articulatory movements via their acoustic correlates, they can be considered tools for indirectly visualizing articulation. As an example, consider Figure 1 below.

Figure 1.

Figure 1.

Waveforms (top signal) and spectrograms (bottom signal) for [a] and [i]

Figure 1 shows a sample waveform and spectrogram and illustrates the potential utility of spectrographic displays for providing indirect visual articulation feedback to L2 learners. The basic premise of this method is that learners could potentially improve their pronunciation by comparing spectrograms representing their own speech with speech produced by native speakers. In Figure 1, the spectrogram depicts formants, dark bands signifying vocal tract resonances that provide acoustic correlates of articulatory configuration. The first formant (F1) correlates inversely with tongue height; in Figure 1, F1 is higher for [a] than for [i], signalling that the tongue body is lower for [a] than for [i]. The second formant (F2) correlates inversely with tongue backness; in Figure 1, F2 is slightly lower for [a] than for [i], signalling that the tongue is positioned slightly further back in the mouth for [a] than for [i].

In an early contribution to the literature, Jenson and Westermeier (1968) attempted to use the characteristics of sound waves via an oscilloscope – similar to the waveforms exemplified in the upper part of Figure 1 – to provide visual feedback to L2 learners on their articulations. However, their efforts were aborted because the visualization was not salient for the learners, as it presented waveforms without the spectral information that conveys vocal tract shape. Jenson and Westermeier conclude that using spectrograms may be a fruitful avenue for providing visual articulation feedback to language learners, but that waveforms alone are not useful.

Molholt (1988, 1990) similarly recommends the use of spectrograms arguing that by using spectrographic displays on a split screen to compare learner and model speaker pronunciations, instructors can shift the focus from teaching technical vocabulary about places and manners of articulation to a visually richer feedback system that allows students to see speech patterns. Molholt goes on to say that “because students are immediately able to see patterns of their speech, they seem to be able to associate their kinesthetic feelings of production with the patterns on the screen. Eventually many seem to be able to associate these feelings with the target sounds. Then they no longer need the intermediary step of seeing the patterns” (1990, p. 83). Through this association of kinesthetic feelings and speech patterns, Molholt argues, spectrographic displays enable students to monitor their own progress in mastering the patterns. Similarly, Olson (2014a) note that spectrographic displays can allow students to conceptualize speech, and to self-analyse and self-monitor their productions in comparison with those of native speakers. However, despite reports that using spectrograms can lead to improved pronunciation of challenging segments (Olson, 2014a, b), it has been argued that spectrograms can be difficult for a learner to interpret, particularly without dedicated training in phonetics (Wilson, 2014).

In terms of the methodology for using spectrograms as a tool for visual feedback, some of the researchers point to the benefits of providing learners with detailed training in how to use the technology to generate and interpret spectrograms (e.g., Olson, 2014a, b; Patten & Edmonds, 2015; Quintana-Lara, 2014). Olson (2014a, b) notes that the objectives of teaching paradigms including this type of feedback need not simply be to improve pronunciation, but can also be to encourage students to approach the language they are learning analytically. Whether this type of analytical approach yields better outcomes in terms of achieving target-like pronunciation is yet unclear. For example, Olson (2014a) reports that students found the activities unique and useful but not necessarily beneficial to their own learning. Olson (2014b), however, reports a more controlled experiment in which two sections of a third-semester Spanish course received different types of pronunciation instruction, focusing on the fricatives [ð], [ɣ], and [β]. The experimental group engaged in activities in which they generated and analyzed spectrographic displays, while the control group received verbal feedback only on their productions. Both groups received the same amount of instruction and feedback, but only the experimental group showed marked improvements, suggesting that, although incorporating visual articulatory feedback using spectrograms requires training the students in spectral analysis, it can have a positive effect on their learning of new speech contrasts. Notably, however, it is not clear whether the improvement observed with the experimental group is due to the nature of the information, the analysis process, or the greater engagement and interaction of the experimental group with the material.

Beyond the presentation of raw spectrograms, more simplified visualizations of acoustic information have been used in visual articulatory feedback in L2 teaching, such as formant plot- style displays and visualizations based on vocal tract resonances. Unlike spectrograms, which show the raw properties of the speech signal, formant displays can use simplified and abstract representations by extracting resonance peaks from raw spectral data, which may be interpreted more quickly and easily by learners (Kartushina et al., 2015). An example is given in Figure 2.

Figure 2.

Figure 2.

Formant display of [a] and [i] vowels

Like Figure 1, this figure presents an illustrative example modelled after those reported in other studies, and as such has not itself been used to provide feedback in L2 learning contexts. Figure 2 uses the same acoustic information (taken from the same recording) as Figure 1, but presents it in a different format, with F1 and F2 plotted against each other to create a schematic of a vowel space, corresponding loosely to the location of the tongue in the oral cavity during vowel production. In Figure 2, the [i] vowel is higher and further left than the [a] vowel; this is meant to represent [i] being pronounced with a tongue position that is higher and further to the front of the mouth than [a].

An early approach using abstract formant displays was by Kalikow and Swets (1972), who attempted to capitalize on the correlation between formant values and tongue position to build a visual feedback system that generated formant plots corresponding to utterances recorded by Spanish learners of English, and compared them with those of native English speakers. They note that “if the salient features of any needed change are presented quickly and unambiguously to the student, he will be able to act appropriately. The system thus simulates the kind of individual attention that trained language teachers can give to limited numbers of students” (p. 24). However, computational speed proved to be a limiting factor in this study, as the system could not handle individual differences in vocal tracts.

Later studies using formant displays have shown mixed results; Carey (2004) reports no major effect of using visual feedback with formant displays for Korean learners of English, and notes that a limiting factor may be the fact that, although there are correlations, because learners have no formal experience with acoustic analysis, they therefore need to learn how to interpret the visual displays. On the other hand, Kartushina et al. (2015) report a more positive result: they found that French-speaking learners of Danish performed better in production tests following exposure to formant displays than did a control group with no such exposure. Kartushina et al. note that their methods of evaluation address limitations of previous studies by including a control group and using objective acoustic measures to assess improvements in learners’ pronunciation.

Dowd, Smith, and Wolfe (1997) advocate for using what they term vocal tract resonance imagery over traditional formant displays as a means of providing visual articulation feedback to language learners. They distinguish between formants and resonances, claiming that their method of using an acoustic impedance spectrometer (playing broadband noise into speakers’ open mouths and recording the output) allows resonances to be measured more accurately than using the glottal source, as the vocal tract shape information carried in normally voiced vowels is limited to fixed multiples of the fundamental frequency measured at the glottis. Dowd et al. (1997) tested the use of this method of using vocal tract resonance imagery via an imitation study with English-speaking learners of French, and found that visual feedback improved the learners’ pronunciations of the vowels [e, ɛ; a, ɑ; u, y] more than auditory feedback alone; they suggest that this finding is somewhat surprising, given that “the vocal tract feedback method is new and requires an eye- mouth coordination which is quite unfamiliar….The auditory feedback method, on the other hand, is a major component of the method used by most children to learn to speak their native language, and the subjects have had regular though informal exposure to auditory feedback ever since infancy” (p. 18). While this method does provide very accurate information about resonances, it has not been widely used because of the practical limitations that it requires a precise laboratory setup and was designed to provide information about vocal tract shapes only during sustained, open vowels with speakers holding the vowel position while not vocalizing, greatly reducing its effectiveness as a tool for feedback that can be used to inform L2 acquisition.

3.2. Direct feedback using visualized articulatory information

In contrast with the indirect feedback methods described in the preceding section, direct feedback by-passes acoustic information to give learners visual feedback on the position and movements of their articulators during speech. In this section we describe two modes of delivery for direct feedback: (i) ultrasound, and (ii) intra-oral techniques, such as EMA and EPG.

3.2.1. Ultrasound-based feedback

There is a long history of using ultrasound visualization technology in the field of articulatory phonetics as a research tool, speech therapy device, and language teaching aid (e.g., Kelsey, Minifie, & Hixon 1969; Stone, 2005). With ultrasound, high frequency sound is emitted through a transducer (the ‘probe’) that can be held against the neck so that the sound can travel through the tongue and be reflected back to the transducer, creating a 2-dimensional image of the tongue. Because it does not image through bone or air, ultrasound does not allow visualization of the palate, jaw, or rear pharyngeal wall; thus, it is typically used to image only the tongue in speech applications. Sample ultrasound images are presented in Figure 3.

Figure 3.

Figure 3.

Ultrasound images of [a] and [i]

The C-shaped curves at the bottom of the images in Figure 3 correspond to the probe (held against the neck), and the white line in each image corresponds to the tongue, with the tip of the tongue appearing towards the left and the root appearing towards the right. The two vowels, [a] and [i], have distinct tongue shapes, with lowering of the back of the tongue for [a] and raising of the front of the tongue for [i].

Inspired by work using ultrasound in clinical interventions with deaf and hard-of-hearing learners of spoken language (e.g., Bernhardt, Gick, Bacsfalvi, & Adler-Bock 2005; Bernhardt, Gick, Bacsfalvi, & Ashdown 2003), ultrasound-based visual feedback has been increasingly used in teaching and learning of difficult second language sounds. While no controlled studies have been conducted comparing ultrasound-based feedback with other kinds of computer-assisted visual articulation feedback, several potential advantages of ultrasound-based feedback have been identified in terms of the system itself, including safety, non-invasiveness, and versatility (see Gick, Bernhardt, Bacsfalvi, & Wilson, 2008; Wilson, 2014). The physical set-up for an ultrasound feedback session (shown in Figure 4) is minimal; portable ultrasound machines are comparable in size to a laptop computer, and the only discomfort for the learners is the application of conductive gel on the chin, which enhances the ultrasound image but may feel sticky on the skin.

Figure 4.

Figure 4.

Portable ultrasound set-up, showing probe position (L) and ultrasound display (R)

The increasing affordability and portability of ultrasound systems, and the relative speed and ease with which these systems can be set up and used, have also made ultrasound use outside of a lab or clinical setting more feasible. In the specific case of pronunciation teaching, Wilson and Gick (2006, p. 148) note that the kinds of acoustically-focused methods described in the previous section “risk a lack of understanding on the part of the learner as to how to map the acoustic information onto articulatory movements. If learners are able to see directly the articulators, then they probably have an improved perception of the articulatory adjustments needed to improve their pronunciation.” Moreover, Wilson (2014, p. 287) suggests that systems that assess a learner’s articulation based on an acoustic signal “have the danger of drawing false inferences and incorrectly judging the learner’s pronunciation as right when it is wrong, or as wrong when it is right.” Systems like ultrasound that measure a learner’s articulation directly, on the other hand, are in principle less prone to these kinds of errors. By visualizing the articulators, ultrasound “has the potential to contribute to the teaching of pronunciation through both a top-down method (i.e., by shedding more light on underlying articulatory setting) and a bottom-up method (i.e., by enabling learners to view real-time images of their tongues as they produce individual sounds)” (Gick et al., 2008, p. 313). In terms of interpretation of the visual feedback, Wilson (2014) suggests that a visual representation of their own tongue movement in real time, or even of a teacher’s or model’s correct tongue movement, is much easier for a phonetically naïve L2 learner to interpret than a spectrogram.

Ultrasound-based visual feedback is not ideal for all articulations and all learning settings. Certain sounds are particularly amenable to ultrasound feedback: for example, vowel articulations, or sounds for which the timing of movements of different parts of the tongue is crucial to successful pronunciation, such as laterals and rhotics (Wilson & Gick, 2006). As noted, ultrasound is not able to visualize bone, which prevents identification of hard structures such as the teeth or the hard palate in the resulting images. As well, ultrasound technology is only just beginning to be explored in the context of in large-group settings like classrooms or language laboratories. This is due in part to the infeasibility of having multiple ultrasounds available in order to have a number of students working on them at the same time. However, this may change with the development of handheld ultrasound devices (e.g., Bruce et al., 2000; Clarius, 2016; Wojtczak & Bonadonna, 2013). Given existing technology, Gick et al. (2008) suggest that ultrasound intervention is particularly amenable to a ‘single participant design,’ in which one learner’s particular needs, variations and challenges can be addressed by an instructor.

Accordingly, most studies incorporating ultrasound-based visual feedback have targeted a small number of learners, typically 10 or fewer (e.g., Gick et al., 2008; Pillot-Loiseau, Kamiyama, & Kocjančič Antolík, 2015; Tateishi & Winters, 2013; Tsui, 2012; Wu, Gendrot, Hallé, & Adda- Decker, 2015; for studies with larger participant groups see Cleland, Scobbie, Nakai, & Wrench, 2015; Ouni, 2014). The small number of participants in most of these studies limits the generalizations that can be drawn. However, for each participant, the training with and exposure to ultrasound imagery was time-intensive; sessions ranged between 30 and 60 minutes, and most studies incorporated multiple sessions. With multiple training sessions, it is perhaps not surprising that learners improved; Tsui (2012) reports that, following four 45-minute sessions, Japanese- speaking learners of English were rated by experts as having more accurate productions of /l/ and /ɹ/, and Tateishi and Winters (2013) report similar results following five 30-minute sessions. Pillot- Loiseau et al. (2015) also employed a multiple-session design in their study of Japanese learners of the French [u]/[y] contrast. They report that, after three 45-minute sessions in which the learners repeated items after a French native speaker with ultrasound feedback, and then moved on to practicing the vowels in isolation, in words, and in sentences, they showed significant improvements in their productions of [u] and [y] compared to a control group that only received conventional pronunciation training.

Even with only one training session using ultrasound, learners show improvements. Gick et al. (2008) performed a small-scale intervention with three Japanese learners of English to improve their /ɹ/ and /l/ pronunciation. Each participant received a single hour-long session that included a pre-assessment of their productions of the target sounds, training with the ultrasound (about 30 minutes in length), and a post-assessment. During the pre-assessment, the participants’ productions were recorded using the ultrasound; their best and worst productions of /ɹ/ and /l/ from that assessment were then shown to them as part of the training block. The post-assessment indicated that all the participants were able to produce the problem segments more successfully than they had before the ultrasound training.

Notably, it is unclear whether the positive results reported in these studies are in fact due to the learners’ ability to interpret and respond to the visual articulation feedback provided by ultrasound, or whether it was simply the “wow factor” of the technology, i.e., that, being a novel tool in language learning contexts, it is attention-orienting. Therefore, further research is needed to determine the broader efficacy of ultrasound-based feedback.

3.2.2. Feedback using intra-oral techniques

In addition to ultrasound, intra-oral techniques using electropalatography (EPG) or electromagnetic articulography (EMA) are methods for providing direct articulatory feedback to language learners under Kartushina et al.’s (2015) classification. EPG tracks points of contact between the tongue and the palate by way of an artificial palate (custom-made for a speaker) equipped with electrodes that send signals about the location and timing of articulatory movements to an external recorder. EMA, on the other hand, tracks articulatory movements by way of sensor coils on a speaker’s tongue and in other parts of their mouth and induction coils around the speaker’s head. The induction coils produce an electromagnetic field that creates a current in the mouth sensors that can be tracked and recorded by an external device.

While EPG has been used to give feedback in speech therapy contexts (e.g., Bernhardt et al., 2003), it has not to our knowledge been tested as a means of providing visual articulatory feedback in L2 learning contexts. Ballard et al. (2012) employ EPG in training English speakers to produce Russian trilled [r], but do not report on the use of feedback. On the other hand, EMA has been employed for this purpose several times in recent years. Levitt and Katz (2007) tested monolingual adult speakers of American English on the acquisition of the Japanese post-alveolar flap using EMA-augmented visual feedback as a means to assess whether this technique is feasible for studies involving individuals with speech disorders. All participants had eight 20-minute training sessions in the production of the flap sound; an experimental group received visual feedback using EMA, and a control group received verbal feedback. The experimental group showed a significantly higher degree of improvement in their productions following the training compared with the control group. Tilsen, Das, and McKee (2015) similarly report improvements in productions following EMA feedback, but only for one participant (who was also a trained phonetician). Katz and Mehta (2015) also report improvements in productions; in this case the feedback was delivered using the Opti-Speech system (Katz et al., 2014), a 3D animated tongue driven by EMA data. Likewise, Suemitsu, Dang, Ito, and Tiede (2015) report improvements in productions following EMA feedback. While Katz and Mehta’s study only included five participants and no control group, Suemitsu et al.’s study included twenty-one participants divided into three conditions (audio, visual, and audiovisual training); those who received visual or audiovisual training using EMA showed greater improvements than those who received audio feedback alone.

While these intra-oral techniques may be useful for experimental studies of pronunciation learning, there are serious impediments to their being used in larger-scale pronunciation studies or practical language-learning applications. Unlike ultrasound, which has a wide range of applications and is increasingly affordable and portable, neither EPG nor EMA systems are widely available as they have little market outside of speech research. They tend to be costly and time-consuming to use, either because custom palates have to be designed and manufactured for each participant (EPG) or because the system itself is expensive and generally requires the assistance of a medical professional to attach the transceiver pellets (EMA). Moreover, these techniques are invasive, with objects being placed in the mouth that interfere somewhat with normal pronunciation (for descriptions and pictures of the apparatus and experimental set-up for EPG and EMA, see Ballard et al., 2012 and Katz et al., 2014, respectively).

3.3. Simulation Approaches

In addition to the direct and indirect feedback techniques described above, there is also the possibility of using virtual or augmented reality to deliver visual articulatory feedback. While audiovisual speech simulations are increasingly common (Mattheyses & Verhelst, 2015), their application in pronunciation instruction has not been widely explored (see Badin, Ben Youssef, Bailly, Elisei, & Hueber, 2010 for a short review; see also Demenko, Wagner, & Cylwick 2010 for a simulation program incorporating suprasegmental features of pronunciation). We classify simulation approaches as neither direct nor indirect visual articulation feedback, because while they provide visual correction based on a learner’s productions, they do not directly visualize the output of those productions.

One study that uses audiovisual speech simulation as a means for visual articulatory feedback in L2 learning is conducted by Massaro and Light (2003). They tested the effectiveness of using the computer-animated talking head “Baldi” for teaching English /r/ and /l/ to native Japanese speakers. Because Baldi’s “skin” could be rendered transparent, it was hypothesized that highlighting the relevant articulators would help the learners to achieve target-like productions.

11 participants received pronunciation training over a six-day period, with half the group receiving articulatory training with transparent skin and visible articulators, and the other half receiving the same training without the transparent skin condition. These conditions switched after three days of training; the group who saw the visible articulators for the first three days did not for the final three, and vice versa. Regarding the use of feedback, Baldi not only modelled the correct pronunciations, but by means of a speech recognition module, was also able to provide corrective feedback to the learners on their pronunciations. Groups were tested on their perception and production of [r] and [l] after the first three days of each training type, and the group that had received the articulatory training with Baldi showed a significant improvement over the group that did not. However, there was no effect for the transparent skin condition.

Engwall (2012) describes another study that employs simulation technology, using the ARTUR virtual articulation teacher, a simulation whose tongue shape and vocal tract is modelled based on statistical analyses of MRI and EMA measurements. In this study, ARTUR was used in teaching seven French speakers to produce the Swedish alveolar trill [r] and the velar fricative [ɧ]. The virtual teacher showed either a front view of the face or an augmented reality (AR) view showing a mid-sagittal cutaway of articulations showing tongue movements. (A video depicting ARTUR is available at https://www.youtube.com/watch?v=09cZRyOS-OU, courtesy of Engwall.) Based on the learners’ pronunciation attempts, a human operator selected the most appropriate feedback to display (i.e., words and AR animations appropriate to the sounds in question). As such, this study does not give feedback in the same way as the direct visualization techniques discussed in this paper; the learners are getting feedback that is informed by their own productions, rather than feedback that actually shows their own productions (see also Katz et al., 2014). Regardless, all seven subjects revised their pronunciations of [r] and [ɧ] in the right direction based on the feedback given, though not always reaching target-like pronunciation. Engwall (2012) concludes that, even when the pronunciation was not correct, “their articulatory changes are still relevant, since this study concerns how learners respond to articulatory instructions. Therefore, it is a positive indication that the subjects are aware of how to change the articulation not only when they are able to achieve the correct pronunciation, but also when they follow the articulatory instructions, even if the acoustic result is incorrect in the short term” (2012, p. 59).

4. Criteria for the Effective Use of Visual Articulation Feedback

In the preceding section, we saw that generally feedback is either indirect or direct, depending on whether it derives articulatory information from visualizations of the acoustic signal (indirect) or provides immediate and dynamic visualizations of articulatory mechanisms (direct). Both positive and negative results are reported for both types of studies. Why is visual articulation feedback successful in some cases and not others? In this section, we consider some criteria for the effective use of visual articulation feedback in L2 pronunciation training, and comment on how the studies considered in this review meet these criteria. Inspired by discussions by Öster (1997) and Carey (2004) on the ingredients of an effective visual feedback tool, we include the following traits as criteria: (i) natural and logical, (ii) understandable, (iii) immediate, (iv) able to facilitate comparison, (v) flexible, (vi) enriching, and (vii) affordable.

Regarding the first criterion, the visual pattern depicted to the learners must be natural and logical. In principle, this criterion could be met using either direct or indirect feedback methods; Öster (1997) suggests that using visual cues that manipulate the vertical or horizontal space of the display, or the size or colour of an object on the display, may be suitable as abstract representations of changes in the acoustic signal. Notably, this level of abstractness is precisely what leads Wilson (2014) to suggest that direct feedback, such as ultrasound imagery, is favourable, as it avoids the potential problem of learners misinterpreting the visual representation and drawing false inferences. However, ultrasound does not image through bone, and consequently learners do not see all of the articulators, some of which could act as landmarks for improving articulation. Carey (2004) expands on this point, noting that feedback should be explicit and lucidly expressed. Waveform displays such as those used by Jenson and Westermeier (1968) do not provide useful visual feedback because the “information they display is only the intensity of the speech signal” (Carey, 2004, p. 573). Kartushina et al. (2015) issue a similar critique for studies using indirect feedback such as Aliaga-García and Mora (2009) and Dowd et al. (1997), stating that “the type of feedback that has been used is not always easy for L2 speakers to use rapidly and to interpret during training” (p. 820).

The next criterion is understandability. Öster (1997) notes that feedback needs to be comprehensible and easy to handle, and Carey (2004) similarly notes that it should be user-friendly and easy to negotiate. Studies that make use of spectrographic displays and/or technological platforms such as Praat involve an extra investment in terms of learning to interpret and interact with the feedback to improve pronunciation. Olson (2014a,b) notes that the primary learning objective for his students is the ability to analyse their productions, rather than improve them. However, if the goal is improvement in pronunciation, are the current visual articulation feedback mechanisms learnable? Arguably, visual representations of the articulators and their movements is easier to interpret than indirect feedback, e.g., from spectrograms (see Wilson, 2014). Of the methods that have been used, ultrasound is relatively straightforward, requiring minimal training to operate and interpret (Noguchi et al., 2015). Methods such as EPG and EMA, on the other hand, require significantly more training. It remains an open question whether ease of interpretation directly corresponds to ease of learnability.

Regardless of what kind of feedback is delivered and in what form, it should be immediate (Öster, 1997). As Carey (2004) notes, real-time or immediate feedback is “crucial because if time elapses between the learner’s utterance and the appearance of feedback to it, the opportunity to make fine adjustments to the articulators and to reflect upon learning is lost” (p. 580). Lambacher’s (1999) description of a software system implemented in a language lab is an example of how to deliver immediate feedback in a visualized acoustic feedback system. Lambacher notes that “the transference of data is in real time, which enables learners to get immediate feedback about their errors and progress from the teacher” (1999, p. 141). Direct feedback methods, including those using ultrasound, EPG and EMA, also have the advantage of providing immediate feedback generated in real-time.

Another criterion is the ability to facilitate comparison; the visual depiction of the learner’s articulation and that of a native speaker are best shown simultaneously (Öster, 1997). In some studies this is facilitated by either a split screen display (e.g., Lambacher, 1999; Molholt 1988, 1990) or a visual representation that incorporates both articulations (e.g., Akahane-Yamada et al., 1998; Kartushina et al., 2015). With technologies that do not permit a split screen display, multiple monitors may be used to facilitate comparison, or a sequential comparison method may be used. For instance, in ultrasound-based studies learners can be asked to view images of both their own articulations and those of native speakers to make explicit comparisons (e.g., Gick et al., 2008; Tateishi & Winters, 2013).

Flexibility is also important for visual articulation feedback. Carey (2004) approaches flexibility from the design perspective, noting that platforms that facilitate feedback should be authorable, allowing for linguistic variation across populations and regions. Öster (1997) stresses flexibility from the delivery perspective; the mode of delivery to the students should be flexible enough to permit individualization (see also Dowd et al., 1997). Direct feedback systems satisfy this criterion by default. But indirect systems can satisfy this criterion as well. Patten and Edmonds (2015) note that feedback using spectrographic displays is more nuanced than more traditional types of pronunciation instruction, allowing for variability in learners’ articulations. They use the example of /ɹ/, which can be articulated differently by different speakers, stating “spectrographic visual feedback training for /ɹ/ circumvents the primary challenges with the two main types of traditional training: articulator placement instruction and repetition. Participants can experiment freely with their own articulatory hypotheses and receive real-time feedback about the relative success of each attempt, and immediately revise for the next trial” (2015, p. 256).

With respect to the segments that visual articulation feedback methods can target, most are flexible. While indirect feedback methods using formant plot displays may best lend themselves to instruction on vowels (see Carey, 2004; Kalikow & Swets, 1972; Kartushina et al., 2015), spectrographic displays have been used not only for vowels, but also fricatives (Olson, 2014a,b), liquids (Akahane-Yamada et al. 1998; Lambacher 1999; Patten and Edmonds 2015), and stops (Aliaga-García & Mora, 2009). Ultrasound-based methods have tested learners on /ɹ/ and /l/ sounds, as well as vowels (Cleland et al. 2015; Pillot-Loiseau et al., 2015), but it cannot image articulations produced at the far front of the mouth (e.g., dental or labial sounds), voicing, or other laryngeal contrasts (though see Moisik et al., 2011). EPG can image detailed tongue-palate contact patterns but is unable to measure articulations elsewhere (lips, larynx, pharynx) and does not provide a more general view of tongue shape. EMA can provide measures of the lips and anterior tongue but cannot image the posterior tongue or larynx. Compared with any other single method, simulation techniques have the potential to be more flexible in terms of the range of sounds they can target.

The technology for providing visual articulation feedback should also be enriching, providing something additional that cannot be provided by a (human) instructor. Carey (2004, p. 575) notes that the problem with things like articulatory diagrams is that “learners cannot see inside their own heads”. Visual articulation feedback has two ingredients that stimulate meaningful learning: it presents material multi- modally, using auditory and visual stimuli, and it provides customized input to learners, based on their own articulations (Neri et al. 2002).

Finally, technology for providing visual articulation feedback should be affordable, adding to traditional instruction in a cost-effective and efficient manner. The types of visual information used to provide indirect articulation feedback, such as spectrograms or formant plot displays, can be generated using freely available software such as Praat (Boersma, 2001) and R (R Core Team, 2014). On the other hand, direct articulation feedback often requires the use of specialized – and costly – equipment, which may be infeasible for many teachers and learners. Ultrasound technology, once prohibitively expensive, is becoming increasingly affordable and widely available (e.g., Wojtczak & Bonadonna, 2013). However, the same is not yet true for intra-oral techniques such as EPG and EMA. Particularly with EPG, custom-made artificial palates must be created for each learner. Simulation approaches such as Baldi (Massaro and Light, 2003), while cost-intensive to develop, can avoid the same challenge through the inclusion of technologies such as speech recognition that enable customized experiences for each learner.

These criteria presented above can and will interact with each other; there may be trade-offs whereby one criterion is satisfied to the detriment of another. For example, if ability to facilitate comparison is achieved through multiple monitors rather than through multiple views on a single display, then flexibility of use in different locations may be lost and cost might be increased.

5. Methods of Evaluation and Reporting

This section shifts the focus from the types and effectiveness of visual articulation feedback strategies that are reported in the literature to the studies themselves, in order to provide commentary on the methods of evaluation that were used to assess the various feedback strategies, as well as the reporting of the methods and results.

In their review of 75 studies on L2 pronunciation instruction, Thomson and Derwing (2014) list the ingredients for a “good pronunciation study” (p. 327) which was quantitative in nature. For Thomson and Derwing, a good study (i) allows replication by reporting enough detail about participants and procedures, (ii) has a large enough sample size to allow statistical analyses, and (iii) includes some way to ascertain that improvement is the result of instruction (i.e., a control group). In what follows, we comment on how well these criteria are met in the quantitative studies looking at visual articulation feedback. The majority of the studies included in this review have a quantitative component to them; exceptions include Molholt (1988, 1990), Olson (2014a), and Öster (1997).

With respect to replicability, we share Thomson and Derwing’s (2014) concern that, in many studies, details regarding the nature of instruction are either under-reported or not rigorously controlled for, making it difficult interpret the results and/or replicate the study. To give just a few illustrative examples, Quintana-Lara (2014) provides sufficient details on the procedure followed for the experimental group, but not for the control group. In Patten and Edmonds’ (2015) study, it is unclear how much pronunciation training the participants received, and how often. These are both examples of under-reporting. Olson (2014a) provides sufficient detail of his procedure, but because some of the training and recording was carried out as homework, it is impossible to determine whether the conditions were the same for all participants.

Regarding sample size, as noted by Gick et al. (2008), direct visual articulation feedback methods such as ultrasound lend themselves best to a single participant design. This factor certainly impacts the number of participants, as it is time-intensive to run an experiment with a large number of participants one at a time. Many of the studies considered in this review had under fifteen participants; the exceptions are all recent papers, including Cleland et al., 2015, Kartushina et al., 2015, Olson, 2014b, Ouni, 2014, and Quintana-Lara, 2014. Notably, the studies with small sample sizes are not designed for implementation in a classroom environment; they can rather be viewed as multiple case studies and should not be discounted because of their small numbers of participants.

As for Thomson and Derwing’s (2014) criterion of being able to isolate feedback as the factor that led to improved pronunciation, Kartushina et al. (2015) note that, of the visual articulation feedback studies they reviewed, few included a control group; those that did include one failed to ensure that the experimental and control groups were equally matched. Notably, the question of whether to include a control group in pedagogical studies is controversial, as it may be deemed unethical to expose only some students and not others to a potentially beneficial instructional aid like visual feedback (e.g., Cook, 1986). Assuming that it is appropriate to use a control group, there is still the issue of ensuring that the groups are equally matched not only in terms of their profiles but also in terms of their exposure to instruction. Not all studies take such measures. For instance, Saito’s (2007) study of the use of Praat as a device for providing visual feedback had unbalanced groups (4 experimental participants; 2 control participants) and unbalanced instruction (only the experimental group received instruction). Thus, it is impossible to conclude that the visual feedback itself was responsible for the improvements observed with the experimental group; there is no way to rule out the null hypothesis that any type of instruction would have been effective.

Beyond specifying criteria for a good quantitative study, Thomson and Derwing (2014) also comment specifically on what they deem to be effective L2 pronunciation studies, noting that assessment should include activities that reflect natural communication, such as measures of extemporaneous or spontaneous speech, and that assessment should also incorporate a delayed post-test to determine whether the effect is lasting or not. In the context of this review, we note that few studies meet these specific criteria. Exceptions include Patten and Edmonds (2015), who included a reading task as part of assessment, and Olson (2014b), who included a delayed post- test 3.5 weeks after the final training session.

6. Summary and Conclusions

In this paper, we have surveyed the literature on various types of computer-assisted visual articulation feedback in second language pronunciation training. Although it may be superficially perceived as a small area within the context of L2 research, there is growing interest in computer- assisted visual articulation feedback, as evidenced by the 41 studies addressing the topic cited in this review, and the spate of related studies in recent years. We anticipate that, as technologies continue to evolve, the number of studies in this area will continue to grow. Developments towards handheld ultrasound technology (Bruce et al., 2000; Clarius, 2016; Wojtczak & Bonadonna, 2013) and improvements in audiovisual speech simulations (Mattheyses & Verhelst, 2015) will likely provide fruitful avenues for language instructors, researchers and industry to develop new methods for delivering and evaluating visual articulation feedback in L2 learning environments. In addition to technological developments, we hope to see the field continue to evolve towards better meeting the criteria for the effective use of visual articulation feedback (e.g., Carey, 2004; Öster, 1997), as well as improving methods of evaluation and reporting (see Thomson & Derwing, 2014).

Acknowledgments

* We wish to thank three anonymous reviewers and the editors for their helpful feedback on this manuscript. This work has been supported by a Banting Postdoctoral Fellowship awarded to the first author, an NIH grant DC-02717 awarded to Haskins Laboratories, and the UBC Teaching and Learning Enhancement Fund.

References

  1. Abberton E, & Fourcin AJ (1975). Visual feedback and the acquisition of intonation. In Lenneberg EH & Lenneberg E. (Eds.). Foundations of language development: A multidisciplinary approach, vol. 2 (pp. 157–165). Paris: UNESCO. [Google Scholar]
  2. Akahane-Yamada R, McDermott E, Adaichi T, Kawahara H, & Pruitt JS (1998). Computer-based second language production training by using spectrographic representation and HMM-based speech recognition scores. Paper presented at the 1998 International Conference on Spoken Language Processing, Sydney, Australia. Retrieved on 16 December 2015 from http://www.mirlab.org/conference_papers/International_Conference/ICSLP%201998/PDF/AUTHOR/SL980429.PDF. [Google Scholar]
  3. Aliaga-García C, & Mora JC (2009). Assessing the effects of phonetic training on L2 sound perception and production. In Watkins MA, Rauber AS, & Baptista BO (Eds.). Recent research in second language phonetics/phonology: Perception and production (pp. 2–31). Newcastle upon Tyne, UK: Cambridge Scholars Publishing. [Google Scholar]
  4. Anderson F. (1960). An experimental pitch indicator for training deaf scholars. Journal of the Acoustical Society of America, 32(8), 1065–1074. [Google Scholar]
  5. Badin P, Ben Youssef A, Bailly G, Elisei F & Hueber T. (2010). Visual articulatory feedback for phonetic correction in second language learning. Proceedings of the Workshop on Second Language Studies: Acquisition, Learning, Education, and Technology, 1–10. [Google Scholar]
  6. Ballard KJ, Smith HD, Paramatmuni D, McCabe P, Theodoros DG, & Murdoch BE (2012). Amount of kinematic feedback affects learning of speech motor skills. Motor Control, 16, 106–119. [DOI] [PubMed] [Google Scholar]
  7. Bernhardt B, Gick B, Bacsfalvi P, & Ashdown J. (2003). Speech habilitation of hard of hearing adolescents using electropalatography and ultrasound as evaluated by trained listeners. Clinical Linguistics & Phonetics, 17(3), 199–216. [DOI] [PubMed] [Google Scholar]
  8. Bernhardt B, Gick B, Bacsfalvi P, & Adler-Bock M. (2005). Ultrasound in speech therapy with adolescents and adults. Clinical Linguistics & Phonetics, 19(6/7), 605–617. [DOI] [PubMed] [Google Scholar]
  9. Boersma P. (2001). Praat, a system for doing phonetics by computer. Glot International, 5(9/10), 341–345. [Google Scholar]
  10. Bruce CJ, Spittell PC, Montgomery SC, Bailey KR, Tajik AJ, & Seward JB (2000). ultrasound imager: abdominal aortic aneurysm screening. Journal of the American Society of Echocardiography, 13, 674–679. [DOI] [PubMed] [Google Scholar]
  11. Carey M. (2004). CALL visual feedback for pronunciation of vowels: Kay Sona-Match. CALICO Journal, 21(3), 571–601. [Google Scholar]
  12. Catford JC & Pisoni DB (1970). Auditory versus articulatory training in exotic sounds. The Modern Language Journal, 54(7), 477–481. [Google Scholar]
  13. Celce-Murcia M, Brinton DM, & Goodwin JM (1996). Teaching pronunciation: a reference for teachers of English to speakers of other languages. Cambridge: Cambridge University Press. [Google Scholar]
  14. Chun DM (1989). Teaching tone and intonation with microcomputers. CALICO Journal, 7(1), 21–46. [Google Scholar]
  15. Chun DM (1998). Signal analysis software for teaching discourse intonation. Language Learning & Technology, 2(1), 61–77. [Google Scholar]
  16. Chun DM (2002). Discourse intonation in L2: From theory and research to practice. Amsterdam: Benjamins. [Google Scholar]
  17. Clarius. (2016). Wireless, handheld ultrasound for iOS and Android debuts. [Press release]. Retrieved from https://www.clarius.me/aium-debut-pr/.
  18. Cleland J, Scobbie JM, Nakai S, & Wrench A. (2015). Helping children learn non-native articulations: the implications for ultrasound-based clinical intervention. Paper presented at the 2015 International Conference of Phonetic Sciences, Glasgow, Scotland. Retrieved on 12 August 2015 from http://www.icphs2015.info/pdfs/Papers/ICPHS0698.pdf. [Google Scholar]
  19. Cook V. (Ed.). (1986). Experimental approaches to second language learning. Oxford: Pergamon. [Google Scholar]
  20. de Bot CLJ (1980). The role of feedback and feedforward in the teaching of pronunciation. System, 8, 35–45. [Google Scholar]
  21. Demenko G, Wagner A, & Cylwik N. (2010). The use of speech technology in foreign language pronunciation training. Archives of Acoustics, 35(3), 309–329. [Google Scholar]
  22. Dowd A, Smith J, & Wolfe J. (1997). Learning to pronounce vowel sounds in a foreign language using acoustic measurements of the vocal tract as feedback in real time. Language and Speech, 41(1), 1–20. [Google Scholar]
  23. Engwall O. (2012). Analysis of and feedback on phonetic features in pronunciation training with a virtual teacher. Computer Assisted Language Learning, 25(1), 37–64. [Google Scholar]
  24. Gick B, Bernhardt B, Bacsfalvi P, & Wilson I. (2008). Ultrasound imaging applications in second language acquisition. In Hansen Edwards JG & Zampini ML (Eds.). Phonology and second language acquisition (pp. 309–322). Amsterdam: John Benjamins. [Google Scholar]
  25. Hardison DM (2004). Generalization of computer-assisted prosody training: Quantitative and qualitative findings. Language Learning & Technology, 8, 34–52. [Google Scholar]
  26. Hincks R. (2015). Technology and learning pronunciation. In Reed M & Levis JM (Eds.) The Handbook of English Pronunciation (pp. 505–519). New Jersey: Wiley and Sons. [Google Scholar]
  27. Jenson PG, & Westermeier FX (1968). The effect of visual feedback on pronunciation in foreign language learning. Retrieved on 29 August 2015 from http://files.eric.ed.gov/fulltext/ED015689.pdf.
  28. Kalikow DN, & Swets JA (1972). Experiments with computer-controlled displays in second-language learning. IEEE Transactions on Audio and Electroacoustics, AU-20(1), 23–28. [Google Scholar]
  29. Kartushina N, Hervais-Adelman A, Frauenfelder UH, & Golestani N. (2015). The effect of phonetic production training with visual feedback on the perception and production of foreign speech sounds. Journal of the Acoustical Society of America, 138(2), 817–832. [DOI] [PubMed] [Google Scholar]
  30. Katz W, Campbell T, Wang J, Farrar E, Eubanks JC, Balasubramanian A, Prabhakaran B, & Rennaker R. (2014). Opti-Speech: a real-time, 3D visual feedback system for speech training. Proceedings of Interspeech 2014, Singapore, 1174–1178. Retrieved on 22 January 2016 from https://www.utdallas.edu/~wangjun/paper/Interspeech14_opti-speech.pdf. [Google Scholar]
  31. Katz WF, & Mehta S. (2015). Visual feedback of tongue movement for novel speech sound learning. Frontiers in Human Neuroscience, 9, 612. DOI: 10.3389/fnhum.2015.00612. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kelsey CA, Minifie FD, & Hixon TJ (1969). Applications of ultrasound in speech research. Journal of Speech, Language, and Hearing Research, 12(3), 564–575 [DOI] [PubMed] [Google Scholar]
  33. Lambacher S. (1999). A CALL tool for improving second language acquisition of English consonants by Japanese learners. Computer Assisted Language Learning, 12(2), 137–156. [Google Scholar]
  34. Lee J, Jang J, & Plonksy L. (2015). The effectiveness of second language pronunciation instruction: A meta-analysis. Applied Linguistics, 36(3), 345–355. [Google Scholar]
  35. Léon PR, & Martin P. (1972). Applied linguistics and the teaching of intonation. The Modern Language Journal, 56(3), 139–144. [Google Scholar]
  36. Levis JM & Pickering L. (2004). Teaching intonation in discourse using speech visualization technology. System, 32(4), 505–524. [Google Scholar]
  37. Levitt JS, & Katz WF (2007). Augmented visual feedback in second language learning: training Japanese post-alveolar flaps to American English speakers. Journal of the Acoustical Society of America, 122(5), 2996. [Google Scholar]
  38. Massaro DW, & Light J. (2003). Read my tongue movements: bimodal learning to perceive and produce non-native speech /r/ and /l/. Proceedings of the 8th European Conference on Speech Communication and Technology. [Google Scholar]
  39. Mattheyses W & Verhelst W. (2015). Audiovisual speech synthesis: An overview of the state- of-the-art. Speech Communication, 66, 182–217. [Google Scholar]
  40. Moisik SR (2013). The epilarynx in speech. Unpublished doctoral dissertation, University of Victoria. [Google Scholar]
  41. Moisik SR, Esling JH, Bird S, & Lin H. (2011). Evaluating laryngeal ultrasound to study larynx state and height. In Lee WS & Zee E. (Eds.), Proceedings of the 17th International Congress of Phonetic Sciences Hong Kong, 136–139. [Google Scholar]
  42. Molholt G. (1988). Computer-assisted instruction in pronunciation for Chinese speakers of American English. TESOL Quarterly, 22(1), 91–111. [Google Scholar]
  43. Molholt G. (1990). Spectrographic analysis and patterns in pronunciation. Computers and the Humanities, 24(1/2), 81–92. [Google Scholar]
  44. Navarra J, & Soto-Faraco S. (2007) Hearing lips in a second language: visual articulatory information enables the perception of second language sounds. Psychological Research 71, 4–12. [DOI] [PubMed] [Google Scholar]
  45. Neri A, Cucchiarini C, Strik H, & Boves L. (2002). The pedagogy-technology interface in computer-assisted pronunciation training. Computer-Assisted Language Learning, 21(5), 393–408. [Google Scholar]
  46. Noguchi M, Yamane N, Tsuda A, Kazama M, Kim B, & Gick B. (2015). Towards protocols for L2 pronunciation training using ultrasound imaging. Poster presentation at the 7th annual Pronunciation in Second Language Learning and Teaching (PSLLT) Conference. Dallas, Texas: October 2015. [Google Scholar]
  47. Olson DJ (2014a). Phonetics and technology in the classroom: a practical approach to using speech analysis software in second-language pronunciation instruction. Hispania, 97(1), 47–68. [Google Scholar]
  48. Olson DJ (2014b). Benefits of visual feedback on segmental production in the L2 classroom. Language Learning and Technology, 18(3), 173–192. [Google Scholar]
  49. Öster A-M (1997). Auditory and visual feedback in spoken L2 teaching. Reports from the Department of Phonetics, Umeå University (PHONUM), 4, 145–148. [Google Scholar]
  50. Ouni S. (2014). Tongue control and its implication in pronunciation training. Computer Assisted Language Learning, 27(5), 439–453. [Google Scholar]
  51. Patten I, & Edmonds LA (2015). Effect of training Japanese L1 speakers in the production of American English /r/ using spectrographic visual feedback. Computer Assisted Language Learning, 28(3), 241–259. [Google Scholar]
  52. Pillot-Loiseau C, Kamiyama T, & Kocjančič Antolík T. (2015). French /y/-/u/ contrast in Japanese learners with/without ultrasound feedback: vowels, non-words and words. Paper presented at the 2015 International Conference of Phonetic Sciences, Glasgow, Scotland. Retrieved on 12 August 2015 from http://www.icphs2015.info/pdfs/Papers/ICPHS0485.pdf. [Google Scholar]
  53. Quintana-Lara M. (2014). Effect of acoustic spectrographic instruction on production of English /i/ and /ɪ/ by Spanish pre-service English teachers. Computer Assisted Language Learning, 27(3), 207–227. [Google Scholar]
  54. R Core Team (2014). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. [Google Scholar]
  55. Saito K. (2007). The influence of explicit pronunciation instruction on pronunciation in EFL settings: the case of English vowels and Japanese learners of English. The Linguistics Journal, 3(3), 16–40. [Google Scholar]
  56. Schwartz B. (1993). On explicit and negative data effecting and affecting competence and linguistic behavior. Studies in Second Language Acquisition, 15, 147–163. [Google Scholar]
  57. Stone M. (2005). Preface to the special issue on ultrasound imaging of the tongue. Clinical Linguistics & Phonetics, 19(6–7), 453–454. [Google Scholar]
  58. Suemitsu A, Dang J, Ito T, & Tiede M. (2015). A real-time articulatory visual feedback approach with target presentation for second language pronunciation learning. Journal of the Acoustical Society of America, 138(4), EL382–EL387. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Tateishi M, & Winters S. (2013). Does ultrasound training lead to improved perception of a non-native sound contrast? Evidence from Japanese learners of English. Paper presented at the 2013 meeting of the Canadian Linguistic Association, Victoria, BC, Canada. Retrieved on 12 August 2015 from http://homes.chass.utoronto.ca/~cla-acl/actes2013/Tateishi_and_Winters-2013.pdf [Google Scholar]
  60. Thomson R, & Derwing T. (2014). The effectiveness of L2 pronunciation instruction: A narrative review. Applied Linguistics, 36(3): 326–344. [Google Scholar]
  61. Tilsen S, Das D, & McKee B. (2015). Real-time articulatory biofeedback with electromagnetic articulography. Linguistics Vanguard, 1(1), 39–55. DOI: 10.1515/lingvan-2014-1006. [DOI] [Google Scholar]
  62. Truscott J. (2007). The effect of error correction on learners’ ability to write accurately. System, 16: 255–272. [Google Scholar]
  63. Tsui HM (2012). Ultrasound speech training for Japanese adults learning English as a second language. Unpublished MSc thesis, University of British Columbia. [Google Scholar]
  64. Vardanian RM (1964). Teaching English intonation through oscilloscope displays. Language Learning, 14(3–4), 109–117. [Google Scholar]
  65. Wilson I. (2014). Using ultrasound for teaching and researching articulation. Acoustical Science and Technology, 35(6), 285–289. [Google Scholar]
  66. Wilson I, & Gick B. (2006). Ultrasound technology and second language acquisition research. In Grantham O’Brien M, Shea C, & Archibald J. (Eds.). Proceedings of the 8th Generative Approaches to Second Language Acquisition Conference (GASLA 2006) (pp. 148–152). Somerville, MA: Cascadilla Proceedings Project. [Google Scholar]
  67. Wojtczak J, & Bonadonna P. (2013). Pocket mobile smartphone system for the point-of-care submandibular ultrasonography. The American Journal of Emergency Medicine, 31, 573–577. [DOI] [PubMed] [Google Scholar]
  68. Wu Y, Gendrot C, Hallé P, & Adda-Decker M. (2015). On improving the pronunciation of French /r/ in Chinese learners by using real-time ultrasound visualization. Paper presented at the 2015 International Conference of Phonetic Sciences, Glasgow, Scotland. Retrieved 12 August 2015 from http://www.icphs2015.info/pdfs/Papers/ICPHS0786.pdf. [Google Scholar]

RESOURCES