Skip to main content
Journal of Speech, Language, and Hearing Research : JSLHR logoLink to Journal of Speech, Language, and Hearing Research : JSLHR
. 2023 Jan 12;66(8 Suppl):3038–3051. doi: 10.1044/2022_JSLHR-22-00322

Factors Affecting Nonnative Consonant Cluster Learning

Adam Buchwald a,, Hung-Shao Cheng a
PMCID: PMC10555463  PMID: 36634242

Abstract

Purpose:

Nonnative consonant cluster learning has become a useful experimental approach for learning about speech motor learning, and we sought to enhance our understanding of this area and to establish best practices for this type of research.

Method:

One hundred twenty individuals completed a nonnative consonant cluster learning task within a speech motor learning paradigm. Following a brief prepractice, participants then practiced the production of eight word-initial nonnative consonant clusters embedded in bisyllabic nonwords (e.g., GD in /gdivu/). The clusters ranged in difficulty according to linguistic typology and sonority sequencing. Acquisition was operationalized as the change across the practice section and learning was assessed with two retention sessions (R1: 30 min after practice; R2: 2 days after practice). We evaluated changes in accuracy as well as in the acoustic details of the cluster production at each time point.

Results:

Overall, participants improved in their production of the consonant clusters. Accuracy increased, and duration measures decreased in specific measures associated with cluster production. The change in coordination measured in the acoustics changed both for clusters that were incorrectly produced and for those that were correctly produced, indicating continued motor learning even in accurate tokens.

Conclusions:

These results aid our understanding of the complexity of nonnative consonant cluster learning. In particular, both factors related to both phonological and speech motor control properties affect the learning of novel speech sequences.

Supplemental Material:

https://doi.org/10.23641/asha.21844185


Accounts of spoken production have typically focused either on the set of processes that occur prior to phonetic/motor planning as in psycholinguistic accounts (e.g., Levelt et al., 1999) or on processes that occur after specifying the sounds to be produced as in accounts of speech motor control (e.g., Guenther, 2016). This distinction has arisen because the types of representations and processes that take place in the psycholinguistic accounts are primarily language-based whereas speech motor control accounts need to explain speech-based processes that interface directly with the motor system and the speech articulators. While speech motor control frameworks broadly incorporate elements of linguistic structure (e.g., GODIVA [Gradient Order Directions Into Velocities of Articulators] model: Bohland et al., 2010; Guenther, 2016), these frameworks frequently treat learning to produce novel consonant clusters as a type of sequence learning (akin to finger tapping), with no formal consideration of differences among specific consonant sequences (Segawa et al., 2015, 2019). In addition, the GODIVA-based accounts of cluster learning as a type of sequence learning define cluster learning as forming a chunk in which specific sound combinations are stored, retrieved, and produced as a single unit. In this article, we present nonnative cluster learning data from a motor learning paradigm that we have used in several previous studies (Buchwald et al., 2019; Cheng & Buchwald, 2021; Lowe & Buchwald, 2017) that indicates that the specific combination of sounds in the cluster are critical to learning, and that production refinement continues even once clusters are produced accurately. We suggest that GODIVA-based accounts can capture our core findings by refining the specification of the speech sound map to include additional details of phonological structure.

In this introduction, we begin with a detailed description of the type of phonological information that underlies the representation of sound structure as has been demonstrated in linguistics. We then consider the relationship between this information and current accounts of speech motor sequence learning. Finally, we introduce our study that identifies specific aspects of sound structure representation that are critical to understanding how people learn to produce nonnative consonant clusters.

Consonant Clusters in Linguistic Accounts

The notion of a consonant cluster in linguistics is primarily based on the concept of the syllable (Kahn, 1976). Sound sequences are hierarchically organized into syllables consisting of an onset, nucleus, and coda, and a consonant cluster is defined as multiple consonants sharing the onset or coda position. At a broad level, languages differ typologically with respect to the syllable structure they permit (Greenberg, 1978), with some languages like English allowing up to three consonants in the onset (e.g., spleen, /splin/) and four in the coda (e.g., sixths, /sɪksθs/) and other languages permitting a maximum of one single onset consonant (e.g., Japanese) or no coda consonants (e.g., Fijian).

Additionally, all languages that permit consonant clusters have constraints on which sounds can be combined in clusters. These constraints are often described with respect to the notion of sonority, and more specifically the sonority sequencing principle (Clements, 1990). Sonority is a phonological construct that broadly corresponds to the openness of the vocal tract, so that stop consonants have the lowest sonority due to complete closure, followed by fricatives, nasals, liquids, and glides, with vowels having the highest sonority. The sonority sequencing principle holds that syllables should increase in sonority from the margin (first and last sounds) to the peak (the nucleus of the syllable, typically a vowel). Languages that allow consonant clusters can differ with respect to what type of increase in sonority is permitted between the two consonants. For example, English permits stop-liquid clusters in onset position (e.g., /kl/) that have a relatively large increase, but not stop-nasal clusters (e.g., */kn/), stop-fricative clusters (e.g., */ks/), or stop-stop clusters (e.g., */kt/). In contrast, German allows both stop-liquid, stop-nasal, and stop-fricative clusters but not the stop-stop clusters that have a sonority plateau, and Russian allows all of these cluster types. Thus, the general sonority sequencing principle has different setting across languages (see Broselow & Finer, 1991, for a discussion of cross-linguistic differences and Blevins, 1995, for a more detailed discussion of sonority hierarchy and syllable structure).

In addition to providing an account of the inventory of clusters across languages, this principle helps account for systematic differences between onset clusters and coda clusters within a language as well. For example, onset clusters increase in sonority (e.g., stop-liquid clusters: /pl/ as in please) whereas coda clusters decrease in sonority (liquid-stop clusters: /lp/ as in kelp), but neither of these clusters can appear in the other syllable position. This also demonstrates that syllable position is a critical consideration for consonant clusters, as many sound sequences can be either an onset cluster or a coda cluster but not both (cf., discussion in Segawa et al., 2019, regarding possible transfer from onset cluster learning to coda cluster).

With respect to speech production, syllable structure plays a key role at the interface between phonology and phonetic/motor planning in most accounts (Browman & Goldstein, 1988). For example, we have long known that syllable structure affects the coordination patterns among articulators, with differences between onset consonants produced as singletons versus clusters (e.g., /m/ in mall longer than /m/ in small; Klatt, 1975). The key role of syllable structure in the interface of phonology and speech motor control also helps us understand why some sequences are allowed in some syllable positions but not others. For example, English does not allow stop-stop sequences (e.g., /gd/, /kt/) in syllable onset position. However, these sequences appear in both word-medial position (e.g., Baghdad, doctor) and in coda position (e.g., jogged, pact) in English. This is easily explained with respect to syllable structure. Words with medial stop-stop clusters have a syllable boundary between the two consonants, and so the two stop consonants do not share a syllable position. Similarly, these clusters are permitted as coda clusters even though they are not permitted in onset. There is also an articulatory basis of these differences, as the onset clusters require the most complex coordination between the two sounds, followed by the coda cluster and then the heterosyllabic sequence (Byrd, 1995, 1996).

Taken together, this section has illustrated several key points. First, the identity of sounds plays a role in whether they are allowed as consonant clusters in a language, and there are certain principles that help account for these differences. There is a key corollary to this point: Some clusters that are not allowed in a language may still be “better” than others based on these properties. This idea helps shape the structure of the stimuli used in this study. Second, syllable structure helps define a fundamental difference between different types of clusters (onset vs. coda) that are crucial for understanding sound structure representation. This point indicates the importance of including a representation of syllable structure in the sound structure representations used in accounts of speech motor control. The next section will turn to the types of representations currently posited in accounts of speech motor control with respect to the discussion in this section.

Consonant Cluster Representations in Accounts of Speech Motor Control

Accounts of speech motor control seek to explain how individuals control their articulators during speech production and the factors that influence that control (for a recent overview of these accounts and their differences, see Parrell et al., 2019). The current article focuses on GODIVA (Guenther, 2016) because it is the most well-specified and used account of this type that has considered how to incorporate linguistic structure, and recent discussions of speech motor sequence learning have referred to this framework. Within GODIVA, the desired output is referred to as a speech sound map. The speech sound map has been described in terms of phonemes, syllables, and/or words. This consideration of syllabic representations further indicates that the speech sound map within GODIVA can be specified to account for some of the findings discussed above. There is widespread agreement that speech motor control consists of both feedforward and feedback control. Feedforward commands initiate speech production by specifying a desired spoken output and a series of commands for producing that output. In addition to feedforward control, accounts of speech production also posit feedback control that monitors the somatosensory and acoustic output signal and sends commands to accomplish error correction in cases of mismatch between the expected and actual signals.

This article focuses on the domain of consonant cluster learning, where GODIVA has formed the basis of interpreting the results of cluster learning studies in several populations (Masapollo et al., 2021; Segawa et al., 2015, 2019). Segawa et al. (2019) presented data from a cluster learning experiment in which participants practiced producing consonant–consonant–vowel–consonant–consonant (CCVCC) nonwords with nonnative consonant clusters in both onset and coda (e.g., SHKEPF /ʃkɛpf/). The stimuli were recorded by a native English speaker and presented auditorily and orthographically to participants during two training sessions. Each consonant cluster was practiced (during training) in a single stimulus word and then tested in the trained word and an additional item. Accuracy was measured perceptually and whole word duration was measured from the acoustic record. The participants improved in their accuracy and shortened their duration in the practiced syllables, with the effect of shortened duration generalizing to novel syllables with the practiced clusters. Segawa et al. argued that the initial representation of these novel sequences CCVCC includes a sequence of five segments. As participants practice the novel clusters in the larger novel syllables, they then form representations (working memory chunks) for both the whole cluster and the whole syllable. These chunks are part of the speech sound map and allow for faster and more accurate retrieval and production in the future. Segawa et al. did not distinguish between onset and coda clusters, but only compared onsets to onsets and codas to codas. As we indicated earlier, we believe that there is sufficient evidence to warrant a clear distinction between the two cluster types, given that many onset clusters are not legal in coda position and vice versa, in addition to differences in articulatory coordination between the consonants and the vowel in these two syllable positions (Marin & Pouplier, 2010).

In our view, the essence of this explanation of how consonant clusters are learned is broadly accurate, but requires refinement to account for the full set of findings on cluster learning. First, our previous work has indicated that learning on one cluster type (e.g., voiced stop-stop clusters) can transfer to improved production of different but similar clusters (e.g., voiceless stop-stop clusters; Cheng & Buchwald, 2021). This suggests that learning occurs over representations that are not limited to specific phonemes; that is, if training on /gd/ generalizes to /kt/, then what is learned must be defined differently than just the /gd/ chunk. Second, the proposal does not formally distinguish between clusters based on the identity of the sounds, and thus cannot capture that some clusters may be easier to produce accurately and to learn. In the previous section, we described a number of systematic differences across languages in what clusters are legal. This article further examines these systematic differences in learning based on different types of consonant cluster sequences. Finally, the GODIVA-based account also appears to treat the formation of the chunk as binary (e.g., there is or is not a memory chunk). Our previous work (Buchwald et al., 2019) had reported that errored tokens get closer to the target during learning (described more below). In this article, we further probed the extent to which individuals improve their cluster productions prior to being accurate, and also included analyses focusing on whether phonemically accurate production is the endpoint of learning.

Considerations in Studying Consonant Cluster Production Learning

This study was designed to identify aspects of linguistic phonological representations as well as phonetic details that must be considered in a robust account of nonnative consonant cluster learning. As in our other recent work (Buchwald et al., 2019; Cheng & Buchwald, 2021; Lowe & Buchwald, 2017), our experimental training of nonnative consonant clusters is grounded in a speech motor learning paradigm. This paradigm is part of a larger direction to build an understanding of motor learning in the speech domain with respect to domain-general accounts of motor learning (see Maas et al., 2008, for a comprehensive review). Although this paradigm focuses on motor learning, learning to produce consonant clusters is at the interface between phonological and motor processing and we believe that there are several considerations that need to be taken into account in designing these studies to remove potential experimental confounds that can complicate interpretation. Therefore, several aspects of our design are specific to the details of learning to produce novel speech sound sequences, and build from large bodies of research in linguistics and psycholinguistics.

There is a large body of evidence indicating that native speakers of a language without a cluster have difficulty accurately perceiving those clusters (Berent & Lennertz, 2007; Davidson & Shaw, 2012; Dupoux et al., 1999, 2011; Guevara-Rukoz et al., 2017). These perceptual difficulties influence our approach in three ways that we believe should be standard in this area. First, when presenting speakers with an auditory stimulus to repeat that includes a nonnative cluster (e.g., /gdivu/), we also include an orthographic transcription of the work (e.g., GDEEVOO) to ensure that production errors do not arise from perceptual limitations. Second, we measure both accuracy and presence of an inserted vowel based exclusively on the acoustic record (see Method section); although other studies have evaluated accuracy of nonnative clusters based on perceptual analyses, there is too much evidence indicating that the perception of nonnative cluster accuracy is unreliable (Davidson, 2007; Davidson & Shaw, 2012). Finally, our auditory models come from speakers of a language that contains the trained clusters. There is evidence that the precise details of the auditory model (e.g., stop burst amplitude) can influence the accuracy of individuals repeating words with nonnative clusters (Wilson et al., 2014) and we, therefore, select speakers who can avoid artifacts that may come from the unnaturalness of the productions. In addition, we ensure that the model productions meet an acoustic definition of being produced accurately as has been previously used in other studies (Buchwald et al., 2019; Cheng & Buchwald, 2021; Wilson et al., 2014).

In addition to perceptual limitations with perceiving nonnative clusters, there is a large body of linguistic work devoted to understanding the cross-linguistic typology of legal syllable structure. This includes the focus on sonority sequencing as discussed above, as well as general work identifying attested and unattested cluster types. This literature led us to generate a specific range of clusters to examine, as well as to avoid two specific consonant cluster sequences in our work due to the complications they provide for the analyses. First, we avoid including mixed voicing obstruent–obstruent clusters, such as /sb/ or /kz/, as these are largely unattested in the world's languages and are considered marked. One related phonetic property of these sequences that may be related to the unlikelihood of finding them in languages is the difficulty of timing the laryngeal gesture for one but not the other consonant, particularly when both are produced with a relatively closed vocal tract (Lindblom, 1990; Lombardi, 1999). Furthermore, even in languages where these clusters are attested in the description of the phonology, they are often produced with the same voicing specification for each consonant anyway (Kreitman, 2010). Second, we avoid consonant sequences that can be produced as a complex single segment or phoneme, such as an affricate. Although English only includes two affricates in the inventory (i.e., /ʧ/ and /ʤ/), we avoid all homorganic stop-fricative clusters that are often produced as affricates in other languages (e.g., /ʦ/ and /ʣ/) as well as in English loanwords (tse-tse). While some languages such as Polish distinguish between the affricate and the cluster composed of the same sounds, this is rare cross-linguistically and difficult to measure in the context of these learning studies.

Taken together, the perceptual limitations and differences among specific consonant clusters demonstrate that consonant cluster learning is a complex phenomenon with considerations that differ from other forms of studying sequence learning such as finger tapping. Given the importance of generating replicable results, as well as the recognition of how small changes in experimental setups can alter result patterns for reasons unrelated to the question being asked (Wisler et al., 2022), we believe that the use of consonant cluster learning to address questions about motor control must include strong consideration of the properties of the clusters being learned.

This Study

In this study, we explore two overarching questions about the nature of learning nonnative consonant clusters. First, we are interested in understanding how different classes of clusters are learned and how much the baseline accuracy of a cluster predicts overall learning. To address this, we examined three types of word-initial onset cluster based on sonority sequencing: fricative-nasal (/fm/, /fn/, /vm/, and /vn/), fricative-stop (/zb/ and /zg/), and stop-stop (/gd/ and /pt/). Previous work examining English speakers producing these sequences has indicated that stop-stop clusters are the most difficult followed by fricative-stop and then fricative-nasal (Davidson, 2010), but also that the voiced clusters are much more difficult than the voiceless (Davidson, 2006). We predicted that those sequences that are more difficult to produce at baseline are also more difficult to learn and, thus, that overall the stop-stop clusters would be the least accurate followed by fricative-stop and then fricative-nasal. Furthermore, we predicted that within the two groups with voicing distinctions, the voiced clusters would be less accurate than the voiced clusters both at baseline as well as after learning. Thus, we treated the cluster stimuli as containing five classes of consonant clusters: voiceless fricative-nasals (fN), voiced fricative-nasals (Vn), voiced fricative-stops (Zd), voiceless stop-stops (pt), and voiced stop-stops (gd).

Second, we are interested in understanding the trajectory of learning. In previous work (Buchwald et al., 2019), we examined vowel epenthesis errors in individuals learning to produce /gd/ clusters (e.g., /gdivu/ ➔ [gədivu]) and found that the vowels produced in error got shorter (i.e., closer to the target) during learning even when they were produced incorrectly. This suggests that during learning, individuals come successively closer to the target of learning even in trials that contain errors. In this article, we sought to replicate and extend that finding by examining a different measure of stop-stop cluster production with higher interrater reliability (burst-to-burst duration), and asking whether we see a decrease in duration on both /gd/ and /pt/ clusters through training. The sounds in the stop-stop clusters differ in place of articulation to ensure two distinct releases where coordination can change over time. In addition, we ask whether clusters that are produced accurately based on the acoustic absence of errors continue to change and be produced with a more native-like coordination pattern during learning. For this question, we consider /fn/ and /fm/ clusters that have been found to be produced with high accuracy rates by native English speakers (Buchwald et al., 2019), and ask whether the nasal duration is shorter in accurately produced clusters at retention.

Method

Participants

One hundred twenty participants (range: 18–39 years, M = 23.9 years) who were recruited from advertisements and flyers in the New York University (NYU) community completed the study. Exclusion criteria included a history of speech or language impairment, familiarity with languages that contained any of the specific consonant clusters being trained (e.g., Russian, Polish, Greek, Hebrew), as well as phonetic training through academic coursework. All participants reported normal or corrected-to-normal vision, and informed consent was obtained according to the NYU Langone Medical Center Institutional Review Board. Participants were recruited as part of a larger study on how noninvasive brain stimulation affects speech motor learning, and two thirds of the participants here were part of a previous write-up on how neuromodulation can affect speech motor learning (Buchwald et al., 2019). In order to make use of this large data set to focus on the aspects of nonnative consonant cluster learning, the analyses in this article include data from all participants and are collapsed across neuromodulation conditions.

Speech Stimuli

The training and testing stimuli were trochaic disyllabic nonwords with word-initial nonnative native consonant clusters (e.g., GDEEVOO; FNEEGDWOP). We used trochaic disyllabic words to be more “word-like” than monosyllabic items, including both stressed and unstressed syllables. The stimuli began with eight nonnative clusters (see Table 1). Each nonword consisted of legal English sounds and sequences other than the initial consonant cluster, but they differed in syllable shape (e.g., GDEEVOO /gdivu/ has the shape [CCV.CV]; FNEEGDWOP /fnigdwɑp/ has the shape [CCVC.CCVC]; these were balanced across clusters). Four nonwords for each cluster were used during the practice session, and four were saved for retention session testing (see Procedure section).

Table 1.

Consonant clusters and classes of clusters used in the experiment.

Cluster class Cluster C1 C2
fN fm, fn voiceless fricative nasal
vN vm, vn voiced fricative nasal
zD zb, zg voiced fricative voiced stop
pt pt voiceless stop voiceless stop
gd gd voiced stop voiced stop

All auditory stimuli were recorded by a phonetically trained Russian–English bilingual speaker who was instructed to produce the clusters as they would in Russian but the rest of the word as they would in English. As discussed in the introduction, this speaker was selected because they could produce the clusters accurately as a native speaker of a language containing these sequences. All cluster stimuli were verified as containing an accurately produced cluster based on the same criteria used for scoring the data, described below. The recordings were made with a Shure SM-10 head-mounted microphone attached to a Marantz PMD660 digital recorder. The files were then spliced to leave 10 ms of silence at the onset and offset of each item using Praat (Boersma & Weenink, 2022), and the stimulus amplitude was normalized. Orthographic versions of the nonwords were also presented to participants to ensure that errors in cluster production did not arise from misperception, as discussed in the introduction.

Speech Motor Learning Paradigm

Testing was conducted in a sound-attenuated testing room. Participants sat in front of a computer and their productions were recorded using a Shure BETA 58A microphone in a desktop microphone stand connected to the Marantz PMD660.

Prepractice

The task began with a short prepractice session to ensure that individuals understood the task of trying to produce consonant clusters. In this session, we had participants produce two monosyllabic nonwords (/fnɪt/ and /fteɪk/) that we used because English listeners' perception of these particular clusters is typically more accurate than of the other clusters in the study (Davidson, 2010), allowing us to provide accurate feedback. Feedback was provided regarding whether the target was produced accurately (knowledge of results) and also what specific aspects of the production were wrong in cases where there was an error (knowledge of performance). All participants were also instructed to attend to the consonant cluster at word onset and to try not to produce a vowel before or in between the two consonants. Prepractice lasted approximately 2 min.

Practice

During the practice session, participants produced 160 nonwords with nonnative clusters. For each of the eight clusters, participants produced four nonwords 5 times each in pseudorandom presentation orders; the nonwords with nonnative onset clusters were interspersed with 62 additional filler items and then modified to ensure that no stimulus was presented twice in succession. Items were balanced across participants such that half of the participants were trained on one half of the nonwords and the other half on the other nonwords. For each stimulus, participants were presented with both auditory and orthographic versions simultaneously. The practice session was structured to include a large amount of practice of the clusters presented randomly and in varied contexts, in line with structuring the practice session to optimize learning (Maas et al., 2008). No online feedback was provided given the difficulty of perceiving the accuracy of these clusters. The practice session lasted approximately 18 min.

Retention

Learning in speech motor learning paradigm is evaluated with retention sessions that occur after the practice phase. We used a shorter-term retention (R1) and longer term retention (R2). The sessions were identical, with R1 beginning approximately 30 min after the end of practice, and R2 taking place in a separate session 2 days later. For each cluster, participants produced the four trained nonwords as well as four untrained nonwords beginning with the cluster. Each stimulus was produced 3 times (192 total cluster stimuli) during the retention sessions, with an additional 110 filler items not containing clusters. The stimuli were randomized and presented electronically using the E-Prime 3.0 software (Psychology Software Tools Inc., 2016). The retention sessions lasted approximately 20 min.

Data Coding

Cluster Accuracy

As discussed in the introduction, perceptual analyses are not reliable for examining the presence versus absence of an error in the consonant clusters (i.e., distinguishing between /gdivu/ and /gədivu/; Davidson, 2010). Therefore, all recorded productions containing consonant clusters were transcribed using a combination of the acoustic waveform, spectrogram, and perception. For the most common error type—vowel insertion in the cluster—the presence or absence of a vowel was established using acoustic criteria based on Wilson et al. (2014). In particular, a token was considered to have a vowel epenthesis if there was both formant structure in the spectrogram (particularly higher formants F2, F3, as these are not confusable with f o), and the waveform had included a vocalic (periodic) portion with at least two cycles. We note here, and return to in the discussion, that clusters with two voiceless sounds (which is only /pt/ in this study) may be less likely to contain these characteristics and thus may appear to be produced without error even if the same supralaryngeal gestures are produced because of the absence of voicing (similar to the voiceless vowel produced during production of a word such as potato in fast speech).

Even though cluster accuracy was the primary dependent variable for accuracy, the whole word was transcribed. Other errors that occurred within the cluster such as consonant deletion (e.g., /gdivu/ → [divu]) and substitution (e.g., /gdivu/ → [glivu]) were identified perceptually and verified in the acoustic record. The remaining sounds in the nonwords (e.g., the [ivu] portion of /gdivu/) were transcribed perceptually. Each participant was fully scored by a single coder who was blinded to the session that each token came from during coding. All raters were trained based on the same stimulus set and had to achieve a high level of accuracy prior to analyzing participant data. Interrater reliability on the accuracy measure was evaluated based on 20% of the participants (n = 24) who were fully coded completely by two separate raters (point-to-point interrater agreement: 87.1%).

Burst-to-Burst Duration

For all stop-stop clusters (items beginning with /gd/ and /pt/) that were either produced correctly or produced with an epenthetic vowel, we measured the duration from the acoustic burst of the first stop to the burst of the second stop as a means of evaluating changes in articulatory coordination. This was done to provide a more fine-grained measure of learning than can be obtained from accuracy alone, and to determine whether there are changes in articulatory coordination in the absence of improvement in accuracy. Each burst was identified in the waveform as the first zero crossing point after the first trough of the acoustic burst using Praat's function of going to the next zero crossing. For the /gd/ stimuli that contained velar stops that commonly have more than one visible burst (Repp & Lin, 1989), we measured from the final visible burst. Interrater reliability was evaluated on 20% of the data coded by two independent raters, with agreement evaluated based on whether the two measurements were within 10 ms (point-to-point interrater agreement: 96%).

Outlier trials were identified using median absolute deviation (MAD; Leys et al., 2013). Each participant's production was compared against the median duration value of their own productions, and trials that fell two MADs away from the median were excluded. This resulted in the removal of 189 trials (less than 1.5% of the total 13,852 in the statistical analyses).

Nasal Duration

In order to evaluate whether there are changes in articulatory coordination based on learning even among items that are already correct, we measured the duration of the nasal consonant in each of the correctly produced /fn/ and /fm/ clusters based on the acoustic record. The onset of the nasal was identified based on changes from the fricative to the nasal in the spectrogram (e.g., offset of frication, onset of voicing, onset of antiformant structure) and the waveform (e.g., higher amplitude and more periodic than fricative), and the offset was established based on changes from the nasal to the vowel in the spectrogram (e.g., onset of vowel formants) and in the waveform (e.g., increase of amplitude in vowel). All measures were again taken at zero crossings using Praat, and interrater agreement was established based on measures within 5 ms (point-to-point inter-rater agreement: 94%). The same outlier removal approach described above was applied to nasal duration as well, again resulting in approximately 1.2% of trials being removed (161 removed, 13,628 retained).

Statistical Analysis

All data as well as reproducible scripts for statistical analyses and data visualization can be found on the OSF repository at https://osf.io/b6gvk/. To address the question of whether different classes of consonant clusters are learned differently, we examined the accuracy of all items during the practice session as well as at the two retention time points. The practice session was split into five equal quintiles (as each item was produced 5 times), and the first quintile (Q1) served as the baseline for comparing to the retention timepoints in order to evaluate learning. Note that our interest was in the comparison between this baseline time point (Q1) and the retention sessions that informed several choices in our analysis. Each statistical analysis was performed in R (R Core Team, 2017). We used a logistic mixed-effects models with cluster accuracy as the dependent variable and treatment coding for the categorical variables. This coding approach allowed us to easily examine simple effects of interest by examining the model with different reference groups, and allowed us not to run tests (and correct for tests) comparing the different retention groups. Our primary interest was in learning that was evaluated by comparing the performance at each retention timepoint. To evaluate this, we calculated a model with fixed effects of class (fN, vN, zD, pt, gd; see Table 1) and session (Q1, R[etention]1, R2), the interaction term of class and session and random intercepts for both participant and item. In addition, given that these data came from a larger study that included several transcranial direct current stimulation conditions, we included stimulation type as an independent fixed-effect predictor to control for any effects that could come from that, in R pseudosyntax: ClusterAccuracy ~ Class × Session + StimCond + (1|Participant) + (1|Item). As this makes the model more complex and was not part of the goal of this article, we then used the Bayesian Information Criterion (BIC; Schwarz, 1978) to compare models with and without both the random intercepts and stimulation type to decide whether to include them in the final model (Harel & McAllister, 2019). Given the treatment coding structure, we looked at the model with each class as the reference level to determine whether performance in that class improved from Q1 to the retentions, and to examine differences in accuracy within the Class × Session interaction. The reference level for session was always set to Q1.

As a secondary analysis, we considered whether trained nonwords differed from untrained nonwords. We did this as a secondary and separate analysis as the untrained nonwords were not part of the practice session, so there were no items during Q1. Performing these analyses separately allowed us to simply ask whether there was a difference between trained and untrained nonwords. We calculated a logistic mixed-effects model with cluster accuracy as the dependent variable and cluster type and train type (trained vs. untrained), as well as their interaction, as fixed effects. We also included the random intercepts for participant and item, as well as stimulation condition as a fixed effect, and used BIC to determine the final model. Because we ran two separate models that examined differences between the five cluster classes, we used a more conservative alpha level of .01 for interpreting the outcomes of both analyses (based on 0.05/5).

Burst-to-Burst Duration and Nasal Duration

For both of the acoustic analyses, we examined the duration changes from Q1 to each of the retention sessions using linear mixed-effects models with the acoustic measure as the dependent variable, and cluster and session as well as their interaction as fixed effects, and random intercepts for participant and item. We included cluster as a fixed effect as there are systematic differences between the two clusters for each measure: /pt/ has a longer burst-to-burst measure than /gd/ because of the aspiration in the voiceless clusters, and the duration of /m/ in /fm/ is systematically longer than /n/ in /fn/; by including cluster as a factor, we are able to examine each separately within the same model so that these systematic differences obviate the effects of interest. As in the accuracy analysis, we used Q1 as the reference level for session and changed the reference level for cluster to examine the effects of session for each cluster independently. Additionally, we followed each of the procedures discussed above: We ran models with and without stimulation condition and the random effects and used BIC to choose the final model. We also ran a secondary model for each analysis based on the retention data that removed the variable of session and added the fixed effect of training type (trained vs. untrained) and the interaction of training type and cluster. For this secondary analysis, we also used BIC to determine the final model. Because we ran two separate models that examined differences between two cluster classes for each of these measures, we used a more conservative alpha level of .025 for interpreting the outcomes of both analyses (based on 0.05/2).

Results

Cluster Accuracy

The cluster accuracy model was built as described above, and included 3,644 items in Q1, 22,567 items in R1, and 22,516 items in R2, and we used an alpha-level of .05. The full accuracy results for Q1, R1, and R2 sorted by consonant cluster class are presented in Figure 1 (Supplemental Material S1 depicts change during acquisition), and Table 2 presents the simple main effects of improvement for each class of consonant cluster. The final model did not include stimulation type, but did include the random intercepts for participant and item. The model revealed a significant improvement in cluster accuracy for the fN cluster class between Q1 and each of the retention sessions (R1: β = 0.38, SE = 0.11, p = .0005; R2: β = 0.58, SE = 0.11, p < .0001). For both vN and zD classes, significant improvement was only observed in the change from Q1 to R2, indicating that the extra time between the practice and retention led to an improvement in performance. For /gd/ clusters, a significant improvement in accuracy was observed between Q1 and R1, but there was no significant improvement at R2. In addition, there was no significant improvement in cluster accuracy for pt between Q1 and each of the retention sessions.

Figure 1.

Figure 1.

Overall accuracy for each consonant cluster class at each time point in the cluster accuracy analysis. Lines indicating baseline performance are provided to help visualize the change. Error bars are standard error. Q1 = first quantile; R1 = short-term retention; R2 = long-term retention.

Table 2.

Simple main effects for improvement for each cluster type from the first quintile (Q1) to the two retention sessions.

Comparison Consonant cluster type
fN
vN
zD
pt
gd
β SE p β SE p β SE p β SE p β SE p
Q1 vs. R1 0.38 0.11 < .0005 0.12 0.08 .13 0.15 0.08 .06 0.24 0.17 .15 0.3 0.13 .03
Q1 vs. R2 0.58 0.11 < .0001 0.23 0.08 .007 0.37 0.08 < .0001 0.14 0.17 .39 0.13 0.14 .36

Note. Bold text reflects significance at an alpha-level of .05. fN = voiceless fricative-nasals; vN = voiced fricative-nasals; zD = voiced fricative-stops; pt = voiceless stop-stops; gd = voiced stop-stops; R1 = short-term retention; R2 = long-term retention.

The interaction term in the model was used to examine differences in magnitude of improvement among classes. The model revealed significantly larger improvement between Q1 and R2 for the fN class compared with that for vN (β = −0.36, SE = 0.14, p = .009) and gd (β = −0.46, SE = 0.17, p = .008), with a trend toward a difference between fN and /pt/ (β = −0.44, SE = 0.20, p = .03). No other comparisons were significant.

In our secondary analysis, the final model preferred by BIC included the random intercepts for participant and item, and excluded stimulation condition. The model revealed no significant differences between trained and untrained items using an alpha level of .01, although there was a trend toward significant difference between trained and untrained items for zD clusters (β = −0.09, SE = 0.04, p = .03) with clusters in trained items produced more accurately. The interaction term in the model was used to examine differences in magnitude of differences among classes. Again, there were no statistical differences but there were two trends in the data, where the difference between trained and untrained items for zD was larger than the difference for vN (β = 0.13, SE = 0.06, p = .03), and pt (β = 0.21, SE = 0.08, p = .012).

Burst-to-Burst Duration

The burst-to-burst duration model was built as described above, and included 670 items in Q1, 4,872 items in R1, and 4,866 items in R2. On average, there was a numerical decrease in burst-to-burst duration from Q1 (M = 110 ms, SE = 1.42) to R1 (M = 106 ms, SE = 0.44) and to R2 (M = 104 ms, SE = 0.45) in the pt clusters, and the same pattern was observed for the gd clusters (Q1: M = 120 ms, SE = 1.20; R1: M = 118 ms, SE = 0.49; R2: M = 117 ms, SE = 0.46). The model revealed that there was a significant decrease for the burst-to-burst duration for pt clusters between Q1 and R2 (β = −3.83, SE = 1.13, p = .0007) but not Q1 and R1 (β = −2.04, SE = 1.13, p = .07) and between Q1 and R2. For the gd cluster, there was not a significant decrease in duration between Q1 and R1 (R1: β = −1.45, SE = 0.87, p = .10) but there was one between Q1 and R2 (β = −3.07, SE = 0.87, p = .0004). The coefficients for the interaction terms between session and cluster were not significant (R1: β = 0.59, SE = 1.42, p = .67; R2: β = 0.76, SE = 1.42, p = .60). This suggests that no difference in the magnitude of decrease was found between the two stop-stop clusters. These data are depicted in Figure 2 (Supplemental Material S2 depicts change during acquisition).

Figure 2.

Figure 2.

Overall burst-to-burst duration in /gd/ and /pt/ clusters at each time point in the analysis. Error bars represent standard error. Q1 = first quantile; R1 = short-term retention; R2 = long-term retention.

In our secondary analysis, the final model preferred by BIC included the random intercepts for participant and item, and excluded stimulation condition. The model revealed a significant difference in burst-to-burst duration between trained and untrained items for /gd/ clusters (β = 1.19, SE = 0.47, p = .01) with clusters in trained items produced with shorter burst-to-burst durations; this difference was not observed for /pt/, and there was no interaction between class and training type.

Nasal Duration

The nasal duration model was built as described above, and included 750 items in Q1, 4,802 items in R1, and 4,864 items in R2. The final model preferred by BIC included the random intercepts for participant and item, and excluded stimulation condition. The average /m/ duration in correctly produced /fm/ clusters was higher during Q1 (M = 74.2 ms, SE = 1.34) compared with each retention session (R1: M = 72.9 ms, SE = 0.52; R2: M = 72.9 ms, SE = 0.49). Similarly, the average duration for the /n/ decreased from Q1 (M = 66.2 ms, SE = 1.15) to R1 (M = 63.1 ms, SE = 0.46) and to R2 (M = 63.2 ms, SE = 0.46). The model revealed that there was a significant decrease in [m] duration between Q1 and R1 (β = −2.56, SE = 0.92, p = .005) and between Q1 and R2 (β = −2.92, SE = 0.92, p = .001). Likewise, there was a significant decrease in [n] duration between Q1 and each of the retention sessions (R1: β = −2.20, SE = 0.98, p = .02; R2: β = −2.75, SE = 0.98, p = .005). No significant interaction between session and cluster was found (R1: β = −0.36, SE = 1.34, p = .79; R2: β = −0.17, SE = 1.34, p = .90), suggesting that the magnitude of change did not differ between the two nasals. These data are depicted in Figure 3 (Supplemental Material S3 depicts change during acquisition). The secondary analysis for nasal duration did not reveal any significant differences in nasal duration at retention between trained and untrained items.

Figure 3.

Figure 3.

Overall nasal duration in correctly produced voiceless fricative-nasals (fN) cluster at each time point in the analysis. Error bars represent standard error. Q1 = first quantile; R1 = short-term retention; R2 = long-term retention.

Discussion

This article reported on an onset consonant cluster learning study with a large number of participants (N = 120) in a speech motor learning paradigm with a 20-min practice session, a same-day retention (R1) and a 2-day retention (R2). We examined two main questions. First, we asked about how consonant cluster classes differ in their accuracy and their change in accuracy following structured practice. Overall, we found that most of the cluster classes we examined (voiceless fricative-nasal, voiced fricative-nasal, fricative-stop, and voiced stop-stop) exhibited an improvement in performance from the beginning of practice to at least one of the retention sessions, whereas one did not exhibit a significant change (voiceless stop-stop), although we note that the voiced stop-stop clusters are the most difficult to analyze for accuracy, a topic we return to later in the discussion. With respect to differences among the clusters, the voiceless fN clusters started as the most accurately produced and also showed significantly greater improvement than vN and the voiced stop-stop cluster (/gd/) at the 2-day retention time point. No other differences across cluster types were observed. Second, we investigated whether we can observe acoustic evidence of learning in clusters that are produced with errors and in clusters that are produced correctly. On the first point, building on our previous work (Buchwald et al., 2019; Cheng & Buchwald, 2021), we examined changes in burst-to-burst duration in stop-stop clusters and observed significant decreases from the beginning of practice to the retention sessions, suggesting that the coordination among the articulators required for accurate onset cluster production was improving despite relatively modest accuracy changes. In addition, we observed decreases in nasal duration within correctly produced fN clusters from practice to the retention sessions, indicating that there are changes in articulatory coordination occurring even once a cluster is being produced accurately.

In the remainder of the discussion, we focus on these findings and their implications for our understanding of consonant cluster learning within speech science. We also consider the clinical implications of our findings and the limitations of our approach.

Changes in Cluster Accuracy

Our findings replicated several phenomena discussed in the phonetics literature. First, we found robust differences in the accuracy of cluster among classes of consonant cluster in overall accuracy, as is predicted by the sonority hierarchy. That is, voiceless fricative-nasal clusters were the most accurately produced and voiced stop-stop were the least accurate, replicating previous findings (e.g., Davidson, 2006, 2010) and consistent with general predictions of the relationship between sonority and syllable structure. In addition, voiceless nonnative clusters were produced more accurately than their voiced counterparts (fN vs. vN; /pt/ vs. /gd/). To the best of our knowledge, this was the first study to examine whether there are differences in a learning task among these different cluster types. In that respect, we found only one clear difference, which was that accuracy on fN clusters improved more than three of the other groups at R2 (vN, /pt/, and /gd/). Only zD showed improvement that did not differ from fN (although it also did not differ from any of the other cluster classes). This suggests that the clusters that are easiest for speakers to produce are also easiest for individuals to improve on in a cluster learning paradigm. We hasten to note that this was a very short training paradigm (a single 20-min session) and that these patterns may not generalize to longer and more focused training paradigms.

While there were few group differences that emerged as significant, it is worth noting that performance improved from the beginning of practice to at least one retention session for all clusters except for /pt/. As we have noted in previous work (Cheng & Buchwald, 2021), the voiceless stop-stop sequences present a particular difficulty for using acoustic evidence to detect the presence of a vowel, as these sequences are often produced with a voiceless vowel in typical and fast speech (e.g., potato). The absence of a clear improvement in producing these sequences may therefore reflect this difficulty, as the acoustic record does not easily distinguish between the presence versus absence of a voiceless vowel. We note that an approach that relied on perception would have similar difficulty with these clusters and would also complicate the analysis of the other seven clusters in this study.

An additional noteworthy result is that several patterns did not emerge until the 2-day retention time point. In particular, participants showed numerical improvement in their production of both vN and zD clusters at R1, but did not exhibit a significant improvement until R2. Similarly, the differences in improvement across cluster classes did not reach significance until that point. It has been suggested that sleeping after learning promotes consolidation of learned material (Earle & Myers, 2015; Rasch & Born, 2013), which could be part of the explanation for these findings. Interestingly, some of the results in the acoustic data reflecting changes in articulatory coordination emerged at R1, suggesting that perhaps the more phonological aspects of learning represented by cluster accuracy are more sensitive to the consolidation period than the more fine-tuning aspects of motor control represented by the nasal duration analysis.

These patterns highlight the importance of considering the identity of the consonants within the clusters in an explanatory account of cluster learning, and further suggests that the linguistics and psycholinguistics literatures have already identified some of the key aspects of consonants that needs to be considered. As discussed in the introduction, the GODIVA framework (Bohland et al., 2010; Guenther, 2016) has already been used to explain aspects of consonant sequence learning (Segawa et al., 2015, 2019); by adding aspects of linguistic and phonological structure to the speech sound maps, it should be readily adaptable to account for the findings discussed above.

Changes in Cluster Coordination

We also found clear evidence of cluster production improvement in aspects of the acoustics that relate to the production of the cluster itself. We observed these changes in two ways. First, we observed a decrease in burst-to-burst duration in both /gd/ and /pt/ tokens from the beginning of practice to at least one of the retention time points. While others have examined whole nonword duration in similar studies (Segawa et al., 2015, 2019), we examined the timing within the cluster itself to ensure that differences in duration came from the aspect of the nonword that was the focus of the learning study. This decrease in duration reflects a more compact coordination between the two consonants in the cluster. Moreover, it suggests that there is some change due to the learning paradigm that is not reflected in accuracy changes, perhaps because the required coordination for the stop-stop sequences is too difficult to learn to produce in the context of such a short learning study. We note that there were bigger learning effects in our previous work that focused on stop-stop clusters (Cheng & Buchwald, 2021) but that study had participants producing stop-stop sequences for the entire practice session whereas they only comprised 25% of the stimuli within this study. In addition, we also observed an effect of training in this study, where the nonwords beginning with /gd/ that were trained showed a shorter burst-to-burst duration than those that were untrained. This same analysis had yielded no difference in Cheng and Buchwald's (2021) article, suggesting again that the additional practice of stop-stop clusters also allowed the changes in consonant coordination to generalize to other words containing the same clusters. This further suggests that increasing the amount of practice may impact the changes within the learning task. We note that the finding that /gd/ is produced with a tighter coordination in trained nonwords aligns with the finding reported in the work of Segawa et al. (2019) that the overall duration of trained nonwords was shorter than untrained nonwords, even when those untrained nonwords were composed of sub-lexical sequences that were trained. However, the overall finding that there are improvements in coordination in the absence of large accuracy gains suggests that the learning process is perhaps best considered as a gradual process rather than a binary measure of accuracy.

This interpretation is further supported by the changes in nasal duration in the accurately produced fN sequences. Based on the finding that the duration of nasal consonants in sN clusters (/sn/ and /sm/) is shorter than singletons (Buchwald & Miozzo, 2012; Klatt, 1975), we examined whether the learning paradigm would lead participants to produce shorter nasals in the fN clusters and found significant decreases in both /fn/ and /fm/ clusters at both retention time points. These analyses examined only correctly produced items, and therefore the decrease in this critical aspect of coordination in producing the cluster occurred in the clusters that were already produced without an intrusive vowel. This suggests that the accurate production is not necessarily the endpoint of the learning process, as individuals continue to change their articulatory coordination of clusters even after that has been achieved. With respect to frameworks such as GODIVA that contend that there is a working memory “chunk” that is learned in these tasks, this suggests that the formation of the chunk may not be the endpoint of learning.

Limitations and Clinical Implications

This work is part of a larger goal to understand motor learning in the domain of speech production. As has been previous discussed (Maas et al., 2008), there are many similarities between speech and other domains but there are also many differences. One key difference is the interaction between the speech and language processing, and we believe that examining nonnative consonant cluster learning is a possible means of studying that interaction given that these sequences are novel and difficult for both phonological and motoric reasons. Our approach here was most clearly limited by the relatively short practice session given the complexity of the task. While it can be more difficult to implement, it is clear that including multiple training sessions would potentially lead to greater improvement and provide more space to determine whether there are differences in a more protracted learning study. In addition, there may be aspects of the production that we are unable to observe in the absence of articulatory data (Marin & Pouplier, 2010; Pouplier et al., 2020) that was not collected as part of this larger study but is being examined as part of a separate ongoing project (Cheng et al., 2021).

The implications of this work are most clear with respect to the application of speech motor learning paradigms to individuals with impairment, such as apraxia of speech (Austermann Hula et al., 2008; Buchwald et al., 2020; Knock et al., 2000; Maas et al., 2008). It should be noted that treatment work typically involves a large amount of practice spread out among a large number of treatment sessions; thus, the limitation of a single session of practice extends to our ability to extrapolate from these findings to research on the apraxia of speech treatment literature. Nevertheless, much of that work has focused on specific articulatory details of consonants or consonant sequences that are being learned, and also highlight the importance of considering cluster learning as more multidimensional than a traditional sequences learning approach would do.

Conclusions

This article presented a large-scale study of nonnative onset consonant cluster learning designed to understand how aspects of phonological structure affect learning and the extent to which acoustic analyses of productions can reflect aspects of learning that are not evident in traditional accuracy analyses. We found clear evidence that consonant cluster learning is not simply sequence learning as is examined in the limb-based motor learning, and that the details of the sequences have clear impact on difficulty of production, with some evidence that these differences impact ease of learning within a short task. We also found evidence of learning changes within the acoustics of cluster production both for clusters that are difficult to produce accurately and show small or no accuracy improvement as well as for clusters that are already produced accurately. Taken together, these findings enhance our understanding of how individuals learn to produce nonnative consonant clusters that is both of practical importance for second language acquisition as well as of theoretical importance for understanding the interaction of speech and language processing systems.

Author Contributions

Adam Buchwald: Conceptualization (Lead), Data curation (Equal), Formal analysis (Supporting), Funding acquisition (Lead), Project administration (Lead), Visualization (Supporting), Writing – original draft (Lead), Writing – review & editing (Lead). Hung-Shao Cheng: Data curation (Equal), Formal analysis: (Lead), Project administration (Supporting), Validation (Supporting), Writing – original draft (Supporting), Writing – review & editing: (Supporting).

Data Availability Statement

All coded data as well as reproducible scripts for statistical analyses and data visualization can be found on the OSF repository at https://osf.io/b6gvk/.

Supplementary Material

Supplemental Material S1. Overall accuracy for each consonant cluster class at each time point, including performance across acquisition (practice session). Lines indicating baseline performance are provided to help visualize the change. Error bars are standard error.
Supplemental Material S2. Overall burst-to-burst duration in /gd/ and /pt/ clusters at each time point, including performance across acquisition (practice session). Error bars represent standard error.
Supplemental Material S3. Overall nasal duration in correctly produced fN clusters at each time point, including performance across acquisition (practice session). Error bars represent standard error.

Acknowledgments

This work was funded by two grants from the National Institute on Deafness and Other Communication Disorders to the first author (A.B.: K01DC014298 and R01DC018589). The authors also acknowledge the help of Chiara Repetti-Ludlow in data organization.

Funding Statement

This work was funded by two grants from the National Institute on Deafness and Other Communication Disorders to the first author (A.B.: K01DC014298 and R01DC018589).

References

  1. Austermann Hula, S. N. , Robin, D. A. , Maas, E. , Ballard, K. J. , & Schmidt, R. A. (2008). Effects of feedback frequency and timing on acquisition, retention, and transfer of speech skills in acquired apraxia of speech. Journal of Speech, Language, and Hearing Research, 51(5), 1088–1113. 10.1044/1092-4388(2008/06-0042) [DOI] [PubMed] [Google Scholar]
  2. Berent, I. , & Lennertz, T. (2007). What we know about what we have never heard before: Beyond phonetics: Reply to Peperkamp. Cognition, 104(3), 638–643. 10.1016/j.cognition.2007.01.006 [DOI] [PubMed] [Google Scholar]
  3. Blevins, J. (1995). The syllable in phonological theory. In Goldsmith J. (Ed.), Handbook of phonological theory (Chapter 6). Blackwell. [Google Scholar]
  4. Boersma, P. , & Weenink, D. (2022). Praat: Doing phonetics by computer (version 6.2.14) [Computer program] . http://www.praat.org/
  5. Bohland, J. W. , Bullock, D. , & Guenther, F. H. (2010). Neural representations and mechanisms for the performance of simple speech sequences. Journal of Cognitive Neuroscience, 22(7), 1504–1529. 10.1162/jocn.2009.21306 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Broselow, E. , & Finer, D. (1991). Parameter setting in second language phonology and syntax. Second Language Research, 7(1), 35–59. 10.1177/026765839100700102 [DOI] [Google Scholar]
  7. Browman, C. P. , & Goldstein, L. (1988). Some notes on syllable structure in articulatory phonology. Phonetica, 45(2–4), 140–155. 10.1159/000261823 [DOI] [PubMed] [Google Scholar]
  8. Buchwald, A. , Calhoun, H. , Rimikis, S. , Lowe, M. S. , Wellner, R. , & Edwards, D. J. (2019). Using tDCS to facilitate motor learning in speech production: The role of timing. Cortex, 111, 274–285. 10.1016/j.cortex.2018.11.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Buchwald, A. , Khosa, N. , Rimikis, S. , & Duncan, E. (2020). Behavioral and neurological effects of tDCS on speech motor recovery: A single-subject intervention study. Brain and Language, 210, Article 104849. 10.1016/j.bandl.2020.104849 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Buchwald, A. , & Miozzo, M. (2012). Phonological and motor errors in individuals with acquired sound production impairment. Journal of Speech, Language, and Hearing Research, 55(5), S1573–S1586. 10.1044/1092-4388(2012/11-0200) [DOI] [PubMed] [Google Scholar]
  11. Byrd, D. (1995). C-centers revisited. Phonetica, 52(4), 285–306. 10.1159/000262183 [DOI] [Google Scholar]
  12. Byrd, D. (1996). Influences on articulatory timing in consonant sequences. Journal of Phonetics, 24(2), 209–244. 10.1006/jpho.1996.0012 [DOI] [Google Scholar]
  13. Cheng, H.-S. , & Buchwald, A. (2021). Does voicing affect patterns of transfer in nonnative cluster learning? Journal of Speech, Language, and Hearing Research, 64(6S), 2103–2120. 10.1044/2021_JSLHR-20-00240 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Cheng, H.-S. , Masapollo, M. , Hagedorn, C. , & Buchwald, A. (2021). Effects of phonotactic legality on gestural coordination in consonant clusters: An electromagnetic articulography study. The Journal of the Acoustical Society of America, 150(4), A189–A189. 10.1121/10.0008078 [DOI] [Google Scholar]
  15. Clements, G. N. (1990). The role of the sonority cycle in core syllabification. In Beckman M. & Kingston J. (Eds.), Papers in laboratory phonology 1. Cambridge University Press. 10.1017/CBO9780511627736.017 [DOI] [Google Scholar]
  16. Davidson, L. (2006). Phonotactics and articulatory coordination interact in phonology: Evidence from non-native production. Cognitive Science, 30(5), 837–862. 10.1207/s15516709cog0000_73 [DOI] [PubMed] [Google Scholar]
  17. Davidson, L. (2007). The relationship between the perception of non-native phonotactics and loanword adaptation. Phonology, 24(2), 261–286. 10.1017/S0952675707001200 [DOI] [Google Scholar]
  18. Davidson, L. (2010). Phonetic bases of similarities in cross-language production: Evidence from English and Catalan. Journal of Phonetics, 38(2), 272–288. 10.1016/j.wocn.2010.01.001 [DOI] [Google Scholar]
  19. Davidson, L. , & Shaw, J. A. (2012). Sources of illusion in consonant cluster perception. Journal of Phonetics, 40(2), 234–248. 10.1016/j.wocn.2011.11.005 [DOI] [Google Scholar]
  20. Dupoux, E. , Kakehi, K. , Hirose, Y. , Pallier, C. , & Mehler, J. (1999). Epenthetic vowels in Japanese: A perceptual illusion? Journal of Experimental Psychology: Human Perception and Performance, 25(6), 1568–1578. 10.1037/0096-1523.25.6.1568 [DOI] [Google Scholar]
  21. Dupoux, E. , Parlato, E. , Frota, S. , Hirose, Y. , & Peperkamp, S. (2011). Where do illusory vowels come from? Journal of Memory and Language, 64(3), 199–210. 10.1016/j.jml.2010.12.004 [DOI] [Google Scholar]
  22. Earle, F. S. , & Myers, E. B. (2015). Overnight consolidation promotes generalization across talkers in the identification of nonnative speech sounds. The Journal of the Acoustical Society of America, 137(1), El91–El97. 10.1121/1.4903918 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Greenberg, J. H. (1978). Typology and cross-linguistic generalizations. In Greenberg J. H. (Ed.), Universals of human language (Vol. I, pp. 33–59). Stanford University Press. [Google Scholar]
  24. Guenther, F. H. (2016). Neural control of speech. MIT Press. 10.7551/mitpress/10471.001.0001 [DOI] [Google Scholar]
  25. Guevara-Rukoz, A. , Lin, I. , Morii, M. , Minagawa, Y. , Dupoux, E. , & Peperkamp, S. (2017). Which epenthetic vowel? Phonetic categories versus acoustic detail in perceptual vowel epenthesis. The Journal of the Acoustical Society of America, 142(2), EL211–EL217. 10.1121/1.4998138 [DOI] [PubMed] [Google Scholar]
  26. Harel, D. , & McAllister, T. (2019). Multilevel models for communication sciences and disorders. Journal of Speech, Language, and Hearing Research, 62(4), 783–801. 10.1044/2018_JSLHR-S-18-0075 [DOI] [PubMed] [Google Scholar]
  27. Kahn, D. (1976). Syllable-based generalizations in English Phonology (Ph.D. dissertation) . Massachusetts Institute of Technology. [Google Scholar]
  28. Klatt, D. H. (1975). Voice onset time, frication, and aspiration in word-initial consonant clusters. Journal of Speech and Hearing Research, 18(4), 686–706. 10.1044/jshr.1804.686 [DOI] [PubMed] [Google Scholar]
  29. Knock, T. R. , Ballard, K. J. , Robin, D. A. , & Schmidt, R. A. (2000). Influence of order of stimulus presentation on speech motor learning: A principled approach to treatment for apraxia of speech. Aphasiology, 14(5–6), 653–668. 10.1080/026870300401379 [DOI] [Google Scholar]
  30. Kreitman, R. (2010). Mixed voicing word-initial onset clusters. In Cécile F., Barbara K., Mariapaola D. I., & Nathalie V. (Eds.), Laboratory phonology (Vol. 10, pp. 169–200). [Google Scholar]
  31. Levelt, W. J. M. , Roelofs, A. , & Meyer, A. S. (1999). A theory of lexical access in speech production. Behavioral and Brain Sciences, 22(1), 1–75. 10.1017/S0140525X99001776 [DOI] [PubMed] [Google Scholar]
  32. Leys, C. , Ley, C. , Klein, O. , Bernard, P. , & Licata, L. (2013). Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median. Journal of Experimental Social Psychology, 49(4), 764–766. 10.1016/j.jesp.2013.03.013 [DOI] [Google Scholar]
  33. Lindblom, B. (1990). Explaining phonetic variation: A sketch of the H&H Theory. In Hardcastle W. J. & Marchal A. (Eds.), Speech production and speech modeling (pp. 403–439). Kluwer. 10.1007/978-94-009-2037-8_16 [DOI] [Google Scholar]
  34. Lombardi, L. (1999). Positional faithfulness and voicing assimilation in optimality theory. Natural Language and Linguistic Theory, 17(2), 267–302. 10.1023/A:1006182130229 [DOI] [Google Scholar]
  35. Lowe, M. S. , & Buchwald, A. (2017). The impact of feedback frequency on performance in a novel speech motor learning task. Journal of Speech, Language, and Hearing Research, 60(6S), 1712–1725. 10.1044/2017_JSLHR-S-16-0207 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Maas, E. , Robin, D. A. , Hula, S. N. A. , Freedman, S. E. , Wulf, G. , Ballard, K. J. , & Schmidt, R. A. (2008). Principles of motor learning in treatment of motor speech disorders. American Journal of Speech-Language Pathology, 17(3), 277–298. 10.1044/1058-0360(2008/025) [DOI] [PubMed] [Google Scholar]
  37. Marin, S. , & Pouplier, M. (2010). Temporal organization of complex onsets and codas in American English: Testing the predictions of a gestural coupling model. Motor Control, 14(3), 380–407. 10.1123/mcj.14.3.380 [DOI] [PubMed] [Google Scholar]
  38. Masapollo, M. , Segawa, J. A. , Beal, D. S. , Tourville, J. A. , Nieto-Castañón, A. , Heyne, M. , Frankford, S. A. , & Guenther, F. H. (2021). Behavioral and neural correlates of speech motor sequence learning in stuttering and neurotypical speakers: An fMRI investigation. Neurobiology of Language, 2(1), 106–137. 10.1162/nol_a_00027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Parrell, B. , Lammert, A. C. , Ciccarelli, G. , & Quatieri, T. F. (2019). Current models of speech motor control: A control-theoretic overview of architectures and properties. The Journal of the Acoustical Society of America, 145(3), 1456–1481. 10.1121/1.5092807 [DOI] [PubMed] [Google Scholar]
  40. Pouplier, M. , Lentz, T. O. , Chitoran, I. , & Hoole, P. (2020). The imitation of coarticulatory timing patterns in consonant clusters for phonotactically familiar and unfamiliar sequences. Journal of Laboratory Phonology, 11(1). 10.5334/labphon.195 [DOI] [Google Scholar]
  41. Psychology Software Tools Inc. (2016). Eprime 3.0. http://www.pstnet.com
  42. R Core Team. (2017). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/ [Google Scholar]
  43. Rasch, B. , & Born, J. (2013). About sleep's role in memory. Physiological Reviews, 93(2), 681–766. 10.1152/physrev.00032.2012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Repp, B. H. , & Lin, H.-B. (1989). Acoustic properties and perception of stop consonant release transients. Journal of the Acoustical Society of America, 85(1), 379–396. 10.1121/1.397689 [DOI] [PubMed] [Google Scholar]
  45. Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 6(2), 461–464. 10.1214/aos/1176344136 [DOI] [Google Scholar]
  46. Segawa, J. , Masapollo, M. , Tong, M. , Smith, D. J. , & Guenther, F. H. (2019). Chunking of phonological units in speech sequencing. Brain and Language, 195, 104636–104636. 10.1016/j.bandl.2019.05.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Segawa, J. A. , Tourville, J. A. , Beal, D. S. , & Guenther, F. H. (2015). The neural correlates of speech motor sequence learning. Journal of Cognitive Neuroscience, 27(4), 819–831. 10.1162/jocn_a_00737 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Wilson, C. , Davidson, L. , & Martin, S. (2014). Effects of acoustic–phonetic detail on cross-language speech production. Journal of Memory and Language, 77, 1–24. 10.1016/j.jml.2014.08.001 [DOI] [Google Scholar]
  49. Wisler, A. , Goffman, L. , Zhang, L. , & Wang, J. (2022). Influences of methodological decisions on assessing the spatiotemporal stability of speech movement sequences. Journal of Speech, Language, and Hearing Research, 65(2), 538–554. 10.1044/2021_JSLHR-21-00298 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material S1. Overall accuracy for each consonant cluster class at each time point, including performance across acquisition (practice session). Lines indicating baseline performance are provided to help visualize the change. Error bars are standard error.
Supplemental Material S2. Overall burst-to-burst duration in /gd/ and /pt/ clusters at each time point, including performance across acquisition (practice session). Error bars represent standard error.
Supplemental Material S3. Overall nasal duration in correctly produced fN clusters at each time point, including performance across acquisition (practice session). Error bars represent standard error.

Data Availability Statement

All coded data as well as reproducible scripts for statistical analyses and data visualization can be found on the OSF repository at https://osf.io/b6gvk/.


Articles from Journal of Speech, Language, and Hearing Research : JSLHR are provided here courtesy of American Speech-Language-Hearing Association

RESOURCES