Abstract
Purpose:
According to theoretical models, perceptual learning is modulated by the level of acoustic predictably in the speech signal. This study investigated whether the variable speech of children is learnable in a perceptual learning paradigm and examined whether speech variability, as quantified by the acoustic spatiotemporal index (STI), predicted perceptual learning outcomes.
Method:
Speech samples were elicited from 40 typically developing children talkers (ages 3–8 years) and presented to 410 adult listeners in a structured perceptual learning experiment. Listeners transcribed phrases before (pretest) and after (posttest) a lexically guided familiarization phase. Intelligibility improvements following familiarization were analyzed using paired t tests, and regression analyses examined whether acoustic STI values predicted perceptual learning outcomes.
Results:
Listeners demonstrated significant perceptual learning of children's speech, with an average intelligibility gain of 6.4% following familiarization (p < .001). Regression analyses revealed a significant interaction between STI and pretest intelligibility: Lower STI values (indicating more predictable speech) were associated with greater intelligibility gains but only when pretest intelligibility was sufficiently low.
Conclusions:
This study provides the first empirical evidence that children's speech is learnable in a structured perceptual learning paradigm. Although the acoustic STI affords some insight into predictability, it does not fully account for variability in perceptual learning. Future work should explore measures that capture segmental predictability to refine models of speech learnability.
Speech perception is a seemingly effortless process by which acoustic signals are processed and mapped onto stored linguistic representations in the mind with very little demand on cognitive processing. This process of cue-to-category mapping can be viewed as statistical pattern matching, assuming that the linguistic (phonetic) categories onto which the speech signal is being mapped are relatively stable and encompass the natural variability observed in the speech signal (Eisner & McQueen, 2005; Kleinschmidt & Jaeger, 2015; Norris et al., 2003). When deviations in the speech signal exceed what the linguistic categories allow for, speech perception becomes more challenging. Such is the case with speech that is foreign accented, disordered, or noise vocoded, for example. These types of noncanonical speech make speech perception much more cognitively demanding because they require plasticity of the phonetic categories and result in reduced speech intelligibility (e.g., Heffner & Myers, 2021; Lansford et al., 2023). Despite the initial difficulty in perceiving noncanonical speech, with experience, it is possible to adapt to the novel variation and accurately perceive this speech (e.g., Borrie & Lansford, 2021; Bradlow & Bent, 2008; Davis et al., 2005; Sidaras et al., 2009). This learning to adapt phonetic categories to accommodate noncanonical speech is a phenomenon known as perceptual learning.
Perceptual learning has been studied extensively with the aforementioned types of noncanonical speech: Listeners are able to adapt to foreign-accented speech (e.g., Bradlow & Bent, 2008; Sidaras et al., 2009), disordered speech (e.g., Borrie et al., 2017; Lansford et al., 2023), and noise-vocoded speech (e.g., Davis et al., 2005; Hervais-Adelman et al., 2008), among others, as evidenced by improved intelligibility gains after a training or familiarization period. For example, a recent study showed that adult listeners were initially able to understand a dysarthric speaker with approximately 66% intelligibility, but after lexically guided listener training, the intelligibility increased to around 78%, showing the effectiveness of listener training (Borrie et al., 2023).
Perceptual learning is theorized to be modulated by the predictability of the speech signal, as listeners can use statistical regularities in the acoustic signal to reshape their phonetic categories for accurate perception of noncanonical speech (Kleinschmidt & Jaeger, 2015, 2016). In other words, each speaker's specific way of talking can be formalized as the distribution of acoustic cues that they use to produce an exemplar for each phonetic category, and listeners are able to infer these novel cue-to-category mappings if the cues are sufficiently consistent, predictable, and statistically regular. This hypothesis that predictability facilitates perceptual learning is supported by empirical evidence that hyperkinetic dysarthria, a dysarthria type characterized by highly unpredictable speech distortions, has not shown the same perceptual learning intelligibility gains as other types of dysarthria that result in more predictable degradations of the speech signal (Borrie et al., 2018; Lansford et al., 2019, 2020). Thus, with both theoretical claims and empirical evidence from dysarthria, it has been advanced that noncanonical speech that is more unpredictable in its deviations will be less learnable.
Children's speech, even of typically developing children, is characterized by acoustic features that differ significantly from adult speech (Huber et al., 1999; Lee et al., 1999). Compared to adults, children's productions are more variable in both temporal (e.g., segment duration, prosodic timing) and spectral (e.g., pitch, formant frequency) dimensions (Chermak & Schneiderman, 1986; DiSimoni, 1974a, 1974b; Lee et al., 1999; Tingley & Allen, 1975). These differences are largely attributable to immature motor control and incomplete phonological acquisition, which affect articulatory precision, timing, and coarticulation (Kent & Forner, 1980; Lee et al., 1999). Importantly, phoneme acquisition follows a predictable developmental sequence. For example, sounds like /b/ and /p/ are typically acquired early (around 2–3 years of age), whereas sounds like /l/ and /r/ are acquired later (around 5–7 years of age; Crowe & McLeod, 2018). Thus, variability in children's speech reflects systematic developmental processes rather than truly random deviations from adult targets. However, from the perspective of naive adult listeners, especially those unfamiliar with child speech, this variability may appear unpredictable or inconsistent, potentially posing challenges for perception.
In fact, it is the case that this variability in children's speech leads to reduced intelligibility of children's versus adults' speech (Sosa, 2015). Children's speech actually shows similar intelligibility deficits to individuals with dysarthria (Hustad et al., 2021): 4-year-old children in the 50th percentile of speech intelligibility have only 78.6% intelligibility in multiword utterances and even lower (74%) intelligibility for isolated words. This increases to 96% and 90.2%, respectively, for single- and multiword utterances by 7 years of age, but there is a great amount of variation between children (Hustad et al., 2021). Despite this reduced intelligibility of children's speech, parents seemingly adapt to their child's speech effortlessly. However, parents themselves have rated their child's intelligibility as high for themselves but significantly lower for strangers (McLeod et al., 2015). In other words, parents believe that their child is more intelligible to them than the child is to a naive (i.e., unfamiliar) listener/communication partner. Nevertheless, the fact that parents are able to understand their child's speech suggests that the highly variable speech of children is learnable, although this has never been directly tested. Furthermore, there is a large degree of interchild variability in the speech of children, and thus, it is reasonable to believe that the degradations of children's speech may vary in terms of acoustic predictability. As such, children's speech forms a good test case to measure speech signal predictability and its effects on perceptual learning. Thus, the primary goal of this study was to examine if children's speech, which is typically characterized by high amounts of variability, is learnable in an experimental perceptual learning paradigm.
In past work, the notion that less predictable speech is less learnable has been based on clinical perceptual judgments of whether the speech signal was predictable (Borrie et al., 2018; Lansford et al., 2019, 2020). Although perceptual judgments of speech are the gold standard for clinical evaluations of motor speech disorder, the notion of acoustic predictability, just like any other perceptual feature, can be subjective (Stipancic et al., 2023). As such, it has previously been posited that objective measurement procedures that accurately quantify predictability in the speech signal would be a valuable addition to the perceptual learning literature (Borrie & Lansford, 2021). Clinically, an objective measure of speech signal predictability would support speaker candidacy decisions (i.e., which speakers would be learnable) for a perceptual learning treatment approach.
There are currently no objective quantitative measures that directly capture predictability of speech. Nevertheless, we can hypothesize that the learnability of speech depends on some degree of consistency in the deviations from typical speech patterns. Thus, for example, if a speaker makes a consistent shift in the cue parameter (e.g., F1) for a given speech sound (say, /a/), it is relatively straightforward for a listener to adapt to this change. However, if the cue parameter is subject to a less predictable shift, where it may be either increased or decreased each time the phoneme is produced, adaptation now becomes increasingly more difficult, if not impossible. Within the existing literature, the measures best suited to quantifying this consistency are measures that quantify the variability in speech patterns across multiple repetitions of the same token. Several such measures have been previously explored within the existing literature. One common approach is to measure perceptual variability by calculating variation in word- or phoneme-level transcriptions of the same token (Goffman et al., 2007; Macrae et al., 2014; McNeill et al., 2022; Preston & Koenig, 2011). Alternatively, some studies have looked at variability in acoustic measures (such as formant, word durations, or voice onset times) in place of transcription-based measures (Koenig, 2001; Preston & Koenig, 2011; Whiteside et al., 2003). Whereas the previous categories have been primarily employed to measure speech within short tokens, words, or phonemes, measures of suprasegmental variability across phrase-level stimuli have also been considered. Historically, these measures have primarily focused on the analysis of kinematic signals, with the best known being the spatiotemporal index (STI; A. Smith et al., 1995). As the STI across the duration of the signal is sensitive to both prosodic and articulatory deviations occurring at any point during the production, it has been most commonly used to measure the movement of speech articulations by attaching sensors to the lips, jaw, and tongue. Prior research has shown inflated STI values relative to healthy controls in a range of conditions including: Parkinson's disease (A. Anderson et al., 2008), developmental language disorder (Brumbach & Goffman, 2014; Goffman, 1999; Saletta et al., 2018), and childhood apraxia of speech (Grigos et al., 2015; Moss & Grigos, 2012; Vuolo & Wisler, 2024; Wisler et al., 2025). More recently, an acoustic STI, which quantifies predictability in the speech by measuring the amplitude envelope, as in Howell et al. (2009), has also been developed. This acoustic STI has the benefit of not requiring kinematic sensor data and has shown moderate-to-high correlations with the kinematic STI in the speech of typically developing children (Benham et al., 2023).
This Study
The first goal of the present study was to evaluate if the speech of young children is learnable in a structured perceptual learning paradigm. The second goal of this study, contingent on at least some subset of the child speakers being learnable through this paradigm, was to determine whether speech variability measured by the acoustic STI could be used to inform the degree of learnability across children. We hypothesized that (a) children's speech would be learnable in a perceptual learning paradigm (i.e., listeners would be able to successfully adapt to the noncanonical acoustic signals of child speech, reflected in intelligibility improvements following training) and (b) speech with lower STI values (i.e., more predictable) would yield greater intelligibility improvements than speech with higher STI values (i.e., less predictable).
Method
Participants
Overview
This study involved collecting data from two groups of participants: We collected speech stimuli from child speakers, and we collected perceptual data from adult listeners. This study was approved by the institutional review board at Utah State University (Protocols 11380 and 14439), and all participants gave their informed consent prior to beginning study procedures.
Child Speakers
The stimuli for this study were collected from a cohort of 40 children between the ages of 3 and 8 years (M = 5.6, SD = 1.34). This cohort included one speaker who was 3 years old, six who were 4 years old, 17 who were 5 years old, five who were 6 years old, six who were 7 years old, and five who were 8 years old. Twenty-two (55%) of the children speakers were female. All children were monolingual native speakers of American English living in northern Utah or southern Minnesota, regions that are commonly considered to reflect relatively standard American English accents, with minimal regional markedness in phonology. The children had no parent-reported diagnosis of a speech, language, or hearing disorder. All children speakers were compensated with a gift card for their time.
Speech Stimuli
The speech stimuli used in this study consisted of audio-recorded productions of testing phrases and a familiarization/training passage produced by the 40 different children, described previously. The stimuli for the testing phrases included a list of 100 two-word phrases ranging from two to six syllables (see the Appendix). The phrases had low interword predictability so that the listener would be less likely to use higher level cognitive–linguistic information to deduce the words. The word list was developed to reflect items and concepts commonly encountered in early childhood (ages 3–8 years). Many of the nouns and adjectives were drawn from materials in prior developmental research (e.g., child-directed vocabulary studies, picture naming tasks). Although the specific source of each word varies, the overall list includes high-frequency concrete nouns (e.g., banana, cat, spoon) and adjectives (happy, cold, soft) that are well documented as being within the productive and receptive vocabularies of children in this age range (e.g., C. Anderson & Cohen, 2012; Bates et al., 2003). Although individual familiarity with the vocabulary may vary, the overall selection is appropriate for generalization across child participants. Following the children speech elicitation procedures of Lee et al. (2014) and Hustad et al. (2021), to elicit these audio recordings, the children were presented with an audio recording with an accompanying picture of each phrase one at a time and were asked to repeat the phrase while being recorded with a Zoom (Version H5) recorder in a quiet room.1 To obtain speech for the familiarization passage, the children repeated the story Hedgehugs by Steve Wilson and Lucy Tapper (2015). This story was selected for its age-appropriate content, engaging narrative, and phonetic richness. It consisted of 36 short phrases (four to 12 syllables each), segmented by natural prosodic boundaries. All English phonemes except /ʒ/ were represented (/ʒ/ was also absent in testing stimuli), supporting broad perceptual exposure during training. Children heard a recording of the story in short phrases and were asked to repeat what they heard, one phrase at a time, while also looking at the pictures of the story. For the measurement of STI, children were asked to repeat the phrases “buy Bobby a puppy,” which has been widely used in studies examining STI in children (Sadagopan & Smith, 2008; A. Smith et al., 1995; Walsh & Smith, 2002; among others), and “beeping broccoli,” 15 times each at their normal speech rate to ensure that at least 10 fluent repetitions were available for analysis; if entire syllables were deleted or added, then the stimuli were considered nonfluent and thus not included for the acoustic STI measure. Following the removal of nonfluent stimuli, each child had an average of 13.375 fluent repetitions (per stimuli), and no child had fewer than 10 fluent repetitions of either phrase. The audio recordings of the speaker that the children heard and were asked to repeat after were produced by a 29-year-old woman who is a native speaker of American English with no history of speech, language, or voice disorder.
Listener Data Collection
The experimental paradigm included three phases: pretest, lexically guided familiarization, and posttest, which have been used extensively in work examining perceptual learning of dysarthric speech (e.g., Borrie et al., 2017, 2023; Lansford et al., 2019). Fifty of the target phrases were used for the pretest baseline intelligibility and 50 for the posttest after the training phase, and all listeners received the same pretest and posttest stimuli sets. The pretest and posttest word lists were balanced for number of syllables, number of phonemes, and phonetic complexity as measured by the word complexity measure (Stoel-Gammon, 2010); this measure assigns numeric values based on phonological features such as consonant clusters, syllable structure, and the presence of later-acquired sounds (e.g., liquids), providing a standardized index of articulatory difficulty. A higher score reflects greater articulatory complexity. Balancing stimuli on this dimension ensured that any changes in listener performance could not be attributed to differences in word-level production between the test phases. For the familiarization/training phase, the contextual passage, the Hedgehugs story, was presented phrase by phrase with accompanying subtitles of the intended targets for the phrases.
The experiment was programmed in Gorilla, an online platform used to create and host behavioral experiments (https://gorilla.sc) and was administered through Prolific, an online site designed to crowdsource suitable participants for online questionnaires and experiments (https://prolific.com), using standard sampling. A total of 410 adult listeners partook in this study. There were 10 adult listeners for each child speaker, with the exception of one child whose speech had 20 listeners. We assessed the intelligibility results of using 10 versus 20 listeners for this child talker and observed no difference in mean percent words correct at pretest or posttest. As such, the decision was made to collect data from 10 listeners for the remaining child talkers.2 This is in line with recent work by Ziegler et al. (2021), which showed that nine crowdsourced listener participants are sufficient and reliable for such study designs. Each adult listener was exposed to only one child's speech (speaker-specific perceptual learning). Listeners were instructed to wear headphones and were able to adjust the volume of the audio prior to starting the experimental task. All participants were English-speaking monolinguals from and living in the United States. The mean age of listener participants was 35.93 years (SD = 11.12). Of those listeners, 237 were female (57.8%), 171 were male (41.7%), and two preferred not to say. In terms of ethnicity, 16 (7.6%) self-identified as Asian, 75 (35.7%) as Black or African American, 23 (11.0%) as mixed, 294 (71.7%) as White or Caucasian, and two preferred not to say. The listeners had no self-reported history of speech, language, or hearing disorders. Listener participants were excluded for poor task engagement, operationally defined as nonresponses for 20% or more of the testing stimuli in either phase (pretest or posttest); we excluded 36 participants for this reason, resulting in a final pool of 371 listeners in the final analysis. The experimental paradigm took an average of 30 min to complete, and listeners were compensated for their participation monetarily through the Prolific platform.
Transcript Analysis
The data from the perceptual learning paradigm consisted of orthographic transcriptions of what each listener thought the children were saying during the pretest and posttest phases. Transcriptions were scored for intelligibility, a measure of percent words correct, using Autoscore (Barrett et al., 2019; Borrie et al., 2019), an open-source application used for automatically scoring orthographic transcripts. Words were considered correct if they matched the intended target exactly or differed only by plurality; homophones and obvious typos or spelling errors were scored as correct using a list of common misspellings in the testing stimuli created by the authors. A percent words correct score was then calculated for both pretest and posttest for each listener, to evaluate the amount of learning following training.
Acoustic STI Analysis
Calculation of the acoustic STI begins with the complete set of fluent repetitions for each stimuli production (“beeping broccoli” or “buy Bobby a puppy”) as described in the Speech Stimuli section. The onsets and offsets for each repetition are then selected using both the waveform and spectrogram as visual cues to identify where the /b/s began and where the /i/s ended. From the parsed signals, the amplitude envelope was extracted using the same general procedure used in Benham et al. (2023) and Howell et al. (2009). This was done by first adding a 50-Hz high-pass filter before halfwave, rectifying the acoustic signal by setting all negative values equal to zero, and then passing the result through a fourth-order 15-Hz Butterworth low-pass filter. Once we had extracted the amplitude envelope, calculation of the acoustic STI used the same four-step procedure as for the kinematic STI. The first step was to amplitude normalize the signal using a z-scoring procedure, which mitigates broad differences in speech intensity and allows the STI to better capture relative differences in speech patterning. Then, each production was time aligned by resampling all signals to a fixed 1,000-point length. This process allowed for comparison of equal length signal vectors and mitigates the influence of differences in overall duration across repetitions. Note that the resampling procedure employs a spline interpolation method; however, because the initial length of the audio signal is typically much greater than the length being resampled to (due to the higher sampling rate), the interpolation method is far less critical for the acoustic STI than it is for the kinematic STI. Once all repetitions had been time and amplitude normalized, standard deviations were calculated across repetitions at each 2% time interval, and the final STI was calculated as the sum of these 50 standard deviations. The value resulting from this process should fall between 0 and 50, where a value of 0 indicates perfect consistency (no differences across repetitions) and a value of 50 indicates that the repetitions were entirely random with no underlying pattern shared across repetitions. As a final step in the calculation of the STI, we applied the bias-correction formula described in Wisler et al. (2022, 2024) to remove bias from inconsistencies in the numbers of repetitions included in the analysis.
Statistical Analysis
To address the first research question, whether children's speech is learnable through a structured perceptual learning paradigm, we conducted a paired t test comparing pretest and posttest intelligibility scores. To examine the second research question, which aimed to identify if a current measure that may reflect speech signal predictability (i.e., acoustic STI) influenced the success of this perceptual learning paradigm, we conducted separate regression models for each of the STI stimuli: “beeping broccoli” and “buy Bobby a puppy.” Note that the use of separate regression models is motivated by our expectation that STI values for the two stimuli will be highly correlated across and present collinearity issues for the regression. Each model uses intelligibility gain as the response variable and includes effects for pretest intelligibility and acoustic STI, as well as their interaction. Additional covariates for participant sex and age were also included in the models. It should be noted that sex was implemented using dummy coding (with female as the base level), and all other variables were implemented numerically without normalization. Assumptions and fit of the models were assessed for violations, including normality of residuals, collinearity, and influential outliers. To assess the interaction further, we used the Johnson–Neyman technique, useful for interactions of continuous variables, to determine the values of the moderator (i.e., pretest scores) where the relationship between STI and posttest intelligibility is significantly different from zero. In addition to the t test and regression models, we also visualized pairwise comparisons using the GGally package in R (Version 4.4.3; Emerson et al., 2013) for the seven variables of interest in our analysis: pretest intelligibility, posttest intelligibility, intelligibility gain, sex, age, acoustic STI (“beeping broccoli”), and acoustic STI (“buy Bobby a puppy”).
Results
Pretest and posttest intelligibility means for each child-specific perceptual learning program are shown in Figure 1. The paired t test comparing pretest and posttest intelligibility scores revealed a statistically significant increase in intelligibility in the posttest condition (t = 7.89, p < .001). The average increase in intelligibility following training was 6.40% (95% confidence interval [4.76%, 8.04%]) across participants, and individual intelligibility improvements for the different children speakers ranged from −2.7% to +20.8%. Improvement at posttest was observed on a consistent basis with only three of the 40 children speakers programs showing lower intelligibility at posttest than pretest conditions.
Figure 1.
Average intelligibility across listeners for each participant both before (pretest) and after (posttest) the perceptual learning protocol. Each line represents an individual pretest to posttest change, and boxplots depict the overall distribution of intelligibility values in each condition.
The results of the regression models are compiled in Table 1. Both models, “beeping broccoli” and “buy Bobby a puppy,” showed significant negative effects for STI (β = −2.789, p = .02; and β = −2.708, p = .01, respectively) with positive effects for the interaction between STI and pretest intelligibility (β = .035, p = .03; and β = .033, p = .015, respectively). To assess the interaction further, the Johnson–Neyman technique determined that, when pretest intelligibility is below 71.0, the relationship between STI (for “beeping broccoli”) and posttest intelligibility is significantly different from zero. For “buy Bobby a puppy,” when pretest intelligibility is below 73.1, the relationship between STI and posttest intelligibility is significantly different from zero. Figure 2 visualizes these interactions for each of the two interaction terms (STI and pretest intelligibility) for the minimum and maximum values of the interacting variable.
Table 1.
Regression results for intelligibility of two sentences as predicted by STI, Pretest intelligibility, their interaction, sex and age.
| Variable | Regression model |
|||||||
|---|---|---|---|---|---|---|---|---|
| “Beeping broccoli” |
“Buy Bobby a puppy” |
|||||||
| Estimate | SE | t value | p value | Estimate | SE | t value | p value | |
| (Intercept) | 113.499 | 38.490 | 2.949 | .006 | 103.242 | 30.749 | 3.358 | .002 |
| STI | −2.789 | 1.139 | −2.449 | .020 | −2.708 | 0.991 | −2.732 | .010 |
| IntPre | −1.442 | 0.557 | −2.589 | .014 | −1.277 | 0.423 | −3.019 | .005 |
| SexM | 0.412 | 1.563 | 0.264 | .794 | 0.142 | 1.486 | 0.096 | .924 |
| Age | 0.999 | 0.738 | 1.352 | .185 | 0.850 | 0.665 | 1.277 | .210 |
| STI:IntPre | 0.035 | 0.016 | 2.244 | .031 | 0.033 | 0.013 | 2.550 | .015 |
| Residual | 4.537 | 4.492 | ||||||
| R 2 | .317 | .331 | ||||||
| AIC | 241.999 | 241.205 | ||||||
| BIC | 253.822 | 253.027 | ||||||
Note. SE = standard error; STI = spatiotemporal index; AIC = Akaike information criterion; BIC = Bayesian information criterion.
Figure 2.
Plot illustrating the interaction between acoustic spatiotemporal index (STI) and pretest intelligibility in their effects on intelligibility gains for the “buy Bobby a puppy” stimuli. The x-axis represents pretest intelligibility, with separate lines indicating the anticipated intelligibility gain associated with different pretest intelligibilities when STI is at its minimum (14.62) and maximum (39.84). The vertical dashed line depicts the point at which the effects of acoustic STI are no longer statistically significant based on the Johnson–Neyman technique. IntPre = Pretest intelligibility; SexM = Sex - Male.
Figure 3 displays pairwise plots for each pair of variables in the analysis. Despite our expectation that acoustic STI values will be highly correlated across the two stimuli, we observed only a moderate correlation between acoustic STI values (r = .407). Additionally, the correlation analysis finds age to be correlated with intelligibility in both pretest and posttest conditions (r = .524 and r = .488, respectively) but not with the difference between the two.
Figure 3.
Pairwise variable comparison plots showing the relationships between sex, age, and percent words correct (PWC): pretest, posttest, and gained, and acoustic spatiotemporal index (STI) for both the “buy Bobby a puppy” (BBAP) and “beeping broccoli” (BB) stimuli. Corr = correlation; Pre = pretest; Post = posttest.
Discussion
Children's speech is highly variable in its baseline intelligibility; the average baseline intelligibility at pretest of the children in our study, aged 3–8 years, was 74.2%, with a range spanning from 41.7% to 94.8%. These values fall in line with that past research on children's speech intelligibility and its variability (Hustad et al., 2021; Sosa, 2015). Despite this variability between children, the novel contribution of this study is the finding that children's speech is highly learnable. That is, following a brief structured familiarization experience with the child's speech, naive adult listeners perceptually adapted to the speech signal, reflected in significant better understanding of that particular child's speech in subsequent encounters (i.e., posttest). This finding is robust, as we examined learnability of 40 child talkers. With structured, speaker-specific familiarization, intelligibility significantly improved by an average of 6.4%; however, intelligibility improvements of up to 20.8% were found for some children speakers. This finding supported our primary hypothesis that children's speech would be learnable.
To contextualize the observed intelligibility gains, we compare our results to perceptual learning studies with adult speakers who have disordered speech. Previous work has shown that intelligibility gains in dysarthria range from approximately 5% to 23% following structured familiarization (see Borrie & Lansford, 2021, for a review). Our observed average gain of 6.4% with improvements of up to over 20% for some child speakers matches this range. Furthermore, work by Stipancic et al. (2025) has proposed that intelligibility gains of 6% to 9% may constitute a clinically meaningful improvement in understanding individuals with motor speech disorders. By this criterion, the average improvement observed in the present study with children's speech is indeed clinically meaningful, and for a subset of children, the learning gains substantially exceeded this threshold. This reinforces the conclusion that despite the unique variability of children's speech, it is learnable, comparably to other forms of noncanonical speech.
Given the learning observed, we can assume that these children's speech signals contained sufficient acoustic regularities in the segmental and suprasegmental information to be learnable. Even though baseline intelligibility was well below ceiling in the vast majority of the children speakers, with training, the listeners were able to adapt their phonetic categories to accommodate the novel cues of that specific child talker. Similar results have been seen with other types of noncanonical speech including time-compressed (Dupoux & Green, 1997; Golomb et al., 2007), accented (Sidaras et al., 2009), and disordered (Borrie et al., 2023; Borrie & Lansford, 2021) speech. Our data suggest that children's speech is not unique in its learnability but behaves similarly to many other types of noncanonical speech in that, with experience, listeners are able to adapt to it. With children specifically, this may explain how often times a child's parents can understand what they are saying, whereas a stranger cannot: The parent has already adapted to their child's unique phonetic cue distribution and can effortlessly map those noncanonical cues to their stored phonetic categories for accurate speech perception.
In order to attempt to quantify this predictability that is likely responsible for perceptual learning, we turned to the acoustic STI. Given that the acoustic STI was designed to quantify predictability in the speech signal, we hypothesized that a lower acoustic STI (i.e., a more predictable speech signal) would correspond to greater intelligibility improvements following familiarization. The statistical analysis supports this idea, as acoustic STI predicted intelligibility improvements. However, the relationship between the acoustic STI and perceptual learning appears to be slightly more complicated. Looking at Models 3a/b and 4a/b together, the relationship between intelligibility gains and the acoustic STI is only observable when the interaction between pretest intelligibility and STI is accounted for. This presents the idea that perceptual learning requires both regularities and adequate room for learning (indicated by pretest intelligibility being significantly below 100%). This second effect is not at all surprising in the extreme—if a participant achieves 100% intelligibility in the pretest, then there is no room for improvement with experience. However, its significance in affecting intelligibility improvements within a data set in which all but one participant had < 90% pretest intelligibility and in mediating the effect of the acoustic STI was not anticipated. The interaction plots presented in Figure 2 tell a relatively clear story that the anticipated negative relationship between intelligibility improvements and the acoustic STI is strong for participants with low pretest intelligibility (< 50%) but trends toward zero as pretest intelligibility nears 100%. In other words, when pretest intelligibility was low, a lower acoustic STI (indicating greater consistency in the speech) was associated with greater perceptual learning outcomes; however, as pretest intelligibility increased, the predictive effect of the acoustic STI on intelligibility improvements diminished, meaning that, for speech that was already relatively intelligible, predictability (as measured by the acoustic STI) did not significantly influence perceptual learning outcomes. The results of the Johnson–Neyman analysis revealed a cutoff point of ~72% (pretest intelligibility), above which the effect of the STI on intelligibility gains was no longer measurable as statistically significant. Note that this does not necessarily imply that no effect for the acoustic STI exists above this threshold; it is just that none could be observed from these data (at the p < .05 significance threshold).
Reviewing the summary statistics for the regression models, it is important not to overstate how strong of a relationship we observed between STI and intelligibility gains. Although pretest intelligibility and the acoustic STI were found to be significant predictors of intelligibility improvements in both models, these models both explained less than one third of the variance in intelligibility improvements across the listener participants. Thus, although the results support the ideas that acoustic regularity is a necessary factor for perceptual learning and that it can be quantified (to some degree) by the acoustic STI, there remains significant variation in perceptual learning outcomes that is not explained by this combination of measures. This leaves some uncertainty on the degree to which the acoustic STI offers sufficient predictive power of perceptual learning outcomes to help guide the direction of treatment in a clinical context. It is also unclear whether the unexplained variance can be attributed to limitations in speech regularity as a component of perceptual learning or limitations in the acoustic STI in its ability to quantify said speech regularity.
Considering the other independent variables, neither age nor sex was predictive of intelligibility improvements in this population. The lack of effect for age was somewhat surprising, as we expected greater acoustic regularity to coincide with increased age in this population. It is also worth noting that age was not significantly correlated with either measures of the acoustic STI. Whereas this is inconsistent with prior literature on the kinematic STI (Holm et al., 2010; A. Smith & Goffman, 1998; B. L. Smith, 1992), it is consistent with recent studies on the acoustic STI, which have also failed to show age-based differences in acoustic variability (Benham et al., 2023; Wang & Grigos, 2024).
Returning to our discussion of the acoustic STI, there are several potential reasons for why this measure may have not fully predicted perceptual learning outcomes in this study. As noted previously, limitations in the capabilities of the regression models to explain intelligibility improvements could simply reflect the existence of other characteristics of speech (beyond regularity), which are themselves dictating perceptual learning outcomes and which have not been quantified in our experiments. Although this possibility cannot be dismissed, as this study has only considered one measure of speech regularity, it is worth first considering whether there exist other measures of variability that could better quantify the regularity in speech that is necessary for perceptual learning. Along these lines, we should note that, although prior research has shown moderate or greater correlations between the acoustic and kinematic STI measures, the acoustic STI itself remains underexplored. As such, it remains possible that measurement of the kinematic STI (where kinematic data available) could better explain the variation in perceptual learning outcomes observed in this study. One indicator of possible limitations in the acoustic STI observed in this study is our failure to observe age-related changes in the acoustic STI in this study (i.e., older children did not have lower acoustic STI values), whereas the kinematic STI has been consistently found to decrease with age in similar populations (Howell et al., 2009). Perhaps the most salient explanation for the limited observable relationship between the acoustic STI and perceptual learning is that the STI primarily captures prosodic or suprasegmental variability. The amplitude envelope on which the acoustic STI is based captures temporally local extrema but does not characterize any spectral properties, which are known to differ between speech sounds. A major part of the hypothesized relationship between predictability and learnability is that consistent substitutions of phonemes (such as the common substitution of /ɹ/ with /w/ in children's speech) will be learnable, whereas inconsistent substitutions would not be. However, as these phonetic differences are not well represented by the amplitude envelope, the acoustic STI likely fails to directly capture this type of inconsistency. As a result, the acoustic STI, and perhaps the kinematic STI as well, may not be the most effective measure to capture the type of variability in speech production that influences perceptual learning.
Although the findings of this study lend partial support the efficacy of the acoustic STI as a measure capable of forecasting perceptual learning outcomes, there remains a need to better understand the characteristics of individual's speech that do influence these outcomes. Measures of segmental variability that directly characterize variability in the perceptual categorization of speech sounds might be better suited toward capturing the type of variability that influences perceptual learning, particularly in children's speech. Furthermore, although we might expect these different forms of variability to be linked to one another, prior research has found these measures to be relatively uncorrelated with one another (Vuolo & Goffman, 2017). Future work could also explore modifications to the acoustic STI or the development of new variability measures better suited to this task. Note that there have already been several modifications to the STI proposed in the existing literature that were not explored in this study. These include a nonlinear version of the STI proposed in Lucero et al. (1997) and recently proposed robust version of the STI introduced in Wisler et al. (2025). Although these modified measures might lead to some improvement on the traditional STI, neither addresses what we believe to be the root problem with the acoustic STI in this context, which is a lack of sensitivity to acoustic properties that differentiate speech on a segmental (rather than suprasegmental) level. Thus, a more important modification would be to move beyond the amplitude envelope and explore measures of variability based on short-term spectral characteristics of the speech signal. Another avenue for future research would be to perform a more fine-grained analysis of the relationship between speech variability and perceptual learning at the phoneme level. Our current design focuses on phrase-level intelligibility but does not capture how learning might vary as a function of within-speaker variability across specific speech sounds. For example, if a child produces /a/ consistently but /r/ inconsistently, learning may be stronger for the more stable phoneme. Because analyzing this relationship at the phoneme or even word level would require a large number of repeated tokens for each phoneme across different contexts, we were unable to pursue the analysis with the current data set. Nevertheless, future work should address this limitation by systematically targeting specific phonemes with controlled repetition across stimuli.
Along these lines, it would be of interest to explore both manual and automated measures of segmental variability. Manual measures could include comparison of phonetic transcriptions across repetitions of the same stimuli, as well as looking at the spectral qualities of sounds across repetitions (e.g., formant and/or spectral analysis). Alternatively, the use of automatic speech recognition (ASR) systems could be explored. One could evaluate how ASR accuracy and the types of inaccuracies output by ASR reflect variability in the speech signal. Furthermore, some ASR systems offer phonetic transcription (e.g., wav2vec 2.0 from HuggingFace), which would allow for the investigation of segmental-level errors, which researchers could then compare across productions to quantify segmental deviations as a mirror for predictability.
In conclusion, this study is the first to show that adult listeners can adapt to and improve their understanding of the highly variable speech of children. It is also the first study to show that the acoustic STI is predictive of the degree of learning achieved for different child participants. Although the results of this study provide support for the hypothesized relationship between the acoustic STI and the degree perceptual learning, the observed relationship was not as strong as initially anticipated. Thus, these findings motivate the consideration of other measures of speech variability, particularly those that can capture segmental variability, as predictors of perceptual learning outcomes. Therefore, future work is needed to explore the role of segmental predictability in perceptual learning, both in children's speech and across other populations.
Data Availability Statement
Listener data and statistical code for this study can be found at https://osf.io/ekuj8/.
Acknowledgments
This research was supported by National Institute on Deafness and Other Communication Disorders Grants R01DC020930 and R01DC020713 (awarded to Stephanie A. Borrie), as well as by a Research Catalyst Grant from the Utah State University Office of Research (awarded to Alan Wisler).
Appendix
A. Pretest Phrases
| Phrase | Number of syllables | Number of phonemes | Word complexity measure |
|---|---|---|---|
| Angry money | 4 | 9 | 6 |
| Beautiful thumb | 4 | 11 | 6 |
| Beeping broccoli | 4 | 11 | 9 |
| Better macaroni | 6 | 11 | 5 |
| Blue temperature | 4 | 11 | 6 |
| Breakable drawing | 5 | 13 | 9 |
| Bumpy snake | 3 | 9 | 6 |
| Calm slipper | 3 | 9 | 6 |
| Chubby turtle | 4 | 9 | 7 |
| Clever shark | 3 | 9 | 10 |
| Crazy trampoline | 5 | 14 | 12 |
| Cruel butterfly | 4 | 11 | 10 |
| Dark snail | 2 | 8 | 8 |
| Dizzy fountain | 4 | 10 | 7 |
| Easy ladder | 4 | 7 | 6 |
| Empty guitar | 4 | 10 | 8 |
| Funny river | 4 | 8 | 7 |
| Great grasshopper | 4 | 12 | 12 |
| Grumpy watch | 3 | 9 | 7 |
| Happy sock | 3 | 7 | 4 |
| Horrible squirrel | 5 | 13 | 11 |
| Hungry pajamas | 5 | 13 | 10 |
| Jealous stomach | 4 | 11 | 10 |
| Key dandelion | 5 | 11 | 5 |
| Large ghost | 2 | 8 | 10 |
| Light parrot | 3 | 8 | 5 |
| Lonely cauliflower | 6 | 13 | 11 |
| Long lollipop | 5 | 10 | 7 |
| Narrow jelly | 4 | 8 | 6 |
| Nervous puzzle | 4 | 10 | 11 |
| Normal chair | 3 | 9 | 8 |
| Old coffee | 3 | 7 | 7 |
| Plain toilet | 3 | 9 | 6 |
| Polite basket | 4 | 11 | 8 |
| Proud cookie | 3 | 8 | 5 |
| Red mouse | 2 | 6 | 4 |
| Rich sunglasses | 4 | 12 | 12 |
| Round roof | 2 | 7 | 5 |
| Sick microwave | 4 | 11 | 9 |
| Silver sponge | 3 | 10 | 13 |
| Small farmer | 3 | 9 | 8 |
| Smart noodle | 3 | 10 | 7 |
| Soft brush | 2 | 8 | 7 |
| Squishy plate | 3 | 10 | 8 |
| Stinky bubble | 4 | 11 | 8 |
| Tall frog | 2 | 7 | 5 |
| Thoughtful blender | 4 | 12 | 11 |
| Ugly ribbon | 4 | 9 | 7 |
| Wild thermometer | 5 | 12 | 8 |
| Yellow umbrella | 4 | 11 | 8 |
B. Posttest Phrases
| Phrase | Number of syllables | Number of phonemes | Word complexity measure |
|---|---|---|---|
| Annoyed hamburger | 5 | 11 | 7 |
| Big spaghetti | 4 | 10 | 6 |
| Brave spoon | 2 | 8 | 7 |
| Brown telescope | 4 | 12 | 9 |
| Bubbly fish | 3 | 8 | 5 |
| Busy parachute | 5 | 11 | 7 |
| Careful alligator | 6 | 13 | 9 |
| Clear banana | 4 | 10 | 6 |
| Cold knife | 2 | 7 | 5 |
| Crispy towel | 4 | 10 | 9 |
| Cute dinosaur | 4 | 11 | 6 |
| Dirty train | 3 | 8 | 7 |
| Dry motorcycle | 5 | 12 | 9 |
| Embarrassed scooter | 5 | 13 | 11 |
| Expensive lemon | 5 | 14 | 14 |
| Fancy crocodile | 5 | 13 | 10 |
| Fit caterpillar | 5 | 11 | 9 |
| Good kangaroo | 4 | 9 | 7 |
| Green cat | 2 | 7 | 7 |
| Hairy duck | 3 | 7 | 4 |
| Helpful monkey | 4 | 12 | 8 |
| Huge balloon | 3 | 9 | 7 |
| Itchy tomato | 5 | 9 | 4 |
| Joyful wolf | 3 | 9 | 9 |
| Kind rabbit | 3 | 9 | 5 |
| Lazy spider | 4 | 9 | 9 |
| Lively feather | 4 | 9 | 11 |
| Lumpy honey | 4 | 9 | 4 |
| Nasty computer | 5 | 13 | 8 |
| North helicopter | 5 | 13 | 11 |
| Open pudding | 4 | 9 | 4 |
| Orange tractor | 3 | 11 | 11 |
| Pink refrigerator | 6 | 14 | 13 |
| Pleasant airplane | 4 | 13 | 11 |
| Pretty kitchen | 4 | 10 | 7 |
| Purple shovel | 4 | 10 | 11 |
| Quiet scarf | 3 | 10 | 9 |
| Sad blueberries | 4 | 11 | 9 |
| Scary barn | 3 | 9 | 8 |
| Short leaf | 2 | 7 | 6 |
| Silly sandwich | 4 | 11 | 8 |
| Skinny book | 3 | 8 | 6 |
| Smooth glove | 2 | 8 | 10 |
| Sour garbage | 3 | 9 | 10 |
| Sticky car | 3 | 8 | 7 |
| Sweet carpet | 3 | 10 | 8 |
| Tiny flower | 4 | 8 | 6 |
| Wet sausage | 3 | 8 | 7 |
| Wise mermaid | 3 | 8 | 7 |
| Yummy picture | 4 | 9 | 7 |
Funding Statement
This research was supported by National Institute on Deafness and Other Communication Disorders Grants R01DC020930 and R01DC020713 (awarded to Stephanie A. Borrie), as well as by a Research Catalyst Grant from the Utah State University Office of Research (awarded to Alan Wisler).
Footnotes
Because children repeated phrases after hearing a model, this approach may have reduced some natural variability in production and introduced individual differences related to imitation ability (e.g., children who were stronger imitators may have produced phrases with more adultlike prosody or amplitude envelope); however, this decision was made to ensure consistency between the child talkers because many children in this age range are not yet able to reliably read stimulus phrases.
For two of the child talkers, only eight and nine listeners had perception data due to a problem with the Gorilla.sc server not saving their responses.
References
- Anderson, A., Lowit, A., & Howell, P. (2008). Temporal and spatial variability in speakers with Parkinson's disease and Friedreich's ataxia. Journal of Medical Speech-Language Pathology, 16(4), Article 173. [PMC free article] [PubMed] [Google Scholar]
- Anderson, C., & Cohen, W. (2012). Measuring word complexity in speech screening: Single-word sampling to identify phonological delay/disorder in preschool children. International Journal of Language & Communication Disorders, 47(5), 534–541. 10.1111/j.1460-6984.2012.00163.x [DOI] [PubMed] [Google Scholar]
- Barrett, T. S., Borrie, S. A., & Yoho, S. E. (2019). Automating with autoscore: Introducing an R package for automating the scoring of orthographic transcripts. https://github.com/autoscore/autoscore
- Bates, E., Dale, P. S., & Thal, D. (2003). Individual differences and their implications for theories of language development. In Bornstein M. & Lamb M. (Eds.), Developmental science (pp. 96–151). Erlbaum. https://www.researchgate.net/profile/Philip-Dale-3/publication/373338544_Individual_Differences_and_their_Implications_for_Theories_of_Language_Development/links/6560f52cb1398a779dad5f05/Individual-Differences-and-their-Implications-for-Theories-of-Language-Development.pdf [PDF] [Google Scholar]
- Benham, S., Wisler, A., Berlin, J., Wang, J., & Goffman, L. (2023). Acoustic and kinematic methods of indexing spatiotemporal stability in children with developmental language disorder. Journal of Speech, Language, and Hearing Research, 66(8S), 3026–3037. 10.1044/2022_JSLHR-22-00290 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Borrie, S. A., Barrett, T. S., & Yoho, S. E. (2019). Autoscore: An open-source automated tool for scoring listener perception of speech. The Journal of the Acoustical Society of America, 145(1), 392–399. 10.1121/1.5087276 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Borrie, S. A., Hepworth, T. J., Wynn, C. J., Hustad, K. C., Barrett, T. S., & Lansford, K. L. (2023). Perceptual learning of dysarthria in adolescence. Journal of Speech, Language, and Hearing Research, 66(10), 3791–3803. 10.1044/2023_JSLHR-23-00231 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Borrie, S. A., & Lansford, K. L. (2021). A perceptual learning approach for dysarthria remediation: An updated review. Journal of Speech, Language, and Hearing Research, 64(8), 3060–3073. 10.1044/2021_JSLHR-21-00012 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Borrie, S. A., Lansford, K. L., & Barrett, T. S. (2017). Rhythm perception and its role in perception and learning of dysrhythmic speech. Journal of Speech, Language, and Hearing Research, 60(3), 561–570. 10.1044/2016_JSLHR-S-16-0094 [DOI] [PubMed] [Google Scholar]
- Borrie, S. A., Lansford, K. L., & Barrett, T. S. (2018). Understanding dysrhythmic speech: When rhythm does not matter and learning does not happen. The Journal of the Acoustical Society of America, 143(5), EL379–EL385. 10.1121/1.5037620 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bradlow, A. R., & Bent, T. (2008). Perceptual adaptation to non-native speech. Cognition, 106(2), 707–729. 10.1016/j.cognition.2007.04.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Brumbach, A. C. D., & Goffman, L. (2014). Interaction of language processing and motor skill in children with specific language impairment. Journal of Speech, Language, and Hearing Research, 57(1), 158–171. 10.1044/1092-4388(2013/12-0215) [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chermak, G. D., & Schneiderman, C. R. (1986). Speech timing variability of children and adults. Journal of Phonetics, 13(4), 477–480. 10.1016/S0095-4470(19)30799-5 [DOI] [Google Scholar]
- Crowe, K., & McLeod, S. (2018). Children's consonant acquisition in 27 languages: A cross-linguistic review. American Journal of Speech-Language Pathology, 29(1), 17–42. 10.1044/2018_AJSLP-17-0100 [DOI] [PubMed] [Google Scholar]
- Davis, M. H., Johnsrude, I. S., Hervais-Adelman, A., Taylor, K., & McGettigan, C. (2005). Lexical information drives perceptual learning of distorted speech: Evidence from the comprehension of noise-vocoded sentences. Journal of Experimental Psychology: General, 134(2), 222–241. 10.1037/0096-3445.134.2.222 [DOI] [PubMed] [Google Scholar]
- DiSimoni, F. G. (1974a). Effect of vowel environment on the duration of consonants in the speech of three-, six-, and nine-year-old children. Journal of the Acoustical Society of America, 55(2), 360–361. 10.1121/1.1914513 [DOI] [PubMed] [Google Scholar]
- DiSimoni, F. G. (1974b). Influence of consonant environment on duration of vowels in the speech of three-, six-, and nine-year-old children. The Journal of the Acoustical Society of America, 55(2), 362–363. 10.1121/1.1914514 [DOI] [PubMed] [Google Scholar]
- Dupoux, E., & Green, K. (1997). Perceptual adjustment to highly compressed speech: Effects of talker and rate changes. Journal of Experimental Psychology: Human Perception and Performance, 23(3), 914–927. 10.1037//0096-1523.23.3.914 [DOI] [PubMed] [Google Scholar]
- Eisner, F., & McQueen, J. M. (2005). The specificity of perceptual learning in speech processing. Perception & Psychophysics, 67(2), 224–238. 10.3758/BF03206487 [DOI] [PubMed] [Google Scholar]
- Emerson, J. W., Green, W. A., Schloerke, B., Crowley, J., Cook, D., Hofmann, H., & Wickham, H. (2013). The generalized pairs plot. Journal of Computational and Graphical Statistics, 22(1), 79–91. 10.1080/10618600.2012.694762 [DOI] [Google Scholar]
- Goffman, L. (1999). Prosodic influences on speech production in children with specific language impairment and speech deficits: Kinematic, acoustic, and transcription evidence. Journal of Speech, Language, and Hearing Research, 42(6), 1499–1517. 10.1044/jslhr.4206.1499 [DOI] [PubMed] [Google Scholar]
- Goffman, L., Gerken, L., & Lucchesi, J. (2007). Relations between segmental and motor variability in prosodically complex nonword sequences. Journal of Speech, Language, and Hearing Research, 50(2), 444–458. 10.1044/1092-4388(2007/031) [DOI] [PubMed] [Google Scholar]
- Golomb, J. D., Peelle, J. E., & Wingfield, A. (2007). Effects of stimulus variability and adult aging on adaptation to time-compressed speech. The Journal of the Acoustical Society of America, 121(3), 1701–1708. 10.1121/1.2436635 [DOI] [PubMed] [Google Scholar]
- Grigos, M. I., Moss, A., & Lu, Y. (2015). Oral articulatory control in childhood apraxia of speech. Journal of Speech, Language, and Hearing Research, 58(4), 1103–1118. 10.1044/2015_JSLHR-S-13-0221 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Heffner, C. C., & Myers, E. B. (2021). Individual differences in phonetic plasticity across native and nonnative contexts. Journal of Speech, Language, and Hearing Research, 64(10), 3720–3733. 10.1044/2021_JSLHR-21-00004 [DOI] [PubMed] [Google Scholar]
- Hervais-Adelman, A., Davis, M. H., Johnsrude, I. S., & Carlyon, R. P. (2008). Perceptual learning of noise vocoded words: Effects of feedback and lexicality. Journal of Experimental Psychology: Human Perception and Performance, 34(2), 460–474. 10.1037/0096-1523.34.2.460 [DOI] [PubMed] [Google Scholar]
- Holm, A., Crosbie, S., & Dodd, B. (2010). Differentiating normal variability from inconsistency in children's speech: Normative data. International Journal of Language & Communication Disorders, 42(4), 467–486. 10.1080/13682820600988967 [DOI] [PubMed] [Google Scholar]
- Howell, P., Anderson, A. J., Bartrip, J., & Bailey, E. (2009). Comparison of acoustic and kinematic approaches to measuring utterance-level speech variability. Journal of Speech, Language, and Hearing Research, 52(4), 1088–1096. 10.1044/1092-4388(2009/07-0167) [DOI] [PMC free article] [PubMed] [Google Scholar]
- Huber, J. E., Stathopoulos, E. T., Curione, G. M., Ash, T. A., & Johnson, K. (1999). Formants of children, women, and men: The effects of vocal intensity variation. The Journal of the Acoustical Society of America, 106(3), 1532–1542. 10.1121/1.427150 [DOI] [PubMed] [Google Scholar]
- Hustad, K. C., Mahr, T. J., Natzke, P., & Rathouz, P. J. (2021). Speech development between 30 and 119 months in typical children I: Intelligibility growth curves for single-word and multiword productions. Journal of Speech, Language, and Hearing Research, 64(10), 3707–3719. 10.1044/2021_JSLHR-21-00142 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kent, R. D., & Forner, L. L. (1980). Speech segment durations in sentence recitations by children and adults. Journal of Phonetics, 8(2), 157–168. 10.1016/S0095-4470(19)31460-3 [DOI] [Google Scholar]
- Kleinschmidt, D. F., & Jaeger, T. F. (2015). Robust speech perception: Recognize the familiar, generalize to the similar, and adapt to the novel. Psychological Review, 122(2), 148–203. 10.1037/a0038695 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kleinschmidt, D. F., & Jaeger, T. F. (2016). What do you expect from an unfamiliar talker? Proceedings from the Annual Meeting of the Cognitive Science Society, 2351–2356. https://escholarship.org/content/qt7h31v1x7/qt7h31v1x7_noSplash_e25ede3de707c1a688ef62129ba0b57e.pdf [PDF] [Google Scholar]
- Koenig, L. L. (2001). Distributional characteristics of VOT in children's voiceless aspirated stops and interpretation of developmental trends. Journal of Speech, Language, and Hearing Research, 44(5), 1058–1068. 10.1044/1092-4388(2001/084) [DOI] [PubMed] [Google Scholar]
- Lansford, K. L., Barrett, T. S., & Borrie, S. A. (2023). Cognitive predictors of perception and adaptation to dysarthric speech in young adult listeners. Journal of Speech, Language, and Hearing Research, 66(1), 30–47. 10.1044/2022_JSLHR-22-00391 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lansford, K. L., Borrie, S. A., & Barrett, T. S. (2019). Regularity matters: Unpredictable speech degradation inhibits adaptation to dysarthric speech. Journal of Speech, Language, and Hearing Research, 62(12), 4282–4290. 10.1044/2019_JSLHR-19-00055 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lansford, K. L., Borrie, S. A., Barrett, T. S., & Flechaus, C. (2020). When additional training isn't enough: Further evidence that unpredictable speech inhibits adaptation. Journal of Speech, Language, and Hearing Research, 63(6), 1700–1711. 10.1044/2020_JSLHR-19-00380 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lee, J., Hustad, K. C., & Weismer, G. (2014). Predicting speech intelligibility with a multiple speech subsystems approach in children with cerebral palsy. Journal of Speech, Language, and Hearing Research, 57(5), 1666–1678. 10.1044/2014_JSLHR-S-13-0292 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lee, S., Potamianos, A., & Narayanan, S. (1999). Acoustics of children's speech: Developmental changes of temporal and spectral parameters. The Journal of the Acoustical Society of America, 105(3), 1455–1468. 10.1121/1.426686 [DOI] [PubMed] [Google Scholar]
- Lucero, J. C., Munhall, K. G., Gracco, V. L., & Ramsay, J. O. (1997). On the registration of time and the patterning of speech movements. Journal of Speech, Language, and Hearing Research, 40(5), 1111–1117. 10.1044/jslhr.4005.1111 [DOI] [PubMed] [Google Scholar]
- Macrae, T., Tyler, A. A., & Lewis, K. E. (2014). Lexical and phonological variability in preschool children with speech sound disorder. American Journal of Speech-Language Pathology, 23(1), 27–35. 10.1044/1058-0360(2013/12-0037) [DOI] [PubMed] [Google Scholar]
- McLeod, S., Crowe, K., & Shahaeian, A. (2015). Intelligibility in context scale: Normative and validation data for English-speaking preschoolers. Language, Speech, and Hearing Services in Schools, 46(3), 266–276. 10.1044/2015_LSHSS-14-0120 [DOI] [PubMed] [Google Scholar]
- McNeill, B., McIlraith, A. L., Macrae, T., Gath, M., & Gillon, G. (2022). Predictors of speech severity and inconsistency over time in children with token-to-token inconsistency. Journal of Speech, Language, and Hearing Research, 65(7), 2459–2473. 10.1044/2022_JSLHR-21-00611 [DOI] [PubMed] [Google Scholar]
- Moss, A., & Grigos, M. I. (2012). Interarticulatory coordination of the lips and jaw in childhood apraxia of speech. Journal of Medical Speech-Language Pathology, 20(4), Article 127. [PMC free article] [PubMed] [Google Scholar]
- Norris, D., McQueen, J. M., & Cutler, A. (2003). Perceptual learning in speech. Cognitive Psychology, 47(2), 204–238. 10.1016/S0010-0285(03)00006-9 [DOI] [PubMed] [Google Scholar]
- Preston, J. L., & Koenig, L. L. (2011). Phonetic variability in residual speech sound disorders: Exploration of subtypes. Topics in Language Disorders, 31(2), 168–184. 10.1097/TLD.0b013e318217b875 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sadagopan, N., & Smith, A. (2008). Developmental changes in the effects of utterance length and complexity on speech movement variability. Journal of Speech, Language, and Hearing Research, 51(5), 1138–1151. 10.1044/1092-4388(2008/06-0222) [DOI] [PubMed] [Google Scholar]
- Saletta, M., Goffman, L., Ward, C., & Oleson, J. (2018). Influence of language load on speech motor skill in children with specific language impairment. Journal of Speech, Language, and Hearing Research, 61(3), 675–689. 10.1044/2017_JSLHR-L-17-0066 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sidaras, S. K., Alexander, J. E., & Nygaard, L. C. (2009). Perceptual learning of systematic variation in Spanish-accented speech. The Journal of the Acoustical Society of America, 125(5), 3306–3316. 10.1121/1.3101452 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith, A., & Goffman, L. (1998). Stability and patterning of speech movement sequences in children and adults. Journal of Speech, Language, and Hearing Research, 41(1), 18–30. 10.1044/jslhr.4101.18 [DOI] [PubMed] [Google Scholar]
- Smith, A., Goffman, L., Zelaznik, H. N., Ying, G., & McGillem, C. (1995). Spatiotemporal stability and patterning of speech movement sequences. Experimental Brain Research, 104, 493–501. 10.1007/BF00231983 [DOI] [PubMed] [Google Scholar]
- Smith, B. L. (1992). Relationships between duration and temporal variability in children's speech. The Journal of the Acoustical Society of America, 91(4), 2165–2174. 10.1121/1.403675 [DOI] [PubMed] [Google Scholar]
- Sosa, A. V. (2015). Intraword variability in typical speech development. American Journal of Speech-Language Pathology, 24(1), 24–35. 10.1044/2014_AJSLP-13-0148 [DOI] [PubMed] [Google Scholar]
- Stipancic, K. L., Golzy, M., Zhao, Y., Pinkerton, L., Rohl, A., & Kuruvilla-Dugdale, M. (2023). Improving perceptual speech ratings: The effects of auditory training on judgments of dysarthric speech. Journal of Speech, Language, and Hearing Research, 66(11), 4236–4258. 10.1044/2023_JSLHR-23-00322 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stipancic, K. L., van Brenk, F., Qiu, M., & Tjaden, K. (2025). Progress toward estimating the minimal clinically important difference of intelligibility: A crowdsourced perceptual experiment. Journal of Speech, Language, and Hearing Research, 68(7S), 3480–3494. 10.1044/2024_JSLHR-24-00354 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stoel-Gammon, C. (2010). The word complexity measure: Description and application to developmental phonology and disorders. Clinical Linguistics & Phonetics, 24(4–5), 271–282. 10.3109/02699200903581059 [DOI] [PubMed] [Google Scholar]
- Tingley, B. M., & Allen, G. D. (1975). Development of speech timing control in children. Child Development, 46(1), 186–194. 10.2307/1128847 [DOI] [Google Scholar]
- Vuolo, J., & Goffman, L. (2017). An exploratory study of the influence of load and practice on segmental and articulatory variability in children with speech sound disorders. Clinical Linguistics & Phonetics, 31(5), 331–350. 10.1080/02699206.2016.1261184 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vuolo, J., & Wisler, A. (2024). Acoustic analysis of spatiotemporal variability in children with childhood apraxia of speech. Journal of Speech, Language, and Hearing Research, 67(10), 3536–3548. 10.1044/2024_JSLHR-24-00079 [DOI] [PubMed] [Google Scholar]
- Walsh, B., & Smith, A. (2002). Articulatory movements in adolescents: Evidence for protracted development of speech motor control processes. Journal of Speech, Language, and Hearing Research, 45(6), 1119–1133. 10.1044/1092-4388(2002/090) [DOI] [PubMed] [Google Scholar]
- Wang, E. W., & Grigos, M. I. (2024). Effects of speaking rate changes on speech motor variability in adults. Language and Speech, 68(1), 141–161. 10.1177/00238309241252983 [DOI] [PubMed] [Google Scholar]
- Whiteside, S. P., Dobbin, R., & Henry, L. (2003). Patterns of variability in voice onset time: A developmental study of motor speech skills in humans. Neuroscience Letters, 347(1), 29–32. 10.1016/S0304-3940(03)00598-6 [DOI] [PubMed] [Google Scholar]
- Wilson, S., & Tapper, L. (2015). Hedgehugs. Henry Holt and Co. [Google Scholar]
- Wisler, A., Goffman, L., Zhang, L., & Wang, J. (2022). Influences of methodological decisions on assessing the spatiotemporal stability of speech movement sequences. Journal of Speech, Language, and Hearing Research, 65(2), 538–554. 10.1044/2021_JSLHR-21-00298 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wisler, A., Teplansky, K., Berlin, J., Wang, J., & Goffman, L. (2024). Validating the influences of methodological decisions on assessing the spatiotemporal stability of speech movement sequences using children's speech data. Journal of Speech, Language, and Hearing Research, 67(12), 4585–4597. 10.1044/2024_JSLHR-24-00190 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wisler, A., Vuolo, J., & Fletcher, A. (2025). The benefits of robustness in measures of spatiotemporal stability: An investigation in childhood apraxia of speech. Journal of Speech, Language, and Hearing Research, 68(7S), 3495–3506. 10.1044/2024_JSLHR-24-00360 [DOI] [PubMed] [Google Scholar]
- Ziegler, W., Lehner, K., Klonowski, M., Geißler, N., Ammer, F., Kurfeß, C., Grötzbach, H., Mandl, A., Knorr, F., & Strecker, K. (2021). Crowdsourcing as a tool in the clinical assessment of intelligibility in dysarthria: How to deal with excessive variation. Journal of Communication Disorders, 93, 106–135. 10.1016/j.jcomdis.2021.106135 [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Listener data and statistical code for this study can be found at https://osf.io/ekuj8/.



