Abstract
Purpose:
The aims of the current study were to determine age-related changes to the phonatory and articulatory subsystems and to investigate an exploratory model of intelligibility for healthy aging based on phonatory and articulatory measures.
Method:
Fifteen healthy, older adults (55–81 years) and 15 younger adults (20–35 years) participated in instrumental assessments of the phonatory (aerodynamic, acoustic) and articulatory (kinematic) subsystems. Speech intelligibility was determined by five listeners during multi-talker babble.
Results:
Older adults displayed shorter maximum phonation time, greater airflow during sentence reading, and lower cepstral peak prominence (CPP) and CPP SD. Additionally, older adults had slower tongue movement speed than younger adults. Speech intelligibility was also significantly reduced in the older group. A generalized estimating equations model combining phonatory and articulatory measures showed that CPP SD, low/high (L/H) spectral ratio mean and SD, Cepstral Spectral Index of Dysphonia (CSID), and maximum tongue movement speed were significant contributors to intelligibility changes in older individuals. While L/H mean and SD and CSID displayed an inverse relationship with intelligibility, CPP SD and maximum tongue speed displayed a direct relationship with intelligibility.
Discussion:
Aging affects the phonatory and articulatory subsystems with implications for speech intelligibility. Phonatory cepstral/spectral measures except for mean CPP were associated with speech intelligibility changes, suggesting that changes in voice quality may contribute to reduced intelligibility in older adults. Pertaining to articulation, slower tongue movement speed likely contributed to reduced intelligibility in older individuals.
Keywords: Aging, Voice, Articulation, Speech intelligibility
1. Introduction
Speech production is altered as part of the normal aging process because of structural and physiologic changes in each speech subsystem, which can begin as early as 50 years (Hixon, Weismer, & Hoit, 2014). Among the speech subsystems, age-related changes to the respiratory (Hoit & Hixon, 1987; Huber & Spruill, 2008; Zraick, Smith-Olinde, & Shotts, 2012) and phonatory1 subsystems (Awan, 2006; Baker, Ramig, Sapir, Luschei, & Smith, 2001; Gregory, Chandran, Lurie, & Sataloff, 2012; Thomas, Harrison, & Stemple, 2008; Watts, Ronshaugen, & Saenz, 2015; Zraick et al., 2012) have been studied to a greater extent than the articulatory subsystem (Liss, Weismer, & Rosenbek, 1990; Wohlert & Smith, 1998). Although older individuals appear to maintain functional communication (Hooper & Cralidis, 2009), the impact of subsystem declines on speech intelligibility in the aging population is poorly understood. For neurological disorders like amyotrophic lateral sclerosis (ALS) and cerebral palsy (CP), researchers are beginning to use a multiple speech subsystem approach to establish a predictive model of speech intelligibility (Lee, Hustad, & Weismer, 2014; Rong et al., 2016). Because bulbar decline in diseases like ALS and Parkinson’s disease (PD) is often overlaid on aging-related changes, it is important to differentiate healthy aging from disease processes (Dromey, Boyce, & Channell, 2014; Duffy, 2013).
As a first step towards understanding the impact of subsystem dysfunction on intelligibility, the current study focused on the phonatory and articulatory subsystems to investigate an exploratory model of speech intelligibility for healthy aging. The phonatory subsystem was included because age-related structural and physiological changes leading to presbyphonia are often observed among the elderly (Gregory et al., 2012), along with self-reports of reduced loudness, greater vocal effort, and/or altered voice quality (Etter, Stemple, & Howell, 2013; Roy, Stemple, Merrill, & Thomas, 2007) that may impact intelligibility (Ishikawa et al., 2018). Articulatory subsystem measures were included because compared to other subsystem measures (e.g., respiratory and resonatory) they were most sensitive to the intelligibility deficit in individuals with ALS and CP (Lee et al., 2014; Rong et al., 2016).
1.1. Aging of the Phonatory Subsystem
Laryngeal structure and function deteriorate during the aging process resulting in presbylarynx and presbyphonia (Kendall, 2007). The structural changes include ossification of hyaline laryngeal cartilages (Hixon et al., 2014; Kendall, 2007), calcification of elastic cartilages such as the epiglottis and parts of the arytenoid cartilages (Hixon et al., 2014), muscle fiber loss leading to vocal fold atrophy and bowing (Mau, Jacobson, & Garrett, 2010; Thomas et al., 2008), and loss of vocal fold tissue viscoelasticity (Mueller, Sweeney, & Baribeau, 1984; Sato, Hirano, & Nakashima, 2002). Further, presbylaryngis may be linked to increased motor unit durations, denervation-like changes in neuromuscular junctions, increase in glycolytic metabolism, increase in mitochondrial abnormalities, and decrease in sex hormone receptors (Takeda, Thomas, & Ludlow, 2000; Thomas et al., 2008). Overall, aging leads to slower, weaker, and less fatigue-resistant vocal fold muscles and reduced elasticity of the vocal fold mucosa (Thomas et al., 2008).
Although there is collective agreement that perceptual characteristics as well as aerodynamic and acoustic phonatory function deteriorate with age, there are conflicting reports of the extent to which these changes occur. Salient perceptual characteristics such as pitch alterations, reduced loudness, altered voice quality, and vocal instability (Baker et al., 2001; Kendall, 2007) are not always observed in older adults (Hooper & Cralidis, 2009). Phonatory aerodynamic function testing including maximum phonation time and laryngeal efficiency measures (mean airflow during voicing, mean peak air pressure, airway resistance) did not yield significant age differences (Zraick et al., 2012). Acoustic studies have also shown mixed results for aging differences. One study using perturbation measures found significantly greater jitter, shimmer, and noise-to-harmonic ratio in a sample of older compared to younger adults (Xue & Deliyski, 2001). While traditional time-based perturbation measures requiring sustained phonations have limitations, cepstral/spectral measures can be used with connected speech and are ecologically more valid (Awan, Roy, Jetté, Meltzner, & Hillman, 2010). The Cepstral Spectral Index of Dysphonia (CSID; Awan, 2011) is a multivariate estimate of dysphonia severity based on a weighted formula that contains means and SDs for cepstral peak prominence (CPP) and low/high (L/H) spectral ratio. CPP, whose relative amplitude of the cepstral peak corresponds to dominant harmonic energy in a person’s voice, is the strongest individual contributor to the formula (Awan, 2011; Lowell, 2012). In a recent study, Awan and colleagues found that values on the CSID were not significantly different between older and younger adults, but that mean CPP and L/H spectral ratio SD were significantly lower in older than younger adults (Awan, Acompanado, Connors, & Fanelli, 2015). The results are consistent with the fact that CPP contributes most to the CSID formula and that dysphonia is characterized by smaller CPP values (Awan, 2011). Relevant to the present study, CPP was used recently to investigate the link between dysphonia severity and intelligibility, for which the harmonic structure of the acoustic signal is important (Ishikawa et al., 2018). The authors noted that CPP moderately predicted intelligibility of dysphonic voices in noise, in addition other acoustic features, especially spectral measures, were also thought to contribute to intelligibility (Ishikawa et al., 2018).
1.2. Aging of the Articulatory Subsystem
Documented changes to skeletal structures include degeneration of the articulator surface of the mandibular condyle (Ishibashi, Takenoshita, Ishibashi, & Oka, 1995), loss of capsular ligament elasticity in the temporomandibular joint (Kahane, 1987), and degradation of the alveolar bone in the maxilla and mandible (Zemlin, 1998). Aging-induced muscular changes to the tongue include decreased thickness of the lingual epithelium and reduced muscle fiber diameter (Nakayama, 1991). Interestingly, increased fatty tissue deposits are thought to compensate for this tissue loss and help maintain the size and shape of the aging tongue (Bässler, 1987). For jaw muscles, unique fiber type changes were observed for each muscle. For example, for the masseter, fiber type changes included a decrease in type I fibers and an increase in type IM and type II fibers. For the lateral pterygoid, however, only type IIA fibers increase (Monemi, Eriksson, Eriksson, & Thornell, 1998). Denervation of type I fibers due to the loss of slow-twitch motor neurons and reinnervation by fast-twitch motor neurons may be responsible for the increased variability in fiber diameter and atrophy with age (Monemi et al., 1998; Monemi, Thornell, & Eriksson, 1999). Alterations to bony structures, muscle fiber type, and cellular composition may underlie the articulatory subsystem changes observed in the elderly.
As described for the phonatory system, evidence varies for age-induced changes in articulatory function (Goozée, Stephenson, Murdoch, Darnell, & Lapointe, 2005). In some aging studies, changes in motor function are demonstrated through findings such as decreased tongue and lip strength (Robbins, Levine, Wood, Roecker, & Luschei, 1995; Wohlert & Smith, 1998) and altered perceptual characteristics including consonant imprecision (Ryan & Burk, 1974), slow speaking and reading rates (Amerman & Parnell, 1992; Flanagan & Dembowski, 2002), and increased utterance durations (Parnell & Amerman, 1996; B. L. Smith, Wasowicz, & Preston, 1987). In contrast, other studies suggest that tongue endurance (Crow & Ship, 1996), tongue strength (Youmans, Stierwalt, & Clark, 2002), tongue speed (Flanagan & Dembowski, 2002), and range of motion (Flanagan & Dembowski, 2002) are maintained with age. Kinematic studies in support of age-related differences in articulatory motor performance demonstrate reduced movement pattern consistency of the lower lip (Wohlert & Smith, 1998), reduced tongue retraction (Sonies, Stone, & Shawker, 1984), and smaller decreases in distance traveled by the tongue during fast speech rates (Goozée et al., 2005) in older adults compared to younger adults. Despite the existing literature, there are no comprehensive investigations of the tongue that include individual spatial and temporal measures as well as composite spatiotemporal indices extracted from multiple stimuli to determine aging effects on speech kinematics. Such a comprehensive approach will help capture age-related articulatory changes both at gross and fine kinematic levels, which is unlikely if fewer measures are used. For example, the spatiotemporal variability index (STI) provides information about overall variability in speech movement patterns across an entire utterance whereas peak speed measurements provide insights into speed generation capabilities associated with specific segments within an utterance. Moreover, STI is known to be relatively less sensitive to clinical changes compared to peak speed (Kuruvilla-Dugdale & Mefferd, 2017), which is a sensitive indicator of preclinical speech decline in conditions like ALS (Green, Yunusova, et al., 2013).
Even with the subsystem alterations described above, it is possible that intelligibility is spared in the elderly (McAuliffe, Wilding, Rickard, & O’Beirne, 2012; Shuey, 1989). However, no studies to date have systematically addressed the impact of age-related subsystem impairments on system-level measures like speech intelligibility.
1.3. Impact of Speech Subsystem Dysfunction on Speech Intelligibility
One approach to establishing a model of intelligibility is to use physiologic measures from multiple subsystems (respiratory, phonatory, resonatory, and articulatory) that are sensitive to age-related changes. Studies to date have used auditory-perceptual (De Bodt, Hernandez-Diaz Huici, & Van De Heyning, 2002), acoustic (Kim, Kent, & Weismer, 2011; Lee et al., 2014), or a combination of acoustic, aerodynamic, and kinematic (Rong et al., 2016) measures to represent the speech subsystems. An early study focused on a predictive model of intelligibility for dysarthria used auditory-perceptual judgments and found that features related to articulation had the strongest correlation with intelligibility (De Bodt et al., 2002). While measures based on perceptual ratings provide valuable clinical insights, they are subjective and are influenced by factors unrelated to the speaker such as the listener and/or environment (Kent, 1996).
Lee, Hustad, and Weismer (2014) used acoustic variables representing different subsystems to develop a predictive model of intelligibility for CP. Of the acoustic features, vowel space, vowel duration, and F2 slope were found to have the largest impact on intelligibility. Similarly, Kim et al. (2011) found that F2 slope was the most sensitive acoustic measure for predicting speech intelligibility across several types of dysarthria caused by various etiologies (e.g., PD, stroke). The use of acoustic measures to represent articulatory, resonatory, and phonatory subsystems provides objective data and also allows for clinical application; however, they “do not unambiguously represent the status of individual speech subsystems” (Rong et al., 2016, p. 22). This is especially true for subsystems like articulation, where there may not be a direct correspondence between articulatory adjustments and acoustic events (Mefferd & Green, 2010; Stevens, 1972, 1989). Therefore, acoustic measures related to the articulatory system were not included in the current study.
So far, only one study has used acoustic, aerodynamic, and kinematic measures to establish a predictive model of intelligibility decline for ALS (Rong et al., 2016). By using longitudinal subsystem data, these authors found that articulatory movement speed had the most substantial contribution to intelligibility decline over time, accounting for more than half of the variance. Their study demonstrated that an instrumentation-based multi-subsystem approach shows promise for predicting intelligibility decline in the ALS population. A limitation is that the phonatory subsystem only included maximum fundamental frequency but not cepstral/spectral measures, which would allow inferences about speech intelligibility in connected speech. Although prior studies suggest that the articulatory subsystem has the most significant impact on intelligibility, populations in which the phonatory subsystem is primarily affected, for example PD, may demonstrate that phonatory measures are most sensitive to the intelligibility deficit. Because the speech subsystems can be differentially impaired depending on dysarthria type and severity, it is imperative to develop dysarthria-specific models of intelligibility in the long-term. For similar reasons, age-specific models of intelligibility are also needed as speech subsystem impairments can vary with age progression.
1.4. Aims and Hypotheses
The purpose of this study was to obtain phonatory (aerodynamic, acoustic), and articulatory (kinematic) data as part of a multiple subsystem approach to investigate an exploratory model of speech intelligibility for healthy aging. Thus, the first aim of the current study was to determine age-related changes to the phonatory and articulatory subsystems. We hypothesized that older adults would display significantly lower mean CPP values than younger adults (Awan et al., 2015). Regarding the articulatory system, significantly greater movement pattern variability was hypothesized for older adults relative to younger adults, based on prior research (Wohlert & Smith, 1998). Because movement speed is a sensitive indicator of preclinical bulbar dysfunction (Green, Yunusova, et al., 2013; Kuruvilla, Green, Yunusova, & Hanford, 2012; Yunusova et al., 2010), significant between-age group differences in tongue speed were also expected.
The second aim of the current study was to investigate whether speech intelligibility changes would be observed in healthy older adults and if so, to investigate an exploratory model of intelligibility for healthy aging based on phonatory and articulatory measures. Hypotheses about subsystem contributions to intelligibility variance were based on existing dysarthria studies that used a similar multi-subsystem approach (Lee et al., 2014; Rong et al., 2016) as well as the extant aging literature that showed age group differences for phonatory and articulatory measures. Among phonatory measures, acoustic measures were expected to contribute significantly to speech intelligibility variance, specifically mean CPP (Awan et al., 2015; Ishikawa et al., 2018). For articulatory kinematics, movement speed was expected to be strongly associated with intelligibility (Green, Yunusova, et al., 2013).
2. Method
2.1. Participants
2.1.1. Speakers.
Fifteen healthy, older adults (11 women, 4 men) and fifteen healthy, younger adults (10 women, 5 men) participated in the study. The goal was to focus on age effects rather than sex-related effects among the elderly. Therefore, despite there being more females than males in both groups, the near-equal sex distribution across the two age groups was considered appropriate for the purposes of the study. The mean age of the older adults (OA) was 67.9 years (SD = 6.7, age range: 55–81 years) and the mean age of the younger adults (YA) was 24.4 years (SD = 4.1, age range: 20–35 years). In the OA group, only one individual was above 75 years of age and therefore, the age range for both the OA and YA groups was almost similar, i.e., 15-year range for both groups. Fifty-five years was selected as the lower age cut-off for the OA group because structural and physiologic changes to speech subsystems can begin as early as 50 years (Hixon et al., 2014). All participants were native speakers of English who had no history of voice, speech, language, or neurological disorders, and had no metal implants in the head and/or upper body, per self-report. The OA group had a mean Voice Handicap Index (VHI; Jacobson et al., 1997) score of 5.8 (SD = 5.7, range: 0–17) and the YA group a mean score of 8.6 (SD = 13.6, range: 0–52), which is within the normal range. The YA group included one extreme outlier (VHI = 52 > 3 SD), however, the individual did not report a voice disorder and voice quality was perceived to be within normal limits. When the outlier was removed, the YA group demonstrated a mean VHI score of 5.5 (SD = 6.6, range: 0–24). Thus, all individuals fell within the normal range apart from the outlier. As a reference, patients with mild voice disorders had a mean VHI score of 33.7 (SD = 5.6) in the validation study (Jacobson et al., 1997).
All older adults underwent a hearing screening at 1, 2, and 4 kHz in a laboratory setting (Maico, model MA40, Eden Prairie, MN). Majority of the older adults were able to detect pure tones at 35 dB HL at the tested frequencies, except for two participants who were only able to detect 4 kHz pure tones at 40 and 50 dB HL, respectively, in one ear. Further, one participant was only able to detect 1 and 2 kHz pure tones at 40 dB HL in one ear. No hearing screening was completed for the younger adults, however, these participants had normal hearing per self-report and were able to follow instructions and conversation at normal loudness levels with no signs of hearing problems. Participants who scored below 26 on the Montreal Cognitive Assessment (Nasreddine et al., 2005) were excluded. The study was approved by the University of Missouri Institutional Review Board. All speakers provided written consent and were compensated for their participation.
2.1.2. Listeners.
Five students at the University of Missouri who were unfamiliar with the test materials and speakers transcribed the sentence intelligibility samples. The inclusion criteria for listeners were: (i) be a native speaker of American English, (ii) be between 18 and 30 years of age, (iii) have normal hearing at 25 dB HL at 1, 2, and 4 kHz bilaterally, and (iv) have no history of speech, language, learning, and/or cognitive disabilities based on self-report. Listeners were compensated for their time.
2.2. Experimental Stimuli
The same set of connected speech and single word stimuli were used for aerodynamic, acoustic, and kinematic data collection (Table 1). To capture tongue movement along the vertical axis, four target words (e.g., ‘cake’ and ‘tell’) comprising alveolar (e.g., /t/) and velar (e.g., /k/) consonants were selected from the Multiple Word Intelligibility Test (Kent, Weismer, Kent, & Rosenbek, 1989). These monosyllabic words were embedded into a carrier phrase (e.g., “Say ____ again”) to ensure naturalness of production and minimize variability of tongue positions at word onset and offset. Each target word, along with two foil words, was repeated 10 times.
Table 1.
Protocols and Outcome Measures for Each Speech Subsystem Obtained from Each Stimulus Type
| Stimulus Type | Task | Subsystem | Signal | Outcome Measure |
|---|---|---|---|---|
| Connected speech | Harvard sentences:
|
Phonatory | Aerodynamic | Mean airflow during voicing (L/s), intensity (dB) |
| Phonatory | Acoustic | CPP (M, SD in dB), L/H ratio (M, SD in dB) | ||
| Articulatory | Kinematic | Speed (mm/s), distance (mm), duration (s), STI | ||
ADSV sentences:
|
Phonatory | Acoustic | CPP (M, SD in dB), L/H ratio (M, SD in dB), CSID | |
| Single words | Multiple Word Intelligibility Test:
|
Phonatory | Aerodynamic | Mean airflow during voicing (L/s), intensity (dB) |
| Phonatory | Acoustic | CPP (M, SD in dB), L/H ratio (M, SD in dB) | ||
| Articulatory | Kinematic | Speed (mm/s), distance (mm), duration (s), STI | ||
| Syllable | PAS protocol: voicing efficiency, /pi/ | Phonatory | Aerodynamic | Mean peak air pressure (cm H2O), mean airflow during voicing (L/s), airway resistance (air pressure/airflow), intensity (dB) |
| Vowel | PAS protocol: maximum sustained phonation, /a/ | Phonatory | Aerodynamic | Duration (s) |
Note. ADSV = Analysis of Dysphonia in Speech and Voice, CPP = cepstral peak prominence, L/H ratio = low/high spectral ratio, CSID = Cepstral Spectral Index of Dysphonia, PAS = Phonatory Aerodynamic System, STI = spatiotemporal variability index.
Similar to the word stimuli, three target sentences comprised predominantly of alveolar and velar consonants were selected from the Harvard sentences (Rothauser et al., 1969). The three target sentences were produced 10 times along with 60 foil sentences wherein two different foil sentences were presented in a list with each repetition of a target sentence. Stimulus order was pseudo-randomized such that words and sentences were presented alternatingly and different target stimuli were presented consecutively. Certain speech stimuli were used only for aerodynamic (maximum phonation time, laryngeal efficiency measures) and acoustic (cepstral/spectral) data collection as part of recommended clinical protocols for instrumental voice assessment (Patel et al., 2018) (Table 1).
Sentences used to estimate speech intelligibility were taken from the computerized Speech Intelligibility Test (SIT; Yorkston, Beukelman, Hakel, & Dorsey, 2007). Each participant produced 11 sentences ranging in length from 5–15 words that were randomly generated by the SIT software. The majority of the sentences were semantically unpredictable and they differed from speaker to speaker in order to reduce listener bias when completing intelligibility transcriptions. In addition to the SIT sentences, the four CAPE-V sentences and the second sentence of the Rainbow Passage were also included so that intelligibility estimates could be obtained from some of the same stimuli used in the experimental tasks.
2.3. Data Acquisition and Data Analysis
Each participant attended two sessions within one week of the other. For all participants, the first session was used to obtain phonatory aerodynamic and acoustic data, as well as scores on the VHI. The second session was used to record articulatory kinematics as well as speech intelligibility samples.
2.3.1. Phonatory-aerodynamic.
Phonatory data acquisition took place in a sound-attenuating booth (IAC Acoustics, North Aurora, IL). Aerodynamic data were collected during sustained vowel, syllable, single word, and sentence productions using the Phonatory Aerodynamic System (PAS, Model 6600; PENTAX Medical, Montvale, NJ). To assess maximum phonation time (MPT), participants took a deep breath, placed the face mask firmly on their face, followed by a sustained /a/ for as long as possible. This task was completed thrice. To assess laryngeal efficiency, participants took a breath, placed the mask firmly on their face, and produced the syllable /pi/ five times in a row on one breath at 90 bpm with comfortable pitch and loudness. The syllable string was repeated three separate times in one recording.
Aerodynamic data for the phonatory system were analyzed through the PAS protocols Maximum Sustained Phonation (MPT), Voicing Efficiency, and Running Speech. For MPT, the measurement was automated once markers were manually placed around the voiced section of the signal. The best value of three trials was used for analysis. For laryngeal efficiency measures, the middle three /pi/ productions out of five /pi/ syllables from each of the three sets that met quality criteria were averaged. Quality criteria included airflow at zero during plosive productions, air pressure returning to zero during vowel productions, and slanted or flat pressure peaks produced in a similar manner and intensity (Solomon, 2011; Solomon & Helou, 2013). The measures included mean peak air pressure (estimate of subglottal pressure via intraoral tube during /p/ production), mean airflow during voicing (during /i/ production or speech), and airway resistance (air pressure/airflow). If quality criteria were not met, then a minimum of five acceptable syllable productions was used. Markers were manually placed to identify air pressure peaks and airflows after which trials were automatically analyzed and averaged. For airflow measurements during reading of Harvard sentences, markers were manually placed around the spoken section of the signal and then the measurement was automated. The airflow values from the first three repetitions were averaged. Air pressure data cannot be collected with the PAS Running Speech protocol, therefore, air pressure data were not collected from the Harvard sentences.
2.3.2. Phonatory-acoustic.
Acoustic data were collected during single word (Multiple Word Intelligibility Test) and sentence productions split into Harvard, CAPE-V (Kempster, Gerratt, Verdolini Abbott, Barkmeier-Kraemer, & Hillman, 2009) and Rainbow Passage sentences (Fairbanks, 1960) using the Analysis of Dysphonia in Speech and Voice (ADSV; Awan, 2011) program as part of the Computerized Speech Lab (CSL, Model 4500, PENTAX Medical, Montvale, NJ). A headset microphone (AKG C520, Vienna, Austria) was distanced 4 cm from the corner of the participants’ mouth. The ADSV sampling rate for audio signals was 25 kHz.
Cepstral/spectral data from connected speech productions were analyzed using the ADSV program. The measurement was automated once markers for voice onset (0.05 seconds before signal onset) and offset were manually placed for analysis. CSID values were only available for ADSV tasks, either automatically (CAPE-V sentences) or manually for the second and third sentences of the Rainbow Passage using the formula provided in Awan, Roy, and Dromey (2009). Of note, while the CSID formula for sentences includes means and SDs for both CPP and L/H spectral ratio, the CSID formula for the Rainbow Passage does not include CPP SD. In addition, cepstral/spectral measures from the first three repetitions of Harvard sentences were averaged. If the first three productions did not meet measurement criteria (i.e., the voice onsets or offsets were not clearly separable from productions of surrounding target sentences), then successive productions of that target, which met measurement criteria, were used instead.
2.3.3. Articulatory-kinematic.
Tongue and jaw kinematic data were collected during single word (Multiple Word Intelligibility Test) and sentence productions (Harvard sentences) using a 3D electromagnetic articulograph (Wave Speech Research System, NDI, Waterloo, Canada). Orofacial sensors were attached along the mid-sagittal plane to the tongue tip and tongue back. The first sensor was placed 1 cm from the tongue tip and the second sensor was placed 4 cm from the tongue tip. Two sensors were also attached to the mandibular gingiva under the lateral incisors on each side. All of the 5DOF (degrees of freedom) sensors placed in and around the mouth were attached using a non-toxic dental adhesive (PeriAcryl®90, Glustitch, Delta, Canada). Movement from each orofacial sensor was expressed relative to local x, y, and z coordinates determined by a 6DOF head reference sensor (Green, Wang, & Wilson, 2013). The head sensor was attached to an adjustable headband to avoid skin motion artifacts (Green, Wang, et al., 2013). Movement data were collected at a sampling rate of 400 Hz. Audio signals were recorded at a sampling rate of 22 kHz using a solid state recorder (Marantz PMD670, Eindhoven, Netherlands) and high quality condenser microphone (Shure PG42, Niles, IL) placed 20 cm from each participant’s mouth. Tongue and jaw movements were not decoupled because our goal was to examine aging effects on naturally occurring coupled tongue-jaw movements during speech; therefore, tongue movements included contributions of the jaw.
To estimate articulatory kinematics, each target word and sentence was parsed using SMASH (Green, Wang, et al., 2013), which is a custom written MATLAB tool (The MathWorks, 2012b, Natick, MA). Onsets and offsets were determined using vertical displacement time-histories corresponding to the primary place of articulation for consonants at the beginning and end of each target stimulus. For example, onset and offset for the word ‘cake’ was based on vocal tract constriction for the word initial and final consonants, which coincides with the peak displacement of the tongue back sensor.
Maximum speed (mm/second) was calculated as the maximum value of the first-order derivative of the vertical position time-history for each sensor. Only the maximum value between the initial release and the final constriction was chosen for analysis because of our interest in the speed generating capacity of talkers regardless of individual opening and closing movements within the utterance. Average movement speed was calculated as the first-order derivative of each sensor’s distance time-history. Distance was calculated as the straight-line path between movement onset and offset for each sensor and each parsed utterance. Duration was calculated as the time in seconds between movement onset and offset for each utterance. To calculate the spatiotemporal variability index (STI), the vertical tongue displacement trajectory for each word and sentence repetition was time and amplitude normalized. Standard deviations (SD) were then calculated for the normalized displacement trajectories from all repetitions of an utterance at fixed 2% intervals to estimate STI, which is the sum of 50 SD (Smith et al., 1995).
2.3.4. Intelligibility.
SIT sentences were recorded using the same microphone and solid state recorder described above for articulatory data collection. From each speaker, five (out of 11) SIT sentences were randomly selected for the sentence transcription task. In addition, the four CAPE-V sentences and the second sentence of the Rainbow Passage were also used to estimate intelligibility. While the five SIT sentences were different for each speaker, the four CAPE-V sentences and the second sentence of the Rainbow Passage were identical across all the speakers. In sum, 10 sentences from each speaker were transcribed by five listeners to obtain average intelligibility scores along with interrater reliability, which was used to determine consistency across the listeners.
For the listening experiment, a modified multi-talker babble protocol was used (Ishikawa et al., 2018). Speech was mixed with freely available multi-talker babble (https://people.kth.se/~e99_ehe/project.html) using Praat version 6.0.30 (Boersma & Weenink, 2019). First, the multi-talker babble was resampled to match the sampling rate of the speech files. Then, both the multi-talker babble and speech recordings were normalized to 60 dB SPL using a Praat script (Yoon) with a mixing result of approximately 63 dB SPL. Intensity was normalized across speech recordings to focus the listening experiment on potential group differences in voice quality and not audibility. As the final step, babble was added to the speech files with the signal-to-noise ratio (SNR) set to +1 dB using a Praat script (McCloy). An SNR of +1 dB was considered optimal during pilot testing to avoid ceiling and floor effects.
For the transcription task, sentences from all speakers were pooled and the listening order was pseudo-randomized such that sentences from the same speaker or same age group were not presented consecutively. Further, the authors ensured that the same CAPE-V and Rainbow Passage sentences were not presented consecutively. Listeners were blind to the participant’s group and heard the speech samples in a sound-field within a sound-attenuating booth (IAC Acoustics, North Aurora, IL) at 50 dB HL, that is, at conversational level (65 dB SPL). Samples were played via speakers connected to a CD-player (Tascam CD-200, Montebello, CA) and an audiometer (Audiostar pro, GSI Grason-Stadler, Eden Prairie, MN). In accordance with SIT instructions, listeners were allowed to listen to each sentence no more than two times and were asked to transcribe exactly what they heard. The sentence transcription task took approximately 1.5–2 hours and was self-paced. A trained graduate student calculated speech intelligibility as the percentage of words correctly transcribed out of the total number of words attempted. For each speaker, the percent speech intelligibility score was averaged across the five listeners. For intelligibility estimates, misspellings and homonyms as well as tense and capitalization errors were accepted as correct.
2.4. Statistical Analysis
To investigate age-related changes in phonatory and articulatory subsystems repeated measures analyses of covariance (ANCOVA) were used. Assumptions for ANCOVA were tested using the Brown-Forsythe test (α = .05) for homogeneity of variance for between-subjects ANCOVAs and both Box’s M (α = .001) and Mauchly’s test of sphericity (α = .05) for homogeneity of variance and covariance (compound symmetry) for any mixed ANCOVA. The Huynh-Feldt adjustment was used if either test for compound symmetry was significant. Normality was tested using the Shapiro-Wilk test (α = .05), but no actions were taken because ANCOVA is robust against violations (Glass & Hopkins, 1996). Extreme outliers (> 3 SD) were identified and sex was used as a covariate. The statistical results are reported along with the effect size partial η2 (ηp2), where 0.01 may be interpreted as small, 0.06 as medium, and 0.14 as large (Cohen, 1988). The significance level was set at α = .05. All analyses were performed using SPSS version 25 (IBM SPSS, Armonk, NY).
To investigate an exploratory model of intelligibility for healthy aging based on phonatory and articulatory measures, a generalized estimating equation (GEE) model (Hardin & Hilbe, 2012) was fit for data from the older adults. For the phonatory system, maximum phonation time was excluded from the model, to prioritize laryngeal efficiency measures based on airflow and air pressure as well as cepstral/spectral measures because these are recommended clinical tasks in voice clinics (Patel et al., 2018). Similarly, for the articulatory system, STI was excluded to prioritize individual spatial and temporal kinematic parameters over STI, which is a composite index of spatiotemporal movement.
Before fitting the model, multicollinearity among the independent variables from the two subsystems was determined by examining the variance inflation factor (VIF). VIF values below 10 indicate that the multiple regression assumption regarding multicollinearity is not violated (Cohen, Cohen, West, & Aiken, 2003). Of the 11 measures across the two subsystems (i.e., eight phonatory and three articulatory measures), Psub, airflow, maximum speed, and distance had VIF values greater than 10. Between the two phonatory variables, VIF of airflow was greater than that of Psub and between the two articulatory variables, VIF for distance was greater than that of maximum speed. Therefore, both airflow and distance were removed and when VIF values were reexamined, all remaining variables had VIF values below 4. The final step before fitting the full (initial) model was to remove variables with missing data, which resulted in the exclusion of Psub and Rlaw.
To take into account the dependency of observations (intelligibility percentages from five raters on the same person), a GEE population-average model with an exchangeable working correlation structure was fit using 375 observations (i.e., 15 speakers × 5 stimuli × 5 raters). In general, 10 observations per predictor are considered sufficient for acceptable error (Harrell, Lee, Califf, Pryor, & Rosati, 1984); more recent research states that 2–3 observations per predictor are sufficient for acceptable error (Austin & Steyerberg, 2015). The GEE regression parameters can be interpreted similarly to regression models, but because GEE models are quasi-likelihood based, traditional goodness of fit measures such as r-squared (R2) are not available. Instead, the quasi-likelihood criteria (QIC) and (QICu) are used as goodness of fit-like measures i.e., if the QIC and QICu are close, the model is appropriately specified. Details about model selection and evaluation such as the structure of the full and final models (i.e., independent variables, dependent variable, covariate) and the strength of the models (indicated by QIC and QICu values) are provided in the results section. The significance level was set at α = .05. All analyses were performed using SAS version 9.4 (SAS Institute, Cary, NC).
3. Results
Phonatory subsystem results do not include word-level data in order to focus on tasks for which norms exist such as syllable (laryngeal efficiency measures) and sentence (cepstral/spectral analyses) tasks. Moreover, the sentence-level tasks are more ecologically valid than word-level tasks. For the articulatory subsystem, both word- and sentence-level results are included.
3.1. Age-related Subsystem Differences
3.1.1. Phonatory-aerodynamic.
A one-way between-subjects ANCOVA was performed on the dependent variable (DV): MPT as a function of the independent variable (IV): group (OA vs. YA). The mean was 18.43 s (SD = 8.36) for the OA group and 26.18 s (SD = 8.69) for the YA group. There was a significant main effect of group [F(1, 27) = 10.73, p = .003, ηp2 = .28], wherein the OA group had shorter MPTs than the YA group.
A multivariate analysis of covariance (MANCOVA) was performed on the DVs: mean peak air pressure, mean airflow during voicing, and airway resistance as a function of the IV: group (OA vs. YA). Data from four participants in the OA group (one male, four females) were excluded because the data did not meet quality criteria (disconnected syllable production, vowel not well sustained) and one data set was missing due to experiment error. There were no significant group differences for any of the measures (Table 2). The descriptive data are presented in Figure 1. There was one extreme outlier for airway resistance in the YA group (>3 SD; not same as VHI outlier) and excluding the outlier did not change results. Mean SPL during voicing was 79.8 dB (SD = 4.4) for the OA group and 79.6 dB (SD = 4.7) for the YA group.
Table 2.
Repeated Measures Multivariate (Syllable) and Univariate (Harvard Sentences) ANCOVA for Phonatory Aerodynamic Measures
| Stimulus Type | Effect | Mean Airflow During Voicing (L/s) | Mean Peak Air Pressurec (cm H2O) | Airway Resistancec (cm H2O/[L/s]) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| df | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | ||||
| Syllable /pi/ | Groupa | 1 | 0.86 | .04 | .363 | 0.38 | .02 | .546 | 1.61 | .07 | .217 | ||
| Connected Speech (Harvard Sentences) | Groupb | 1 | 7.17 | .21 | .012 | ||||||||
| Sentence | 2 | 10.65 | .28 | < .001 | |||||||||
| Grp*Sent | 2 | 2.03 | .07 | .142 | |||||||||
Note. Grp = group, Sent = sentence, df = degrees of freedom, F = F statistic, p = significance level, ηp2 = partial eta squared.
Bold text= p < .05.
n = 15 for the older group and n = 15 for the younger adult group.
n = 10 for the older group and n = 15 for the younger adult group because data from five participants did not meet quality criteria for analysis.
Harvard sentences were recorded with the Phonatory Aerodynamic System’s Running Speech protocol, for which mean peak air pressure and airway resistance measures are not available.
Figure 1.

A. Mean airflow during voicing (L/s), B. mean peak air pressure (cm H2O), and C. airway resistance (cm H2O/[L/s]) in older and younger adults during production of the laryngeal efficiency task consisting of repetitions of /pi/ using the PENTAX Medical Phonatory Aerodynamic System. Error bars represent standard errors of the mean. For the older adults, four datasets did not meet quality criteria and another one was missing.
A 3×2 two-way mixed ANCOVA was performed on the DV: mean airflow during voicing as a function of the between-subjects factor: group (IV: OA vs. YA) and the within-subjects factor: stimuli (IV: Harvard sentences). While there was not a significant main effect of group on any of the laryngeal efficiency measures using the syllable task, there was one on airflow for Harvard sentences (p = .012, ηp2 = .21) as shown in Table 2. The OA group had greater airflow values than the YA group during sentence reading (Table 3). Mean SPL during voicing averaged across the sentences was 74.0 dB (SD = 2.4) for the OA group and 73.0 dB (SD = 3.2) for the YA group.
Table 3.
Descriptive Statistics for Phonatory Aerodynamic and Acoustic Variables Derived from Connected Speech (Harvard Sentences)
| Task | Groupsa | Airflow During Voicing (L/s)b | CPP M (dB) | CPP SD (dB) | L/H M (dB) | L/H SD (dB) | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean (SD) | Mean (SD) | Mean (SD) | Mean (SD) | Mean (SD) | ||||||
| Sentence 1 | OA | 0.173 (0.057) | 5.11 (0.95) | 3.69 (0.42) | 29.79 (2.23) | 13.94 (1.49) | ||||
| YA | 0.133 (0.050) | 5.13 (1.04) | 3.72 (0.47) | 29.13 (2.79) | 13.88 (1.68) | |||||
| Sentence 2 | OA | 0.177 (0.055) | 5.61 (0.96) | 4.07 (0.54) | 30.02 (2.85) | 16.28 (1.83) | ||||
| YA | 0.143 (0.055) | 6.05 (0.83) | 4.23 (0.33) | 29.23 (2.49) | 16.18 (1.84) | |||||
| Sentence 3 | OA | 0.211 (0.062) | 4.05 (0.87) | 3.58 (0.58) | 27.07 (2.48) | 15.71 (1.95) | ||||
| YA | 0.161 (0.065) | 4.34 (0.99) | 3.72 (0.43) | 26.64 (2.66) | 15.82 (2.03) |
Note. OA = older adults, YA = younger adults, CPP = cepstral peak prominence, L/H ratio = low/high spectral ratio, Sentence 1 = Cats and dogs each hate the other, Sentence 2 = The grass curled around the fence post, Sentence 3 = The cup cracked and spilled its contents.
n = 15 for the older group and n = 15 for the younger adult group.
Airflow during voicing was recorded with the Phonatory Aerodynamic System’s Running Speech protocol, for which mean peak air pressure and airway resistance measures are not available.
3.1.2. Phonatory-acoustic.
A two-way mixed MANCOVA was performed on the acoustic DVs: CPP (M, SD), L/H spectral ratio (M, SD), and CSID as a function of the between-subjects factor: age (IV: OA vs. YA) and the within-subjects factor: stimuli (IV: five ADSV sentences). Another two-way mixed MANCOVA was performed on the acoustic DVs: CPP (M, SD) and L/H spectral ratio (M, SD) as a function of the between-subjects factor: age (IV: OA vs. YA) and the within-subjects factor: stimuli (IV: three Harvard sentences). Separate MANCOVAs were carried out for ADSV and Harvard sentences because the CSID DV only existed for ADSV but not Harvard sentences. Significant main effects of group were found on mean CPP (p = .041, ηp2 = .15) and CPP SD (p = .033, ηp2 = .16) across ADSV sentences (Table 4), wherein the OA group had lower CPP mean and SD values than the YA group (Figure 2). There was no significant main effect of group on acoustic measures for Harvard sentences (Table 4). The descriptive data for Harvard sentences are presented in Table 3.
Table 4.
Two-Way Mixed MANCOVAs for Phonatory Acoustic Measures
| Stimulus Type | Effect | CPP M (dB) | CPP SD (dB) | L/H M (dB) | L/H SD (dB) | CSIDb | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| df | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | ||||||
| Connected Speech (ADSV) | Groupa | 1 | 4.59 | .15 | .041 | 5.02 | .16 | .033 | 0.54 | .02 | .468 | 1.33 | .05 | .259 | 3.14 | .10 | .088 | ||||
| Sentence | 4 | 8.46 | .24 | < .001 | 2.00 | .07 | .099 | 5.19 | .16 | .001 | 1.26 | .04 | .292 | 3.28 | .11 | .014 | |||||
| Grp*Sent | 4 | 0.96 | .03 | .435 | 0.65 | .02 | .630 | 1.16 | .04 | .332 | 2.02 | .07 | .096 | 0.55 | .02 | .700 | |||||
| Connected Speech (Harvard Sentences) | Group | 1 | 0.62 | .02 | .438 | 0.69 | .03 | .412 | 1.18 | .04 | .287 | 0.00 | .00 | .983 | |||||||
| Sentence | 2 | 5.18 | .16 | .009 | 1.21 | .04 | .306 | 8.46 | .24 | .001 | 2.49 | .09 | .092 | ||||||||
| Grp*Sent | 2 | 1.51 | .05 | .229 | 0.40 | .02 | .672 | 0.34 | .01 | .712 | 0.19 | .01 | .824 | ||||||||
Note. Grp = group, Sent = sentence, CPP = cepstral peak prominence, L/H = low to high spectral ratio, CSID = Cepstral Spectral Index of Dysphonia, df = degrees of freedom, F = F statistic, p = significance level, ηp2 = partial eta squared.
Bold text= p < .05.
n = 15 for the older group and n = 15 for the younger adult group.
A CSID formula is not available for Harvard sentences.
Figure 2.

A. Cepstral and B. spectral measures as well as C. Cepstral Spectral Index of Dysphonia (CSID) values for older (OA) and younger (YA) adults during the production of reading tasks from the Analysis of Dysphonia in Speech and Voice program (ADSV, PENTAX Medical). Error bars represent standard errors of the mean. CPP = cepstral peak prominence, L/H = low/high spectral ratio.
3.1.3. Articulatory-kinematic.
A three-way mixed MANCOVA was performed on the kinematic DVs: maximum speed, average speed, distance, duration, and STI as a function of the between-subjects factor: group (IV: OA vs. YA) and two within-subjects factors: articulator (IV: jaw, tongue tip, tongue back) and stimuli (IV: Multiple Word Intelligibility Test words or Harvard sentences). Separate MANCOVAs were carried out for words and sentences, and sex was used as a covariate in both analyses. Descriptive data for words and sentences are in Tables 5 and 6, respectively. Main and interaction effects for words and sentences are in Table 7. Pairwise comparisons were carried out for significant interaction effects, adjusting the family-wise Type I error rate to .05 using the Sidak correction. Posthoc results are provided below.
Table 5.
Descriptive Statistics for Articulatory Kinematic Variables Derived from Target Words
| Factors | Levels | Groupsa | Max Speed (mm/sec) |
Average Speed (mm/sec) |
Distance (mm) |
Duration (sec) |
STI | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean (SE) | Mean (SE) | Mean (SE) | Mean (SE) | Mean (SE) | |||||||
| Articulator | Jaw | OA | 48.50 (4.02) | 18.07 (1.32) | 7.25 (1.22) | 0.41 (0.06) | 14.57 (1.22) | ||||
| YA | 53.73 (4.16) | 19.74 (1.37) | 8.38 (1.26) | 0.43 (0.06) | 12.19 (1.22) | ||||||
| Anterior Tongue | OA | 79.30 (6.35) | 31.07 (2.10) | 12.06 (2.06) | 0.41 (0.06) | 17.60 (1.37) | |||||
| YA | 90.71 (6.57) | 33.33 (2.17) | 14.04 (2.13) | 0.43 (0.06) | 16.74 (1.37) | ||||||
| Posterior Tongue | OA | 76.50 (4.02) | 30.31 (1.66) | 12.00 (1.82) | 0.41 (0.06) | 18.44 (1.23) | |||||
| YA | 85.19 (4.16) | 32.84 (1.72) | 13.68 (1.89) | 0.43 (0.06) | 16.46 (1.23) | ||||||
| Stimuli | Word 1 | OA | 47.98 (4.60) | 18.99 (1.24) | 7.29 (1.54) | 0.38 (0.06) | 18.43 (1.62) | ||||
| YA | 50.09 (4.76) | 18.88 (1.29) | 8.27 (1.59) | 0.41 (0.06) | 15.33 (1.61) | ||||||
| Word 2 | OA | 83.66 (4.97) | 30.44 (1.86) | 12.76 (1.92) | 0.44 (0.06) | 16.45 (0.95) | |||||
| YA | 103.78 (5.14) | 36.89 (1.92) | 15.72 (1.99) | 0.45 (0.06) | 14.26 (0.95) | ||||||
| Word 3 | OA | 81.04 (5.15) | 34.75 (2.57) | 12.39 (1.77) | 0.37 (0.06) | 13.89 (1.63) | |||||
| YA | 88.74 (5.33) | 36.06 (2.66) | 13.47 (1.83) | 0.40 (0.06) | 12.53 (1.63) | ||||||
| Word 4 | OA | 59.74 (5.46) | 21.76 (1.65) | 9.30 (1.64) | 0.44 (0.06) | 18.71 (1.43) | |||||
| YA | 63.56 (5.65) | 22.73 (1.71) | 10.68 (1.70) | 0.46 (0.06) | 18.40 (1.43) |
Note. OA = older adults, YA = younger adults, SE = standard error, Word 1 = Ache, Word 2 = Tell, Word 3 = Ate, Word 4 = Cake.
n = 15 for the older group and n = 14 for the younger adult group.
Table 6.
Descriptive Statistics for Articulatory Kinematic Variables Derived from Target Sentences
| Factors | Levels | Groupsa | Max Speed (mm/sec) | Average Speed (mm/sec) | Distance (mm) |
Duration (sec) |
STI | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean (SE) | Mean (SE) | Mean (SE) | Mean (SE) | Mean (SE) | |||||||
| Articulator | Jaw | OA | 85.77 (5.71) | 19.48 (1.30) | 48.35 (3.79) | 3.44 (0.72) | 24.40 (1.07) | ||||
| YA | 75.92 (7.29) | 17.51 (1.66) | 38.30 (4.84) | 2.19 (0.92) | 25.53 (1.07) | ||||||
| Anterior Tongue | OA | 155.60 (7.92) | 32.94 (1.54) | 80.43 (3.85) | 3.43 (0.72) | 24.07 (1.03) | |||||
| YA | 192.83 (10.11) | 38.51 (1.97) | 83.33 (4.91) | 2.19 (0.92) | 24.66 (1.03) | ||||||
| Posterior Tongue | OA | 152.15 (9.02) | 34.25 (1.37) | 83.94 (4.03) | 3.43 (0.72) | 23.91 (1.02) | |||||
| YA | 187.78 (11.52) | 39.80 (1.75) | 86.51 (5.14) | 2.16 (0.92) | 23.07 (1.02) | ||||||
| Stimuli | Sentence 1 | OA | 132.52 (6.62) | 29.64(1.20) | 67.57 (3.45) | 5.23 (2.19) | 22.23 (1.10) | ||||
| YA | 142.46 (8.46) | 30.95 (1.53) | 62.20 (4.40) | 2.02 (2.80) | 22.61 (1.10) | ||||||
| Sentence 2 | OA | 126.07 (6.33) | 28.62 (1.28) | 68.19 (3.53) | 2.37 (0.06) | 24.55 (1.41) | |||||
| YA | 162.36 (8.07) | 33.80 (1.63) | 72.02 (4.50) | 2.14 (0.08) | 25.75 (1.41) | ||||||
| Sentence 3 | OA | 134.94 (6.69) | 28.41 (1.01) | 76.96 (3.52) | 2.70 (0.09) | 25.59 (1.15) | |||||
| YA | 151.71 (8.54) | 31.07 (1.29) | 73.91 (4.50) | 2.39 (0.11) | 24.90 (1.15) |
Note. OA = older adults, YA = younger adults, SE = standard error, Sentence 1 = Cats and dogs each hate the other, Sentence 2 = The grass curled around the fence post, Sentence 3 = The cup cracked and spilled its contents.
n = 15 for the older group and n = 14 for the younger adult group.
Table 7.
Three-Way Mixed MANCOVAs for Articulatory Kinematic Measures
| Stimulus Type | Effect | Max Speed (mm/sec) | Average Speed (mm/sec) |
Distance (mm) |
Duration (sec) |
STI | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| df | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | F | ηp2 | p | ||||||
| Words | Groupa | 1 | 1.87 | .07 | .183 | 1.01 | .04 | .325 | 0.42 | .02 | .523 | .08 | .01 | .777 | 1.21 | .05 | .282 | ||||
| Articulator | 2 | 8.36 | .24 | .001 | 12.97 | .33 | .000 | 1.57 | .06 | .219 | 1.83 | .07 | .171 | 2.58 | .09 | .086 | |||||
| Word | 3 | 8.90 | .26 | .000 | 8.26 | .24 | .000 | 6.70 | .21 | .000 | 8.37 | .24 | .000 | 5.77 | .19 | .001 | |||||
| Grp*Art | 2 | 0.52 | .02 | .600 | 0.09 | .01 | .914 | 0.29 | .01 | .748 | 1.88 | .07 | .162 | 0.65 | .03 | .528 | |||||
| Grp*Word | 3 | 3.27 | .11 | .025 | 2.39 | .08 | .075 | 1.87 | .07 | .141 | 0.30 | .01 | .829 | 0.69 | .03 | .562 | |||||
| Grp*Art*Word | 6 | 0.35 | .01 | .912 | 0.20 | .01 | .978 | 0.34 | .01 | .914 | 1.37 | .05 | .230 | 1.25 | .05 | .284 | |||||
| Sentences | Groupa | 1 | 5.43 | .20 | .029 | 2.67 | .11 | .117 | 0.14 | .01 | .709 | 1.38 | .06 | .253 | 0.05 | .01 | .822 | ||||
| Articulator | 2 | 7.82 | .26 | .001 | 7.49 | .25 | .002 | 7.24 | .25 | .002 | 0.35 | .02 | .705 | 2.08 | .08 | .135 | |||||
| Sentence | 2 | 0.26 | .01 | .770 | 1.91 | .08 | .160 | 1.03 | .05 | .367 | 0.67 | .03 | .519 | 2.67 | .10 | .079 | |||||
| Grp*Art | 2 | 5.78 | .21 | .006 | 4.70 | .18 | .014 | 2.71 | .11 | .078 | 1.27 | .06 | .290 | 2.10 | .08 | .133 | |||||
| Grp*Sent | 2 | 3.19 | .13 | .051 | 3.70 | .14 | .033 | 4.10 | .16 | .023 | 0.83 | .04 | .444 | 0.52 | .02 | .596 | |||||
| Grp*Art*Sent | 4 | 5.95 | .21 | .000 | 3.94 | .15 | .005 | 3.33 | .13 | .014 | 1.61 | .07 | .180 | 1.26 | .05 | .292 | |||||
Note. Grp = group, Art = articulator, Sent = sentence, df = degrees of freedom, F = F statistic, p = significance level, ηp2 = partial eta squared, STI = spatiotemporal variability index. Bold text= p < .05.
n = 15 for the older group and n = 14 for the younger adult group.
For maximum speed, pairwise comparisons of the significant group × stimuli interaction (p = .025, ηp2 = .11) revealed significantly reduced maximum tongue movement speed for older adults relative to younger adults for the word ‘tell’ (p = .009, ηp2 = .24). For sentences, pairwise comparisons of the significant group × articulator × stimuli interaction (p < .001, ηp2 = .21) revealed significantly reduced maximum speed for the older compared to younger adults for both the anterior (p = .001, ηp2 = .42) and posterior tongue (p = .008, ηp2 = .28) for the sentence ‘The grass curled around the fence post.’
The sentence-level results for average speed were similar to those for maximum speed. Specifically, pairwise comparisons of the significant group × articulator × stimuli interaction (p = .07, ηp2 = .12), showed significantly reduced average speed values for the older adults relative to younger adults for both the anterior tongue (p = .002, ηp2 = .22) and posterior tongue (p = .004, ηp2 = .33) for the sentence ‘The grass curled around the fence post.’ No significant interaction effects were observed for word-level data.
For sentence-level distance results, post hoc analysis of the significant group × articulator × stimuli interaction (p = .014, ηp2 = .13) revealed no significant between-group differences, but only significant within-group differences, which are not detailed here as it is not within the scope of this study. Similar observations were noted for posthoc comparisons of the significant group × stimuli interaction (p = .023, ηp2 = .16). No significant interaction effects were observed for word-level distance data. Similarly, for duration and STI, no significant interaction effects were observed for either words or sentences.
3.1.4. Intermeasurer Reliability.
Intermeasurer reliability was assessed by having the second author (M.D.) analyze 20% of the samples subjected to aerodynamic (laryngeal efficiency) and acoustic (CSID, all voiced sentence) analyses using Pearson correlation. Reliability was extremely strong for airflow (r = .96) and air pressure (r = .97) and moderately strong for airway resistance (r = .72). Reliability for CSID (r = .96) was extremely strong due to the semi-automated nature of analysis.
Similarly, for the articulatory kinematic measures, Pearson correlations were used to determine intermeasurer reliability for 20% of the single word samples analyzed by two undergraduate students. Reliability was extremely strong for maximum speed (r = .95), distance (r = .99), and duration (r = .99), and moderately strong for STI (r = .77). Although these measures were extracted automatically through SMASH, the data analysts could have varied in their placement of onset and offset points, which in turn affects the measurements.
3.2. Speech Intelligibility in Healthy Older Adults
3.2.1. Age effects on speech intelligibility.
A univariate ANOVA was used to determine the effect of the IVs: age (OA vs. YA) and sentence type (SIT vs. ADSV/Rainbow sentences) on the DV: speech intelligibility. Findings revealed significant between-group differences in intelligibility [F(1, 57) = 10.36, p = .002]. A significant main effect for sentence type was also observed [F(1, 57) = 16.78, p < .001], but the group × sentence type interaction was not significant [F(1, 57) = 0.102, p = .751]. The following were the mean intelligibility scores for younger adults: M = 82.32 (SD = 10.56, range: 77.8–86.8%) and older adults: M = 71.97 (SD = 16.89, range: 67.4–76.6%).
Interrater judgments of sentence intelligibility were assessed using ICC estimates (Landers, 2015; Shrout & Fleiss, 1979) to establish the overall consistency of intelligibility scores among the five listeners. A two-way random effects model was selected because listeners were chosen from a larger population. In addition to the average ICC (agreement on average), the single ICC (agreement on a per-item basis) is also consistent with the literature on speech intelligibility: average ICC [(2,5) = .900, 95% CI = .881–.917, p < .001]; single ICC [(2,5) = .642, 95% CI = .596–.688, p < .001]. Rater reliability for the present study is comparable to values reported elsewhere for intelligibility (Tjaden, Sussman, & Wilding, 2014).
3.2.2. Subsystem measures associated with reduced speech intelligibility in older adults.
After excluding four independent variables for multicollinearity and missing data, the full (initial) multiple regression model included seven IVs: CPP, CPP SD, L/H, L/H SD, and CSID for the phonatory system and maximum speed and duration for the articulatory system. The DV: intelligibility and covariate: sex were also part of the full model (see Table 8a for parameter estimates and VIF values). Sex, CPP, and average duration were removed from the full model using stepwise backward regression because these three variables displayed nonsignificant parameter estimates, and the final (reduced) model with only significant parameter estimates was obtained (see Table 8b). For the full model the goodness of fit-like measures QIC and QICu were 449.24 and 384.00, respectively whereas for the final model, the QIC and QICu values were closer (i.e., 416.31 and 381.00, respectively), which indicates that the model is appropriately specified. Based on the final model, for every one unit decrease in CPP SD (i.e., 1 dB), intelligibility decreased by 6.40% (95% CI = 1.70, 11.10, p = .008). For every one unit increase in L/H (i.e., 1 dB), intelligibility decreased by 2.01% (95% CI = −2.95, −1.07), and, for every one unit increase in L/H SD (i.e., 1 dB), intelligibility decreased by 3.20% (95% CI = −4.81, −1.60, p < .001). Further, for every one unit increase in CSID, intelligibility decreased by 0.21% (95% CI = −0.33, −0.10, p < .001). Finally, for every one unit decrease in maximum speed (i.e., 1 mm/sec), intelligibility decreased by 0.13% (95% CI = 0.07, 0.19, p < .001).
Table 8a.
Generalized Estimating Equations Full Model for Estimating Intelligibility Using Phonatory and Articulatory Subsystem Measures for Older Adults
| Measure | Estimate (SE) | 95% Confidence Interval | p-value | Variance Inflation Factor |
|---|---|---|---|---|
| Sex | −4.47 (6.75) | (−17.69, 8.75) | .507 | 1.49 |
| CPP | −0.08 (1.12) | (−2.27, 2.11) | .945 | 2.62 |
| CPP SD | 6.26 (2.37) | (1.62, 10.90) | .008 | 1.53 |
| L/H | −1.97 (0.71) | (−3.35, −0.59) | .005 | 2.46 |
| L/H SD | −3.02 (0.99) | (−4.96, −1.07) | .002 | 2.31 |
| CSID | −0.23 (0.09) | (−0.40, −0.07) | .006 | 1.76 |
| Maximum Speed | 0.11 (0.04) | (0.04, 0.19) | .002 | 3.39 |
| Duration | 1.15 (1.81) | (−2.40, 4.69) | .526 | 3.61 |
Note. CPP = cepstral peak prominence, SD = standard deviation, L/H = low to high spectral ratio, CSID = Cepstral Spectral Index of Dysphonia. QIC and QICu values for the full model were 449.24 and 384.00, respectively.
Table 8b.
Generalized Estimating Equations Final Reduced Model for Estimating Intelligibility Using Phonatory and Articulatory Subsystem Measures for Older Adults
| Measure | Estimate (SE) | 95% Confidence Interval | p-value |
|---|---|---|---|
| CPP SD | 6.40 (2.40) | (1.70, 11.10) | .008 |
| L/H | −2.01 (0.48) | (−2.95, −1.07) | <.001 |
| L/H SD | −3.20 (0.82) | (−4.81, −1.60) | <.001 |
| CSID | −0.21 (0.06) | (−0.33, −0.10) | <.001 |
| Maximum Speed | 0.13 (0.03) | (0.07, 0.19) | <.001 |
Note. CPP = cepstral peak prominence, SD = standard deviation, L/H = low to high spectral ratio, CSID = Cepstral Spectral Index of Dysphonia. QIC and QICu values for the final model were 416.31 and 381.00, respectively. Stepwise backward regression was conducted to arrive at the final model, whereby sex, CPP, and duration were removed sequentially from the full model to determine the significant predictors of intelligibility change in healthy aging.
4. Discussion
As hypothesized for the first aim, significant differences were observed for mean CPP with OA having lower amplitudes of the cepstral peak than YA. As expected for the articulatory system, tongue movement speed was slower for OA than YA, but in contrast to our predictions, token-to-token variability in tongue movement patterns was not affected by age in the present study. Overall, most of our results are consistent with previous aging studies and further demonstrate that subsystem changes are indeed a part of the healthy aging process (Awan et al., 2015; Flanagan & Dembowski, 2002).
With regard to the second aim, older adults displayed significantly reduced speech intelligibility scores compared to younger adults. Regarding subsystem contributions to intelligibility variance in older adults, our hypothesis about mean CPP was unconfirmed. Congruent with our hypothesis for the articulatory subsystem, maximum speed emerged as a significant contributor to intelligibility variance in older adults. Some findings within this aim are discussed using relevant research on disordered populations due to the lack of research on aging in this area.
4.1. Age-related Subsystem Differences
4.1.1. Phonatory subsystem.
Older adults had shorter MPT values than younger adults in the present study. Age-related differences for this respiratory-laryngeal measure have been found in a study with older female adults by Awan (2006). Further, our lack of significant laryngeal aerodynamic findings is congruent with findings by Zraick et al. (2012). However, significantly greater airflow during voicing in OA compared with YA emerged during sentence reading, which is a more ecologically valid task than syllable productions. Taken together, these findings suggest that OA may have compromised laryngeal valving activity compared with YA leading to greater escape of air during speaking.
Mean CPP, the cepstral measure that primarily contributes to CSID (Awan, 2011), was found to be significantly lower in OA in the present study with no significant difference in the broader CSID measure. This finding is consistent with a recent aging voice study by Awan et al. (2015). The fact that Watts et al. (2015) found a significant age difference in the CSID measure might be related to sex differences. Watts et al. (2015) used only male participants while our study and Awan et al. (2015) used both females and males. CPP is a measure that represents the degree of periodicity and dominant harmonic energy carried largely by the voice fundamental frequency compared against the background noise in the signal, which would be degraded in dysphonic voices resulting in lower CPP averages and lower variability in connected speech (Watts, Awan, & Maryn, 2016). Perceptually, lower CPP values have been associated more with a breathy than rough voice quality and discriminate well between normal and dysphonic voice (Awan et al., 2009; Hillenbrand & Houde, 1996). Our acoustic finding suggests that older adults have voices with lower harmonic energy, which are breathier and do not carry as well, a finding which is consistent with literature on presbyphonia and presbylaryngis (Kendall, 2007). Thus, mean CPP may be a sensitive acoustic measure for age-related voice changes. The group mean CPPs for the Rainbow Passage (OA M = 5.26, SD = 3.66; YA, M = 5.71, SD = 3.79) were on either side of the cutoff of 5.53 dB suggested by Sauder et al. (2017) to predict voice disorder status.
4.1.2. Articulatory subsystem.
Lower maximum and average movement speed of the anterior and posterior tongue suggests that there is an aging-induced decrease in speed generating capacity as well as in speed per se, which likely contributes to slower speech in older individuals. These findings are in line with previous research by Wohlert and Smith (1998) who also reported a slower rate among older adults relative to younger adults during habitual, fast, and slow speech. Wohlert and colleagues attributed this age-related slowing to the loss of peripheral sensorimotor function due to the loss of fast twitch muscles and declining physical status (Doherty, Vandervoort, & Brown, 1993; Ramig, 1983). They also suggested that slow speech may be a compensatory strategy which allows for more accurate speech movements and consequently, improves intelligibility (Wohlert & Smith, 1998). In fact, several other researchers have suggested slow rate as a compensatory strategy to improve intelligibility because it allows talkers more time to make precise articulatory contacts, improve coordination, and facilitate speech motor control (McHenry, 2003; Weismer, Laures, Jeng, Kent, & Kent, 2000; Yorkston, Hakel, Beukelman, & Fager, 2007). Conversely, some speech motor control studies show significantly greater movement pattern variability indicative of poorer motor control during slow speech in both healthy and disordered speakers (Kleinow, Smith, & Ramig, 2001; Kuruvilla-Dugdale & Mefferd, 2017; Mefferd, Pattee, & Green, 2014; A. Smith, Goffman, Zelaznik, Ying, & McGillem, 1995; Wohlert & Smith, 1998). Our findings partly support the prior aging literature by showing that older talkers have lower tongue movement speeds.
In contrast to prior literature (Wohlert & Smith, 1998), trial-to-trial variability in articulatory movement patterns was comparable between older and younger adults in our study. One important factor likely contributing to the differing STI findings is the age of the participants, which was considerably higher in the previous study than the current one. A second factor is that different articulators (lip versus tongue) were tested in each study. In the limb literature, similar inconsistencies exist regarding gait variability, with some studies showing greater variability among the elderly (Callisaya, Blizzard, Schmidt, McGinley, & Srikanth, 2010; Grabiner, Biswas, & Grabiner, 2001) and others reporting no age effects on gait variability (Gabell & Nayak, 1984; Stolze, Friedrich, Steinauer, & Vieregge, 2000). Interestingly, when researchers tested the contribution of slower walking speed to gait variability in older adults, they observed greater variability irrespective of speed and concluded that age-induced gait variability may result more from diminished strength and flexibility (Kang & Dingwell, 2008).
Overall, it appears that aging-induced changes to the articulatory subsystem are similar to those observed during preclinical stages of dysarthria in conditions like ALS. Although this finding suggests that speed parameters may not be helpful for distinguishing between healthy and pathological aging, it also suggests that regardless of etiology, the articulatory mechanisms that contribute to mild speech changes may be similar. Moreover, it is known that the magnitude of these articulatory changes will increase as individuals transition from healthy to pathological aging and from one disease stage to the next (Kuruvilla-Dugdale & Mefferd, 2017; Kuruvilla et al., 2012).
4.2. Age Effects on Speech Intelligibility
The significant group difference in speech intelligibility suggests that an everyday communication context high in background noise like multi-talker babble can interfere with a listener’s ability to understand aging speech. When intelligibility judgments were carried out without multitalker babble with ALS participants presymptomatic for bulbar decline, elevated intelligibility scores suggestive of compensatory adjustments were noted for this group, whereas measures like tongue movement speed showed decrements consistent with bulbar decline (Green, Yunusova, et al., 2013). Because individuals may use compensatory inter- or intra-articulator adjustments in an attempt to improve their intelligibility (DePaul & Brooks, 1993; Green, Yunusova, et al., 2013; Yunusova et al., 2010), judging intelligibility in suboptimal listening environments may not only be a workaround for possible speaker- and listener-related factors, but could also provide estimates of functional speech from settings that mimic challenging daily communication contexts.
4.3. Contribution of Selected Subsystem Variables to Speech Intelligibility
4.3.1. Phonatory subsystem.
CPP SD was the strongest phonatory-acoustic contributor to intelligibility changes where for every one dB unit decrease in CPP variability, intelligibility decreased by 6.40%. Variability in CPP is known to decrease in dysphonic voices during continuous speech (Awan et al., 2010; Watts & Awan, 2011). Continuous speech is marked by prosody and requires agile transitions between highly periodic vowel productions and aperiodic consonant productions, which are degraded in dysphonic voices (Awan et al., 2010). Moreover, this phonatory flexibility decreases in the aging voice and CPP SD may be a sensitive indicator of this change (Lowell, 2012; Lowell, Colton, Kelley, & Mizia, 2013). A review of previous research investigating the effects of a monotone voice on intelligibility showed that speech with minimal fo variation or flattened fo contours was less intelligible (Klopfenstein, 2009).
Differential effects of CAPE-V sentences on cepstral/spectral measures in dysphonic speakers (Awan, 2011; Watts, 2015) provide an explanation as to why mean CPP did not contribute to intelligibility changes. Similar to Watts (2015), our results showed a significant main effect of sentence type for mean CPP and L/H (Table 4). Mean CPP is known to be a robust measure in the all voiced sentence task (Awan, 2011; Watts, 2015) as can be seen in Figure 2. Thus, an all voiced sentence is the preferred choice for auditory-perceptual ratings of voice quality because compared with other ADSV sentences correlations with dysphonia severity are strongest (Awan, 2011). On the other hand, CPP SD is a measure that is least affected by sentence type (Watts, 2015) and this makes CPP SD a potentially clinically useful acoustic marker for reduced intelligibility as shown in our results. Of note, as natural connected speech is phonetically variable and our model of intelligibility is exploratory, we chose not to focus on a particular ADSV sentence and sentence type was not included as a variable in our GEE model.
In contrast to mean CPP, mean L/H ratio was a significant contributor to intelligibility changes where for every one dB unit increase in L/H, intelligibility decreased by 2.01%. The measurement differences likely played a role. Whereas mean CPP focuses on the spectral energy of the fundamental frequency, L/H ratio calculates the relative distribution of low- versus high-frequency spectral energy (below/above 4 kHz) so that low spectral energy also includes the first couple of formants (Awan, 2011). Both mean CPP and mean L/H ratio have been correlated with perceived breathiness (Hillenbrand & Houde, 1996), but Awan et al. (Awan, 2011; Awan et al., 2009) concluded that L/H ratio may be more important for the prediction of the severity of breathiness (greater spectral noise above 2–3 kHz). Increased breathiness can be a feature in the aging voice (Gregory et al., 2012) and its influence on speech intelligibility may be better captured with L/H ratio than mean CPP.
Variability in L/H spectral ratio also emerged as a contributor to speech intelligibility differences, but the direction was reversed. Our results show that for every one dB unit increase in L/H SD, intelligibility decreased by 3.20%. The result is paradoxical because normal voices demonstrate increased variability in L/H due to normal transitions between vowel and consonant productions, while dysphonic voices are more consistently unstable resulting in reduced variability in L/H (Awan et al., 2010). Again, our interpretation may be complicated by the fact that L/H SD was averaged across ADSV tasks. Although in the current study the main effect of ADSV sentence type on L/H SD was nonsignificant, research by Watts (2015) did find one. Focusing only on the Rainbow Passage, Lowell et al. (2013) found that group differences in L/H SD were minimal between individuals with normal, breathy, and rough voices. Nonetheless, Awan et al. (2015) found lower L/H SD during the production of Rainbow Passage sentences to differentiate between older and younger adults. L/H SD thus deserves greater attention in future research.
Finally, regarding the composite CSID measure, which reflects overall dysphonia severity, for every one unit increase in CSID, intelligibility decreased by 0.21%. CSID weights the input of CPP and L/H spectral ratio (means and SDs). CSID values are correlated with the 100 mm CAPE-V rating scale (Awan, 2011) and a CAPE-V rating of 30, representing a mild dysphonia, would already be associated with a 6% drop in intelligibility. Of note, the two prior studies investigating a multiple subsystem approach to intelligibility did not include cepstral/spectral measures (Lee et al., 2014; Rong et al., 2016). Additionally, Ishikawa et al. (2018) only focused on mean CPP in their study on the relationship between dysphonia and intelligibility for which they found a moderate effect. Our data show that other cepstral/spectral measures are valuable to elucidate the relationship between the aging voice and intelligibility as suggested by Ishikawa et al. (2018) as a future direction.
4.3.2. Articulatory subsystem.
Among all the articulatory measures, maximum movement speed emerged as the strongest contributor to intelligibility changes in older adults and is congruent with prior research where maximum speed was reported to predict intelligibility decline in the ALS population (Rong et al., 2016). Based on the estimates from the current study, for every one mm/sec unit decrease in lingual maximum speed, intelligibility dropped by 0.13%, suggesting that slowing of tongue movements, which is an early indicator of age-induced change, likely contributes to reduced speech intelligibility in the elderly. In neurological populations like ALS and CP, slowing of articulatory movements is also known to have a detrimental effect on intelligibility (Ball, Willis, Beukelman, & Pattee, 2001; Lee et al., 2014; Rong et al., 2016; Yunusova et al., 2012). In fact, acoustic measures representing articulatory speed like average F2 slope are considered an index of the severity of the speech motor deficit (Lee et al., 2014). At the kinematic level, (Rong et al., 2016) found that slowed lip and jaw movements resulted in intelligibility declines for individuals with ALS. Although tongue data were not included in their study, these authors recommended including the tongue in future studies to help strengthen the relationship between articulatory function and intelligibility.
Between the two subsystems, a greater number of statistically significant associations with speech intelligibility were observed for phonatory measures than articulatory measures. A preliminary study on the prevalence of voice disorders in the healthy aging population indicated a point prevalence of 29% and a lifetime prevalence of 47% in non-treatment seeking adults over the age of 65 who complained of voice problems broadly defined as “any time the voice did not work, perform, or sound as it normally should so that it interfered with communication” (Roy et al., 2007, p. 629, p. 629). Common symptoms reported by participants in that study included hoarseness, difficulty projecting, and voice tiring or changing quality, all of which may contribute to reduced intelligibility, as supported by the findings of the current study. The fact that voice complaints are reported in the aging population while articulatory problems are rarely reported may indicate that voice problems are more prevalent in older adults. Although prior dysarthria studies suggest that the articulatory subsystem has the most significant impact on intelligibility, phonatory subsystem measures may be most sensitive to the intelligibility deficit in populations with predominant phonatory involvement. The current findings reinforce the need to develop disease- and age-specific models of intelligibility because speech subsystems can be differentially impaired with disease and age progression.
4.4. Clinical Implications
By studying healthy aging and investigating an exploratory age-based model of speech intelligibility, we were able to identify subsystem measures that differentiate healthy older and younger adults’ speech with the potential for application to dysarthria. From a clinical standpoint, developing a model of speech intelligibility for dysarthria will improve diagnostic and prognostic accuracy, allowing patients more time to make personal, financial, and/or treatment decisions. From a clinician’s perspective, developing a model for intelligibility will not only allow the clinician to present treatment options in a timely manner, but will also help generate effective and appropriate therapy goals, to best serve clients’ communication needs. Importantly, intelligibility models can help focus treatment on subsystem problems that interfere most with intelligibility (Lee et al., 2014).
4.5. Limitations and Future Directions
Certain limitations need to be considered when interpreting the results from the present study such as the small sample size and disproportionate age representation per decade in the older adults, where majority of the older adults were between 55–75 years and only one person was above 80 years (i.e., 81 years old). Another limitation is the absence of measures representing the respiratory and resonatory subsystems. Even though vital capacity data was collected from all the participants, it was not included in the study because this single measure does not adequately represent respiratory function during speech. Resonatory subsystem measures were not included due to the lack of appropriate instrumentation (e.g., nasometer) and the inadequacy of acoustic measures to represent resonatory function. Regarding kinematic analysis, admittedly, the sentence-level signals are dynamic, and therefore, could benefit from being segmented into smaller units, to produce more representative measures of articulatory movement across all gestures in an utterance. Future studies should consider using techniques such as the stroke segmentation technique (Tasko & Westbury, (2002) along with dispersion measures based on convex hull and generalized variance calculations of sentence-level kinematic data.
Despite well-documented reports of sex-based differences in the aging voice (Lenell, Sandage, & Johnson, 2019), the two groups in the present study were not carefully controlled for sex. However, sex-based differences were not the focus of our study, because a much larger sample size with balanced groups would have been necessary to adequately account for sex. Instead, our groups were nearly balanced for sex and sex was used as a covariate in the statistical analysis. With regard to hearing status, three older individuals had elevated hearing thresholds, which may have confounded their speech production data. Of note, the hearing screenings were conducted in a laboratory setting rather than in a sound-attenuating booth, which may have affected the hearing thresholds slightly. For speech intelligibility estimates, the same CAPE-V and Rainbow Passage sentences were used across talkers, which can inflate intelligibility scores because of listener familiarity with the test stimuli. However, the group by sentence type (i.e., SIT versus CAPE-V/Rainbow) interaction was not significant for intelligibility, which suggests that repeated exposure to the stimuli did not result in significant between-group differences in intelligibility scores.
In future studies, laryngeal videostroboscopies should be completed to document the presence or absence of glottal insufficiency or laryngeal pathology. Additionally, previously established acoustic measures such as F2 slope should also be included in a model of speech intelligibility related to healthy aging.
4.6. Conclusion
The study findings confirm that there are age-related changes to the phonatory and articulatory subsystems, which likely contributes to the reduction in speech intelligibility. For the phonatory system, older adults demonstrated shorter phonation durations and voices with lower harmonic energy. For the articulatory subsystem, older adults displayed slower tongue movement speed compared to younger adults. Phonatory cepstral/spectral measures such as CPP SD, L/H mean and SD, and CSID, as well as maximum tongue speed emerged as significant contributors to speech intelligibility changes in older adults. The suggested clinical significance is that changes in voice quality contribute to reduced intelligibility. Similarly, pertaining to the articulatory system, slower tongue movement speed contributes to reduced intelligibility in older individuals. Our study is a first step toward identifying subsystem measures that are clinically relevant for detecting differences in healthy older and younger adults’ speech with the potential for application to progressive and nonprogressive dysarthrias.
Highlights.
Phonatory age-related changes included greater airflow and lower cepstral measures
Aging-induced slowing of tongue movements was observed for the articulatory system
Cepstral/spectral measures, except CPP predicted age-related intelligibility change
Maximum speed was the strongest contributor to age-related intelligibility change
Acknowledgments:
The manuscript is based on research from the Master’s thesis by Jacob McKinley who consented to the authorship order. This research was supported by Students Pursuing an Academic and Research Career (SPARC) Award from the American Speech-Language-Hearing Association. Research reported in this publication was also partially supported by the National Institute on Deafness and Other Communication Disorders of the National Institutes of Health under Award Numbers R15DC015335 and R15DC01638. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Thanks to Keiko Ishikawa for feedback regarding the listening experiment. Thanks to Kristina Adler, Ashton Bernskoetter, Adrienne Cameron, Alexandra Dent, Kelly Fousek, Haley Harding, Jessica Lisenbee, Haley McCabe, Julie Meyer, Erin Nichols, Katie Nielsen, Mary Salazar, Katie Threlkeld, and Allison Walker for assistance with data collection and analysis.
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
For the rest of this manuscript we will refer to the function of the vocal folds as purely phonatory, following the convention in clinical literature.
References
- Amerman JD, & Parnell MM (1992). Speech timing strategies in elderly adults. Journal of Phonetics, 20(1), 65–76. [Google Scholar]
- Austin PC, & Steyerberg EW (2015). The number of subjects per variable required in linear regression analyses. Journal of Clinical Epidemiology, 68(6), 627–636. doi: 10.1016/j.jclinepi.2014.12.014 [DOI] [PubMed] [Google Scholar]
- Awan SN (2006). The aging female voice: Acoustic and respiratory data. Clinical Linguistics & Phonetics, 20(2–3), 171–180. doi: 10.1080/02699200400026918 [DOI] [PubMed] [Google Scholar]
- Awan SN (2011). Analysis of Dysphonia in Speech and Voice (ADSV): An application guide. Montvale, NJ: KayPENTAX. [Google Scholar]
- Awan SN, Acompanado J, Connors E, & Fanelli K (2015). The aging voice: A comprehensive analysis. Paper presented at the American Speech-Language-Hearing Association Convention, Denver, CO. [Google Scholar]
- Awan SN, Roy N, & Dromey C (2009). Estimating dysphonia severity in continuous speech: Application of a multi-parameter spectral/cepstral model. Clinical Linguistics & Phonetics, 23(11), 825–841. doi: 10.3109/02699200903242988 [DOI] [PubMed] [Google Scholar]
- Awan SN, Roy N, Jetté ME, Meltzner GS, & Hillman RE (2010). Quantifying dysphonia severity using a spectral/cepstral-based acoustic index: Comparisons with auditory-perceptual judgements from the CAPE-V. Clinical Linguistics & Phonetics, 24(9), 742–758. doi: 10.3109/02699206.2010.492446 [DOI] [PubMed] [Google Scholar]
- Baker KK, Ramig L, Sapir S, Luschei ES, & Smith ME (2001). Control of vocal loudness in young and old adults. Journal of Speech, Language, and Hearing Research, 44, 297–305. [DOI] [PubMed] [Google Scholar]
- Ball LJ, Willis A, Beukelman DR, & Pattee GL (2001). A protocol for identification of early bulbar signs in amyotrophic lateral sclerosis. Journal of the Neurological Sciences, 191(1–2), 43–53. [DOI] [PubMed] [Google Scholar]
- Bässler R (1987). Histopathology of different types of atrophy of the human tongue. Pathology, Research and Practice, 182(1), 87–97. doi: 10.1016/S0344-0338(87)80147-4 [DOI] [PubMed] [Google Scholar]
- Boersma P, & Weenink D (2019). Praat: Doing phonetics by computer (Version 6.0.30). Retrieved from http://www.praat.org
- Callisaya ML, Blizzard L, Schmidt MD, McGinley JL, & Srikanth VK (2010). Ageing and gait variability--a population-based study of older people. Age and Ageing, 39(2), 191–197. doi: 10.1093/ageing/afp250 [DOI] [PubMed] [Google Scholar]
- Cohen J (1988). Statistical power analysis for the behavioral sciences (2 ed.). Hillsdale, NJ: Lawrence Erlbaum Associates. [Google Scholar]
- Cohen J, Cohen P, West SG, & Aiken LS (2003). Applied multiple regression/correlation analysis for the behavioral sciences (3 ed.). Mahwah, NJ: Lawrence Erlbaum Associates. [Google Scholar]
- De Bodt MS, Hernandez-Diaz Huici ME, & Van De Heyning PH (2002). Intelligibility as a linear combination of dimensions in dysarthric speech. Journal of Communication Disorders, 35, 284–292. [DOI] [PubMed] [Google Scholar]
- DePaul R, & Brooks BR (1993). Multiple orofacial indices in amyotrophic lateral sclerosis. Journal of Speech and Hearing Research, 36(6), 1158–1167. [DOI] [PubMed] [Google Scholar]
- Doherty TJ, Vandervoort AA, & Brown WF (1993). Effects of ageing on the motor unit: A brief review. Canadian Journal of Applied Physiology = Revue Canadienne de Physiologie Appliquee, 18(4), 331–358. [DOI] [PubMed] [Google Scholar]
- Dromey C, Boyce K, & Channell R (2014). Effects of age and syntactic complexity on speech motor performance. Journal of Speech, Language, and Hearing Research, 57(6), 2142–2151. doi: 10.1044/2014_JSLHR-S-13-0327 [DOI] [PubMed] [Google Scholar]
- Duffy JR (2013). Motor speech disorders: Substrates, differential diagnosis, and management. St. Louis, MO: Elsevier Health Sciences. [Google Scholar]
- Etter NM, Stemple JC, & Howell DM (2013). Defining the lived experience of older adults with voice disorders. Journal of Voice, 27(1), 61–67. doi: 10.1016/j.jvoice.2012.07.002 [DOI] [PubMed] [Google Scholar]
- Fairbanks G (1960). Voice and articulation drill book (2nd ed.). New York: Harper and Row. [Google Scholar]
- Flanagan KP, & Dembowski JS (2002). Kinematics of normal lingual diadokokinesis. Journal of the Acoustical Society of America, 111(5), 2476–2476. [Google Scholar]
- Gabell A, & Nayak US (1984). The effect of age on variability in gait. Journal of Gerontology, 39(6), 662–666. [DOI] [PubMed] [Google Scholar]
- Glass GV, & Hopkins KD (1996). Statistical methods in education and psychology (3 ed.). Boston: Allyn and Bacon. [Google Scholar]
- Goozée JV, Stephenson DK, Murdoch BE, Darnell RE, & Lapointe LL (2005). Lingual kinematic strategies used to increase speech rate: Comparison between younger and older adults. Clinical Linguistics & Phonetics, 19(4), 319–334. doi: 10.1080/02699200420002268862 [DOI] [PubMed] [Google Scholar]
- Grabiner PC, Biswas ST, & Grabiner MD (2001). Age-related changes in spatial and temporal gait variables. Archives of Physical Medicine and Rehabilitation, 82(1), 31–35. doi: 10.1053/apmr.2001.182919 [DOI] [PubMed] [Google Scholar]
- Green JR, Wang J, & Wilson DL (2013). Smash: A tool for articulatory data processing and analysis In Bimbot F, Cerisara C, Fougeron G, Gravier G, Lamel L, Pellegrino F & Perrier P (Eds.), Interspeech 2013—14th annual conference of the International Speech Communication Association (pp. 1331–1335). Lyon, France: International Speech Communication Association. [Google Scholar]
- Green JR, Yunusova Y, Kuruvilla MS, Wang J, Pattee GL, Synhorst L, … Berry JD (2013). Bulbar and speech motor assessment in ALS: Challenges and future directions. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 14, 494–500. doi: 10.3109/21678421.2013.817585 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gregory ND, Chandran S, Lurie D, & Sataloff RT (2012). Voice disorders in the elderly. Journal of Voice, 26(2), 254–258. doi: 10.1016/j.jvoice.2010.10.024 [DOI] [PubMed] [Google Scholar]
- Harrell FE Jr., Lee KL, Califf RM, Pryor DB, & Rosati RA (1984). Regression modelling strategies for improved prognostic prediction. Statistics in Medicine, 3(2), 143–152. doi: 10.1002/sim.4780030207 [DOI] [PubMed] [Google Scholar]
- Hillenbrand J, & Houde RA (1996). Acoustic correlates of breathy vocal quality: Dysphonic voices and continuous speech. Journal of Speech and Hearing Research, 39(2), 311–321. [DOI] [PubMed] [Google Scholar]
- Hixon TJ, Weismer G, & Hoit JD (2014). Preclinical speech science: Anatomy, physiology, acoustics, perception (2nd ed.). San Diego: Plural. [Google Scholar]
- Hoit JD, & Hixon TJ (1987). Age and speech breathing. Journal of Speech and Hearing Research, 30(3), 351–366. [DOI] [PubMed] [Google Scholar]
- Hooper CR, & Cralidis A (2009). Normal changes in the speech of older adults: You’ve still got what it takes; it just takes a little longer! Perspectives on Gerontology, 14(2), 47–56. [Google Scholar]
- Huber JE, & Spruill JI (2008). Age-related changes to speech breathing with increased vocal loudness. Journal of Speech, Language, and Hearing Research, 51(3), 651–668. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ishibashi H, Takenoshita Y, Ishibashi K, & Oka M (1995). Age-related changes in the human mandibular condyle: A morphologic, radiologic, and histologic study. Journal of Oral and Maxillofacial Surgery, 53(9), 1016–1023. [DOI] [PubMed] [Google Scholar]
- Ishikawa K, de Alarcon A, Khosla S, Kelchner L, Silbert N, & Boyce S (2018). Predicting intelligibility deficit in dysphonic speech with cepstral peak prominence. Annals of Otology, Rhinology and Laryngology, 127(2), 69–78. doi: 10.1177/0003489417743518 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jacobson BH, Johnson A, Grywalski C, Silbergleit A, Jacobson G, Benninger MS, & Newman CW (1997). The Voice Handicap Index (VHI): Development and validation. American Journal of Speech-Language Pathology, 6(3), 66–70. doi: 1058-0360/97/0603-006 [Google Scholar]
- Kahane J (1987). Connective tissue changes in the larynx and their effects on voice. Journal of Voice, 1, 27–30. [Google Scholar]
- Kang HG, & Dingwell JB (2008). Separating the effects of age and walking speed on gait variability. Gait and Posture, 27(4), 572–577. doi: 10.1016/j.gaitpost.2007.07.009 [DOI] [PubMed] [Google Scholar]
- Kempster GB, Gerratt BR, Verdolini Abbott K, Barkmeier-Kraemer J, & Hillman RE (2009). Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol. American Journal of Speech Language Pathology, 18(2), 124–132. doi: 1058-0360/09/1802-012 [DOI] [PubMed] [Google Scholar]
- Kendall K (2007). Presbyphonia: A review. Current Opinion in Otolaryngology & Head and Neck Surgery, 15, 137–140. [DOI] [PubMed] [Google Scholar]
- Kent RD (1996). Hearing and believing: Some limits to the auditory-perceptual assessment of speech and voice disorders. American Journal of Speech-Language Pathology, 5(3), 7–23. [Google Scholar]
- Kent RD, Weismer G, Kent JF, & Rosenbek JC (1989). Toward phonetic intelligibility testing in dysarthria. Journal of Speech and Hearing Disorders, 54(4), 482–499. [DOI] [PubMed] [Google Scholar]
- Kim Y, Kent RD, & Weismer G (2011). An acoustic study of the relationships among neurologic disease, dysarthria type, and severity of dysarthria. Journal of Speech, Language, and Hearing Research, 54(2), 417–429. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kleinow J, Smith A, & Ramig LO (2001). Speech motor stability in IPD: Effects of rate and loudness manipulations. Journal of Speech, Language, and Hearing Research, 44(5), 1041–1051. doi: 10.1044/1092-4388(2001/082) [DOI] [PubMed] [Google Scholar]
- Klopfenstein M (2009). Interaction between prosody and intelligibility. International Journal of Speech-Language Pathology, 11(4), 326–331. doi: 10.1080/17549500903003094 [DOI] [Google Scholar]
- Kuruvilla-Dugdale M, & Mefferd A (2017). Spatiotemporal movement variability in ALS: Speaking rate effects on tongue, lower lip, and jaw motor control. Journal of Communication Disorders, 67, 22–34. doi: 10.1016/j.jcomdis.2017.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kuruvilla MS, Green JR, Yunusova Y, & Hanford K (2012). Spatiotemporal coupling of the tongue in amyotrophic lateral sclerosis. Journal of Speech, Language, and Hearing Research, 55(6), 1897–1909. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Landers R (2015). Computing intraclass correlations (ICC) as estimates of interrater reliability in SPSS. The Winnower. doi: 10.15200/winn.143518.81744 [DOI] [Google Scholar]
- Lee J, Hustad KC, & Weismer G (2014). Predicting speech intelligibility with a multiple speech subsystems approach in children with cerebral palsy. Journal of Speech, Language, and Hearing Research, 57(5), 1666–1678. doi: 10.1044/2014_JSLHR-S-13-0292 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lenell C, Sandage MJ, & Johnson AM (2019). A tutorial of the effects of sex hormones on laryngeal senescence and neuromuscular response to exercise. Journal of Speech, Language, and Hearing Research, 62(3), 602–610. doi: 10.1044/2018_JSLHR-S-18-0179 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liss JM, Weismer G, & Rosenbek JC (1990). Selected acoustic characteristics of speech production in very old males. Journal of Gerontology, 45(2), 35–45. doi: 10.1093/geronj/45.2.P35 [DOI] [PubMed] [Google Scholar]
- Lowell SY (2012). The acoustic assessment of voice in continuous speech. Perspectives on Voice and Voice Disorders, 22(2), 57–63. doi: 10.1044/vvd22.2.57 [DOI] [Google Scholar]
- Lowell SY, Colton RH, Kelley RT, & Mizia SA (2013). Predictive value and discriminant capacity of cepstral- and spectral-based measures during continuous speech. Journal of Voice, 27(4), 393–400. doi: 10.1016/j.jvoice.2013.02.005 [DOI] [PubMed] [Google Scholar]
- Mau T, Jacobson BH, & Garrett CG (2010). Factors associated with voice therapy outcomes in the treatment of presbyphonia. Laryngoscope, 120(6), 1181–1187. doi: 10.1002/lary.20890 [DOI] [PubMed] [Google Scholar]
- McAuliffe MJ, Wilding PJ, Rickard NA, & O’Beirne GA (2012). Effect of speaker age on speech recognition and perceived listening effort in older adults with hearing loss. Journal of Speech, Language, and Hearing Research, 55(3), 838–847. doi: 10.1044/1092-4388(2011/11-0101) [DOI] [PubMed] [Google Scholar]
- McCloy D Mix speech with noise. Retrieved from http://groups.linguistics.northwestern.edu/speech_comm_group/documents/praat%20scripts/MixSpeechNoise.praat
- McHenry MA (2003). The effect of pacing strategies on the variability of speech movement sequences in dysarthria. Journal of Speech, Language, and Hearing Research, 46(3), 702–710. [DOI] [PubMed] [Google Scholar]
- Mefferd AS, & Green JR (2010). Articulatory-to-acoustic relations in response to speaking rate and loudness manipulations. Journal of Speech, Language, and Hearing Research, 53(5), 1206–1219. doi: 10.1044/1092-4388(2010/09-0083) [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mefferd AS, Pattee GL, & Green JR (2014). Speaking rate effects on articulatory pattern consistency in talkers with mild ALS. Clinical Linguistics & Phonetics, 28(11), 799–811. doi: 10.3109/02699206.2014.908239 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Monemi M, Eriksson PO, Eriksson A, & Thornell LE (1998). Adverse changes in fibre type composition of the human masseter versus biceps brachii muscle during aging. Journal of the Neurological Sciences, 154(1), 35–48. doi: 10.1016/S0022-510x(97)00208-6 [DOI] [PubMed] [Google Scholar]
- Monemi M, Thornell LE, & Eriksson PO (1999). Diverse changes in fibre type composition of the human lateral pterygoid and digastric muscles during aging. Journal of the Neurological Sciences, 171(1), 38–48. [DOI] [PubMed] [Google Scholar]
- Mueller PB, Sweeney RJ, & Baribeau LJ (1984). Acoustic and morphologic study of the senescent voice. Ear, Nose, and Throat Journal, 63(6), 292–295. [PubMed] [Google Scholar]
- Nakayama M (1991). Histological study on aging changes in the human tongue. Nihon Jibiinkoka Gakkai Kaiho, 94(4), 541–555. [DOI] [PubMed] [Google Scholar]
- Nasreddine ZS, Phillips NA, Bédirian V, Charbonneau S, Whitehead V, Collin I, … Chertkow H (2005). The Montreal Cognitive Assessment, MoCA: A brief screening tool for mild cognitive impairment. Journal of the American Geriatrics Society, 53(4), 659–699. doi: 10.1111/j.1532-5415.2005.53221.x [DOI] [PubMed] [Google Scholar]
- Parnell MM, & Amerman JD (1996). An 11-year follow-up of motor speech abilities in elderly adults. Clinical Linguistics & Phonetics, 10(2), 103–118. doi: 10.3109/02699209608985165 [DOI] [Google Scholar]
- Patel RR, Awan SN, Barkmeier-Kraemer J, Courey M, Deliyski D, Eadie T, … Hillman R (2018). Recommended protocols for instrumental assessment of voice: American Speech-Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function. American Journal of Speech-Language Pathology, 27(3), 887–905. doi: 10.1044/2018_ajslp-17-0009 [DOI] [PubMed] [Google Scholar]
- Ramig LA (1983). Effects of physiological aging on speaking and reading rates. Journal of Communication Disorders, 16(3), 217–226. [DOI] [PubMed] [Google Scholar]
- Robbins JA, Levine R, Wood J, Roecker EB, & Luschei E (1995). Age effects on lingual pressure generation as a risk factor for dysphagia. The Journals of Gerontology Series A: Biological Sciences and Medical Sciences, 50(5), M257–M262. [DOI] [PubMed] [Google Scholar]
- Rong P, Yunusova Y, Wang J, Zinman L, Pattee GL, Berry JD, … Green JR (2016). Predicting speech intelligibility decline in amyotrophic lateral sclerosis based on the deterioration of individual speech subsystems. PloS One, 11(5), e0154971. doi: 10.1371/journal.pone.0154971 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rothauser EH, Chapman WD, Guttman N, Hecker MHL, Nordby KS, Silbiger HR, … Weinstock M (1969). IEEE recommended pratice for speech quality measurements. IEEE Transactions on Audio and Electroacoustics, 17(3), 225–246. [Google Scholar]
- Roy N, Stemple J, Merrill RM, & Thomas L (2007). Epidemiology of voice disorders in the elderly: Preliminary findings. Laryngoscope, 117(4), 628–633. doi: 10.1097/MLG.0b013e3180306da1 [DOI] [PubMed] [Google Scholar]
- Ryan WJ, & Burk KW (1974). Perceptual and acoustic correlates of aging in the speech of males. Journal of Communication Disorders, 7(2), 181–192. doi: 0021-9924(74)90030-6 [pii] [DOI] [PubMed] [Google Scholar]
- Sato K, Hirano M, & Nakashima T (2002). Age-related changes of collagenous fibers in the human vocal fold mucosa. Annals of Otology, Rhinology and Laryngology, 111(1), 15–20. doi: 10.1177/000348940211100103 [DOI] [PubMed] [Google Scholar]
- Sauder C, Bretl M, & Eadie T (2017). Predicting voice disorder status from smoothed measures of cepstral peak prominence using Praat and Analysis of Dysphonia in Speech and Voice (ADSV). Journal of Voice, 31(5), 557–566. doi: 10.1016/j.jvoice.2017.01.006 [DOI] [PubMed] [Google Scholar]
- Shrout PE, & Fleiss JL (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. doi: 10.1037//0033-2909.86.2.420 [DOI] [PubMed] [Google Scholar]
- Shuey EM (1989). Intelligibility of older versus younger adults’ cvc productions. Journal of Communication Disorders, 22(6), 437–444. doi: 10.1016/0021-9924(89)90036-1 [DOI] [PubMed] [Google Scholar]
- Smith A, Goffman L, Zelaznik HN, Ying G, & McGillem C (1995). Spatiotemporal stability and patterning of speech movement sequences. Experimental Brain Research, 104(3), 493–501. [DOI] [PubMed] [Google Scholar]
- Smith BL, Wasowicz J, & Preston J (1987). Temporal characteristics of the speech of normal elderly adults. Journal of Speech and Hearing Research, 30(4), 522–529. [DOI] [PubMed] [Google Scholar]
- Solomon NP (2011). Assessment of laryngeal airway resistance and phonation threshold pressure: Glottal enterprises In Ma EP-M & Yiu EM-L (Eds.), Handbook of voice assessments (pp. 31–49). San Diego: Plural. [Google Scholar]
- Solomon NP, & Helou LB (2013). Aerodynamic assessment of phonation: Avoiding common mistakes In Scherer R& Verdolini Abbott K (Eds.), The continuing influence of Ingo R. Titze on voice, science, and music: A Festschrift collection (pp. 49–56). Salt Lake City, Utah: National Center for Voice and Speech, University of Utah. [Google Scholar]
- Sonies BC, Stone M, & Shawker T (1984). Speech and swallowing in the elderly. Gerodontology, 3(2), 115–123. [DOI] [PubMed] [Google Scholar]
- Stevens KN (1972). The quantal nature of speech: Evidence from articulatory-acoustic data In Denes PB & David EE Jr (Eds.), Human communication: A unified view (pp. 51–66). New York: McGraw-Hill. [Google Scholar]
- Stevens KN (1989). On the quantal nature of speech. Journal of Phonetics, 17, 3–45. [Google Scholar]
- Stolze H, Friedrich HJ, Steinauer K, & Vieregge P (2000). Stride parameters in healthy young and old women--measurement variability on a simple walkway. Experimental Aging Research, 26(2), 159–168. doi: 10.1080/036107300243623 [DOI] [PubMed] [Google Scholar]
- Takeda N, Thomas GR, & Ludlow CL (2000). Aging effects on motor units in the human thyroarytenoid muscle. Laryngoscope, 110(6), 1018–1025. doi: 10.1097/00005537-200006000-00025 [DOI] [PubMed] [Google Scholar]
- Tasko SM, & Westbury JR (2002). Defining and measuring speech movement events. Journal of Speech, Language, and Hearing Research, 45(1), 127–142. doi: 10.1044/1092-4388(2002/010) [DOI] [PubMed] [Google Scholar]
- Thomas LB, Harrison AL, & Stemple JC (2008). Aging thyroarytenoid and limb skeletal muscle: Lessons in contrast. Journal of Voice, 22(4), 430–450. doi: 10.1016/j.jvoice.2006.11.006 [DOI] [PubMed] [Google Scholar]
- Tjaden K, Sussman JE, & Wilding GE (2014). Impact of clear, loud, and slow speech on scaled intelligibility and speech severity in Parkinson’s disease and multiple sclerosis. Journal of Speech, Language, and Hearing Research, 57(3), 779–792. doi: 10.1044/2014_JSLHR-S-12-0372 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Watts CR (2015). The effect of CAPE-V sentences on cepstral/spectral acoustic measures in dysphonic speakers. Folia Phoniatrica et Logopedica, 67(1), 15–20. doi: 10.1159/000371656 [DOI] [PubMed] [Google Scholar]
- Watts CR, & Awan SN (2011). Use of spectral/cepstral analyses for differentiating normal from hypofunctional voices in sustained vowel and continuous speech contexts. Journal of Speech, Language, and Hearing Research, 54(6), 1525–1537. doi: 10.1044/1092-4388(2011/10-0209) [DOI] [PubMed] [Google Scholar]
- Watts CR, Awan SN, & Maryn Y (2016). A comparison of cepstral peak prominence measures from two acoustic analysis programs. Journal of Voice, 31(3), 387.e381–387.e310. doi: 10.1016/j.jvoice.2016.09.012 [DOI] [PubMed] [Google Scholar]
- Watts CR, Ronshaugen R, & Saenz D (2015). The effect of age and vocal task on cepstral/spectral measures of vocal function in adult males. Clinical Linguistics & Phonetics, 29(6), 415–423. [DOI] [PubMed] [Google Scholar]
- Weismer G, Laures JS, Jeng JY, Kent RD, & Kent JF (2000). Effect of speaking rate manipulations on acoustic and perceptual aspects of the dysarthria in amyotrophic lateral sclerosis. Folia Phoniatrica et Logopaedica, 52(5), 201–219. [DOI] [PubMed] [Google Scholar]
- Wohlert AB, & Smith A (1998). Spatiotemporal stability of lip movements in older adult speakers. Journal of Speech, Language, and Hearing Research, 41, 41–50. [DOI] [PubMed] [Google Scholar]
- Xue SA, & Deliyski D (2001). Effects of aging on selected acoustic voice parameters: Preliminary normative data and educational implications. Educational Gerontology, 27, 159–168. [Google Scholar]
- Yoon K Normalize-intensity-db.Praat (computer program).
- Yorkston KM, Beukelman DR, Hakel M, & Dorsey M (2007). Speech intelligibility test. Lincoln, NE: Institute for Rehabilitation Science and Engineering at the Madonna Rehabilitation Hospital. [Google Scholar]
- Yorkston KM, Hakel M, Beukelman DR, & Fager S (2007). Evidence for effectiveness of treatment of loudness, rate, or prosody in dysarthria: A systematic review. Journal of Medical Speech-Language Pathology, 15(2), xi–xxxvi. [Google Scholar]
- Youmans SR, Stierwalt JAG, & Clark HM (2002). Measures of tongue function in healthy adults. Paper presented at the American Speech-Language-Hearing Association, Atlanta, GA. [Google Scholar]
- Yunusova Y, Green JR, Greenwood L, Wang J, Pattee GL, & Zinman L (2012). Tongue movements and their acoustic consequences in amyotrophic lateral sclerosis. Folia Phoniatrica Et Logopaedica, 64(2), 94–102. doi: 10.1159/000336890 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yunusova Y, Green JR, Lindstrom MJ, Ball LJ, Pattee GL, & Zinman L (2010). Kinematics of disease progression in bulbar ALS. Journal of Communication Disorders, 43(1), 6–20. doi: 10.1016/j.jcomdis.2009.07.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zemlin WR (1998). Speech and hearing science: Anatomy and physiology (4 ed.). Boston, MA: Allyn and Bacon. [Google Scholar]
- Zraick RI, Smith-Olinde L, & Shotts LL (2012). Adult normative data for the KayPENTAX Phonatory Aerodynamic System model 6600… [corrected] [published erratum appears in Journal of Voice 2013; 27(1) 2. Journal of Voice, 26(2), 164–176 113p. doi: 10.1016/j.jvoice.2011.01.006 [DOI] [PubMed] [Google Scholar]
